
From juha.hakala@helsinki.fi  Wed Jan  5 02:08:25 2011
Return-Path: <juha.hakala@helsinki.fi>
X-Original-To: urn@core3.amsl.com
Delivered-To: urn@core3.amsl.com
Received: from localhost (localhost [127.0.0.1]) by core3.amsl.com (Postfix) with ESMTP id B149B3A6B5E for <urn@core3.amsl.com>; Wed,  5 Jan 2011 02:08:25 -0800 (PST)
X-Virus-Scanned: amavisd-new at amsl.com
X-Spam-Flag: NO
X-Spam-Score: -2.524
X-Spam-Level: 
X-Spam-Status: No, score=-2.524 tagged_above=-999 required=5 tests=[AWL=0.075,  BAYES_00=-2.599]
Received: from mail.ietf.org ([64.170.98.32]) by localhost (core3.amsl.com [127.0.0.1]) (amavisd-new, port 10024) with ESMTP id DOvSx+3pDSlE for <urn@core3.amsl.com>; Wed,  5 Jan 2011 02:08:24 -0800 (PST)
Received: from smtp-rs1.it.helsinki.fi (smtp-rs1-vallila2.fe.helsinki.fi [128.214.173.75]) by core3.amsl.com (Postfix) with ESMTP id 4C1C03A6A1A for <urn@ietf.org>; Wed,  5 Jan 2011 02:08:22 -0800 (PST)
Received: from [128.214.91.90] (kkkl25.lib.helsinki.fi [128.214.91.90]) by smtp-rs1.it.helsinki.fi (8.13.1/8.13.1) with ESMTP id p05AAQwK014742 (version=TLSv1/SSLv3 cipher=DHE-RSA-AES256-SHA bits=256 verify=NOT); Wed, 5 Jan 2011 12:10:26 +0200
Message-ID: <4D244392.6000701@helsinki.fi>
Date: Wed, 05 Jan 2011 12:10:26 +0200
From: Juha Hakala <juha.hakala@helsinki.fi>
User-Agent: Thunderbird 2.0.0.24 (Windows/20100228)
MIME-Version: 1.0
To: Mykyta Yevstifeyev <evnikita2@gmail.com>
References: <4D131B01.3070502@helsinki.fi> <4D132D59.7020906@gmail.com> <4D133CF7.7010302@helsinki.fi> <4D136A1D.2060905@gmail.com>
In-Reply-To: <4D136A1D.2060905@gmail.com>
Content-Type: text/plain; charset=ISO-8859-1; format=flowed
Content-Transfer-Encoding: 7bit
Cc: urn@ietf.org
Subject: Re: [urn] Comments to the draft-ietf-urnbis-rfc2141bis-urn-00.txt
X-BeenThere: urn@ietf.org
X-Mailman-Version: 2.1.9
Precedence: list
List-Id: Discussions about possible revisions to the definition of Uniform Resource Names <urn.ietf.org>
List-Unsubscribe: <https://www.ietf.org/mailman/listinfo/urn>, <mailto:urn-request@ietf.org?subject=unsubscribe>
List-Archive: <http://www.ietf.org/mail-archive/web/urn>
List-Post: <mailto:urn@ietf.org>
List-Help: <mailto:urn-request@ietf.org?subject=help>
List-Subscribe: <https://www.ietf.org/mailman/listinfo/urn>, <mailto:urn-request@ietf.org?subject=subscribe>
X-List-Received-Date: Wed, 05 Jan 2011 10:08:25 -0000

Hello Mykyta,

See comments below.

Mykyta Yevstifeyev wrote:
> Jula,
> 
> Some responds to your comments:

>>>> Persistence (additional considerations): URN is typically related to 
>>>> a single manifestation of a resource (work). URN will have a longer 
>>>> life time than any such manifestation, which means that over long 
>>>> period of time resolution is only possible if URNs of successive 
>>>> manifestations of a resource are linked to one another, allowing the 
>>>> user to access a more modern version of the resource. The technical 
>>>> implementation of linking these URNs to one another is namespace / 
>>>> implementation specific and discussed in e.g. namespace registration 
>>>> documents.
>>
>>> I think this issue is in the scope of implementators who will develop 
>>> the URN resolution mechanisms.
>>
>> Yes and no. From the point of view of the URN system, I think this 
>> means that it should be possible to request a list of URNs related to 
>> the URN at hand. Thus, if I know one manifestation, I would be able to 
>> find all other manifestations as well, and the description of the work 
>> itself. This service and its inbuilt URN-URN linking would also make 
>> systems resilient in the sense that if the resolution service can not 
>> locate the manifestation the URN refers to, it could still deliver 
>> another, more modern manifestation.
> But in the way you describe it, it seems to be that different URNs can 
> refer to the same resource. But URNs are *unique* and it won't be OK, I 
> think, if more than 1 URN refers to the same resource.

Our different understanding of the term "resource" causes some 
confusion. Because it is important to understand how (bibliographic) 
identifiers work, let me elaborate this point a bit.

In libraries the current thinking is that the top level thing is work, 
such as Hamlet (the novel). There may be related work, such as Hamlet 
(the movie). There may also be different expressions of the Hamlet (the 
novel), such as translations to Finnish, Russian etc. Each expression 
will in turn have one or more manifestations, such as (in the case of 
Hamlet (the novel) hard-cover and paperback book, and electronic 
versions in e.g. plain text, Word 2003, PDF/A and EPUB.

Identifier communities have in theory clear guidelines on how to deal 
with these manifestations. For instance, every manifestation of a book 
should get its own ISBN. (As an aside, there is a pressure within the 
ISBN community to cut corners in this respect, which I think is a bad 
idea.) AFAIK whenever two manifestations of the same work no longer have 
the same intellectual content, it becomes vitally important that they do 
not have the same identifier. And when digital resources are migrated to 
new file formats, changes in look and feel and eventually also in the 
actual content are almost inevitable. Over time these changes will 
probably get more pronounced.

So my view is that each manifestation of a work should have one but only 
one identifier. When a new manifestation is produced via migration, it 
should get a new identifier, because migration may have changed the look 
and feel and/or the content of the resource. So eventually there will be 
a set of identifiers, consisting of one identifier referring to each 
manifestation of the work, and probably also one  belonging to the work 
itself (and its expressions, if any).

All manifestations of the same work should be linked to one another and 
to the metadata record describing the work itself using a persistent 
identifier, since a user should be made aware of the fact that there are 
other variants of the work. As long term preservation of digital 
resources will usually be based on migration strategy, there will 
eventually be a large number of different manifestations available for 
each properly preserved work. We have no control over which 
manifestation the permanent link has been made to. If the linked 
manifestation is outdated, a user following the link may not be able to 
use the document with the applications he/she has available; in such a 
case it is important to know that there are more modern versions 
available. On the other hand, the user may be a digital archeologist who 
wants to find the oldest manifestation so as to avoid all the changes 
migrations have caused to the look and feel and the intellectual content.

Short lifetime of digital manifestations is a headache for persistent 
linking, when "persistent" means at least several decades and often 
centuries. One solution would be to make the permanent link to work / 
expression; that is, to metadata record which is never changed apart 
from new links added whenever there are new manifestations available.

>>>>
>>>> Query and fragment
>>>>
>>>> The I-D proposes two different approaches for supporting <fragment>: 
>>>> individual assignment of fragments on document-to-document basis, 
>>>> and/or creation of a specific set of fragment identifiers which are 
>>>> generally applicable to all resources encompassed by a given URN 
>>>> namespace. Within a namespace only one of these approaches would be 
>>>> acceptable. While it is likely that both methods can be applied, we 
>>>> are not sure if these options are mutually exclusive within one 
>>>> namespace.
>>
>>> If we use fragments, I think it should apply to the whole NS. 
>>> Generally URN NSs specify resources of one type and structure so 
>>> one-type fragments for whole NS is a possibility.
>>
>> While I agree that most namespaces are homogeneous (ISSN namespace 
>> deals with serials, ISBN with books) there are some that are broad. 
>> NBN (National bibliography number) encompasses anything that the 
>> national libraries hold in their collections, which means all kinds of 
>> published cultural heritage, and even some non-published materials).

> But in this way, fragments should not be allowed at all, as assigning 
> fragments to all (or not all) elements in the namespace (when there 
> could thousands of them) does not seem very great perspective to me.

We both agree that this kind of random usage of <fragment> would not be 
sensible.

When the namespace belongs to a well-specified standard such as ISBN, 
the namespace registration could / should specify whether fragments are 
allowed at all, and if they are, how. With more vague namespaces such as 
for instance NBN it is harder to specify how fragments will be used, 
since we know neither the syntax of the identifier nor the resources 
NBNs will be applied to. If RFC2141bis allows the usage of fragments, 
NBN namespace registration should specify at least in broad terms some 
principles for the usage of fragments within that particular namespace.

Juha
-- 

  Juha Hakala
  Senior advisor, standardisation and IT

  The National Library of Finland
  P.O.Box 15 (Unioninkatu 36, room 503), FIN-00014 Helsinki University
  Email juha.hakala@helsinki.fi, tel +358 50 382 7678

From evnikita2@gmail.com  Sat Jan  8 08:07:07 2011
Return-Path: <evnikita2@gmail.com>
X-Original-To: urn@core3.amsl.com
Delivered-To: urn@core3.amsl.com
Received: from localhost (localhost [127.0.0.1]) by core3.amsl.com (Postfix) with ESMTP id 8336A28C117 for <urn@core3.amsl.com>; Sat,  8 Jan 2011 08:07:07 -0800 (PST)
X-Virus-Scanned: amavisd-new at amsl.com
X-Spam-Flag: NO
X-Spam-Score: -3.301
X-Spam-Level: 
X-Spam-Status: No, score=-3.301 tagged_above=-999 required=5 tests=[AWL=0.298,  BAYES_00=-2.599, RCVD_IN_DNSWL_LOW=-1]
Received: from mail.ietf.org ([64.170.98.32]) by localhost (core3.amsl.com [127.0.0.1]) (amavisd-new, port 10024) with ESMTP id 14K4O4IDZUjL for <urn@core3.amsl.com>; Sat,  8 Jan 2011 08:07:06 -0800 (PST)
Received: from mail-bw0-f44.google.com (mail-bw0-f44.google.com [209.85.214.44]) by core3.amsl.com (Postfix) with ESMTP id CA1B328C115 for <urn@ietf.org>; Sat,  8 Jan 2011 08:07:05 -0800 (PST)
Received: by bwz12 with SMTP id 12so17991597bwz.31 for <urn@ietf.org>; Sat, 08 Jan 2011 08:09:13 -0800 (PST)
DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=gamma; h=domainkey-signature:received:received:message-id:date:from :user-agent:mime-version:to:cc:subject:references:in-reply-to :content-type:content-transfer-encoding; bh=Oj6a0WcBbnEz9bnvR/pr2SsgSvkU97z/soAFse7G/8Q=; b=NtoN5220ErOXVyS08x9I+8uL2wP5zN17kpdBE38APdA/6x67bpxShH1O5B/rPj+brn JeVX6DcTTueXE1xmt1t8gCSw6D7RBKgIIxTTz6iKSeFS01evHpTY1bKwhuuZfpalXJTJ 5yBhj/YF/LJcydpmgB/gglvH1mk5bLTeRZg5w=
DomainKey-Signature: a=rsa-sha1; c=nofws; d=gmail.com; s=gamma; h=message-id:date:from:user-agent:mime-version:to:cc:subject :references:in-reply-to:content-type:content-transfer-encoding; b=pjdFBeA1HVc39Ih85696ph+JRSuWmXJExaRyQvpMTj5lnWXe98qKlyH+o3WxhKw5sK 4FH4tcfxoJydppYgy0Tni5ulPG9rw7O7IlxiYIIKL+JIHZv9ot2zreeUYNYxkLFlv1AE MfJz87Kqfi6fmyrbx2MZ/u+VsvhBjfO8NgNhY=
Received: by 10.204.116.5 with SMTP id k5mr1097474bkq.73.1294502953443; Sat, 08 Jan 2011 08:09:13 -0800 (PST)
Received: from [127.0.0.1] ([195.191.104.134]) by mx.google.com with ESMTPS id v1sm14774492bkt.17.2011.01.08.08.09.11 (version=SSLv3 cipher=RC4-MD5); Sat, 08 Jan 2011 08:09:12 -0800 (PST)
Message-ID: <4D288C39.30406@gmail.com>
Date: Sat, 08 Jan 2011 18:09:29 +0200
From: Mykyta Yevstifeyev <evnikita2@gmail.com>
User-Agent: Mozilla/5.0 (Windows; U; Windows NT 5.1; ru; rv:1.9.2.13) Gecko/20101207 Thunderbird/3.1.7
MIME-Version: 1.0
To: Juha Hakala <juha.hakala@helsinki.fi>
References: <4D131B01.3070502@helsinki.fi> <4D132D59.7020906@gmail.com> <4D133CF7.7010302@helsinki.fi> <4D136A1D.2060905@gmail.com> <4D244392.6000701@helsinki.fi>
In-Reply-To: <4D244392.6000701@helsinki.fi>
Content-Type: text/plain; charset=ISO-8859-1; format=flowed
Content-Transfer-Encoding: 7bit
Cc: urn@ietf.org
Subject: Re: [urn] Comments to the draft-ietf-urnbis-rfc2141bis-urn-00.txt
X-BeenThere: urn@ietf.org
X-Mailman-Version: 2.1.9
Precedence: list
List-Id: Discussions about possible revisions to the definition of Uniform Resource Names <urn.ietf.org>
List-Unsubscribe: <https://www.ietf.org/mailman/listinfo/urn>, <mailto:urn-request@ietf.org?subject=unsubscribe>
List-Archive: <http://www.ietf.org/mail-archive/web/urn>
List-Post: <mailto:urn@ietf.org>
List-Help: <mailto:urn-request@ietf.org?subject=help>
List-Subscribe: <https://www.ietf.org/mailman/listinfo/urn>, <mailto:urn-request@ietf.org?subject=subscribe>
X-List-Received-Date: Sat, 08 Jan 2011 16:07:07 -0000

Juha,

Please find some comments below.

05.01.2011 12:10, Juha Hakala wrote:
> Hello Mykyta,
>
> See comments below.
>
> Mykyta Yevstifeyev wrote:
>> Jula,
>>
>> Some responds to your comments:
>
>>>>> Persistence (additional considerations): URN is typically related 
>>>>> to a single manifestation of a resource (work). URN will have a 
>>>>> longer life time than any such manifestation, which means that 
>>>>> over long period of time resolution is only possible if URNs of 
>>>>> successive manifestations of a resource are linked to one another, 
>>>>> allowing the user to access a more modern version of the resource. 
>>>>> The technical implementation of linking these URNs to one another 
>>>>> is namespace / implementation specific and discussed in e.g. 
>>>>> namespace registration documents.
>>>
>>>> I think this issue is in the scope of implementators who will 
>>>> develop the URN resolution mechanisms.
>>>
>>> Yes and no. From the point of view of the URN system, I think this 
>>> means that it should be possible to request a list of URNs related 
>>> to the URN at hand. Thus, if I know one manifestation, I would be 
>>> able to find all other manifestations as well, and the description 
>>> of the work itself. This service and its inbuilt URN-URN linking 
>>> would also make systems resilient in the sense that if the 
>>> resolution service can not locate the manifestation the URN refers 
>>> to, it could still deliver another, more modern manifestation.
>> But in the way you describe it, it seems to be that different URNs 
>> can refer to the same resource. But URNs are *unique* and it won't be 
>> OK, I think, if more than 1 URN refers to the same resource.
>
> Our different understanding of the term "resource" causes some 
> confusion. Because it is important to understand how (bibliographic) 
> identifiers work, let me elaborate this point a bit.
>
> In libraries the current thinking is that the top level thing is work, 
> such as Hamlet (the novel). There may be related work, such as Hamlet 
> (the movie). There may also be different expressions of the Hamlet 
> (the novel), such as translations to Finnish, Russian etc. Each 
> expression will in turn have one or more manifestations, such as (in 
> the case of Hamlet (the novel) hard-cover and paperback book, and 
> electronic versions in e.g. plain text, Word 2003, PDF/A and EPUB.
>
> Identifier communities have in theory clear guidelines on how to deal 
> with these manifestations. For instance, every manifestation of a book 
> should get its own ISBN. (As an aside, there is a pressure within the 
> ISBN community to cut corners in this respect, which I think is a bad 
> idea.) AFAIK whenever two manifestations of the same work no longer 
> have the same intellectual content, it becomes vitally important that 
> they do not have the same identifier. And when digital resources are 
> migrated to new file formats, changes in look and feel and eventually 
> also in the actual content are almost inevitable. Over time these 
> changes will probably get more pronounced.
>
> So my view is that each manifestation of a work should have one but 
> only one identifier. When a new manifestation is produced via 
> migration, it should get a new identifier, because migration may have 
> changed the look and feel and/or the content of the resource. So 
> eventually there will be a set of identifiers, consisting of one 
> identifier referring to each manifestation of the work, and probably 
> also one  belonging to the work itself (and its expressions, if any).
I agree with you here. But if you say that we need the identifier for 
each version (or you say 'manifestation') of resource, why do we need 
the generic one?
>
> All manifestations of the same work should be linked to one another 
> and to the metadata record describing the work itself using a 
> persistent identifier, since a user should be made aware of the fact 
> that there are other variants of the work. As long term preservation 
> of digital resources will usually be based on migration strategy, 
> there will eventually be a large number of different manifestations 
> available for each properly preserved work. We have no control over 
> which manifestation the permanent link has been made to. If the linked 
> manifestation is outdated, a user following the link may not be able 
> to use the document with the applications he/she has available; in 
> such a case it is important to know that there are more modern 
> versions available. On the other hand, the user may be a digital 
> archeologist who wants to find the oldest manifestation so as to avoid 
> all the changes migrations have caused to the look and feel and the 
> intellectual content.
Here I agree with you too.
>
> Short lifetime of digital manifestations is a headache for persistent 
> linking, when "persistent" means at least several decades and often 
> centuries. One solution would be to make the permanent link to work / 
> expression; that is, to metadata record which is never changed apart 
> from new links added whenever there are new manifestations available.
But the question is who will record metadata? Different URN resolution 
services will have different metadata databases. Who or what will 
synchronize them?
>
>>>>>
>>>>> Query and fragment
>>>>>
>>>>> The I-D proposes two different approaches for supporting 
>>>>> <fragment>: individual assignment of fragments on 
>>>>> document-to-document basis, and/or creation of a specific set of 
>>>>> fragment identifiers which are generally applicable to all 
>>>>> resources encompassed by a given URN namespace. Within a namespace 
>>>>> only one of these approaches would be acceptable. While it is 
>>>>> likely that both methods can be applied, we are not sure if these 
>>>>> options are mutually exclusive within one namespace.
>>>
>>>> If we use fragments, I think it should apply to the whole NS. 
>>>> Generally URN NSs specify resources of one type and structure so 
>>>> one-type fragments for whole NS is a possibility.
>>>
>>> While I agree that most namespaces are homogeneous (ISSN namespace 
>>> deals with serials, ISBN with books) there are some that are broad. 
>>> NBN (National bibliography number) encompasses anything that the 
>>> national libraries hold in their collections, which means all kinds 
>>> of published cultural heritage, and even some non-published materials).
>
>> But in this way, fragments should not be allowed at all, as assigning 
>> fragments to all (or not all) elements in the namespace (when there 
>> could thousands of them) does not seem very great perspective to me.
>
> We both agree that this kind of random usage of <fragment> would not 
> be sensible.
>
> When the namespace belongs to a well-specified standard such as ISBN, 
> the namespace registration could / should specify whether fragments 
> are allowed at all, and if they are, how. With more vague namespaces 
> such as for instance NBN it is harder to specify how fragments will be 
> used, since we know neither the syntax of the identifier nor the 
> resources NBNs will be applied to. If RFC2141bis allows the usage of 
> fragments, NBN namespace registration should specify at least in broad 
> terms some principles for the usage of fragments within that 
> particular namespace.
In this case we should define smth like this: "The URN specification MAY 
define the usage of <fragmets>. In this way it SHALL either specify 
their generic syntax or define procedures for assignment fragments to 
segregate resources."

Mykyta Yevstifeyev
>
> Juha


From duerst@it.aoyama.ac.jp  Sun Jan  9 17:47:07 2011
Return-Path: <duerst@it.aoyama.ac.jp>
X-Original-To: urn@core3.amsl.com
Delivered-To: urn@core3.amsl.com
Received: from localhost (localhost [127.0.0.1]) by core3.amsl.com (Postfix) with ESMTP id 478F028C0E3 for <urn@core3.amsl.com>; Sun,  9 Jan 2011 17:47:07 -0800 (PST)
X-Virus-Scanned: amavisd-new at amsl.com
X-Spam-Flag: NO
X-Spam-Score: -100.399
X-Spam-Level: 
X-Spam-Status: No, score=-100.399 tagged_above=-999 required=5 tests=[AWL=-0.609, BAYES_00=-2.599, HELO_EQ_JP=1.244, HOST_EQ_JP=1.265, MIME_8BIT_HEADER=0.3, USER_IN_WHITELIST=-100]
Received: from mail.ietf.org ([64.170.98.32]) by localhost (core3.amsl.com [127.0.0.1]) (amavisd-new, port 10024) with ESMTP id 8Fuqu+hB4Nci for <urn@core3.amsl.com>; Sun,  9 Jan 2011 17:47:05 -0800 (PST)
Received: from scintmta01.scbb.aoyama.ac.jp (scintmta01.scbb.aoyama.ac.jp [133.2.253.33]) by core3.amsl.com (Postfix) with ESMTP id 5578328C0D7 for <urn@ietf.org>; Sun,  9 Jan 2011 17:47:04 -0800 (PST)
Received: from scmse02.scbb.aoyama.ac.jp ([133.2.253.231]) by scintmta01.scbb.aoyama.ac.jp (secret/secret) with SMTP id p0A1nFUx001762 for <urn@ietf.org>; Mon, 10 Jan 2011 10:49:16 +0900
Received: from (unknown [133.2.206.133]) by scmse02.scbb.aoyama.ac.jp with smtp id 741a_1408_d4944322_1c5b_11e0_8d8b_001d096c5782; Mon, 10 Jan 2011 10:49:15 +0900
Received: from [IPv6:::1] ([133.2.210.1]:51894) by itmail.it.aoyama.ac.jp with [XMail 1.22 ESMTP Server] id <S14B0910> for <urn@ietf.org> from <duerst@it.aoyama.ac.jp>; Mon, 10 Jan 2011 10:49:15 +0900
Message-ID: <4D2A6588.6000309@it.aoyama.ac.jp>
Date: Mon, 10 Jan 2011 10:48:56 +0900
From: =?ISO-8859-1?Q?=22Martin_J=2E_D=FCrst=22?= <duerst@it.aoyama.ac.jp>
Organization: Aoyama Gakuin University
User-Agent: Mozilla/5.0 (Windows; U; Windows NT 6.0; en-US; rv:1.9.1.9) Gecko/20100722 Eudora/3.0.4
MIME-Version: 1.0
To: Juha Hakala <juha.hakala@helsinki.fi>
References: <4D131B01.3070502@helsinki.fi> <4D132D59.7020906@gmail.com>	<4D133CF7.7010302@helsinki.fi> <4D136A1D.2060905@gmail.com> <4D244392.6000701@helsinki.fi>
In-Reply-To: <4D244392.6000701@helsinki.fi>
Content-Type: text/plain; charset=ISO-8859-1; format=flowed
Content-Transfer-Encoding: 8bit
Cc: urn@ietf.org
Subject: Re: [urn] Comments to the draft-ietf-urnbis-rfc2141bis-urn-00.txt
X-BeenThere: urn@ietf.org
X-Mailman-Version: 2.1.9
Precedence: list
List-Id: Discussions about possible revisions to the definition of Uniform Resource Names <urn.ietf.org>
List-Unsubscribe: <https://www.ietf.org/mailman/listinfo/urn>, <mailto:urn-request@ietf.org?subject=unsubscribe>
List-Archive: <http://www.ietf.org/mail-archive/web/urn>
List-Post: <mailto:urn@ietf.org>
List-Help: <mailto:urn-request@ietf.org?subject=help>
List-Subscribe: <https://www.ietf.org/mailman/listinfo/urn>, <mailto:urn-request@ietf.org?subject=subscribe>
X-List-Received-Date: Mon, 10 Jan 2011 01:47:07 -0000

Hello Juha, Mykyta,

May I remind you of the basics of fragment identifiers?

The meaning of fragment identifiers, which I would understand to include 
their presence or absence, is determined by the MIME media type of the 
representation returned when resolving an URI. So in general, very, very 
little if anything will be needed in terms of considerations for 
fragment identifiers in URN namespace specs. As for the URN spec itself, 
the most important thing to say is just what's said above.

Regards,   Martin.

On 2011/01/05 19:10, Juha Hakala wrote:
> Hello Mykyta,

>>>>> Query and fragment
>>>>>
>>>>> The I-D proposes two different approaches for supporting
>>>>> <fragment>: individual assignment of fragments on
>>>>> document-to-document basis, and/or creation of a specific set of
>>>>> fragment identifiers which are generally applicable to all
>>>>> resources encompassed by a given URN namespace. Within a namespace
>>>>> only one of these approaches would be acceptable. While it is
>>>>> likely that both methods can be applied, we are not sure if these
>>>>> options are mutually exclusive within one namespace.
>>>
>>>> If we use fragments, I think it should apply to the whole NS.
>>>> Generally URN NSs specify resources of one type and structure so
>>>> one-type fragments for whole NS is a possibility.
>>>
>>> While I agree that most namespaces are homogeneous (ISSN namespace
>>> deals with serials, ISBN with books) there are some that are broad.
>>> NBN (National bibliography number) encompasses anything that the
>>> national libraries hold in their collections, which means all kinds
>>> of published cultural heritage, and even some non-published materials).
>
>> But in this way, fragments should not be allowed at all, as assigning
>> fragments to all (or not all) elements in the namespace (when there
>> could thousands of them) does not seem very great perspective to me.
>
> We both agree that this kind of random usage of <fragment> would not be
> sensible.
>
> When the namespace belongs to a well-specified standard such as ISBN,
> the namespace registration could / should specify whether fragments are
> allowed at all, and if they are, how. With more vague namespaces such as
> for instance NBN it is harder to specify how fragments will be used,
> since we know neither the syntax of the identifier nor the resources
> NBNs will be applied to. If RFC2141bis allows the usage of fragments,
> NBN namespace registration should specify at least in broad terms some
> principles for the usage of fragments within that particular namespace.
>
> Juha

-- 
#-# Martin J. Dürst, Professor, Aoyama Gakuin University
#-# http://www.sw.it.aoyama.ac.jp   mailto:duerst@it.aoyama.ac.jp

From juha.hakala@helsinki.fi  Mon Jan 10 02:54:48 2011
Return-Path: <juha.hakala@helsinki.fi>
X-Original-To: urn@core3.amsl.com
Delivered-To: urn@core3.amsl.com
Received: from localhost (localhost [127.0.0.1]) by core3.amsl.com (Postfix) with ESMTP id B1E4528C137 for <urn@core3.amsl.com>; Mon, 10 Jan 2011 02:54:48 -0800 (PST)
X-Virus-Scanned: amavisd-new at amsl.com
X-Spam-Flag: NO
X-Spam-Score: -2.539
X-Spam-Level: 
X-Spam-Status: No, score=-2.539 tagged_above=-999 required=5 tests=[AWL=0.060,  BAYES_00=-2.599]
Received: from mail.ietf.org ([64.170.98.32]) by localhost (core3.amsl.com [127.0.0.1]) (amavisd-new, port 10024) with ESMTP id vufFPF5mJg7G for <urn@core3.amsl.com>; Mon, 10 Jan 2011 02:54:47 -0800 (PST)
Received: from smtp-rs1.it.helsinki.fi (smtp-rs1-vallila2.fe.helsinki.fi [128.214.173.75]) by core3.amsl.com (Postfix) with ESMTP id 40DA528C143 for <urn@ietf.org>; Mon, 10 Jan 2011 02:54:45 -0800 (PST)
Received: from [128.214.91.90] (kkkl25.lib.helsinki.fi [128.214.91.90]) by smtp-rs1.it.helsinki.fi (8.13.1/8.13.1) with ESMTP id p0AAusxs016895 (version=TLSv1/SSLv3 cipher=DHE-RSA-AES256-SHA bits=256 verify=NOT); Mon, 10 Jan 2011 12:56:54 +0200
Message-ID: <4D2AE5F5.1090200@helsinki.fi>
Date: Mon, 10 Jan 2011 12:56:53 +0200
From: Juha Hakala <juha.hakala@helsinki.fi>
User-Agent: Thunderbird 2.0.0.24 (Windows/20100228)
MIME-Version: 1.0
To: Mykyta Yevstifeyev <evnikita2@gmail.com>
References: <4D131B01.3070502@helsinki.fi> <4D132D59.7020906@gmail.com> <4D133CF7.7010302@helsinki.fi> <4D136A1D.2060905@gmail.com> <4D244392.6000701@helsinki.fi> <4D288C39.30406@gmail.com>
In-Reply-To: <4D288C39.30406@gmail.com>
Content-Type: text/plain; charset=ISO-8859-1; format=flowed
Content-Transfer-Encoding: 7bit
Cc: urn@ietf.org
Subject: Re: [urn] Comments to the draft-ietf-urnbis-rfc2141bis-urn-00.txt
X-BeenThere: urn@ietf.org
X-Mailman-Version: 2.1.9
Precedence: list
List-Id: Discussions about possible revisions to the definition of Uniform Resource Names <urn.ietf.org>
List-Unsubscribe: <https://www.ietf.org/mailman/listinfo/urn>, <mailto:urn-request@ietf.org?subject=unsubscribe>
List-Archive: <http://www.ietf.org/mail-archive/web/urn>
List-Post: <mailto:urn@ietf.org>
List-Help: <mailto:urn-request@ietf.org?subject=help>
List-Subscribe: <https://www.ietf.org/mailman/listinfo/urn>, <mailto:urn-request@ietf.org?subject=subscribe>
X-List-Received-Date: Mon, 10 Jan 2011 10:54:48 -0000

Dear Mykyta (and Martin),

See some answers and comments below.

Mykyta Yevstifeyev wrote:
> Juha,
> 
> Please find some comments below.

<snip>

>> So my view is that each manifestation of a work should have one but 
>> only one identifier. When a new manifestation is produced via 
>> migration, it should get a new identifier, because migration may have 
>> changed the look and feel and/or the content of the resource. So 
>> eventually there will be a set of identifiers, consisting of one 
>> identifier referring to each manifestation of the work, and probably 
>> also one  belonging to the work itself (and its expressions, if any).

> I agree with you here. But if you say that we need the identifier for 
> each version (or you say 'manifestation') of resource, why do we need 
> the generic one?

The generic identifier is for the work itself, which library community 
regards to be separate from its physical manifestations. Embodiment of 
the work is a metadata record describing it. There are applications 
which are capable of extracting work level metadata from manifestation 
level records.

As a persistent identifier system URN is in a way in an advantageous 
position compared with for instance DOI and ARK because URN is based on 
existing identifiers (specified as namespaces). Therefore we do not need 
to think the scope of the URN; scope decisions have been made in 
identifier communities. URN does cover books because there is a 
namespace for ISBN, but it does not cover textual works yet since there 
is no namespace for the ISTC (International Standard Text Code).

For the time being none of the ISO work level identifiers (ISWC, ISAN, 
ISTC) have URN namespaces, but IMHO they would qualify. From the point 
of view of the library community tere is a need to identify the work 
itself separately from its manifestations; work level metadata record 
should have a unique access key, and such record is also an ideal place 
for links to all the manifestations related to the work. One might argue 
that if a link has to be really persistent, it should be made to the 
work level metadata record, since such a simple resource we should be 
able to preserve for centuries, and from there there will be links to 
those manifestations that are available at that point of time.

 From the preservation point of view, providing links between works and 
expressions may become useful in the long run (in the national 
libraries, the time scale being centuries, not decades). For instance, 
in a distant future we may no longer have Gone with the wind (the 
movie), but Gone with the wind (the novel) could still give an idea of 
what the movie was about. Or we may not have been able to preserve a 
digital copy of the Finnish translation of the novel, in which case we 
should be able to direct the user to the printed book (another 
manifestation of the resource) or to the original text in English 
(another expression of the work) which may at that point be available in 
the Web.

>> Short lifetime of digital manifestations is a headache for persistent 
>> linking, when "persistent" means at least several decades and often 
>> centuries. One solution would be to make the permanent link to work / 
>> expression; that is, to metadata record which is never changed apart 
>> from new links added whenever there are new manifestations available.

> But the question is who will record metadata? Different URN resolution 
> services will have different metadata databases. Who or what will 
> synchronize them?

There is no single answer to this, but for certain kind of resources we 
have candidates. The basic requirement is persistence, not superior 
technical skills.

National libraries are responsible of creating national bibliographies, 
which contain metadata about books, serials, etc published in the 
country in question. This data is routinely shared with other libraries 
and union catalogue hosts. In a few years' time, these (fairly 
persistent) databases will contain work level metadata, linked to the 
manifestation level metadata records.

There are already national libraries applying URNs in their national 
bibliographies and other databases, and the number of libraries using 
them may grow in the future.

There are many URN namespaces which the national libraries are not 
using. These identifier systems may or may not share the "world view" 
the libraries have. URN system as such is flexible and can accommodate 
different approaches. But URN services specified should allow the 
different communities to do the things they need, as long as they are 
not in conflict with URN basics.


>>>>>> Query and fragment

>> When the namespace belongs to a well-specified standard such as ISBN, 
>> the namespace registration could / should specify whether fragments 
>> are allowed at all, and if they are, how. With more vague namespaces 
>> such as for instance NBN it is harder to specify how fragments will be 
>> used, since we know neither the syntax of the identifier nor the 
>> resources NBNs will be applied to. If RFC2141bis allows the usage of 
>> fragments, NBN namespace registration should specify at least in broad 
>> terms some principles for the usage of fragments within that 
>> particular namespace.

> In this case we should define smth like this: "The URN specification MAY 
> define the usage of <fragmets>. In this way it SHALL either specify 
> their generic syntax or define procedures for assignment fragments to 
> segregate resources."

Something like this, yes. We still need to consider how to formulate this.

Martin Duerst said:

> May I remind you of the basics of fragment identifiers?
> 
> The meaning of fragment identifiers, which I would understand to include their presence or absence, is determined by the MIME media type of the representation returned when resolving an URI. So in general, very, very little if anything will be needed in terms of considerations for fragment identifiers in URN namespace specs. As for the URN spec itself, the most important thing to say is just what's said above. 

I am not sure that this approach is sufficient. When a URN / URI is 
resolved the representation returned may be a complex resource 
consisting of several files in different MIME types, all incorporated 
within a (METS) container. Thus for us a fragment could be a JPEG 2000 
image which is a part of a digitized article which is a part of a serial 
issue which is in its entirety encapsulated in a single METS container.

Best regards,

Juha
> 
> Mykyta Yevstifeyev
>>
>> Juha
> 
> 

-- 

  Juha Hakala
  Senior advisor, standardisation and IT

  The National Library of Finland
  P.O.Box 15 (Unioninkatu 36, room 503), FIN-00014 Helsinki University
  Email juha.hakala@helsinki.fi, tel +358 50 382 7678

From evnikita2@gmail.com  Sun Jan 23 19:29:55 2011
Return-Path: <evnikita2@gmail.com>
X-Original-To: urn@core3.amsl.com
Delivered-To: urn@core3.amsl.com
Received: from localhost (localhost [127.0.0.1]) by core3.amsl.com (Postfix) with ESMTP id 7B7C83A6A26 for <urn@core3.amsl.com>; Sun, 23 Jan 2011 19:29:55 -0800 (PST)
X-Virus-Scanned: amavisd-new at amsl.com
X-Spam-Flag: NO
X-Spam-Score: -3.217
X-Spam-Level: 
X-Spam-Status: No, score=-3.217 tagged_above=-999 required=5 tests=[AWL=0.382,  BAYES_00=-2.599, RCVD_IN_DNSWL_LOW=-1]
Received: from mail.ietf.org ([64.170.98.32]) by localhost (core3.amsl.com [127.0.0.1]) (amavisd-new, port 10024) with ESMTP id ytzZNJTLVt4J for <urn@core3.amsl.com>; Sun, 23 Jan 2011 19:29:53 -0800 (PST)
Received: from mail-fx0-f44.google.com (mail-fx0-f44.google.com [209.85.161.44]) by core3.amsl.com (Postfix) with ESMTP id 8ED533A689B for <urn@ietf.org>; Sun, 23 Jan 2011 19:29:52 -0800 (PST)
Received: by fxm9 with SMTP id 9so3920669fxm.31 for <urn@ietf.org>; Sun, 23 Jan 2011 19:32:45 -0800 (PST)
DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=gamma; h=domainkey-signature:message-id:date:from:user-agent:mime-version:to :cc:subject:references:in-reply-to:content-type :content-transfer-encoding; bh=BYnHe3gvRBe0MkMw40bldDVD+2PtXGUWizZXTSpsD3A=; b=pVCP0Pr5UT1IpNzRIxzJgwt+tKm0/uQxpHgqpWuy4gbxFvH3YBqq9M5KNEIEm2EN1a 3WuHLq3mu/hv1udtDPcKxv0wfO7f2J/WMDPYPPZQ41cLwAGkbTdu1/4I+J28ztX0kseF GdPaoLiAsnDogbjO3vVzPNH5nGGbnJXZQWe3U=
DomainKey-Signature: a=rsa-sha1; c=nofws; d=gmail.com; s=gamma; h=message-id:date:from:user-agent:mime-version:to:cc:subject :references:in-reply-to:content-type:content-transfer-encoding; b=qFM9kt+d4Rf8nZfV39FWh6ZchMAXgs1OfB/vSpAj1KtyP/FPcJ6i54RjFyQe9FGEqI 5V3nc7PyYXbzHKuKP3i4WQaG2GwH0mDikqLtSnUDUUr9nVvwikaQjEv+goCNonZJvhOg vPwS5hqWEaWiONaU0IIVkuheoKPHlKQkyyae0=
Received: by 10.223.79.6 with SMTP id n6mr3585962fak.122.1295839965470; Sun, 23 Jan 2011 19:32:45 -0800 (PST)
Received: from [127.0.0.1] ([195.191.104.134]) by mx.google.com with ESMTPS id c11sm4354117fav.26.2011.01.23.19.32.43 (version=SSLv3 cipher=RC4-MD5); Sun, 23 Jan 2011 19:32:44 -0800 (PST)
Message-ID: <4D3CF2F2.1080407@gmail.com>
Date: Mon, 24 Jan 2011 05:33:06 +0200
From: Mykyta Yevstifeyev <evnikita2@gmail.com>
User-Agent: Mozilla/5.0 (Windows; U; Windows NT 5.1; ru; rv:1.9.2.13) Gecko/20101207 Thunderbird/3.1.7
MIME-Version: 1.0
To: Juha Hakala <juha.hakala@helsinki.fi>
References: <4D131B01.3070502@helsinki.fi> <4D132D59.7020906@gmail.com> <4D133CF7.7010302@helsinki.fi> <4D136A1D.2060905@gmail.com> <4D244392.6000701@helsinki.fi> <4D288C39.30406@gmail.com> <4D2AE5F5.1090200@helsinki.fi>
In-Reply-To: <4D2AE5F5.1090200@helsinki.fi>
Content-Type: text/plain; charset=ISO-8859-1; format=flowed
Content-Transfer-Encoding: 7bit
Cc: urn@ietf.org
Subject: Re: [urn] Comments to the draft-ietf-urnbis-rfc2141bis-urn-00.txt
X-BeenThere: urn@ietf.org
X-Mailman-Version: 2.1.9
Precedence: list
List-Id: Discussions about possible revisions to the definition of Uniform Resource Names <urn.ietf.org>
List-Unsubscribe: <https://www.ietf.org/mailman/listinfo/urn>, <mailto:urn-request@ietf.org?subject=unsubscribe>
List-Archive: <http://www.ietf.org/mail-archive/web/urn>
List-Post: <mailto:urn@ietf.org>
List-Help: <mailto:urn-request@ietf.org?subject=help>
List-Subscribe: <https://www.ietf.org/mailman/listinfo/urn>, <mailto:urn-request@ietf.org?subject=subscribe>
X-List-Received-Date: Mon, 24 Jan 2011 03:29:55 -0000

Hello Juha,

Sorry for late answer and find some comments below.

10.01.2011 12:56, Juha Hakala wrote:
> Dear Mykyta (and Martin),
>
> See some answers and comments below.
>
> Mykyta Yevstifeyev wrote:
>> Juha,
>>
>> Please find some comments below.
>
> <snip>
>
>>> So my view is that each manifestation of a work should have one but 
>>> only one identifier. When a new manifestation is produced via 
>>> migration, it should get a new identifier, because migration may 
>>> have changed the look and feel and/or the content of the resource. 
>>> So eventually there will be a set of identifiers, consisting of one 
>>> identifier referring to each manifestation of the work, and probably 
>>> also one  belonging to the work itself (and its expressions, if any).
>
>> I agree with you here. But if you say that we need the identifier for 
>> each version (or you say 'manifestation') of resource, why do we need 
>> the generic one?
>
> The generic identifier is for the work itself, which library community 
> regards to be separate from its physical manifestations. Embodiment of 
> the work is a metadata record describing it. There are applications 
> which are capable of extracting work level metadata from manifestation 
> level records.
>
> As a persistent identifier system URN is in a way in an advantageous 
> position compared with for instance DOI and ARK because URN is based 
> on existing identifiers (specified as namespaces). Therefore we do not 
> need to think the scope of the URN; scope decisions have been made in 
> identifier communities. URN does cover books because there is a 
> namespace for ISBN, but it does not cover textual works yet since 
> there is no namespace for the ISTC (International Standard Text Code).
>
> For the time being none of the ISO work level identifiers (ISWC, ISAN, 
> ISTC) have URN namespaces, but IMHO they would qualify. From the point 
> of view of the library community tere is a need to identify the work 
> itself separately from its manifestations; work level metadata record 
> should have a unique access key, and such record is also an ideal 
> place for links to all the manifestations related to the work. One 
> might argue that if a link has to be really persistent, it should be 
> made to the work level metadata record, since such a simple resource 
> we should be able to preserve for centuries, and from there there will 
> be links to those manifestations that are available at that point of 
> time.
>
> From the preservation point of view, providing links between works and 
> expressions may become useful in the long run (in the national 
> libraries, the time scale being centuries, not decades). For instance, 
> in a distant future we may no longer have Gone with the wind (the 
> movie), but Gone with the wind (the novel) could still give an idea of 
> what the movie was about. Or we may not have been able to preserve a 
> digital copy of the Finnish translation of the novel, in which case we 
> should be able to direct the user to the printed book (another 
> manifestation of the resource) or to the original text in English 
> (another expression of the work) which may at that point be available 
> in the Web.
See below.
>
>>> Short lifetime of digital manifestations is a headache for 
>>> persistent linking, when "persistent" means at least several decades 
>>> and often centuries. One solution would be to make the permanent 
>>> link to work / expression; that is, to metadata record which is 
>>> never changed apart from new links added whenever there are new 
>>> manifestations available.
>
>> But the question is who will record metadata? Different URN 
>> resolution services will have different metadata databases. Who or 
>> what will synchronize them?
>
> There is no single answer to this, but for certain kind of resources 
> we have candidates. The basic requirement is persistence, not superior 
> technical skills.
>
> National libraries are responsible of creating national 
> bibliographies, which contain metadata about books, serials, etc 
> published in the country in question. This data is routinely shared 
> with other libraries and union catalogue hosts. In a few years' time, 
> these (fairly persistent) databases will contain work level metadata, 
> linked to the manifestation level metadata records.
>
> There are already national libraries applying URNs in their national 
> bibliographies and other databases, and the number of libraries using 
> them may grow in the future.
>
> There are many URN namespaces which the national libraries are not 
> using. These identifier systems may or may not share the "world view" 
> the libraries have. URN system as such is flexible and can accommodate 
> different approaches. But URN services specified should allow the 
> different communities to do the things they need, as long as they are 
> not in conflict with URN basics.
However if we want to establish the database for URN-referenced 
resources' metalinks, that should be smth. that will have the access to 
all URN-referenced resources and smbd. who should have enough energy to 
perform this type of work.

The national libraries you speak about does not seem to be a candidate 
for synchronizing all the metadata and metalinks of all the resources 
linked to via URNs.  And I can hardly imagine that some organization 
will perform this work in the nearest future.  Therefore, despite this 
idea is quite good and interesting, but not now, IMO.
>
>
>>>>>>> Query and fragment
>
>>> When the namespace belongs to a well-specified standard such as 
>>> ISBN, the namespace registration could / should specify whether 
>>> fragments are allowed at all, and if they are, how. With more vague 
>>> namespaces such as for instance NBN it is harder to specify how 
>>> fragments will be used, since we know neither the syntax of the 
>>> identifier nor the resources NBNs will be applied to. If RFC2141bis 
>>> allows the usage of fragments, NBN namespace registration should 
>>> specify at least in broad terms some principles for the usage of 
>>> fragments within that particular namespace.
>
>> In this case we should define smth like this: "The URN specification 
>> MAY define the usage of <fragmets>. In this way it SHALL either 
>> specify their generic syntax or define procedures for assignment 
>> fragments to segregate resources."
>
> Something like this, yes. We still need to consider how to formulate 
> this.
If we consider this appropriate, we will have to modify the URN 
Namespaces IDs registry and add the value 'Use of Fragments/Queries' 
(separately).  And if we do in this way, we should investigate what URN 
namespaces already use these parts.

All the best,
Mykyta Yevstifeyev
>
> Martin Duerst said:
>
>> May I remind you of the basics of fragment identifiers?
>>
>> The meaning of fragment identifiers, which I would understand to 
>> include their presence or absence, is determined by the MIME media 
>> type of the representation returned when resolving an URI. So in 
>> general, very, very little if anything will be needed in terms of 
>> considerations for fragment identifiers in URN namespace specs. As 
>> for the URN spec itself, the most important thing to say is just 
>> what's said above. 
>
> I am not sure that this approach is sufficient. When a URN / URI is 
> resolved the representation returned may be a complex resource 
> consisting of several files in different MIME types, all incorporated 
> within a (METS) container. Thus for us a fragment could be a JPEG 2000 
> image which is a part of a digitized article which is a part of a 
> serial issue which is in its entirety encapsulated in a single METS 
> container.
>
> Best regards,
>
> Juha
>>
>> Mykyta Yevstifeyev
>>>
>>> Juha
>>
>>
>

