From discuss-bounces@apps.ietf.org Thu Feb 01 01:19:50 2007
Received: from [127.0.0.1] (helo=stiedprmman1.va.neustar.com)
	by megatron.ietf.org with esmtp (Exim 4.43)
	id 1HCVHY-00004v-As; Thu, 01 Feb 2007 01:18:52 -0500
Received: from [10.91.34.44] (helo=ietf-mx.ietf.org)
	by megatron.ietf.org with esmtp (Exim 4.43) id 1HCVHX-00004q-ER
	for discuss@apps.ietf.org; Thu, 01 Feb 2007 01:18:51 -0500
Received: from nz-out-0506.google.com ([64.233.162.234])
	by ietf-mx.ietf.org with esmtp (Exim 4.43) id 1HCVHW-0004tx-7c
	for discuss@apps.ietf.org; Thu, 01 Feb 2007 01:18:51 -0500
Received: by nz-out-0506.google.com with SMTP id z3so389333nzf
	for <discuss@apps.ietf.org>; Wed, 31 Jan 2007 22:18:49 -0800 (PST)
DomainKey-Signature: a=rsa-sha1; c=nofws; d=gmail.com; s=beta;
	h=received:message-id:date:from:sender:to:subject:cc:in-reply-to:mime-version:content-type:content-transfer-encoding:content-disposition:references:x-google-sender-auth;
	b=IXtQ1nM81jMzD1x6MusCigzYBdhYISfewTpRbwN71yVAaf07+uBaR68fOSwNSzEL4mEgYrNxuJD1hFdcgTHTbAKIdtPV86wPpPvHNEhZw+Ij9C9OY1I29ZDh8T09TOD1cb6boqosMa3BSeeO6mgjJ1UwQjMTmmR03/D/z/d8nI0=
Received: by 10.35.84.16 with SMTP id m16mr3401929pyl.1170310728852;
	Wed, 31 Jan 2007 22:18:48 -0800 (PST)
Received: by 10.35.71.14 with HTTP; Wed, 31 Jan 2007 22:18:48 -0800 (PST)
Message-ID: <517bf110701312218l5b8525b8p3a72e48ad81d3038@mail.gmail.com>
Date: Wed, 31 Jan 2007 22:18:48 -0800
From: "Tim Bray" <tbray@textuality.com>
To: "John C Klensin" <john-ietf@jck.com>
Subject: Re: New draft (Was: I-D ACTION:draft-klensin-unicode-escapes-00.txt
In-Reply-To: <E4790BD63A92B0F55375CE85@p3.JCK.COM>
MIME-Version: 1.0
Content-Type: text/plain; charset=ISO-8859-1; format=flowed
Content-Transfer-Encoding: 7bit
Content-Disposition: inline
References: <E4790BD63A92B0F55375CE85@p3.JCK.COM>
X-Google-Sender-Auth: 279c9c582547ecbc
X-Spam-Score: 0.0 (/)
X-Scan-Signature: 7d33c50f3756db14428398e2bdedd581
Cc: discuss@apps.ietf.org
X-BeenThere: discuss@apps.ietf.org
X-Mailman-Version: 2.1.5
Precedence: list
List-Id: general discussion of application-layer protocols
	<discuss.apps.ietf.org>
List-Unsubscribe: <https://www1.ietf.org/mailman/listinfo/discuss>,
	<mailto:discuss-request@apps.ietf.org?subject=unsubscribe>
List-Post: <mailto:discuss@apps.ietf.org>
List-Help: <mailto:discuss-request@apps.ietf.org?subject=help>
List-Subscribe: <https://www1.ietf.org/mailman/listinfo/discuss>,
	<mailto:discuss-request@apps.ietf.org?subject=subscribe>
Errors-To: discuss-bounces@apps.ietf.org

On 1/31/07, John C Klensin <john-ietf@jck.com> wrote:

> While I think I agree with you about your second proposed
> paragraph above ("New protocols..."), I think my instructions
> with this document is to keep it narrow and to focus on escapes,
> not on general advice to protocol designers about Unicode or
> internationalization more broadly.   So I don't want to go so
> far as to make specific (or even specific-sounding) suggestions.

I hadn't read 2277 in years; having done so, I think that it says what
I was trying to say quite effectively.  De facto, this spec is really
only usable for text (in the 2277 sense) when internationalizing
existing protocols.

> This effort, and some others, have convinced me that we are
> getting closer to the time at which RFC 2277/ BCP 18 needs to be
> reopened, reviewed, and updated, but this document isn't the
> right place to do it, at least IMO.

Really? Having just re-read that, I found little to disagree with or
want to change.  My pain point is 2223, but everyone knows that. -Tim




From discuss-bounces@apps.ietf.org Thu Feb 01 11:10:06 2007
Received: from [127.0.0.1] (helo=stiedprmman1.va.neustar.com)
	by megatron.ietf.org with esmtp (Exim 4.43)
	id 1HCeUZ-0005DR-5W; Thu, 01 Feb 2007 11:08:55 -0500
Received: from [10.91.34.44] (helo=ietf-mx.ietf.org)
	by megatron.ietf.org with esmtp (Exim 4.43) id 1HCeUX-0005DL-Bw
	for discuss@apps.ietf.org; Thu, 01 Feb 2007 11:08:53 -0500
Received: from ns.jck.com ([209.187.148.211] helo=bs.jck.com)
	by ietf-mx.ietf.org with esmtp (Exim 4.43) id 1HCeUU-0006vv-0c
	for discuss@apps.ietf.org; Thu, 01 Feb 2007 11:08:53 -0500
Received: from [127.0.0.1] (helo=p3.JCK.COM)
	by bs.jck.com with esmtp (Exim 4.34)
	id 1HCeUT-0002G5-2a; Thu, 01 Feb 2007 11:08:49 -0500
Date: Thu, 01 Feb 2007 11:08:48 -0500
From: John C Klensin <john-ietf@jck.com>
To: Tim Bray <tbray@textuality.com>
Subject: Re: New draft (Was: I-D
 ACTION:draft-klensin-unicode-escapes-00.txt
Message-ID: <E836EBB0222DFF648B2B9A1C@p3.JCK.COM>
In-Reply-To: <517bf110701312218l5b8525b8p3a72e48ad81d3038@mail.gmail.com>
References: <E4790BD63A92B0F55375CE85@p3.JCK.COM>
	<517bf110701312218l5b8525b8p3a72e48ad81d3038@mail.gmail.com>
X-Mailer: Mulberry/4.0.7 (Win32)
MIME-Version: 1.0
Content-Type: text/plain; charset=us-ascii
Content-Transfer-Encoding: 7bit
Content-Disposition: inline
X-Spam-Score: 0.0 (/)
X-Scan-Signature: a7d6aff76b15f3f56fcb94490e1052e4
Cc: discuss@apps.ietf.org
X-BeenThere: discuss@apps.ietf.org
X-Mailman-Version: 2.1.5
Precedence: list
List-Id: general discussion of application-layer protocols
	<discuss.apps.ietf.org>
List-Unsubscribe: <https://www1.ietf.org/mailman/listinfo/discuss>,
	<mailto:discuss-request@apps.ietf.org?subject=unsubscribe>
List-Post: <mailto:discuss@apps.ietf.org>
List-Help: <mailto:discuss-request@apps.ietf.org?subject=help>
List-Subscribe: <https://www1.ietf.org/mailman/listinfo/discuss>,
	<mailto:discuss-request@apps.ietf.org?subject=subscribe>
Errors-To: discuss-bounces@apps.ietf.org



--On Wednesday, 31 January, 2007 22:18 -0800 Tim Bray
<tbray@textuality.com> wrote:

> On 1/31/07, John C Klensin <john-ietf@jck.com> wrote:
> 
>> While I think I agree with you about your second proposed
>> paragraph above ("New protocols..."), I think my instructions
>> with this document is to keep it narrow and to focus on
>> escapes, not on general advice to protocol designers about
>> Unicode or internationalization more broadly.   So I don't
>> want to go so far as to make specific (or even
>> specific-sounding) suggestions.
> 
> I hadn't read 2277 in years; having done so, I think that it
> says what
> I was trying to say quite effectively.  De facto, this spec is
> really
> only usable for text (in the 2277 sense) when
> internationalizing existing protocols.

That is more or less what the text says now... watch for -02
probably sometime next week.

>> This effort, and some others, have convinced me that we are
>> getting closer to the time at which RFC 2277/ BCP 18 needs to
>> be reopened, reviewed, and updated, but this document isn't
>> the right place to do it, at least IMO.
> 
> Really? Having just re-read that, I found little to disagree
> with or
> want to change.  My pain point is 2223, but everyone knows
> that. -Tim

To give just one example, I think there is a case to be made
about making/keeping text NFC complaint or at least examining
that question on a per-protocol basis.

    john






From discuss-bounces@apps.ietf.org Fri Feb 02 07:22:42 2007
Received: from [127.0.0.1] (helo=stiedprmman1.va.neustar.com)
	by megatron.ietf.org with esmtp (Exim 4.43)
	id 1HCxQ5-0000JK-H7; Fri, 02 Feb 2007 07:21:33 -0500
Received: from [10.91.34.44] (helo=ietf-mx.ietf.org)
	by megatron.ietf.org with esmtp (Exim 4.43) id 1HCxQ4-0000I3-89
	for discuss@apps.ietf.org; Fri, 02 Feb 2007 07:21:32 -0500
Received: from anchor-internal-1.mail.demon.net ([195.173.56.100])
	by ietf-mx.ietf.org with esmtp (Exim 4.43) id 1HCxQ2-0000gV-Rp
	for discuss@apps.ietf.org; Fri, 02 Feb 2007 07:21:32 -0500
Received: from finch-staff-1.server.demon.net (finch-staff-1.server.demon.net [193.195.224.1])
	by anchor-internal-1.mail.demon.net with ESMTPœ id l12Bl83R024325Fri, 2 Feb 2007 11:47:08 GMT
Received: from clive by finch-staff-1.server.demon.net with local (Exim 3.36
	#1) id 1HCwsc-0009c1-00; Fri, 02 Feb 2007 11:46:58 +0000
Date: Fri, 2 Feb 2007 11:46:58 +0000
From: "Clive D.W. Feather" <clive@demon.net>
To: John C Klensin <john-ietf@jck.com>
Subject: Re: New draft (Was: I-D ACTION:draft-klensin-unicode-escapes-00.txt
Message-ID: <20070202114658.GX7742@finch-staff-1.thus.net>
References: <875A124D75A8B481E176CF06@p3.JCK.COM>
	<uppsr2hs59srbd7eufbcul5a1ekl7i09nl@hive.bjoern.hoehrmann.de>
	<EF59DA6FD89C4F19750C68C3@p3.JCK.COM>
Mime-Version: 1.0
Content-Type: text/plain; charset=us-ascii
Content-Disposition: inline
In-Reply-To: <EF59DA6FD89C4F19750C68C3@p3.JCK.COM>
User-Agent: Mutt/1.5.3i
X-Spam-Score: 0.0 (/)
X-Scan-Signature: 7655788c23eb79e336f5f8ba8bce7906
Cc: discuss@apps.ietf.org
X-BeenThere: discuss@apps.ietf.org
X-Mailman-Version: 2.1.5
Precedence: list
List-Id: general discussion of application-layer protocols
	<discuss.apps.ietf.org>
List-Unsubscribe: <https://www1.ietf.org/mailman/listinfo/discuss>,
	<mailto:discuss-request@apps.ietf.org?subject=unsubscribe>
List-Post: <mailto:discuss@apps.ietf.org>
List-Help: <mailto:discuss-request@apps.ietf.org?subject=help>
List-Subscribe: <https://www1.ietf.org/mailman/listinfo/discuss>,
	<mailto:discuss-request@apps.ietf.org?subject=subscribe>
Errors-To: discuss-bounces@apps.ietf.org

>> In section 5.2 it is said HTML uses the &#xNNNN; form and that
>> this form has a clear terminator. This is not really false but
>> HTML allows to omit the terminator if it is not needed, for
>> example <p>Bj&#xf6rn</p> is also valid. I would suggest to
>> mention only XML or note that HTML's mechanism is similar to
>> that of XML.

If you check the HTML specification (section 5.3), it says that SGML
allows the semicolon to be omitted in certain contexts, but "strongly
suggest" not to do that.

In particular, the example
    <p>Bj&#xf6rn</p>
is not valid because it lies in the middle of a word. A permitted case
would be:
    <p>Bj&#xf6</p>
where the tag begin symbol < ends the entity.

-- 
Clive D.W. Feather  | Work:  <clive@demon.net>   | Tel:    +44 20 8495 6138
Internet Expert     | Home:  <clive@davros.org>  | Fax:    +44 870 051 9937
Demon Internet      | WWW: http://www.davros.org | Mobile: +44 7973 377646
THUS plc            |                            |




From discuss-bounces@apps.ietf.org Fri Feb 02 07:41:36 2007
Received: from [127.0.0.1] (helo=stiedprmman1.va.neustar.com)
	by megatron.ietf.org with esmtp (Exim 4.43)
	id 1HCxih-0002r8-IM; Fri, 02 Feb 2007 07:40:47 -0500
Received: from [10.91.34.44] (helo=ietf-mx.ietf.org)
	by megatron.ietf.org with esmtp (Exim 4.43) id 1HCxih-0002r3-8q
	for discuss@apps.ietf.org; Fri, 02 Feb 2007 07:40:47 -0500
Received: from anchor-internal-1.mail.demon.net ([195.173.56.100])
	by ietf-mx.ietf.org with esmtp (Exim 4.43) id 1HCxif-00065f-SL
	for discuss@apps.ietf.org; Fri, 02 Feb 2007 07:40:47 -0500
Received: from finch-staff-1.server.demon.net (finch-staff-1.server.demon.net [193.195.224.1])
	by anchor-internal-1.mail.demon.net with ESMTPœ id l12Bd3k9020071Fri, 2 Feb 2007 11:39:03 GMT
Received: from clive by finch-staff-1.server.demon.net with local (Exim 3.36
	#1) id 1HCwkn-0009Ow-00; Fri, 02 Feb 2007 11:38:53 +0000
Date: Fri, 2 Feb 2007 11:38:53 +0000
From: "Clive D.W. Feather" <clive@demon.net>
To: John C Klensin <john-ietf@jck.com>
Subject: Re: New draft (Was: I-D ACTION:draft-klensin-unicode-escapes-00.txt
Message-ID: <20070202113853.GW7742@finch-staff-1.thus.net>
References: <875A124D75A8B481E176CF06@p3.JCK.COM>
Mime-Version: 1.0
Content-Type: text/plain; charset=us-ascii
Content-Disposition: inline
In-Reply-To: <875A124D75A8B481E176CF06@p3.JCK.COM>
User-Agent: Mutt/1.5.3i
X-Spam-Score: 0.0 (/)
X-Scan-Signature: bb8f917bb6b8da28fc948aeffb74aa17
Cc: discuss@apps.ietf.org
X-BeenThere: discuss@apps.ietf.org
X-Mailman-Version: 2.1.5
Precedence: list
List-Id: general discussion of application-layer protocols
	<discuss.apps.ietf.org>
List-Unsubscribe: <https://www1.ietf.org/mailman/listinfo/discuss>,
	<mailto:discuss-request@apps.ietf.org?subject=unsubscribe>
List-Post: <mailto:discuss@apps.ietf.org>
List-Help: <mailto:discuss-request@apps.ietf.org?subject=help>
List-Subscribe: <https://www1.ietf.org/mailman/listinfo/discuss>,
	<mailto:discuss-request@apps.ietf.org?subject=subscribe>
Errors-To: discuss-bounces@apps.ietf.org

John C Klensin said:
> I've just submitted draft-klensin-unicode-escapes-01.txt and
> assume it will show up in the posting directory today or
> tomorrow.  

Some comments for you.

* In 1.1, rather than saying that Unicode occupies "two or more octets",
wouldn't it be better to say "21 bits - rather than the 7 bits of ASCII -"?

* Somewhere in the last two paragraphs of 1.1 you should be talking about
mini-languages (e.g. Cosmogol) as well as protocols and UIs.

* In 3, you're inconsistent between "U+NNN[N[N]]" and "NNN...". Indeed,
shouldn't the former actually be "U+[[N]N]NNNN"? (Note both the order and
the number of Ns.) I would suggest that better wording might be:

    ... U+NN syntax for code point references specified in the Unicode
    Standard, where NN is between four and six hexadecimal digits.

* In 4, second bullet, "string terminators" should be "string delimiters".

* In 5.2, you've said "generally considered ugly and awkward" but I'm not
aware of anyone else who's made that complaint.

* In 6 you need to copy in all the security stuff from Unicode; the stuff
that says that you must use shortest-form UTF-8 (so not using %xC1.A1 for
'A') because of the problems of filters and firewalls not spotting longer
forms.

-- 
Clive D.W. Feather  | Work:  <clive@demon.net>   | Tel:    +44 20 8495 6138
Internet Expert     | Home:  <clive@davros.org>  | Fax:    +44 870 051 9937
Demon Internet      | WWW: http://www.davros.org | Mobile: +44 7973 377646
THUS plc            |                            |




From discuss-bounces@apps.ietf.org Fri Feb 02 07:51:17 2007
Received: from [127.0.0.1] (helo=stiedprmman1.va.neustar.com)
	by megatron.ietf.org with esmtp (Exim 4.43)
	id 1HCxs8-00021G-9N; Fri, 02 Feb 2007 07:50:32 -0500
Received: from [10.91.34.44] (helo=ietf-mx.ietf.org)
	by megatron.ietf.org with esmtp (Exim 4.43) id 1HCxs7-0001wh-B6
	for discuss@apps.ietf.org; Fri, 02 Feb 2007 07:50:31 -0500
Received: from mx2.nic.fr ([192.134.4.11])
	by ietf-mx.ietf.org with esmtp (Exim 4.43) id 1HCxs5-0007zb-2H
	for discuss@apps.ietf.org; Fri, 02 Feb 2007 07:50:31 -0500
Received: from localhost (localhost.localdomain [127.0.0.1])
	by mx2.nic.fr (Postfix) with ESMTP id AF44D26C20D
	for <discuss@apps.ietf.org>; Fri,  2 Feb 2007 13:50:12 +0100 (CET)
Received: from relay2.nic.fr (relay2.nic.fr [192.134.4.163])
	by mx2.nic.fr (Postfix) with ESMTP id 94FB926C173
	for <discuss@apps.ietf.org>; Fri,  2 Feb 2007 13:50:12 +0100 (CET)
Received: from bortzmeyer.nic.fr (batilda.nic.fr [192.134.4.69])
	by relay2.nic.fr (Postfix) with ESMTP id 8750A58ECE7
	for <discuss@apps.ietf.org>; Fri,  2 Feb 2007 13:50:12 +0100 (CET)
Date: Fri, 2 Feb 2007 13:50:12 +0100
From: Stephane Bortzmeyer <bortzmeyer@nic.fr>
To: discuss@apps.ietf.org
Subject: Re: New draft (Was: I-D ACTION:draft-klensin-unicode-escapes-00.txt
Message-ID: <20070202125012.GA18307@nic.fr>
References: <875A124D75A8B481E176CF06@p3.JCK.COM>
	<20070202113853.GW7742@finch-staff-1.thus.net>
MIME-Version: 1.0
Content-Type: text/plain; charset=us-ascii
Content-Disposition: inline
In-Reply-To: <20070202113853.GW7742@finch-staff-1.thus.net>
X-Operating-System: Debian GNU/Linux 4.0
X-Kernel: Linux 2.6.17-2-686 i686
Organization: NIC France
X-URL: http://www.nic.fr/
User-Agent: Mutt/1.5.13 (2006-08-11)
X-Spam-Score: 0.0 (/)
X-Scan-Signature: 1ac7cc0a4cd376402b85bc1961a86ac2
X-BeenThere: discuss@apps.ietf.org
X-Mailman-Version: 2.1.5
Precedence: list
List-Id: general discussion of application-layer protocols
	<discuss.apps.ietf.org>
List-Unsubscribe: <https://www1.ietf.org/mailman/listinfo/discuss>,
	<mailto:discuss-request@apps.ietf.org?subject=unsubscribe>
List-Post: <mailto:discuss@apps.ietf.org>
List-Help: <mailto:discuss-request@apps.ietf.org?subject=help>
List-Subscribe: <https://www1.ietf.org/mailman/listinfo/discuss>,
	<mailto:discuss-request@apps.ietf.org?subject=subscribe>
Errors-To: discuss-bounces@apps.ietf.org

On Fri, Feb 02, 2007 at 11:38:53AM +0000,
 Clive D.W. Feather <clive@demon.net> wrote 
 a message of 35 lines which said:

> * Somewhere in the last two paragraphs of 1.1 you should be talking
> about mini-languages (e.g. Cosmogol) as well as protocols and UIs.

This is clearly a difficult issue since RFC 2277 is clear on:

* text carried by a protocol (i18n is necessary)
* protocol elements (i18n is optional)

but does not mention formats, mini-languages and so on. RFC 4234 is a
good example of a format whose i18n rules are unclear.




From discuss-bounces@apps.ietf.org Fri Feb 02 08:34:55 2007
Received: from [127.0.0.1] (helo=stiedprmman1.va.neustar.com)
	by megatron.ietf.org with esmtp (Exim 4.43)
	id 1HCyYE-0007MJ-Eb; Fri, 02 Feb 2007 08:34:02 -0500
Received: from [10.91.34.44] (helo=ietf-mx.ietf.org)
	by megatron.ietf.org with esmtp (Exim 4.43) id 1HCyYD-0007ME-2c
	for discuss@apps.ietf.org; Fri, 02 Feb 2007 08:34:01 -0500
Received: from main.gmane.org ([80.91.229.2] helo=ciao.gmane.org)
	by ietf-mx.ietf.org with esmtp (Exim 4.43) id 1HCyYA-0001Cb-On
	for discuss@apps.ietf.org; Fri, 02 Feb 2007 08:34:01 -0500
Received: from list by ciao.gmane.org with local (Exim 4.43)
	id 1HCyXs-0005ZF-E3
	for discuss@apps.ietf.org; Fri, 02 Feb 2007 14:33:40 +0100
Received: from d252203.dialin.hansenet.de ([80.171.252.203])
	by main.gmane.org with esmtp (Gmexim 0.1 (Debian))
	id 1AlnuQ-0007hv-00
	for <discuss@apps.ietf.org>; Fri, 02 Feb 2007 14:33:40 +0100
Received: from nobody by d252203.dialin.hansenet.de with local (Gmexim 0.1
	(Debian)) id 1AlnuQ-0007hv-00
	for <discuss@apps.ietf.org>; Fri, 02 Feb 2007 14:33:40 +0100
X-Injected-Via-Gmane: http://gmane.org/
To: discuss@apps.ietf.org
From: Frank Ellermann <nobody@xyzzy.claranet.de>
Subject: I-D.klensin-unicode-escapes (was: New Draft)
Date: Fri, 02 Feb 2007 14:30:52 +0100
Organization: <URL:http://purl.net/xyzzy>
Lines: 29
Message-ID: <45C33D0C.7BF@xyzzy.claranet.de>
References: <875A124D75A8B481E176CF06@p3.JCK.COM>
	<20070202113853.GW7742@finch-staff-1.thus.net>
Mime-Version: 1.0
Content-Type: text/plain; charset=us-ascii
Content-Transfer-Encoding: 7bit
X-Complaints-To: usenet@sea.gmane.org
X-Gmane-NNTP-Posting-Host: d252203.dialin.hansenet.de
X-Mailer: Mozilla 3.0 (OS/2; U)
X-Spam-Score: 0.0 (/)
X-Scan-Signature: ea4ac80f790299f943f0a53be7e1a21a
X-BeenThere: discuss@apps.ietf.org
X-Mailman-Version: 2.1.5
Precedence: list
List-Id: general discussion of application-layer protocols
	<discuss.apps.ietf.org>
List-Unsubscribe: <https://www1.ietf.org/mailman/listinfo/discuss>,
	<mailto:discuss-request@apps.ietf.org?subject=unsubscribe>
List-Post: <mailto:discuss@apps.ietf.org>
List-Help: <mailto:discuss-request@apps.ietf.org?subject=help>
List-Subscribe: <https://www1.ietf.org/mailman/listinfo/discuss>,
	<mailto:discuss-request@apps.ietf.org?subject=subscribe>
Errors-To: discuss-bounces@apps.ietf.org

Clive D.W. Feather wrote:

> In 3, you're inconsistent between "U+NNN[N[N]]" and "NNN...". Indeed,
> shouldn't the former actually be "U+[[N]N]NNNN"? (Note both the order
> and the number of Ns.)

+1

>     ... U+NN syntax for code point references specified in the Unicode
>     Standard, where NN is between four and six hexadecimal digits.

No, folks could misinterpret U+NN as "anything up to 6 digits".

> In 5.2, you've said "generally considered ugly and awkward" but I'm
> not aware of anyone else who's made that complaint.

+1  Obviously John hates it, that would justify "often".  Others don't
like backslash-U for various reasons, not only ugly and awkward, also
confusing (due to various conventions), unclear (lack of delimiter),
and a royal PITA in conjunction with <quoted-string>, when it results
in multiple backslashes.  "Harmful" is worse than "ugly and awkward".

> In 6 you need to copy in all the security stuff from Unicode

IMO not "all", folks are supposed to know RFC 3629, it's a STD.  So far
all attacks on the "three steps" model fortunately failed, STD is STD.

Frank






From discuss-bounces@apps.ietf.org Fri Feb 02 08:34:55 2007
Received: from [127.0.0.1] (helo=stiedprmman1.va.neustar.com)
	by megatron.ietf.org with esmtp (Exim 4.43)
	id 1HCyYG-0007OO-Id; Fri, 02 Feb 2007 08:34:04 -0500
Received: from [10.91.34.44] (helo=ietf-mx.ietf.org)
	by megatron.ietf.org with esmtp (Exim 4.43) id 1HCyYF-0007Mb-0E
	for discuss@apps.ietf.org; Fri, 02 Feb 2007 08:34:03 -0500
Received: from main.gmane.org ([80.91.229.2] helo=ciao.gmane.org)
	by ietf-mx.ietf.org with esmtp (Exim 4.43) id 1HCyYD-0001Cb-Lu
	for discuss@apps.ietf.org; Fri, 02 Feb 2007 08:34:02 -0500
Received: from list by ciao.gmane.org with local (Exim 4.43)
	id 1HCyCi-0000QJ-5h
	for discuss@apps.ietf.org; Fri, 02 Feb 2007 14:11:48 +0100
Received: from d252203.dialin.hansenet.de ([80.171.252.203])
	by main.gmane.org with esmtp (Gmexim 0.1 (Debian))
	id 1AlnuQ-0007hv-00
	for <discuss@apps.ietf.org>; Fri, 02 Feb 2007 14:11:48 +0100
Received: from nobody by d252203.dialin.hansenet.de with local (Gmexim 0.1
	(Debian)) id 1AlnuQ-0007hv-00
	for <discuss@apps.ietf.org>; Fri, 02 Feb 2007 14:11:48 +0100
X-Injected-Via-Gmane: http://gmane.org/
To: discuss@apps.ietf.org
From: Frank Ellermann <nobody@xyzzy.claranet.de>
Subject: I-D.klensin-unicode-escapes (was: New Draft)
Date: Fri, 02 Feb 2007 14:05:34 +0100
Organization: <URL:http://purl.net/xyzzy>
Lines: 34
Message-ID: <45C3371E.330F@xyzzy.claranet.de>
References: <875A124D75A8B481E176CF06@p3.JCK.COM>
	<uppsr2hs59srbd7eufbcul5a1ekl7i09nl@hive.bjoern.hoehrmann.de>
	<EF59DA6FD89C4F19750C68C3@p3.JCK.COM>
	<20070202114658.GX7742@finch-staff-1.thus.net>
Mime-Version: 1.0
Content-Type: text/plain; charset=us-ascii
Content-Transfer-Encoding: 7bit
X-Complaints-To: usenet@sea.gmane.org
X-Gmane-NNTP-Posting-Host: d252203.dialin.hansenet.de
X-Mailer: Mozilla 3.0 (OS/2; U)
X-Spam-Score: 0.0 (/)
X-Scan-Signature: bb8f917bb6b8da28fc948aeffb74aa17
X-BeenThere: discuss@apps.ietf.org
X-Mailman-Version: 2.1.5
Precedence: list
List-Id: general discussion of application-layer protocols
	<discuss.apps.ietf.org>
List-Unsubscribe: <https://www1.ietf.org/mailman/listinfo/discuss>,
	<mailto:discuss-request@apps.ietf.org?subject=unsubscribe>
List-Post: <mailto:discuss@apps.ietf.org>
List-Help: <mailto:discuss-request@apps.ietf.org?subject=help>
List-Subscribe: <https://www1.ietf.org/mailman/listinfo/discuss>,
	<mailto:discuss-request@apps.ietf.org?subject=subscribe>
Errors-To: discuss-bounces@apps.ietf.org

Clive D.W. Feather wrote:

> If you check the HTML specification (section 5.3), it says that SGML
> allows the semicolon to be omitted in certain contexts, but "strongly
> suggest" not to do that.

Yes, never ever mention that HTML exists, it's horrible.  The [Charmod]
bible requires no (SGML) nonsense in http://www.w3.org/TR/charmod/#C044

The I-D should IMO adopt and cite [Charmod] C042 up to C048 verbatim.

A few other conformance criteria in [Charmod] might be also interesting:
http://www.w3.org/TR/charmod/#C070  Don't exclude arbitrary code points
http://www.w3.org/TR/charmod/#C077  Don't allow anything above U+10FFFF
http://www.w3.org/TR/charmod/#C078  Don't (ab)use surrogates
http://www.w3.org/TR/charmod/#C079  Don't (ab)use non-characters

http://www.w3.org/TR/charmod/#C015  n/a (covered by the better RFC 2277)
http://www.w3.org/TR/charmod/#C016  n/a (covered by the better RFC 2277)
http://www.w3.org/TR/charmod/#C017  Stick to working encoding rules
http://www.w3.org/TR/charmod/#C018  n/a (covered by the better RFC 2277)

http://www.w3.org/TR/charmod/#C049  n/a (for the I-D US-ASCII is given)
http://www.w3.org/TR/charmod/#C026  n/a (covered by the better RFC 2277)

Etc.  The "better RFC 2277" idea is a single default UTF-8, instead of a
choice between UTF-8, UTF-16, UTF-16LE, UTF16-BE, UTF-32, UTF-32LE, and
UTF-32BE in [Charmod], let alone hypothetical UTF-32 "2143" or "3412".

Probably the I-D should mention that one famous exception from its rule
to avoid encoded UTF-8 is the URL form of IRIs.

Frank




From discuss-bounces@apps.ietf.org Fri Feb 02 09:16:02 2007
Received: from [127.0.0.1] (helo=stiedprmman1.va.neustar.com)
	by megatron.ietf.org with esmtp (Exim 4.43)
	id 1HCzC1-0007aK-SW; Fri, 02 Feb 2007 09:15:10 -0500
Received: from [10.91.34.44] (helo=ietf-mx.ietf.org)
	by megatron.ietf.org with esmtp (Exim 4.43) id 1HCzBz-0007YX-HH
	for discuss@apps.ietf.org; Fri, 02 Feb 2007 09:15:07 -0500
Received: from main.gmane.org ([80.91.229.2] helo=ciao.gmane.org)
	by ietf-mx.ietf.org with esmtp (Exim 4.43) id 1HCz9f-0001Gh-KX
	for discuss@apps.ietf.org; Fri, 02 Feb 2007 09:12:45 -0500
Received: from list by ciao.gmane.org with local (Exim 4.43)
	id 1HCz8L-0005bG-Al
	for discuss@apps.ietf.org; Fri, 02 Feb 2007 15:11:21 +0100
Received: from d252203.dialin.hansenet.de ([80.171.252.203])
	by main.gmane.org with esmtp (Gmexim 0.1 (Debian))
	id 1AlnuQ-0007hv-00
	for <discuss@apps.ietf.org>; Fri, 02 Feb 2007 15:11:21 +0100
Received: from nobody by d252203.dialin.hansenet.de with local (Gmexim 0.1
	(Debian)) id 1AlnuQ-0007hv-00
	for <discuss@apps.ietf.org>; Fri, 02 Feb 2007 15:11:21 +0100
X-Injected-Via-Gmane: http://gmane.org/
To: discuss@apps.ietf.org
From: Frank Ellermann <nobody@xyzzy.claranet.de>
Subject: ABNF (was: New draft)
Date: Fri, 02 Feb 2007 15:07:37 +0100
Organization: <URL:http://purl.net/xyzzy>
Lines: 28
Message-ID: <45C345A9.1589@xyzzy.claranet.de>
References: <875A124D75A8B481E176CF06@p3.JCK.COM>
	<20070202113853.GW7742@finch-staff-1.thus.net>
	<20070202125012.GA18307@nic.fr>
Mime-Version: 1.0
Content-Type: text/plain; charset=us-ascii
Content-Transfer-Encoding: 7bit
X-Complaints-To: usenet@sea.gmane.org
X-Gmane-NNTP-Posting-Host: d252203.dialin.hansenet.de
X-Mailer: Mozilla 3.0 (OS/2; U)
X-Spam-Score: 0.0 (/)
X-Scan-Signature: 9466e0365fc95844abaf7c3f15a05c7d
X-BeenThere: discuss@apps.ietf.org
X-Mailman-Version: 2.1.5
Precedence: list
List-Id: general discussion of application-layer protocols
	<discuss.apps.ietf.org>
List-Unsubscribe: <https://www1.ietf.org/mailman/listinfo/discuss>,
	<mailto:discuss-request@apps.ietf.org?subject=unsubscribe>
List-Post: <mailto:discuss@apps.ietf.org>
List-Help: <mailto:discuss-request@apps.ietf.org?subject=help>
List-Subscribe: <https://www1.ietf.org/mailman/listinfo/discuss>,
	<mailto:discuss-request@apps.ietf.org?subject=subscribe>
Errors-To: discuss-bounces@apps.ietf.org

Stephane Bortzmeyer wrote:

> RFC 4234 is a good example of a format whose i18n rules are unclear.

| NOTE:
|
|     ABNF strings are case-insensitive and the character set for these
|     strings is us-ascii.

That's clear.  A tricky part could be <name> in chapter 2.2, because...

|     rulename       =  ALPHA *(ALPHA / DIGIT / "-")

...in chapter 4 could be interpreted as different from <name>.  Now I've
found a typo in 4234 chapter 2.4:

   although Appendix A (Core) provides definitions for a 7-bit US-ASCII
   environment as has been common to much of the Internet.

It's Appendix B, not Appendix A (Acknowledgements).  Chapter 4 is based
on Appendix B, a <rulename> is clearly ASCII.  IMO RFC 4234 is fine, its
LWSP is an exception, FWS as in RFC 2822 (excl. obs-FWS) would be better.

Unlike an utter dubious variant in RFCs 2068, 2069, 2616, 2617, and 2831,
where nobody sees the potential damage caused by <LWS> hidden in a #rule.

Frank






From discuss-bounces@apps.ietf.org Fri Feb 02 13:26:10 2007
Received: from [127.0.0.1] (helo=stiedprmman1.va.neustar.com)
	by megatron.ietf.org with esmtp (Exim 4.43)
	id 1HD368-0001Ft-Ng; Fri, 02 Feb 2007 13:25:20 -0500
Received: from [10.91.34.44] (helo=ietf-mx.ietf.org)
	by megatron.ietf.org with esmtp (Exim 4.43) id 1HD367-0001Fm-PU
	for discuss@apps.ietf.org; Fri, 02 Feb 2007 13:25:19 -0500
Received: from anchor-internal-1.mail.demon.net ([195.173.56.100])
	by ietf-mx.ietf.org with esmtp (Exim 4.43) id 1HD366-0006JB-BY
	for discuss@apps.ietf.org; Fri, 02 Feb 2007 13:25:19 -0500
Received: from finch-staff-1.server.demon.net (finch-staff-1.server.demon.net [193.195.224.1])
	by anchor-internal-1.mail.demon.net with ESMTPœ id l12IPGMF024213Fri, 2 Feb 2007 18:25:16 GMT
Received: from clive by finch-staff-1.server.demon.net with local (Exim 3.36
	#1) id 1HD35u-000KZL-00; Fri, 02 Feb 2007 18:25:06 +0000
Date: Fri, 2 Feb 2007 18:25:06 +0000
From: "Clive D.W. Feather" <clive@demon.net>
To: John C Klensin <john-ietf@jck.com>
Subject: Re: New draft (Was: I-D ACTION:draft-klensin-unicode-escapes-00.txt
Message-ID: <20070202182506.GF68544@finch-staff-1.thus.net>
References: <875A124D75A8B481E176CF06@p3.JCK.COM>
Mime-Version: 1.0
Content-Type: text/plain; charset=us-ascii
Content-Disposition: inline
In-Reply-To: <875A124D75A8B481E176CF06@p3.JCK.COM>
User-Agent: Mutt/1.5.3i
X-Spam-Score: 0.0 (/)
X-Scan-Signature: 52e1467c2184c31006318542db5614d5
Cc: discuss@apps.ietf.org
X-BeenThere: discuss@apps.ietf.org
X-Mailman-Version: 2.1.5
Precedence: list
List-Id: general discussion of application-layer protocols
	<discuss.apps.ietf.org>
List-Unsubscribe: <https://www1.ietf.org/mailman/listinfo/discuss>,
	<mailto:discuss-request@apps.ietf.org?subject=unsubscribe>
List-Post: <mailto:discuss@apps.ietf.org>
List-Help: <mailto:discuss-request@apps.ietf.org?subject=help>
List-Subscribe: <https://www1.ietf.org/mailman/listinfo/discuss>,
	<mailto:discuss-request@apps.ietf.org?subject=subscribe>
Errors-To: discuss-bounces@apps.ietf.org

John C Klensin said:
> 	* I have not touched the ABNF associated with the \u /
> 	\U case.  I have inserted an explicit placeholder but,
> 	as discussed on this list, I think we need to figure out
> 	what we want to do and then go back and adjust the
> 	metalanguage productions.   In particular, there has
> 	been one strong suggestion, with which I agree, that we
> 	not take the obvious approach of substituting %x5C.75
> 	for "\u", since the intent is a character string
> 	abstraction (independent of the implementation character
> 	set) rather than specific octets.

I asked Paul Overell about this and got the following answer (in part):

>        "By separating external encoding from the syntax, it is intended
>        that alternate encoding environments can be used for the same
>        syntax."
>
> So although "\" means us-ascii %x5C, the same ABNF may still be used to
> specify the syntax of strings expressed in a different character set by
> specifying the mapping between %x5C and the encoding used, but that is
> outside the scope of ABNF.
   
In other words, you write the ABNF as if the target encoding was ASCII, but
then state somewhere in the document that other encodings may be used and
the ABNF is meant to represent the abstract characters, not specific octet
values.

Or else we either drop ABNF (wrong, I think) or state explicitly that the
notation is "ABNF except for case-sensitivity".

-- 
Clive D.W. Feather  | Work:  <clive@demon.net>   | Tel:    +44 20 8495 6138
Internet Expert     | Home:  <clive@davros.org>  | Fax:    +44 870 051 9937
Demon Internet      | WWW: http://www.davros.org | Mobile: +44 7973 377646
THUS plc            |                            |




From discuss-bounces@apps.ietf.org Fri Feb 02 13:48:42 2007
Received: from [127.0.0.1] (helo=stiedprmman1.va.neustar.com)
	by megatron.ietf.org with esmtp (Exim 4.43)
	id 1HD3Rl-0005Td-Au; Fri, 02 Feb 2007 13:47:41 -0500
Received: from [10.91.34.44] (helo=ietf-mx.ietf.org)
	by megatron.ietf.org with esmtp (Exim 4.43) id 1HD3Rj-0005Rh-Om
	for discuss@apps.ietf.org; Fri, 02 Feb 2007 13:47:39 -0500
Received: from anchor-internal-1.mail.demon.net ([195.173.56.100])
	by ietf-mx.ietf.org with esmtp (Exim 4.43) id 1HD3Ri-000147-CX
	for discuss@apps.ietf.org; Fri, 02 Feb 2007 13:47:39 -0500
Received: from finch-staff-1.server.demon.net (finch-staff-1.server.demon.net [193.195.224.1])
	by anchor-internal-1.mail.demon.net with ESMTPœ id l12Ilb1Z004475Fri, 2 Feb 2007 18:47:37 GMT
Received: from clive by finch-staff-1.server.demon.net with local (Exim 3.36
	#1) id 1HD3RX-000L1K-00; Fri, 02 Feb 2007 18:47:27 +0000
Date: Fri, 2 Feb 2007 18:47:27 +0000
From: "Clive D.W. Feather" <clive@demon.net>
To: Frank Ellermann <nobody@xyzzy.claranet.de>
Subject: Re: I-D.klensin-unicode-escapes (was: New Draft)
Message-ID: <20070202184727.GG68544@finch-staff-1.thus.net>
References: <875A124D75A8B481E176CF06@p3.JCK.COM>
	<uppsr2hs59srbd7eufbcul5a1ekl7i09nl@hive.bjoern.hoehrmann.de>
	<EF59DA6FD89C4F19750C68C3@p3.JCK.COM>
	<20070202114658.GX7742@finch-staff-1.thus.net>
	<45C3371E.330F@xyzzy.claranet.de>
Mime-Version: 1.0
Content-Type: text/plain; charset=us-ascii
Content-Disposition: inline
In-Reply-To: <45C3371E.330F@xyzzy.claranet.de>
User-Agent: Mutt/1.5.3i
X-Spam-Score: 0.0 (/)
X-Scan-Signature: 856eb5f76e7a34990d1d457d8e8e5b7f
Cc: discuss@apps.ietf.org
X-BeenThere: discuss@apps.ietf.org
X-Mailman-Version: 2.1.5
Precedence: list
List-Id: general discussion of application-layer protocols
	<discuss.apps.ietf.org>
List-Unsubscribe: <https://www1.ietf.org/mailman/listinfo/discuss>,
	<mailto:discuss-request@apps.ietf.org?subject=unsubscribe>
List-Post: <mailto:discuss@apps.ietf.org>
List-Help: <mailto:discuss-request@apps.ietf.org?subject=help>
List-Subscribe: <https://www1.ietf.org/mailman/listinfo/discuss>,
	<mailto:discuss-request@apps.ietf.org?subject=subscribe>
Errors-To: discuss-bounces@apps.ietf.org

Frank Ellermann said:
> The I-D should IMO adopt and cite [Charmod] C042 up to C048 verbatim.

C042 would require &#x1234; rather than allowing us to invent \u'1234'.

> A few other conformance criteria in [Charmod] might be also interesting:
> http://www.w3.org/TR/charmod/#C070  Don't exclude arbitrary code points
> http://www.w3.org/TR/charmod/#C077  Don't allow anything above U+10FFFF
> http://www.w3.org/TR/charmod/#C078  Don't (ab)use surrogates
> http://www.w3.org/TR/charmod/#C079  Don't (ab)use non-characters

Those are worth including, I think.

-- 
Clive D.W. Feather  | Work:  <clive@demon.net>   | Tel:    +44 20 8495 6138
Internet Expert     | Home:  <clive@davros.org>  | Fax:    +44 870 051 9937
Demon Internet      | WWW: http://www.davros.org | Mobile: +44 7973 377646
THUS plc            |                            |




From discuss-bounces@apps.ietf.org Fri Feb 02 13:52:02 2007
Received: from [127.0.0.1] (helo=stiedprmman1.va.neustar.com)
	by megatron.ietf.org with esmtp (Exim 4.43)
	id 1HD3VL-0002IN-RY; Fri, 02 Feb 2007 13:51:23 -0500
Received: from [10.91.34.44] (helo=ietf-mx.ietf.org)
	by megatron.ietf.org with esmtp (Exim 4.43) id 1HD3VK-0002IC-5U
	for discuss@apps.ietf.org; Fri, 02 Feb 2007 13:51:22 -0500
Received: from anchor-internal-1.mail.demon.net ([195.173.56.100])
	by ietf-mx.ietf.org with esmtp (Exim 4.43) id 1HD3VH-0001di-NI
	for discuss@apps.ietf.org; Fri, 02 Feb 2007 13:51:22 -0500
Received: from finch-staff-1.server.demon.net (finch-staff-1.server.demon.net [193.195.224.1])
	by anchor-internal-1.mail.demon.net with ESMTPœ id l12IpIhD006349Fri, 2 Feb 2007 18:51:19 GMT
Received: from clive by finch-staff-1.server.demon.net with local (Exim 3.36
	#1) id 1HD3V2-000L5A-00; Fri, 02 Feb 2007 18:51:04 +0000
Date: Fri, 2 Feb 2007 18:51:04 +0000
From: "Clive D.W. Feather" <clive@demon.net>
To: Frank Ellermann <nobody@xyzzy.claranet.de>
Subject: Re: I-D.klensin-unicode-escapes (was: New Draft)
Message-ID: <20070202185104.GH68544@finch-staff-1.thus.net>
References: <875A124D75A8B481E176CF06@p3.JCK.COM>
	<20070202113853.GW7742@finch-staff-1.thus.net>
	<45C33D0C.7BF@xyzzy.claranet.de>
Mime-Version: 1.0
Content-Type: text/plain; charset=us-ascii
Content-Disposition: inline
In-Reply-To: <45C33D0C.7BF@xyzzy.claranet.de>
User-Agent: Mutt/1.5.3i
X-Spam-Score: 0.0 (/)
X-Scan-Signature: 8abaac9e10c826e8252866cbe6766464
Cc: discuss@apps.ietf.org
X-BeenThere: discuss@apps.ietf.org
X-Mailman-Version: 2.1.5
Precedence: list
List-Id: general discussion of application-layer protocols
	<discuss.apps.ietf.org>
List-Unsubscribe: <https://www1.ietf.org/mailman/listinfo/discuss>,
	<mailto:discuss-request@apps.ietf.org?subject=unsubscribe>
List-Post: <mailto:discuss@apps.ietf.org>
List-Help: <mailto:discuss-request@apps.ietf.org?subject=help>
List-Subscribe: <https://www1.ietf.org/mailman/listinfo/discuss>,
	<mailto:discuss-request@apps.ietf.org?subject=subscribe>
Errors-To: discuss-bounces@apps.ietf.org

Frank Ellermann said:
>>     ... U+NN syntax for code point references specified in the Unicode
>>     Standard, where NN is between four and six hexadecimal digits.
> No, folks could misinterpret U+NN as "anything up to 6 digits".

Um, the wording I've used is almost identical to that in the charmod
document (section 1.3).

>> In 5.2, you've said "generally considered ugly and awkward" but I'm
>> not aware of anyone else who's made that complaint.
> +1  Obviously John hates it, that would justify "often".  Others don't
> like backslash-U for various reasons, not only ugly and awkward, also
> confusing (due to various conventions), unclear (lack of delimiter),

Solvable if we include a delimiter.

> and a royal PITA in conjunction with <quoted-string>, when it results
> in multiple backslashes.

Um, every scheme has that problem, surely? See "&amp;#x1234;".

>> In 6 you need to copy in all the security stuff from Unicode
> IMO not "all", folks are supposed to know RFC 3629, it's a STD.

Okay, then the security section needs to explicitly point at the security
section of 3629; it's not enough to say "people should know it".

-- 
Clive D.W. Feather  | Work:  <clive@demon.net>   | Tel:    +44 20 8495 6138
Internet Expert     | Home:  <clive@davros.org>  | Fax:    +44 870 051 9937
Demon Internet      | WWW: http://www.davros.org | Mobile: +44 7973 377646
THUS plc            |                            |




From discuss-bounces@apps.ietf.org Fri Feb 02 13:59:51 2007
Received: from [127.0.0.1] (helo=stiedprmman1.va.neustar.com)
	by megatron.ietf.org with esmtp (Exim 4.43)
	id 1HD3cr-0003lT-NM; Fri, 02 Feb 2007 13:59:09 -0500
Received: from [10.91.34.44] (helo=ietf-mx.ietf.org)
	by megatron.ietf.org with esmtp (Exim 4.43) id 1HD3cq-0003lL-KO
	for discuss@apps.ietf.org; Fri, 02 Feb 2007 13:59:08 -0500
Received: from anchor-internal-1.mail.demon.net ([195.173.56.100])
	by ietf-mx.ietf.org with esmtp (Exim 4.43) id 1HD3cp-0002Yt-8P
	for discuss@apps.ietf.org; Fri, 02 Feb 2007 13:59:08 -0500
Received: from finch-staff-1.server.demon.net (finch-staff-1.server.demon.net [193.195.224.1])
	by anchor-internal-1.mail.demon.net with ESMTPœ id l12Ix64l010236Fri, 2 Feb 2007 18:59:06 GMT
Received: from clive by finch-staff-1.server.demon.net with local (Exim 3.36
	#1) id 1HD3ce-000LPX-00; Fri, 02 Feb 2007 18:58:56 +0000
Date: Fri, 2 Feb 2007 18:58:56 +0000
From: "Clive D.W. Feather" <clive@demon.net>
To: Frank Ellermann <nobody@xyzzy.claranet.de>
Subject: Re: ABNF (was: New draft)
Message-ID: <20070202185856.GI68544@finch-staff-1.thus.net>
References: <875A124D75A8B481E176CF06@p3.JCK.COM>
	<20070202113853.GW7742@finch-staff-1.thus.net>
	<20070202125012.GA18307@nic.fr> <45C345A9.1589@xyzzy.claranet.de>
Mime-Version: 1.0
Content-Type: text/plain; charset=us-ascii
Content-Disposition: inline
In-Reply-To: <45C345A9.1589@xyzzy.claranet.de>
User-Agent: Mutt/1.5.3i
X-Spam-Score: 0.0 (/)
X-Scan-Signature: 9182cfff02fae4f1b6e9349e01d62f32
Cc: discuss@apps.ietf.org
X-BeenThere: discuss@apps.ietf.org
X-Mailman-Version: 2.1.5
Precedence: list
List-Id: general discussion of application-layer protocols
	<discuss.apps.ietf.org>
List-Unsubscribe: <https://www1.ietf.org/mailman/listinfo/discuss>,
	<mailto:discuss-request@apps.ietf.org?subject=unsubscribe>
List-Post: <mailto:discuss@apps.ietf.org>
List-Help: <mailto:discuss-request@apps.ietf.org?subject=help>
List-Subscribe: <https://www1.ietf.org/mailman/listinfo/discuss>,
	<mailto:discuss-request@apps.ietf.org?subject=subscribe>
Errors-To: discuss-bounces@apps.ietf.org

Frank Ellermann said:
> A tricky part could be <name> in chapter 2.2, because...
> 
> |     rulename       =  ALPHA *(ALPHA / DIGIT / "-")
> 
> ...in chapter 4 could be interpreted as different from <name>.

That's stretching it somewhat.

> Now I've
> found a typo in 4234 chapter 2.4:
[...]

I pointed this out to Paul a while ago.

> IMO RFC 4234 is fine, its
> LWSP is an exception, FWS as in RFC 2822 (excl. obs-FWS) would be better.

I'm not sure what point you're trying to make here.

-- 
Clive D.W. Feather  | Work:  <clive@demon.net>   | Tel:    +44 20 8495 6138
Internet Expert     | Home:  <clive@davros.org>  | Fax:    +44 870 051 9937
Demon Internet      | WWW: http://www.davros.org | Mobile: +44 7973 377646
THUS plc            |                            |




From discuss-bounces@apps.ietf.org Fri Feb 02 16:15:53 2007
Received: from [127.0.0.1] (helo=stiedprmman1.va.neustar.com)
	by megatron.ietf.org with esmtp (Exim 4.43)
	id 1HD5kb-0002f6-Ef; Fri, 02 Feb 2007 16:15:17 -0500
Received: from [10.91.34.44] (helo=ietf-mx.ietf.org)
	by megatron.ietf.org with esmtp (Exim 4.43) id 1HD5ka-0002eO-Dz
	for discuss@apps.ietf.org; Fri, 02 Feb 2007 16:15:16 -0500
Received: from main.gmane.org ([80.91.229.2] helo=ciao.gmane.org)
	by ietf-mx.ietf.org with esmtp (Exim 4.43) id 1HD5kY-0007mD-GM
	for discuss@apps.ietf.org; Fri, 02 Feb 2007 16:15:16 -0500
Received: from list by ciao.gmane.org with local (Exim 4.43)
	id 1HD5kW-0007aE-Mx
	for discuss@apps.ietf.org; Fri, 02 Feb 2007 22:15:13 +0100
Received: from 212.82.251.96 ([212.82.251.96])
	by main.gmane.org with esmtp (Gmexim 0.1 (Debian))
	id 1AlnuQ-0007hv-00
	for <discuss@apps.ietf.org>; Fri, 02 Feb 2007 22:15:12 +0100
Received: from nobody by 212.82.251.96 with local (Gmexim 0.1 (Debian))
	id 1AlnuQ-0007hv-00
	for <discuss@apps.ietf.org>; Fri, 02 Feb 2007 22:15:12 +0100
X-Injected-Via-Gmane: http://gmane.org/
To: discuss@apps.ietf.org
From: Frank Ellermann <nobody@xyzzy.claranet.de>
Subject: Re: ABNF
Date: Fri, 02 Feb 2007 22:08:19 +0100
Organization: <URL:http://purl.net/xyzzy>
Lines: 48
Message-ID: <45C3A843.10A7@xyzzy.claranet.de>
References: <875A124D75A8B481E176CF06@p3.JCK.COM>
	<20070202113853.GW7742@finch-staff-1.thus.net>
	<20070202125012.GA18307@nic.fr> <45C345A9.1589@xyzzy.claranet.de>
	<20070202185856.GI68544@finch-staff-1.thus.net>
Mime-Version: 1.0
Content-Type: text/plain; charset=us-ascii
Content-Transfer-Encoding: 7bit
X-Complaints-To: usenet@sea.gmane.org
X-Gmane-NNTP-Posting-Host: 212.82.251.96
X-Mailer: Mozilla 3.0 (OS/2; U)
X-Spam-Score: 1.6 (+)
X-Scan-Signature: 50a516d93fd399dc60588708fd9a3002
X-BeenThere: discuss@apps.ietf.org
X-Mailman-Version: 2.1.5
Precedence: list
List-Id: general discussion of application-layer protocols
	<discuss.apps.ietf.org>
List-Unsubscribe: <https://www1.ietf.org/mailman/listinfo/discuss>,
	<mailto:discuss-request@apps.ietf.org?subject=unsubscribe>
List-Post: <mailto:discuss@apps.ietf.org>
List-Help: <mailto:discuss-request@apps.ietf.org?subject=help>
List-Subscribe: <https://www1.ietf.org/mailman/listinfo/discuss>,
	<mailto:discuss-request@apps.ietf.org?subject=subscribe>
Errors-To: discuss-bounces@apps.ietf.org

Clive D.W. Feather wrote:

> I pointed this out to Paul a while ago.

Thanks, apparently not yet noted as RFC 4234 "erratum".

But something happened, they've finally processed Ned's
simple 2045 erratum - as "unverified", now what's that, 
do they try to verify that the submitter is the author ?

>> IMO RFC 4234 is fine, its LWSP is an exception, FWS
>> as in RFC 2822 (excl. obs-FWS) would be better.

> I'm not sure what point you're trying to make here.

| LWSP =  *(WSP / CRLF WSP)

That allows multiple empty lines consisting only of WSP.

For FWS minus obs-FWS that's impossible (and a loophole
in CFWS is covered by a MUST NOT in the RFC 2822 prose):

| FWS  = ([*WSP CRLF] 1*WSP) /   ; Folding white space

The real damage is in the #rule in RFC 2068 ff. up to 2831bis:

| a rule such as "( *LWS element *( *LWS "," *LWS element )) "
| can be shown as "1#element".
[...]
| implied *LWS
|  The grammar described by this specification is word-based. Except
|  where noted otherwise, linear whitespace (LWS) can be included
|  between any two adjacent words (token or quoted-string), and
|  between adjacent tokens and delimiters (tspecials), without
|  changing the interpretation of a field. At least one delimiter
|  (tspecials) must exist between any two tokens, since they would
|  otherwise be interpreted as a single token.
[...]
| LWS   = [CRLF] 1*( SP | HT )

The same issue as in LWSP, but in RFC 2068 etc. it's hidden in the
#rule and the odd "implied *LWS":  It allows to insert constructs
like <CRLF><SP><CRLF><HT> anywhere.  RFC 2822 tries to eliminate
this possibility in its obs-FWS chapter (4.2), and in the USEFOR
I-Ds obs-FWS is explicitly verboten.  Trailing white space is evil.

Frank






From discuss-bounces@apps.ietf.org Fri Feb 02 16:38:20 2007
Received: from [127.0.0.1] (helo=stiedprmman1.va.neustar.com)
	by megatron.ietf.org with esmtp (Exim 4.43)
	id 1HD66F-0006rb-3A; Fri, 02 Feb 2007 16:37:39 -0500
Received: from [10.91.34.44] (helo=ietf-mx.ietf.org)
	by megatron.ietf.org with esmtp (Exim 4.43) id 1HD66D-0006nu-AB
	for discuss@apps.ietf.org; Fri, 02 Feb 2007 16:37:37 -0500
Received: from ns.jck.com ([209.187.148.211] helo=bs.jck.com)
	by ietf-mx.ietf.org with esmtp (Exim 4.43) id 1HD65g-0002cH-L1
	for discuss@apps.ietf.org; Fri, 02 Feb 2007 16:37:07 -0500
Received: from [127.0.0.1] (helo=p3.JCK.COM)
	by bs.jck.com with esmtp (Exim 4.34)
	id 1HD65e-000FiN-4w; Fri, 02 Feb 2007 16:37:02 -0500
Date: Fri, 02 Feb 2007 16:37:01 -0500
From: John C Klensin <john-ietf@jck.com>
To: "Clive D.W. Feather" <clive@demon.net>,
	Frank Ellermann <nobody@xyzzy.claranet.de>
Subject: draft-klensin-unicode-escapes-01 (was: New Draft)
Message-ID: <B7F8733D73E8CC7227785A69@p3.JCK.COM>
In-Reply-To: <20070202184727.GG68544@finch-staff-1.thus.net>
References: <875A124D75A8B481E176CF06@p3.JCK.COM>
	<uppsr2hs59srbd7eufbcul5a1ekl7i09nl@hive.bjoern.hoehrmann.de>
	<EF59DA6FD89C4F19750C68C3@p3.JCK.COM>
	<20070202114658.GX7742@finch-staff-1.thus.net>
	<45C3371E.330F@xyzzy.claranet.de>
	<20070202184727.GG68544@finch-staff-1.thus.net>
X-Mailer: Mulberry/4.0.7 (Win32)
MIME-Version: 1.0
Content-Type: text/plain; charset=us-ascii
Content-Transfer-Encoding: 7bit
Content-Disposition: inline
X-Spam-Score: 0.0 (/)
X-Scan-Signature: 86f85b2f88b0d50615aed44a7f9e33c7
Cc: discuss@apps.ietf.org
X-BeenThere: discuss@apps.ietf.org
X-Mailman-Version: 2.1.5
Precedence: list
List-Id: general discussion of application-layer protocols
	<discuss.apps.ietf.org>
List-Unsubscribe: <https://www1.ietf.org/mailman/listinfo/discuss>,
	<mailto:discuss-request@apps.ietf.org?subject=unsubscribe>
List-Post: <mailto:discuss@apps.ietf.org>
List-Help: <mailto:discuss-request@apps.ietf.org?subject=help>
List-Subscribe: <https://www1.ietf.org/mailman/listinfo/discuss>,
	<mailto:discuss-request@apps.ietf.org?subject=subscribe>
Errors-To: discuss-bounces@apps.ietf.org

In order to reduce the number of threads consisting of
incremental micro-messages, I'm going to try an omnibus
response....

--On Friday, 02 February, 2007 18:47 +0000 "Clive D.W. Feather"
<clive@demon.net> wrote:

> Frank Ellermann said:
>> The I-D should IMO adopt and cite [Charmod] C042 up to C048
>> verbatim.
> 
> C042 would require &#x1234; rather than allowing us to invent
> \u'1234'.
> 
>> A few other conformance criteria in [Charmod] might be also
>> interesting: http://www.w3.org/TR/charmod/#C070  Don't
>> exclude arbitrary code points
>> http://www.w3.org/TR/charmod/#C077  Don't allow anything
>> above U+10FFFF http://www.w3.org/TR/charmod/#C078  Don't
>> (ab)use surrogates http://www.w3.org/TR/charmod/#C079  Don't
>> (ab)use non-characters
> 
> Those are worth including, I think.

Folks, someday, someone will write a document called "using
Unicode on the Internet".  To some extent, the now-suspended
draft-klensin-net-utf8 is a step in that direction (Mike and I
will get back to it, but, right now, I've got some other things
on my plate).  

But _this_ document is about escaping characters, or really code
points, not about what characters should be permitted or how
they should be used.  Or one of you might take some of this
energy and divert it into an RFC-ish version of "CharMod".   But
for this, just taking the above list as examples:

  C070: irrelevant.   if it is in range, this works.
  C077: needs to be fixed somewhere else.  This is just about a
syntax for escapes
  C078: Already prohibited
  C079: Precisely the reason why one might want escapes is to be
able to deal with non-characters in a sensible way.  Whether
they should be used or not depends on the relevant protocol.



--On Friday, 02 February, 2007 14:05 +0100 Frank Ellermann
<nobody@xyzzy.claranet.de> wrote:

> Yes, never ever mention that HTML exists, it's horrible.  The
> [Charmod] bible requires no (SGML) nonsense in
> http://www.w3.org/TR/charmod/#C044

Sorry, but, IMO, for a document like this which is merely giving
examples (at this point), "very widely deployed and used" trumps
"horrible".
 
>...
> Probably the I-D should mention that one famous exception from
> its rule to avoid encoded UTF-8 is the URL form of IRIs.

Why?  I don't see that (and a few other cases) as "exceptions"
(famous or not), but as mistakes from which we should learn and,
I hope, have learned.   The document is reasonably clear that it
is not a proposal to retrofit any existing protocol.



--On Friday, 02 February, 2007 11:38 +0000 "Clive D.W. Feather"
<clive@demon.net> wrote:

> John C Klensin said:
>> I've just submitted draft-klensin-unicode-escapes-01.txt and
>> assume it will show up in the posting directory today or
>> tomorrow.  
> 
> Some comments for you.

Thanks

> * In 1.1, rather than saying that Unicode occupies "two or
> more octets", wouldn't it be better to say "21 bits - rather
> than the 7 bits of ASCII -"?

No, because of net-ascii and some other issues, probably not.
Again, because the important issue here is that this stuff is
about escapes, you are picking nits that belong elsewhere (see
rant at the end).

> * Somewhere in the last two paragraphs of 1.1 you should be
> talking about mini-languages (e.g. Cosmogol) as well as
> protocols and UIs.

Why?  In principle I could talk about Japanese business cards
and all sorts of other things too.  But that is introductory
material to help the reader understand the context.  Trying to
create an exhaustive list would add nothing to the document
except more pages and making it harder to read.

> * In 3, you're inconsistent between "U+NNN[N[N]]" and
> "NNN...". Indeed, shouldn't the former actually be
> "U+[[N]N]NNNN"? (Note both the order and the number of Ns.) I
> would suggest that better wording might be:

Partially fixed in -02.   U+NNNN[N[N]] versus U+[[N]N]NNNN is a
matter of taste.  Clearly, in that pseudo-notation, one "N" is
as good as another.   See rant at end.

>     ... U+NN syntax for code point references specified in the
> Unicode     Standard, where NN is between four and six
> hexadecimal digits.

I agree with another comment -- too easy to misunderstand.

> * In 4, second bullet, "string terminators" should be "string
> delimiters".

I was deliberately trying to be general.  If one has a form like
H'nnnn', one is clearly talking about "string delimiters".
However, if one has, e.g., &#xNNNN;, it is clear that ";" is a
string terminator, but whether there is a starting delimiter at
all depends on how one defines things in metalanguage or words.
E.g., using BNF (_not_ ABNF), one could reasonably have
   <XML-like-Unicode-escape-string> ::=  <type-introducer>
<value> <terminator>
   <type-introducer> ::= "&#x"
   <value> ::= ....
   <escape-terminator> ::= ";"
or you could construct it in other ways.  Matter of taste, see
rant at the bottom.

> * In 5.2, you've said "generally considered ugly and awkward"
> but I'm not aware of anyone else who's made that complaint.

I could trundle out several others, but it is probably more
efficient to write off to editor's privilege.   See rant at end.

> * In 6 you need to copy in all the security stuff from
> Unicode; the stuff that says that you must use shortest-form
> UTF-8 (so not using %xC1.A1 for 'A') because of the problems
> of filters and firewalls not spotting longer forms.

Absolutely not.  Again, this document is about escapes for
strings of Unicode characters, not about the general use and
appropriateness of Unicode.   And shortest-form UTF-8 is
especially irrelevant because this document is an attempt to
prohibit escapes for UTF-8 entirely where that is still
possible.  The business about different, almost-equivalent,
forms of UTF-8 arguably should go into Section 1 or 2.1 as
further evidence that escaped UTF-8 is a bad idea, but I think
the point has been made.

---------------
<rant>
There is a lot of work to be done in the internationalization
area in the IETF.  Whether one likes what RFC 2277 says or not,
it has become fairly clear that it is not as close to the last
word on the subject as many of us assumed it would be when it
was written.   The suggestions about about CharMod, the
shortest-form string issues that caused us to need to revise the
UTF-8 specs, and the recent work on comparators and comparator
registries, are just the beginning of a very long list.

Speaking personally, I'd be really pleased if I were doing a
much smaller percentage of the document-writing in this area.
But, if I'm going to do it, then my editorial judgment and
preferences are going to prevail... at least until the document
gets far enough along that I have to start arm-wrestling with
the RFC Editor and _their_ judgment and preferences.  I really
appreciate comments about how to make a document more clear but
an argument about, e.g., U+NNNN[N[N]] versus U+[[N]N]NNNN is
only about taste and is hence a waste of everyone's time.

If you don't like my writing style -- and many people don't --
please take on these efforts yourselves and let me complain
about your style (or not) some of the time.  Sniping and
nit-picking is easy and may be fun, but it tends to block,
rather than contribute to, progress.   If you are not going to
pick up some of the writing work, unless your purpose is to
introduce delays and lay down obstacles -- which I assume it is
not-- can we please concentrate on substantive issues and
document changes that would have a clear positive impact on
clarity?
</rant>

thanks,
     john





From discuss-bounces@apps.ietf.org Fri Feb 02 17:27:20 2007
Received: from [127.0.0.1] (helo=stiedprmman1.va.neustar.com)
	by megatron.ietf.org with esmtp (Exim 4.43)
	id 1HD6rE-0000cA-0u; Fri, 02 Feb 2007 17:26:12 -0500
Received: from [10.91.34.44] (helo=ietf-mx.ietf.org)
	by megatron.ietf.org with esmtp (Exim 4.43) id 1HD6rC-0000c1-6o
	for discuss@apps.ietf.org; Fri, 02 Feb 2007 17:26:10 -0500
Received: from main.gmane.org ([80.91.229.2] helo=ciao.gmane.org)
	by ietf-mx.ietf.org with esmtp (Exim 4.43) id 1HD6rA-0002Df-T7
	for discuss@apps.ietf.org; Fri, 02 Feb 2007 17:26:10 -0500
Received: from list by ciao.gmane.org with local (Exim 4.43)
	id 1HD6r2-0002xw-5f
	for discuss@apps.ietf.org; Fri, 02 Feb 2007 23:26:00 +0100
Received: from 212.82.251.96 ([212.82.251.96])
	by main.gmane.org with esmtp (Gmexim 0.1 (Debian))
	id 1AlnuQ-0007hv-00
	for <discuss@apps.ietf.org>; Fri, 02 Feb 2007 23:26:00 +0100
Received: from nobody by 212.82.251.96 with local (Gmexim 0.1 (Debian))
	id 1AlnuQ-0007hv-00
	for <discuss@apps.ietf.org>; Fri, 02 Feb 2007 23:26:00 +0100
X-Injected-Via-Gmane: http://gmane.org/
To: discuss@apps.ietf.org
From: Frank Ellermann <nobody@xyzzy.claranet.de>
Subject: Re: I-D.klensin-unicode-escapes
Date: Fri, 02 Feb 2007 23:23:27 +0100
Organization: <URL:http://purl.net/xyzzy>
Lines: 35
Message-ID: <45C3B9DF.6DA@xyzzy.claranet.de>
References: <875A124D75A8B481E176CF06@p3.JCK.COM>
	<20070202113853.GW7742@finch-staff-1.thus.net>
	<45C33D0C.7BF@xyzzy.claranet.de>
	<20070202185104.GH68544@finch-staff-1.thus.net>
Mime-Version: 1.0
Content-Type: text/plain; charset=us-ascii
Content-Transfer-Encoding: 7bit
X-Complaints-To: usenet@sea.gmane.org
X-Gmane-NNTP-Posting-Host: 212.82.251.96
X-Mailer: Mozilla 3.0 (OS/2; U)
X-Spam-Score: 2.7 (++)
X-Scan-Signature: 52e1467c2184c31006318542db5614d5
X-BeenThere: discuss@apps.ietf.org
X-Mailman-Version: 2.1.5
Precedence: list
List-Id: general discussion of application-layer protocols
	<discuss.apps.ietf.org>
List-Unsubscribe: <https://www1.ietf.org/mailman/listinfo/discuss>,
	<mailto:discuss-request@apps.ietf.org?subject=unsubscribe>
List-Post: <mailto:discuss@apps.ietf.org>
List-Help: <mailto:discuss-request@apps.ietf.org?subject=help>
List-Subscribe: <https://www1.ietf.org/mailman/listinfo/discuss>,
	<mailto:discuss-request@apps.ietf.org?subject=subscribe>
Errors-To: discuss-bounces@apps.ietf.org

Clive D.W. Feather wrote:

 [U+NN]
> Um, the wording I've used is almost identical to that in the charmod
> document (section 1.3).

The wording is okay, but IMO your reversed mnemonic U+[[N]N]NNNN is
better than only U+NN.  With readers you never know, some like me never
look into the prose if the ABNF is apparently clear, while others also
including me look for the examples before ever reading a single word of
the prose or ABNF, and if I understood John correctly his approach is
more like the opposite, he looks into the ABNF if prose and examples
are hopeless... ;-)  

>> a royal PITA in conjunction with <quoted-string>, when it results
>> in multiple backslashes.
 
> Um, every scheme has that problem, surely? See "&amp;#x1234;".

Yes, but it doesn't have to fight with putting <quoted-string>s into
MIME parameter values and similar horrors, compare the RFC 3696 errata.

Escaping backslashes is a pain, the USEFOR WG needed some months^Wtime
to figure this out.  And for 2831bis it strikes again.  Probably it's
a matter of taste, I recall times when I desperately tried \\ or \\\\
or worse with sh or csh scripts.

> the security section needs to explicitly point at the security
> section of 3629; it's not enough to say "people should know it".

Yes.  Only copying the same old UTF-8 security considerations again and
again is boring, a distraction from "real" (specific and fresh) issues.

Frank






From discuss-bounces@apps.ietf.org Fri Feb 02 18:02:55 2007
Received: from [127.0.0.1] (helo=stiedprmman1.va.neustar.com)
	by megatron.ietf.org with esmtp (Exim 4.43)
	id 1HD7Q8-0007YB-VP; Fri, 02 Feb 2007 18:02:16 -0500
Received: from [10.91.34.44] (helo=ietf-mx.ietf.org)
	by megatron.ietf.org with esmtp (Exim 4.43) id 1HD7Q8-0007Y1-Cs
	for discuss@apps.ietf.org; Fri, 02 Feb 2007 18:02:16 -0500
Received: from ns.jck.com ([209.187.148.211] helo=bs.jck.com)
	by ietf-mx.ietf.org with esmtp (Exim 4.43) id 1HD7Q6-000753-Vp
	for discuss@apps.ietf.org; Fri, 02 Feb 2007 18:02:16 -0500
Received: from [127.0.0.1] (helo=p3.JCK.COM)
	by bs.jck.com with esmtp (Exim 4.34)
	id 1HD7Q6-000GHZ-AY; Fri, 02 Feb 2007 18:02:14 -0500
Date: Fri, 02 Feb 2007 18:02:13 -0500
From: John C Klensin <john-ietf@jck.com>
To: Frank Ellermann <nobody@xyzzy.claranet.de>, discuss@apps.ietf.org
Subject: Re: I-D.klensin-unicode-escapes
Message-ID: <6459D49DFD9478F24C914416@p3.JCK.COM>
In-Reply-To: <45C3B9DF.6DA@xyzzy.claranet.de>
References: <875A124D75A8B481E176CF06@p3.JCK.COM>
	<20070202113853.GW7742@finch-staff-1.thus.net>
	<45C33D0C.7BF@xyzzy.claranet.de>
	<20070202185104.GH68544@finch-staff-1.thus.net>
	<45C3B9DF.6DA@xyzzy.claranet.de>
X-Mailer: Mulberry/4.0.7 (Win32)
MIME-Version: 1.0
Content-Type: text/plain; charset=us-ascii
Content-Transfer-Encoding: 7bit
Content-Disposition: inline
X-Spam-Score: 1.1 (+)
X-Scan-Signature: e8a67952aa972b528dd04570d58ad8fe
Cc: 
X-BeenThere: discuss@apps.ietf.org
X-Mailman-Version: 2.1.5
Precedence: list
List-Id: general discussion of application-layer protocols
	<discuss.apps.ietf.org>
List-Unsubscribe: <https://www1.ietf.org/mailman/listinfo/discuss>,
	<mailto:discuss-request@apps.ietf.org?subject=unsubscribe>
List-Post: <mailto:discuss@apps.ietf.org>
List-Help: <mailto:discuss-request@apps.ietf.org?subject=help>
List-Subscribe: <https://www1.ietf.org/mailman/listinfo/discuss>,
	<mailto:discuss-request@apps.ietf.org?subject=subscribe>
Errors-To: discuss-bounces@apps.ietf.org



--On Friday, 02 February, 2007 23:23 +0100 Frank Ellermann
<nobody@xyzzy.claranet.de> wrote:

> Clive D.W. Feather wrote:
> 
>  [U+NN]
>> Um, the wording I've used is almost identical to that in the
>> charmod document (section 1.3).
> 
> The wording is okay, but IMO your reversed mnemonic
> U+[[N]N]NNNN is better than only U+NN.  With readers you never
> know, some like me never look into the prose if the ABNF is
> apparently clear, while others also including me look for the
> examples before ever reading a single word of the prose or
> ABNF, and if I understood John correctly his approach is more
> like the opposite, he looks into the ABNF if prose and examples
> are hopeless... ;-)  

Pretty close.  I'll look at formal and semi-formal definitions
first, but only if they cover semantics in addition to syntax.
If they don't, I tend to focus on semantics and worry about the
syntax details later.  Probably too many years spent worrying
about formal definitions of programming languages.

>>> a royal PITA in conjunction with <quoted-string>, when it
>>> results in multiple backslashes.
>  
>> Um, every scheme has that problem, surely? See "&amp;#x1234;".
> 
> Yes, but it doesn't have to fight with putting
> <quoted-string>s into MIME parameter values and similar
> horrors, compare the RFC 3696 errata.
> 
> Escaping backslashes is a pain, the USEFOR WG needed some
> months^Wtime to figure this out.  And for 2831bis it strikes
> again.  Probably it's a matter of taste, I recall times when I
> desperately tried \\ or \\\\ or worse with sh or csh scripts.

I have to confess that an early pre-posting draft of
unicode-escapes-00 had a section that was supposed to be titled
something like
   Recommendation for \UNNNNNNNN
Processing various combinations of slashes, escapes, and named
characters got two slashes in the output, occasionally three,
and occasionally the escapes themselves, but never one.  I
imagine, given enough time, that I could figure it out, but I'm
really sympathetic your concern above.

>> the security section needs to explicitly point at the security
>> section of 3629; it's not enough to say "people should know
>> it".
> 
> Yes.  Only copying the same old UTF-8 security considerations
> again and again is boring, a distraction from "real" (specific
> and fresh) issues.

See other note about the distinction between a spec about the
use of Unicode characters and one about escapes.

        john





From discuss-bounces@apps.ietf.org Fri Feb 02 21:08:54 2007
Received: from [127.0.0.1] (helo=stiedprmman1.va.neustar.com)
	by megatron.ietf.org with esmtp (Exim 4.43)
	id 1HDAJg-0007tG-79; Fri, 02 Feb 2007 21:07:48 -0500
Received: from [10.91.34.44] (helo=ietf-mx.ietf.org)
	by megatron.ietf.org with esmtp (Exim 4.43) id 1HDAJe-0007tA-Qv
	for discuss@apps.ietf.org; Fri, 02 Feb 2007 21:07:46 -0500
Received: from main.gmane.org ([80.91.229.2] helo=ciao.gmane.org)
	by ietf-mx.ietf.org with esmtp (Exim 4.43) id 1HDAJd-0003xD-Bk
	for discuss@apps.ietf.org; Fri, 02 Feb 2007 21:07:46 -0500
Received: from list by ciao.gmane.org with local (Exim 4.43)
	id 1HDAJV-0002GA-Tw
	for discuss@apps.ietf.org; Sat, 03 Feb 2007 03:07:37 +0100
Received: from 212.82.251.96 ([212.82.251.96])
	by main.gmane.org with esmtp (Gmexim 0.1 (Debian))
	id 1AlnuQ-0007hv-00
	for <discuss@apps.ietf.org>; Sat, 03 Feb 2007 03:07:37 +0100
Received: from nobody by 212.82.251.96 with local (Gmexim 0.1 (Debian))
	id 1AlnuQ-0007hv-00
	for <discuss@apps.ietf.org>; Sat, 03 Feb 2007 03:07:37 +0100
X-Injected-Via-Gmane: http://gmane.org/
To: discuss@apps.ietf.org
From: Frank Ellermann <nobody@xyzzy.claranet.de>
Subject: Re: draft-klensin-unicode-escapes-01
Date: Sat, 03 Feb 2007 03:03:48 +0100
Organization: <URL:http://purl.net/xyzzy>
Lines: 80
Message-ID: <45C3ED84.70C@xyzzy.claranet.de>
References: <875A124D75A8B481E176CF06@p3.JCK.COM>
	<uppsr2hs59srbd7eufbcul5a1ekl7i09nl@hive.bjoern.hoehrmann.de>
	<EF59DA6FD89C4F19750C68C3@p3.JCK.COM>
	<20070202114658.GX7742@finch-staff-1.thus.net>
	<45C3371E.330F@xyzzy.claranet.de>
	<20070202184727.GG68544@finch-staff-1.thus.net>
	<B7F8733D73E8CC7227785A69@p3.JCK.COM>
Mime-Version: 1.0
Content-Type: text/plain; charset=us-ascii
Content-Transfer-Encoding: 7bit
X-Complaints-To: usenet@sea.gmane.org
X-Gmane-NNTP-Posting-Host: 212.82.251.96
X-Mailer: Mozilla 3.0 (OS/2; U)
X-Spam-Score: 2.7 (++)
X-Scan-Signature: d185fa790257f526fedfd5d01ed9c976
X-BeenThere: discuss@apps.ietf.org
X-Mailman-Version: 2.1.5
Precedence: list
List-Id: general discussion of application-layer protocols
	<discuss.apps.ietf.org>
List-Unsubscribe: <https://www1.ietf.org/mailman/listinfo/discuss>,
	<mailto:discuss-request@apps.ietf.org?subject=unsubscribe>
List-Post: <mailto:discuss@apps.ietf.org>
List-Help: <mailto:discuss-request@apps.ietf.org?subject=help>
List-Subscribe: <https://www1.ietf.org/mailman/listinfo/discuss>,
	<mailto:discuss-request@apps.ietf.org?subject=subscribe>
Errors-To: discuss-bounces@apps.ietf.org

John C Klensin wrote:

> Sorry, but, IMO, for a document like this which is merely giving
> examples (at this point), "very widely deployed and used" trumps
> "horrible".

When I mentioned hex. NCRs I meant XML, not SGML and its many ways
to save keystrokes.

>> Probably the I-D should mention that one famous exception from
>> its rule to avoid encoded UTF-8 is the URL form of IRIs.

> Why?  I don't see that (and a few other cases) as "exceptions"
> (famous or not), but as mistakes from which we should learn and,
> I hope, have learned.

The overall RFC 2277 rule is IMO "if you don't know what it is or
can't say what it is, it's UTF-8, and if that theory fails it's
UNKNOWN-8BIT".  And unlike RFC 2231 an URL can't say what it is.

>> wouldn't it be better to say "21 bits - rather than the 7 bits
>> of ASCII -"?

> No, because of net-ascii and some other issues, probably not.
> Again, because the important issue here is that this stuff is
> about escapes, you are picking nits that belong elsewhere

That nit is closely related to escape mechanisms providing for
31 or 32 bits, and attempts to get rid of leading zeros in these
mechanisms, which could fail without explicit delimiters.

> U+NNNN[N[N]] versus U+[[N]N]NNNN is a matter of taste.

If the idea is to reflect appendix A of Unicode 5, it talks about
stripping leading zeros until four digits are left.  In table A.1
it has U+HHHH vs. U-HHHHHHHHH, saying that this is the same as
\uHHHH vs. \UHHHHHHHH.  I haven't seen U-HHHHHHHH before, and I
won't miss it if you don't want to talk about it in the draft.

> my editorial judgment and preferences are going to prevail...
> at least until the document gets far enough along that I have
> to start arm-wrestling with the RFC Editor and _their_ judgment
> and preferences.

Oops, sorry, I thought the intended status was BCP, "an attempt
to prohibit escapes for UTF-8 entirely".

> If you don't like my writing style -- and many people don't --
> please take on these efforts yourselves and let me complain
> about your style (or not) some of the time.

I like it, your drafts are almost always very interesting.  I was
really surprised when I stumbled over RFC 2345 some day ago, it
supports UTF-8 for whois.  So that wasn't a DeNIC invention after
all, and it's clearly older than RFC 3912 chapter 4.

But "interesting" sometimes includes "controversial", and then the
intended status is relevant.  As you say it makes no sense to nit-
pick your personal preferences.

> nit-picking is easy and may be fun, but it tends to block,
> rather than contribute to, progress.

I don't read 99% of all drafts, and I don't try to contribute to
99.9%.  The remaining 0.1% somehow attracted my attention, often
because I like them, the opposite is also possible.  No idea how
easy it generally is, but this article alone took me about three
hours, because I tried to find out what "net ascii" really is, why
your "net utf-8" mentions an RFC that's not more available, where
the Unicode 5 notational conventions are, etc.

> unless your purpose is to introduce delays and lay down obstacles

Escaping that with <rant> and adding a qualifier doesn't make it
better.  If you don't want a discussion about the draft it's okay,
as you say we're free to submit our own drafts about this and / or
related topics.

Frank






From discuss-bounces@apps.ietf.org Sat Feb 03 12:09:02 2007
Received: from [127.0.0.1] (helo=stiedprmman1.va.neustar.com)
	by megatron.ietf.org with esmtp (Exim 4.43)
	id 1HDOLw-0000B2-Jk; Sat, 03 Feb 2007 12:07:04 -0500
Received: from [10.91.34.44] (helo=ietf-mx.ietf.org)
	by megatron.ietf.org with esmtp (Exim 4.43) id 1HDOLv-00008o-8M
	for discuss@apps.ietf.org; Sat, 03 Feb 2007 12:07:03 -0500
Received: from ns.jck.com ([209.187.148.211] helo=bs.jck.com)
	by ietf-mx.ietf.org with esmtp (Exim 4.43) id 1HDOLq-00033b-Nf
	for discuss@apps.ietf.org; Sat, 03 Feb 2007 12:07:03 -0500
Received: from [127.0.0.1] (helo=p3.JCK.COM)
	by bs.jck.com with esmtp (Exim 4.34)
	id 1HDOLp-000O7W-To; Sat, 03 Feb 2007 12:06:58 -0500
Date: Sat, 03 Feb 2007 12:06:56 -0500
From: John C Klensin <john-ietf@jck.com>
To: Frank Ellermann <nobody@xyzzy.claranet.de>, discuss@apps.ietf.org
Subject: Re: draft-klensin-unicode-escapes-01
Message-ID: <AF334D6BB0BFF3037B0DE609@p3.JCK.COM>
X-Mailer: Mulberry/4.0.7 (Win32)
MIME-Version: 1.0
Content-Type: text/plain; charset=us-ascii
Content-Transfer-Encoding: 7bit
Content-Disposition: inline
X-Spam-Score: 1.1 (+)
X-Scan-Signature: 963faf56c3a5b6715f0b71b66181e01a
Cc: 
X-BeenThere: discuss@apps.ietf.org
X-Mailman-Version: 2.1.5
Precedence: list
List-Id: general discussion of application-layer protocols
	<discuss.apps.ietf.org>
List-Unsubscribe: <https://www1.ietf.org/mailman/listinfo/discuss>,
	<mailto:discuss-request@apps.ietf.org?subject=unsubscribe>
List-Post: <mailto:discuss@apps.ietf.org>
List-Help: <mailto:discuss-request@apps.ietf.org?subject=help>
List-Subscribe: <https://www1.ietf.org/mailman/listinfo/discuss>,
	<mailto:discuss-request@apps.ietf.org?subject=subscribe>
Errors-To: discuss-bounces@apps.ietf.org



--On Saturday, 03 February, 2007 03:03 +0100 Frank Ellermann
<nobody@xyzzy.claranet.de> wrote:

> John C Klensin wrote:
> 
>> Sorry, but, IMO, for a document like this which is merely
>> giving examples (at this point), "very widely deployed and
>> used" trumps "horrible".
> 
> When I mentioned hex. NCRs I meant XML, not SGML and its many
> ways to save keystrokes.

And I was commenting only on the suggestion that all reference
to HTML be removed, however horrible it is.  XML is, as far as I
know, already covered.  If it isn't covered adequately, I don't
understand your suggestion.

>>> Probably the I-D should mention that one famous exception
>>> from its rule to avoid encoded UTF-8 is the URL form of IRIs.
> 
>> Why?  I don't see that (and a few other cases) as "exceptions"
>> (famous or not), but as mistakes from which we should learn
>> and, I hope, have learned.
> 
> The overall RFC 2277 rule is IMO "if you don't know what it is
> or can't say what it is, it's UTF-8, and if that theory fails
> it's UNKNOWN-8BIT".  And unlike RFC 2231 an URL can't say what
> it is.

I think I've understood both of those things.  I just haven't
seen the justification or requirement to start exploring
existing protocols in this document, famous or not.  I'm willing
to be convinced but note that every time I put something into a
document that is not strictly necessary I get attacked for
excessive length, etc.

>>> wouldn't it be better to say "21 bits - rather than the 7
>>> bits of ASCII -"?
> 
>> No, because of net-ascii and some other issues, probably not.
>> Again, because the important issue here is that this stuff is
>> about escapes, you are picking nits that belong elsewhere
> 
> That nit is closely related to escape mechanisms providing for
> 31 or 32 bits, and attempts to get rid of leading zeros in
> these mechanisms, which could fail without explicit delimiters.

Sure.  But we routinely express ASCII in terms of octets.  We
don't use the "7-bit" or "21-bit" language very often.  And
wiring it into protocols and conventions just gets us into
trouble with other standards bodies change their minds and
decide that however many bits they though was certainly enough,
wasn't.   Consider the evolution from "6 bits is enough" (with
BCD and other code sets) to "7 bits is enough" (with ASCII and
ISO 646), to "8 bits is enough" (with 8859 (including the
unfortunately-named ASCII-8), EBCDIC, and others), to "16 bits
is certainly enough" (with Unicode 1.0), to "we will never need
more than 21" (current Unicode).   Of course, the IETF and its
predecessors have made the same errors, starting with the
conviction that we would never need more than 255 network nodes,
to 32 bits with IPv4, to 128 bits with IPv6, to some suspicions
that, depending on allocation policies, the latter may turn out
to not be enough in a world with sensor networks, multiple
connectivity paths, and many, many small IP-connected devices.

I don't think I disagree with your point -- it is certainly
factual-- but don't yet see the  need to open this topic up in
this document (see the comment about length, etc., above).  I
have taken \uNNNN and \UNNNNNNNN out of their special status.
I've inserted an explicit discussion about why explicit
delimiters (or terminators) are desirable.  And, with the
working draft for -02, I've made additional comments, as an
extra bullet in Section 4, about the inability to tell the
difference, just by looking, between \uNNNN as a short,
BMP-only, form of the \uNNNN and \UNNNNNNNN pair, a short form
for \uNNNNN or \uNNNNNN, and \uNNNN as an octet encoding for
UTF-16 as reasons why none of those "\u" form (without
delimiters) are attractive.   What else do you suggest and why?
Text please.

>> U+NNNN[N[N]] versus U+[[N]N]NNNN is a matter of taste.
 
> If the idea is to reflect appendix A of Unicode 5, it talks
> about stripping leading zeros until four digits are left.

Which defines the contents and semantics of the NNNN... string,
not the syntax.  I can argue this either way (as I presume you
can).  I just find the second form harder to read.

>  In
> table A.1 it has U+HHHH vs. U-HHHHHHHHH, saying that this is
> the same as \uHHHH vs. \UHHHHHHHH.  I haven't seen U-HHHHHHHH
> before, and I won't miss it if you don't want to talk about it
> in the draft.

Thanks.  I don't intend to talk about it unless someone makes a
very persuasive argument.

>> my editorial judgment and preferences are going to prevail...
>> at least until the document gets far enough along that I have
>> to start arm-wrestling with the RFC Editor and _their_
>> judgment and preferences.
> 
> Oops, sorry, I thought the intended status was BCP, "an attempt
> to prohibit escapes for UTF-8 entirely".

Well, we clearly can't do that without making the spec
retroactive on existing protocols, especially the famous IRI
situation.  But, yes, the target is a standards track document
of some favor... see below.

>> If you don't like my writing style -- and many people don't --
>> please take on these efforts yourselves and let me complain
>> about your style (or not) some of the time.
> 
> I like it, your drafts are almost always very interesting.  I
> was really surprised when I stumbled over RFC 2345 some day
> ago, it supports UTF-8 for whois.  So that wasn't a DeNIC
> invention after all, and it's clearly older than RFC 3912
> chapter 4.

For whatever it is worth, I've become convinced, as I have
delved further into the history of telnet-based protocols, that
it was a mistake.  And I have no reason to believe that DENIC
looked at that document rather than inventing the idea
independently.  If they did draw inspiration from it, I think
I'd be suitably chagrined :-(

> But "interesting" sometimes includes "controversial", and then
> the intended status is relevant.  As you say it makes no sense
> to nit- pick your personal preferences.

In this case, and probably more generally, the comment was
strictly about editorial preferences, not substantive issues of
protocol design or recommendations.    See below.

>> nit-picking is easy and may be fun, but it tends to block,
>> rather than contribute to, progress.
> 
> I don't read 99% of all drafts, and I don't try to contribute
> to 99.9%.  The remaining 0.1% somehow attracted my attention,
> often because I like them, the opposite is also possible.  No
> idea how easy it generally is, but this article alone took me
> about three hours, because I tried to find out what "net
> ascii" really is, why your "net utf-8" mentions an RFC that's
> not more available, where the Unicode 5 notational conventions
> are, etc.

My apologies.  And I'm flattered that you thought it worth the
time.  I'm typically willing to answer questions about comments
of mine that seem obscure if that is more efficient for you.
 
>> unless your purpose is to introduce delays and lay down
>> obstacles
> 
> Escaping that with <rant> and adding a qualifier doesn't make
> it better.  If you don't want a discussion about the draft
> it's okay, as you say we're free to submit our own drafts
> about this and / or related topics.

Frank, I welcome, and want, comments and suggestions.  The fact
that I completely revamped the documents and what it suggested
between -00 and and -01 in response to comments critical of the
recommendation about \u and \U should be evidence of that.  On
this document, virtually every substantive suggestion or comment
has resulted in some changes to the text.  I count "this
sentence is incomprehensible and needs to be fixed" as
substantive in that regard.  As with open source programs and
programming, I believe that having multiple critical eyes on one
of these things makes it better, and that certainly has been the
case here.  

What I've objected to, here and elsewhere, has been what seemed
to be high-frequency nit-picking about editorial style and
preferences.  From my point of view, that is irritating and it
does little or nothing to advance the work.  More important,
especially when I'm working on multiple documents in parallel, I
often start feeling like a Bear of Very Little Brain (apologies
if that reference is not obvious to you): while long words don't
bother me,  I start getting scared that I will miss something
substantive and important among flurries of editorial quibbling.

      john






From discuss-bounces@apps.ietf.org Mon Feb 05 11:01:58 2007
Received: from [127.0.0.1] (helo=stiedprmman1.va.neustar.com)
	by megatron.ietf.org with esmtp (Exim 4.43)
	id 1HE6H7-0008TV-B6; Mon, 05 Feb 2007 11:01:01 -0500
Received: from [10.91.34.44] (helo=ietf-mx.ietf.org)
	by megatron.ietf.org with esmtp (Exim 4.43) id 1HE6H6-0008TM-Io
	for discuss@apps.ietf.org; Mon, 05 Feb 2007 11:01:00 -0500
Received: from ns.jck.com ([209.187.148.211] helo=bs.jck.com)
	by ietf-mx.ietf.org with esmtp (Exim 4.43) id 1HE6H5-0007Os-9A
	for discuss@apps.ietf.org; Mon, 05 Feb 2007 11:01:00 -0500
Received: from [127.0.0.1] (helo=p3.JCK.COM)
	by bs.jck.com with esmtp (Exim 4.34) id 1HE5fG-000F87-8v
	for discuss@apps.ietf.org; Mon, 05 Feb 2007 10:21:54 -0500
Date: Mon, 05 Feb 2007 10:21:53 -0500
From: John C Klensin <john-ietf@jck.com>
To: discuss@apps.ietf.org
Subject: draft-klensin-unicode-escapes-02.txt
Message-ID: <74711BCF624DBEC4F2C000C5@p3.JCK.COM>
X-Mailer: Mulberry/4.0.7 (Win32)
MIME-Version: 1.0
Content-Type: text/plain; charset=us-ascii
Content-Transfer-Encoding: 7bit
Content-Disposition: inline
X-Spam-Score: 0.0 (/)
X-Scan-Signature: ea4ac80f790299f943f0a53be7e1a21a
X-BeenThere: discuss@apps.ietf.org
X-Mailman-Version: 2.1.5
Precedence: list
List-Id: general discussion of application-layer protocols
	<discuss.apps.ietf.org>
List-Unsubscribe: <https://www1.ietf.org/mailman/listinfo/discuss>,
	<mailto:discuss-request@apps.ietf.org?subject=unsubscribe>
List-Post: <mailto:discuss@apps.ietf.org>
List-Help: <mailto:discuss-request@apps.ietf.org?subject=help>
List-Subscribe: <https://www1.ietf.org/mailman/listinfo/discuss>,
	<mailto:discuss-request@apps.ietf.org?subject=subscribe>
Errors-To: discuss-bounces@apps.ietf.org

Hi.

I've just placed another version of the "unicode escapes"
document into the posting queue.  It should be announced to this
list when posted.

I believe I have incorporated all of the changes suggested so
far that haven't gotten pushback, including adjusting the ABNF
that has been controversial (I even changed 4*4 -> 4; others
seem to be more concerned about that syntax preference than I
am).

It occurs to me now --after I sent the document off-- that the
\u / \U form is the only one for which there is ABNF.  This is
the legacy from its previous featured role.  Recommendations
from others as to whether I should just drop that ABNF, leave
things as they are, or add ABNF to the other forms would be
appreciated.  Anyone who prefers the latter should please send
the ABNF they would like to see.

The thing I have _not_ done is to try to expand this document
into making general suggestions or requirements on the use of
Unicode.  It assumes that the strings that one might want to
escape are valid and reasonable and that the definition of
"valid and reasonable" is the province of other documents.

More comments welcome, but I hope we are converging.

     john





From discuss-bounces@apps.ietf.org Mon Feb 05 17:16:23 2007
Received: from [127.0.0.1] (helo=stiedprmman1.va.neustar.com)
	by megatron.ietf.org with esmtp (Exim 4.43)
	id 1HEC7S-0001Al-HQ; Mon, 05 Feb 2007 17:15:26 -0500
Received: from [10.91.34.44] (helo=ietf-mx.ietf.org)
	by megatron.ietf.org with esmtp (Exim 4.43) id 1HEC7Q-0001Ag-W6
	for discuss@apps.ietf.org; Mon, 05 Feb 2007 17:15:24 -0500
Received: from mx2.nic.fr ([192.134.4.11])
	by ietf-mx.ietf.org with esmtp (Exim 4.43) id 1HEC7O-0007jr-Na
	for discuss@apps.ietf.org; Mon, 05 Feb 2007 17:15:24 -0500
Received: from localhost (localhost.localdomain [127.0.0.1])
	by mx2.nic.fr (Postfix) with ESMTP
	id B1E5126C1D9; Mon,  5 Feb 2007 23:15:03 +0100 (CET)
X-Virus-Scanned: by amavisd-new at mx2.nic.fr
Received: from relay2.nic.fr (relay2.nic.fr [192.134.4.163])
	by mx2.nic.fr (Postfix) with ESMTP
	id 3C53826C192; Mon,  5 Feb 2007 23:15:03 +0100 (CET)
Received: from bortzmeyer.nic.fr (batilda.nic.fr [192.134.4.69])
	by relay2.nic.fr (Postfix) with ESMTP id 2F76358E9F0;
	Mon,  5 Feb 2007 23:15:03 +0100 (CET)
Date: Mon, 5 Feb 2007 23:15:03 +0100
From: Stephane Bortzmeyer <bortzmeyer@nic.fr>
To: Frank Ellermann <nobody@xyzzy.claranet.de>
Subject: Re: I-D.klensin-unicode-escapes (was: New Draft)
Message-ID: <20070205221503.GA11623@nic.fr>
References: <875A124D75A8B481E176CF06@p3.JCK.COM>
	<uppsr2hs59srbd7eufbcul5a1ekl7i09nl@hive.bjoern.hoehrmann.de>
	<EF59DA6FD89C4F19750C68C3@p3.JCK.COM>
	<20070202114658.GX7742@finch-staff-1.thus.net>
	<45C3371E.330F@xyzzy.claranet.de>
MIME-Version: 1.0
Content-Type: text/plain; charset=us-ascii
Content-Disposition: inline
In-Reply-To: <45C3371E.330F@xyzzy.claranet.de>
X-Operating-System: Debian GNU/Linux 4.0
X-Kernel: Linux 2.6.17-2-686 i686
Organization: NIC France
X-URL: http://www.nic.fr/
User-Agent: Mutt/1.5.13 (2006-08-11)
X-Spam-Score: 0.0 (/)
X-Scan-Signature: 08e48e05374109708c00c6208b534009
Cc: discuss@apps.ietf.org
X-BeenThere: discuss@apps.ietf.org
X-Mailman-Version: 2.1.5
Precedence: list
List-Id: general discussion of application-layer protocols
	<discuss.apps.ietf.org>
List-Unsubscribe: <https://www1.ietf.org/mailman/listinfo/discuss>,
	<mailto:discuss-request@apps.ietf.org?subject=unsubscribe>
List-Post: <mailto:discuss@apps.ietf.org>
List-Help: <mailto:discuss-request@apps.ietf.org?subject=help>
List-Subscribe: <https://www1.ietf.org/mailman/listinfo/discuss>,
	<mailto:discuss-request@apps.ietf.org?subject=subscribe>
Errors-To: discuss-bounces@apps.ietf.org

On Fri, Feb 02, 2007 at 02:05:34PM +0100,
 Frank Ellermann <nobody@xyzzy.claranet.de> wrote 
 a message of 34 lines which said:

> Probably the I-D should mention that one famous exception from its
> rule to avoid encoded UTF-8 is the URL form of IRIs.

And the protocol described in RFC 2324 :-)




From discuss-bounces@apps.ietf.org Mon Feb 05 19:36:14 2007
Received: from [127.0.0.1] (helo=stiedprmman1.va.neustar.com)
	by megatron.ietf.org with esmtp (Exim 4.43)
	id 1HEEId-0004mc-8G; Mon, 05 Feb 2007 19:35:07 -0500
Received: from [10.91.34.44] (helo=ietf-mx.ietf.org)
	by megatron.ietf.org with esmtp (Exim 4.43) id 1HEEIc-0004mX-FZ
	for discuss@apps.ietf.org; Mon, 05 Feb 2007 19:35:06 -0500
Received: from ns.jck.com ([209.187.148.211] helo=bs.jck.com)
	by ietf-mx.ietf.org with esmtp (Exim 4.43) id 1HEEIa-0006YJ-5Z
	for discuss@apps.ietf.org; Mon, 05 Feb 2007 19:35:06 -0500
Received: from [127.0.0.1] (helo=p3.JCK.COM)
	by bs.jck.com with esmtp (Exim 4.34)
	id 1HEEIV-000IW4-U0; Mon, 05 Feb 2007 19:35:00 -0500
Date: Mon, 05 Feb 2007 19:34:59 -0500
From: John C Klensin <john-ietf@jck.com>
To: Stephane Bortzmeyer <bortzmeyer@nic.fr>,
	Frank Ellermann <nobody@xyzzy.claranet.de>
Subject: Re: I-D.klensin-unicode-escapes (was: New Draft)
Message-ID: <9FD11BDC5B1307B16FE353FB@p3.JCK.COM>
In-Reply-To: <20070205221503.GA11623@nic.fr>
References: <875A124D75A8B481E176CF06@p3.JCK.COM>
	<uppsr2hs59srbd7eufbcul5a1ekl7i09nl@hive.bjoern.hoehrmann.de>
	<EF59DA6FD89C4F19750C68C3@p3.JCK.COM>
	<20070202114658.GX7742@finch-staff-1.thus.net>
	<45C3371E.330F@xyzzy.claranet.de> <20070205221503.GA11623@nic.fr>
X-Mailer: Mulberry/4.0.7 (Win32)
MIME-Version: 1.0
Content-Type: text/plain; charset=us-ascii
Content-Transfer-Encoding: 7bit
Content-Disposition: inline
X-Spam-Score: 0.0 (/)
X-Scan-Signature: 9466e0365fc95844abaf7c3f15a05c7d
Cc: discuss@apps.ietf.org
X-BeenThere: discuss@apps.ietf.org
X-Mailman-Version: 2.1.5
Precedence: list
List-Id: general discussion of application-layer protocols
	<discuss.apps.ietf.org>
List-Unsubscribe: <https://www1.ietf.org/mailman/listinfo/discuss>,
	<mailto:discuss-request@apps.ietf.org?subject=unsubscribe>
List-Post: <mailto:discuss@apps.ietf.org>
List-Help: <mailto:discuss-request@apps.ietf.org?subject=help>
List-Subscribe: <https://www1.ietf.org/mailman/listinfo/discuss>,
	<mailto:discuss-request@apps.ietf.org?subject=subscribe>
Errors-To: discuss-bounces@apps.ietf.org



--On Monday, 05 February, 2007 23:15 +0100 Stephane Bortzmeyer
<bortzmeyer@nic.fr> wrote:

> On Fri, Feb 02, 2007 at 02:05:34PM +0100,
>  Frank Ellermann <nobody@xyzzy.claranet.de> wrote 
>  a message of 34 lines which said:
> 
>> Probably the I-D should mention that one famous exception
>> from its rule to avoid encoded UTF-8 is the URL form of IRIs.
> 
> And the protocol described in RFC 2324 :-)

Now _that_ is clearly the important one, probably sufficiently
so to make an exception to my principle that we should not start
enumerating protocols that might or might not need escapes and
might or might not use encoded UTF-8.

On the other hand, perhaps it is time to start working on
RFC2324bis to permit it to accept other escape characters and to
fully generalize its device control interface formats.

Thanks.  I needed that.

     john







From discuss-bounces@apps.ietf.org Tue Feb 06 14:59:22 2007
Received: from [127.0.0.1] (helo=stiedprmman1.va.neustar.com)
	by megatron.ietf.org with esmtp (Exim 4.43)
	id 1HEWRt-0003kg-Ph; Tue, 06 Feb 2007 14:57:53 -0500
Received: from [10.91.34.44] (helo=ietf-mx.ietf.org)
	by megatron.ietf.org with esmtp (Exim 4.43) id 1HEWRt-0003kb-AS
	for discuss@apps.ietf.org; Tue, 06 Feb 2007 14:57:53 -0500
Received: from main.gmane.org ([80.91.229.2] helo=ciao.gmane.org)
	by ietf-mx.ietf.org with esmtp (Exim 4.43) id 1HEWRq-0001sd-Qu
	for discuss@apps.ietf.org; Tue, 06 Feb 2007 14:57:53 -0500
Received: from list by ciao.gmane.org with local (Exim 4.43)
	id 1HEWRT-0006IL-V3
	for discuss@apps.ietf.org; Tue, 06 Feb 2007 20:57:27 +0100
Received: from d255146.dialin.hansenet.de ([80.171.255.146])
	by main.gmane.org with esmtp (Gmexim 0.1 (Debian))
	id 1AlnuQ-0007hv-00
	for <discuss@apps.ietf.org>; Tue, 06 Feb 2007 20:57:27 +0100
Received: from nobody by d255146.dialin.hansenet.de with local (Gmexim 0.1
	(Debian)) id 1AlnuQ-0007hv-00
	for <discuss@apps.ietf.org>; Tue, 06 Feb 2007 20:57:27 +0100
X-Injected-Via-Gmane: http://gmane.org/
To: discuss@apps.ietf.org
From: Frank Ellermann <nobody@xyzzy.claranet.de>
Subject: Re: draft-klensin-unicode-escapes-01
Date: Tue, 06 Feb 2007 20:53:01 +0100
Organization: <URL:http://purl.net/xyzzy>
Lines: 89
Message-ID: <45C8DC9D.3D61@xyzzy.claranet.de>
References: <AF334D6BB0BFF3037B0DE609@p3.JCK.COM>
Mime-Version: 1.0
Content-Type: text/plain; charset=us-ascii
Content-Transfer-Encoding: 7bit
X-Complaints-To: usenet@sea.gmane.org
X-Gmane-NNTP-Posting-Host: d255146.dialin.hansenet.de
X-Mailer: Mozilla 3.0 (OS/2; U)
X-Spam-Score: 1.1 (+)
X-Scan-Signature: f66b12316365a3fe519e75911daf28a8
X-BeenThere: discuss@apps.ietf.org
X-Mailman-Version: 2.1.5
Precedence: list
List-Id: general discussion of application-layer protocols
	<discuss.apps.ietf.org>
List-Unsubscribe: <https://www1.ietf.org/mailman/listinfo/discuss>,
	<mailto:discuss-request@apps.ietf.org?subject=unsubscribe>
List-Post: <mailto:discuss@apps.ietf.org>
List-Help: <mailto:discuss-request@apps.ietf.org?subject=help>
List-Subscribe: <https://www1.ietf.org/mailman/listinfo/discuss>,
	<mailto:discuss-request@apps.ietf.org?subject=subscribe>
Errors-To: discuss-bounces@apps.ietf.org

John C Klensin wrote:

>> When I mentioned hex. NCRs I meant XML, not SGML and its many
>> ways to save keystrokes.

> And I was commenting only on the suggestion that all reference
> to HTML be removed, however horrible it is.

That's one of those mail communication breakdowns, my remark here
about the wonders of SGML (as seen or not in most browsers trying
to display HTML) was in reply to Clive.  It was no objection to or
suggestion for your I-D.  Of course I support CharMod C044 with
explicit delimiters (as in XML but not SGML) and CharMod C043 with
hex. escapes (something XML and SGML couldn't do, but your I-D and
RFC 4646 got it right.).

 [unlike RFC 2231 an URL can't say what it is]
> I think I've understood both of those things.  I just haven't
> seen the justification or requirement to start exploring existing
> protocols in this document, famous or not.

IIRC you have a SHOULD.  One accepted justification to violate a
SHOULD is "your new rule came too late for my old implementation",
and so far it's unnecessary to talk about it.  But IMO there can
be also reasons to violate this SHOULD in future protocols, if it's
in a context remotely related to IRIs.  Or similar situations where
say using B64-encoded UTF-8 is better than ASCII with hex. NCRs.

If you think that's obvious it's okay.  Sometimes folks ask why a
SHOULD is "only" a SHOULD, and want to know what a _good_ reason to
violate it could be (apart from the clear "too late"), and for that
I thought the IRI example might help.

 From a "protocol lawyer" POV, RFC 2324 is "only" informational <eg>

> every time I put something into a document that is not strictly
> necessary I get attacked for excessive length, etc.

If you think that an example for this SHOULD is unnecessary it's
fine.  With Murphy somebody will attack you later claiming that the
potential exceptions have to be spelled out.

 [21 bits vs. 7 bits]
> Sure.  But we routinely express ASCII in terms of octets.  We
> don't use the "7-bit" or "21-bit" language very often.

Yes, the matter of 21 vs. 31 bits was recently discussed on the
Unicode list in conjunction with a (hypothetical) "UTF-21", maybe
Clive had that discussion in mind.  I'm also fascinated by such
charset encoding details.  One reason that I've not yet published
an "UTF-4" I-D was RFC 4042 with its UTF-9 and UTF-18 "nonets".
Your "net UTF-8" I-D also mentions "nonets" (not using that name).

Of course you don't need to mention that matter in the "escapes"
I-D, unless you want to explain why old conventions demand _eight_
hex. digits where (today) _six_ should be good enough.  I only
tried to state that Clive's remark wasn't off topic or something.

> I don't think I disagree with your point -- it is certainly
> factual-- but don't yet see the  need to open this topic up in
> this document (see the comment about length, etc., above).

It's perfectly okay if you stick to the "octet layer" in this I-D.

FWIW, in theory "UTF-4" (like the "old" UTF-8) could be extented
to 31 bits (again), but of course it will never happen.  It would
break UTF-16, BOCU-1, and all implementations of STD 66.  A lame
excuse is that those aliens are supposed to bring their own kind
of "Intergalacode" when they need more bits.

 [about RFC 2345, unrelated to the Unicode escapes]
> I've become convinced, as I have delved further into the history
> of telnet-based protocols, that it was a mistake.

IBTD, if that's about using UTF-8 in whois.  UTF-8 doesn't need
some of the "critical" (wrt telnet) octets, especially no 0xFF.

RFC 3912 was a huge victory for any anti-1591 cabal, but the whois-
battlefield in the war on spam isn't completely lost yet... <beg>

 [about net-UTF-8, unrelated to the Unicode escapes]
> I'm typically willing to answer questions about comments of mine
> that seem obscure if that is more efficient for you.

The rest can wait for net-utf8-03.  If you have by chance an old
copy of RFC 97, it's AWOL in all RFC collections I've heard of.

Frank






From discuss-bounces@apps.ietf.org Wed Feb 07 12:51:13 2007
Received: from [127.0.0.1] (helo=stiedprmman1.va.neustar.com)
	by megatron.ietf.org with esmtp (Exim 4.43)
	id 1HEqvt-0000lo-4V; Wed, 07 Feb 2007 12:50:13 -0500
Received: from [10.91.34.44] (helo=ietf-mx.ietf.org)
	by megatron.ietf.org with esmtp (Exim 4.43) id 1HEqvs-0000lW-98
	for discuss@apps.ietf.org; Wed, 07 Feb 2007 12:50:12 -0500
Received: from anchor-internal-1.mail.demon.net ([195.173.56.100])
	by ietf-mx.ietf.org with esmtp (Exim 4.43) id 1HEqvm-0006mr-Lp
	for discuss@apps.ietf.org; Wed, 07 Feb 2007 12:50:12 -0500
Received: from finch-staff-1.server.demon.net (finch-staff-1.server.demon.net [193.195.224.1])
	by anchor-internal-1.mail.demon.net with ESMTPœ id l17Ho27J021912Wed, 7 Feb 2007 17:50:03 GMT
Received: from clive by finch-staff-1.server.demon.net with local (Exim 3.36
	#1) id 1HEqvN-000I35-00; Wed, 07 Feb 2007 17:49:41 +0000
Date: Wed, 7 Feb 2007 17:49:41 +0000
From: "Clive D.W. Feather" <clive@demon.net>
To: John C Klensin <john-ietf@jck.com>
Subject: Re: draft-klensin-unicode-escapes-01 (was: New Draft)
Message-ID: <20070207174941.GA64818@finch-staff-1.thus.net>
References: <875A124D75A8B481E176CF06@p3.JCK.COM>
	<uppsr2hs59srbd7eufbcul5a1ekl7i09nl@hive.bjoern.hoehrmann.de>
	<EF59DA6FD89C4F19750C68C3@p3.JCK.COM>
	<20070202114658.GX7742@finch-staff-1.thus.net>
	<45C3371E.330F@xyzzy.claranet.de>
	<20070202184727.GG68544@finch-staff-1.thus.net>
	<B7F8733D73E8CC7227785A69@p3.JCK.COM>
Mime-Version: 1.0
Content-Type: text/plain; charset=us-ascii
Content-Disposition: inline
In-Reply-To: <B7F8733D73E8CC7227785A69@p3.JCK.COM>
User-Agent: Mutt/1.5.3i
X-Spam-Score: 0.0 (/)
X-Scan-Signature: ed68cc91cc637fea89623888898579ba
Cc: Frank Ellermann <nobody@xyzzy.claranet.de>, discuss@apps.ietf.org
X-BeenThere: discuss@apps.ietf.org
X-Mailman-Version: 2.1.5
Precedence: list
List-Id: general discussion of application-layer protocols
	<discuss.apps.ietf.org>
List-Unsubscribe: <https://www1.ietf.org/mailman/listinfo/discuss>,
	<mailto:discuss-request@apps.ietf.org?subject=unsubscribe>
List-Post: <mailto:discuss@apps.ietf.org>
List-Help: <mailto:discuss-request@apps.ietf.org?subject=help>
List-Subscribe: <https://www1.ietf.org/mailman/listinfo/discuss>,
	<mailto:discuss-request@apps.ietf.org?subject=subscribe>
Errors-To: discuss-bounces@apps.ietf.org

John C Klensin said:
>>> A few other conformance criteria in [Charmod] might be also
>>> interesting: http://www.w3.org/TR/charmod/#C070  Don't
>>> exclude arbitrary code points
[...]

> But _this_ document is about escaping characters, or really code
> points, not about what characters should be permitted or how
> they should be used.

Agreed.

>   C070: irrelevant.   if it is in range, this works.

Okay.

>   C077: needs to be fixed somewhere else.  This is just about a
> syntax for escapes

Mumble. Anything above U+10FFFF is, by definition, not Unicode (okay, not
Unicode today). But I'm not going to fight this one any more.

However, it does bring something else to mind: whatever escape mechanism is
chosen, should there be a limit on the length of the hex string? Should an
implementation be expected to correctly handle:

  \u'000000000000000000000000000000000000000000000000000000000000000001234'
or
  \u'000000000000000000000008000000000000000000000000000000000000000001234'
?

What is "correct" in the latter case?

>   C078: Already prohibited

Okay.

>   C079: Precisely the reason why one might want escapes is to be
> able to deal with non-characters in a sensible way.

Very true, and a point I should have realized.

>> Yes, never ever mention that HTML exists, it's horrible.  The
>> [Charmod] bible requires no (SGML) nonsense in
>> http://www.w3.org/TR/charmod/#C044
> Sorry, but, IMO, for a document like this which is merely giving
> examples (at this point), "very widely deployed and used" trumps
> "horrible".

Um, I'm not sure what Frank was thinking, but I would say:
- it's fine to mention HTML;
- we should *not* mention the construct "<B>&aring</B>" (note the missing
  semicolon), which SGML allows and HTML says SHOULD NOT be used.

>>...
>> Probably the I-D should mention that one famous exception from
>> its rule to avoid encoded UTF-8 is the URL form of IRIs.
> Why?  I don't see that (and a few other cases) as "exceptions"
> (famous or not), but as mistakes from which we should learn and,
> I hope, have learned.   The document is reasonably clear that it
> is not a proposal to retrofit any existing protocol.

As someone else asked, what about new protocols based on IRIs?

>> * In 1.1, rather than saying that Unicode occupies "two or
>> more octets", wouldn't it be better to say "21 bits - rather
>> than the 7 bits of ASCII -"?
> No, because of net-ascii and some other issues, probably not.

"net-ascii"? The only reference I can find is in RFC 1350, where it appears
to mean "ASCII with top bit clear, plus some codes from RFC 764 which have
the top bit set".

My point is that Unicode is not octet-based at all - encodings like UTF-8
are, but Unicode isn't. It's a numbering of characters from 0 to 0x10FFFF
just like ASCII is a numnbering from 0 to 127. Neither are octet based.

> Again, because the important issue here is that this stuff is
> about escapes, you are picking nits that belong elsewhere (see
> rant at the end).

Nits have to be picked at some point, and sometimes earlier is better than
later.

>> * Somewhere in the last two paragraphs of 1.1 you should be
>> talking about mini-languages (e.g. Cosmogol) as well as
>> protocols and UIs.
> Why?  In principle I could talk about Japanese business cards
> and all sorts of other things too.

Because they are something that RFCs often use or contain, and for which
this is highly relevant. Unlike Japanese business cards.

> But that is introductory
> material to help the reader understand the context.  Trying to
> create an exhaustive list would add nothing to the document
> except more pages and making it harder to read.

I'm not suggesting exhaustive; I'm suggesting that mini-languages are a
relevant target for this specification.

>> * In 3, you're inconsistent between "U+NNN[N[N]]" and
>> "NNN...". Indeed, shouldn't the former actually be
>> "U+[[N]N]NNNN"? (Note both the order and the number of Ns.) I
>> would suggest that better wording might be:
> Partially fixed in -02.   U+NNNN[N[N]] versus U+[[N]N]NNNN is a
> matter of taste.  Clearly, in that pseudo-notation, one "N" is
> as good as another.

I don't agree, but ....

>> * In 4, second bullet, "string terminators" should be "string
>> delimiters".
> I was deliberately trying to be general.  If one has a form like
> H'nnnn', one is clearly talking about "string delimiters".
> However, if one has, e.g., &#xNNNN;, it is clear that ";" is a
> string terminator,

No, it's a delimiter. A string terminator marks the end of the *string* -
in C, for example, it's the terminating " in the source code or the zero
byte at run-time. A delimiter is something that separates one part of the
string from another.

> but whether there is a starting delimiter at
> all depends on how one defines things in metalanguage or words.

Delimiters don't have to be in pairs.

>> * In 6 you need to copy in all the security stuff from
>> Unicode; the stuff that says that you must use shortest-form
>> UTF-8 (so not using %xC1.A1 for 'A') because of the problems
>> of filters and firewalls not spotting longer forms.
> Absolutely not.  Again, this document is about escapes for
> strings of Unicode characters, not about the general use and
> appropriateness of Unicode.   And shortest-form UTF-8 is
> especially irrelevant because this document is an attempt to
> prohibit escapes for UTF-8 entirely where that is still
> possible.

You've completely missed my point.

The security section needs wording along the lines of:

    An escape mechanism such as the one specified in this document can
    allow characters to be represented in more than one way. Where
    software interprets the escaped form, there is a risk that security
    checks are done at the wrong point.

    For example, a security system might prohibit the substring
    "/../" within certain strings. If so, an attacker could attempt to
    avoid the test by sending "/\u'002E'\u'002E'/" instead. If the
    security check is made before interpretation of escaped characters,
    the attack will be successful.

> Speaking personally, I'd be really pleased if I were doing a
> much smaller percentage of the document-writing in this area.
> But, if I'm going to do it, then my editorial judgment and
> preferences are going to prevail... at least until the document
> gets far enough along that I have to start arm-wrestling with
> the RFC Editor and _their_ judgment and preferences.

Or until Last Call shows that you're in a very small minority.

> I really
> appreciate comments about how to make a document more clear but
> an argument about, e.g., U+NNNN[N[N]] versus U+[[N]N]NNNN is
> only about taste and is hence a waste of everyone's time.

No, it isn't. It's showing the *semantics* of the omission.

> If you don't like my writing style -- and many people don't --
> please take on these efforts yourselves and let me complain
> about your style (or not) some of the time.

Oh no. The last time I accepted that challenge, I ended up authoring a 125
page RFC.

> Sniping and
> nit-picking is easy and may be fun, but it tends to block,
> rather than contribute to, progress.

Excuse me, I'm not sniping (at least, not deliberately; if I'm giving that
impression, I apologise). Yes, I'm nit-picking sometimes, but that's
something that should be done if the resulting document is to be clear,
precise, and useful.

Note, for example, that I'm not nit-picking about what I consider to be
wrong spelling and grammar where I'm aware that it's a dialectic variant.

> If you are not going to
> pick up some of the writing work,

Where I have text to offer - be it 3 words or 3 pages - be assured that I
will. See above, for example. But sometimes it's a lot easier to make the
general point and let the author apply it.

> unless your purpose is to
> introduce delays and lay down obstacles -- which I assume it is
> not--

It is not.

> can we please concentrate on substantive issues and
> document changes that would have a clear positive impact on
> clarity?

If I didn't think my suggestions had a positive impact, I wouldn't make
them. Equally, sometimes I am asking questions or trying to get an idea of
what people think. Perhaps *I*'m in the minority.

> I'm willing
> to be convinced but note that every time I put something into a
> document that is not strictly necessary I get attacked for
> excessive length, etc.                             

Not by me. Examples and explanations are often a good thing. I could have
made RFC 3977 about 40 to 50 pages shorter if I'd taken that approach.

Frank wrote:
> Yes, the matter of 21 vs. 31 bits was recently discussed on the
> Unicode list in conjunction with a (hypothetical) "UTF-21", maybe
> Clive had that discussion in mind.

No. I just object to people introducing octets (or, even worse, "bytes")
when they're unnecessary and - to my mind - confuse the issue.

I note, for example, that the entire POSIX process didn't understand the
difference between the two until I pointed it out, and pointed out some of
the wording changes needed to deal with it. At which point they decided
that POSIX would have to be limited to implementations where they are the
same.

-- 
Clive D.W. Feather  | Work:  <clive@demon.net>   | Tel:    +44 20 8495 6138
Internet Expert     | Home:  <clive@davros.org>  | Fax:    +44 870 051 9937
Demon Internet      | WWW: http://www.davros.org | Mobile: +44 7973 377646
THUS plc            |                            |




From discuss-bounces@apps.ietf.org Wed Feb 07 16:26:41 2007
Received: from [127.0.0.1] (helo=stiedprmman1.va.neustar.com)
	by megatron.ietf.org with esmtp (Exim 4.43)
	id 1HEuIb-0004Lb-7f; Wed, 07 Feb 2007 16:25:53 -0500
Received: from [10.91.34.44] (helo=ietf-mx.ietf.org)
	by megatron.ietf.org with esmtp (Exim 4.43) id 1HEuIa-0004LW-6g
	for discuss@apps.ietf.org; Wed, 07 Feb 2007 16:25:52 -0500
Received: from ns.jck.com ([209.187.148.211] helo=bs.jck.com)
	by ietf-mx.ietf.org with esmtp (Exim 4.43) id 1HEuIY-0005T3-Kv
	for discuss@apps.ietf.org; Wed, 07 Feb 2007 16:25:52 -0500
Received: from [127.0.0.1] (helo=p3.JCK.COM)
	by bs.jck.com with esmtp (Exim 4.34)
	id 1HEuIS-0008rK-Fr; Wed, 07 Feb 2007 16:25:45 -0500
Date: Wed, 07 Feb 2007 16:25:42 -0500
From: John C Klensin <john-ietf@jck.com>
To: "Clive D.W. Feather" <clive@demon.net>
Subject: Re: draft-klensin-unicode-escapes-01 (was: New Draft)
Message-ID: <CABD1699B87DF7916AE364E7@p3.JCK.COM>
In-Reply-To: <20070207174941.GA64818@finch-staff-1.thus.net>
References: <875A124D75A8B481E176CF06@p3.JCK.COM>
	<uppsr2hs59srbd7eufbcul5a1ekl7i09nl@hive.bjoern.hoehrmann.de>
	<EF59DA6FD89C4F19750C68C3@p3.JCK.COM>
	<20070202114658.GX7742@finch-staff-1.thus.net>
	<45C3371E.330F@xyzzy.claranet.de>
	<20070202184727.GG68544@finch-staff-1.thus.net>
	<B7F8733D73E8CC7227785A69@p3.JCK.COM>
	<20070207174941.GA64818@finch-staff-1.thus.net>
X-Mailer: Mulberry/4.0.7 (Win32)
MIME-Version: 1.0
Content-Type: text/plain; charset=us-ascii
Content-Transfer-Encoding: 7bit
Content-Disposition: inline
X-Spam-Score: 0.0 (/)
X-Scan-Signature: d49da3f50144c227c0d2fac65d3953e6
Cc: Frank Ellermann <nobody@xyzzy.claranet.de>, discuss@apps.ietf.org
X-BeenThere: discuss@apps.ietf.org
X-Mailman-Version: 2.1.5
Precedence: list
List-Id: general discussion of application-layer protocols
	<discuss.apps.ietf.org>
List-Unsubscribe: <https://www1.ietf.org/mailman/listinfo/discuss>,
	<mailto:discuss-request@apps.ietf.org?subject=unsubscribe>
List-Post: <mailto:discuss@apps.ietf.org>
List-Help: <mailto:discuss-request@apps.ietf.org?subject=help>
List-Subscribe: <https://www1.ietf.org/mailman/listinfo/discuss>,
	<mailto:discuss-request@apps.ietf.org?subject=subscribe>
Errors-To: discuss-bounces@apps.ietf.org



--On Wednesday, 07 February, 2007 17:49 +0000 "Clive D.W.
Feather" <clive@demon.net> wrote:

> However, it does bring something else to mind: whatever escape
> mechanism is chosen, should there be a limit on the length of
> the hex string? Should an implementation be expected to
> correctly handle:
> 
> \u'00000000000000000000000000000000000000000000000000000000000
> 0000001234' or
> 
> \u'00000000000000000000000800000000000000000000000000000000000
> 0000001234' ?
> 
> What is "correct" in the latter case?

I have assumed that compatibility with existing practice and
good sense suggests that the string should be at least four and
no more than eight hex digits in length.  Obviously, I didn't
write that down.  This is probably a strong argument for
including ABNF for all of the even slightly-recommended forms
(see previous note).   It is not fixed in -02, which has been in
the posting queue since Monday morning, but I'll try to get it
in -03.

>>> Yes, never ever mention that HTML exists, it's horrible.  The
>>> [Charmod] bible requires no (SGML) nonsense in
>>> http://www.w3.org/TR/charmod/#C044
>> Sorry, but, IMO, for a document like this which is merely
>> giving examples (at this point), "very widely deployed and
>> used" trumps "horrible".
> 
> Um, I'm not sure what Frank was thinking, but I would say:
> - it's fine to mention HTML;
> - we should *not* mention the construct "<B>&aring</B>" (note
> the missing   semicolon), which SGML allows and HTML says
> SHOULD NOT be used.

It is not mentioned.   Even &aring; is not mentioned.  That is a
whole different sort of abstraction than what the document talks
about.  Perhaps it should be mentioned; if you think so, text
would be helpful.

>>> ...
>>> Probably the I-D should mention that one famous exception
>>> from its rule to avoid encoded UTF-8 is the URL form of IRIs.
>> Why?  I don't see that (and a few other cases) as "exceptions"
>> (famous or not), but as mistakes from which we should learn
>> and, I hope, have learned.   The document is reasonably clear
>> that it is not a proposal to retrofit any existing protocol.
> 
> As someone else asked, what about new protocols based on IRIs?

Words here won't help, IMO.  Put differently, this is not the
right place to put the words.  If we can get reasonable
consensus on some variation of this and a few other things
wrapped up, I think the right approach is to revisit IRI itself
and see if we want to make changes, narrow scope, insert
warnings, or whatever.  And, in the meantime, I trust that these
conversations have raised the issues with the right group to
prevent new work from going forward without at least their being
considered.

I just don't know what else to suggest.  I have personally never
been happy about aspects of the IRI spec but, up through and
after the time that it was in Last Call, most of my concerns
were just intuitions that there were things in it that would get
us in trouble in the long term -- intuitions I wasn't able to
explain in terms of, e.g., concrete examples.   I also have
concerns that IRIs addressed a problem that didn't really need
solving and ignored the important ones.   Today, I could
probably do somewhat better at explanation and examples and
others reading this could probably do better than I can.   But
I'm still not ready to try to assemble either "IRIbis" or "IRIs
considered harmful".  If someone else is, he or she should go to
it.

>>> * In 1.1, rather than saying that Unicode occupies "two or
>>> more octets", wouldn't it be better to say "21 bits - rather
>>> than the 7 bits of ASCII -"?
>> No, because of net-ascii and some other issues, probably not.
> 
> "net-ascii"? The only reference I can find is in RFC 1350,
> where it appears to mean "ASCII with top bit clear, plus some
> codes from RFC 764 which have the top bit set".

Long discussion.  But "ASCII with top bit clear", etc., implies
7-bit ASCII in an 8-bit unit.  If one were to transmit ASCII in
7-bit units, there is no top bit to be clear or set.  Some
systems, of course, actually did that, even though ASCII on the
network was defined in octet terms.

> My point is that Unicode is not octet-based at all - encodings
> like UTF-8 are, but Unicode isn't. It's a numbering of
> characters from 0 to 0x10FFFF just like ASCII is a numnbering
> from 0 to 127. Neither are octet based.

On the other hand, UCS-4 was certainly octet based, and UTF-32
is now presented as new terminology for UCS-4.   So the
situation isn't as clear, or as pure, as your comment above
suggests.  A bit more on this below.

>> Again, because the important issue here is that this stuff is
>> about escapes, you are picking nits that belong elsewhere (see
>> rant at the end).
> 
> Nits have to be picked at some point, and sometimes earlier is
> better than later.

Ok.   It is just wearing me out on a document I somewhat
accidentally volunteered to assemble but don't feel strongly
committed to.   My problem, not yours, unless I give up and you
do care about it.   Co-authors would be welcome at this point.
 
>>> * Somewhere in the last two paragraphs of 1.1 you should be
>>> talking about mini-languages (e.g. Cosmogol) as well as
>>> protocols and UIs.
>> Why?  In principle I could talk about Japanese business cards
>> and all sorts of other things too.
> 
> Because they are something that RFCs often use or contain, and
> for which this is highly relevant. Unlike Japanese business
> cards.

I still don't see what the stopping rule is.  And, while I may
not be looking in the right places, I don't see Cosmogol (for
example) playing the role in IETF protocol definitions that,
e.g., ABNF or ASN.1 do.
 
>> But that is introductory
>> material to help the reader understand the context.  Trying to
>> create an exhaustive list would add nothing to the document
>> except more pages and making it harder to read.
> 
> I'm not suggesting exhaustive; I'm suggesting that
> mini-languages are a relevant target for this specification.

If you suggest text that explains why they are relevant and what
the impact is, and others agree that it is important enough to
justify the added length, I'll happily drop that text in.

>>> * In 3, you're inconsistent between "U+NNN[N[N]]" and
>>> "NNN...". Indeed, shouldn't the former actually be
>>> "U+[[N]N]NNNN"? (Note both the order and the number of Ns.) I
>>> would suggest that better wording might be:
>> Partially fixed in -02.   U+NNNN[N[N]] versus U+[[N]N]NNNN is
>> a matter of taste.  Clearly, in that pseudo-notation, one "N"
>> is as good as another.
> 
> I don't agree, but ....

Maybe this is connected to our slightly different views of
octets/ bytes (and perhaps "nibbles") above.  Or maybe not.
And, fwiw, if there were a requirement along the lines of "no
leading zeros in strings unless they are needed to make the
string at least four digits long", then I would certainly want
to write
   U+[[1]M]NNNN
with N in the range 0..F and
M in the range 1..F
or something like that.

But, at least so far, we don't have that rule and, so far, I'm
happy to leave deciding on whether or not it is needed to those
specifying particular escapes for particular protocols.

>>> * In 4, second bullet, "string terminators" should be "string
>>> delimiters".
>> I was deliberately trying to be general.  If one has a form
>> like H'nnnn', one is clearly talking about "string
>> delimiters". However, if one has, e.g., &#xNNNN;, it is clear
>> that ";" is a string terminator,
> 
> No, it's a delimiter. A string terminator marks the end of the
> *string* - in C, for example, it's the terminating " in the
> source code or the zero byte at run-time. A delimiter is
> something that separates one part of the string from another.

Sigh.  I think that "terminator" is still right.  If I read your
definition, whether it is correct or not depends on the
definition of strings and substrings. And part of that
definitional question goes back to discussions long ago that
I'll discuss only over appropriate beverages (coming to Prague?).

However, I don't think this is worth shedding blood over so,
unless someone else objects, -03 will use "delimiter" throughout.
 
>>> * In 6 you need to copy in all the security stuff from
>>> Unicode; the stuff that says that you must use shortest-form
>>> UTF-8 (so not using %xC1.A1 for 'A') because of the problems
>>> of filters and firewalls not spotting longer forms.
>> Absolutely not.  Again, this document is about escapes for
>> strings of Unicode characters, not about the general use and
>> appropriateness of Unicode.   And shortest-form UTF-8 is
>> especially irrelevant because this document is an attempt to
>> prohibit escapes for UTF-8 entirely where that is still
>> possible.
> 
> You've completely missed my point.
> 
> The security section needs wording along the lines of:
> 
>     An escape mechanism such as the one specified in this
> document can     allow characters to be represented in more
> than one way. Where     software interprets the escaped form,
> there is a risk that security     checks are done at the wrong
> point.

While I do not believe this is necessary or the right place to
put these sorts of warnings (IMO, they belong in something like
2277bis or a "safely using Unicode" doc), it is at worst
harmless, so I'm invoking the "not willing to shed blood"
principle.  Text included in -03 (with an added comment about
checks for minimal or normalized forms).
 
>     For example, a security system might prohibit the substring
>     "/../" within certain strings. If so, an attacker could
> attempt to     avoid the test by sending "/\u'002E'\u'002E'/"
> instead. If the     security check is made before
> interpretation of escaped characters,     the attack will be
> successful.

See above.

>> Speaking personally, I'd be really pleased if I were doing a
>> much smaller percentage of the document-writing in this area.
>> But, if I'm going to do it, then my editorial judgment and
>> preferences are going to prevail... at least until the
>> document gets far enough along that I have to start
>> arm-wrestling with the RFC Editor and _their_ judgment and
>> preferences.
> 
> Or until Last Call shows that you're in a very small minority.

Which, in the case of _purely_ editorial preferences, involves
some probability of the author saying "I don't need this; find
someone else to finish the document or let it drop".

>> I really
>> appreciate comments about how to make a document more clear
>> but an argument about, e.g., U+NNNN[N[N]] versus U+[[N]N]NNNN
>> is only about taste and is hence a waste of everyone's time.
> 
> No, it isn't. It's showing the *semantics* of the omission.

Not without more words or a more complete definition, IMO.  See
the comments above about U+[1[M]]NNNN.  Do you think that is
worth belaboring?  Do others?

As I think about it, our difference in view may rest in my
viewing U+NNNN[N[N]] pretty much as a name for a concept that I
assumed everyone reading the document was likely to already
understand.  I could, equally comfortably, have just used U+NNNN
or U+NNN... and put in some words, once, about actual lengths.
In fact, that may have been where the original use of U+NNN came
from.   You are viewing it, I guess, as a piece of semi-formal
metanotation with associated semantics, etc.  If I had seen it
that way, I would probably have gone all the way to ABNF on the
theory that U+[[N]N]NNNN does little more than give a strong
hint.

No commitment yet, but, if no one but Clive and I have strong
opinions about this, I am likely to invoke the "no bloodshed"
rule and change this too... but I'd really like to hear from
others about whether to go to
   U+[[N]N]NNNN
or something along the lines of 
   U+[[1]M]NNNN
or to giving this critter a name and writing ABNF, thereby
ridding ourselves of trying to communicate a lot of information
in a 12-character notational form.  If we do go to ABNF, I also
need guidance as to whether to write a leading zero rule into it
(preferably in the form of the ABNF that people would prefer).

>> If you don't like my writing style -- and many people don't --
>> please take on these efforts yourselves and let me complain
>> about your style (or not) some of the time.
> 
> Oh no. The last time I accepted that challenge, I ended up
> authoring a 125 page RFC.

And here I assumed that 2821 would never qualify for a brevity
award in the messaging space.  Time to go add 50 pages to
2821bis :-)

>> Sniping and
>> nit-picking is easy and may be fun, but it tends to block,
>> rather than contribute to, progress.
> 
> Excuse me, I'm not sniping (at least, not deliberately; if I'm
> giving that impression, I apologise). Yes, I'm nit-picking
> sometimes, but that's something that should be done if the
> resulting document is to be clear, precise, and useful.
> 
> Note, for example, that I'm not nit-picking about what I
> consider to be wrong spelling and grammar where I'm aware that
> it's a dialectic variant.

Ok.  For various historical reasons, I often actually prefer
what I assume is your dialect to my native one.  The RFC Editor
generally doesn't.  If you are so inclined, we should have an
offline discussion and then I'll send you the XML and let you
have at it -- part of what is going on here is that this
document is a very marginal activity for me right now and the
time needed to insert trivial and low-priority changes feels
burdensome.  But, if you felt inclined to make a stylistic
editing pass, we should chat to find out whether we can agree on
criteria and limits.

>> If you are not going to
>> pick up some of the writing work,
> 
> Where I have text to offer - be it 3 words or 3 pages - be
> assured that I will. See above, for example. But sometimes
> it's a lot easier to make the general point and let the author
> apply it.
>...

Ok.   As long as we can both understand that there are limited
here, even if we don't precisely agree about where they are,
keep those cards and letters coming.  Just for planning
purposes, I'll probably get -03, containing the changes
discussed above and whatever else comes up (including the
results of corrections or arguments about things I've put into
-02, which was finally announced about five minutes ago), out in
the next week or two. But, if more versions are needed after
that, I will have to put the thing aside until post-Prague.

> Not by me. Examples and explanations are often a good thing. I
> could have made RFC 3977 about 40 to 50 pages shorter if I'd
> taken that approach.

See snide comment about its length relative to that of 2821
above.  We probably mostly agree, but several people in the
community don't... or are willing to insist on their favorite
example and then complain about length vis-a-vis everyone else's.
 
> No. I just object to people introducing octets (or, even
> worse, "bytes") when they're unnecessary and - to my mind -
> confuse the issue.

And most of us who first encountered the Internet or ARPANET, or
computer systems generally, at a time when "characters" could
come in 5, 6, 7, 8, 9, or 12 bit units (and maybe some others),
with or without padding, are sensitive to that issue too.  I
just don't know whether this document is the right place to make
that point, partially because of the UCS-4 problem mentioned
above (and the corresponding original definition of ISO 10646 as
a 32-bit character set) and partially because I have no real
confidence that the 21-bit limit will last (any more than seven
or eight bit limits on characters or eight or 32 bit limits on
address have lasted).  So, in practical terms, I'm strongly
inclined to say that, if we have a variable-length delimited
string, it can be up to eight hex digits long and folks need to
be prepared to parse that much although individual protocols
adopting escapes can impose a leading zero rule.
 
> I note, for example, that the entire POSIX process didn't
> understand the difference between the two until I pointed it
> out, and pointed out some of the wording changes needed to
> deal with it. At which point they decided that POSIX would
> have to be limited to implementations where they are the same.

Sigh.  Unfortunately, not the only thing wrong in the POSIX
process.  We should have a drink sometime and swap stories.

best,
    john








From discuss-bounces@apps.ietf.org Thu Feb 08 00:18:04 2007
Received: from [127.0.0.1] (helo=stiedprmman1.va.neustar.com)
	by megatron.ietf.org with esmtp (Exim 4.43)
	id 1HF1eX-0004XR-Vh; Thu, 08 Feb 2007 00:17:01 -0500
Received: from [10.91.34.44] (helo=ietf-mx.ietf.org)
	by megatron.ietf.org with esmtp (Exim 4.43) id 1HF1eW-0004XJ-Sq
	for discuss@apps.ietf.org; Thu, 08 Feb 2007 00:17:00 -0500
Received: from mxout-03.mxes.net ([216.86.168.178])
	by ietf-mx.ietf.org with esmtp (Exim 4.43) id 1HF1eV-0005ox-Ky
	for discuss@apps.ietf.org; Thu, 08 Feb 2007 00:17:00 -0500
Received: from [127.0.0.1] (unknown [216.145.48.13])
	(using TLSv1 with cipher AES128-SHA (128/128 bits))
	(No client certificate requested)
	by smtp.mxes.net (Postfix) with ESMTP id 8634051934
	for <discuss@apps.ietf.org>; Thu,  8 Feb 2007 00:16:52 -0500 (EST)
Mime-Version: 1.0 (Apple Message framework v752.2)
References: <1D9A5FE2-23DC-431B-B096-B9B38179DBC1@mnot.net>
Content-Type: text/plain; charset=US-ASCII; delsp=yes; format=flowed
Message-Id: <265CD844-23A8-4377-8F06-8DE1B43A153E@mnot.net>
Content-Transfer-Encoding: 7bit
X-Image-Url: http://www.mnot.net/personal/MarkNottingham.jpg
From: Mark Nottingham <mnot@mnot.net>
Subject: Fwd: Informal get-together in Prague
Date: Thu, 8 Feb 2007 16:16:49 +1100
To: Apps Discuss <discuss@apps.ietf.org>
X-Mailer: Apple Mail (2.752.2)
X-Spam-Score: 0.0 (/)
X-Scan-Signature: ea4ac80f790299f943f0a53be7e1a21a
X-BeenThere: discuss@apps.ietf.org
X-Mailman-Version: 2.1.5
Precedence: list
List-Id: general discussion of application-layer protocols
	<discuss.apps.ietf.org>
List-Unsubscribe: <https://www1.ietf.org/mailman/listinfo/discuss>,
	<mailto:discuss-request@apps.ietf.org?subject=unsubscribe>
List-Post: <mailto:discuss@apps.ietf.org>
List-Help: <mailto:discuss-request@apps.ietf.org?subject=help>
List-Subscribe: <https://www1.ietf.org/mailman/listinfo/discuss>,
	<mailto:discuss-request@apps.ietf.org?subject=subscribe>
Errors-To: discuss-bounces@apps.ietf.org

FYI.

Begin forwarded message:

> Resent-From: ietf-http-wg@w3.org
> From: Mark Nottingham <mnot@mnot.net>
> Date: 7 February 2007 4:21:49 PM
> To: "ietf-http-wg@w3.org Group" <ietf-http-wg@w3.org>
> Subject: Informal get-together in Prague
> X-Archived-At: http://www.w3.org/mid/1D9A5FE2-23DC-431B-B096- 
> B9B38179DBC1@mnot.net
>
>
> Julian, Yves, I and at least one other contributor will be meeting  
> on Sunday, March 18 in Prague, to do some editorial work, discuss  
> the issues list, the path forward, and whatever other HTTP-related  
> things come up.
>
> This is intentionally co-located with the IETF meeting, so if  
> anyone wants to come along, please drop me a line so that we have  
> enough room.
>
> More details to follow; also, any suggestions for discussion topics  
> would be welcome.
>
> Cheers,

--
Mark Nottingham     http://www.mnot.net/





From discuss-bounces@apps.ietf.org Fri Feb 09 09:25:56 2007
Received: from [127.0.0.1] (helo=stiedprmman1.va.neustar.com)
	by megatron.ietf.org with esmtp (Exim 4.43)
	id 1HFWgD-00014Z-HE; Fri, 09 Feb 2007 09:24:49 -0500
Received: from [10.91.34.44] (helo=ietf-mx.ietf.org)
	by megatron.ietf.org with esmtp (Exim 4.43) id 1HFWgD-00014U-4F
	for discuss@apps.ietf.org; Fri, 09 Feb 2007 09:24:49 -0500
Received: from anchor-internal-1.mail.demon.net ([195.173.56.100])
	by ietf-mx.ietf.org with esmtp (Exim 4.43) id 1HFWgA-0000n1-8E
	for discuss@apps.ietf.org; Fri, 09 Feb 2007 09:24:49 -0500
Received: from finch-staff-1.server.demon.net (finch-staff-1.server.demon.net [193.195.224.1])
	by anchor-internal-1.mail.demon.net with ESMTPœ id l19EOhAc009068Fri, 9 Feb 2007 14:24:44 GMT
Received: from clive by finch-staff-1.server.demon.net with local (Exim 3.36
	#1) id 1HFWfh-000AaJ-00; Fri, 09 Feb 2007 14:24:17 +0000
Date: Fri, 9 Feb 2007 14:24:17 +0000
From: "Clive D.W. Feather" <clive@demon.net>
To: John C Klensin <john-ietf@jck.com>
Subject: Re: draft-klensin-unicode-escapes-01 (was: New Draft)
Message-ID: <20070209142417.GK18441@finch-staff-1.thus.net>
References: <875A124D75A8B481E176CF06@p3.JCK.COM>
	<uppsr2hs59srbd7eufbcul5a1ekl7i09nl@hive.bjoern.hoehrmann.de>
	<EF59DA6FD89C4F19750C68C3@p3.JCK.COM>
	<20070202114658.GX7742@finch-staff-1.thus.net>
	<45C3371E.330F@xyzzy.claranet.de>
	<20070202184727.GG68544@finch-staff-1.thus.net>
	<B7F8733D73E8CC7227785A69@p3.JCK.COM>
	<20070207174941.GA64818@finch-staff-1.thus.net>
	<CABD1699B87DF7916AE364E7@p3.JCK.COM>
Mime-Version: 1.0
Content-Type: text/plain; charset=us-ascii
Content-Disposition: inline
In-Reply-To: <CABD1699B87DF7916AE364E7@p3.JCK.COM>
User-Agent: Mutt/1.5.3i
X-Spam-Score: 0.0 (/)
X-Scan-Signature: e367d58950869b6582535ddf5a673488
Cc: Frank Ellermann <nobody@xyzzy.claranet.de>, discuss@apps.ietf.org
X-BeenThere: discuss@apps.ietf.org
X-Mailman-Version: 2.1.5
Precedence: list
List-Id: general discussion of application-layer protocols
	<discuss.apps.ietf.org>
List-Unsubscribe: <https://www1.ietf.org/mailman/listinfo/discuss>,
	<mailto:discuss-request@apps.ietf.org?subject=unsubscribe>
List-Post: <mailto:discuss@apps.ietf.org>
List-Help: <mailto:discuss-request@apps.ietf.org?subject=help>
List-Subscribe: <https://www1.ietf.org/mailman/listinfo/discuss>,
	<mailto:discuss-request@apps.ietf.org?subject=subscribe>
Errors-To: discuss-bounces@apps.ietf.org

John C Klensin said:
>> However, it does bring something else to mind: whatever escape
>> mechanism is chosen, should there be a limit on the length of
>> the hex string?

> I have assumed that compatibility with existing practice and
> good sense suggests that the string should be at least four and
> no more than eight hex digits in length.  Obviously, I didn't
> write that down.  This is probably a strong argument for
> including ABNF for all of the even slightly-recommended forms
> (see previous note).   It is not fixed in -02, which has been in
> the posting queue since Monday morning, but I'll try to get it
> in -03.

Okay.

>> Um, I'm not sure what Frank was thinking, but I would say:
>> - it's fine to mention HTML;
>> - we should *not* mention the construct "<B>&aring</B>" (note
>> the missing   semicolon), which SGML allows and HTML says
>> SHOULD NOT be used.
> It is not mentioned.   Even &aring; is not mentioned.  That is a
> whole different sort of abstraction than what the document talks
> about.

Sorry, I shouldn't have mentioned &aring; in my example. I see that -02
talks about the missing semicolon, so I'm happy that this be closed.


>> My point is that Unicode is not octet-based at all - encodings
>> like UTF-8 are, but Unicode isn't. It's a numbering of
>> characters from 0 to 0x10FFFF just like ASCII is a numnbering
>> from 0 to 127. Neither are octet based.
> On the other hand, UCS-4 was certainly octet based, and UTF-32
> is now presented as new terminology for UCS-4.   So the
> situation isn't as clear, or as pure, as your comment above
> suggests.

Not quite, no.

>> Nits have to be picked at some point, and sometimes earlier is
>> better than later.
> Ok.   It is just wearing me out on a document I somewhat
> accidentally volunteered to assemble but don't feel strongly
> committed to.   My problem, not yours, unless I give up and you
> do care about it.

Sorry, I didn't mean to stick a huge load on you by my comments.

> Co-authors would be welcome at this point.

Not at the moment, I'm afraid. If things change at work, then perhaps.

>>>> * Somewhere in the last two paragraphs of 1.1 you should be
>>>> talking about mini-languages (e.g. Cosmogol) as well as
>>>> protocols and UIs.

> I still don't see what the stopping rule is.  And, while I may
> not be looking in the right places, I don't see Cosmogol (for
> example) playing the role in IETF protocol definitions that,
> e.g., ABNF or ASN.1 do.

Well, ABNF is another mini-language that might benefit from a way to handle
Unicode.

> If you suggest text that explains why they are relevant and what
> the impact is, and others agree that it is important enough to
> justify the added length, I'll happily drop that text in.

In -02, change the last paragraph of 1.1 to:

    In addition to the protocol contexts addressed in this specification,
    escapes to represent Unicode characters could also be useful in formal
    languages (such as ABNF and Cosmogol) and in presentations to users
    (i.e. user interfaces). The formats specified in, and the reasoning of,
    this document may be applicable to these contexts as well, but this is
    not a proposal to standardise them.

That would satisfy me.

> And, fwiw, if there were a requirement along the lines of "no
> leading zeros in strings unless they are needed to make the
> string at least four digits long", then I would certainly want
> to write
>    U+[[1]M]NNNN
> with N in the range 0..F and
> M in the range 1..F
> or something like that.

Something like that, yes, since as written you've ruled out U+10xxxx.

That's also, perhaps, going too far in putting the detail in the text.

>>>> * In 4, second bullet, "string terminators" should be "string
>>>> delimiters".
[...]
> Sigh.  I think that "terminator" is still right.  If I read your
> definition, whether it is correct or not depends on the
> definition of strings and substrings. And part of that
> definitional question goes back to discussions long ago that
> I'll discuss only over appropriate beverages (coming to Prague?).

Regrettably not - I can't justify the time. But if you're ever in London or
Cambridge ....

>> The security section needs wording along the lines of:
>> 
>>     An escape mechanism such as the one specified in this
>> document can     allow characters to be represented in more
>> than one way. Where     software interprets the escaped form,
>> there is a risk that security     checks are done at the wrong
>> point.
> 
> While I do not believe this is necessary or the right place to
> put these sorts of warnings (IMO, they belong in something like
> 2277bis or a "safely using Unicode" doc), it is at worst
> harmless, so I'm invoking the "not willing to shed blood"
> principle.

Thanks.

I can agree with "not necessary"; I just think it's a good idea (see other
discussions on minimalism).

>>> But, if I'm going to do it, then my editorial judgment and
>>> preferences are going to prevail...
>> Or until Last Call shows that you're in a very small minority.
> Which, in the case of _purely_ editorial preferences, involves
> some probability of the author saying "I don't need this; find
> someone else to finish the document or let it drop".

True.

> As I think about it, our difference in view may rest in my
> viewing U+NNNN[N[N]] pretty much as a name for a concept that I
> assumed everyone reading the document was likely to already
> understand.  I could, equally comfortably, have just used U+NNNN
> or U+NNN... and put in some words, once, about actual lengths.
> In fact, that may have been where the original use of U+NNN came
> from.   You are viewing it, I guess, as a piece of semi-formal
> metanotation with associated semantics, etc.

Or, more precisely, I'm concerned that others may read it that way. As you
say, you and I know what it's supposed to mean.

[My day job involves doing this sort of analysis on draft legislation.
You'd be amazed just how badly people can misinterpret even the slightest
ambiguity.]

> No commitment yet, but, if no one but Clive and I have strong
> opinions about this, I am likely to invoke the "no bloodshed"
> rule and change this too... but I'd really like to hear from
> others about whether to go to
>    U+[[N]N]NNNN
> or something along the lines of 
>    U+[[1]M]NNNN
> or to giving this critter a name and writing ABNF, thereby
> ridding ourselves of trying to communicate a lot of information
> in a 12-character notational form.

I think that the ABNF approach is probably a good one.

> If we do go to ABNF, I also
> need guidance as to whether to write a leading zero rule into it
> (preferably in the form of the ABNF that people would prefer).

    unicode-notation = "U+" code-point
    code-point = bmp-code-point / extended-code-point
    bmp-code-point = 4hex-digit
    extended-code-point = (non-zero-hex-digit / "10") 4hex-digit
    hex-digit = "0" / non-zero-hex-digit
    non-zero-hex-digit = "1" / "2" / "3" / "4" / "5" / "6" / "7" / "8" /
        "9" / "A" / "B" / "C" / "D" / "E" / "F"

>> Note, for example, that I'm not nit-picking about what I
>> consider to be wrong spelling and grammar where I'm aware that
>> it's a dialectic variant.
> Ok.  For various historical reasons, I often actually prefer
> what I assume is your dialect to my native one.  The RFC Editor
> generally doesn't.  If you are so inclined, we should have an
> offline discussion and then I'll send you the XML and let you
> have at it

No thanks. I actually view this as being very much an author's privilege.
Throughout 3977 I used *my* choices as to correct spelling and grammar, and
fought the RFC Editor when necessary. You're the author, so these are your
choice IMAO.

>> No. I just object to people introducing octets (or, even
>> worse, "bytes") when they're unnecessary and - to my mind -
>> confuse the issue.
> And most of us who first encountered the Internet or ARPANET, or
> computer systems generally, at a time when "characters" could
> come in 5, 6, 7, 8, 9, or 12 bit units (and maybe some others),

[5 and a half in my case: I learned my programming on a system with 39 bit
words, and a character encoding with 5 bits with two shift states - there
were two separate packing systems, one of 7x5 with explicit shifts, the
other of 6x6 with shift bits on each character.]

> partially because I have no real
> confidence that the 21-bit limit will last

Point.

> So, in practical terms, I'm strongly
> inclined to say that, if we have a variable-length delimited
> string, it can be up to eight hex digits long and folks need to
> be prepared to parse that much although individual protocols
> adopting escapes can impose a leading zero rule.

Okay.

> Sigh.  Unfortunately, not the only thing wrong in the POSIX
> process.  We should have a drink sometime and swap stories.

See above.

-- 
Clive D.W. Feather  | Work:  <clive@demon.net>   | Tel:    +44 20 8495 6138
Internet Expert     | Home:  <clive@davros.org>  | Fax:    +44 870 051 9937
Demon Internet      | WWW: http://www.davros.org | Mobile: +44 7973 377646
THUS plc            |                            |




From discuss-bounces@apps.ietf.org Fri Feb 09 09:39:48 2007
Received: from [127.0.0.1] (helo=stiedprmman1.va.neustar.com)
	by megatron.ietf.org with esmtp (Exim 4.43)
	id 1HFWu7-0008Gf-NG; Fri, 09 Feb 2007 09:39:11 -0500
Received: from [10.91.34.44] (helo=ietf-mx.ietf.org)
	by megatron.ietf.org with esmtp (Exim 4.43) id 1HFWu6-0008GX-VW
	for discuss@apps.ietf.org; Fri, 09 Feb 2007 09:39:10 -0500
Received: from anchor-internal-1.mail.demon.net ([195.173.56.100])
	by ietf-mx.ietf.org with esmtp (Exim 4.43) id 1HFWu5-0002R6-GY
	for discuss@apps.ietf.org; Fri, 09 Feb 2007 09:39:10 -0500
Received: from finch-staff-1.server.demon.net (finch-staff-1.server.demon.net [193.195.224.1])
	by anchor-internal-1.mail.demon.net with ESMTPœ id l19Ed8BV018349Fri, 9 Feb 2007 14:39:08 GMT
Received: from clive by finch-staff-1.server.demon.net with local (Exim 3.36
	#1) id 1HFWte-000Azz-00; Fri, 09 Feb 2007 14:38:42 +0000
Date: Fri, 9 Feb 2007 14:38:42 +0000
From: "Clive D.W. Feather" <clive@demon.net>
To: John C Klensin <john-ietf@jck.com>
Subject: Re: draft-klensin-unicode-escapes-02.txt
Message-ID: <20070209143842.GL18441@finch-staff-1.thus.net>
References: <74711BCF624DBEC4F2C000C5@p3.JCK.COM>
Mime-Version: 1.0
Content-Type: text/plain; charset=us-ascii
Content-Disposition: inline
In-Reply-To: <74711BCF624DBEC4F2C000C5@p3.JCK.COM>
User-Agent: Mutt/1.5.3i
X-Spam-Score: 0.0 (/)
X-Scan-Signature: 5a9a1bd6c2d06a21d748b7d0070ddcb8
Cc: discuss@apps.ietf.org
X-BeenThere: discuss@apps.ietf.org
X-Mailman-Version: 2.1.5
Precedence: list
List-Id: general discussion of application-layer protocols
	<discuss.apps.ietf.org>
List-Unsubscribe: <https://www1.ietf.org/mailman/listinfo/discuss>,
	<mailto:discuss-request@apps.ietf.org?subject=unsubscribe>
List-Post: <mailto:discuss@apps.ietf.org>
List-Help: <mailto:discuss-request@apps.ietf.org?subject=help>
List-Subscribe: <https://www1.ietf.org/mailman/listinfo/discuss>,
	<mailto:discuss-request@apps.ietf.org?subject=subscribe>
Errors-To: discuss-bounces@apps.ietf.org

John C Klensin said:
> I've just placed another version of the "unicode escapes"
> document into the posting queue.  It should be announced to this
> list when posted.

One typo: in section 3, you have "NNMN".

The abstract says "makes a proposal for general use". But you don't any
more, do you? You simply propose rules for consideration.

> The thing I have _not_ done is to try to expand this document
> into making general suggestions or requirements on the use of
> Unicode.  It assumes that the strings that one might want to
> escape are valid and reasonable and that the definition of
> "valid and reasonable" is the province of other documents.

Okay. That paragraph might be worth adding to the text.

> More comments welcome, but I hope we are converging.

There's something I found in my email from last year when this topic was
first being discussed.

Consider an explicitly delimited form like \u(NNNN). Is it worth saying
that "customers" MAY allow the use of multiple code-points within the
delimiters. In other words:

    \u(1234,5678,109ABC)

is an acceptable abbreviation for:

    \u(1234)\u(5678)\u(109ABC)

The ABNF for this would be:

    unicode-escape = escape-prefix "(" escape-body ")"
    escape-prefix = "\u"   ; I'm ignore case-sensitivity issues
    escape-body = required-body / permitted-body
    required-body = code-point
    permitted-body = code-point 1*( "," code-point )

See previous email for code-point, though perhaps an implementation MAY
allow leading zeros to be omitted.

[The other point from that old email, though if you say it's too esoteric
to mention I won't complain, is that this also allows other notations to be
mixed in if an implementer wants. For example, explicit character names, so
that:
    \u('e acute')
would be another way to write
    \u(00E9)

-- 
Clive D.W. Feather  | Work:  <clive@demon.net>   | Tel:    +44 20 8495 6138
Internet Expert     | Home:  <clive@davros.org>  | Fax:    +44 870 051 9937
Demon Internet      | WWW: http://www.davros.org | Mobile: +44 7973 377646
THUS plc            |                            |




From discuss-bounces@apps.ietf.org Mon Feb 19 14:57:25 2007
Received: from [127.0.0.1] (helo=stiedprmman1.va.neustar.com)
	by megatron.ietf.org with esmtp (Exim 4.43)
	id 1HJEcR-0003zx-B0; Mon, 19 Feb 2007 14:56:15 -0500
Received: from [10.91.34.44] (helo=ietf-mx.ietf.org)
	by megatron.ietf.org with esmtp (Exim 4.43) id 1HJEWL-0002ZZ-2c
	for discuss@apps.ietf.org; Mon, 19 Feb 2007 14:49:57 -0500
Received: from gateout01.mbox.net ([165.212.64.21])
	by ietf-mx.ietf.org with esmtp (Exim 4.43) id 1HJEWH-0003yZ-3n
	for discuss@apps.ietf.org; Mon, 19 Feb 2007 14:49:57 -0500
Received: from gateout01.mbox.net (gateout01.mbox.net [165.212.64.21])
	by gateout01.mbox.net (Postfix) with ESMTP id 7E79019E8;
	Mon, 19 Feb 2007 19:49:39 +0000 (GMT)
Received: from GW2.EXCHPROD.USA.NET [165.212.116.254] by gateout01.mbox.net
	via smtad (C8.MAIN.3.34P) 
	with ESMTP id XID767LBsTxn2626Xo1; Mon, 19 Feb 2007 19:49:39 -0000
X-USANET-Source: 165.212.116.254 IN ldusseault@commerce.net
	GW2.EXCHPROD.USA.NET
X-USANET-MsgId: XID767LBsTxn2626Xo1
Received: from [192.168.1.100] ([69.181.78.47]) by GW2.EXCHPROD.USA.NET over
	TLS secured channel with Microsoft SMTPSVC(6.0.3790.1830); 
	Mon, 19 Feb 2007 12:49:23 -0700
Mime-Version: 1.0 (Apple Message framework v752.2)
To: Apps Discuss <discuss@apps.ietf.org>,
	HTTP Working Group <ietf-http-wg@w3.org>
Message-Id: <EC0B32D0-39FC-4985-8A51-7D57C68C344E@commerce.net>
Content-Type: multipart/alternative; boundary=Apple-Mail-26-333833273
References: <45D94EA0.2050706@alvestrand.no>
From: Lisa Dusseault <ldusseault@commerce.net>
Subject: Fwd: [Standards-discuss] [HTML in email]
Date: Mon, 19 Feb 2007 11:49:21 -0800
X-Mailer: Apple Mail (2.752.2)
X-OriginalArrivalTime: 19 Feb 2007 19:49:23.0761 (UTC)
	FILETIME=[0E0B8210:01C7545F]
X-Spam-Score: 0.0 (/)
X-Scan-Signature: 7f3fa64b9851a63d7f3174ef64114da7
X-Mailman-Approved-At: Mon, 19 Feb 2007 14:56:13 -0500
Cc: 
X-BeenThere: discuss@apps.ietf.org
X-Mailman-Version: 2.1.5
Precedence: list
List-Id: general discussion of application-layer protocols
	<discuss.apps.ietf.org>
List-Unsubscribe: <https://www1.ietf.org/mailman/listinfo/discuss>,
	<mailto:discuss-request@apps.ietf.org?subject=unsubscribe>
List-Post: <mailto:discuss@apps.ietf.org>
List-Help: <mailto:discuss-request@apps.ietf.org?subject=help>
List-Subscribe: <https://www1.ietf.org/mailman/listinfo/discuss>,
	<mailto:discuss-request@apps.ietf.org?subject=subscribe>
Errors-To: discuss-bounces@apps.ietf.org


--Apple-Mail-26-333833273
Content-Transfer-Encoding: 7bit
Content-Type: text/plain;
	charset=US-ASCII;
	delsp=yes;
	format=flowed

More forwarding/ FYI

Lisa

Begin forwarded message:

> From: Harald Alvestrand <harald@alvestrand.no>
> Date: February 18, 2007 11:15:44 PM PST
> To: mhtml@segate.sunet.se, ietf-822@imc.org
> Subject: [Fwd: [Standards-discuss] [HTML in email]]
>
> (sending again, with a From address that has a chance of getting  
> through....)
>
> I guess that if the lists above still exist, there may be people  
> with an
> interest in this proposal for a new W3C activity on them....
>
>                 Harald
>
>
> From: Daniel Glazman <daniel@glazman.org>
> Date: February 15, 2007 6:58:30 AM PST
> To: daniel@glazman.org
> Subject: HTML in email
>
>
> People,
>
> Following a few personal discussions with friends here in France,
> and a thread in a W3C Members-only mailing-list, the W3C
> has decided to launch a new mailing-list for issues related to HTML
> in email: editing, rendering, interoperability, security, ...
> The idea is - for instance - to take advantage of the work done by the
> future HTML WG to have an email-safe profile of HTML. That's only one
> of the possibilities, just to give you an example.
> BTW, that includes HTML emails sent from and received by mobile  
> devices,
> of course...
>
> The mailing-list is public-html-mail@w3.org
>
>   http://lists.w3.org/Archives/Public/public-html-mail/
>
> We also think of having an official W3C workshop on this topic in the
> coming weeks or months, and we'll ask for position papers.
>
> Could you please kindly forward this invitation to join the mailing-
> list to the key people working on HTML email editing and rendering
> in your organization ? Please also forward to anyone working in this
> scope, even outside of your own org, who can provide useful  
> thoughts and
> help on this subject.
>
> I'll send another message to companies involved in direct email-
> based marketing, because their input and help is needed here.
>
> Best regards,
>
> Daniel Glazman
>
> -- 
> Best Regards,
> --raman
>
> Title:  Research Scientist      Email:  raman@google.com
> WWW:    http://emacspeak.sf.net/raman/
> Google: tv+raman GTalk:  raman@google.com, tv.raman.tv@gmail.com
> PGP:    http://emacspeak.sf.net/raman/raman-almaden.asc
>
>
>
> _______________________________________________
> Standards-discuss mailing list
> Standards-discuss@google.com
> https://mailman.corp.google.com/mailman/listinfo/standards-discuss
>
>


--Apple-Mail-26-333833273
Content-Transfer-Encoding: quoted-printable
Content-Type: text/html;
	charset=ISO-8859-1

<HTML><BODY style=3D"word-wrap: break-word; -khtml-nbsp-mode: space; =
-khtml-line-break: after-white-space; ">More forwarding/ FYI<DIV><BR =
class=3D"khtml-block-placeholder"></DIV><DIV>Lisa<BR><DIV><BR><DIV>Begin =
forwarded message:</DIV><BR =
class=3D"Apple-interchange-newline"><BLOCKQUOTE type=3D"cite"><DIV =
style=3D"margin-top: 0px; margin-right: 0px; margin-bottom: 0px; =
margin-left: 0px; "><FONT face=3D"Helvetica" size=3D"3" color=3D"#000000" =
style=3D"font: 12.0px Helvetica; color: #000000"><B>From: =
</B></FONT><FONT face=3D"Helvetica" size=3D"3" style=3D"font: 12.0px =
Helvetica">Harald Alvestrand &lt;<A =
href=3D"mailto:harald@alvestrand.no">harald@alvestrand.no</A>&gt;</FONT></=
DIV><DIV style=3D"margin-top: 0px; margin-right: 0px; margin-bottom: =
0px; margin-left: 0px; "><FONT face=3D"Helvetica" size=3D"3" =
color=3D"#000000" style=3D"font: 12.0px Helvetica; color: =
#000000"><B>Date: </B></FONT><FONT face=3D"Helvetica" size=3D"3" =
style=3D"font: 12.0px Helvetica">February 18, 2007 11:15:44 PM =
PST</FONT></DIV><DIV style=3D"margin-top: 0px; margin-right: 0px; =
margin-bottom: 0px; margin-left: 0px; "><FONT face=3D"Helvetica" =
size=3D"3" color=3D"#000000" style=3D"font: 12.0px Helvetica; color: =
#000000"><B>To: </B></FONT><FONT face=3D"Helvetica" size=3D"3" =
style=3D"font: 12.0px Helvetica"><A =
href=3D"mailto:mhtml@segate.sunet.se">mhtml@segate.sunet.se</A>, <A =
href=3D"mailto:ietf-822@imc.org">ietf-822@imc.org</A></FONT></DIV><DIV =
style=3D"margin-top: 0px; margin-right: 0px; margin-bottom: 0px; =
margin-left: 0px; "><FONT face=3D"Helvetica" size=3D"3" color=3D"#000000" =
style=3D"font: 12.0px Helvetica; color: #000000"><B>Subject: =
</B></FONT><FONT face=3D"Helvetica" size=3D"3" style=3D"font: 12.0px =
Helvetica"><B>[Fwd: [Standards-discuss] [HTML in =
email]]</B></FONT></DIV><DIV style=3D"margin-top: 0px; margin-right: =
0px; margin-bottom: 0px; margin-left: 0px; min-height: 14px; =
"><BR></DIV> <DIV style=3D"margin-top: 0px; margin-right: 0px; =
margin-bottom: 0px; margin-left: 0px; ">(sending again, with a =46rom =
address that has a chance of getting through....)</DIV><DIV =
style=3D"margin-top: 0px; margin-right: 0px; margin-bottom: 0px; =
margin-left: 0px; min-height: 14px; "><BR></DIV><DIV style=3D"margin-top: =
0px; margin-right: 0px; margin-bottom: 0px; margin-left: 0px; ">I guess =
that if the lists above still exist, there may be people with =
an</DIV><DIV style=3D"margin-top: 0px; margin-right: 0px; margin-bottom: =
0px; margin-left: 0px; ">interest in this proposal for a new W3C =
activity on them....</DIV><DIV style=3D"margin-top: 0px; margin-right: =
0px; margin-bottom: 0px; margin-left: 0px; min-height: 14px; =
"><BR></DIV><DIV style=3D"margin-top: 0px; margin-right: 0px; =
margin-bottom: 0px; margin-left: 0px; "><SPAN =
class=3D"Apple-converted-space">=A0 =A0 =A0 =A0 =A0 =A0 =A0 =A0 =
</SPAN>Harald</DIV><DIV style=3D"margin-top: 0px; margin-right: 0px; =
margin-bottom: 0px; margin-left: 0px; min-height: 14px; "><BR></DIV><DIV =
style=3D"margin-top: 0px; margin-right: 0px; margin-bottom: 0px; =
margin-left: 37px; text-indent: -37px; font: normal normal normal =
12px/normal Helvetica; color: rgb(0, 0, 0); min-height: 14px; =
"><B></B><BR></DIV><DIV style=3D"margin-top: 0px; margin-right: 0px; =
margin-bottom: 0px; margin-left: 37px; text-indent: -37px; "><FONT =
face=3D"Helvetica" size=3D"3" color=3D"#000000" style=3D"font: 12.0px =
Helvetica; color: #000000"><B>From: </B></FONT><FONT face=3D"Helvetica" =
size=3D"3" style=3D"font: 12.0px Helvetica">Daniel Glazman &lt;<A =
href=3D"mailto:daniel@glazman.org">daniel@glazman.org</A>&gt;</FONT></DIV>=
<DIV style=3D"margin-top: 0px; margin-right: 0px; margin-bottom: 0px; =
margin-left: 34px; text-indent: -34px; "><FONT face=3D"Helvetica" =
size=3D"3" color=3D"#000000" style=3D"font: 12.0px Helvetica; color: =
#000000"><B>Date: </B></FONT><FONT face=3D"Helvetica" size=3D"3" =
style=3D"font: 12.0px Helvetica">February 15, 2007 6:58:30 AM =
PST</FONT></DIV><DIV style=3D"margin-top: 0px; margin-right: 0px; =
margin-bottom: 0px; margin-left: 21px; text-indent: -21px; "><FONT =
face=3D"Helvetica" size=3D"3" color=3D"#000000" style=3D"font: 12.0px =
Helvetica; color: #000000"><B>To: </B></FONT><FONT face=3D"Helvetica" =
size=3D"3" style=3D"font: 12.0px Helvetica"><A =
href=3D"mailto:daniel@glazman.org">daniel@glazman.org</A></FONT></DIV><DIV=
 style=3D"margin-top: 0px; margin-right: 0px; margin-bottom: 0px; =
margin-left: 50px; text-indent: -50px; "><FONT face=3D"Helvetica" =
size=3D"3" color=3D"#000000" style=3D"font: 12.0px Helvetica; color: =
#000000"><B>Subject: </B></FONT><FONT face=3D"Helvetica" size=3D"3" =
style=3D"font: 12.0px Helvetica"><B>HTML in email</B></FONT></DIV><DIV =
style=3D"margin-top: 0px; margin-right: 0px; margin-bottom: 0px; =
margin-left: 0px; min-height: 14px; "><BR></DIV><DIV style=3D"margin-top: =
0px; margin-right: 0px; margin-bottom: 0px; margin-left: 0px; =
min-height: 14px; "><BR></DIV><DIV style=3D"margin-top: 0px; =
margin-right: 0px; margin-bottom: 0px; margin-left: 0px; =
">People,</DIV><DIV style=3D"margin-top: 0px; margin-right: 0px; =
margin-bottom: 0px; margin-left: 0px; min-height: 14px; "><BR></DIV><DIV =
style=3D"margin-top: 0px; margin-right: 0px; margin-bottom: 0px; =
margin-left: 0px; ">Following a few personal discussions with friends =
here in France,</DIV><DIV style=3D"margin-top: 0px; margin-right: 0px; =
margin-bottom: 0px; margin-left: 0px; ">and a thread in a W3C =
Members-only mailing-list, the W3C</DIV><DIV style=3D"margin-top: 0px; =
margin-right: 0px; margin-bottom: 0px; margin-left: 0px; ">has decided =
to launch a new mailing-list for issues related to HTML</DIV><DIV =
style=3D"margin-top: 0px; margin-right: 0px; margin-bottom: 0px; =
margin-left: 0px; ">in email: editing, rendering, interoperability, =
security, ...</DIV><DIV style=3D"margin-top: 0px; margin-right: 0px; =
margin-bottom: 0px; margin-left: 0px; ">The idea is - for instance - to =
take advantage of the work done by the</DIV><DIV style=3D"margin-top: =
0px; margin-right: 0px; margin-bottom: 0px; margin-left: 0px; ">future =
HTML WG to have an email-safe profile of HTML. That's only one</DIV><DIV =
style=3D"margin-top: 0px; margin-right: 0px; margin-bottom: 0px; =
margin-left: 0px; ">of the possibilities, just to give you an =
example.</DIV><DIV style=3D"margin-top: 0px; margin-right: 0px; =
margin-bottom: 0px; margin-left: 0px; ">BTW, that includes HTML emails =
sent from and received by mobile devices,</DIV><DIV style=3D"margin-top: =
0px; margin-right: 0px; margin-bottom: 0px; margin-left: 0px; ">of =
course...</DIV><DIV style=3D"margin-top: 0px; margin-right: 0px; =
margin-bottom: 0px; margin-left: 0px; min-height: 14px; "><BR></DIV><DIV =
style=3D"margin-top: 0px; margin-right: 0px; margin-bottom: 0px; =
margin-left: 0px; ">The mailing-list is <A =
href=3D"mailto:public-html-mail@w3.org">public-html-mail@w3.org</A></DIV><=
DIV style=3D"margin-top: 0px; margin-right: 0px; margin-bottom: 0px; =
margin-left: 0px; min-height: 14px; "><BR></DIV><DIV style=3D"margin-top: =
0px; margin-right: 0px; margin-bottom: 0px; margin-left: 0px; "><SPAN =
class=3D"Apple-converted-space">=A0 </SPAN><A =
href=3D"http://lists.w3.org/Archives/Public/public-html-mail/">http://list=
s.w3.org/Archives/Public/public-html-mail/</A></DIV><DIV =
style=3D"margin-top: 0px; margin-right: 0px; margin-bottom: 0px; =
margin-left: 0px; min-height: 14px; "><BR></DIV><DIV style=3D"margin-top: =
0px; margin-right: 0px; margin-bottom: 0px; margin-left: 0px; ">We also =
think of having an official W3C workshop on this topic in the</DIV><DIV =
style=3D"margin-top: 0px; margin-right: 0px; margin-bottom: 0px; =
margin-left: 0px; ">coming weeks or months, and we'll ask for position =
papers.</DIV><DIV style=3D"margin-top: 0px; margin-right: 0px; =
margin-bottom: 0px; margin-left: 0px; min-height: 14px; "><BR></DIV><DIV =
style=3D"margin-top: 0px; margin-right: 0px; margin-bottom: 0px; =
margin-left: 0px; ">Could you please kindly forward this invitation to =
join the mailing-</DIV><DIV style=3D"margin-top: 0px; margin-right: 0px; =
margin-bottom: 0px; margin-left: 0px; ">list to the key people working =
on HTML email editing and rendering</DIV><DIV style=3D"margin-top: 0px; =
margin-right: 0px; margin-bottom: 0px; margin-left: 0px; ">in your =
organization ? Please also forward to anyone working in this</DIV><DIV =
style=3D"margin-top: 0px; margin-right: 0px; margin-bottom: 0px; =
margin-left: 0px; ">scope, even outside of your own org, who can provide =
useful thoughts and</DIV><DIV style=3D"margin-top: 0px; margin-right: =
0px; margin-bottom: 0px; margin-left: 0px; ">help on this =
subject.</DIV><DIV style=3D"margin-top: 0px; margin-right: 0px; =
margin-bottom: 0px; margin-left: 0px; min-height: 14px; "><BR></DIV><DIV =
style=3D"margin-top: 0px; margin-right: 0px; margin-bottom: 0px; =
margin-left: 0px; ">I'll send another message to companies involved in =
direct email-</DIV><DIV style=3D"margin-top: 0px; margin-right: 0px; =
margin-bottom: 0px; margin-left: 0px; ">based marketing, because their =
input and help is needed here.</DIV><DIV style=3D"margin-top: 0px; =
margin-right: 0px; margin-bottom: 0px; margin-left: 0px; min-height: =
14px; "><BR></DIV><DIV style=3D"margin-top: 0px; margin-right: 0px; =
margin-bottom: 0px; margin-left: 0px; ">Best regards,</DIV><DIV =
style=3D"margin-top: 0px; margin-right: 0px; margin-bottom: 0px; =
margin-left: 0px; min-height: 14px; "><BR></DIV><DIV style=3D"margin-top: =
0px; margin-right: 0px; margin-bottom: 0px; margin-left: 0px; ">Daniel =
Glazman</DIV><DIV style=3D"margin-top: 0px; margin-right: 0px; =
margin-bottom: 0px; margin-left: 0px; min-height: 14px; "><BR></DIV><DIV =
style=3D"margin-top: 0px; margin-right: 0px; margin-bottom: 0px; =
margin-left: 0px; ">--<SPAN =
class=3D"Apple-converted-space">=A0</SPAN></DIV><DIV style=3D"margin-top: =
0px; margin-right: 0px; margin-bottom: 0px; margin-left: 0px; ">Best =
Regards,</DIV><DIV style=3D"margin-top: 0px; margin-right: 0px; =
margin-bottom: 0px; margin-left: 0px; ">--raman</DIV><DIV =
style=3D"margin-top: 0px; margin-right: 0px; margin-bottom: 0px; =
margin-left: 0px; min-height: 14px; "><BR></DIV><DIV style=3D"margin-top: =
0px; margin-right: 0px; margin-bottom: 0px; margin-left: 0px; =
">Title:<SPAN class=3D"Apple-converted-space">=A0 </SPAN>Research =
Scientist<SPAN class=3D"Apple-converted-space">=A0 =A0 =A0 =
</SPAN>Email:<SPAN class=3D"Apple-converted-space">=A0 </SPAN><A =
href=3D"mailto:raman@google.com">raman@google.com</A></DIV><DIV =
style=3D"margin-top: 0px; margin-right: 0px; margin-bottom: 0px; =
margin-left: 0px; ">WWW:<SPAN class=3D"Apple-converted-space">=A0 =A0 =
</SPAN><A =
href=3D"http://emacspeak.sf.net/raman/">http://emacspeak.sf.net/raman/</A>=
</DIV><DIV style=3D"margin-top: 0px; margin-right: 0px; margin-bottom: =
0px; margin-left: 0px; ">Google: tv+raman GTalk:<SPAN =
class=3D"Apple-converted-space">=A0 </SPAN><A =
href=3D"mailto:raman@google.com">raman@google.com</A>, <A =
href=3D"mailto:tv.raman.tv@gmail.com">tv.raman.tv@gmail.com</A></DIV><DIV =
style=3D"margin-top: 0px; margin-right: 0px; margin-bottom: 0px; =
margin-left: 0px; ">PGP:<SPAN class=3D"Apple-converted-space">=A0 =A0 =
</SPAN><A =
href=3D"http://emacspeak.sf.net/raman/raman-almaden.asc">http://emacspeak.=
sf.net/raman/raman-almaden.asc</A></DIV><DIV style=3D"margin-top: 0px; =
margin-right: 0px; margin-bottom: 0px; margin-left: 0px; min-height: =
14px; "><BR></DIV><DIV style=3D"margin-top: 0px; margin-right: 0px; =
margin-bottom: 0px; margin-left: 0px; min-height: 14px; "><BR></DIV><DIV =
style=3D"margin-top: 0px; margin-right: 0px; margin-bottom: 0px; =
margin-left: 0px; min-height: 14px; "><BR></DIV><DIV style=3D"margin-top: =
0px; margin-right: 0px; margin-bottom: 0px; margin-left: 0px; =
">_______________________________________________</DIV><DIV =
style=3D"margin-top: 0px; margin-right: 0px; margin-bottom: 0px; =
margin-left: 0px; ">Standards-discuss mailing list</DIV><DIV =
style=3D"margin-top: 0px; margin-right: 0px; margin-bottom: 0px; =
margin-left: 0px; "><A =
href=3D"mailto:Standards-discuss@google.com">Standards-discuss@google.com<=
/A></DIV><DIV style=3D"margin-top: 0px; margin-right: 0px; =
margin-bottom: 0px; margin-left: 0px; "><A =
href=3D"https://mailman.corp.google.com/mailman/listinfo/standards-discuss=
">https://mailman.corp.google.com/mailman/listinfo/standards-discuss</A></=
DIV><DIV style=3D"margin-top: 0px; margin-right: 0px; margin-bottom: =
0px; margin-left: 0px; min-height: 14px; "><BR></DIV><DIV =
style=3D"margin-top: 0px; margin-right: 0px; margin-bottom: 0px; =
margin-left: 0px; min-height: 14px; "><BR></DIV> =
</BLOCKQUOTE></DIV><BR></DIV></BODY></HTML>=

--Apple-Mail-26-333833273--




From discuss-bounces@apps.ietf.org Mon Feb 19 18:11:33 2007
Received: from [127.0.0.1] (helo=stiedprmman1.va.neustar.com)
	by megatron.ietf.org with esmtp (Exim 4.43)
	id 1HJHek-0005PP-PU; Mon, 19 Feb 2007 18:10:50 -0500
Received: from [10.91.34.44] (helo=ietf-mx.ietf.org)
	by megatron.ietf.org with esmtp (Exim 4.43) id 1HJHek-0005PK-Ao
	for discuss@apps.ietf.org; Mon, 19 Feb 2007 18:10:50 -0500
Received: from main.gmane.org ([80.91.229.2] helo=ciao.gmane.org)
	by ietf-mx.ietf.org with esmtp (Exim 4.43) id 1HJHeh-0005J1-Iy
	for discuss@apps.ietf.org; Mon, 19 Feb 2007 18:10:50 -0500
Received: from list by ciao.gmane.org with local (Exim 4.43)
	id 1HJGRY-0002Dt-UM
	for discuss@apps.ietf.org; Mon, 19 Feb 2007 22:53:09 +0100
Received: from 212.82.251.170 ([212.82.251.170])
	by main.gmane.org with esmtp (Gmexim 0.1 (Debian))
	id 1AlnuQ-0007hv-00
	for <discuss@apps.ietf.org>; Mon, 19 Feb 2007 22:53:08 +0100
Received: from nobody by 212.82.251.170 with local (Gmexim 0.1 (Debian))
	id 1AlnuQ-0007hv-00
	for <discuss@apps.ietf.org>; Mon, 19 Feb 2007 22:53:08 +0100
X-Injected-Via-Gmane: http://gmane.org/
To: discuss@apps.ietf.org
From: Frank Ellermann <nobody@xyzzy.claranet.de>
Subject: Re: draft-klensin-unicode-escapes-02.txt
Date: Mon, 19 Feb 2007 22:34:44 +0100
Organization: <URL:http://purl.net/xyzzy>
Lines: 41
Message-ID: <45DA17F4.4857@xyzzy.claranet.de>
References: <74711BCF624DBEC4F2C000C5@p3.JCK.COM>
Mime-Version: 1.0
Content-Type: text/plain; charset=us-ascii
Content-Transfer-Encoding: 7bit
X-Complaints-To: usenet@sea.gmane.org
X-Gmane-NNTP-Posting-Host: 212.82.251.170
X-Mailer: Mozilla 3.0 (OS/2; U)
X-Spam-Score: 1.6 (+)
X-Scan-Signature: a7d6aff76b15f3f56fcb94490e1052e4
X-BeenThere: discuss@apps.ietf.org
X-Mailman-Version: 2.1.5
Precedence: list
List-Id: general discussion of application-layer protocols
	<discuss.apps.ietf.org>
List-Unsubscribe: <https://www1.ietf.org/mailman/listinfo/discuss>,
	<mailto:discuss-request@apps.ietf.org?subject=unsubscribe>
List-Post: <mailto:discuss@apps.ietf.org>
List-Help: <mailto:discuss-request@apps.ietf.org?subject=help>
List-Subscribe: <https://www1.ietf.org/mailman/listinfo/discuss>,
	<mailto:discuss-request@apps.ietf.org?subject=subscribe>
Errors-To: discuss-bounces@apps.ietf.org

John C Klensin wrote:

> Anyone who prefers the latter should please send the ABNF they would
> like to see.

Here's what I'd do with the existing ABNF:

   EmbeddedUnicodeChar =  BMP-form / Full-form
   BMP-form  =  %x5C.75 4HEXDIG   ; starting with lower case "\u"
   Full-form =  %x5C.55 8HEXDIG   ; starting with upper case "\U"

IOW no more <Hex-quad> because you didn't use it elsewhere, and
replacing <HexDigit> by <HEXDIG>, because the latter is already
defined in RFC 4234.

For the XML version I propose to adopt and fix the RFC 4646 ABNF:

   UNICHAR    = %x26.23.78 2*6HEXDIG ";"    ; starts with "&#x"

With a remark in the prose, that a literal "&" can be expressed by
&#x26; when using this style.  Otherwise folks could be tempted to
use &amp; - but we don't want them to try that.

You don't need ABNF in 5.4, it's in essence the same as in 5.1.

For 5.3 (perl) I don't know the correct syntax, if it's like 5.2:

  UNICODEPOINT = %x5C.78 "{" 2*6HEXDIG "}"  ; starts with "\x"

If the x is case insensitive it's simply:

  UNICODEPOINT = "\x{" 2*6HEXDIG "}"

The name "UNICODEPOINT" is horrible, please find something better, I
tried to use a new name.

Please add the [CharMod] reference (informative) with a note about
its C042..C048 conformance criteria in ch. 4.6 (character escaping).

Frank






From discuss-bounces@apps.ietf.org Mon Feb 19 19:37:55 2007
Received: from [127.0.0.1] (helo=stiedprmman1.va.neustar.com)
	by megatron.ietf.org with esmtp (Exim 4.43)
	id 1HJJ03-000415-Sy; Mon, 19 Feb 2007 19:36:55 -0500
Received: from [10.91.34.44] (helo=ietf-mx.ietf.org)
	by megatron.ietf.org with esmtp (Exim 4.43) id 1HJJ02-0003wE-Lr
	for discuss@apps.ietf.org; Mon, 19 Feb 2007 19:36:54 -0500
Received: from smtp-out.sendmail.com ([209.246.26.45] helo=foon.sendmail.com)
	by ietf-mx.ietf.org with esmtp (Exim 4.43) id 1HJJ01-000096-AC
	for discuss@apps.ietf.org; Mon, 19 Feb 2007 19:36:54 -0500
Received: from [10.201.0.245] (adsl-64-58-1-252.mho.net [64.58.1.252] (may be
	forged)) (authenticated bits=0)
	by foon.sendmail.com (Switch-3.2.5/Switch-3.2.0) with ESMTP id
	l1K0aYBc000809
	(version=TLSv1/SSLv3 cipher=DHE-RSA-AES256-SHA bits=256 verify=NO)
	for <discuss@apps.ietf.org>; Mon, 19 Feb 2007 16:36:37 -0800
X-DKIM: Sendmail DKIM Filter v0.5.1 foon.sendmail.com l1K0aYBc000809
DKIM-Signature: a=rsa-sha1; c=relaxed/simple; d=sendmail.com; s=tls.dkim;
	t=1171931798; bh=QShmx0CmqsyD6OnJI+FpAWKDjNQ=; h=X-DomainKeys:
	DomainKey-Signature:Date:From:X-X-Sender:To:Subject:In-Reply-To:
	Message-ID:References:MIME-Version:Content-Type; b=nt40W7vCrNtfOZw8
	cMOBVT0Jlfhfc24w3CxD8L8sZOtrx8M92taywCfqzHh3eN/Nb9s/tbzUzotFSKz5/c2
	1PE3eB3Me/W+Ij2VrgYJRZEp5UeY29NduqQJX3xBtYrZpz/Y+wwfMnP6MoHi+pyvAuY
	qeVyKXix8rG3RNwJppFkE=
X-DomainKeys: Sendmail DomainKeys Filter v0.4.1 foon.sendmail.com
	l1K0aYBc000809
DomainKey-Signature: a=rsa-sha1; s=tls; d=sendmail.com; c=nofws; q=dns;
	h=date:from:x-x-sender:to:subject:in-reply-to:message-id:
	references:mime-version:content-type;
	b=WG5xYeI04+dIoBdcONf4H0fGA9BqVCxeTOvLnD3gNe/G1UHj/8CG3vxbtyMuLm3IU
	GSr98WEcb89I18+roNltka8yjM8m5jnDJc0svuXYWLbNhCK2xJj41IrhiEJS5p0iteN
	eQrK9sROC/S1Wnjt5iyhdQVPcPKhgTks0w+CM+A=
Date: Mon, 19 Feb 2007 17:36:30 -0700
From: Philip Guenther <guenther+ietf@sendmail.com>
X-X-Sender: guenther@vanye.mho.net
To: discuss@apps.ietf.org
Subject: Re: draft-klensin-unicode-escapes-02.txt
In-Reply-To: <45DA17F4.4857@xyzzy.claranet.de>
Message-ID: <Pine.BSO.4.64.0702191627560.12052@vanye.mho.net>
References: <74711BCF624DBEC4F2C000C5@p3.JCK.COM>
	<45DA17F4.4857@xyzzy.claranet.de>
MIME-Version: 1.0
Content-Type: TEXT/PLAIN; charset=US-ASCII; format=flowed
X-Spam-Score: 0.0 (/)
X-Scan-Signature: 0bc60ec82efc80c84b8d02f4b0e4de22
X-BeenThere: discuss@apps.ietf.org
X-Mailman-Version: 2.1.5
Precedence: list
List-Id: general discussion of application-layer protocols
	<discuss.apps.ietf.org>
List-Unsubscribe: <https://www1.ietf.org/mailman/listinfo/discuss>,
	<mailto:discuss-request@apps.ietf.org?subject=unsubscribe>
List-Post: <mailto:discuss@apps.ietf.org>
List-Help: <mailto:discuss-request@apps.ietf.org?subject=help>
List-Subscribe: <https://www1.ietf.org/mailman/listinfo/discuss>,
	<mailto:discuss-request@apps.ietf.org?subject=subscribe>
Errors-To: discuss-bounces@apps.ietf.org

On Mon, 19 Feb 2007, Frank Ellermann wrote:
...
> For 5.3 (perl) I don't know the correct syntax, if it's like 5.2:
>
>  UNICODEPOINT = %x5C.78 "{" 2*6HEXDIG "}"  ; starts with "\x"

That looks good to me.  Perl accepts values that don't match that syntax, 
but we don't want to get into trying to define perl's corner cases (or, 
indeed, those of *any* language not originally defined by RFC).  This 
should be a 'generate' syntax, not an 'accept' syntax, IMHO.


> If the x is case insensitive it's simply:
...


The 'x' is case sensitive in perl.  \X has no meaning in plain strings, 
but in regexps it is a built in pattern of sorts.  To quote the 
perlunicode(1) manpage:

        o   The special pattern "\X" matches any extended Unicode
            sequence--"a combining character sequence" in Stan-
            dardese--where the first character is a base character
            and subsequent characters are mark characters that
            apply to the base character.  <...>


Philip Guenther




From discuss-bounces@apps.ietf.org Wed Feb 21 03:46:07 2007
Received: from [127.0.0.1] (helo=stiedprmman1.va.neustar.com)
	by megatron.ietf.org with esmtp (Exim 4.43)
	id 1HJn6E-0002l7-Rh; Wed, 21 Feb 2007 03:45:18 -0500
Received: from [10.91.34.44] (helo=ietf-mx.ietf.org)
	by megatron.ietf.org with esmtp (Exim 4.43) id 1HJn6D-0002kP-Mx
	for discuss@apps.ietf.org; Wed, 21 Feb 2007 03:45:17 -0500
Received: from anchor-internal-1.mail.demon.net ([195.173.56.100])
	by ietf-mx.ietf.org with esmtp (Exim 4.43) id 1HJn69-00069J-B9
	for discuss@apps.ietf.org; Wed, 21 Feb 2007 03:45:17 -0500
Received: from finch-staff-1.server.demon.net (finch-staff-1.server.demon.net [193.195.224.1])
	by anchor-internal-1.mail.demon.net with ESMTPœ id l1L8j9aY014801Wed, 21 Feb 2007 08:45:09 GMT
Received: from clive by finch-staff-1.server.demon.net with local (Exim 3.36
	#1) id 1HJn5F-000Oh3-00; Wed, 21 Feb 2007 08:44:17 +0000
Date: Wed, 21 Feb 2007 08:44:17 +0000
From: "Clive D.W. Feather" <clive@demon.net>
To: Frank Ellermann <nobody@xyzzy.claranet.de>
Subject: Re: draft-klensin-unicode-escapes-02.txt
Message-ID: <20070221084417.GA93361@finch-staff-1.thus.net>
References: <74711BCF624DBEC4F2C000C5@p3.JCK.COM>
	<45DA17F4.4857@xyzzy.claranet.de>
Mime-Version: 1.0
Content-Type: text/plain; charset=us-ascii
Content-Disposition: inline
In-Reply-To: <45DA17F4.4857@xyzzy.claranet.de>
User-Agent: Mutt/1.5.3i
X-Spam-Score: 0.0 (/)
X-Scan-Signature: 79899194edc4f33a41f49410777972f8
Cc: discuss@apps.ietf.org
X-BeenThere: discuss@apps.ietf.org
X-Mailman-Version: 2.1.5
Precedence: list
List-Id: general discussion of application-layer protocols
	<discuss.apps.ietf.org>
List-Unsubscribe: <https://www1.ietf.org/mailman/listinfo/discuss>,
	<mailto:discuss-request@apps.ietf.org?subject=unsubscribe>
List-Post: <mailto:discuss@apps.ietf.org>
List-Help: <mailto:discuss-request@apps.ietf.org?subject=help>
List-Subscribe: <https://www1.ietf.org/mailman/listinfo/discuss>,
	<mailto:discuss-request@apps.ietf.org?subject=subscribe>
Errors-To: discuss-bounces@apps.ietf.org

Frank Ellermann said:
> Here's what I'd do with the existing ABNF:
> 
>    EmbeddedUnicodeChar =  BMP-form / Full-form
>    BMP-form  =  %x5C.75 4HEXDIG   ; starting with lower case "\u"
>    Full-form =  %x5C.55 8HEXDIG   ; starting with upper case "\U"
> 
> IOW no more <Hex-quad> because you didn't use it elsewhere,

I would prefer this as well. The <Hex-quad> comes from the C Standard,
where we didn't have the <number><token> notation and so would have had to
write the equivalent of:

    %x5C.55 HEXDIG HEXDIG HEXDIG HEXDIG HEXDIG HEXDIG HEXDIG HEXDIG

-- 
Clive D.W. Feather  | Work:  <clive@demon.net>   | Tel:    +44 20 8495 6138
Internet Expert     | Home:  <clive@davros.org>  | Fax:    +44 870 051 9937
Demon Internet      | WWW: http://www.davros.org | Mobile: +44 7973 377646
THUS plc            |                            |




From discuss-bounces@apps.ietf.org Wed Feb 21 13:44:35 2007
Received: from [127.0.0.1] (helo=stiedprmman1.va.neustar.com)
	by megatron.ietf.org with esmtp (Exim 4.43)
	id 1HJwRV-0001nd-T3; Wed, 21 Feb 2007 13:43:53 -0500
Received: from [10.91.34.44] (helo=ietf-mx.ietf.org)
	by megatron.ietf.org with esmtp (Exim 4.43) id 1HJwRU-0001nH-9A
	for discuss@apps.ietf.org; Wed, 21 Feb 2007 13:43:52 -0500
Received: from shu.cs.utk.edu ([160.36.56.39])
	by ietf-mx.ietf.org with esmtp (Exim 4.43) id 1HJwRR-0000SN-2W
	for discuss@apps.ietf.org; Wed, 21 Feb 2007 13:43:52 -0500
Received: from localhost (localhost [127.0.0.1])
	by shu.cs.utk.edu (Postfix) with ESMTP id 048ED56411;
	Wed, 21 Feb 2007 13:43:45 -0500 (EST)
X-Virus-Scanned: by amavisd-new with ClamAV and SpamAssasin at cs.utk.edu
Received: from shu.cs.utk.edu ([127.0.0.1])
	by localhost (shu.cs.utk.edu [127.0.0.1]) (amavisd-new, port 10024)
	with ESMTP id PtiFf8Z+IEN2; Wed, 21 Feb 2007 13:43:40 -0500 (EST)
Received: from [192.168.0.2] (user-119b1dm.biz.mindspring.com [66.149.133.182])
	by shu.cs.utk.edu (Postfix) with ESMTP id 677FB56470;
	Wed, 21 Feb 2007 13:43:21 -0500 (EST)
Message-ID: <45DC92C3.6040607@cs.utk.edu>
Date: Wed, 21 Feb 2007 13:43:15 -0500
From: Keith Moore <moore@cs.utk.edu>
User-Agent: Thunderbird 1.5.0.9 (Macintosh/20061207)
MIME-Version: 1.0
To: Lisa Dusseault <ldusseault@commerce.net>
Subject: Re: Fwd: [Standards-discuss] [HTML in email]
References: <45D94EA0.2050706@alvestrand.no>
	<EC0B32D0-39FC-4985-8A51-7D57C68C344E@commerce.net>
In-Reply-To: <EC0B32D0-39FC-4985-8A51-7D57C68C344E@commerce.net>
Content-Type: text/plain; charset=ISO-8859-1; format=flowed
Content-Transfer-Encoding: 7bit
X-Spam-Score: 0.0 (/)
X-Scan-Signature: 8ac499381112328dd60aea5b1ff596ea
Cc: Apps Discuss <discuss@apps.ietf.org>,
	HTTP Working Group <ietf-http-wg@w3.org>
X-BeenThere: discuss@apps.ietf.org
X-Mailman-Version: 2.1.5
Precedence: list
List-Id: general discussion of application-layer protocols
	<discuss.apps.ietf.org>
List-Unsubscribe: <https://www1.ietf.org/mailman/listinfo/discuss>,
	<mailto:discuss-request@apps.ietf.org?subject=unsubscribe>
List-Post: <mailto:discuss@apps.ietf.org>
List-Help: <mailto:discuss-request@apps.ietf.org?subject=help>
List-Subscribe: <https://www1.ietf.org/mailman/listinfo/discuss>,
	<mailto:discuss-request@apps.ietf.org?subject=subscribe>
Errors-To: discuss-bounces@apps.ietf.org

seems like a worthwhile discussion.  pity they want to make it a 
members-only deal; a lot of the relevant expertise isn't concentrated in 
the w3c membership.

Keith





From discuss-bounces@apps.ietf.org Wed Feb 21 14:36:40 2007
Received: from [127.0.0.1] (helo=stiedprmman1.va.neustar.com)
	by megatron.ietf.org with esmtp (Exim 4.43)
	id 1HJxFs-0007is-0E; Wed, 21 Feb 2007 14:35:56 -0500
Received: from [10.91.34.44] (helo=ietf-mx.ietf.org)
	by megatron.ietf.org with esmtp (Exim 4.43) id 1HJxFq-0007ib-L7
	for discuss@apps.ietf.org; Wed, 21 Feb 2007 14:35:54 -0500
Received: from shu.cs.utk.edu ([160.36.56.39])
	by ietf-mx.ietf.org with esmtp (Exim 4.43) id 1HJxFp-0004vl-FD
	for discuss@apps.ietf.org; Wed, 21 Feb 2007 14:35:54 -0500
Received: from localhost (localhost [127.0.0.1])
	by shu.cs.utk.edu (Postfix) with ESMTP id 982CC56437;
	Wed, 21 Feb 2007 14:35:52 -0500 (EST)
X-Virus-Scanned: by amavisd-new with ClamAV and SpamAssasin at cs.utk.edu
Received: from shu.cs.utk.edu ([127.0.0.1])
	by localhost (shu.cs.utk.edu [127.0.0.1]) (amavisd-new, port 10024)
	with ESMTP id rkbXejY1S4FI; Wed, 21 Feb 2007 14:35:44 -0500 (EST)
Received: from [192.168.0.2] (user-119b1dm.biz.mindspring.com [66.149.133.182])
	by shu.cs.utk.edu (Postfix) with ESMTP id B69C456468;
	Wed, 21 Feb 2007 14:35:38 -0500 (EST)
Message-ID: <45DC9F19.8070007@cs.utk.edu>
Date: Wed, 21 Feb 2007 14:35:53 -0500
From: Keith Moore <moore@cs.utk.edu>
User-Agent: Thunderbird 1.5.0.9 (Macintosh/20061207)
MIME-Version: 1.0
To: Keith Moore <moore@cs.utk.edu>
Subject: Re: Fwd: [Standards-discuss] [HTML in email]
References: <45D94EA0.2050706@alvestrand.no>	<EC0B32D0-39FC-4985-8A51-7D57C68C344E@commerce.net>
	<45DC92C3.6040607@cs.utk.edu>
In-Reply-To: <45DC92C3.6040607@cs.utk.edu>
Content-Type: text/plain; charset=ISO-8859-1; format=flowed
Content-Transfer-Encoding: 7bit
X-Spam-Score: 0.0 (/)
X-Scan-Signature: 08e48e05374109708c00c6208b534009
Cc: Apps Discuss <discuss@apps.ietf.org>,
	HTTP Working Group <ietf-http-wg@w3.org>,
	Lisa Dusseault <ldusseault@commerce.net>
X-BeenThere: discuss@apps.ietf.org
X-Mailman-Version: 2.1.5
Precedence: list
List-Id: general discussion of application-layer protocols
	<discuss.apps.ietf.org>
List-Unsubscribe: <https://www1.ietf.org/mailman/listinfo/discuss>,
	<mailto:discuss-request@apps.ietf.org?subject=unsubscribe>
List-Post: <mailto:discuss@apps.ietf.org>
List-Help: <mailto:discuss-request@apps.ietf.org?subject=help>
List-Subscribe: <https://www1.ietf.org/mailman/listinfo/discuss>,
	<mailto:discuss-request@apps.ietf.org?subject=subscribe>
Errors-To: discuss-bounces@apps.ietf.org

sorry, I misread the earlier message.  it's a public list.
> seems like a worthwhile discussion.  pity they want to make it a 
> members-only deal; a lot of the relevant expertise isn't concentrated 
> in the w3c membership.
>
> Keith
>
>




From discuss-bounces@apps.ietf.org Fri Feb 23 12:35:54 2007
Received: from [127.0.0.1] (helo=stiedprmman1.va.neustar.com)
	by megatron.ietf.org with esmtp (Exim 4.43)
	id 1HKeJ8-0001Iy-77; Fri, 23 Feb 2007 12:34:10 -0500
Received: from [10.91.34.44] (helo=ietf-mx.ietf.org)
	by megatron.ietf.org with esmtp (Exim 4.43) id 1HKeJ5-0001HT-Uv
	for discuss@apps.ietf.org; Fri, 23 Feb 2007 12:34:09 -0500
Received: from ns.jck.com ([209.187.148.211] helo=bs.jck.com)
	by ietf-mx.ietf.org with esmtp (Exim 4.43) id 1HKeJ2-0002PA-JU
	for discuss@apps.ietf.org; Fri, 23 Feb 2007 12:34:05 -0500
Received: from [127.0.0.1] (helo=p3.JCK.COM)
	by bs.jck.com with esmtp (Exim 4.34)
	id 1HKeIv-0006S3-Ef; Fri, 23 Feb 2007 12:33:57 -0500
Date: Fri, 23 Feb 2007 12:33:56 -0500
From: John C Klensin <john-ietf@jck.com>
To: Frank Ellermann <nobody@xyzzy.claranet.de>,
	"Clive D.W. Feather" <clive@demon.net>
Subject: draft-klensin-unicode-escapes-03.txt (was: Re:
	draft-klensin-unicode-escapes-02.txt)
Message-ID: <754B21F623BA398D14CCDE97@p3.JCK.COM>
In-Reply-To: <45DA17F4.4857@xyzzy.claranet.de>
References: <74711BCF624DBEC4F2C000C5@p3.JCK.COM>
	<45DA17F4.4857@xyzzy.claranet.de>
X-Mailer: Mulberry/4.0.7 (Win32)
MIME-Version: 1.0
Content-Type: text/plain; charset=us-ascii
Content-Transfer-Encoding: 7bit
Content-Disposition: inline
X-Spam-Score: 0.0 (/)
X-Scan-Signature: 4d87d2aa806f79fed918a62e834505ca
Cc: discuss@apps.ietf.org
X-BeenThere: discuss@apps.ietf.org
X-Mailman-Version: 2.1.5
Precedence: list
List-Id: general discussion of application-layer protocols
	<discuss.apps.ietf.org>
List-Unsubscribe: <https://www1.ietf.org/mailman/listinfo/discuss>,
	<mailto:discuss-request@apps.ietf.org?subject=unsubscribe>
List-Post: <mailto:discuss@apps.ietf.org>
List-Help: <mailto:discuss-request@apps.ietf.org?subject=help>
List-Subscribe: <https://www1.ietf.org/mailman/listinfo/discuss>,
	<mailto:discuss-request@apps.ietf.org?subject=subscribe>
Errors-To: discuss-bounces@apps.ietf.org



--On Monday, 19 February, 2007 22:34 +0100 Frank Ellermann
<nobody@xyzzy.claranet.de> wrote:

> John C Klensin wrote:
> 
>> Anyone who prefers the latter should please send the ABNF
>> they would like to see.
> 
> Here's what I'd do with the existing ABNF:
> 
>    EmbeddedUnicodeChar =  BMP-form / Full-form
>    BMP-form  =  %x5C.75 4HEXDIG   ; starting with lower case
> "\u"    Full-form =  %x5C.55 8HEXDIG   ; starting with upper
> case "\U"
> 
> IOW no more <Hex-quad> because you didn't use it elsewhere, and
> replacing <HexDigit> by <HEXDIG>, because the latter is already
> defined in RFC 4234.
>...

Done, including removal of <HexDigit>

> Please add the [CharMod] reference (informative) with a note
> about its C042..C048 conformance criteria in ch. 4.6
> (character escaping).

A reference to [CharMod] has been included, but as part of a
comment (per discussion with Charles on 9 February) about what
this document does not try to address.

I think most (and hope all) of the other comments since -02 was
posted have been picked up, with one major exception.  I believe
that complexity is the enemy of this type of work.  So, while it
is clearly possible to figure out how to write, e.g.,
   \u(NNNN, NNNNN, NNNN)
rather than
   \u(NNNN)\u(NNNNN)\u(NNNN)
or to incorporate named characters into escaped strings, I don't
think doing that is wise.   YMMD, of course, and, if you, I'm
sure I will hear about it.

-03 is now in the posting queue.  I just realized that I didn't
update Section 8 (the change log) for either this version or
-02.  The major changes are the elimination of a specific
recommendation in -02 and cleaning up of a lot of text and
fixing the ABNF in -03.

      john






From discuss-bounces@apps.ietf.org Fri Feb 23 20:57:55 2007
Received: from [127.0.0.1] (helo=stiedprmman1.va.neustar.com)
	by megatron.ietf.org with esmtp (Exim 4.43)
	id 1HKm8r-0006Mx-4k; Fri, 23 Feb 2007 20:56:05 -0500
Received: from [10.91.34.44] (helo=ietf-mx.ietf.org)
	by megatron.ietf.org with esmtp (Exim 4.43) id 1HKm6n-0005t8-Uc
	for discuss@apps.ietf.org; Fri, 23 Feb 2007 20:53:57 -0500
Received: from ns.jck.com ([209.187.148.211] helo=bs.jck.com)
	by ietf-mx.ietf.org with esmtp (Exim 4.43) id 1HKm3B-0004zb-Lk
	for discuss@apps.ietf.org; Fri, 23 Feb 2007 20:50:15 -0500
Received: from [127.0.0.1] (helo=p2) by bs.jck.com with esmtp (Exim 4.34)
	id 1HKm2y-000Act-Em
	for discuss@apps.ietf.org; Fri, 23 Feb 2007 20:50:01 -0500
Date: Fri, 23 Feb 2007 20:49:57 -0500
From: John C Klensin <john-ietf@jck.com>
To: discuss@apps.ietf.org
Subject: FWD: I-D ACTION:draft-klensin-unicode-escapes-03.txt
Message-ID: <B91C111FCFE3BD7A4D65967C@[192.168.1.110]>
X-Mailer: Mulberry/4.0.7 (Win32)
MIME-Version: 1.0
Content-Type: multipart/mixed;
	boundary="==========A6B7712B9747A00683EF=========="
X-Spam-Score: 0.1 (/)
X-Scan-Signature: b045c2b078f76b9f842d469de8a32de3
X-BeenThere: discuss@apps.ietf.org
X-Mailman-Version: 2.1.5
Precedence: list
List-Id: general discussion of application-layer protocols
	<discuss.apps.ietf.org>
List-Unsubscribe: <https://www1.ietf.org/mailman/listinfo/discuss>,
	<mailto:discuss-request@apps.ietf.org?subject=unsubscribe>
List-Post: <mailto:discuss@apps.ietf.org>
List-Help: <mailto:discuss-request@apps.ietf.org?subject=help>
List-Subscribe: <https://www1.ietf.org/mailman/listinfo/discuss>,
	<mailto:discuss-request@apps.ietf.org?subject=subscribe>
Errors-To: discuss-bounces@apps.ietf.org

--==========A6B7712B9747A00683EF==========
Content-Type: text/plain; charset=us-ascii; format=flowed
Content-Transfer-Encoding: 7bit
Content-Disposition: inline

As discussed in my note earlier today...

    john


------------ Forwarded Message ------------
Date: Friday, February 23, 2007 6:50 PM -0500
From: Internet-Drafts@ietf.org
To: i-d-announce@ietf.org
Subject: I-D ACTION:draft-klensin-unicode-escapes-03.txt

A New Internet-Draft is available from the on-line
Internet-Drafts  directories.


	Title		: ASCII Escaping of Unicode Characters
	Author(s)	: J. Klensin
	Filename	: draft-klensin-unicode-escapes-03.txt
	Pages		: 12
	Date		: 2007-2-23
	
There are a number of circumstances in which an escape mechanism
is    needed in conjunction with a protocol to encode characters
that    cannot be represented or transmitted directly.  With
ASCII coding the    traditional escape has been either the
decimal or hexadecimal offset    of the character, written in a
variety of different ways.  The move    to Unicode, where
characters occupy two or more octets and may be    coded in
several different forms, has further complicated the    question
of escapes.  This document discusses some options now in use
and discusses considerations for selecting one for use in new
IETF    protocols and protocols that are now being
internationalized.

A URL for this Internet-Draft is:
http://www.ietf.org/internet-drafts/draft-klensin-unicode-escape
s-03.txt

To remove yourself from the I-D Announcement list, send a
message to  i-d-announce-request@ietf.org with the word
unsubscribe in the body of  the message.
You can also visit
https://www1.ietf.org/mailman/listinfo/I-D-announce  to change
your subscription settings.

Internet-Drafts are also available by anonymous FTP. Login with
the  username "anonymous" and a password of your e-mail address.
After  logging in, type "cd internet-drafts" and then
"get draft-klensin-unicode-escapes-03.txt".

A list of Internet-Drafts directories can be found in
http://www.ietf.org/shadow.html
or ftp://ftp.ietf.org/ietf/1shadow-sites.txt

Internet-Drafts can also be obtained by e-mail.

Send a message to:
	mailserv@ietf.org.
In the body type:
	"FILE /internet-drafts/draft-klensin-unicode-escapes-03.txt".
	
NOTE:	The mail server at ietf.org can return the document in
	MIME-encoded form by using the "mpack" utility.  To use this
	feature, insert the command "ENCODING mime" before the "FILE"
	command.  To decode the response(s), you will need "munpack" or
	a MIME-compliant mail reader.  Different MIME-compliant mail
readers 	exhibit different behavior, especially when dealing with
	"multipart" MIME messages (i.e. documents which have been split
	up into multiple messages), so check your local documentation on
	how to manipulate these messages.

Below is the data which will enable a MIME compliant mail reader
implementation to automatically retrieve the ASCII version of the
Internet-Draft.

---------- End Forwarded Message ----------




--==========A6B7712B9747A00683EF==========
Content-Type: message/rfc822;
	name="I-D ACTION:draft-klensin-unicode-escapes-03.txt"

Return-path: <i-d-announce-bounces@ietf.org>
Envelope-to: klensin+ietf@jck.com
Delivery-date: Fri, 23 Feb 2007 18:54:22 -0500
Received: from [156.154.16.145] (helo=megatron.ietf.org)
	by bs.jck.com with esmtp (Exim 4.34) id 1HKkF3-0009Xa-Sn
	for klensin+ietf@jck.com; Fri, 23 Feb 2007 18:54:22 -0500
Received: from [127.0.0.1] (helo=stiedprmman1.va.neustar.com)
	by megatron.ietf.org with esmtp (Exim 4.43)
	id 1HKkB2-00052o-K3; Fri, 23 Feb 2007 18:50:12 -0500
Received: from [10.91.34.44] (helo=ietf-mx.ietf.org)
	by megatron.ietf.org with esmtp (Exim 4.43) id 1HKkAt-0004wY-7R
	for i-d-announce@ietf.org; Fri, 23 Feb 2007 18:50:03 -0500
Received: from ns0.neustar.com ([156.154.16.158])
	by ietf-mx.ietf.org with esmtp (Exim 4.43) id 1HKkAs-0005Gg-QX
	for i-d-announce@ietf.org; Fri, 23 Feb 2007 18:50:03 -0500
Received: from stiedprstage1.ietf.org (stiedprstage1.va.neustar.com
	[10.31.47.10]) by ns0.neustar.com (Postfix) with ESMTP id A556A328F1
	for <i-d-announce@ietf.org>; Fri, 23 Feb 2007 23:50:02 +0000 (GMT)
Received: from ietf by stiedprstage1.ietf.org with local (Exim 4.43)
	id 1HKkAs-00024l-IB
	for i-d-announce@ietf.org; Fri, 23 Feb 2007 18:50:02 -0500
Content-Type: Multipart/Mixed; Boundary="NextPart"
Mime-Version: 1.0
To: i-d-announce@ietf.org
Cc: 
From: Internet-Drafts@ietf.org
Message-Id: <E1HKkAs-00024l-IB@stiedprstage1.ietf.org>
Date: Fri, 23 Feb 2007 18:50:02 -0500
X-Spam-Score: -2.5 (--)
X-Scan-Signature: f66b12316365a3fe519e75911daf28a8
Subject: I-D ACTION:draft-klensin-unicode-escapes-03.txt 
X-BeenThere: i-d-announce@ietf.org
X-Mailman-Version: 2.1.5
Precedence: list
Reply-To: internet-drafts@ietf.org
List-Id: i-d-announce.ietf.org
List-Unsubscribe: <https://www1.ietf.org/mailman/listinfo/i-d-announce>,
	<mailto:i-d-announce-request@ietf.org?subject=unsubscribe>
List-Archive: <http://www1.ietf.org/pipermail/i-d-announce>
List-Post: <mailto:i-d-announce@ietf.org>
List-Help: <mailto:i-d-announce-request@ietf.org?subject=help>
List-Subscribe: <https://www1.ietf.org/mailman/listinfo/i-d-announce>,
	<mailto:i-d-announce-request@ietf.org?subject=subscribe>
Errors-To: i-d-announce-bounces@ietf.org

--NextPart

A New Internet-Draft is available from the on-line Internet-Drafts 
directories.


	Title		: ASCII Escaping of Unicode Characters
	Author(s)	: J. Klensin
	Filename	: draft-klensin-unicode-escapes-03.txt
	Pages		: 12
	Date		: 2007-2-23
	
There are a number of circumstances in which an escape mechanism is
   needed in conjunction with a protocol to encode characters that
   cannot be represented or transmitted directly.  With ASCII coding the
   traditional escape has been either the decimal or hexadecimal offset
   of the character, written in a variety of different ways.  The move
   to Unicode, where characters occupy two or more octets and may be
   coded in several different forms, has further complicated the
   question of escapes.  This document discusses some options now in use
   and discusses considerations for selecting one for use in new IETF
   protocols and protocols that are now being internationalized.

A URL for this Internet-Draft is:
http://www.ietf.org/internet-drafts/draft-klensin-unicode-escapes-03.txt

To remove yourself from the I-D Announcement list, send a message to 
i-d-announce-request@ietf.org with the word unsubscribe in the body of 
the message. 
You can also visit https://www1.ietf.org/mailman/listinfo/I-D-announce 
to change your subscription settings.

Internet-Drafts are also available by anonymous FTP. Login with the 
username "anonymous" and a password of your e-mail address. After 
logging in, type "cd internet-drafts" and then 
"get draft-klensin-unicode-escapes-03.txt".

A list of Internet-Drafts directories can be found in
http://www.ietf.org/shadow.html 
or ftp://ftp.ietf.org/ietf/1shadow-sites.txt

Internet-Drafts can also be obtained by e-mail.

Send a message to:
	mailserv@ietf.org.
In the body type:
	"FILE /internet-drafts/draft-klensin-unicode-escapes-03.txt".
	
NOTE:	The mail server at ietf.org can return the document in
	MIME-encoded form by using the "mpack" utility.  To use this
	feature, insert the command "ENCODING mime" before the "FILE"
	command.  To decode the response(s), you will need "munpack" or
	a MIME-compliant mail reader.  Different MIME-compliant mail readers
	exhibit different behavior, especially when dealing with
	"multipart" MIME messages (i.e. documents which have been split
	up into multiple messages), so check your local documentation on
	how to manipulate these messages.

Below is the data which will enable a MIME compliant mail reader
implementation to automatically retrieve the ASCII version of the
Internet-Draft.

--NextPart
Content-Type: Multipart/Alternative; Boundary="OtherAccess"

--OtherAccess
Content-Type: Message/External-body; access-type="mail-server";
	server="mailserv@ietf.org"

Content-Type: text/plain
Content-ID: <2007-2-23165529.I-D@ietf.org>

ENCODING mime
FILE /internet-drafts/draft-klensin-unicode-escapes-03.txt

--OtherAccess
Content-Type: Message/External-body;
	name="draft-klensin-unicode-escapes-03.txt"; site="ftp.ietf.org";
	access-type="anon-ftp"; directory="internet-drafts"

Content-Type: text/plain
Content-ID: <2007-2-23165529.I-D@ietf.org>


--OtherAccess--

--NextPart
Content-Type: text/plain; charset="us-ascii"
MIME-Version: 1.0
Content-Transfer-Encoding: 7bit
Content-Disposition: inline

_______________________________________________
I-D-Announce mailing list
I-D-Announce@ietf.org
https://www1.ietf.org/mailman/listinfo/i-d-announce

--NextPart--



--==========A6B7712B9747A00683EF==========--





From discuss-bounces@apps.ietf.org Sat Feb 24 10:58:46 2007
Received: from [127.0.0.1] (helo=stiedprmman1.va.neustar.com)
	by megatron.ietf.org with esmtp (Exim 4.43)
	id 1HKzGi-0006q0-WC; Sat, 24 Feb 2007 10:57:05 -0500
Received: from [10.91.34.44] (helo=ietf-mx.ietf.org)
	by megatron.ietf.org with esmtp (Exim 4.43) id 1HKzGi-0006kj-1H
	for discuss@apps.ietf.org; Sat, 24 Feb 2007 10:57:04 -0500
Received: from main.gmane.org ([80.91.229.2] helo=ciao.gmane.org)
	by ietf-mx.ietf.org with esmtp (Exim 4.43) id 1HKzGf-0002hQ-NO
	for discuss@apps.ietf.org; Sat, 24 Feb 2007 10:57:04 -0500
Received: from list by ciao.gmane.org with local (Exim 4.43)
	id 1HKzGZ-00049q-NZ
	for discuss@apps.ietf.org; Sat, 24 Feb 2007 16:56:55 +0100
Received: from 212.82.251.32 ([212.82.251.32])
	by main.gmane.org with esmtp (Gmexim 0.1 (Debian))
	id 1AlnuQ-0007hv-00
	for <discuss@apps.ietf.org>; Sat, 24 Feb 2007 16:56:55 +0100
Received: from nobody by 212.82.251.32 with local (Gmexim 0.1 (Debian))
	id 1AlnuQ-0007hv-00
	for <discuss@apps.ietf.org>; Sat, 24 Feb 2007 16:56:55 +0100
X-Injected-Via-Gmane: http://gmane.org/
To: discuss@apps.ietf.org
From: Frank Ellermann <nobody@xyzzy.claranet.de>
Subject: Re: draft-klensin-unicode-escapes-03.txt
Date: Sat, 24 Feb 2007 16:56:03 +0100
Organization: <URL:http://purl.net/xyzzy>
Lines: 37
Message-ID: <45E06013.4BCA@xyzzy.claranet.de>
References: <74711BCF624DBEC4F2C000C5@p3.JCK.COM>
	<45DA17F4.4857@xyzzy.claranet.de> <754B21F623BA398D14CCDE97@p3.JCK.COM>
Mime-Version: 1.0
Content-Type: text/plain; charset=us-ascii
Content-Transfer-Encoding: 7bit
X-Complaints-To: usenet@sea.gmane.org
X-Gmane-NNTP-Posting-Host: 212.82.251.32
X-Mailer: Mozilla 3.0 (OS/2; U)
X-Spam-Score: 1.6 (+)
X-Scan-Signature: 69a74e02bbee44ab4f8eafdbcedd94a1
X-BeenThere: discuss@apps.ietf.org
X-Mailman-Version: 2.1.5
Precedence: list
List-Id: general discussion of application-layer protocols
	<discuss.apps.ietf.org>
List-Unsubscribe: <https://www1.ietf.org/mailman/listinfo/discuss>,
	<mailto:discuss-request@apps.ietf.org?subject=unsubscribe>
List-Post: <mailto:discuss@apps.ietf.org>
List-Help: <mailto:discuss-request@apps.ietf.org?subject=help>
List-Subscribe: <https://www1.ietf.org/mailman/listinfo/discuss>,
	<mailto:discuss-request@apps.ietf.org?subject=subscribe>
Errors-To: discuss-bounces@apps.ietf.org

John C Klensin wrote:

>> Here's what I'd do with the existing ABNF:

>>    EmbeddedUnicodeChar =  BMP-form / Full-form
>>    BMP-form  =  %x5C.75 4HEXDIG   ; starting with lower case "\u"
>>    Full-form =  %x5C.55 8HEXDIG   ; starting with upper case "\U"
[...]
> Done, including removal of <HexDigit>

You kept the multiline comment after <BMP-form>, and Bill's parser
doesn't like this:

BMP-form =  %x5C.75 4HEXDIG ; starting with lower case "\u"
   ; In both this case and the one above, note that the encodings are
   considered to be abstractions for the relevant characters, not
   designations of specific octets.

With more semicolons it's okay:

BMP-form =  %x5C.75 4HEXDIG ; starting with lower case "\u"
   ; In both this case and the one above, note that the encodings are
   ; considered to be abstractions for the relevant characters, not
   ; designations of specific octets.

But actually I think this ABNF comment is unnecessary - or if not it
could be moved into the prose.  Then the contrast \u vs. \U would be
also more directly visible:

BMP-form  =  %x5C.75 4HEXDIG   ; starting with lower case "\u"
Full-form =  %x5C.55 8HEXDIG   ; starting with upper case "\U"

Lillyguilding => ready, maybe add "intended status: BCP" and set the
"PubReq" flag while authors are still allowed to manage this flag. ;-)

Frank






From discuss-bounces@apps.ietf.org Wed Feb 28 00:05:23 2007
Received: from [127.0.0.1] (helo=stiedprmman1.va.neustar.com)
	by megatron.ietf.org with esmtp (Exim 4.43)
	id 1HMGyW-0000EL-U5; Wed, 28 Feb 2007 00:03:36 -0500
Received: from [10.91.34.44] (helo=ietf-mx.ietf.org)
	by megatron.ietf.org with esmtp (Exim 4.43) id 1HMGyW-0000EC-2a
	for discuss@apps.ietf.org; Wed, 28 Feb 2007 00:03:36 -0500
Received: from send01.jprs.co.jp ([202.11.17.113])
	by ietf-mx.ietf.org with esmtp (Exim 4.43) id 1HMGyU-0004Oc-CZ
	for discuss@apps.ietf.org; Wed, 28 Feb 2007 00:03:36 -0500
Received: from send01.jprs.co.jp (localhost [127.0.0.1])
	by send01.jprs.co.jp (8.12.10+Sun/8.12.11) with SMTP id l1S3p5R9016718
	for <discuss@apps.ietf.org>; Wed, 28 Feb 2007 12:51:06 +0900 (JST)
Received: (from NOTE233 [172.18.4.45])
	by send01.jprs.co.jp (SMSSMTP 4.0.4.64) with SMTP id
	M2007022812510502187
	for <discuss@apps.ietf.org>; Wed, 28 Feb 2007 12:51:05 +0900
Date: Wed, 28 Feb 2007 12:51:04 +0900
From: Yoshiro YONEYA <yone@jprs.co.jp>
To: discuss@apps.ietf.org
Subject: Fw: I-D ACTION:draft-yoneya-iri-recognition-00.txt
Message-Id: <20070228125104.ab7c9ae3.yone@jprs.co.jp>
X-Mailer: Sylpheed 2.3.1 (GTK+ 2.10.7; i686-pc-mingw32)
Mime-Version: 1.0
Content-Type: multipart/mixed;
	boundary="Multipart=_Wed__28_Feb_2007_12_51_04_+0900_og7m8JwAE/K6PDNU"
X-Spam-Score: 0.0 (/)
X-Scan-Signature: 43317e64100dd4d87214c51822b582d1
X-BeenThere: discuss@apps.ietf.org
X-Mailman-Version: 2.1.5
Precedence: list
List-Id: general discussion of application-layer protocols
	<discuss.apps.ietf.org>
List-Unsubscribe: <https://www1.ietf.org/mailman/listinfo/discuss>,
	<mailto:discuss-request@apps.ietf.org?subject=unsubscribe>
List-Post: <mailto:discuss@apps.ietf.org>
List-Help: <mailto:discuss-request@apps.ietf.org?subject=help>
List-Subscribe: <https://www1.ietf.org/mailman/listinfo/discuss>,
	<mailto:discuss-request@apps.ietf.org?subject=subscribe>
Errors-To: discuss-bounces@apps.ietf.org

This is a multi-part message in MIME format.

--Multipart=_Wed__28_Feb_2007_12_51_04_+0900_og7m8JwAE/K6PDNU
Content-Type: text/plain; charset=US-ASCII
Content-Transfer-Encoding: 7bit

Hi, all.

I wrote an I-D which discusses how IRIs should be recognized and how to 
be handled.  The background is that many of applications can recognize 
URIs in unstructured text data, but IRIs are exception although use of 
IDNs/IRIs is getting popular.

If you have interested in, please take a look and give your comments.  
Also, I'd like to have time slot at APPAREA BoF in Prague to explain 
the I-D.

Thanks,

-- 
Yoshiro YONEYA <yone@jprs.co.jp>

--Multipart=_Wed__28_Feb_2007_12_51_04_+0900_og7m8JwAE/K6PDNU
Content-Type: message/rfc822
Content-Disposition: inline
Content-Transfer-Encoding: 7bit

Return-Path: <i-d-announce-bounces@ietf.org>
Received: from msgmgr01.jprs.co.jp (msgmgr01.jprs.co.jp [172.18.8.17])
	by spool01.jprs.co.jp (8.12.10+Sun/8.12.11) with ESMTP id
	l1RNG6EX000366
	for <yone@spool01.jprs.co.jp>; Wed, 28 Feb 2007 08:16:06 +0900 (JST)
Received: from mx01.jprs.co.jp (mx01.jprs.co.jp [202.11.17.111])
	by msgmgr01.jprs.co.jp (8.12.10+Sun/8.12.11) with SMTP id
	l1RNFwUS014889
	for <yone@jprs.co.jp>; Wed, 28 Feb 2007 08:16:06 +0900 (JST)
X-Brightmail-Tracker: AAAAAQAAA+k=
X-Language-Identified: TRUE
Received: from mx0.jprs.co.jp ([127.0.0.1])
	by mx01.jprs.co.jp (SMSSMTP 4.1.12.43) with SMTP id
	M2007022808160411385
	for <yone@jprs.co.jp>; Wed, 28 Feb 2007 08:16:04 +0900
Received: from megatron.ietf.org (odin.ietf.ORG [156.154.16.145])
	by mx0.jprs.co.jp (8.13.6/8.13.6) with ESMTP id l1RNG465003203
	for <yone@jprs.co.jp>; Wed, 28 Feb 2007 08:16:05 +0900 (JST)
Authentication-Results: mx0.jprs.co.jp from=Internet-Drafts@ietf.org;
	sender-id=neutral; spf=neutral
Received: from [127.0.0.1] (helo=stiedprmman1.va.neustar.com)
	by megatron.ietf.org with esmtp (Exim 4.43)
	id 1HM9IV-0002pX-EO; Tue, 27 Feb 2007 15:51:43 -0500
Received: from [10.90.34.44] (helo=chiedprmail1.ietf.org)
	by megatron.ietf.org with esmtp (Exim 4.43) id 1HM9Ht-0001zn-BS
	for i-d-announce@ietf.org; Tue, 27 Feb 2007 15:51:05 -0500
Received: from ns3.neustar.com ([156.154.24.138])
	by chiedprmail1.ietf.org with esmtp (Exim 4.43) id 1HM9Hs-0000U5-RO
	for i-d-announce@ietf.org; Tue, 27 Feb 2007 15:51:05 -0500
Received: from stiedprstage1.ietf.org (stiedprstage1.va.neustar.com
	[10.31.47.10]) by ns3.neustar.com (Postfix) with ESMTP id 3BF2F1767C
	for <i-d-announce@ietf.org>; Tue, 27 Feb 2007 20:50:03 +0000 (GMT)
Received: from ietf by stiedprstage1.ietf.org with local (Exim 4.43)
	id 1HM9Gs-0001ep-Vp
	for i-d-announce@ietf.org; Tue, 27 Feb 2007 15:50:02 -0500
Content-Type: Multipart/Mixed; Boundary="NextPart"
Mime-Version: 1.0
To: i-d-announce@ietf.org
Cc: 
From: Internet-Drafts@ietf.org
Message-Id: <E1HM9Gs-0001ep-Vp@stiedprstage1.ietf.org>
Date: Tue, 27 Feb 2007 15:50:02 -0500
X-Spam-Score: -2.5 (--)
X-Scan-Signature: a2c12dacc0736f14d6b540e805505a86
Subject: I-D ACTION:draft-yoneya-iri-recognition-00.txt 
X-BeenThere: i-d-announce@ietf.org
X-Mailman-Version: 2.1.5
Precedence: list
Reply-To: internet-drafts@ietf.org
List-Id: i-d-announce.ietf.org
List-Unsubscribe: <https://www1.ietf.org/mailman/listinfo/i-d-announce>,
	<mailto:i-d-announce-request@ietf.org?subject=unsubscribe>
List-Archive: <http://www1.ietf.org/pipermail/i-d-announce>
List-Post: <mailto:i-d-announce@ietf.org>
List-Help: <mailto:i-d-announce-request@ietf.org?subject=help>
List-Subscribe: <https://www1.ietf.org/mailman/listinfo/i-d-announce>,
	<mailto:i-d-announce-request@ietf.org?subject=subscribe>
Errors-To: i-d-announce-bounces@ietf.org
X-UIDL: C`R"!VL-!!,_R!!2$D!!
Status: RO

--NextPart

A New Internet-Draft is available from the on-line Internet-Drafts 
directories.


	Title		: IRI recognition in Applications
	Author(s)	: Y. Yoneya
	Filename	: draft-yoneya-iri-recognition-00.txt
	Pages		: 7
	Date		: 2007-2-27
	
   Nowadays access to the Internet is a part of daily life.  Users see
   URIs written in various ways on various media, recognize them as "the
   Internet Addresses", and use them to access to the Internet.  Many
   application programs recognize URIs automatically and make links to
   them, so the users can access to the URIs very easily.  But, at this
   moment, most of application programs can't recognize
   Internationalized Domain Name (IDN) and Internationalized URI (IRI)
   correctly, so users will feel stress when using IDNs and IRIs.

   Utilization of the IDNs and the IRIs are getting higher.  Therefore,
   improvement of the application programs are highly recommended.  This
   document is intended to be an application developpers' guideline for
   recognizing and corresponding to IDNs/IRIs correctly.


A URL for this Internet-Draft is:
http://www.ietf.org/internet-drafts/draft-yoneya-iri-recognition-00.txt

To remove yourself from the I-D Announcement list, send a message to 
i-d-announce-request@ietf.org with the word unsubscribe in the body of 
the message. 
You can also visit https://www1.ietf.org/mailman/listinfo/I-D-announce 
to change your subscription settings.

Internet-Drafts are also available by anonymous FTP. Login with the 
username "anonymous" and a password of your e-mail address. After 
logging in, type "cd internet-drafts" and then 
"get draft-yoneya-iri-recognition-00.txt".

A list of Internet-Drafts directories can be found in
http://www.ietf.org/shadow.html 
or ftp://ftp.ietf.org/ietf/1shadow-sites.txt

Internet-Drafts can also be obtained by e-mail.

Send a message to:
	mailserv@ietf.org.
In the body type:
	"FILE /internet-drafts/draft-yoneya-iri-recognition-00.txt".
	
NOTE:	The mail server at ietf.org can return the document in
	MIME-encoded form by using the "mpack" utility.  To use this
	feature, insert the command "ENCODING mime" before the "FILE"
	command.  To decode the response(s), you will need "munpack" or
	a MIME-compliant mail reader.  Different MIME-compliant mail readers
	exhibit different behavior, especially when dealing with
	"multipart" MIME messages (i.e. documents which have been split
	up into multiple messages), so check your local documentation on
	how to manipulate these messages.

Below is the data which will enable a MIME compliant mail reader
implementation to automatically retrieve the ASCII version of the
Internet-Draft.

--NextPart
Content-Type: Multipart/Alternative; Boundary="OtherAccess"

--OtherAccess
Content-Type: Message/External-body; access-type="mail-server";
	server="mailserv@ietf.org"

Content-Type: text/plain
Content-ID: <2007-2-27135834.I-D@ietf.org>

ENCODING mime
FILE /internet-drafts/draft-yoneya-iri-recognition-00.txt

--OtherAccess
Content-Type: Message/External-body;
	name="draft-yoneya-iri-recognition-00.txt"; site="ftp.ietf.org";
	access-type="anon-ftp"; directory="internet-drafts"

Content-Type: text/plain
Content-ID: <2007-2-27135834.I-D@ietf.org>


--OtherAccess--

--NextPart
Content-Type: text/plain; charset="us-ascii"
MIME-Version: 1.0
Content-Transfer-Encoding: 7bit
Content-Disposition: inline

_______________________________________________
I-D-Announce mailing list
I-D-Announce@ietf.org
https://www1.ietf.org/mailman/listinfo/i-d-announce

--NextPart--




--Multipart=_Wed__28_Feb_2007_12_51_04_+0900_og7m8JwAE/K6PDNU--





