From owner-idn@ops.ietf.org  Tue Apr  1 11:18:20 2003
Received: from psg.com (mailnull@psg.com [147.28.0.62])
	by ietf.org (8.9.1a/8.9.1a) with ESMTP id LAA25682
	for <idn-archive@lists.ietf.org>; Tue, 1 Apr 2003 11:18:20 -0500 (EST)
Received: from lserv by psg.com with local (Exim 3.36 #1)
	id 190OKD-000ATy-00
	for idn-data@psg.com; Tue, 01 Apr 2003 08:09:25 -0800
Received: from w1309.hostcentric.net ([66.40.78.254])
	by psg.com with smtp (Exim 3.36 #1)
	id 190OK6-000ATa-00
	for idn@ops.ietf.org; Tue, 01 Apr 2003 08:09:18 -0800
Received: (qmail 8743 invoked by alias); 1 Apr 2003 16:09:11 -0000
Received: from unknown (HELO DAVIS1) (12.234.231.178)
  by 0 with SMTP; 1 Apr 2003 16:09:11 -0000
Message-ID: <004d01c2f868$fbfcc850$7900a8c0@DAVIS1>
From: "Mark Davis" <mark.davis@jtcsv.com>
To: <idn@ops.ietf.org>
References: <0057BC9383E9AA4ABC1749790A6E201A0118A59E@PSG-MAIL.psglaw.dk>
Subject: [idn] Error in IDN eamples for testing
Date: Tue, 1 Apr 2003 08:08:50 -0800
MIME-Version: 1.0
Content-Type: text/plain;
	charset="iso-8859-1"
Content-Transfer-Encoding: 8bit
X-Priority: 3
X-MSMail-Priority: Normal
X-Mailer: Microsoft Outlook Express 6.00.2800.1106
X-MimeOLE: Produced By Microsoft MimeOLE V6.00.2800.1106
X-Spam-Status: No, hits=-19.5 required=5.0
	tests=BAYES_10,ORIGINAL_MESSAGE,QUOTED_EMAIL_TEXT,REFERENCES,
	      WEIRD_QUOTING
	autolearn=ham	version=2.50
X-Spam-Checker-Version: SpamAssassin 2.50 (1.173-2003-02-20-exp)
Sender: owner-idn@ops.ietf.org
Precedence: bulk
Content-Transfer-Encoding: 8bit

One of our engineers took a look at the IDN samples for testing, and found a
problem in the data. Apparently some of the UTF-8 is illegal (using 2
three-byte sequences for a supplementary character).

Mark

========

Hi Mark,
The following are the offending members of structure in Appendix A of
draft-josefsson-idn-test-vectors. The data claiming to be UTF-8 is actually
CESU-8.

struct stringprep
{
char *comment;
char *in;
char *out;
char *profile;
int flags;
int rc;
}
strprep[] =
{
..........
{
"Surrogate code U+DF42",
"\xED\xBD\x82", NULL, "Nameprep", 0,
STRINGPREP_CONTAINS_PROHIBITED
},
...........
{
"Larger test (shrinking)",
"X\xC2\xAD\xC3\xDF\xC4\xB0\xE2\x84\xA1\x6a\xcc\x8c\xc2\xa0\xc2"
"\xaa\xce\xb0\xe2\x80\x80", "xssi\xcc\x87""tel\xc7\xb0 a\xce\xb0 ",
"Nameprep"
},
{
"Larger test (expanding)",
"X\xC3\xDF\xe3\x8c\x96\xC4\xB0\xE2\x84\xA1\xE2\x92\x9F\xE3\x8c\x80",
"xss\xe3\x82\xad\xe3\x83\xad\xe3\x83\xa1\xe3\x83\xbc\xe3\x83\x88"
"\xe3\x83\xab""i\xcc\x87""tel\x28""d\x29\xe3\x82\xa2\xe3\x83\x91"
"\xe3\x83\xbc\xe3\x83\x88"
},
}
...

----- Original Message -----
From: "Peter Gustav Olson - pgo" <pgo@psglaw.dk>
To: "vinton g. cerf" <vinton.g.cerf@wcom.com>; "Paul Hoffman / IMC"
<phoffman@imc.org>; <idn@ops.ietf.org>
Sent: Tuesday, March 25, 2003 05:37
Subject: RE: [idn] IDN eamples for testing




> -----Original Message-----
> From: vinton g. cerf [mailto:vinton.g.cerf@wcom.com]
> Sent: Tuesday, March 25, 2003 1:07 PM
> To: Peter Gustav Olson - pgo; Paul Hoffman / IMC; idn@ops.ietf.org
> Subject: RE: [idn] IDN eamples for testing
>
>
> Peter,
>
> thanks for pointing out the cybersquatter problem - I
> certainly agree that we don't want to adopt policies that
> will encourage or support cybersquatting!


For companies like Ben & Jerry's, Bang & Olufsen, Nestlé, L'Oreal, Kookaï,
etc. the IDN will be much more significant than the introduction of .biz and
.info. Let's make it significant in a positive sense, rather than a
significant new anti-cybersquatting expenditure.

Yours Sincerely,

Peter Gustav Olson

PLESNER SVANE GRØNBORG
Law firm

34, Esplanaden
DK-1263  Copenhagen K.

Tel: +45 33 12 11 33
Fax: +45 33 12 00 14
E-mail: pgo@psglaw.dk
Website: www.psglaw.dk


This e-mail and any attachment are confidential and may also be privileged.
If you are not the intended recipient, please notify us immediately and then
delete this e-mail and any attachment without retaining copies or disclosing
the contents thereof to any other person.  Thank you.







From owner-idn@ops.ietf.org  Tue Apr  1 11:18:48 2003
Received: from psg.com (mailnull@psg.com [147.28.0.62])
	by ietf.org (8.9.1a/8.9.1a) with ESMTP id LAA25701
	for <idn-archive@lists.ietf.org>; Tue, 1 Apr 2003 11:18:48 -0500 (EST)
Received: from lserv by psg.com with local (Exim 3.36 #1)
	id 190OJT-000ASK-00
	for idn-data@psg.com; Tue, 01 Apr 2003 08:08:39 -0800
Received: from tux.w3.org ([18.29.0.27])
	by psg.com with esmtp (Exim 3.36 #1)
	id 190OJQ-000AS8-00
	for idn@ops.ietf.org; Tue, 01 Apr 2003 08:08:36 -0800
Received: from enoshima (IDENT:root@tux.w3.org [18.29.0.27])
	by tux.w3.org (8.12.8/8.12.8) with ESMTP id h31G8VrC026868;
	Tue, 1 Apr 2003 11:08:34 -0500
Message-Id: <4.2.0.58.J.20030331172631.028c7f18@localhost>
X-Sender: duerst@localhost
X-Mailer: QUALCOMM Windows Eudora Pro Version 4.2.0.58.J 
Date: Mon, 31 Mar 2003 17:32:54 -0500
To: Paul Hoffman / IMC <phoffman@imc.org>, idn@ops.ietf.org
From: Martin Duerst <duerst@w3.org>
Subject: Re: [idn] Fwd: I-D
  ACTION:draft-josefsson-idn-test-vectors-00.txt
In-Reply-To: <p05210605baaa1f8218bb@[63.202.92.152]>
Mime-Version: 1.0
Content-Type: text/plain; charset="us-ascii"; format=flowed
X-Spam-Status: No, hits=-25.9 required=5.0
	tests=BAYES_01,CLICK_BELOW,DATE_IN_PAST_12_24,EMAIL_ATTRIBUTION,
	      IN_REP_TO,QUOTED_EMAIL_TEXT,REPLY_WITH_QUOTES
	autolearn=ham	version=2.50
X-Spam-Checker-Version: SpamAssassin 2.50 (1.173-2003-02-20-exp)
Sender: owner-idn@ops.ietf.org
Precedence: bulk

This is clearly extremely valuable. But it also (once again)
shows the limitation of the current Internet-Draft format.

Are the inputs of this test suite available as actual UTF-8,
e.g. as one input per line, or one input line followed by
one output line? I don't want to write a C program and
run the C compiler and execute a program just to be able
to use the data for my purposes.

Regards,    Martin.

At 07:55 03/03/28 -0800, Paul Hoffman / IMC wrote:
>There is a new draft that is of very direct interest to this mailing list. 
>It would be great if people who have implemented IDNA could check their 
>results against this draft. And thank you to Simon for doing the work on 
>this! These kinds of documents are very useful in the IETF and can help 
>IDNA progress on standards track.
>
>--Paul Hoffman
>
>>To: IETF-Announce: ;
>>From: Internet-Drafts@ietf.org
>>Reply-to: Internet-Drafts@ietf.org
>>Subject: I-D ACTION:draft-josefsson-idn-test-vectors-00.txt
>>Date: Fri, 28 Mar 2003 07:22:21 -0500
>>Sender: owner-ietf-announce@ietf.org
>>
>>
>>
>>A New Internet-Draft is available from the on-line Internet-Drafts 
>>directories.
>>
>>
>>         Title           : Nameprep and IDNA Test Vectors
>>         Author(s)       : S. Josefsson
>>         Filename        : draft-josefsson-idn-test-vectors-00.txt
>>         Pages           : 41
>>         Date            : 2003-3-27
>>
>>This document contains test vectors for Nameprep and IDNA
>>
>>A URL for this Internet-Draft is:
>>http://www.ietf.org/internet-drafts/draft-josefsson-idn-test-vectors-00.txt
>>
>>To remove yourself from the IETF Announcement list, send a message to
>>ietf-announce-request with the word unsubscribe in the body of the message.
>>
>>Internet-Drafts are also available by anonymous FTP. Login with the username
>>"anonymous" and a password of your e-mail address. After logging in,
>>type "cd internet-drafts" and then
>>         "get draft-josefsson-idn-test-vectors-00.txt".
>>
>>A list of Internet-Drafts directories can be found in
>>http://www.ietf.org/shadow.html
>>or ftp://ftp.ietf.org/ietf/1shadow-sites.txt
>>
>>
>>Internet-Drafts can also be obtained by e-mail.
>>
>>Send a message to:
>>         mailserv@ietf.org.
>>In the body type:
>>         "FILE /internet-drafts/draft-josefsson-idn-test-vectors-00.txt".
>>
>>NOTE:   The mail server at ietf.org can return the document in
>>         MIME-encoded form by using the "mpack" utility.  To use this
>>         feature, insert the command "ENCODING mime" before the "FILE"
>>         command.  To decode the response(s), you will need "munpack" or
>>         a MIME-compliant mail reader.  Different MIME-compliant mail readers
>>         exhibit different behavior, especially when dealing with
>>         "multipart" MIME messages (i.e. documents which have been split
>>         up into multiple messages), so check your local documentation on
>>         how to manipulate these messages.
>>
>>
>>Below is the data which will enable a MIME compliant mail reader
>>implementation to automatically retrieve the ASCII version of the
>>Internet-Draft.
>>
>>
>>[The following attachment must be fetched by mail. Command-click the URL 
>>below and send the resulting message to get the attachment.]
>><mailto:mailserv@ietf.org?body=ENCODING%20mime%0D%0AFILE%20/internet-draft 
>>s/draft-josefsson-idn-test-vectors-00.txt>
>>[The following attachment must be fetched by ftp.  Command-click the URL 
>>below to ask your ftp client to fetch it.]
>><ftp://ftp.ietf.org/internet-drafts/draft-josefsson-idn-test-vectors-00.txt>




From owner-idn@ops.ietf.org  Mon Apr  7 15:21:46 2003
Received: from psg.com (mailnull@psg.com [147.28.0.62])
	by ietf.org (8.9.1a/8.9.1a) with ESMTP id PAA02852
	for <idn-archive@lists.ietf.org>; Mon, 7 Apr 2003 15:21:46 -0400 (EDT)
Received: from lserv by psg.com with local (Exim 3.36 #1)
	id 192byU-0000m4-00
	for idn-data@psg.com; Mon, 07 Apr 2003 19:08:10 +0000
Received: from tux.w3.org ([18.29.0.27])
	by psg.com with esmtp (Exim 3.36 #1)
	id 192byH-0000kN-00
	for idn@ops.ietf.org; Mon, 07 Apr 2003 12:07:57 -0700
Received: from enoshima (IDENT:root@tux.w3.org [18.29.0.27])
	by tux.w3.org (8.12.9/8.12.9) with ESMTP id h37J7tsj006934
	for <idn@ops.ietf.org>; Mon, 7 Apr 2003 15:07:55 -0400
Message-Id: <4.2.0.58.J.20030331173417.027fb080@localhost>
X-Sender: duerst@localhost
X-Mailer: QUALCOMM Windows Eudora Pro Version 4.2.0.58.J 
Date: Mon, 07 Apr 2003 15:07:49 -0400
To: idn@ops.ietf.org
From: Martin Duerst <duerst@w3.org>
Subject: [idn] Challenge: longest UTF-8 with valid domain name
Mime-Version: 1.0
Content-Type: text/plain; charset="us-ascii"; format=flowed
X-Spam-Status: No, hits=-6.6 required=5.0
	tests=BAYES_01
	version=2.50
X-Spam-Checker-Version: SpamAssassin 2.50 (1.173-2003-02-20-exp)
Sender: owner-idn@ops.ietf.org
Precedence: bulk

Here is a little challenge that might end up in Simon's test suite
and be otherwise valuable for testing:

What is the longest (in terms of bytes) internationalized
domain name (when encoded as UTF-8)? Obviously because there
are characters that are ignored by nameprep, we would have to
ask for the longest one after ToUnicode.

The problem applies both to single labels as well as to
FQDNs.

Regards,    Martin.



From owner-idn@ops.ietf.org  Mon Apr  7 17:52:12 2003
Received: from psg.com (mailnull@psg.com [147.28.0.62])
	by ietf.org (8.9.1a/8.9.1a) with ESMTP id RAA06977
	for <idn-archive@lists.ietf.org>; Mon, 7 Apr 2003 17:52:12 -0400 (EDT)
Received: from lserv by psg.com with local (Exim 3.36 #1)
	id 192eQk-000Dqy-00
	for idn-data@psg.com; Mon, 07 Apr 2003 21:45:30 +0000
Received: from arwen.cs.berkeley.edu ([128.32.132.165] helo=nicemice.net)
	by psg.com with esmtp (Exim 3.36 #1)
	id 192eNY-000DOR-00
	for idn@ops.ietf.org; Mon, 07 Apr 2003 14:42:12 -0700
Received: from amc by nicemice.net with local (Exim 3.35 #1 (Debian))
	id 192eNX-0002Cn-00
	for <idn@ops.ietf.org>; Mon, 07 Apr 2003 14:42:11 -0700
Date: Mon, 7 Apr 2003 21:42:11 +0000
From: "Adam M. Costello" <idn.amc+0@nicemice.net.RemoveThisWord>
To: idn@ops.ietf.org
Subject: Re: [idn] Challenge: longest UTF-8 with valid domain name
Message-ID: <20030407214211.GE7147@nicemice.net>
Reply-To: IETF idn working group <idn@ops.ietf.org>
References: <4.2.0.58.J.20030331173417.027fb080@localhost>
Mime-Version: 1.0
Content-Type: text/plain; charset=us-ascii
Content-Disposition: inline
In-Reply-To: <4.2.0.58.J.20030331173417.027fb080@localhost>
User-Agent: Mutt/1.4i
X-Spam-Status: No, hits=-38.0 required=5.0
	tests=BAYES_10,EMAIL_ATTRIBUTION,IN_REP_TO,QUOTED_EMAIL_TEXT,
	      REFERENCES,REPLY_WITH_QUOTES,USER_AGENT_MUTT
	autolearn=ham	version=2.50
X-Spam-Checker-Version: SpamAssassin 2.50 (1.173-2003-02-20-exp)
Sender: owner-idn@ops.ietf.org
Precedence: bulk

Martin Duerst <duerst@w3.org> wrote:

> What is the longest (in terms of bytes) internationalized domain name
> (when encoded as UTF-8)?
>
> Obviously because there are characters that are ignored by nameprep,
> we would have to ask for the longest one after ToUnicode.

You mean the longest after nameprep.  The output of ToUnicode is not
necessarily nameprepped (because ToUnicode leaves its input untouched if
the input is not an ACE).

> The problem applies both to single labels as well as to FQDNs.

I think the longest nameprepped label is 224 bytes in UTF-8.  Taken any
code point in the range 10000..55931 (hex) and repeat it 56 times.  The
ACE form will be 63 characters.  The nameprepped non-ACE form in UTF-8
will be 56*4 = 224 bytes.

As for the longest IDN, I think that would have four labels, three
of which are maximal, and one of which is just shy of maximal (62
characters in the ACE).  With the four length bytes, that hits the DNS
limit of 255 bytes.

The longest UTF-8 representation of that name (with nameprepped labels)
would use ideographic full stops for dots, and would include the
trailing dot.  So it would be (3*56+55)*4 + 4*3 = 904 bytes.

AMC



From owner-idn@ops.ietf.org  Mon Apr  7 18:43:16 2003
Received: from psg.com (mailnull@psg.com [147.28.0.62])
	by ietf.org (8.9.1a/8.9.1a) with ESMTP id SAA09652
	for <idn-archive@lists.ietf.org>; Mon, 7 Apr 2003 18:43:16 -0400 (EDT)
Received: from lserv by psg.com with local (Exim 3.36 #1)
	id 192fFQ-000Jn9-00
	for idn-data@psg.com; Mon, 07 Apr 2003 22:37:52 +0000
Received: from [66.96.233.180] (helo=amanda.ez-web-hosting.com)
	by psg.com with esmtp (Exim 3.36 #1)
	id 192fFK-000Jm6-00
	for idn@ops.ietf.org; Mon, 07 Apr 2003 15:37:46 -0700
Received: from ip68-14-188-45.pn.at.cox.net ([68.14.188.45] helo=0202012)
	by amanda.ez-web-hosting.com with asmtp (Exim 3.36 #1)
	id 192fEU-0006GN-00
	for idn@ops.ietf.org; Mon, 07 Apr 2003 18:36:54 -0400
Reply-To: <jim@mathies.com>
From: <jim@mathies.com>
To: "'IETF idn working group'" <idn@ops.ietf.org>
Subject: [idn] Question about a ToUnicode step
Date: Mon, 7 Apr 2003 17:38:37 -0500
Message-ID: <000c01c2fd56$6fca7830$2dbc0e44@0202012>
MIME-Version: 1.0
Content-Type: text/plain;
	charset="US-ASCII"
Content-Transfer-Encoding: 7bit
X-Priority: 3 (Normal)
X-MSMail-Priority: Normal
X-Mailer: Microsoft Outlook, Build 10.0.4510
In-Reply-To: <20030407214211.GE7147@nicemice.net>
Importance: Normal
X-MimeOLE: Produced By Microsoft MimeOLE V6.00.2800.1106
X-AntiAbuse: This header was added to track abuse, please include it with any abuse report
X-AntiAbuse: Primary Hostname - amanda.ez-web-hosting.com
X-AntiAbuse: Original Domain - ops.ietf.org
X-AntiAbuse: Originator/Caller UID/GID - [0 0] / [0 0]
X-AntiAbuse: Sender Address Domain - mathies.com
X-Spam-Status: No, hits=-8.0 required=5.0
	tests=BAYES_10,IN_REP_TO,NO_REAL_NAME
	autolearn=ham	version=2.50
X-Spam-Checker-Version: SpamAssassin 2.50 (1.173-2003-02-20-exp)
Sender: owner-idn@ops.ietf.org
Precedence: bulk
Content-Transfer-Encoding: 7bit


Hello,

I'm trying to understand the purpose behind some steps in 
ToUnicode. From the IDNA spec:


1. If all code points in the sequence are in the ASCII range (0..7F)
   then skip to step 3.

2. Perform the steps specified in [NAMEPREP] and fail if there is an
   error. (If step 3 of ToAscii is also performed here, it will not
   affect the overall behavior of ToUnicode, but it is not
   necessary.) The AllowUnassigned flag is used in [NAMEPREP].

3. Verify that the sequence begins with the ACE prefix, and save a
   copy of the sequence.


I'm curious about steps 1 & 2. I don't understand why nameprep
is being applied to ASCII domain labels. This seems to be pointless 
in situations where the input string is 8-bit ASCII. Wouldn't
a simple dns character compatibility check in place of steps 1 
and 2 suffice?

Regards,
Jim Mathies




From owner-idn@ops.ietf.org  Mon Apr  7 19:11:39 2003
Received: from psg.com (mailnull@psg.com [147.28.0.62])
	by ietf.org (8.9.1a/8.9.1a) with ESMTP id TAA10365
	for <idn-archive@lists.ietf.org>; Mon, 7 Apr 2003 19:11:38 -0400 (EDT)
Received: from lserv by psg.com with local (Exim 3.36 #1)
	id 192fhI-000NCh-00
	for idn-data@psg.com; Mon, 07 Apr 2003 23:06:40 +0000
Received: from arwen.cs.berkeley.edu ([128.32.132.165] helo=nicemice.net)
	by psg.com with esmtp (Exim 3.36 #1)
	id 192fhF-000NCV-00
	for idn@ops.ietf.org; Mon, 07 Apr 2003 16:06:37 -0700
Received: from amc by nicemice.net with local (Exim 3.35 #1 (Debian))
	id 192fhE-0002OV-00
	for <idn@ops.ietf.org>; Mon, 07 Apr 2003 16:06:36 -0700
Date: Mon, 7 Apr 2003 23:06:36 +0000
From: "Adam M. Costello" <idn.amc+0@nicemice.net.RemoveThisWord>
To: "'IETF idn working group'" <idn@ops.ietf.org>
Subject: Re: [idn] Question about a ToUnicode step
Message-ID: <20030407230636.GF7147@nicemice.net>
Reply-To: IETF idn working group <idn@ops.ietf.org>
References: <20030407214211.GE7147@nicemice.net> <000c01c2fd56$6fca7830$2dbc0e44@0202012>
Mime-Version: 1.0
Content-Type: text/plain; charset=us-ascii
Content-Disposition: inline
In-Reply-To: <000c01c2fd56$6fca7830$2dbc0e44@0202012>
User-Agent: Mutt/1.4i
X-Spam-Status: No, hits=-38.8 required=5.0
	tests=BAYES_01,EMAIL_ATTRIBUTION,IN_REP_TO,QUOTED_EMAIL_TEXT,
	      REFERENCES,REPLY_WITH_QUOTES,USER_AGENT_MUTT
	autolearn=ham	version=2.50
X-Spam-Checker-Version: SpamAssassin 2.50 (1.173-2003-02-20-exp)
Sender: owner-idn@ops.ietf.org
Precedence: bulk

jim@mathies.com wrote:

> I'm trying to understand the purpose behind some steps in 
> ToUnicode.  From the IDNA spec:
> 
> 
> 1. If all code points in the sequence are in the ASCII range (0..7F)
>    then skip to step 3.
> 
> 2. Perform the steps specified in [NAMEPREP] and fail if there is an
>    error. (If step 3 of ToAscii is also performed here, it will not
>    affect the overall behavior of ToUnicode, but it is not
>    necessary.) The AllowUnassigned flag is used in [NAMEPREP].
> 
> 3. Verify that the sequence begins with the ACE prefix, and save a
>    copy of the sequence.
> 
> 
> I'm curious about steps 1 & 2.  I don't understand why nameprep
> is being applied to ASCII domain labels.

Nameprep is *not* being applied to ASCII labels.  That's what step 1
does, it prevents nameprep from being applied to ASCII labels.

> Wouldn't a simple dns character compatibility check in place of steps
> 1 and 2 suffice?

ToUnicode is not only intended to be applied to DNS-compatible labels.
It is intended to be applied to any internationalized label.

Imagine an application wants to display a domain name.  The name might
contain some non-ACE ASCII labels, some ACE ASCII labels, some non-ACE
non-ASCII labels, and some ACE non-ASCII labels.  (A mixture could arise
if the name is constructed from pieces obtained from various places,
like typed user input, the clipboard, a config file, the network, etc.)
The application can simply apply ToUnicode to every label, and the
result will be an equivalent name containing no ACE labels.

By the way, if you're wondering what an ACE non-ASCII label is, the
typical example arises when an ACE label is manually typed using an
input method that outputs fullwidth characters (which are non-ASCII, but
which would be mapped to ASCII by nameprep).

AMC



From owner-idn@ops.ietf.org  Thu Apr 24 16:51:50 2003
Received: from psg.com (mailnull@psg.com [147.28.0.62])
	by ietf.org (8.9.1a/8.9.1a) with ESMTP id QAA02721
	for <idn-archive@lists.ietf.org>; Thu, 24 Apr 2003 16:51:50 -0400 (EDT)
Received: from lserv by psg.com with local (Exim 3.36 #1)
	id 198nbV-000GO3-00
	for idn-data@psg.com; Thu, 24 Apr 2003 20:46:01 +0000
Received: from arwen.cs.berkeley.edu ([128.32.132.165] helo=nicemice.net)
	by psg.com with esmtp (Exim 3.36 #1)
	id 198nbP-000GNo-00
	for idn@ops.ietf.org; Thu, 24 Apr 2003 20:45:55 +0000
Received: from amc by nicemice.net with local (Exim 3.35 #1 (Debian))
	id 198nbN-0001ca-00
	for <idn@ops.ietf.org>; Thu, 24 Apr 2003 13:45:53 -0700
Date: Thu, 24 Apr 2003 20:45:53 +0000
From: "Adam M. Costello" <idn.amc+0@nicemice.net.RemoveThisWord>
To: IETF idn working group <idn@ops.ietf.org>
Subject: [idn] ToUnicode output can be longer than input
Message-ID: <20030424204553.GA5014@nicemice.net>
Reply-To: IETF idn working group <idn@ops.ietf.org>
Mime-Version: 1.0
Content-Type: text/plain; charset=iso-8859-1
Content-Disposition: inline
Content-Transfer-Encoding: 8bit
User-Agent: Mutt/1.4i
X-Spam-Status: No, hits=-12.1 required=5.0
	tests=BAYES_10,USER_AGENT_MUTT
	autolearn=ham version=2.53
X-Spam-Checker-Version: SpamAssassin 2.53 (1.174.2.15-2003-03-30-exp)
Sender: owner-idn@ops.ietf.org
Precedence: bulk
Content-Transfer-Encoding: 8bit

The IDNA spec contains an incidental statement that was intended to be
helpful, in section 4.2:

    The ToUnicode output never contains more code points than its input.

Oops, that's not true, because Nameprep can cause strings to expand.
For example, consider the input:

x n - - fi fi - a ffl u e n t - s o u ffl - viii - u i c

The spaces are not really there, they just indicate the clusters, which
represent single code points (ligatures and roman numerals: U+FB01,
U+FB04, U+2177).  That's 24 code points.

ToUnicode would apply Nameprep (which expands the ligatures and roman
numerals to their ASCII equivalents), then apply the Punycode decoder,
yielding:

fifi-affluent-soufflé-viii

(For the Latin-1 impaired, the non-ASCII character is e with an acute
accent.)  That's 26 code points.  26 > 24.

So the statement needs to be removed or altered if/when the RFC is
revised.  It would be correct to say that the Punycode decoder cannot
output more code points than it inputs, but Nameprep can, and therefore
ToUnicode can.

AMC



From owner-idn@ops.ietf.org  Fri Apr 25 09:27:56 2003
Received: from psg.com (mailnull@psg.com [147.28.0.62])
	by ietf.org (8.9.1a/8.9.1a) with ESMTP id JAA08082
	for <idn-archive@lists.ietf.org>; Fri, 25 Apr 2003 09:27:51 -0400 (EDT)
Received: from lserv by psg.com with local (Exim 3.36 #1)
	id 19938j-00015J-00
	for idn-data@psg.com; Fri, 25 Apr 2003 13:21:21 +0000
Received: from www.namesbeyond.com ([216.220.34.103] helo=neteka.com)
	by psg.com with smtp (Exim 3.36 #1)
	id 19938b-000139-00
	for idn@ops.ietf.org; Fri, 25 Apr 2003 13:21:13 +0000
Message-ID: <018501c30b2d$86dcb170$0f01a8c0@neteka.inc>
From: "Edmon Chung" <edmon@neteka.com>
To: "IETF idn working group" <idn@ops.ietf.org>
References: <20030424204553.GA5014@nicemice.net>
Subject: Re: [idn] ToUnicode output can be longer than input
Date: Fri, 25 Apr 2003 09:20:43 -0400
MIME-Version: 1.0
Content-Type: text/plain;
	charset="iso-8859-1"
Content-Transfer-Encoding: 7bit
X-Priority: 3
X-MSMail-Priority: Normal
X-Mailer: Microsoft Outlook Express 6.00.2600.0000
X-MimeOLE: Produced By Microsoft MimeOLE V6.00.2600.0000
X-Spam-Status: No, hits=-22.3 required=5.0
	tests=BAYES_01,MAILTO_TO_REMOVE,ORIGINAL_MESSAGE,
	      QUOTED_EMAIL_TEXT,REFERENCES
	autolearn=ham version=2.53
X-Spam-Checker-Version: SpamAssassin 2.53 (1.174.2.15-2003-03-30-exp)
Sender: owner-idn@ops.ietf.org
Precedence: bulk
Content-Transfer-Encoding: 7bit

Hi Adam,

----- Original Message -----
From: "Adam M. Costello" <idn.amc+0@nicemice.net.RemoveThisWord>
> For example, consider the input:
>
> x n - - fi fi - a ffl u e n t - s o u ffl - viii - u i c
>
> The spaces are not really there, they just indicate the clusters, which
> represent single code points (ligatures and roman numerals: U+FB01,
> U+FB04, U+2177).  That's 24 code points.

If I counted it correctly, there are 33 "codepoints" in the above ACE
string. (I agree with your assessment however, please see below, but the
example doesnt seem to illustrate your point...)

> The IDNA spec contains an incidental statement that was intended to be
> helpful, in section 4.2:
>
>     The ToUnicode output never contains more code points than its input.
>
> Oops, that's not true, because Nameprep can cause strings to expand.

I can understand this possibility.
Basically, if the length of the Unicode composition for one or more
characters in the string is longer than the ACE composition and the total
excess for all the characters within the string is more than 4 (compensating
the "xn--"), then the ToUnicode output will be longer than the input.

> So the statement needs to be removed or altered if/when the RFC is
> revised.  It would be correct to say that the Punycode decoder cannot
> output more code points than it inputs, but Nameprep can, and therefore
> ToUnicode can.

Seems reasonable.

Edmon




From owner-idn@ops.ietf.org  Fri Apr 25 15:35:41 2003
Received: from psg.com (mailnull@psg.com [147.28.0.62])
	by ietf.org (8.9.1a/8.9.1a) with ESMTP id PAA19918
	for <idn-archive@lists.ietf.org>; Fri, 25 Apr 2003 15:35:40 -0400 (EDT)
Received: from lserv by psg.com with local (Exim 3.36 #1)
	id 1998uC-000N0H-00
	for idn-data@psg.com; Fri, 25 Apr 2003 19:30:44 +0000
Received: from arwen.cs.berkeley.edu ([128.32.132.165] helo=nicemice.net)
	by psg.com with esmtp (Exim 3.36 #1)
	id 1998uA-000N05-00
	for idn@ops.ietf.org; Fri, 25 Apr 2003 19:30:42 +0000
Received: from amc by nicemice.net with local (Exim 3.35 #1 (Debian))
	id 1998u9-0004LQ-00
	for <idn@ops.ietf.org>; Fri, 25 Apr 2003 12:30:41 -0700
Date: Fri, 25 Apr 2003 19:30:41 +0000
From: "Adam M. Costello" <idn.amc+0@nicemice.net.RemoveThisWord>
To: IETF idn working group <idn@ops.ietf.org>
Subject: Re: [idn] ToUnicode output can be longer than input
Message-ID: <20030425193041.GA16627@nicemice.net>
Reply-To: IETF idn working group <idn@ops.ietf.org>
References: <20030424204553.GA5014@nicemice.net> <018501c30b2d$86dcb170$0f01a8c0@neteka.inc>
Mime-Version: 1.0
Content-Type: text/plain; charset=us-ascii
Content-Disposition: inline
In-Reply-To: <018501c30b2d$86dcb170$0f01a8c0@neteka.inc>
User-Agent: Mutt/1.4i
X-Spam-Status: No, hits=-39.4 required=5.0
	tests=BAYES_01,EMAIL_ATTRIBUTION,IN_REP_TO,QUOTED_EMAIL_TEXT,
	      QUOTE_TWICE_1,REFERENCES,REPLY_WITH_QUOTES,USER_AGENT_MUTT
	autolearn=ham version=2.53
X-Spam-Checker-Version: SpamAssassin 2.53 (1.174.2.15-2003-03-30-exp)
Sender: owner-idn@ops.ietf.org
Precedence: bulk

Edmon Chung <edmon@neteka.com> wrote:

> > x n - - fi fi - a ffl u e n t - s o u ffl - viii - u i c
> >
> > The spaces are not really there, they just indicate the clusters, which
> > represent single code points (ligatures and roman numerals: U+FB01,
> > U+FB04, U+2177).  That's 24 code points.
> 
> If I counted it correctly, there are 33 "codepoints" in the above ACE
> string.

fi represents one code point (U+FB01), ffl represents one code point
(U+FB04), and viii represents one code point (U+2177).  Now if you count
again, you should count 24.  I'm trying to describe a non-ASCII ACE
string containing 24 code points, some of which are ASCII and some of
which are compatibility characters.

AMC



From owner-idn@ops.ietf.org  Sat Apr 26 19:30:33 2003
Received: from psg.com (mailnull@psg.com [147.28.0.62])
	by ietf.org (8.9.1a/8.9.1a) with ESMTP id TAA02513
	for <idn-archive@lists.ietf.org>; Sat, 26 Apr 2003 19:30:33 -0400 (EDT)
Received: from lserv by psg.com with local (Exim 3.36 #1)
	id 199YzP-0009Aw-00
	for idn-data@psg.com; Sat, 26 Apr 2003 23:21:51 +0000
Received: from kathmandu.sun.com ([192.18.98.36])
	by psg.com with esmtp (Exim 3.36 #1)
	id 199YzN-0009Ak-00
	for idn@ops.ietf.org; Sat, 26 Apr 2003 23:21:49 +0000
Received: from bebop.France.Sun.COM ([129.157.174.15])
	by kathmandu.sun.com (8.9.3p2+Sun/8.9.3) with ESMTP id RAA19902
	for <idn@ops.ietf.org>; Sat, 26 Apr 2003 17:21:47 -0600 (MDT)
Received: from localhost (punchin-nordmark.Eng.Sun.COM [192.9.61.11])
	by bebop.France.Sun.COM (8.11.6+Sun/8.10.2/ENSMAIL,v2.2) with SMTP id h3QNLjL24455
	for <idn@ops.ietf.org>; Sun, 27 Apr 2003 01:21:45 +0200 (MEST)
Date: Sun, 27 Apr 2003 01:21:25 +0200 (CEST)
From: Erik Nordmark <Erik.Nordmark@sun.com>
Reply-To: Erik Nordmark <Erik.Nordmark@sun.com>
Subject: Re: [idn] ToUnicode output can be longer than input
To: IETF idn working group <idn@ops.ietf.org>
In-Reply-To: "Your message with ID" <20030424204553.GA5014@nicemice.net>
Message-ID: <Roam.SIMC.2.0.6.1051399285.26791.nordmark@bebop.france>
MIME-Version: 1.0
Content-Type: TEXT/PLAIN; CHARSET=US-ASCII
X-Spam-Status: No, hits=-13.0 required=5.0
	tests=BAYES_01,IN_REP_TO,QUOTED_EMAIL_TEXT
	autolearn=ham version=2.53
X-Spam-Checker-Version: SpamAssassin 2.53 (1.174.2.15-2003-03-30-exp)
Sender: owner-idn@ops.ietf.org
Precedence: bulk

> So the statement needs to be removed or altered if/when the RFC is
> revised.  It would be correct to say that the Punycode decoder cannot
> output more code points than it inputs, but Nameprep can, and therefore
> ToUnicode can.

The RFC editor maintains an errata page. It would make sense to send them
email asking them to add this to their page.

  Erik




From owner-idn@ops.ietf.org  Sun Apr 27 04:18:14 2003
Received: from psg.com (mailnull@psg.com [147.28.0.62])
	by ietf.org (8.9.1a/8.9.1a) with ESMTP id EAA20629
	for <idn-archive@lists.ietf.org>; Sun, 27 Apr 2003 04:18:14 -0400 (EDT)
Received: from lserv by psg.com with local (Exim 3.36 #1)
	id 199hFP-000EjG-00
	for idn-data@psg.com; Sun, 27 Apr 2003 08:10:55 +0000
Received: from mailout10.sul.t-online.com ([194.25.134.21])
	by psg.com with esmtp (Exim 3.36 #1)
	id 199hFJ-000Egw-00
	for idn@ops.ietf.org; Sun, 27 Apr 2003 08:10:49 +0000
Received: from fwd02.sul.t-online.de 
	by mailout10.sul.t-online.com with smtp 
	id 199hFC-0005C6-06; Sun, 27 Apr 2003 10:10:42 +0200
Received: from mira.informatik.hu-berlin.de (03047300346-0001@[217.232.42.8]) by fmrl02.sul.t-online.com
	with esmtp id 199hF2-20CKYKC; Sun, 27 Apr 2003 10:10:32 +0200
Received: from mira.informatik.hu-berlin.de (localhost [127.0.0.1])
	by mira.informatik.hu-berlin.de (8.12.6/8.11.6/SuSE Linux 0.5) with ESMTP id h3R8AbRE001987;
	Sun, 27 Apr 2003 10:10:37 +0200
Received: (from martin@localhost)
	by mira.informatik.hu-berlin.de (8.12.6/8.12.6/Submit) id h3R8AYZ6001984;
	Sun, 27 Apr 2003 10:10:34 +0200
X-Authentication-Warning: mira.informatik.hu-berlin.de: martin set sender to martin@v.loewis.de using -f
To: Marc Blanchet <Marc.Blanchet@viagenie.qc.ca>
Cc: idn@ops.ietf.org
Subject: Re: [idn] implementations list
References: <HEEHIJAAIOLDCMKIFMKLMEFNCDAA.iana@iana.org>
	<21980000.1046183027@classic.hexago.com>
From: martin@v.loewis.de (Martin v. =?iso-8859-15?q?L=F6wis?=)
Date: 27 Apr 2003 10:10:32 +0200
In-Reply-To: <21980000.1046183027@classic.hexago.com>
Message-ID: <m3llxwtl6v.fsf@mira.informatik.hu-berlin.de>
Lines: 22
User-Agent: Gnus/5.09 (Gnus v5.9.0) Emacs/21.2
MIME-Version: 1.0
Content-Type: text/plain; charset=us-ascii
X-Sender: 03047300346-0001@t-dialin.net
X-Spam-Status: No, hits=-38.5 required=5.0
	tests=BAYES_01,EMAIL_ATTRIBUTION,IN_REP_TO,QUOTED_EMAIL_TEXT,
	      RCVD_IN_NJABL,RCVD_IN_OSIRUSOFT_COM,REFERENCES,
	      REPLY_WITH_QUOTES,USER_AGENT_GNUS_UA,X_AUTH_WARNING
	autolearn=ham version=2.53
X-Spam-Checker-Version: SpamAssassin 2.53 (1.174.2.15-2003-03-30-exp)
Sender: owner-idn@ops.ietf.org
Precedence: bulk

Marc Blanchet <Marc.Blanchet@viagenie.qc.ca> writes:

> I would like to rework the http://www.i-d-n.net (which is very outdated, my
> apologies) to reflect the status of the idn work and to start collecting
> information on implementations.

Python 2.3 implements IDNA. Here is the record in the format you are
requesting.

Name: Python 2.3b1
Purpose: library
Programming language: Python
Url: http://www.python.org/2.3/
     http://www.python.org/dev/doc/devel/lib/module-encodings.idna.html
Description: Unicode strings are transparently accepted as host names
 in the socket, ftplib, httplib, and urllib libraries. For conversions
 from ACE, the "idna" codec is provided. The implementation assumes
 query strings, and UseSTD3ASCIIRules is false. Along with the IDNA
 implementation comes a "punycode" codec and a stringprep module.

Regards,
Martin



From owner-idn@ops.ietf.org  Sun Apr 27 10:41:14 2003
Received: from psg.com (mailnull@psg.com [147.28.0.62])
	by ietf.org (8.9.1a/8.9.1a) with ESMTP id KAA26318
	for <idn-archive@lists.ietf.org>; Sun, 27 Apr 2003 10:41:14 -0400 (EDT)
Received: from lserv by psg.com with local (Exim 3.36 #1)
	id 199n6c-0004BA-00
	for idn-data@psg.com; Sun, 27 Apr 2003 14:26:14 +0000
From: "Edmon Chung" <edmon@neteka.com>
To: "IETF idn working group" <idn@ops.ietf.org>
Subject: Re: [idn] ToUnicode output can be longer than input
Date: Fri, 25 Apr 2003 16:35:10 -0400
Sender: owner-idn@ops.ietf.org
Precedence: bulk
Message-Id: <E199n6c-0004BA-00@psg.com>

Hi Adam,

----- Original Message -----
From: "Adam M. Costello" <idn.amc+0@nicemice.net.RemoveThisWord>
> > > x n - - fi fi - a ffl u e n t - s o u ffl - viii - u i c
> > >
> > > The spaces are not really there, they just indicate the clusters,
which
> > > represent single code points (ligatures and roman numerals: U+FB01,
> > > U+FB04, U+2177).  That's 24 code points.
> >
> > If I counted it correctly, there are 33 "codepoints" in the above ACE
> > string.
>
> fi represents one code point (U+FB01), ffl represents one code point
> (U+FB04), and viii represents one code point (U+2177).  Now if you count
> again, you should count 24.  I'm trying to describe a non-ASCII ACE
> string containing 24 code points, some of which are ASCII and some of
> which are compatibility characters.
>

I understand, your intent, however I think it would be better to find an
example that is a valid Punycode string that when ToUnicode is performed
will exceed the number of codepoints of the original.  Right now, the ACE
string provided is not valid because it contains characters beyond A-z,
0-9, -.

Edmon






From owner-idn@ops.ietf.org  Mon Apr 28 16:17:30 2003
Received: from psg.com (mailnull@psg.com [147.28.0.62])
	by ietf.org (8.9.1a/8.9.1a) with ESMTP id QAA13375
	for <idn-archive@lists.ietf.org>; Mon, 28 Apr 2003 16:17:29 -0400 (EDT)
Received: from lserv by psg.com with local (Exim 3.36 #1)
	id 19AEyS-000H4x-00
	for idn-data@psg.com; Mon, 28 Apr 2003 20:11:40 +0000
Received: from arwen.cs.berkeley.edu ([128.32.132.165] helo=nicemice.net)
	by psg.com with esmtp (Exim 3.36 #1)
	id 19AEyP-000H4l-00
	for idn@ops.ietf.org; Mon, 28 Apr 2003 20:11:37 +0000
Received: from amc by nicemice.net with local (Exim 3.35 #1 (Debian))
	id 19AEyP-0004Ih-00
	for <idn@ops.ietf.org>; Mon, 28 Apr 2003 13:11:37 -0700
Date: Mon, 28 Apr 2003 20:11:36 +0000
From: "Adam M. Costello" <idn.amc+0@nicemice.net.RemoveThisWord>
To: IETF idn working group <idn@ops.ietf.org>
Subject: Re: [idn] ToUnicode output can be longer than input
Message-ID: <20030428201136.GB15810@nicemice.net>
Reply-To: IETF idn working group <idn@ops.ietf.org>
References: <E199n6c-0004BA-00@psg.com>
Mime-Version: 1.0
Content-Type: text/plain; charset=us-ascii
Content-Disposition: inline
In-Reply-To: <E199n6c-0004BA-00@psg.com>
User-Agent: Mutt/1.4i
X-Spam-Status: No, hits=-38.0 required=5.0
	tests=BAYES_10,EMAIL_ATTRIBUTION,IN_REP_TO,QUOTED_EMAIL_TEXT,
	      REFERENCES,REPLY_WITH_QUOTES,USER_AGENT_MUTT
	autolearn=ham version=2.53
X-Spam-Checker-Version: SpamAssassin 2.53 (1.174.2.15-2003-03-30-exp)
Sender: owner-idn@ops.ietf.org
Precedence: bulk

Edmon Chung <edmon@neteka.com> wrote:

> Right now, the ACE string provided is not valid because it contains
> characters beyond A-z, 0-9, -.

It is valid.  Some labels are ASCII and some are not.  Some labels are
ACE and some are not.  All four combinations are possible (ASCII ACE,
ASCII non-ACE, non-ASCII ACE, non-ASCII non-ACE).

An ACE label is formally defined as a label that ToUnicode would alter.
A (valid) internationalized label is formally defined as a label to
which ToASCII can be applied without failing.  It can be shown that all
ACE labels are (valid) internationalized labels.

> I think it would be better to find an example that is a valid Punycode
> string that when ToUnicode is performed will exceed the number of
> codepoints of the original.

The Punycode decoder cannot output more code points than it inputs.

If the input of ToUnicode is ASCII, then Nameprep will not be applied,
and therefore the output of ToUnicode cannot contain more code points
than the input.  It's Nameprep that can cause strings to grow, not the
Punycode decoder.

AMC



From owner-idn@ops.ietf.org  Tue Apr 29 03:50:51 2003
Received: from psg.com (mailnull@psg.com [147.28.0.62])
	by ietf.org (8.9.1a/8.9.1a) with ESMTP id DAA09493
	for <idn-archive@lists.ietf.org>; Tue, 29 Apr 2003 03:50:51 -0400 (EDT)
Received: from lserv by psg.com with local (Exim 3.36 #1)
	id 19APoE-000J0D-00
	for idn-data@psg.com; Tue, 29 Apr 2003 07:45:50 +0000
Date: Tue, 29 Apr 2003 08:24:52 +0200 (CEST)
From: Dan Oscarsson <Dan.Oscarsson@kiconsulting.se>
Reply-To: Dan Oscarsson <Dan.Oscarsson@kiconsulting.se>
Subject: Re: [idn] ToUnicode output can be longer than input
To: idn@ops.ietf.org
Cc: idn.amc+0@nicemice.net
Sender: owner-idn@ops.ietf.org
Precedence: bulk
Message-Id: <E19APoE-000J0D-00@psg.com>

Adam M. Costello wrote:

>An ACE label is formally defined as a label that ToUnicode would alter.
>A (valid) internationalized label is formally defined as a label to
>which ToASCII can be applied without failing.  It can be shown that all
>ACE labels are (valid) internationalized labels.

No that is wrong.

A IDNA ACE label is defined as above. Not ACE in general.
An internationalized label might be an ACE label, but is only
defined by ToASCII within IDNA. In general a domain name (including
all so called internationalised) do not require the IDNA ToASCII
to work. There are several domain names that will fail when ToASCII
is used, but are still domain names. They just cannot be handled by IDNA.

While this discussion is focused on IDNA, IDNA do not define thw world
and do not define the basic semantics of ACE or domain names with non-ASCII
characters. IDNA does only define a way to encode domain names so
they can be sent over lagacy ASCII DNS protocol. It does not define
what domain names work in an international context.

   Dan






From owner-idn@ops.ietf.org  Tue Apr 29 04:53:39 2003
Received: from psg.com (mailnull@psg.com [147.28.0.62])
	by ietf.org (8.9.1a/8.9.1a) with ESMTP id DAA09492
	for <idn-archive@lists.ietf.org>; Tue, 29 Apr 2003 03:50:51 -0400 (EDT)
Received: from lserv by psg.com with local (Exim 3.36 #1)
	id 19APqC-000JBy-00
	for idn-data@psg.com; Tue, 29 Apr 2003 07:47:52 +0000
Received: from arwen.cs.berkeley.edu ([128.32.132.165] helo=nicemice.net)
	by psg.com with esmtp (Exim 3.36 #1)
	id 19APq7-000JBZ-00
	for idn@ops.ietf.org; Tue, 29 Apr 2003 07:47:47 +0000
Received: from amc by nicemice.net with local (Exim 3.35 #1 (Debian))
	id 19APq7-0005vj-00
	for <idn@ops.ietf.org>; Tue, 29 Apr 2003 00:47:47 -0700
Date: Tue, 29 Apr 2003 07:47:47 +0000
From: "Adam M. Costello" <idn.amc+0@nicemice.net.RemoveThisWord>
To: idn@ops.ietf.org
Subject: Re: [idn] ToUnicode output can be longer than input
Message-ID: <20030429074747.GL15810@nicemice.net>
Reply-To: IETF idn working group <idn@ops.ietf.org>
References: <200304290624.h3T6OpWl000645@valinor.malmo.kicore.net>
Mime-Version: 1.0
Content-Type: text/plain; charset=us-ascii
Content-Disposition: inline
In-Reply-To: <200304290624.h3T6OpWl000645@valinor.malmo.kicore.net>
User-Agent: Mutt/1.4i
X-Spam-Status: No, hits=-38.6 required=5.0
	tests=BAYES_10,EMAIL_ATTRIBUTION,IN_REP_TO,QUOTED_EMAIL_TEXT,
	      QUOTE_TWICE_1,REFERENCES,REPLY_WITH_QUOTES,USER_AGENT_MUTT
	autolearn=ham version=2.53
X-Spam-Checker-Version: SpamAssassin 2.53 (1.174.2.15-2003-03-30-exp)
Sender: owner-idn@ops.ietf.org
Precedence: bulk

Dan Oscarsson <Dan.Oscarsson@kiconsulting.se> wrote:

> > An ACE label is formally defined as a label that ToUnicode would alter.
> > A (valid) internationalized label is formally defined as a label to
> > which ToASCII can be applied without failing.  It can be shown that all
> > ACE labels are (valid) internationalized labels.
> 
> No that is wrong.
> 
> A IDNA ACE label is defined as above.  Not ACE in general.

When I said "ACE label" above I was obviously talking about an IDNA ACE
label, not some more general ACE concept.  I had started out talking
about whether ToUnicode can output more code points than it inputs,
and Edmon questioned whether the example I gave was a valid ACE.  He
obviously meant IDNA ACE, and my statement above was explaining why it
was in fact a valid IDNA ACE.

I think the term ACE was introduced before I arrived on this mailing
list, so I have no memory of its origin, but I do remember the group
deciding that after a particular ACE was selected, it would be known
simply as "ACE".  Now that IDNA is a proposed standard, I would argue
that "ACE" means "IDNA ACE".

> There are several domain names that will fail when ToASCII is used,
> but are still domain names.  They just cannot be handled by IDNA.

There are non-text domain labels to which ToASCII cannot even be
applied, because ToASCII can be applied only to labels that are text.
But among text labels, ToASCII defines which ones are valid.

> IDNA do not define the world and do not define the basic semantics
> of ACE or domain names with non-ASCII characters.  IDNA does only
> define a way to encode domain names so they can be sent over lagacy
> ASCII DNS protocol.  It does not define what domain names work in an
> international context.

The IDNA spec disagrees.  It says:

    This document defines internationalized domain names (IDNs)...

    If an application wants to use non-ASCII characters in domain names,
    IDNA is the only currently-defined option.

    Applications can also define protocols and interfaces that support
    IDNs directly using non-ASCII representations.  IDNA does not
    prescribe any particular representation for new protocols, but it
    still defines which names are valid and how they are compared.

AMC



From owner-idn@ops.ietf.org  Tue Apr 29 07:36:34 2003
Received: from psg.com (mailnull@psg.com [147.28.0.62])
	by ietf.org (8.9.1a/8.9.1a) with ESMTP id HAA12843
	for <idn-archive@lists.ietf.org>; Tue, 29 Apr 2003 07:36:33 -0400 (EDT)
Received: from lserv by psg.com with local (Exim 3.36 #1)
	id 19ATBO-000Clp-00
	for idn-data@psg.com; Tue, 29 Apr 2003 11:21:58 +0000
Received: from ns.jck.com ([209.187.148.211] helo=bs.jck.com)
	by psg.com with esmtp (Exim 3.36 #1)
	id 19ATBL-000ClS-00
	for idn@ops.ietf.org; Tue, 29 Apr 2003 11:21:55 +0000
Received: from [209.187.148.215] (helo=p3.JCK.COM)
	by bs.jck.com with esmtp (Exim 4.10)
	id 19ATBK-0000S2-00
	for idn@ops.ietf.org; Tue, 29 Apr 2003 06:21:54 -0500
Date: Tue, 29 Apr 2003 07:21:54 -0400
From: John C Klensin <klensin@jck.com>
To: IETF idn working group <idn@ops.ietf.org>
Subject: Re: [idn] ToUnicode output can be longer than input
Message-ID: <61796558.1051600914@p3.JCK.COM>
In-Reply-To: <20030429074747.GL15810@nicemice.net>
References: <200304290624.h3T6OpWl000645@valinor.malmo.kicore.net>
  <20030429074747.GL15810@nicemice.net>
X-Mailer: Mulberry/3.0.3 (Win32)
MIME-Version: 1.0
Content-Type: text/plain; charset=us-ascii; format=flowed
Content-Transfer-Encoding: 7bit
Content-Disposition: inline
X-Spam-Status: No, hits=-25.7 required=5.0
	tests=BAYES_01,IN_REP_TO,MAILTO_TO_REMOVE,QUOTED_EMAIL_TEXT,
	      REFERENCES,REPLY_WITH_QUOTES
	autolearn=ham version=2.53
X-Spam-Checker-Version: SpamAssassin 2.53 (1.174.2.15-2003-03-30-exp)
Sender: owner-idn@ops.ietf.org
Precedence: bulk
Content-Transfer-Encoding: 7bit



--On Tuesday, 29 April, 2003 07:47 +0000 "Adam M. Costello" 
<idn.amc+0@nicemice.net.RemoveThisWord> wrote:

> I think the term ACE was introduced before I arrived on this
> mailing list, so I have no memory of its origin, but I do
> remember the group deciding that after a particular ACE was
> selected, it would be known simply as "ACE".  Now that IDNA is
> a proposed standard, I would argue that "ACE" means "IDNA ACE".

Adam,

Regardless of what the WG may or may not have decided at the 
time, the fact that various "testbed names" and other 
non-standard approaches continue to exist (something that at 
least some of the WG didn't anticipate), makes the use of the 
term "ACE" at least ambiguous.  The WG would be doing itself, 
and the community, a disservice by ignoring the fact that there 
continue to be other "ACE" approaches in active use and 
insisting that only it can define what "ACE" means.

I have been using "IDNA name" or "IDNA ACE" in my writing, and 
would suggest that others adopt that, or some similar, 
convention that precisely identifies what is being discussed.

thanks,
      john




From owner-idn@ops.ietf.org  Tue Apr 29 18:52:32 2003
Received: from psg.com (mailnull@psg.com [147.28.0.62])
	by ietf.org (8.9.1a/8.9.1a) with ESMTP id SAA15690
	for <idn-archive@lists.ietf.org>; Tue, 29 Apr 2003 18:52:31 -0400 (EDT)
Received: from lserv by psg.com with local (Exim 3.36 #1)
	id 19AdsS-0008pH-00
	for idn-data@psg.com; Tue, 29 Apr 2003 22:47:08 +0000
Received: from arwen.cs.berkeley.edu ([128.32.132.165] helo=nicemice.net)
	by psg.com with esmtp (Exim 3.36 #1)
	id 19AdsP-0008p2-00
	for idn@ops.ietf.org; Tue, 29 Apr 2003 22:47:05 +0000
Received: from amc by nicemice.net with local (Exim 3.35 #1 (Debian))
	id 19AdsO-0007fF-00
	for <idn@ops.ietf.org>; Tue, 29 Apr 2003 15:47:04 -0700
Date: Tue, 29 Apr 2003 22:47:04 +0000
From: "Adam M. Costello" <idn.amc+0@nicemice.net.RemoveThisWord>
To: idn@ops.ietf.org
Subject: Re: [idn] ToUnicode output can be longer than input
Message-ID: <20030429224703.GA29027@nicemice.net>
Reply-To: IETF idn working group <idn@ops.ietf.org>
References: <200304291336.h3TDaYt7000864@valinor.malmo.kicore.net>
Mime-Version: 1.0
Content-Type: text/plain; charset=us-ascii
Content-Disposition: inline
In-Reply-To: <200304291336.h3TDaYt7000864@valinor.malmo.kicore.net>
User-Agent: Mutt/1.4i
X-Spam-Status: No, hits=-39.4 required=5.0
	tests=BAYES_01,EMAIL_ATTRIBUTION,IN_REP_TO,QUOTED_EMAIL_TEXT,
	      QUOTE_TWICE_1,REFERENCES,REPLY_WITH_QUOTES,USER_AGENT_MUTT
	autolearn=ham version=2.53
X-Spam-Checker-Version: SpamAssassin 2.53 (1.174.2.15-2003-03-30-exp)
Sender: owner-idn@ops.ietf.org
Precedence: bulk

Dan Oscarsson <Dan.Oscarsson@kiconsulting.se> wrote:

> > The IDNA spec disagrees.  It says:
> >
> >    This document defines internationalized domain names (IDNs)...
> >
> >    If an application wants to use non-ASCII characters in domain
> >    names, IDNA is the only currently-defined option.
> >
> >    Applications can also define protocols and interfaces that
> >    support IDNs directly using non-ASCII representations.  IDNA does
> >    not prescribe any particular representation for new protocols,
> >    but it still defines which names are valid and how they are
> >    compared.
>
> Maybe IDNs, as that is a construct of IDNA, but not domain names in
> general.  Domain names in an international context is not defined by
> IDNA,

The *representation* of domain names in an international context is not
defined by IDNA, but IDNA does define a mechanism for deciding which
names are valid in that context, and does define a way to compare names
in that context.

The IETF could produce a second set of definitions in the future, but
until that happens, the IDNA definitions provide the only standard way
to use non-ASCII characters in domain names, even in an international
context.

> I can see no reason to limit the international world due to limits of
> ASCII and the solution selected by IDNA.

The reason is backward compatibility, so that all domain names can be
accessed by all protocols.  That was the primary motivation behind the
whole IDNA approach.  Even as new protocols are introduced that can
represent IDNs directly without using ACE, old protocols will continue
to be used, and it would be a mess if some names worked with some
protocols and not with other protocols.

AMC



From owner-idn@ops.ietf.org  Tue Apr 29 21:45:41 2003
Received: from psg.com (mailnull@psg.com [147.28.0.62])
	by ietf.org (8.9.1a/8.9.1a) with ESMTP id VAA20884
	for <idn-archive@lists.ietf.org>; Tue, 29 Apr 2003 21:45:40 -0400 (EDT)
Received: from lserv by psg.com with local (Exim 3.36 #1)
	id 19AgWn-000O45-00
	for idn-data@psg.com; Wed, 30 Apr 2003 01:36:57 +0000
Date: Tue, 29 Apr 2003 15:36:34 +0200 (CEST)
From: Dan Oscarsson <Dan.Oscarsson@kiconsulting.se>
Reply-To: Dan Oscarsson <Dan.Oscarsson@kiconsulting.se>
Subject: Re: [idn] ToUnicode output can be longer than input
To: idn@ops.ietf.org
Cc: idn.amc+0@nicemice.net
Sender: owner-idn@ops.ietf.org
Precedence: bulk
Message-Id: <E19AgWn-000O45-00@psg.com>

Adam M. Costello wrote:

>> There are several domain names that will fail when ToASCII is used,
>> but are still domain names.  They just cannot be handled by IDNA.
>
>There are non-text domain labels to which ToASCII cannot even be
>applied, because ToASCII can be applied only to labels that are text.
>But among text labels, ToASCII defines which ones are valid.

Valid for IDNA, not valid in general.

>
>> IDNA do not define the world and do not define the basic semantics
>> of ACE or domain names with non-ASCII characters.  IDNA does only
>> define a way to encode domain names so they can be sent over lagacy
>> ASCII DNS protocol.  It does not define what domain names work in an
>> international context.
>
>The IDNA spec disagrees.  It says:
>
>    This document defines internationalized domain names (IDNs)...
>
>    If an application wants to use non-ASCII characters in domain names,
>    IDNA is the only currently-defined option.
>
>    Applications can also define protocols and interfaces that support
>    IDNs directly using non-ASCII representations.  IDNA does not
>    prescribe any particular representation for new protocols, but it
>    still defines which names are valid and how they are compared.

Maybe IDNs, as that is a construct of IDNA, but not domain names
in general. Domain names in an international context is not defined
by IDNA, IDNA have limitations so that it cannot handle all
domain names that can exist in an international context.
I can see no reason to limit the international world due to limits
of ASCII and the solution selected by IDNA.

When an international context of the DNS protocol is defined, it can
support more domain names than can be handled by IDNA.

   Dan







From owner-idn@ops.ietf.org  Wed Apr 30 00:13:01 2003
Received: from psg.com (mailnull@psg.com [147.28.0.62])
	by ietf.org (8.9.1a/8.9.1a) with ESMTP id AAA24618
	for <idn-archive@lists.ietf.org>; Wed, 30 Apr 2003 00:13:01 -0400 (EDT)
Received: from lserv by psg.com with local (Exim 3.36 #1)
	id 19AipD-00052a-00
	for idn-data@psg.com; Wed, 30 Apr 2003 04:04:07 +0000
Received: from mta01bw.bigpond.com ([139.134.6.78])
	by psg.com with esmtp (Exim 3.36 #1)
	id 19Aip8-00052O-00
	for idn@ops.ietf.org; Wed, 30 Apr 2003 04:04:03 +0000
Received: from BCK1 ([144.135.24.78]) by mta01bw.bigpond.com
          (Netscape Messaging Server 4.15 mta01bw Jul 16 2002 22:47:55)
          with SMTP id HE51YV00.3SB for <idn@ops.ietf.org>; Wed, 30 Apr
          2003 14:04:07 +1000 
Received: from CPE-203-51-155-68.vic.bigpond.net.au ([203.51.155.68]) by bwmam04bpa.bigpond.com(MAM V3.3.2 35/619885); 30 Apr 2003 14:05:15
From: "Jarrod Hollingworth" <jarrod@backslash.com.au>
To: "IETF idn working group" <idn@ops.ietf.org>
Subject: [idn] IDN's with any ASCII character
Date: Wed, 30 Apr 2003 14:03:54 +1000
Message-ID: <EDEAIFIAOJACNIGAMCPBAEBJFKAA.jarrod@backslash.com.au>
MIME-Version: 1.0
Content-Type: text/plain;
	charset="us-ascii"
Content-Transfer-Encoding: 7bit
X-Priority: 3 (Normal)
X-MSMail-Priority: Normal
X-Mailer: Microsoft Outlook IMO, Build 9.0.2416 (9.0.2911.0)
In-Reply-To: <20030429224703.GA29027@nicemice.net>
X-MIMEOLE: Produced By Microsoft MimeOLE V6.00.2800.1106
Importance: Normal
X-Spam-Status: No, hits=-7.7 required=5.0
	tests=IN_REP_TO,MSGID_GOOD_EXCHANGE,RCVD_IN_NJABL
	autolearn=ham version=2.53
X-Spam-Checker-Version: SpamAssassin 2.53 (1.174.2.15-2003-03-30-exp)
Sender: owner-idn@ops.ietf.org
Precedence: bulk
Content-Transfer-Encoding: 7bit

Will IDN's allow encoding of domain names with *any* ASCII character?

For example, let's say that I want to register the domain name
"this&that.com" or "100^10.com".

Will IDN allow this or does it only facilitate international languages?

Regards,
Jarrod Hollingworth





From owner-idn@ops.ietf.org  Wed Apr 30 01:21:33 2003
Received: from psg.com (mailnull@psg.com [147.28.0.62])
	by ietf.org (8.9.1a/8.9.1a) with ESMTP id BAA26191
	for <idn-archive@lists.ietf.org>; Wed, 30 Apr 2003 01:21:33 -0400 (EDT)
Received: from lserv by psg.com with local (Exim 3.36 #1)
	id 19Ajqc-0009CJ-00
	for idn-data@psg.com; Wed, 30 Apr 2003 05:09:38 +0000
Received: from arwen.cs.berkeley.edu ([128.32.132.165] helo=nicemice.net)
	by psg.com with esmtp (Exim 3.36 #1)
	id 19Ajqa-0009C6-00
	for idn@ops.ietf.org; Wed, 30 Apr 2003 05:09:36 +0000
Received: from amc by nicemice.net with local (Exim 3.35 #1 (Debian))
	id 19AjqZ-000060-00
	for <idn@ops.ietf.org>; Tue, 29 Apr 2003 22:09:35 -0700
Date: Wed, 30 Apr 2003 05:09:35 +0000
From: "Adam M. Costello" <idn.amc+0@nicemice.net.RemoveThisWord>
To: IETF idn working group <idn@ops.ietf.org>
Subject: Re: [idn] IDN's with any ASCII character
Message-ID: <20030430050935.GF29027@nicemice.net>
Reply-To: IETF idn working group <idn@ops.ietf.org>
References: <20030429224703.GA29027@nicemice.net> <EDEAIFIAOJACNIGAMCPBAEBJFKAA.jarrod@backslash.com.au>
Mime-Version: 1.0
Content-Type: text/plain; charset=us-ascii
Content-Disposition: inline
In-Reply-To: <EDEAIFIAOJACNIGAMCPBAEBJFKAA.jarrod@backslash.com.au>
User-Agent: Mutt/1.4i
X-Spam-Status: No, hits=-38.8 required=5.0
	tests=BAYES_01,EMAIL_ATTRIBUTION,IN_REP_TO,QUOTED_EMAIL_TEXT,
	      REFERENCES,REPLY_WITH_QUOTES,USER_AGENT_MUTT
	autolearn=ham version=2.53
X-Spam-Checker-Version: SpamAssassin 2.53 (1.174.2.15-2003-03-30-exp)
Sender: owner-idn@ops.ietf.org
Precedence: bulk

Jarrod Hollingworth <jarrod@backslash.com.au> wrote:

> Will IDN's allow encoding of domain names with *any* ASCII character?
>
> For example, let's say that I want to register the domain name
> "this&that.com" or "100^10.com".
>
> Will IDN allow this or does it only facilitate international
> languages?

IDNA allows the addition of non-ASCII characters to domain names.  For
ASCII characters, IDNA adds no new restrictions, but nor does it relax
the old restrictions.  The ASCII characters & and ^ (and every other
ASCII character besides letters, digits, and hyphen) are not allowed in
the "preferred syntax", which is used for domain names that name hosts
and mail exchangers.

It is not merely by fiat that IDNA keeps the old ASCII restrictions,
it follows from the technical details of the encoding.  In IDNA, every
non-ASCII domain label has an ASCII form, where the non-ASCII characters
are encoded using ASCII letters and digits.  But any ASCII characters
that occur in the non-ASCII label are represented literally, not
encoded.  For example, if we want to put an acute accent over the "a"
in this&that.com, the ASCII form will be xn--this&tht-fza.  As you can
see, IDNA does nothing to help you "sneak" the "&" into the name; it is
still there as "&", so you can't use such a name anywhere that "&" is
forbidden.

One thing that is by fiat is the restriction on initial and final
hyphens.  Technically, IDNA could have enabled one to sneak an initial
or final hyphen into a label where initial and final hyphens are
forbidden, but IDNA includes optional checks to prevent initial/final
hyphen from sneaking in where it's not allowed.

AMC



From owner-idn@ops.ietf.org  Wed Apr 30 01:30:50 2003
Received: from psg.com (mailnull@psg.com [147.28.0.62])
	by ietf.org (8.9.1a/8.9.1a) with ESMTP id BAA26361
	for <idn-archive@lists.ietf.org>; Wed, 30 Apr 2003 01:30:50 -0400 (EDT)
Received: from lserv by psg.com with local (Exim 3.36 #1)
	id 19Ak2v-0009xE-00
	for idn-data@psg.com; Wed, 30 Apr 2003 05:22:21 +0000
Received: from mta01bw.bigpond.com ([139.134.6.78])
	by psg.com with esmtp (Exim 3.36 #1)
	id 19Ak2s-0009x2-00
	for idn@ops.ietf.org; Wed, 30 Apr 2003 05:22:18 +0000
Received: from BCK1 ([144.135.24.75]) by mta01bw.bigpond.com
          (Netscape Messaging Server 4.15 mta01bw Jul 16 2002 22:47:55)
          with SMTP id HE55LA00.89T for <idn@ops.ietf.org>; Wed, 30 Apr
          2003 15:22:22 +1000 
Received: from CPE-203-51-155-68.vic.bigpond.net.au ([203.51.155.68]) by bwmam03bpa.bigpond.com(MAM V3.3.2 26/676633); 30 Apr 2003 15:22:12
From: "Jarrod Hollingworth" <jarrod@backslash.com.au>
To: "IETF idn working group" <idn@ops.ietf.org>
Subject: RE: [idn] IDN's with any ASCII character
Date: Wed, 30 Apr 2003 15:22:08 +1000
Message-ID: <EDEAIFIAOJACNIGAMCPBAEBNFKAA.jarrod@backslash.com.au>
MIME-Version: 1.0
Content-Type: text/plain;
	charset="us-ascii"
Content-Transfer-Encoding: 7bit
X-Priority: 3 (Normal)
X-MSMail-Priority: Normal
X-Mailer: Microsoft Outlook IMO, Build 9.0.2416 (9.0.2911.0)
X-MIMEOLE: Produced By Microsoft MimeOLE V6.00.2800.1106
Importance: Normal
X-Spam-Status: No, hits=-14.7 required=5.0
	tests=BAYES_10,MSGID_GOOD_EXCHANGE,QUOTED_EMAIL_TEXT
	autolearn=ham version=2.53
X-Spam-Checker-Version: SpamAssassin 2.53 (1.174.2.15-2003-03-30-exp)
Sender: owner-idn@ops.ietf.org
Precedence: bulk
Content-Transfer-Encoding: 7bit

> For example, let's say that I want to register the domain name
> "this&that.com" or "100^10.com".

And another example the micro (u) symbol - ? or other high ASCII
characters??

Jarrod





From owner-idn@ops.ietf.org  Wed Apr 30 02:31:58 2003
Received: from psg.com (mailnull@psg.com [147.28.0.62])
	by ietf.org (8.9.1a/8.9.1a) with ESMTP id CAA22966
	for <idn-archive@lists.ietf.org>; Wed, 30 Apr 2003 02:31:57 -0400 (EDT)
Received: from lserv by psg.com with local (Exim 3.36 #1)
	id 19Akzk-000E9C-00
	for idn-data@psg.com; Wed, 30 Apr 2003 06:23:08 +0000
Received: from mta1-0.mail.adelphia.net ([64.8.50.175] helo=mta1.adelphia.net)
	by psg.com with esmtp (Exim 3.36 #1)
	id 19Akzh-000E90-00
	for idn@ops.ietf.org; Wed, 30 Apr 2003 06:23:05 +0000
Received: from DouglasEwell.anhmca.adelphia.net ([68.66.67.126])
          by mta1.adelphia.net
          (InterMail vM.5.01.05.32 201-253-122-126-132-20030307) with SMTP
          id <20030430063126.OGTK17125.mta1.adelphia.net@DouglasEwell.anhmca.adelphia.net>;
          Wed, 30 Apr 2003 02:31:26 -0400
Message-ID: <004601c30ee0$af3d5000$7e434244@anhmca.adelphia.net>
From: "Doug Ewell" <dewell@adelphia.net>
To: "IETF idn working group" <idn@ops.ietf.org>
Cc: "Jarrod Hollingworth" <jarrod@backslash.com.au>
References: <EDEAIFIAOJACNIGAMCPBAEBNFKAA.jarrod@backslash.com.au>
Subject: Re: [idn] IDN's with any ASCII character
Date: Tue, 29 Apr 2003 23:21:05 -0700
MIME-Version: 1.0
Content-Type: text/plain;
	charset="utf-8"
Content-Transfer-Encoding: 8bit
X-Priority: 3
X-MSMail-Priority: Normal
X-Mailer: Microsoft Outlook Express 5.50.4807.1700
X-MIMEOLE: Produced By Microsoft MimeOLE V5.50.4807.1700
X-Spam-Status: No, hits=-28.5 required=5.0
	tests=BAYES_10,EMAIL_ATTRIBUTION,QUOTED_EMAIL_TEXT,REFERENCES,
	      REPLY_WITH_QUOTES
	autolearn=ham version=2.53
X-Spam-Checker-Version: SpamAssassin 2.53 (1.174.2.15-2003-03-30-exp)
Sender: owner-idn@ops.ietf.org
Precedence: bulk
Content-Transfer-Encoding: 8bit

Jarrod Hollingworth <jarrod at backslash dot com dot au> wrote:

> And another example the micro (u) symbol - Âµ or other high ASCII
> characters??

There's no such thing as "high ASCII."  ASCII only goes up to 0x7F.
Everything else is Latin-1 or Unicode or GB 18030 or something else, but
it can't be called ASCII.

U+00B5 MICRO SIGN would seem to be legal in an IDNA label, although not
necessarily easy for everyone to type.

-Doug Ewell
 Fullerton, California
 http://users.adelphia.net/~dewell/




From owner-idn@ops.ietf.org  Wed Apr 30 02:51:54 2003
Received: from psg.com (mailnull@psg.com [147.28.0.62])
	by ietf.org (8.9.1a/8.9.1a) with ESMTP id CAA23641
	for <idn-archive@lists.ietf.org>; Wed, 30 Apr 2003 02:51:53 -0400 (EDT)
Received: from lserv by psg.com with local (Exim 3.36 #1)
	id 19AlO9-000FsM-00
	for idn-data@psg.com; Wed, 30 Apr 2003 06:48:21 +0000
Received: from arwen.cs.berkeley.edu ([128.32.132.165] helo=nicemice.net)
	by psg.com with esmtp (Exim 3.36 #1)
	id 19AlO5-000FqL-00
	for idn@ops.ietf.org; Wed, 30 Apr 2003 06:48:17 +0000
Received: from amc by nicemice.net with local (Exim 3.35 #1 (Debian))
	id 19AlO5-0000I7-00
	for <idn@ops.ietf.org>; Tue, 29 Apr 2003 23:48:17 -0700
Date: Wed, 30 Apr 2003 06:48:17 +0000
From: "Adam M. Costello" <idn.amc+0@nicemice.net.RemoveThisWord>
To: IETF idn working group <idn@ops.ietf.org>
Subject: Re: [idn] IDN's with any ASCII character
Message-ID: <20030430064817.GG29027@nicemice.net>
Reply-To: IETF idn working group <idn@ops.ietf.org>
References: <EDEAIFIAOJACNIGAMCPBAEBNFKAA.jarrod@backslash.com.au> <004601c30ee0$af3d5000$7e434244@anhmca.adelphia.net>
Mime-Version: 1.0
Content-Type: text/plain; charset=us-ascii
Content-Disposition: inline
In-Reply-To: <004601c30ee0$af3d5000$7e434244@anhmca.adelphia.net>
User-Agent: Mutt/1.4i
X-Spam-Status: No, hits=-35.6 required=5.0
	tests=BAYES_01,EMAIL_ATTRIBUTION,IN_REP_TO,REFERENCES,
	      REPLY_WITH_QUOTES,USER_AGENT_MUTT
	autolearn=ham version=2.53
X-Spam-Checker-Version: SpamAssassin 2.53 (1.174.2.15-2003-03-30-exp)
Sender: owner-idn@ops.ietf.org
Precedence: bulk

Doug Ewell <dewell@adelphia.net> wrote:

> U+00B5 MICRO SIGN would seem to be legal in an IDNA label

Yes, it is.  Although if you convert the label to ASCII and back again,
this code point will become U+03BC (Greek small letter mu) because of
the NFKC step in Nameprep.

Also note that just because a name is syntactically valid, that
doesn't mean a registry is under any obligation to let anyone register
it.  It is expected that most registries will be cautious and limit
registrations to subsets of Unicode that they are confident they
understand well enough to manage properly.

AMC



From owner-idn@ops.ietf.org  Wed Apr 30 03:30:30 2003
Received: from psg.com (mailnull@psg.com [147.28.0.62])
	by ietf.org (8.9.1a/8.9.1a) with ESMTP id DAA24416
	for <idn-archive@lists.ietf.org>; Wed, 30 Apr 2003 03:30:29 -0400 (EDT)
Received: from lserv by psg.com with local (Exim 3.36 #1)
	id 19AlyR-000J1d-00
	for idn-data@psg.com; Wed, 30 Apr 2003 07:25:51 +0000
Received: from mta04bw.bigpond.com ([139.134.6.87])
	by psg.com with esmtp (Exim 3.36 #1)
	id 19AlyO-000J1Q-00
	for idn@ops.ietf.org; Wed, 30 Apr 2003 07:25:48 +0000
Received: from BCK1 ([144.135.24.84]) by mta04bw.bigpond.com
          (Netscape Messaging Server 4.15 mta04bw Jul 16 2002 22:47:55)
          with SMTP id HE5BA700.926 for <idn@ops.ietf.org>; Wed, 30 Apr
          2003 17:25:19 +1000 
Received: from CPE-203-51-175-31.vic.bigpond.net.au ([203.51.175.31]) by bwmam06bpa.bigpond.com(MailRouter V3.2g 53/14783245); 30 Apr 2003 17:25:19
From: "Jarrod Hollingworth" <jarrod@backslash.com.au>
To: "IETF idn working group" <idn@ops.ietf.org>
Subject: RE: [idn] IDN's with any ASCII character
Date: Wed, 30 Apr 2003 17:25:16 +1000
Message-ID: <EDEAIFIAOJACNIGAMCPBGECAFKAA.jarrod@backslash.com.au>
MIME-Version: 1.0
Content-Type: text/plain;
	charset="us-ascii"
Content-Transfer-Encoding: 7bit
X-Priority: 3 (Normal)
X-MSMail-Priority: Normal
X-Mailer: Microsoft Outlook IMO, Build 9.0.2416 (9.0.2911.0)
Importance: Normal
In-Reply-To: <20030430064817.GG29027@nicemice.net>
X-MimeOLE: Produced By Microsoft MimeOLE V6.00.2800.1106
X-Spam-Status: No, hits=-12.0 required=5.0
	tests=BAYES_20,IN_REP_TO,MSGID_GOOD_EXCHANGE
	autolearn=ham version=2.53
X-Spam-Checker-Version: SpamAssassin 2.53 (1.174.2.15-2003-03-30-exp)
Sender: owner-idn@ops.ietf.org
Precedence: bulk
Content-Transfer-Encoding: 7bit

Thanks Adam and Doug, this has cleared things up.

Regards,

Jarrod Hollingworth





From owner-idn@ops.ietf.org  Wed Apr 30 08:32:07 2003
Received: from psg.com (mailnull@psg.com [147.28.0.62])
	by ietf.org (8.9.1a/8.9.1a) with ESMTP id IAA29787
	for <idn-archive@lists.ietf.org>; Wed, 30 Apr 2003 08:32:07 -0400 (EDT)
Received: from lserv by psg.com with local (Exim 3.36 #1)
	id 19Aqa3-000EWD-00
	for idn-data@psg.com; Wed, 30 Apr 2003 12:20:59 +0000
Received: from ns.jck.com ([209.187.148.211] helo=bs.jck.com)
	by psg.com with esmtp (Exim 3.36 #1)
	id 19AqZp-000EVl-00
	for idn@ops.ietf.org; Wed, 30 Apr 2003 12:20:45 +0000
Received: from [209.187.148.215] (helo=p3.JCK.COM)
	by bs.jck.com with esmtp (Exim 4.10)
	id 19AqZi-0006d3-00; Wed, 30 Apr 2003 07:20:38 -0500
Date: Wed, 30 Apr 2003 08:20:38 -0400
From: John C Klensin <klensin@jck.com>
To: Doug Ewell <dewell@adelphia.net>,
        IETF idn working group <idn@ops.ietf.org>
cc: Jarrod Hollingworth <jarrod@backslash.com.au>
Subject: Re: [idn] IDN's with any ASCII character
Message-ID: <84052330.1051690838@p3.JCK.COM>
In-Reply-To: <004601c30ee0$af3d5000$7e434244@anhmca.adelphia.net>
References: <EDEAIFIAOJACNIGAMCPBAEBNFKAA.jarrod@backslash.com.au>
  <004601c30ee0$af3d5000$7e434244@anhmca.adelphia.net>
X-Mailer: Mulberry/3.0.3 (Win32)
MIME-Version: 1.0
Content-Type: text/plain; charset=us-ascii; format=flowed
Content-Transfer-Encoding: 7bit
Content-Disposition: inline
X-Spam-Status: No, hits=-26.0 required=5.0
	tests=BAYES_01,IN_REP_TO,QUOTED_EMAIL_TEXT,REFERENCES,
	      REPLY_WITH_QUOTES
	autolearn=ham version=2.53
X-Spam-Checker-Version: SpamAssassin 2.53 (1.174.2.15-2003-03-30-exp)
Sender: owner-idn@ops.ietf.org
Precedence: bulk
Content-Transfer-Encoding: 7bit

--On Tuesday, 29 April, 2003 23:21 -0700 Doug Ewell 
<dewell@adelphia.net> wrote:

> There's no such thing as "high ASCII."  ASCII only goes up to
> 0x7F. Everything else is Latin-1 or Unicode or GB 18030 or
> something else, but it can't be called ASCII.

Doug, just in the interest of being precise about this, my 
recollection is that the first US version of "Latin-1" (aka ISO 
8859-1) was formally known as "8 bit ASCII".  The current 
version of "ASCII" -- ANSI/INCITS 4-1986 (formerly X3.4) is 
titled "Information Systems - Coded Character Sets - 7-Bit 
American National Standard Code for Information Interchange 
(7-Bit ASCII)"
The parenthetical note is an artifact of that "... Coded 
Character Sets - 8-Bit American National... document which, if I 
recall, was withdrawn when the US adopted/endorsed ISO 8859-1 
rather than maintaining its own version.

The ANSI (and INCITS) rules about references to withdrawn 
standards are a bit muddy, but to characterize such a reference 
with "can't be called" is excessive.  While, as far as I know, 
"high ASCII" has never been a standard term, "ASCII-8" and 
"8-Bit ASCII" definitely have been.  And both terms are still 
used informally, both inside and outside the US, to refer to the 
Standardized form of Latin-1, i.e., ISO 8859-1.

      john








From owner-idn@ops.ietf.org  Wed Apr 30 08:33:21 2003
Received: from psg.com (mailnull@psg.com [147.28.0.62])
	by ietf.org (8.9.1a/8.9.1a) with ESMTP id IAA29869
	for <idn-archive@lists.ietf.org>; Wed, 30 Apr 2003 08:33:20 -0400 (EDT)
Received: from lserv by psg.com with local (Exim 3.36 #1)
	id 19AqaY-000Ecy-00
	for idn-data@psg.com; Wed, 30 Apr 2003 12:21:30 +0000
Received: from malmo.kicore.net ([217.212.0.10])
	by psg.com with esmtp (Exim 3.36 #1)
	id 19AqaU-000Ecc-00
	for idn@ops.ietf.org; Wed, 30 Apr 2003 12:21:26 +0000
Received: from valinor.malmo.kicore.net (valinor.malmo.kicore.net [217.212.0.20]) by malmo.kicore.net (8.12.9+Sun/KiNet-primary) with ESMTP id h3UCLPma015017 for <idn@ops.ietf.org>; Wed, 30 Apr 2003 14:21:25 +0200 (MEST)
Received: from valinor by valinor.malmo.kicore.net (8.12.8+Sun/TRM-1-KLIENT); Wed, 30 Apr 2003 14:21:24 +0200 (CEST) (MET)
Message-Id: <200304301221.h3UCLOt7008508@valinor.malmo.kicore.net>
Date: Wed, 30 Apr 2003 14:21:24 +0200 (CEST)
From: Dan Oscarsson <Dan.Oscarsson@kiconsulting.se>
Reply-To: Dan Oscarsson <Dan.Oscarsson@kiconsulting.se>
Subject: Re: [idn] ToUnicode output can be longer than input
To: idn@ops.ietf.org
MIME-Version: 1.0
Content-Type: TEXT/plain; charset=us-ascii
Content-MD5: ldxULKd/fDGB2QOuJhbYuw==
X-Mailer: dtmail 1.3.0 @(#)CDE Version 1.5.3_06 SunOS 5.9 sun4u sparc 
X-Spam-Status: No, hits=-15.4 required=5.0
	tests=BAYES_01,EMAIL_ATTRIBUTION,MSG_ID_ADDED_BY_MTA_3,
	      QUOTED_EMAIL_TEXT
	autolearn=ham version=2.53
X-Spam-Checker-Version: SpamAssassin 2.53 (1.174.2.15-2003-03-30-exp)
Sender: owner-idn@ops.ietf.org
Precedence: bulk

Adam M. Costello wrote:

>> Maybe IDNs, as that is a construct of IDNA, but not domain names in
>> general.  Domain names in an international context is not defined by
>> IDNA,
>
>The *representation* of domain names in an international context is not
>defined by IDNA, but IDNA does define a mechanism for deciding which
>names are valid in that context, and does define a way to compare names
>in that context.
>
>The IETF could produce a second set of definitions in the future, but
>until that happens, the IDNA definitions provide the only standard way
>to use non-ASCII characters in domain names, even in an international
>context.

Unless I remember wrong, IDNA defines a way to compare names in ASCII
context because it requires names to be in IDNA ACE format.
Comparing names in an international context must be done using
UCS characters directely.

At the moment, unfortunately, IDNA restricts the way domain names can
be written and what characters can be used, partly due to its way
to handle domain names in legacy protocols. It was possible to handle
it better, but was not chosen so.

I do not want to limit domain names, in an international context, because
of a legacy compatibility issue.

   Dan




From owner-idn@ops.ietf.org  Wed Apr 30 09:05:51 2003
Received: from psg.com (mailnull@psg.com [147.28.0.62])
	by ietf.org (8.9.1a/8.9.1a) with ESMTP id JAA01251
	for <idn-archive@lists.ietf.org>; Wed, 30 Apr 2003 09:05:51 -0400 (EDT)
Received: from lserv by psg.com with local (Exim 3.36 #1)
	id 19ArDg-000J4d-00
	for idn-data@psg.com; Wed, 30 Apr 2003 13:01:56 +0000
Received: from ns.jck.com ([209.187.148.211] helo=bs.jck.com)
	by psg.com with esmtp (Exim 3.36 #1)
	id 19ArDc-000J3n-00
	for idn@ops.ietf.org; Wed, 30 Apr 2003 13:01:53 +0000
Received: from [209.187.148.215] (helo=p3.JCK.COM)
	by bs.jck.com with esmtp (Exim 4.10)
	id 19ArDX-0006pR-00; Wed, 30 Apr 2003 08:01:47 -0500
Date: Wed, 30 Apr 2003 09:01:47 -0400
From: John C Klensin <klensin@jck.com>
To: Dan Oscarsson <Dan.Oscarsson@kiconsulting.se>,
        "idn@ops.ietf.org" <idn@ops.ietf.org>
Subject: Re: [idn] ToUnicode output can be longer than input
Message-ID: <86521110.1051693307@p3.JCK.COM>
In-Reply-To: <200304301221.h3UCLOt7008508@valinor.malmo.kicore.net>
References: <200304301221.h3UCLOt7008508@valinor.malmo.kicore.net
 >
X-Mailer: Mulberry/3.0.3 (Win32)
MIME-Version: 1.0
Content-Type: text/plain; charset=us-ascii; format=flowed
Content-Transfer-Encoding: 7bit
Content-Disposition: inline
X-Spam-Status: No, hits=-13.0 required=5.0
	tests=BAYES_01,IN_REP_TO,QUOTED_EMAIL_TEXT
	autolearn=ham version=2.53
X-Spam-Checker-Version: SpamAssassin 2.53 (1.174.2.15-2003-03-30-exp)
Sender: owner-idn@ops.ietf.org
Precedence: bulk
Content-Transfer-Encoding: 7bit



--On Wednesday, 30 April, 2003 14:21 +0200 Dan Oscarsson 
<Dan.Oscarsson@kiconsulting.se> wrote:

> Unless I remember wrong, IDNA defines a way to compare names
> in ASCII context because it requires names to be in IDNA ACE
> format. Comparing names in an international context must be
> done using UCS characters directely.
>
> At the moment, unfortunately, IDNA restricts the way domain
> names can be written and what characters can be used, partly
> due to its way to handle domain names in legacy protocols. It
> was possible to handle it better, but was not chosen so.
>
> I do not want to limit domain names, in an international
> context, because of a legacy compatibility issue.

Dan,

You've made several comments above with which I agree, and a few 
which I don't.  Fortunately or unfortunately, this isn't the 
right place for the discussion.

	* This list, or various of its participants, may have an
	opinion about the use of the term "ACE", but "invent
	definitive terminology for talking about IDNs (or
	anything else)" is not in the charter of the late WG,
	nor has reactivation and rechartering to do that been
	proposed.
	
	* IETF Standards are not mandatory and IETF has no
	enforcement capability.    Suppose a zone decides to
	adopt naming rules of its own --perhaps taking advantage
	of the assertions of RFC 2181 that any binary string can
	be used in the label of conventional/ traditional RRs in
	Class=IN, or by avoiding the normalization and mapping
	rules of IDNA.  Any problems that creates are between
	the zone administrators and applications developers or
	users who get burned.  The IETF is not involved.  ICANN
	might be, various national authorities might be, and
	assorted [other] lawyers might  be, but IETF is not.

	Some of the issues such actions would raise would
	certainly be relevant in looking at adoption and
	interoperability when someone tries to move IDNA to
	Draft Standard, but that isn't on the agenda right now
	either.

For the record, I personally consider local-zone rules about how 
DNS labels are interpreted as really stupid and an invitation to 
non-interoperability of applications, but, again, consensus (or 
not) in this group, on way or the other, isn't going to make 
that more or less true.

I suggest that, if you are serious about this, the usual IETF 
model is probably more useful than complaining on this list 
about paths not taken or paths that the WG tried (intentionally 
or not) to cut off).  Write a draft that goes into enough detail 
--both about what you propose and about how to make a transition 
from, or interoperate with, IDNA zones-- that it can be 
meaningfully evaluated.   Create a mailing list to discuss it. 
With the draft posted, try to organize a BOF or get some AD to 
create a WG directly, to explore the issues and see if a 
standards can be developed.   And so on.

But here?  My impression is that most of the people who watched 
or particpated in the WG are exhausted about the subject and 
really glad it reached _some_ conclusion and got documents out 
the door.  Those who really like IDNA will defend it, and some 
of us who still have misgivings about certain aspects of it 
would like to give it some time to see if it is adopted and 
workable in practice.  Those who are convinced that IDNA is a 
serious mistake, or that it will cut off important future 
evolution (and I have never been a member of either group) have, 
I think, mostly gone elsewhere.  Neither your ideas for more 
radical, in-DNS, approaches, nor mine ever got any traction in 
the WG and they are still less likely to get traction now, 
especially without solid drafts.

So I suggest taking it elsewhere, leaving this list for 
discussion of issues in implementation and deployment of IDNA.

regards,
     john







From owner-idn@ops.ietf.org  Wed Apr 30 10:57:53 2003
Received: from psg.com (mailnull@psg.com [147.28.0.62])
	by ietf.org (8.9.1a/8.9.1a) with ESMTP id KAA07412
	for <idn-archive@lists.ietf.org>; Wed, 30 Apr 2003 10:57:53 -0400 (EDT)
Received: from lserv by psg.com with local (Exim 3.36 #1)
	id 19AsrG-00010e-00
	for idn-data@psg.com; Wed, 30 Apr 2003 14:46:54 +0000
Received: from hosting.altserver.com ([209.124.80.2])
	by psg.com with esmtp (Exim 3.36 #1)
	id 19AsrD-00010R-00
	for idn@ops.ietf.org; Wed, 30 Apr 2003 14:46:51 +0000
Received: from f12m-4-220.d1.club-internet.fr ([212.195.67.220] helo=mine.jefsey.com)
	by hosting.altserver.com with esmtp (Exim 3.36 #1)
	id 19Asqv-0004YB-00; Wed, 30 Apr 2003 07:46:34 -0700
Message-Id: <5.2.0.9.0.20030430160318.0339b410@mail.jefsey.com>
X-Sender: jefsey+jefsey.com@mail.jefsey.com
X-Mailer: QUALCOMM Windows Eudora Version 5.2.0.9
Date: Wed, 30 Apr 2003 16:53:19 +0200
To: John C Klensin <klensin@jck.com>,
        Dan Oscarsson <Dan.Oscarsson@kiconsulting.se>,
        "idn@ops.ietf.org" <idn@ops.ietf.org>
From: "JFC (Jefsey) Morfin" <jefsey@jefsey.com>
Subject: Re: [idn] ToUnicode output can be longer than input
In-Reply-To: <86521110.1051693307@p3.JCK.COM>
References: <200304301221.h3UCLOt7008508@valinor.malmo.kicore.net>
 <200304301221.h3UCLOt7008508@valinor.malmo.kicore.net >
Mime-Version: 1.0
Content-Type: text/plain; charset="us-ascii"; format=flowed
X-AntiAbuse: This header was added to track abuse, please include it with any abuse report
X-AntiAbuse: Primary Hostname - hosting.altserver.com
X-AntiAbuse: Original Domain - ops.ietf.org
X-AntiAbuse: Originator/Caller UID/GID - [0 0] / [0 0]
X-AntiAbuse: Sender Address Domain - jefsey.com
X-Spam-Status: No, hits=-16.3 required=5.0
	tests=BAYES_01,EMAIL_ATTRIBUTION,IN_REP_TO
	autolearn=ham version=2.53
X-Spam-Checker-Version: SpamAssassin 2.53 (1.174.2.15-2003-03-30-exp)
Sender: owner-idn@ops.ietf.org
Precedence: bulk

Dear John and Oscar,
At 15:01 30/04/03, John C Klensin wrote:
>I suggest that, if you are serious about this, the usual IETF model is 
>probably more useful than complaining on this list about paths not taken 
>or paths that the WG tried (intentionally or not) to cut off).

We saw this approach has probably failed in spite of the efforts and the 
capacity of the involved persons. There is no interest to reuse it. May be 
interest in understanding why it did not work well?

>But here?  My impression is that most of the people who watched or 
>particpated in the WG are exhausted about the subject and really glad it 
>reached _some_ conclusion and got documents out the door.

True.

>Those who really like IDNA will defend it, and some of us who still have 
>misgivings about certain aspects of it would like to give it some time to 
>see if it is adopted and workable in practice.  Those who are convinced 
>that IDNA is a serious mistake, or that it will cut off important future 
>evolution (and I have never been a member of either group) have, I think, 
>mostly gone elsewhere.

I am probably one of them since I think the current solution is a 
bugged?  Sorry, I am still here.

>So I suggest taking it elsewhere, leaving this list for discussion of 
>issues in implementation and deployment of IDNA.

I fully agree. As long as IDNs were under IETF control the rest of the 
partners were blocked. Now the IDNA has been released, even if the proposed 
solution was(is) no good, it is A solution.

You wanted it? you got it? now help us making it work!
Don't disband, that would be too easy.

Now, what has been worked out in three years is a very small part of the 
problem. So it is no big deal to circumvent it if it is realy that bad (and 
may be to globally benefit from that effiort?).

1. wording is ununderstandable? ITU translated "internationalized" into 
"multinational" and others into "mutilateral". Habit takes off to note IDNs 
as "ML.ASCII" and MDNs as ML.ML". That is OK.

2. Texts are unreadable. It realy boils down to adding an European standard 
for "nasty danger" header (xn) to domain names of which non ascii 
characters have been transcoded in using Adam's funy code.

3. the choice of "xx--" as an header format instead of "x--x" creates a 
legally complex babelsquatting?. Great, the GNSO/IPC might help fixing it.

4. TLD Managers, WTO, Edifact, laws, Govs, are details "engineers" did not 
bother considering? May be this gives us more flexiblity in finishingthe 
work? Some are already trying to use it to make the IDNA bug a feature for 
the 5000 languages of the mankind (happily 20 are disapearing a year).


John, I fully agree with you. Let stop disputing about IDNA. Let help ".us" 
to support domain names after the proper spelling of the name of every US 
citizen (USA might be a country where this support might be made a law 
quick? Maybe could we read French law under voting that way?). Or let help 
scSLDs to do it, in an ML.ML format, for non roman T/SLDs.

Not three other years from now, just now.
jfc

























>regards,
>     john
>
>
>
>
>
>
>
>
>
>
>
>---
>Incoming mail is certified Virus Free.
>Checked by AVG anti-virus system (http://www.grisoft.com).
>Version: 6.0.474 / Virus Database: 272 - Release Date: 18/04/03




From owner-idn@ops.ietf.org  Wed Apr 30 22:07:34 2003
Received: from psg.com (mailnull@psg.com [147.28.0.62])
	by ietf.org (8.9.1a/8.9.1a) with ESMTP id WAA02902
	for <idn-archive@lists.ietf.org>; Wed, 30 Apr 2003 22:07:34 -0400 (EDT)
Received: from lserv by psg.com with local (Exim 3.36 #1)
	id 19B3KF-000JaC-00
	for idn-data@psg.com; Thu, 01 May 2003 01:57:31 +0000
Received: from arwen.cs.berkeley.edu ([128.32.132.165] helo=nicemice.net)
	by psg.com with esmtp (Exim 3.36 #1)
	id 19B3KC-000Ja0-00
	for idn@ops.ietf.org; Thu, 01 May 2003 01:57:28 +0000
Received: from amc by nicemice.net with local (Exim 3.35 #1 (Debian))
	id 19B3KB-0002g6-00
	for <idn@ops.ietf.org>; Wed, 30 Apr 2003 18:57:27 -0700
Date: Thu, 1 May 2003 01:57:27 +0000
From: "Adam M. Costello" <idn.amc+0@nicemice.net.RemoveThisWord>
To: idn@ops.ietf.org
Subject: Re: [idn] ToUnicode output can be longer than input
Message-ID: <20030501015727.GD7482@nicemice.net>
Reply-To: IETF idn working group <idn@ops.ietf.org>
References: <200304301221.h3UCLOt7008508@valinor.malmo.kicore.net>
Mime-Version: 1.0
Content-Type: text/plain; charset=us-ascii
Content-Disposition: inline
In-Reply-To: <200304301221.h3UCLOt7008508@valinor.malmo.kicore.net>
User-Agent: Mutt/1.4i
X-Spam-Status: No, hits=-38.8 required=5.0
	tests=BAYES_01,EMAIL_ATTRIBUTION,IN_REP_TO,QUOTED_EMAIL_TEXT,
	      REFERENCES,REPLY_WITH_QUOTES,USER_AGENT_MUTT
	autolearn=ham version=2.53
X-Spam-Checker-Version: SpamAssassin 2.53 (1.174.2.15-2003-03-30-exp)
Sender: owner-idn@ops.ietf.org
Precedence: bulk

Dan Oscarsson <Dan.Oscarsson@kiconsulting.se> wrote:

> IDNA defines a way to compare names in ASCII context because it
> requires names to be in IDNA ACE format.

It requires names to be in ASCII format in IDN-unaware contexts.  It
does not require names to be in ASCII format when they are compared.  It
says:

    Whenever two labels are compared, they MUST be considered to match
    if and only if they are equivalent, that is, their ASCII forms
    (obtained by applying ToASCII) match using a case-insensitive ASCII
    comparison.

That doesn't say you must compare the ASCII forms, it says you must
reach the same answer as if you compared the ASCII forms.  And the rule
doesn't say it applies only in certain contexts, it says "whenever"
two labels are compared.  If that's not clear enough, the point is
underscored in the introduction:

    Applications can also define protocols and interfaces that support
    IDNs directly using non-ASCII representations.  IDNA does not
    prescribe any particular representation for new protocols, but it
    still defines which names are valid and how they are compared.

> Comparing names in an international context must be done using UCS
> characters directly.

I assume that's your opinion.  It's certainly not a requirement of any
standard.

IDNA allows you to perform the comparison any way you
like, provided you get the right answer.  For example,
given any two valid internationalized labels X and Y,
tolower(ToASCII(X)) == tolower(ToASCII(Y)) iff nameprep(ToUnicode(X)) ==
nameprep(ToUnicode(Y)), so you can use either form as a canonical form
for comparisons.

AMC



