From owner-idn@ops.ietf.org  Thu Apr  7 22:02:02 2005
Received: from psg.com (mailnull@psg.com [147.28.0.62])
	by ietf.org (8.9.1a/8.9.1a) with ESMTP id WAA29743
	for <idn-archive@lists.ietf.org>; Thu, 7 Apr 2005 22:02:02 -0400 (EDT)
Received: from majordom by psg.com with local (Exim 4.44 (FreeBSD))
	id 1DJiiX-000GIT-QY
	for idn-data@psg.com; Fri, 08 Apr 2005 01:55:29 +0000
Received: from [207.115.63.98] (helo=pimout4-ext.prodigy.net)
	by psg.com with esmtp (Exim 4.44 (FreeBSD))
	id 1DJiiW-000GIA-Ku
	for idn@ops.ietf.org; Fri, 08 Apr 2005 01:55:28 +0000
Received: from [10.1.1.2] (adsl-64-174-147-206.dsl.sntc01.pacbell.net [64.174.147.206])
	by pimout4-ext.prodigy.net (8.12.10 milter /8.12.10) with ESMTP id j381tL5K222276;
	Thu, 7 Apr 2005 21:55:21 -0400
Message-ID: <4255E488.8010302@vanderpoel.org>
Date: Thu, 07 Apr 2005 18:55:20 -0700
From: Erik van der Poel <erik@vanderpoel.org>
User-Agent: Mozilla Thunderbird 1.0.2 (X11/20050317)
X-Accept-Language: en-us, en
MIME-Version: 1.0
To: Soobok Lee <lsb@lsb.org>
CC: idn@ops.ietf.org
Subject: Re: [idn] space-like unicode char
References: <42181FD5.3070608@lsb.org>
In-Reply-To: <42181FD5.3070608@lsb.org>
Content-Type: text/plain; charset=ISO-8859-1; format=flowed
Content-Transfer-Encoding: 7bit
X-Spam-Checker-Version: SpamAssassin 3.0.1 (2004-10-22) on psg.com
X-Spam-Status: No, score=-2.6 required=5.0 tests=AWL,BAYES_00 autolearn=ham 
	version=3.0.1
Sender: owner-idn@ops.ietf.org
Precedence: bulk
Content-Transfer-Encoding: 7bit

Soobok Lee wrote:
> U+1160 is a space-like char and even stringprep/nameprep does not
> filter it out because the char is not for punctuational purpose.

U+1160 is HANGUL JUNGSEONG FILLER and it is used to transform 
nonstandard syllables into standard ones (Unicode 3.0 section 3.11 (RFC 
3454 refers to Unicode 3.2.0)). However, this transformation is one of 
the additional transformations not considered part of Unicode 
normalization (3.2.0's UAX #15 Annex 10). So this character is not 
generated by Stringprep/Nameprep.

However, it is not prohibited either, so it may occur in the input to 
(and output from) Stringprep/Nameprep. I read some of the sections on 
Hangul in the Unicode book and Web site, but I did not see any rules 
regarding repeated occurrences of U+1160 (as you had in your example, 
not quoted above). I also did not see any rules about what to do when a 
filler is not followed by a Hangul jamo. It would be nice to have these 
rules in Unicode or in Stringprep.

I tried U+1160 followed by a Latin character in MSIE with i-Nav and in 
Firefox with IDN turned on, and it was displayed as a wide space. It is 
unfortunate that both implementations chose to display it as a space 
instead of deleting it.

Erik



From owner-idn@ops.ietf.org  Fri Apr  8 03:11:39 2005
Received: from psg.com (mailnull@psg.com [147.28.0.62])
	by ietf.org (8.9.1a/8.9.1a) with ESMTP id DAA12263
	for <idn-archive@lists.ietf.org>; Fri, 8 Apr 2005 03:11:39 -0400 (EDT)
Received: from majordom by psg.com with local (Exim 4.44 (FreeBSD))
	id 1DJnYH-0000Fq-Tb
	for idn-data@psg.com; Fri, 08 Apr 2005 07:05:13 +0000
Received: from [211.196.150.53] (helo=postel5.postel.co.kr)
	by psg.com with esmtp (Exim 4.44 (FreeBSD))
	id 1DJnYF-0000FO-De
	for idn@ops.ietf.org; Fri, 08 Apr 2005 07:05:11 +0000
Received: from [10.1.1.21] ([61.73.48.22])
	by postel5.postel.co.kr (8.13.0.PreAlpha4/8.13.0.PreAlpha4) with ESMTP id j38758JR024364;
	Fri, 8 Apr 2005 16:05:08 +0900
Message-ID: <42562D22.3090609@lsb.org>
Date: Fri, 08 Apr 2005 16:05:06 +0900
From: Soobok Lee <lsb@lsb.org>
User-Agent: Mozilla Thunderbird 1.0 (Windows/20041206)
X-Accept-Language: en-us, en
MIME-Version: 1.0
To: Erik van der Poel <erik@vanderpoel.org>
CC: idn@ops.ietf.org
Subject: Re: [idn] space-like unicode char
References: <42181FD5.3070608@lsb.org> <4255E488.8010302@vanderpoel.org>
In-Reply-To: <4255E488.8010302@vanderpoel.org>
Content-Type: text/plain; charset=EUC-KR
Content-Transfer-Encoding: 7bit
X-Spam-Checker-Version: SpamAssassin 3.0.1 (2004-10-22) on psg.com
X-Spam-Status: No, score=-2.6 required=5.0 tests=BAYES_00 autolearn=ham 
	version=3.0.1
Sender: owner-idn@ops.ietf.org
Precedence: bulk
Content-Transfer-Encoding: 7bit

Erik van der Poel wrote:

> Soobok Lee wrote:
>
>> U+1160 is a space-like char and even stringprep/nameprep does not
>> filter it out because the char is not for punctuational purpose.
>
>
> U+1160 is HANGUL JUNGSEONG FILLER and it is used to transform
> nonstandard syllables into standard ones (Unicode 3.0 section 3.11
> (RFC 3454 refers to Unicode 3.2.0)). However, this transformation is
> one of the additional transformations not considered part of Unicode
> normalization (3.2.0's UAX #15 Annex 10). 

Exactly. U+1160 is not "touched" by Unicode normalization (NFC).

> So this character is not generated by Stringprep/Nameprep.However, it
> is not prohibited either, so it may occur in the input to (and output
> from) Stringprep/Nameprep.

Yes, it may occur.

> I read some of the sections on Hangul in the Unicode book and Web
> site, but I did not see any rules regarding repeated occurrences of
> U+1160 (as you had in your example, not quoted above). I also did not
> see any rules about what to do when a filler is not followed by a
> Hangul jamo. It would be nice to have these rules in Unicode or in
> Stringprep.

U+1160 problem has been raised 3.5 years ago (you can look into this
huge idn-list archive by keyword search for 1160 or filler)
with some additional hangul jamo problem. One draft has been submitted
by me (you may find that in www.i-d-n.net)
to filter out these invalid char sequences. But the draft had been
discarded . Someone argued that such filtering * complicates *
stringprep algorithms with context-sensitive filtering/prohibiting and
the problem is up to UTC/NFC not to IETF. of course, i couldn't accept that.

Anyway, we can't backtrack into 2002/Dec without giving up backward
compatibility promise of stringprep.


>
> I tried U+1160 followed by a Latin character in MSIE with i-Nav and in
> Firefox with IDN turned on, and it was displayed as a wide space. It
> is unfortunate that both implementations chose to display it as a
> space instead of deleting it.

Yes. Plugins M U S T filter out U+1160 from validated ToUnicode()ed
labels, whether or not IDNA requires that.

Soobok




From owner-idn@ops.ietf.org  Fri Apr  8 03:23:48 2005
Received: from psg.com (mailnull@psg.com [147.28.0.62])
	by ietf.org (8.9.1a/8.9.1a) with ESMTP id DAA13098
	for <idn-archive@lists.ietf.org>; Fri, 8 Apr 2005 03:23:47 -0400 (EDT)
Received: from majordom by psg.com with local (Exim 4.44 (FreeBSD))
	id 1DJno4-0002BI-Uc
	for idn-data@psg.com; Fri, 08 Apr 2005 07:21:32 +0000
Received: from [211.196.150.53] (helo=postel5.postel.co.kr)
	by psg.com with esmtp (Exim 4.44 (FreeBSD))
	id 1DJno3-0002Al-QM
	for idn@ops.ietf.org; Fri, 08 Apr 2005 07:21:32 +0000
Received: from [10.1.1.21] ([61.73.48.22])
	by postel5.postel.co.kr (8.13.0.PreAlpha4/8.13.0.PreAlpha4) with ESMTP id j387LTJR025641;
	Fri, 8 Apr 2005 16:21:29 +0900
Message-ID: <425630F8.5030204@lsb.org>
Date: Fri, 08 Apr 2005 16:21:28 +0900
From: Soobok Lee <lsb@lsb.org>
User-Agent: Mozilla Thunderbird 1.0 (Windows/20041206)
X-Accept-Language: en-us, en
MIME-Version: 1.0
To: Soobok Lee <lsb@lsb.org>
CC: Erik van der Poel <erik@vanderpoel.org>, idn@ops.ietf.org
Subject: [idn] combining marks and  space-like unicode char 
References: <42181FD5.3070608@lsb.org> <4255E488.8010302@vanderpoel.org> <42562D22.3090609@lsb.org>
In-Reply-To: <42562D22.3090609@lsb.org>
Content-Type: text/plain; charset=EUC-KR
Content-Transfer-Encoding: 7bit
X-Spam-Checker-Version: SpamAssassin 3.0.1 (2004-10-22) on psg.com
X-Spam-Status: No, score=-2.6 required=5.0 tests=BAYES_00 autolearn=ham 
	version=3.0.1
Sender: owner-idn@ops.ietf.org
Precedence: bulk
Content-Transfer-Encoding: 7bit



>>I tried U+1160 followed by a Latin character in MSIE with i-Nav and in
>>Firefox with IDN turned on, and it was displayed as a wide space. It
>>is unfortunate that both implementations chose to display it as a
>>space instead of deleting it.
>>    
>>
>
>Yes. Plugins M U S T filter out U+1160 from validated ToUnicode()ed
>labels, whether or not IDNA requires that.
>
>Soobok
>
I will add this: In standard hangul writing system,
U+1160 is meaningful only in some context (surrounded by at least one
jamo char).
But, is standalone U+1160 is illegal ? No, it is NOT illegal.

So, blind filtering of U+1160 is fault. Plugins' filtering should be
context-sensitive.
That is why it would complicate stringprep if it were included into
stringprep. :-)

We can find similar problems in "combining diacritical marks" (U+3xx).
What if
a label with single char 'combining accent or above-dot ' without any
preceding
alphabet? It will combine with its preceding dot delimiter. and that
will produce
confusing looks ( looks like a colon which is a protocol delimiter).

AFAIK, any single standalone combining accent char is not prohibited by
stringprep.

Sooobk



From owner-idn@ops.ietf.org  Fri Apr  8 15:44:18 2005
Received: from psg.com (mailnull@psg.com [147.28.0.62])
	by ietf.org (8.9.1a/8.9.1a) with ESMTP id PAA01382
	for <idn-archive@lists.ietf.org>; Fri, 8 Apr 2005 15:44:18 -0400 (EDT)
Received: from majordom by psg.com with local (Exim 4.44 (FreeBSD))
	id 1DJzHe-000Fso-26
	for idn-data@psg.com; Fri, 08 Apr 2005 19:36:50 +0000
Received: from [207.115.63.98] (helo=pimout4-ext.prodigy.net)
	by psg.com with esmtp (Exim 4.44 (FreeBSD))
	id 1DJzHa-000FsS-JV
	for idn@ops.ietf.org; Fri, 08 Apr 2005 19:36:46 +0000
Received: from [10.1.1.2] (adsl-64-174-147-206.dsl.sntc01.pacbell.net [64.174.147.206])
	by pimout4-ext.prodigy.net (8.12.10 milter /8.12.10) with ESMTP id j38JaP5K214294;
	Fri, 8 Apr 2005 15:36:25 -0400
Message-ID: <4256DD38.3070708@vanderpoel.org>
Date: Fri, 08 Apr 2005 12:36:24 -0700
From: Erik van der Poel <erik@vanderpoel.org>
User-Agent: Mozilla Thunderbird 1.0.2 (X11/20050317)
X-Accept-Language: en-us, en
MIME-Version: 1.0
To: Soobok Lee <lsb@lsb.org>
CC: idn@ops.ietf.org
Subject: Re: [idn] space-like unicode char
References: <42181FD5.3070608@lsb.org> <4255E488.8010302@vanderpoel.org> <42562D22.3090609@lsb.org>
In-Reply-To: <42562D22.3090609@lsb.org>
Content-Type: text/plain; charset=ISO-8859-1; format=flowed
Content-Transfer-Encoding: 7bit
X-Spam-Checker-Version: SpamAssassin 3.0.1 (2004-10-22) on psg.com
X-Spam-Status: No, score=-2.6 required=5.0 tests=AWL,BAYES_00 autolearn=ham 
	version=3.0.1
Sender: owner-idn@ops.ietf.org
Precedence: bulk
Content-Transfer-Encoding: 7bit

Soobok Lee wrote:
> U+1160 problem has been raised 3.5 years ago (you can look into this
> huge idn-list archive by keyword search for 1160 or filler)
> with some additional hangul jamo problem. One draft has been submitted
> by me (you may find that in www.i-d-n.net)
> to filter out these invalid char sequences. But the draft had been
> discarded . Someone argued that such filtering * complicates *
> stringprep algorithms with context-sensitive filtering/prohibiting and
> the problem is up to UTC/NFC not to IETF. of course, i couldn't accept that.

The i-d-n.net name no longer takes you to a real site, but I believe I 
found your draft here:

http://www.watersprings.org/pub/id/draft-ietf-idn-hangeulchar-00.txt

I agree that the U+1160 issues would complicate a spec, and I can see 
why the IETF decided not to include them in the RFCs, but now that we 
have seen that a number of implementations display this character in a 
potentially dangerous way, we should reconsider the specs.

Unicode may not be able to address these issues in the normalization 
spec since they have promised not to make any incompatible changes. 
Unicode might be able to address the issues in other normative or 
informative parts of their book or documents, and the IETF might just 
want to refer to those parts of Unicode.

Alternatively, the IETF can write up its own specifications or 
recommendations. It's not immediately clear to me whether U+1160 ought 
to be addressed in Stringprep or Nameprep. As we have seen, Stringprep 
is used in various protocols, including SASLprep, which is for user 
names and passwords. Some perverse people might suggest that passwords 
ought to allow strange character sequences like multiple consecutive 
U+1160s in order to make it harder to guess the password. I'm new to 
Stringprep, so I don't know how most IETFers feel about this type of thing.

In the meantime, I have added U+1160 and the combining mark issue to my 
list and I have filed a bug report for Mozilla:

http://nameprep.org/#display
https://bugzilla.mozilla.org/show_bug.cgi?id=289588

Erik



