
From nobody Wed Aug  3 09:32:14 2016
Return-Path: <jroatch@gmail.com>
X-Original-To: cbor@ietfa.amsl.com
Delivered-To: cbor@ietfa.amsl.com
Received: from localhost (localhost [127.0.0.1]) by ietfa.amsl.com (Postfix) with ESMTP id C131712D099 for <cbor@ietfa.amsl.com>; Wed,  3 Aug 2016 09:32:12 -0700 (PDT)
X-Virus-Scanned: amavisd-new at amsl.com
X-Spam-Flag: NO
X-Spam-Score: -2.7
X-Spam-Level: 
X-Spam-Status: No, score=-2.7 tagged_above=-999 required=5 tests=[BAYES_00=-1.9, DKIM_SIGNED=0.1, DKIM_VALID=-0.1, DKIM_VALID_AU=-0.1, FREEMAIL_FROM=0.001, RCVD_IN_DNSWL_LOW=-0.7, SPF_PASS=-0.001] autolearn=ham autolearn_force=no
Authentication-Results: ietfa.amsl.com (amavisd-new); dkim=pass (2048-bit key) header.d=gmail.com
Received: from mail.ietf.org ([4.31.198.44]) by localhost (ietfa.amsl.com [127.0.0.1]) (amavisd-new, port 10024) with ESMTP id EGPkzPGHo5DA for <cbor@ietfa.amsl.com>; Wed,  3 Aug 2016 09:32:08 -0700 (PDT)
Received: from mail-qk0-x22a.google.com (mail-qk0-x22a.google.com [IPv6:2607:f8b0:400d:c09::22a]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by ietfa.amsl.com (Postfix) with ESMTPS id 6C09812D151 for <cbor@ietf.org>; Wed,  3 Aug 2016 09:32:08 -0700 (PDT)
Received: by mail-qk0-x22a.google.com with SMTP id x185so14753092qkc.2 for <cbor@ietf.org>; Wed, 03 Aug 2016 09:32:08 -0700 (PDT)
DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20120113;  h=message-id:subject:from:to:cc:date:in-reply-to:references :mime-version:content-transfer-encoding; bh=gCvHgQwMgttNFHBAWi4Ij3buvALkRm4dTAnPA606Hz0=; b=AqDaFkTeHe0SosV5Kwy7QdHvGc3z1gpHOupj19YYfASNcIGV6KEORU88tw1b61IX6M 29qJBBQgdMVqEcMnSHyY1jKV+csB+LLUVKA6VNbKiyQriz88ALkoFLtd4BDl1MLMjHUv 3j/pweLGhTNL97CSvS6d2TlAdmrppErB/cfy3a1ctvtbxhCcxRVJLs7+nhK6r0ed0Wds DUP4pMwCuXWjp/3VD2VL8Wzp02PWIW0DuqDuXRWRlyh60jMgGQybVyOQzy2Ii2icfTqt PPehDd/q41aqor6FYiNKDhKpuqW+1BcjqAymxPreog5NOWbJM0NpaoUUr+QhrmUg3KfP 94WQ==
X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20130820; h=x-gm-message-state:message-id:subject:from:to:cc:date:in-reply-to :references:mime-version:content-transfer-encoding; bh=gCvHgQwMgttNFHBAWi4Ij3buvALkRm4dTAnPA606Hz0=; b=hFr41CobNDse581Xks1jNcBWwNmh/pmZy68FAKRwcIJFxVwclqKBvUGl2iFQcEWTbb d+cy9HmQsQtkeW9cW+nLFaPEXEbREgia9VRsAz5cD5zsRbBjQMEvIHe2U+oLf8TxhB3m DT7W9TpPZXiQ3wtGz5Me5lBFD5YY7vm+VFP5N8NsWy/xQ0BhdQ7y63SPQgtsMNVuSf6q czNRRz7W+P8L9Jfv9nNWwrMDEjoL8K9cDISxhPGVrF7TsHKXygg0R/QUR6YxBElY7NvE y7PPbgERA/dYhJwojqi/wZ9zB6d+ZlL0aD+8Y4nY8qp+IB4gisf48zzLdyAyFtQsIfgZ vjFw==
X-Gm-Message-State: AEkooutxad2ScT5WgOIqfIngu31W5iy1u4Q/v+75l6Ed/+h5vfQ43KFiR26KqtngLB0yEw==
X-Received: by 10.55.201.200 with SMTP id m69mr841765qkl.235.1470241927480; Wed, 03 Aug 2016 09:32:07 -0700 (PDT)
Received: from mindcandy ([172.79.206.138]) by smtp.googlemail.com with ESMTPSA id t1sm4559916qtt.25.2016.08.03.09.32.06 (version=TLS1_2 cipher=ECDHE-RSA-CHACHA20-POLY1305 bits=256/256); Wed, 03 Aug 2016 09:32:06 -0700 (PDT)
Message-ID: <1470241923.1200.99.camel@gmail.com>
From: Johnathan Roatch <jroatch@gmail.com>
To: Sean Leonard <dev+ietf@seantek.com>, Carsten Bormann <cabo@tzi.org>,  glenn_engel@keysight.com
Date: Wed, 03 Aug 2016 12:32:03 -0400
In-Reply-To: <3306456d-17d4-dfbb-9940-e304f5f73c61@seantek.com>
References: <04EFF12F483FA149B07653989B86861F217A069F@wcosexch01k.cos.is.keysight.com> <577416B1.3090308@tzi.org> <577ED54A.3050705@tzi.org> <3306456d-17d4-dfbb-9940-e304f5f73c61@seantek.com>
Content-Type: text/plain; charset="UTF-8"
X-Mailer: Evolution 3.20.4 
Mime-Version: 1.0
Content-Transfer-Encoding: 8bit
Archived-At: <https://mailarchive.ietf.org/arch/msg/cbor/QUECChtkVnSWuEPQCHIOinFIU6w>
Cc: cbor@ietf.org, draft-jroatch-cbor-tags@tools.ietf.org
Subject: Re: [Cbor] draft-jroatch-cbor-tags-04
X-BeenThere: cbor@ietf.org
X-Mailman-Version: 2.1.17
Precedence: list
List-Id: "Concise Binary Object Representation \(CBOR\)" <cbor.ietf.org>
List-Unsubscribe: <https://www.ietf.org/mailman/options/cbor>, <mailto:cbor-request@ietf.org?subject=unsubscribe>
List-Archive: <https://mailarchive.ietf.org/arch/browse/cbor/>
List-Post: <mailto:cbor@ietf.org>
List-Help: <mailto:cbor-request@ietf.org?subject=help>
List-Subscribe: <https://www.ietf.org/mailman/listinfo/cbor>, <mailto:cbor-request@ietf.org?subject=subscribe>
X-List-Received-Date: Wed, 03 Aug 2016 16:32:13 -0000

On Thu, 2016-07-28 at 18:09 -0700, Sean Leonard wrote:
> Section 2. Typed Arrays
> 
> Okay; makes sense. However, the mismatch between [u/s]int8 and
> binary16 
> is jarring. It would seem a lot better if 
> uint16/sint16/binary16..uint128/sint128/binary128.

When I initially made up the scheme, my primary goal was to pack in any
possible gaps. I agree now that byte sizes should definitely be
aligned, especially to include 128-bit integers.

> In contrast, I don't see a lot of value to tagging an array as
> uint8...that would be called a plain old "byte string" by any
> definition of the term.

Good insight here. I confused myself thinking that a JavaScript
ArrayBuffer could been used without a uint8 view. The uint8 tag can
definitely be dropped to make room for byte size alignment.
Implementations should then use un-tagged byte strings for uint8
arrays.

> For the case of uint8 and sint8, I propose that you allocate four
> tags in approximately 0b001111_c_s, where c = clamped, and s =
> signed. 

If the uint8 is gone then all we have left here is sint8. clamped
arithmetic isn't as critical to fit into some bit pattern as it's
really just a hint on how to process numbers and not a number format on
it's own. Now that I think of it, clamped arithmetic can actually be
it's own tag that can apply to any integer or integer array.

> If you want parity with the bit positions in the other tags, I would
> propose rearranging the 0b010_f_s_e_ll to: 0b010_ll_f_e_s

I put float and signed together to form the collective bit field of
number type, and those bits were made the most significant so that
there would be less gaps in the list of tags. We have 3 types of
numbers in a field that can hold 4. Maybe a fourth type of number
should fill in those gaps.

Are the IEEE 754-2008 decimal formats used enough to include them?

After playing around with possible arrangements a bit this would be
what I would do: 0b01_lll_f_s_e.
"lll" is the used as "2**(lll)" for calculating the number of bytes.
All tags of a single byte size, except sint8, are unreserved. This will
make is simple to compute, as all it would take on most platforms is a
bit-mask and a couple shifts. the unreserved space will also be a good
place to put related tags such as Homogeneous and Multi-dimensional
Arrays.
"f_s" is the number format: [uintN, sintN, IEEE 754 binaryN, IEEE
754 decimalN].
"e" is big or little endianness.

In this formulation, uint8 is no tag, uint16-LE is tag 0x49, IEEE 754
binary32-LE is tag 0x55, uint128-LE is 0x61, and I guess sint8 can be
tag 0x42.

> I would also like to suggest using tags in the 2-byte space, such as 
> 0x0220-0x0237 (or 0x021C - 0x0237). 

Since there's already implementations that follow the old drafts, going
to the 2 byte tag space might be necessary for backwards compatibility
of those implementations.

If we are going to reuse the original range, now is the time to have a
breaking change.

> I suppose that you have thought of this already. I am less concerned
> about low-byte "pollution"; 

About "low-byte pollution", I've heard criticism about how endianness
takes up too many tags. While I thought about separating out endianness
to it's own tag, it's doesn't make much sense outside of the context of
typed arrays.

If we include IEEE 754-2008 decimal while fixing up the byte size
aliment as I wrote above, the total amount of tags will increase from
24 to 32. Forcing everything to be little-endian will reduce that
amount to 16. It really depends on how important having big endianness
is.

> it has more to do with the need for expansion of typed arrays in the
> future. If you want to add binary256, int256, etc.,

For now, I think we should limit ourselves to 128 bit arrays.

> Section 7. Security Considerations
> 
> ...
> 
> I consider an irregular-length array to be a coding error; that is
> with my security hat on.

Yes.

Would it be right to say that in the case of such a coding error, the
implementation should either terminate processing or treat the byte
string as if there were no TypeArray tag applied to it (as uint8)?

-- 
Johnathan Roatch <jroatch@gmail.com>
https://jroatch.nfshost.com/


From nobody Wed Aug  3 15:22:13 2016
Return-Path: <cabo@tzi.org>
X-Original-To: cbor@ietfa.amsl.com
Delivered-To: cbor@ietfa.amsl.com
Received: from localhost (localhost [127.0.0.1]) by ietfa.amsl.com (Postfix) with ESMTP id 8683D12D8BE for <cbor@ietfa.amsl.com>; Wed,  3 Aug 2016 15:22:12 -0700 (PDT)
X-Virus-Scanned: amavisd-new at amsl.com
X-Spam-Flag: NO
X-Spam-Score: -4.199
X-Spam-Level: 
X-Spam-Status: No, score=-4.199 tagged_above=-999 required=5 tests=[BAYES_00=-1.9, HTML_MESSAGE=0.001, RCVD_IN_DNSWL_MED=-2.3] autolearn=ham autolearn_force=no
Received: from mail.ietf.org ([4.31.198.44]) by localhost (ietfa.amsl.com [127.0.0.1]) (amavisd-new, port 10024) with ESMTP id DLlVxqTh6lCu for <cbor@ietfa.amsl.com>; Wed,  3 Aug 2016 15:22:09 -0700 (PDT)
Received: from mailhost.informatik.uni-bremen.de (mailhost.informatik.uni-bremen.de [IPv6:2001:638:708:30c9::12]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by ietfa.amsl.com (Postfix) with ESMTPS id 5E06512D8A3 for <cbor@ietf.org>; Wed,  3 Aug 2016 15:22:09 -0700 (PDT)
X-Virus-Scanned: amavisd-new at informatik.uni-bremen.de
Received: from submithost.informatik.uni-bremen.de (submithost.informatik.uni-bremen.de [134.102.201.11]) by mailhost.informatik.uni-bremen.de (8.14.5/8.14.5) with ESMTP id u73MM4Lv025637; Thu, 4 Aug 2016 00:22:04 +0200 (CEST)
Received: from nar-3.local.mail (p5DC7E34C.dip0.t-ipconnect.de [93.199.227.76]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by submithost.informatik.uni-bremen.de (Postfix) with ESMTPSA id 3s4SGz4SZ8zDCjQ; Thu,  4 Aug 2016 00:22:03 +0200 (CEST)
Date: Thu, 4 Aug 2016 00:22:02 +0200
From: Carsten Bormann <cabo@tzi.org>
To: glenn_engel@keysight.com, Johnathan Roatch <jroatch@gmail.com>, Sean Leonard <dev+ietf@seantek.com>
Message-ID: <etPan.57a26e8a.6838aa10.385a@tzi.org>
In-Reply-To: <etPan.57a26cbc.4216f5a2.385a@AirmailxGenerated.am>
References: <04EFF12F483FA149B07653989B86861F217A069F@wcosexch01k.cos.is.keysight.com> <577416B1.3090308@tzi.org> <577ED54A.3050705@tzi.org> <3306456d-17d4-dfbb-9940-e304f5f73c61@seantek.com> <etPan.57a26cbc.4216f5a2.385a@AirmailxGenerated.am>
X-Mailer: Airmail (382)
MIME-Version: 1.0
Content-Type: multipart/alternative; boundary="57a26e8a_4dadee93_385a"
Archived-At: <https://mailarchive.ietf.org/arch/msg/cbor/cqXDtx-nSXXW7gffXwuKDFvOuzo>
Cc: cbor@ietf.org, draft-jroatch-cbor-tags@tools.ietf.org
Subject: Re: [Cbor] draft-jroatch-cbor-tags-04
X-BeenThere: cbor@ietf.org
X-Mailman-Version: 2.1.17
Precedence: list
List-Id: "Concise Binary Object Representation \(CBOR\)" <cbor.ietf.org>
List-Unsubscribe: <https://www.ietf.org/mailman/options/cbor>, <mailto:cbor-request@ietf.org?subject=unsubscribe>
List-Archive: <https://mailarchive.ietf.org/arch/browse/cbor/>
List-Post: <mailto:cbor@ietf.org>
List-Help: <mailto:cbor-request@ietf.org?subject=help>
List-Subscribe: <https://www.ietf.org/mailman/listinfo/cbor>, <mailto:cbor-request@ietf.org?subject=subscribe>
X-List-Received-Date: Wed, 03 Aug 2016 22:22:12 -0000

--57a26e8a_4dadee93_385a
Content-Type: text/plain; charset="utf-8"
Content-Transfer-Encoding: quoted-printable
Content-Disposition: inline

Hmm, I=E2=80=99m not sure I agree with either the criticism or the conclu=
sions below.
As the draft stands, it is closely mirroring the typed array functionalit=
y that has been defined in several environments including JavaScript. =C2=
=A0Why would we be chasing some form of symmetry instead of following wha=
t the users seem to need=3F

uint8/sint8 arrays are arrays of integers, which is different from untagg=
ed byte strings. =C2=A0Hence the tags.
I=E2=80=99m not aware of a use case for IEEE 754 decimal at this point. =C2=
=A0Sure, one can be constructed, but the objective is not to cover everyt=
hing that has been invented under the sun, but what is actually needed.

In summary, I continue to believe that the draft is all good as it stands=
; nothing should be changed.

Gr=C3=BC=C3=9Fe, Carsten

On 3 August 2016 at 18:57:07, Johnathan Roatch (jroatch=40gmail.com) wrot=
e:

On Thu, 2016-07-28 at 18:09 -0700, Sean Leonard wrote: =20
> Section 2. Typed Arrays =20
> =20
> Okay; makes sense. However, the mismatch between =5Bu/s=5Dint8 and =20
> binary16=C2=A0 =20
> is jarring. It would seem a lot better if=C2=A0 =20
> uint16/sint16/binary16..uint128/sint128/binary128. =20

When I initially made up the scheme, my primary goal was to pack in any =20
possible gaps. I agree now that byte sizes should definitely be =20
aligned, especially to include 128-bit integers. =20

> In contrast, I don't see a lot of value to tagging an array as =20
> uint8...that would be called a plain old =22byte string=22 by any =20
> definition=C2=A0of the term. =20

Good insight here. I confused myself thinking that a JavaScript =20
ArrayBuffer could been used without a uint8 view. The uint8 tag can =20
definitely be dropped to make room for byte size alignment. =20
Implementations should then use un-tagged byte strings for uint8 =20
arrays. =20

> =46or the case of uint8 and sint8, I propose that you allocate four =20
> tags=C2=A0in approximately 0b001111=5Fc=5Fs, where c =3D clamped, and s=
 =3D =20
> signed. =20

If the uint8 is gone then all we have left here is sint8. clamped =20
arithmetic isn't as critical to fit into some bit pattern as it's =20
really just a hint on how to process numbers and not a number format on =20
it's own. Now that I think of it, clamped arithmetic can actually be =20
it's own tag that can apply to any integer or integer array. =20

> If you want parity with the bit positions in the other tags, I would =20
> propose rearranging the 0b010=5Ff=5Fs=5Fe=5Fll to: 0b010=5Fll=5Ff=5Fe=5F=
s =20

I put float and signed together to form the collective bit field of =20
number type, and those bits were made the most significant so that =20
there would be less gaps in the list of tags. We have 3 types of =20
numbers in a field that can hold 4. Maybe a fourth type of number =20
should fill in those gaps. =20

Are the IEEE 754-2008 decimal formats used enough to include them=3F =20

After playing around with possible arrangements a bit this would be =20
what I would do: 0b01=5Flll=5Ff=5Fs=5Fe. =20
=22lll=22 is the used as =222**(lll)=22 for calculating the number of byt=
es. =20
All tags of a single byte size, except sint8, are unreserved. This will =20
make is simple to compute, as all it would take on most platforms is a =20
bit-mask and a couple shifts. the unreserved space will also be a good =20
place to put related tags such as Homogeneous and Multi-dimensional =20
Arrays. =20
=22f=5Fs=22 is the number format: =5BuintN, sintN, IEEE 754 binaryN,=C2=A0=
IEEE =20
754=C2=A0decimalN=5D. =20
=22e=22 is big or little endianness. =20

In this formulation, uint8 is no tag, uint16-LE is tag 0x49,=C2=A0IEEE 75=
4 =20
binary32-LE is tag 0x55, uint128-LE is 0x61, and I guess sint8 can be =20
tag 0x42. =20

> I would also like to suggest using tags in the 2-byte space, such as=C2=
=A0 =20
> 0x0220-0x0237 (or 0x021C - 0x0237). =20

Since there's already implementations that follow the old drafts, going =20
to the 2 byte tag space might be necessary for backwards compatibility =20
of those implementations. =20

If we are going to reuse the original range, now is the time to have a =20
breaking change. =20

> I suppose that you have thought of this already. I am less concerned =20
> about low-byte =22pollution=22; =20

About =22low-byte pollution=22, I've heard criticism about how endianness=
 =20
takes up too many tags. While I thought about separating out endianness =20
to it's own tag, it's doesn't make much sense outside of the context of =20
typed arrays. =20

If we include IEEE 754-2008 decimal while fixing up the byte size =20
aliment as I wrote above, the total amount of tags will increase from =20
24 to 32. =46orcing everything to be little-endian will reduce that =20
amount to 16. It really depends on how important having big endianness =20
is. =20

> it has more to do with the need for expansion of typed arrays in the =20
> future. If you want to add binary256, int256, etc., =20

=46or now, I think we should limit ourselves to 128 bit arrays. =20

> Section 7. Security Considerations =20
> =20
> ... =20
> =20
> I consider an irregular-length array to be a coding error; that is =20
> with my security hat on. =20

Yes. =20

Would it be right to say that in the case of such a=C2=A0coding error, th=
e =20
implementation should either terminate processing or treat the byte =20
string as if there were no TypeArray tag applied to it (as uint8)=3F =20

-- =20
Johnathan Roatch <jroatch=40gmail.com> =20
https://jroatch.nfshost.com/ =20


--57a26e8a_4dadee93_385a
Content-Type: text/html; charset="utf-8"
Content-Transfer-Encoding: quoted-printable
Content-Disposition: inline

<html><head><style>body=7Bfont-family:Helvetica,Arial;font-size:13px=7D</=
style></head><body style=3D=22word-wrap: break-word; -webkit-nbsp-mode: s=
pace; -webkit-line-break: after-white-space;=22><div id=3D=22bloop=5Fcust=
omfont=22 style=3D=22font-family:Helvetica,Arial;font-size:13px; color: r=
gba(0,0,0,1.0); margin: 0px; line-height: auto;=22>Hmm, I=E2=80=99m not s=
ure I agree with either the criticism or the conclusions below.</div><div=
 id=3D=22bloop=5Fcustomfont=22 style=3D=22font-family:Helvetica,Arial;fon=
t-size:13px; color: rgba(0,0,0,1.0); margin: 0px; line-height: auto;=22>A=
s the draft stands, it is closely mirroring the typed array functionality=
 that has been defined in several environments including JavaScript. &nbs=
p;Why would we be chasing some form of symmetry instead of following what=
 the users seem to need=3F</div><div id=3D=22bloop=5Fcustomfont=22 style=3D=
=22font-family:Helvetica,Arial;font-size:13px; color: rgba(0,0,0,1.0); ma=
rgin: 0px; line-height: auto;=22><br></div><div id=3D=22bloop=5Fcustomfon=
t=22 style=3D=22font-family:Helvetica,Arial;font-size:13px; color: rgba(0=
,0,0,1.0); margin: 0px; line-height: auto;=22>uint8/sint8 arrays are arra=
ys of integers, which is different from untagged byte strings. &nbsp;Henc=
e the tags.</div><div id=3D=22bloop=5Fcustomfont=22 style=3D=22font-famil=
y:Helvetica,Arial;font-size:13px; color: rgba(0,0,0,1.0); margin: 0px; li=
ne-height: auto;=22>I=E2=80=99m not aware of a use case for IEEE 754 deci=
mal at this point. &nbsp;Sure, one can be constructed, but the objective =
is not to cover everything that has been invented under the sun, but what=
 is actually needed.</div><div id=3D=22bloop=5Fcustomfont=22 style=3D=22f=
ont-family:Helvetica,Arial;font-size:13px; color: rgba(0,0,0,1.0); margin=
: 0px; line-height: auto;=22><br></div><div id=3D=22bloop=5Fcustomfont=22=
 style=3D=22font-family:Helvetica,Arial;font-size:13px; color: rgba(0,0,0=
,1.0); margin: 0px; line-height: auto;=22>In summary, I continue to belie=
ve that the draft is all good as it stands; nothing should be changed.</d=
iv> <br> <div id=3D=22bloop=5Fsign=5F1470262591353765120=22 class=3D=22bl=
oop=5Fsign=22><div style=3D=22font-family:helvetica,arial;font-size:13px=22=
>Gr=C3=BC=C3=9Fe, Carsten</div></div> <br><p class=3D=22airmail=5Fon=22>O=
n 3 August 2016 at 18:57:07, Johnathan Roatch (<a href=3D=22mailto:jroatc=
h=40gmail.com=22>jroatch=40gmail.com</a>) wrote:</p> <blockquote type=3D=22=
cite=22 class=3D=22clean=5Fbq=22><span><div><div></div><div>On Thu, 2016-=
07-28 at 18:09 -0700, Sean Leonard wrote:
<br>&gt; Section 2. Typed Arrays
<br>&gt; =20
<br>&gt; Okay; makes sense. However, the mismatch between =5Bu/s=5Dint8 a=
nd
<br>&gt; binary16&nbsp;
<br>&gt; is jarring. It would seem a lot better if&nbsp;
<br>&gt; uint16/sint16/binary16..uint128/sint128/binary128.
<br>
<br>When I initially made up the scheme, my primary goal was to pack in a=
ny
<br>possible gaps. I agree now that byte sizes should definitely be
<br>aligned, especially to include 128-bit integers.
<br>
<br>&gt; In contrast, I don't see a lot of value to tagging an array as
<br>&gt; uint8...that would be called a plain old =22byte string=22 by an=
y
<br>&gt; definition&nbsp;of the term.
<br>
<br>Good insight here. I confused myself thinking that a JavaScript
<br>ArrayBuffer could been used without a uint8 view. The uint8 tag can
<br>definitely be dropped to make room for byte size alignment.
<br>Implementations should then use un-tagged byte strings for uint8
<br>arrays.
<br>
<br>&gt; =46or the case of uint8 and sint8, I propose that you allocate f=
our
<br>&gt; tags&nbsp;in approximately 0b001111=5Fc=5Fs, where c =3D clamped=
, and s =3D
<br>&gt; signed. =20
<br>
<br>If the uint8 is gone then all we have left here is sint8. clamped
<br>arithmetic isn't as critical to fit into some bit pattern as it's
<br>really just a hint on how to process numbers and not a number format =
on
<br>it's own. Now that I think of it, clamped arithmetic can actually be
<br>it's own tag that can apply to any integer or integer array.
<br>
<br>&gt; If you want parity with the bit positions in the other tags, I w=
ould
<br>&gt; propose rearranging the 0b010=5Ff=5Fs=5Fe=5Fll to: 0b010=5Fll=5F=
f=5Fe=5Fs
<br>
<br>I put float and signed together to form the collective bit field of
<br>number type, and those bits were made the most significant so that
<br>there would be less gaps in the list of tags. We have 3 types of
<br>numbers in a field that can hold 4. Maybe a fourth type of number
<br>should fill in those gaps.
<br>
<br>Are the IEEE 754-2008 decimal formats used enough to include them=3F
<br>
<br>After playing around with possible arrangements a bit this would be
<br>what I would do: 0b01=5Flll=5Ff=5Fs=5Fe.
<br>=22lll=22 is the used as =222**(lll)=22 for calculating the number of=
 bytes.
<br>All tags of a single byte size, except sint8, are unreserved. This wi=
ll
<br>make is simple to compute, as all it would take on most platforms is =
a
<br>bit-mask and a couple shifts. the unreserved space will also be a goo=
d
<br>place to put related tags such as Homogeneous and Multi-dimensional
<br>Arrays.
<br>=22f=5Fs=22 is the number format: =5BuintN, sintN, IEEE 754 binaryN,&=
nbsp;IEEE
<br>754&nbsp;decimalN=5D.
<br>=22e=22 is big or little endianness.
<br>
<br>In this formulation, uint8 is no tag, uint16-LE is tag 0x49,&nbsp;IEE=
E 754
<br>binary32-LE is tag 0x55, uint128-LE is 0x61, and I guess sint8 can be=

<br>tag 0x42.
<br>
<br>&gt; I would also like to suggest using tags in the 2-byte space, suc=
h as&nbsp;
<br>&gt; 0x0220-0x0237 (or 0x021C - 0x0237). =20
<br>
<br>Since there's already implementations that follow the old drafts, goi=
ng
<br>to the 2 byte tag space might be necessary for backwards compatibilit=
y
<br>of those implementations.
<br>
<br>If we are going to reuse the original range, now is the time to have =
a
<br>breaking change.
<br>
<br>&gt; I suppose that you have thought of this already. I am less conce=
rned
<br>&gt; about low-byte =22pollution=22; =20
<br>
<br>About =22low-byte pollution=22, I've heard criticism about how endian=
ness
<br>takes up too many tags. While I thought about separating out endianne=
ss
<br>to it's own tag, it's doesn't make much sense outside of the context =
of
<br>typed arrays.
<br>
<br>If we include IEEE 754-2008 decimal while fixing up the byte size
<br>aliment as I wrote above, the total amount of tags will increase from=

<br>24 to 32. =46orcing everything to be little-endian will reduce that
<br>amount to 16. It really depends on how important having big endiannes=
s
<br>is.
<br>
<br>&gt; it has more to do with the need for expansion of typed arrays in=
 the
<br>&gt; future. If you want to add binary256, int256, etc.,
<br>
<br>=46or now, I think we should limit ourselves to 128 bit arrays.
<br>
<br>&gt; Section 7. Security Considerations
<br>&gt; =20
<br>&gt; ...
<br>&gt; =20
<br>&gt; I consider an irregular-length array to be a coding error; that =
is
<br>&gt; with my security hat on.
<br>
<br>Yes.
<br>
<br>Would it be right to say that in the case of such a&nbsp;coding error=
, the
<br>implementation should either terminate processing or treat the byte
<br>string as if there were no TypeArray tag applied to it (as uint8)=3F
<br>
<br>-- =20
<br>Johnathan Roatch &lt;jroatch=40gmail.com&gt;
<br>https://jroatch.nfshost.com/
<br>
<br></div></div></span></blockquote></body></html>
--57a26e8a_4dadee93_385a--


From nobody Wed Aug  3 16:30:58 2016
Return-Path: <dev+ietf@seantek.com>
X-Original-To: cbor@ietfa.amsl.com
Delivered-To: cbor@ietfa.amsl.com
Received: from localhost (localhost [127.0.0.1]) by ietfa.amsl.com (Postfix) with ESMTP id 37F5812D8F1 for <cbor@ietfa.amsl.com>; Wed,  3 Aug 2016 16:30:57 -0700 (PDT)
X-Virus-Scanned: amavisd-new at amsl.com
X-Spam-Flag: NO
X-Spam-Score: -2.601
X-Spam-Level: 
X-Spam-Status: No, score=-2.601 tagged_above=-999 required=5 tests=[BAYES_00=-1.9, RCVD_IN_DNSWL_LOW=-0.7, SPF_HELO_PASS=-0.001] autolearn=ham autolearn_force=no
Received: from mail.ietf.org ([4.31.198.44]) by localhost (ietfa.amsl.com [127.0.0.1]) (amavisd-new, port 10024) with ESMTP id WUynMq4EkQ6V for <cbor@ietfa.amsl.com>; Wed,  3 Aug 2016 16:30:54 -0700 (PDT)
Received: from mxout-08.mxes.net (mxout-08.mxes.net [216.86.168.183]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by ietfa.amsl.com (Postfix) with ESMTPS id 6356212D8ED for <cbor@ietf.org>; Wed,  3 Aug 2016 16:30:54 -0700 (PDT)
Received: from [192.168.123.7] (unknown [75.83.2.34]) (using TLSv1.2 with cipher DHE-RSA-AES128-SHA (128/128 bits)) (No client certificate requested) by smtp.mxes.net (Postfix) with ESMTPSA id 99F43509B6; Wed,  3 Aug 2016 19:30:52 -0400 (EDT)
To: Johnathan Roatch <jroatch@gmail.com>, Carsten Bormann <cabo@tzi.org>, glenn_engel@keysight.com
References: <04EFF12F483FA149B07653989B86861F217A069F@wcosexch01k.cos.is.keysight.com> <577416B1.3090308@tzi.org> <577ED54A.3050705@tzi.org> <3306456d-17d4-dfbb-9940-e304f5f73c61@seantek.com> <1470241923.1200.99.camel@gmail.com>
From: Sean Leonard <dev+ietf@seantek.com>
Message-ID: <0409ea61-d03f-015e-ef28-abfdcd394709@seantek.com>
Date: Wed, 3 Aug 2016 16:29:20 -0700
User-Agent: Mozilla/5.0 (Windows NT 6.1; WOW64; rv:45.0) Gecko/20100101 Thunderbird/45.2.0
MIME-Version: 1.0
In-Reply-To: <1470241923.1200.99.camel@gmail.com>
Content-Type: text/plain; charset=utf-8; format=flowed
Content-Transfer-Encoding: quoted-printable
Archived-At: <https://mailarchive.ietf.org/arch/msg/cbor/4AUUxks5LMdeTTgNLIf5erJxE14>
Cc: cbor@ietf.org, draft-jroatch-cbor-tags@tools.ietf.org
Subject: Re: [Cbor] draft-jroatch-cbor-tags-04
X-BeenThere: cbor@ietf.org
X-Mailman-Version: 2.1.17
Precedence: list
List-Id: "Concise Binary Object Representation \(CBOR\)" <cbor.ietf.org>
List-Unsubscribe: <https://www.ietf.org/mailman/options/cbor>, <mailto:cbor-request@ietf.org?subject=unsubscribe>
List-Archive: <https://mailarchive.ietf.org/arch/browse/cbor/>
List-Post: <mailto:cbor@ietf.org>
List-Help: <mailto:cbor-request@ietf.org?subject=help>
List-Subscribe: <https://www.ietf.org/mailman/listinfo/cbor>, <mailto:cbor-request@ietf.org?subject=subscribe>
X-List-Received-Date: Wed, 03 Aug 2016 23:30:57 -0000

Hello Johnathan,

Glad to see that my review was helpful. Additionally:

On 8/3/2016 9:32 AM, Johnathan Roatch wrote:
> On Thu, 2016-07-28 at 18:09 -0700, Sean Leonard wrote:
>> Section 2. Typed Arrays
>>
>> Okay; makes sense. However, the mismatch between [u/s]int8 and
>> binary16
>> is jarring. It would seem a lot better if
>> uint16/sint16/binary16..uint128/sint128/binary128.
> When I initially made up the scheme, my primary goal was to pack in any=

> possible gaps. I agree now that byte sizes should definitely be
> aligned, especially to include 128-bit integers.

Great.

>
>> In contrast, I don't see a lot of value to tagging an array as
>> uint8...that would be called a plain old "byte string" by any
>> definition of the term.
> Good insight here. I confused myself thinking that a JavaScript
> ArrayBuffer could been used without a uint8 view. The uint8 tag can
> definitely be dropped to make room for byte size alignment.
> Implementations should then use un-tagged byte strings for uint8
> arrays.

The issues here have to do with code complexity (branches versus=20
jumping) and semantics. When handling JavaScript array buffers, is there =

a meaningful distinction between uint8 arrays and "byte strings", such=20
that the semantic difference needs to be encoded on the wire?

>
>> For the case of uint8 and sint8, I propose that you allocate four
>> tags in approximately 0b001111_c_s, where c =3D clamped, and s =3D
>> signed.
> If the uint8 is gone then all we have left here is sint8. clamped
> arithmetic isn't as critical to fit into some bit pattern as it's
> really just a hint on how to process numbers and not a number format on=

> it's own. Now that I think of it, clamped arithmetic can actually be
> it's own tag that can apply to any integer or integer array.

Yes, it could. I have not formed an opinion yet of whether it should or=20
should not be. There is a technique discussed in=20
draft-bormann-cbor-tags-oid, to stack tags, so that you can have tag1 (=20
tag2 ( cbor ) ), which can have various semantics. In that case, a=20
proposal is that the first tag applies uniformly to the keys in a map;=20
the second tag (inner tag) (if present) applies uniformly to the values. =

But that is only proposed for certain binary identifiers, namely for=20
OIDs and UUIDs.

When I wrote that proposal, I did not intend it to be general semantic=20
advice. However it might be useful.

If you have a wider block of tags (i.e., in the two-byte tag range), you =

can use more bits.

>
>> If you want parity with the bit positions in the other tags, I would
>> propose rearranging the 0b010_f_s_e_ll to: 0b010_ll_f_e_s
> I put float and signed together to form the collective bit field of
> number type, and those bits were made the most significant so that
> there would be less gaps in the list of tags. We have 3 types of
> numbers in a field that can hold 4. Maybe a fourth type of number
> should fill in those gaps.
>
> Are the IEEE 754-2008 decimal formats used enough to include them?

Clearly IEEE 754 decimal formats are important enough that a group of=20
(very talented) academics and engineers got together to write a (very=20
expensive) standard about it.

http://dec64.com/
http://dx.doi.org/10.1109/AICERA-ICMiCR.2013.6575957

floating point decimal appears to have applications in financial,=20
astronomical, or scientific calculations. I would not be surprised if=20
there are chemical applications as well.

These applications make sense. For finance (not accounting), such as=20
financial engineering, preserving base 10 for derivatives etc. makes=20
sense as well as very small fractions of currency units in orders of=20
magnitude (i.e., 10^-x). For very large or very small physical=20
calculations (distances between stars, molecular weights, mols, SI=20
units), preserving base 10 is absolutely essential.

We don't see decimal formats much around IETF because the most=20
math-intensive stuff tends to be about routing (which has to do with=20
graph and set theory), latency/time calculations (which can be measured=20
in fixed units, e.g., nanoseconds), and cryptography (where a single bit =

of rounding error is designed to be fatal). That does not mean they are=20
not useful.

JavaScript Typed Arrays are mainly based on binary (base 2) floating=20
point because a big application is for web graphics, like games,=20
animation, and video and whatnot. Also it doesn't hurt that JavaScript=20
has binary floating point built-in. But CBOR is about far more than the=20
web and far more than graphics.

I say, put a call or email in to one of the IEEE 754 people.

Overall, I think it's worth an extra bit in the major type 6 (tag) space =

(as you propose below). It would be too difficult to add floating point=20
decimal to add it to the major type 7 (simple) space, but I think we=20
kind of all know that.

Maybe it would be good to have a tag to tag a single type 7 (simple)=20
floating-point number as decimal instead of binary. A bit out of scope,=20
though.

>
> After playing around with possible arrangements a bit this would be
> what I would do: 0b01_lll_f_s_e.
> "lll" is the used as "2**(lll)" for calculating the number of bytes.
> All tags of a single byte size, except sint8, are unreserved. This will=

> make is simple to compute, as all it would take on most platforms is a
> bit-mask and a couple shifts. the unreserved space will also be a good
> place to put related tags such as Homogeneous and Multi-dimensional
> Arrays.
> "f_s" is the number format: [uintN, sintN, IEEE 754 binaryN, IEEE
> 754 decimalN].
> "e" is big or little endianness.
>
> In this formulation, uint8 is no tag, uint16-LE is tag 0x49, IEEE 754
> binary32-LE is tag 0x55, uint128-LE is 0x61, and I guess sint8 can be
> tag 0x42.
>
>> I would also like to suggest using tags in the 2-byte space, such as
>> 0x0220-0x0237 (or 0x021C - 0x0237).
> Since there's already implementations that follow the old drafts, going=

> to the 2 byte tag space might be necessary for backwards compatibility
> of those implementations.
>
> If we are going to reuse the original range, now is the time to have a
> breaking change.

(Note: when I wrote 2-byte space, I meant 2 bytes for the integer and=20
one byte for the major type and additional information byte, hence 2+1=3D=
3=20
bytes.)

>
>> I suppose that you have thought of this already. I am less concerned
>> about low-byte "pollution";
> About "low-byte pollution", I've heard criticism about how endianness
> takes up too many tags. While I thought about separating out endianness=

> to it's own tag, it's doesn't make much sense outside of the context of=

> typed arrays.
>
> If we include IEEE 754-2008 decimal while fixing up the byte size
> aliment as I wrote above, the total amount of tags will increase from
> 24 to 32. Forcing everything to be little-endian will reduce that
> amount to 16. It really depends on how important having big endianness
> is.

Network byte order is big-endian. It would be a bit silly not to support =
it.

I'm seeing 32 * 2 =3D 64...

Overall, what is the length of the tagged binary arrays that we are=20
talking about? I am thinking that the number of items in each array will =

significantly dwarf the difference between a 40-255 tag (two bytes) and=20
a 256-65535 tag (three bytes, what I earlier called "2-byte space").

>
>> it has more to do with the need for expansion of typed arrays in the
>> future. If you want to add binary256, int256, etc.,
> For now, I think we should limit ourselves to 128 bit arrays.
>
>> Section 7. Security Considerations
>>
>> ...
>>
>> I consider an irregular-length array to be a coding error; that is
>> with my security hat on.
> Yes.
>
> Would it be right to say that in the case of such a coding error, the
> implementation should either terminate processing or treat the byte
> string as if there were no TypeArray tag applied to it (as uint8)?

To me, "enforce" means fatal (fail processing) if not met.

A low-level, generic CBOR decoder would not enforce anything related to=20
the tags.

Actually I would argue that a "secure" low-level, generic decoder would=20
enforce the simple types (right now, there's not much to enforce...other =

than for types 25-27, 16, 32, and 64 arbitrary bits follow, and for 24,=20
one byte follows), and would enforce uniqueness of keys in a map, pairs=20
of keys/values in a map, and (importantly) UTF-8 conformance. It is not=20
clear whether a "secure" low-level, generic decoder ought to enforce=20
"smallest number of bytes per integer" (aka prohibit overlong=20
encodings). For example, tag 0 can be encoded as C0 or D800; unsigned=20
integer 0 can be encoded as 00, 1800, or even 1A00000000. RFC 7049=20
doesn't say these are erroneous, so I guess they are okay. Actually RFC=20
7049 Section 3.6 specifically discusses this. (Unicode says overlong=20
encodings are bad, hence why a "secure" CBOR decoder needs to enforce it.=
)

I would expect that a generic CBOR decoder that is aware of a tag's=20
semantics, would enforce the tag's semantics. More tags mean more code.=20
Since a large block of tags is getting allocated, it would definitely=20
save on decoder complexity if all the tags are in one contiguous block=20
(hence arguing more for the 2-byte range 256-65535) and therefore a=20
decoder could enforce length multiples based on a small handful of bits=20
and a couple of shift ops. In that case you could allocate a couple of=20
extra spare bits for future growth. Nobody is going to bat an eye at=20
allocating 256 or even 512 or 1024 tags out of the 256-65535 range.

Such a tag-aware CBOR decoder needs to be aware of the overlong encoding =

issue, for any tag that is not in the 33-64 bit range.

Sean


