
From nobody Tue Apr  1 03:07:33 2014
Return-Path: <rohanse2@cisco.com>
X-Original-To: clue@ietfa.amsl.com
Delivered-To: clue@ietfa.amsl.com
Received: from localhost (ietfa.amsl.com [127.0.0.1]) by ietfa.amsl.com (Postfix) with ESMTP id 113671A7D85 for <clue@ietfa.amsl.com>; Tue,  1 Apr 2014 03:07:30 -0700 (PDT)
X-Virus-Scanned: amavisd-new at amsl.com
X-Spam-Flag: NO
X-Spam-Score: -12.71
X-Spam-Level: 
X-Spam-Status: No, score=-12.71 tagged_above=-999 required=5 tests=[BAYES_00=-1.9, DKIM_SIGNED=0.1, DKIM_VALID=-0.1, DKIM_VALID_AU=-0.1, HTML_MESSAGE=0.001, J_CHICKENPOX_111=0.6, J_CHICKENPOX_14=0.6, J_CHICKENPOX_15=0.6, RCVD_IN_DNSWL_HI=-5, SPF_PASS=-0.001, T_RP_MATCHES_RCVD=-0.01, USER_IN_DEF_DKIM_WL=-7.5] autolearn=ham
Received: from mail.ietf.org ([4.31.198.44]) by localhost (ietfa.amsl.com [127.0.0.1]) (amavisd-new, port 10024) with ESMTP id EezATB_2iwT7 for <clue@ietfa.amsl.com>; Tue,  1 Apr 2014 03:07:25 -0700 (PDT)
Received: from rcdn-iport-2.cisco.com (rcdn-iport-2.cisco.com [173.37.86.73]) by ietfa.amsl.com (Postfix) with ESMTP id 067791A0901 for <clue@ietf.org>; Tue,  1 Apr 2014 03:07:24 -0700 (PDT)
DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/simple; d=cisco.com; i=@cisco.com; l=25030; q=dns/txt; s=iport; t=1396346842; x=1397556442; h=from:to:subject:date:message-id:mime-version; bh=1nW0hQ0E5eJONK/L8fc8cYPGKMLgYx0JVM/cRrbQLUM=; b=gcJOoBTGlgxz6O4pjlrn2K0JfMUpSykI7UtqmafIvBaQGYCCvsfEtzKH XVMOkB9CcGuKi4e8La/tsY2wjTHDyn4ANrXJjmgqzvT3Az6Wg/nT6gIo1 Ni2TWopH+ydyW0fXtDgi06YfMciY++pRwlxZFHelW4JeWqpyhlTv2Bkb6 Y=;
X-IronPort-Anti-Spam-Filtered: true
X-IronPort-Anti-Spam-Result: AlsFAFyPOlOtJV2a/2dsb2JhbABZgkJEO1fDL4EbFnSCJwEELUcXASpWJgEEGxOHXp9ksXgXjhgng1yBFASrD4MwgWlC
X-IronPort-AV: E=Sophos;i="4.97,771,1389744000";  d="scan'208,217";a="314232327"
Received: from rcdn-core-3.cisco.com ([173.37.93.154]) by rcdn-iport-2.cisco.com with ESMTP; 01 Apr 2014 10:07:12 +0000
Received: from xhc-rcd-x14.cisco.com (xhc-rcd-x14.cisco.com [173.37.183.88]) by rcdn-core-3.cisco.com (8.14.5/8.14.5) with ESMTP id s31A7BbR016866 (version=TLSv1/SSLv3 cipher=AES128-SHA bits=128 verify=FAIL) for <clue@ietf.org>; Tue, 1 Apr 2014 10:07:11 GMT
Received: from xmb-aln-x07.cisco.com ([169.254.2.162]) by xhc-rcd-x14.cisco.com ([173.37.183.88]) with mapi id 14.03.0123.003; Tue, 1 Apr 2014 05:07:10 -0500
From: "Robert Hansen (rohanse2)" <rohanse2@cisco.com>
To: "clue@ietf.org" <clue@ietf.org>
Thread-Topic: Using BUNDLE with CLUE
Thread-Index: Ac9Nkg5IUPNx7ldUTDK8kHk510O6AQ==
Date: Tue, 1 Apr 2014 10:07:10 +0000
Message-ID: <C6252EA94E00E44EADC3A2FEB59D4402020DF93F@xmb-aln-x07.cisco.com>
Accept-Language: en-GB, en-US
Content-Language: en-US
X-MS-Has-Attach: 
X-MS-TNEF-Correlator: 
x-originating-ip: [10.147.32.14]
Content-Type: multipart/alternative; boundary="_000_C6252EA94E00E44EADC3A2FEB59D4402020DF93Fxmbalnx07ciscoc_"
MIME-Version: 1.0
Archived-At: http://mailarchive.ietf.org/arch/msg/clue/_En2m6bu7Jp-_mUIVRZ1kADiRVU
Subject: [clue] Using BUNDLE with CLUE
X-BeenThere: clue@ietf.org
X-Mailman-Version: 2.1.15
Precedence: list
List-Id: CLUE - ControLling mUltiple streams for TElepresence <clue.ietf.org>
List-Unsubscribe: <https://www.ietf.org/mailman/options/clue>, <mailto:clue-request@ietf.org?subject=unsubscribe>
List-Archive: <http://www.ietf.org/mail-archive/web/clue/>
List-Post: <mailto:clue@ietf.org>
List-Help: <mailto:clue-request@ietf.org?subject=help>
List-Subscribe: <https://www.ietf.org/mailman/listinfo/clue>, <mailto:clue-request@ietf.org?subject=subscribe>
X-List-Received-Date: Tue, 01 Apr 2014 10:07:30 -0000

--_000_C6252EA94E00E44EADC3A2FEB59D4402020DF93Fxmbalnx07ciscoc_
Content-Type: text/plain; charset="us-ascii"
Content-Transfer-Encoding: quoted-printable

Introduction:

A fundamental aspect of CLUE is the sending of multiple streams of media. W=
e are using conventional SDP to specify the encodings we support, and hence=
 each media stream necessitates a separate m-line. For most use-cases, howe=
ver, using a separate port per m-line is suboptimal: it means opening more =
ports, more NAT work, more resources consumed for ICE, etc. As such, a meth=
od to allow multiple m-lines to share the same 5-tuple address is highly de=
sirable.

Until now the signalling work has focused on the demultiplexed case. As suc=
h, this is an attempt to evaluate the use of BUNDLE with CLUE - those with =
a much better understanding of BUNDLE than I will be able to correct the mi=
stakes I'll inevitably make.

We have generally figured that CLUE would not raise any particular issues w=
ith BUNDLE, given that CLUE now uses separate m-lines for encodings, but we=
 need to go through and make sure, as well as providing guidance in the CLU=
E documentation on how BUNDLE interacts with it.

Obviously I'm not going to duplicate the BUNDLE draft here, instead I'll ju=
st make reference to the draft.

Initial offer/answer:

One of the principle concerns of BUNDLE is to ensure that a receiver does n=
ot receive an SDP it considers invalid due to its lack of support for bundl=
ing m-lines. The fact that we recommend that CLUE-controlled m-lines are no=
t included in the initial O/A means that I think that we can combine the ex=
tra O/A that BUNDLE normally adds over an unBUNDLEd call with one of the O/=
As required by CLUE...

Adding CLUE-controlled m-lines during the Bundle Address Synchronization (B=
AS) offer:

Having completed the initial O/A and established BUNDLE support the initial=
 offerer needs to send a new offer to synchronise the BUNDLE addresses and =
make any intermediary devices aware of the addresses in use.

In CLUE, the subsequent offer is also when we would like to start adding CL=
UE-controlled m-lines. However, section 6.4.3. of the BUNDLE specification =
warns that it important that the BAS offer is accepted, and while it makes =
clear that the offerer MAY change the SDP, it warns to avoid changes that c=
ould cause the answerer to reject the new offer.

CLUE definitely needs to provide guidance here. My belief is that adding th=
e CLUE-controlled m-lines at this stage should not increase the chance of t=
he offer as a whole being rejected, so long as they share the same media ty=
pes and attributes as the existing media lines. This would also help resolv=
e the glare issue at the start of a CLUE call: since a BAS is mandatory in =
BUNDLE this provides an obvious 'who should reINVITE first' case for CLUE, =
where the initial offerer has to send an new offer even if they don't have =
any CLUE encodings to add.

If we did feel that adding new m-lines in the BAS offer was too high a risk=
 then BUNDLE and CLUE become sequential: BAS should be done first, and then=
 CLUE-controlled m-lines should be added in a subsequent INVITE. In this ca=
se BUNDLE would include one extra O/A compared to the unBUNDLEd case.

I've included an example below, modifying the example from the BUNDLE draft=
 to show the initial offerer using the BAS offer to also add CLUE-controlle=
d media.

Initial offer example - the offer includes an audio and video line with uni=
que ports, both in the same BUNDLE group (indicating that the offerer suppo=
rts BUNDLE and wants to multiplex these media flows). There is also a data =
channel that will be used for CLUE, which is included in the CLUE group.

a=3Dgroup:BUNDLE foo bar
a=3Dgroup:CLUE zen
m=3Daudio 10000 RTP/AVP 0 8 106
...
a=3Dmid:foo
m=3Dvideo 10002 RTP/AVP 96 97
...
a=3Dmid:bar
m=3Dapplication 10004 SCTP/DTLS 10004
...
a=3Dmid:zen

Initial answer example - the answer picks a local BUNDLE address and (via o=
rdering in the group attribute) selects an address for the offerer.

a=3Dgroup:BUNDLE foo bar
a=3Dgroup:CLUE zen
m=3Daudio 20000 RTP/AVP 106
...
a=3Dmid:foo
m=3Dvideo 20000 RTP/AVP 96
...
a=3Dmid:bar
m=3Dapplication 20002 SCTP/DTLS 20002
...
a=3Dmid:zen

Subsequent offer example - the offerer synchronises addresses and adds CLUE=
-controlled media lines using the same address; these are included in both =
the BUNDLE and CLUE groups.

a=3Dgroup:BUNDLE foo bar enc1 enc2 enc3
a=3Dgroup:CLUE zen enc1 enc2 enc3
m=3Daudio 10000 RTP/AVP 0 8 106
...
a=3Dmid:foo
m=3Dvideo 10000 RTP/AVP 96 97
...
a=3Dmid:bar
m=3Dapplication 10004 SCTP/DTLS 10004
...
a=3Dmid:zen
m=3Dvideo 10000 RTP/AVP 96 97
...
a=3Dmid:enc1
a=3Dlabel:1
m=3Dvideo 10000 RTP/AVP 96 97
...
a=3Dmid:enc2
a=3Dlabel:2
m=3Dvideo 10000 RTP/AVP 96 97
...
a=3Dmid:enc3
a=3Dlabel:3

Answerer rejects BUNDLE:

If the answer does not support BUNDLE then the offerer continues as in the =
disaggregated case, though it may decide to offer fewer streams, or even no=
t do CLUE at all (in the latter case it should reINVITE and remove the CLUE=
 group and data channel).

Effects of multiplexing on CLUE-relevant SDP attributes:

The BUNDLE draft specifies how certain SDP attributes are affected by multi=
plexing the m-lines, and draft-ietf-mmusic-sdp-mux-attributes describes how=
 multiplexing affects many other SDP attributes. I can't see any CLUE-speci=
fic issues here: the 'label' attribute is unaffected by multiplexing, nor i=
s the directionality of the m-lines, and the 'mid' attribute needed by CLUE=
 is also needed by BUNDLE.

Directionality of CLUE-controlled media:

CLUE-controlled m-lines are currently unidirectional. I don't believe this =
raises any specific BUNDLE issues. Christer suggests that potentially the e=
ncodings on each side could be in separate BUNDLE groups - I suggest we use=
 the same group for both sides, which would also make it easier to move to =
bidirectional streams if we ever wanted to do that.

BUNDLE support in CLUE:

While in most cases multiplexing m-lines onto a single 5-tuple will be pref=
erable we do have use cases involving disaggregated media. Because of this,=
 and because my understanding of BUNDLE is that it does not support the agg=
regated to disaggregated use case (where one device send/receives media ass=
ociated with multiple m-lines on a single IP/port, while the other sends/re=
ceives the media associated with multiple m-lines on multiple IP/ports) I d=
on't see a reason to mandate BUNDLE support for CLUE devices that only oper=
ate in disaggregated mode. As such, while CLUE should document how to use B=
UNDLE to multiplex media streams, it shouldn't be mandatory for CLUE device=
s to support it.

Demultiplexing of RTP:

Media flows from the same BUNDLE group all belong to the same RTP session; =
hence a receiver needs a method is needed to separate the various capture e=
ncodings. The method currently defined in BUNDLE is very straightforward: a=
ll of the dynamic payload types in each m-line in a BUNDLE group must be un=
ique; as such, any received RTP packet can be mapped to a specific m-line a=
nd codec.

This method, however, does suffer from the fact that the dynamic payload ty=
pe range is limited. In the case of many m-lines, of audio (with potentiall=
y many codecs) and video bundled together, and/or the increasing use of sca=
lable video codecs with multiple layers and the use of FEC repair flows, th=
is space may be insufficient.

There are other methods that could be used for this demultiplexing. For imp=
lementation with static sources or RTP mixers, streams will have a consiste=
nt SSRC, and hence the a=3Dssrc attribute can be supplied by the sender to =
allow the receiver to differentiate flows via their SSRC. For implementatio=
ns where the streams being sent do not have a static SSRC, such as source p=
rojection mixers, a specific stream-correlator can be used, such as an AppI=
D token specified by the sender in SDP and in an RTP header extension.

How stream demultiplexing is performed with BUNDLE is not a problem specifi=
c to CLUE, and is not a problem that should be solved in a CLUE-specific fa=
shion. However, it is one that we need to ensure is solved in a way consist=
ent with our use-cases, and the number of encodings we envisage allowing. G=
iven that the simultaneous sending of many capture encodings is a primary m=
otivator behind CLUE I believe that demultiplexing via payload will not be =
sufficient, and that we would need additional methods of differentiation.

If not all payload types must be unique then the question of under what cir=
cumstances codecs on two m-lines can share the same payload type becomes re=
levant.

Summary:

I believe BUNDLE is a good solution for how we can multiplex media streams =
when doing CLUE, and I can't see any CLUE-specific issues that would need t=
o be addressed in BUNDLE. I think the main decision CLUE would need to make=
 is whether do first do BUNDLE negotiation (including address synchronisati=
on) and then begin adding CLUE m-lines, or whether the CLUE media negotiati=
on process can begin with the BAS, essentially eliminating the extra O/A 'c=
ost' of BUNDLE.

--_000_C6252EA94E00E44EADC3A2FEB59D4402020DF93Fxmbalnx07ciscoc_
Content-Type: text/html; charset="us-ascii"
Content-Transfer-Encoding: quoted-printable

<html xmlns:v=3D"urn:schemas-microsoft-com:vml" xmlns:o=3D"urn:schemas-micr=
osoft-com:office:office" xmlns:w=3D"urn:schemas-microsoft-com:office:word" =
xmlns:m=3D"http://schemas.microsoft.com/office/2004/12/omml" xmlns=3D"http:=
//www.w3.org/TR/REC-html40">
<head>
<meta http-equiv=3D"Content-Type" content=3D"text/html; charset=3Dus-ascii"=
>
<meta name=3D"Generator" content=3D"Microsoft Word 14 (filtered medium)">
<style><!--
/* Font Definitions */
@font-face
	{font-family:Calibri;
	panose-1:2 15 5 2 2 2 4 3 2 4;}
/* Style Definitions */
p.MsoNormal, li.MsoNormal, div.MsoNormal
	{margin:0cm;
	margin-bottom:.0001pt;
	font-size:11.0pt;
	font-family:"Calibri","sans-serif";
	mso-fareast-language:EN-US;}
a:link, span.MsoHyperlink
	{mso-style-priority:99;
	color:blue;
	text-decoration:underline;}
a:visited, span.MsoHyperlinkFollowed
	{mso-style-priority:99;
	color:purple;
	text-decoration:underline;}
span.EmailStyle17
	{mso-style-type:personal-compose;
	font-family:"Calibri","sans-serif";
	color:windowtext;}
.MsoChpDefault
	{mso-style-type:export-only;
	font-family:"Calibri","sans-serif";
	mso-fareast-language:EN-US;}
@page WordSection1
	{size:612.0pt 792.0pt;
	margin:72.0pt 72.0pt 72.0pt 72.0pt;}
div.WordSection1
	{page:WordSection1;}
--></style><!--[if gte mso 9]><xml>
<o:shapedefaults v:ext=3D"edit" spidmax=3D"1026" />
</xml><![endif]--><!--[if gte mso 9]><xml>
<o:shapelayout v:ext=3D"edit">
<o:idmap v:ext=3D"edit" data=3D"1" />
</o:shapelayout></xml><![endif]-->
</head>
<body lang=3D"EN-GB" link=3D"blue" vlink=3D"purple">
<div class=3D"WordSection1">
<p class=3D"MsoNormal">Introduction:<o:p></o:p></p>
<p class=3D"MsoNormal"><o:p>&nbsp;</o:p></p>
<p class=3D"MsoNormal">A fundamental aspect of CLUE is the sending of multi=
ple streams of media. We are using conventional SDP to specify the encoding=
s we support, and hence each media stream necessitates a separate m-line. F=
or most use-cases, however, using
 a separate port per m-line is suboptimal: it means opening more ports, mor=
e NAT work, more resources consumed for ICE, etc. As such, a method to allo=
w multiple m-lines to share the same 5-tuple address is highly desirable.<o=
:p></o:p></p>
<p class=3D"MsoNormal"><o:p>&nbsp;</o:p></p>
<p class=3D"MsoNormal">Until now the signalling work has focused on the dem=
ultiplexed case. As such, this is an attempt to evaluate the use of BUNDLE =
with CLUE - those with a much better understanding of BUNDLE than I will be=
 able to correct the mistakes I'll
 inevitably make.<o:p></o:p></p>
<p class=3D"MsoNormal"><o:p>&nbsp;</o:p></p>
<p class=3D"MsoNormal">We have generally figured that CLUE would not raise =
any particular issues with BUNDLE, given that CLUE now uses separate m-line=
s for encodings, but we need to go through and make sure, as well as provid=
ing guidance in the CLUE documentation
 on how BUNDLE interacts with it.<o:p></o:p></p>
<p class=3D"MsoNormal"><o:p>&nbsp;</o:p></p>
<p class=3D"MsoNormal">Obviously I'm not going to duplicate the BUNDLE draf=
t here, instead I'll just make reference to the draft.<o:p></o:p></p>
<p class=3D"MsoNormal"><o:p>&nbsp;</o:p></p>
<p class=3D"MsoNormal">Initial offer/answer:<o:p></o:p></p>
<p class=3D"MsoNormal"><o:p>&nbsp;</o:p></p>
<p class=3D"MsoNormal">One of the principle concerns of BUNDLE is to ensure=
 that a receiver does not receive an SDP it considers invalid due to its la=
ck of support for bundling m-lines. The fact that we recommend that CLUE-co=
ntrolled m-lines are not included
 in the initial O/A means that I think that we can combine the extra O/A th=
at BUNDLE normally adds over an unBUNDLEd call with one of the O/As require=
d by CLUE...<o:p></o:p></p>
<p class=3D"MsoNormal"><o:p>&nbsp;</o:p></p>
<p class=3D"MsoNormal">Adding CLUE-controlled m-lines during the Bundle Add=
ress Synchronization (BAS) offer:<o:p></o:p></p>
<p class=3D"MsoNormal"><o:p>&nbsp;</o:p></p>
<p class=3D"MsoNormal">Having completed the initial O/A and established BUN=
DLE support the initial offerer needs to send a new offer to synchronise th=
e BUNDLE addresses and make any intermediary devices aware of the addresses=
 in use.<o:p></o:p></p>
<p class=3D"MsoNormal"><o:p>&nbsp;</o:p></p>
<p class=3D"MsoNormal">In CLUE, the subsequent offer is also when we would =
like to start adding CLUE-controlled m-lines. However, section 6.4.3. of th=
e BUNDLE specification warns that it important that the BAS offer is accept=
ed, and while it makes clear that
 the offerer MAY change the SDP, it warns to avoid changes that could cause=
 the answerer to reject the new offer.<o:p></o:p></p>
<p class=3D"MsoNormal"><o:p>&nbsp;</o:p></p>
<p class=3D"MsoNormal">CLUE definitely needs to provide guidance here. My b=
elief is that adding the CLUE-controlled m-lines at this stage should not i=
ncrease the chance of the offer as a whole being rejected, so long as they =
share the same media types and attributes
 as the existing media lines. This would also help resolve the glare issue =
at the start of a CLUE call: since a BAS is mandatory in BUNDLE this provid=
es an obvious 'who should reINVITE first' case for CLUE, where the initial =
offerer has to send an new offer
 even if they don't have any CLUE encodings to add.<o:p></o:p></p>
<p class=3D"MsoNormal"><o:p>&nbsp;</o:p></p>
<p class=3D"MsoNormal">If we did feel that adding new m-lines in the BAS of=
fer was too high a risk then BUNDLE and CLUE become sequential: BAS should =
be done first, and then CLUE-controlled m-lines should be added in a subseq=
uent INVITE. In this case BUNDLE would
 include one extra O/A compared to the unBUNDLEd case.<o:p></o:p></p>
<p class=3D"MsoNormal"><o:p>&nbsp;</o:p></p>
<p class=3D"MsoNormal">I've included an example below, modifying the exampl=
e from the BUNDLE draft to show the initial offerer using the BAS offer to =
also add CLUE-controlled media.<o:p></o:p></p>
<p class=3D"MsoNormal"><o:p>&nbsp;</o:p></p>
<p class=3D"MsoNormal">Initial offer example - the offer includes an audio =
and video line with unique ports, both in the same BUNDLE group (indicating=
 that the offerer supports BUNDLE and wants to multiplex these media flows)=
. There is also a data channel that
 will be used for CLUE, which is included in the CLUE group.<o:p></o:p></p>
<p class=3D"MsoNormal"><o:p>&nbsp;</o:p></p>
<p class=3D"MsoNormal">a=3Dgroup:BUNDLE foo bar<o:p></o:p></p>
<p class=3D"MsoNormal">a=3Dgroup:CLUE zen<o:p></o:p></p>
<p class=3D"MsoNormal">m=3Daudio 10000 RTP/AVP 0 8 106<o:p></o:p></p>
<p class=3D"MsoNormal">...<o:p></o:p></p>
<p class=3D"MsoNormal">a=3Dmid:foo<o:p></o:p></p>
<p class=3D"MsoNormal">m=3Dvideo 10002 RTP/AVP 96 97<o:p></o:p></p>
<p class=3D"MsoNormal">...<o:p></o:p></p>
<p class=3D"MsoNormal">a=3Dmid:bar<o:p></o:p></p>
<p class=3D"MsoNormal">m=3Dapplication 10004 SCTP/DTLS 10004<o:p></o:p></p>
<p class=3D"MsoNormal">...<o:p></o:p></p>
<p class=3D"MsoNormal">a=3Dmid:zen<o:p></o:p></p>
<p class=3D"MsoNormal"><o:p>&nbsp;</o:p></p>
<p class=3D"MsoNormal">Initial answer example - the answer picks a local BU=
NDLE address and (via ordering in the group attribute) selects an address f=
or the offerer.<o:p></o:p></p>
<p class=3D"MsoNormal"><o:p>&nbsp;</o:p></p>
<p class=3D"MsoNormal">a=3Dgroup:BUNDLE foo bar<o:p></o:p></p>
<p class=3D"MsoNormal">a=3Dgroup:CLUE zen<o:p></o:p></p>
<p class=3D"MsoNormal">m=3Daudio 20000 RTP/AVP 106<o:p></o:p></p>
<p class=3D"MsoNormal">...<o:p></o:p></p>
<p class=3D"MsoNormal">a=3Dmid:foo<o:p></o:p></p>
<p class=3D"MsoNormal">m=3Dvideo 20000 RTP/AVP 96<o:p></o:p></p>
<p class=3D"MsoNormal">...<o:p></o:p></p>
<p class=3D"MsoNormal">a=3Dmid:bar<o:p></o:p></p>
<p class=3D"MsoNormal">m=3Dapplication 20002 SCTP/DTLS 20002<o:p></o:p></p>
<p class=3D"MsoNormal">...<o:p></o:p></p>
<p class=3D"MsoNormal">a=3Dmid:zen<o:p></o:p></p>
<p class=3D"MsoNormal"><o:p>&nbsp;</o:p></p>
<p class=3D"MsoNormal">Subsequent offer example - the offerer synchronises =
addresses and adds CLUE-controlled media lines using the same address; thes=
e are included in both the BUNDLE and CLUE groups.<o:p></o:p></p>
<p class=3D"MsoNormal"><o:p>&nbsp;</o:p></p>
<p class=3D"MsoNormal">a=3Dgroup:BUNDLE foo bar enc1 enc2 enc3<o:p></o:p></=
p>
<p class=3D"MsoNormal">a=3Dgroup:CLUE zen enc1 enc2 enc3<o:p></o:p></p>
<p class=3D"MsoNormal">m=3Daudio 10000 RTP/AVP 0 8 106<o:p></o:p></p>
<p class=3D"MsoNormal">...<o:p></o:p></p>
<p class=3D"MsoNormal">a=3Dmid:foo<o:p></o:p></p>
<p class=3D"MsoNormal">m=3Dvideo 10000 RTP/AVP 96 97<o:p></o:p></p>
<p class=3D"MsoNormal">...<o:p></o:p></p>
<p class=3D"MsoNormal">a=3Dmid:bar<o:p></o:p></p>
<p class=3D"MsoNormal">m=3Dapplication 10004 SCTP/DTLS 10004<o:p></o:p></p>
<p class=3D"MsoNormal">...<o:p></o:p></p>
<p class=3D"MsoNormal">a=3Dmid:zen<o:p></o:p></p>
<p class=3D"MsoNormal">m=3Dvideo 10000 RTP/AVP 96 97<o:p></o:p></p>
<p class=3D"MsoNormal">...<o:p></o:p></p>
<p class=3D"MsoNormal">a=3Dmid:enc1<o:p></o:p></p>
<p class=3D"MsoNormal">a=3Dlabel:1<o:p></o:p></p>
<p class=3D"MsoNormal">m=3Dvideo 10000 RTP/AVP 96 97<o:p></o:p></p>
<p class=3D"MsoNormal">...<o:p></o:p></p>
<p class=3D"MsoNormal">a=3Dmid:enc2<o:p></o:p></p>
<p class=3D"MsoNormal">a=3Dlabel:2<o:p></o:p></p>
<p class=3D"MsoNormal">m=3Dvideo 10000 RTP/AVP 96 97<o:p></o:p></p>
<p class=3D"MsoNormal">...<o:p></o:p></p>
<p class=3D"MsoNormal">a=3Dmid:enc3<o:p></o:p></p>
<p class=3D"MsoNormal">a=3Dlabel:3<o:p></o:p></p>
<p class=3D"MsoNormal"><o:p>&nbsp;</o:p></p>
<p class=3D"MsoNormal">Answerer rejects BUNDLE:<o:p></o:p></p>
<p class=3D"MsoNormal"><o:p>&nbsp;</o:p></p>
<p class=3D"MsoNormal">If the answer does not support BUNDLE then the offer=
er continues as in the disaggregated case, though it may decide to offer fe=
wer streams, or even not do CLUE at all (in the latter case it should reINV=
ITE and remove the CLUE group and
 data channel).<o:p></o:p></p>
<p class=3D"MsoNormal"><o:p>&nbsp;</o:p></p>
<p class=3D"MsoNormal">Effects of multiplexing on CLUE-relevant SDP attribu=
tes:<o:p></o:p></p>
<p class=3D"MsoNormal"><o:p>&nbsp;</o:p></p>
<p class=3D"MsoNormal">The BUNDLE draft specifies how certain SDP attribute=
s are affected by multiplexing the m-lines, and draft-ietf-mmusic-sdp-mux-a=
ttributes describes how multiplexing affects many other SDP attributes. I c=
an't see any CLUE-specific issues
 here: the 'label' attribute is unaffected by multiplexing, nor is the dire=
ctionality of the m-lines, and the 'mid' attribute needed by CLUE is also n=
eeded by BUNDLE.<o:p></o:p></p>
<p class=3D"MsoNormal"><o:p>&nbsp;</o:p></p>
<p class=3D"MsoNormal">Directionality of CLUE-controlled media:<o:p></o:p><=
/p>
<p class=3D"MsoNormal"><o:p>&nbsp;</o:p></p>
<p class=3D"MsoNormal">CLUE-controlled m-lines are currently unidirectional=
. I don't believe this raises any specific BUNDLE issues. Christer suggests=
 that potentially the encodings on each side could be in separate BUNDLE gr=
oups - I suggest we use the same group
 for both sides, which would also make it easier to move to bidirectional s=
treams if we ever wanted to do that.<o:p></o:p></p>
<p class=3D"MsoNormal"><o:p>&nbsp;</o:p></p>
<p class=3D"MsoNormal">BUNDLE support in CLUE:<o:p></o:p></p>
<p class=3D"MsoNormal"><o:p>&nbsp;</o:p></p>
<p class=3D"MsoNormal">While in most cases multiplexing m-lines onto a sing=
le 5-tuple will be preferable we do have use cases involving disaggregated =
media. Because of this, and because my understanding of BUNDLE is that it d=
oes not support the aggregated to
 disaggregated use case (where one device send/receives media associated wi=
th multiple m-lines on a single IP/port, while the other sends/receives the=
 media associated with multiple m-lines on multiple IP/ports) I don't see a=
 reason to mandate BUNDLE support
 for CLUE devices that only operate in disaggregated mode. As such, while C=
LUE should document how to use BUNDLE to multiplex media streams, it should=
n't be mandatory for CLUE devices to support it.<o:p></o:p></p>
<p class=3D"MsoNormal"><o:p>&nbsp;</o:p></p>
<p class=3D"MsoNormal">Demultiplexing of RTP:<o:p></o:p></p>
<p class=3D"MsoNormal"><o:p>&nbsp;</o:p></p>
<p class=3D"MsoNormal">Media flows from the same BUNDLE group all belong to=
 the same RTP session; hence a receiver needs a method is needed to separat=
e the various capture encodings. The method currently defined in BUNDLE is =
very straightforward: all of the dynamic
 payload types in each m-line in a BUNDLE group must be unique; as such, an=
y received RTP packet can be mapped to a specific m-line and codec.<o:p></o=
:p></p>
<p class=3D"MsoNormal"><o:p>&nbsp;</o:p></p>
<p class=3D"MsoNormal">This method, however, does suffer from the fact that=
 the dynamic payload type range is limited. In the case of many m-lines, of=
 audio (with potentially many codecs) and video bundled together, and/or th=
e increasing use of scalable video
 codecs with multiple layers and the use of FEC repair flows, this space ma=
y be insufficient.<o:p></o:p></p>
<p class=3D"MsoNormal"><o:p>&nbsp;</o:p></p>
<p class=3D"MsoNormal">There are other methods that could be used for this =
demultiplexing. For implementation with static sources or RTP mixers, strea=
ms will have a consistent SSRC, and hence the a=3Dssrc attribute can be sup=
plied by the sender to allow the receiver
 to differentiate flows via their SSRC. For implementations where the strea=
ms being sent do not have a static SSRC, such as source projection mixers, =
a specific stream-correlator can be used, such as an AppID token specified =
by the sender in SDP and in an RTP
 header extension.<o:p></o:p></p>
<p class=3D"MsoNormal"><o:p>&nbsp;</o:p></p>
<p class=3D"MsoNormal">How stream demultiplexing is performed with BUNDLE i=
s not a problem specific to CLUE, and is not a problem that should be solve=
d in a CLUE-specific fashion. However, it is one that we need to ensure is =
solved in a way consistent with our
 use-cases, and the number of encodings we envisage allowing. Given that th=
e simultaneous sending of many capture encodings is a primary motivator beh=
ind CLUE I believe that demultiplexing via payload will not be sufficient, =
and that we would need additional
 methods of differentiation.<o:p></o:p></p>
<p class=3D"MsoNormal"><o:p>&nbsp;</o:p></p>
<p class=3D"MsoNormal">If not all payload types must be unique then the que=
stion of under what circumstances codecs on two m-lines can share the same =
payload type becomes relevant.<o:p></o:p></p>
<p class=3D"MsoNormal"><o:p>&nbsp;</o:p></p>
<p class=3D"MsoNormal">Summary:<o:p></o:p></p>
<p class=3D"MsoNormal"><o:p>&nbsp;</o:p></p>
<p class=3D"MsoNormal">I believe BUNDLE is a good solution for how we can m=
ultiplex media streams when doing CLUE, and I can't see any CLUE-specific i=
ssues that would need to be addressed in BUNDLE. I think the main decision =
CLUE would need to make is whether
 do first do BUNDLE negotiation (including address synchronisation) and the=
n begin adding CLUE m-lines, or whether the CLUE media negotiation process =
can begin with the BAS, essentially eliminating the extra O/A 'cost' of BUN=
DLE.<o:p></o:p></p>
</div>
</body>
</html>

--_000_C6252EA94E00E44EADC3A2FEB59D4402020DF93Fxmbalnx07ciscoc_--


From nobody Tue Apr  1 03:15:38 2014
Return-Path: <roberta.presta@unina.it>
X-Original-To: clue@ietfa.amsl.com
Delivered-To: clue@ietfa.amsl.com
Received: from localhost (ietfa.amsl.com [127.0.0.1]) by ietfa.amsl.com (Postfix) with ESMTP id 74D351A7032 for <clue@ietfa.amsl.com>; Tue,  1 Apr 2014 03:15:36 -0700 (PDT)
X-Virus-Scanned: amavisd-new at amsl.com
X-Spam-Flag: NO
X-Spam-Score: -0.031
X-Spam-Level: 
X-Spam-Status: No, score=-0.031 tagged_above=-999 required=5 tests=[BAYES_00=-1.9, HELO_EQ_IT=0.635, HOST_EQ_IT=1.245, SPF_PASS=-0.001, T_RP_MATCHES_RCVD=-0.01] autolearn=no
Received: from mail.ietf.org ([4.31.198.44]) by localhost (ietfa.amsl.com [127.0.0.1]) (amavisd-new, port 10024) with ESMTP id v3gjeMITAKN3 for <clue@ietfa.amsl.com>; Tue,  1 Apr 2014 03:15:33 -0700 (PDT)
Received: from smtp1.unina.it (smtp1.unina.it [192.132.34.61]) by ietfa.amsl.com (Postfix) with ESMTP id 65B5D1A702B for <clue@ietf.org>; Tue,  1 Apr 2014 03:15:33 -0700 (PDT)
Received: from [127.0.0.1] ([143.225.229.123]) (authenticated bits=0) by smtp1.unina.it (8.14.4/8.14.4) with ESMTP id s31AFRIe024592 (version=TLSv1/SSLv3 cipher=DHE-RSA-AES128-SHA bits=128 verify=NO); Tue, 1 Apr 2014 12:15:28 +0200
Message-ID: <533A91BF.7040404@unina.it>
Date: Tue, 01 Apr 2014 12:15:27 +0200
From: Roberta Presta <roberta.presta@unina.it>
User-Agent: Mozilla/5.0 (Windows NT 6.1; WOW64; rv:24.0) Gecko/20100101 Thunderbird/24.4.0
MIME-Version: 1.0
To: Christian Groves <Christian.Groves@nteczone.com>, "clue@ietf.org" <clue@ietf.org>
References: <5318809E.2030204@nteczone.com> <5318AE5E.4050404@alum.mit.edu> <5318B0C9.1050603@nteczone.com> <5318B48E.3090300@alum.mit.edu> <53290C9B.6090106@nteczone.com> <5329F3FF.7090606@alum.mit.edu> <533A16BA.40707@nteczone.com>
In-Reply-To: <533A16BA.40707@nteczone.com>
Content-Type: text/plain; charset=ISO-8859-1; format=flowed
Content-Transfer-Encoding: 7bit
Archived-At: http://mailarchive.ietf.org/arch/msg/clue/DOm9A8rVI2yZc2fwcYZugYe9hrU
Subject: Re: [clue] Participant info/type followup
X-BeenThere: clue@ietf.org
X-Mailman-Version: 2.1.15
Precedence: list
List-Id: CLUE - ControLling mUltiple streams for TElepresence <clue.ietf.org>
List-Unsubscribe: <https://www.ietf.org/mailman/options/clue>, <mailto:clue-request@ietf.org?subject=unsubscribe>
List-Archive: <http://www.ietf.org/mail-archive/web/clue/>
List-Post: <mailto:clue@ietf.org>
List-Help: <mailto:clue-request@ietf.org?subject=help>
List-Subscribe: <https://www.ietf.org/mailman/listinfo/clue>, <mailto:clue-request@ietf.org?subject=subscribe>
X-List-Received-Date: Tue, 01 Apr 2014 10:15:36 -0000

Hi Christian,

In section 7.1.1.X we are dealing with media capture attributes and in 
section 7.1.1.11 we want to define an attribute conveying information 
about participants.
I would say that here we can find both information (*references*, in the 
data model) about "who is represented in the capture" ("captured 
participants"?) and "who is the owner of the generating device" ("owner"?).

Maybe participant metadata, such as Participant Information and 
Participant Type as they are currently defined, should be treated in a 
separate section of the same level of capture scenes or media captures as.
Indeed we showed in London that repeating the vcard and the role of the 
participants in each capture, as if they were capture attributes,  
causes redundancy.

I know that I have a data model definition perspective, but I would 
propose to make a change that is more coherent with what we will 
describe formally.

Cheers,

Roberta





Il 01/04/2014 03:30, Christian Groves ha scritto:
> Hello Paul, all,
>
> If we follow the approach that there is a specific indicating of 
> whether the participant information is based on an explicit indication 
> then I would suggest the following text for the framework:
>
> Clause 7.1.11 Participant information
> (Under the 1st paragraph)
>
> The participant information contains an explicit indication of whether 
> it relates to a participant contained in the capture, from a 
> participants capture device or both. For example a video camera may 
> capture an image containing the participant, or a participant may send 
> a video capture with a presentation that does not depict the participant.
>
> Something similar would be needed under participant type.
>
> Thoughts?
>
> Regards, Christian
>
> On 20/03/2014 6:46 AM, Paul Kyzivat wrote:
>> On 3/18/14 11:18 PM, Christian Groves wrote:
>>> Hello Paul,
>>>
>>> "How" they differ is given by the example bullets below the 
>>> sentence. If
>>> you want something more normative we could remove the "For example".
>>
>> Yeah, I don't believe in specification by example. :-)
>>
>> IMO it is a bit dicey to base this distinction on the type of capture.
>>
>> I'm more comfortable with an explicit syntactic indication of the 
>> distinction, such as proposed by Roberta.
>>
>>     Thanks,
>>     Paul
>>
>>> Regards, Christian
>>>
>>> On 7/03/2014 4:46 AM, Paul Kyzivat wrote:
>>>> On 3/6/14 5:30 PM, Christian Groves wrote:
>>>>> Hello Paul,
>>>>>
>>>>> The text says media type and presentation attribute. Is that the
>>>>> relationship you're talking about?
>>>>
>>>> "How the generated content relates to the entity described in the
>>>> participant info is dependent on media type and and the presentation
>>>> attribute."
>>>>
>>>> I take that to mean that the relationship may be different for
>>>> presentation streams than non-presentation streams. But it doesn't say
>>>> *how* they differ.
>>>>
>>>>     Thanks,
>>>>     Paul
>>>>
>>>>> Regards, Christian
>>>>>
>>>>> On 7/03/2014 4:20 AM, Paul Kyzivat wrote:
>>>>>> Christian,
>>>>>>
>>>>>> I've read the quoted text several times, and I cannot figure out how
>>>>>> to *derive* your example conclusions from it. The text says the
>>>>>> relationship is dependent on the presentation attribute, but not 
>>>>>> how.
>>>>>>
>>>>>> AFAICT I could make a new definition where the a participant 
>>>>>> attached
>>>>>> to a presentation capture means that the participant is shown in the
>>>>>> presentation, and that would be equally compatible with the text.
>>>>>>
>>>>>> ISTM that more text is required to actually specify the 
>>>>>> relationships.
>>>>>>
>>>>>>     Thanks,
>>>>>>     Paul
>>>>>>
>>>>>> On 3/6/14 2:05 PM, Christian Groves wrote:
>>>>>>> Hello all,
>>>>>>>
>>>>>>> To follow up on Jonathon's comments on participant info/type and 
>>>>>>> the
>>>>>>> semantics and particularly how it relates to a presentation. 
>>>>>>> Here's a
>>>>>>> first stab at some text to stimulate some discussions.
>>>>>>>
>>>>>>>
>>>>>>> "The participant info attribute allows a provider to associate
>>>>>>> participant information with the capture source. When used in an
>>>>>>> individual capture it indicates that the captured content (e.g.
>>>>>>> video/audio/text etc.) as opposed to the actual media streams is
>>>>>>> generated from the entity described. How the generated content 
>>>>>>> relates
>>>>>>> to the entity described in the participant info is dependent on 
>>>>>>> media
>>>>>>> type and and the presentation attribute.
>>>>>>>
>>>>>>> For example:
>>>>>>> - a video capture with participant info would indicate that the 
>>>>>>> video
>>>>>>> contains a picture of the entity associated with the information
>>>>>>> provided.
>>>>>>> - a video capture with participant info and the presentation 
>>>>>>> attribute
>>>>>>> would indicate that the presentation video is associated with the
>>>>>>> participant but could contain any video content.
>>>>>>> - a text capture with participant info would indicate that the 
>>>>>>> text is
>>>>>>> generated from the actual participant.
>>>>>>> - a text capture with participant info and the presentation 
>>>>>>> attribute
>>>>>>> would indicate that the text is associated with the participant but
>>>>>>> could contain any text content."
>>>>>>>
>>>>>>> Comments?
>>>>>>>
>>>>>>>
>>>>>>> Regards, Christian
>>>>>>>
>>>>>>> _______________________________________________
>>>>>>> clue mailing list
>>>>>>> clue@ietf.org
>>>>>>> https://www.ietf.org/mailman/listinfo/clue
>>>>>>>
>>>>>>
>>>>>> _______________________________________________
>>>>>> clue mailing list
>>>>>> clue@ietf.org
>>>>>> https://www.ietf.org/mailman/listinfo/clue
>>>>>>
>>>>>
>>>>> _______________________________________________
>>>>> clue mailing list
>>>>> clue@ietf.org
>>>>> https://www.ietf.org/mailman/listinfo/clue
>>>>>
>>>>
>>>>
>>>
>>>
>>
>>
>
> _______________________________________________
> clue mailing list
> clue@ietf.org
> https://www.ietf.org/mailman/listinfo/clue
>


From nobody Tue Apr  1 03:28:01 2014
Return-Path: <christer.holmberg@ericsson.com>
X-Original-To: clue@ietfa.amsl.com
Delivered-To: clue@ietfa.amsl.com
Received: from localhost (ietfa.amsl.com [127.0.0.1]) by ietfa.amsl.com (Postfix) with ESMTP id 961AB1A8028 for <clue@ietfa.amsl.com>; Tue,  1 Apr 2014 03:27:58 -0700 (PDT)
X-Virus-Scanned: amavisd-new at amsl.com
X-Spam-Flag: NO
X-Spam-Score: -1.45
X-Spam-Level: 
X-Spam-Status: No, score=-1.45 tagged_above=-999 required=5 tests=[BAYES_00=-1.9, HELO_EQ_SE=0.35, HTML_MESSAGE=0.001, J_CHICKENPOX_111=0.6, J_CHICKENPOX_14=0.6, J_CHICKENPOX_15=0.6, J_CHICKENPOX_54=0.6, RCVD_IN_DNSWL_MED=-2.3, SPF_PASS=-0.001] autolearn=no
Received: from mail.ietf.org ([4.31.198.44]) by localhost (ietfa.amsl.com [127.0.0.1]) (amavisd-new, port 10024) with ESMTP id Sk5jQG_CV9UV for <clue@ietfa.amsl.com>; Tue,  1 Apr 2014 03:27:53 -0700 (PDT)
Received: from mailgw1.ericsson.se (mailgw1.ericsson.se [193.180.251.45]) by ietfa.amsl.com (Postfix) with ESMTP id 0A1E91A0825 for <clue@ietf.org>; Tue,  1 Apr 2014 03:27:51 -0700 (PDT)
X-AuditID: c1b4fb2d-b7f328e0000012ab-83-533a94a349b1
Received: from ESESSHC008.ericsson.se (Unknown_Domain [153.88.253.124]) by mailgw1.ericsson.se (Symantec Mail Security) with SMTP id 3C.B5.04779.3A49A335; Tue,  1 Apr 2014 12:27:47 +0200 (CEST)
Received: from ESESSMB209.ericsson.se ([169.254.9.213]) by ESESSHC008.ericsson.se ([153.88.183.42]) with mapi id 14.03.0174.001; Tue, 1 Apr 2014 12:27:47 +0200
From: Christer Holmberg <christer.holmberg@ericsson.com>
To: "Robert Hansen (rohanse2)" <rohanse2@cisco.com>, "clue@ietf.org" <clue@ietf.org>
Thread-Topic: Using BUNDLE with CLUE
Thread-Index: Ac9Nkg5IUPNx7ldUTDK8kHk510O6AQAAOGwQ
Date: Tue, 1 Apr 2014 10:27:46 +0000
Message-ID: <7594FB04B1934943A5C02806D1A2204B1D271311@ESESSMB209.ericsson.se>
References: <C6252EA94E00E44EADC3A2FEB59D4402020DF93F@xmb-aln-x07.cisco.com>
In-Reply-To: <C6252EA94E00E44EADC3A2FEB59D4402020DF93F@xmb-aln-x07.cisco.com>
Accept-Language: en-US
Content-Language: en-US
X-MS-Has-Attach: 
X-MS-TNEF-Correlator: 
x-originating-ip: [153.88.183.18]
Content-Type: multipart/alternative; boundary="_000_7594FB04B1934943A5C02806D1A2204B1D271311ESESSMB209erics_"
MIME-Version: 1.0
X-Brightmail-Tracker: H4sIAAAAAAAAA+NgFrrJLMWRmVeSWpSXmKPExsUyM+Jvje7iKVbBBme3aFvsP3WZ2eLT/w/s DkweU35vZPVYsuQnUwBTFJdNSmpOZllqkb5dAlfG/bZJzAVzu5gqdn34wNjA2PSYsYuRk0NC wETi3ZVj7BC2mMSFe+vZuhi5OIQEDjNKzDrRwgjhLGaUOLVqOXMXIwcHm4CFRPc/bRBTRCBM 4sj5SJBeYQFlicWXZoDNERFQkbh37gAzhG0kcarnBNguFqD4p6lb2UBsXgFfiUc79rGA2EIC PhL/9vwGq+EEii94Pgkszgh0z/dTa5hAbGYBcYlbT+YzQdwpILFkz3lmCFtU4uXjf6wQtqLE x1f7GCHq8yWW7/kHtUtQ4uTMJywTGEVmIRk1C0nZLCRlEHEdiQW7P7FB2NoSyxa+Zoaxzxx4 zIQsvoCRfRUje25iZk56ueEmRmD0HNzyW3cH46lzIocYpTlYlMR5P7x1DhISSE8sSc1OTS1I LYovKs1JLT7EyMTBKdXAmOu6Ou21TtPyqS/PnTisFbCwpbOwfnLPq82MkoK9OyW3djv5+dZY L4lmqezt+193s1/gs25a/Po33a4XTM0r09YwP7wwOeDP/Yn9oUVFWVvq3jfIJPDUtnx6b8W0 WzJF4tPaf6J32/l/Sh0+GMTNtmSdQLrwmUe/U5QPNZa+f/x+q0D2W+1MJZbijERDLeai4kQA kiOHO2wCAAA=
Archived-At: http://mailarchive.ietf.org/arch/msg/clue/O1rMBoxsM6-cUd4K6eHeP4B60MA
Subject: Re: [clue] Using BUNDLE with CLUE
X-BeenThere: clue@ietf.org
X-Mailman-Version: 2.1.15
Precedence: list
List-Id: CLUE - ControLling mUltiple streams for TElepresence <clue.ietf.org>
List-Unsubscribe: <https://www.ietf.org/mailman/options/clue>, <mailto:clue-request@ietf.org?subject=unsubscribe>
List-Archive: <http://www.ietf.org/mail-archive/web/clue/>
List-Post: <mailto:clue@ietf.org>
List-Help: <mailto:clue-request@ietf.org?subject=help>
List-Subscribe: <https://www.ietf.org/mailman/listinfo/clue>, <mailto:clue-request@ietf.org?subject=subscribe>
X-List-Received-Date: Tue, 01 Apr 2014 10:27:58 -0000

--_000_7594FB04B1934943A5C02806D1A2204B1D271311ESESSMB209erics_
Content-Type: text/plain; charset="us-ascii"
Content-Transfer-Encoding: quoted-printable

Hi,

Q1: If the remote part has indicated support of BUNDLE in the first answer,=
 then I see no problem in adding the BUNDLE m- lines in the BAS offer. But,=
 of course you can also first send a BAS offer containing the initial m- li=
nes, and after that send a third offer with the BUNDLE m- lines. I am not s=
ure whether we would need to mandate one way or another, but it would proba=
bly be good to describe both.

Q2: Regarding your example, I hope we'll use a media feature tag instead of=
 group:CLUE to indicate support of CLUE. At least nobody has objected to it=
 :)

Q3: Regarding your text on Demultiplexing of RTP, the status has changed.

You CAN re-use the same PT value in multiple m- lines (due to the limited n=
umber range), as long as the codec configuration associated with the PT val=
ue is the same. The plan is that BUNDLE will specify usage of an identifier=
, associated with an m- line, which is inserted into the RTP packet by the =
sender, so that the receiver knows to which m- line the RTP packet belongs.=
 This will reflected in bundle-06, which will be submitted in the near futu=
re.

Q4: Regarding the directionality and number of BUNDLE groups, using a BUNDL=
E group per direction was one option. Another could be to use a BUNDLE grou=
p per media type. Etc. There could e.g. be QoS related reasons why you want=
 to use different BUNDLE groups (read: different media 5-tuples). I guess t=
he question is whether it would be useful for CLUE entities to be able to i=
ndicate which "BUNDLE configurations" they support. "Everything in one BUND=
LE group" could be default.

Regards,

Christer

From: clue [mailto:clue-bounces@ietf.org] On Behalf Of Robert Hansen (rohan=
se2)
Sent: 1. huhtikuuta 2014 13:07
To: clue@ietf.org
Subject: [clue] Using BUNDLE with CLUE

Introduction:

A fundamental aspect of CLUE is the sending of multiple streams of media. W=
e are using conventional SDP to specify the encodings we support, and hence=
 each media stream necessitates a separate m-line. For most use-cases, howe=
ver, using a separate port per m-line is suboptimal: it means opening more =
ports, more NAT work, more resources consumed for ICE, etc. As such, a meth=
od to allow multiple m-lines to share the same 5-tuple address is highly de=
sirable.

Until now the signalling work has focused on the demultiplexed case. As suc=
h, this is an attempt to evaluate the use of BUNDLE with CLUE - those with =
a much better understanding of BUNDLE than I will be able to correct the mi=
stakes I'll inevitably make.

We have generally figured that CLUE would not raise any particular issues w=
ith BUNDLE, given that CLUE now uses separate m-lines for encodings, but we=
 need to go through and make sure, as well as providing guidance in the CLU=
E documentation on how BUNDLE interacts with it.

Obviously I'm not going to duplicate the BUNDLE draft here, instead I'll ju=
st make reference to the draft.

Initial offer/answer:

One of the principle concerns of BUNDLE is to ensure that a receiver does n=
ot receive an SDP it considers invalid due to its lack of support for bundl=
ing m-lines. The fact that we recommend that CLUE-controlled m-lines are no=
t included in the initial O/A means that I think that we can combine the ex=
tra O/A that BUNDLE normally adds over an unBUNDLEd call with one of the O/=
As required by CLUE...

Adding CLUE-controlled m-lines during the Bundle Address Synchronization (B=
AS) offer:

Having completed the initial O/A and established BUNDLE support the initial=
 offerer needs to send a new offer to synchronise the BUNDLE addresses and =
make any intermediary devices aware of the addresses in use.

In CLUE, the subsequent offer is also when we would like to start adding CL=
UE-controlled m-lines. However, section 6.4.3. of the BUNDLE specification =
warns that it important that the BAS offer is accepted, and while it makes =
clear that the offerer MAY change the SDP, it warns to avoid changes that c=
ould cause the answerer to reject the new offer.

CLUE definitely needs to provide guidance here. My belief is that adding th=
e CLUE-controlled m-lines at this stage should not increase the chance of t=
he offer as a whole being rejected, so long as they share the same media ty=
pes and attributes as the existing media lines. This would also help resolv=
e the glare issue at the start of a CLUE call: since a BAS is mandatory in =
BUNDLE this provides an obvious 'who should reINVITE first' case for CLUE, =
where the initial offerer has to send an new offer even if they don't have =
any CLUE encodings to add.

If we did feel that adding new m-lines in the BAS offer was too high a risk=
 then BUNDLE and CLUE become sequential: BAS should be done first, and then=
 CLUE-controlled m-lines should be added in a subsequent INVITE. In this ca=
se BUNDLE would include one extra O/A compared to the unBUNDLEd case.

I've included an example below, modifying the example from the BUNDLE draft=
 to show the initial offerer using the BAS offer to also add CLUE-controlle=
d media.

Initial offer example - the offer includes an audio and video line with uni=
que ports, both in the same BUNDLE group (indicating that the offerer suppo=
rts BUNDLE and wants to multiplex these media flows). There is also a data =
channel that will be used for CLUE, which is included in the CLUE group.

a=3Dgroup:BUNDLE foo bar
a=3Dgroup:CLUE zen
m=3Daudio 10000 RTP/AVP 0 8 106
...
a=3Dmid:foo
m=3Dvideo 10002 RTP/AVP 96 97
...
a=3Dmid:bar
m=3Dapplication 10004 SCTP/DTLS 10004
...
a=3Dmid:zen

Initial answer example - the answer picks a local BUNDLE address and (via o=
rdering in the group attribute) selects an address for the offerer.

a=3Dgroup:BUNDLE foo bar
a=3Dgroup:CLUE zen
m=3Daudio 20000 RTP/AVP 106
...
a=3Dmid:foo
m=3Dvideo 20000 RTP/AVP 96
...
a=3Dmid:bar
m=3Dapplication 20002 SCTP/DTLS 20002
...
a=3Dmid:zen

Subsequent offer example - the offerer synchronises addresses and adds CLUE=
-controlled media lines using the same address; these are included in both =
the BUNDLE and CLUE groups.

a=3Dgroup:BUNDLE foo bar enc1 enc2 enc3
a=3Dgroup:CLUE zen enc1 enc2 enc3
m=3Daudio 10000 RTP/AVP 0 8 106
...
a=3Dmid:foo
m=3Dvideo 10000 RTP/AVP 96 97
...
a=3Dmid:bar
m=3Dapplication 10004 SCTP/DTLS 10004
...
a=3Dmid:zen
m=3Dvideo 10000 RTP/AVP 96 97
...
a=3Dmid:enc1
a=3Dlabel:1
m=3Dvideo 10000 RTP/AVP 96 97
...
a=3Dmid:enc2
a=3Dlabel:2
m=3Dvideo 10000 RTP/AVP 96 97
...
a=3Dmid:enc3
a=3Dlabel:3

Answerer rejects BUNDLE:

If the answer does not support BUNDLE then the offerer continues as in the =
disaggregated case, though it may decide to offer fewer streams, or even no=
t do CLUE at all (in the latter case it should reINVITE and remove the CLUE=
 group and data channel).

Effects of multiplexing on CLUE-relevant SDP attributes:

The BUNDLE draft specifies how certain SDP attributes are affected by multi=
plexing the m-lines, and draft-ietf-mmusic-sdp-mux-attributes describes how=
 multiplexing affects many other SDP attributes. I can't see any CLUE-speci=
fic issues here: the 'label' attribute is unaffected by multiplexing, nor i=
s the directionality of the m-lines, and the 'mid' attribute needed by CLUE=
 is also needed by BUNDLE.

Directionality of CLUE-controlled media:

CLUE-controlled m-lines are currently unidirectional. I don't believe this =
raises any specific BUNDLE issues. Christer suggests that potentially the e=
ncodings on each side could be in separate BUNDLE groups - I suggest we use=
 the same group for both sides, which would also make it easier to move to =
bidirectional streams if we ever wanted to do that.

BUNDLE support in CLUE:

While in most cases multiplexing m-lines onto a single 5-tuple will be pref=
erable we do have use cases involving disaggregated media. Because of this,=
 and because my understanding of BUNDLE is that it does not support the agg=
regated to disaggregated use case (where one device send/receives media ass=
ociated with multiple m-lines on a single IP/port, while the other sends/re=
ceives the media associated with multiple m-lines on multiple IP/ports) I d=
on't see a reason to mandate BUNDLE support for CLUE devices that only oper=
ate in disaggregated mode. As such, while CLUE should document how to use B=
UNDLE to multiplex media streams, it shouldn't be mandatory for CLUE device=
s to support it.

Demultiplexing of RTP:

Media flows from the same BUNDLE group all belong to the same RTP session; =
hence a receiver needs a method is needed to separate the various capture e=
ncodings. The method currently defined in BUNDLE is very straightforward: a=
ll of the dynamic payload types in each m-line in a BUNDLE group must be un=
ique; as such, any received RTP packet can be mapped to a specific m-line a=
nd codec.

This method, however, does suffer from the fact that the dynamic payload ty=
pe range is limited. In the case of many m-lines, of audio (with potentiall=
y many codecs) and video bundled together, and/or the increasing use of sca=
lable video codecs with multiple layers and the use of FEC repair flows, th=
is space may be insufficient.

There are other methods that could be used for this demultiplexing. For imp=
lementation with static sources or RTP mixers, streams will have a consiste=
nt SSRC, and hence the a=3Dssrc attribute can be supplied by the sender to =
allow the receiver to differentiate flows via their SSRC. For implementatio=
ns where the streams being sent do not have a static SSRC, such as source p=
rojection mixers, a specific stream-correlator can be used, such as an AppI=
D token specified by the sender in SDP and in an RTP header extension.

How stream demultiplexing is performed with BUNDLE is not a problem specifi=
c to CLUE, and is not a problem that should be solved in a CLUE-specific fa=
shion. However, it is one that we need to ensure is solved in a way consist=
ent with our use-cases, and the number of encodings we envisage allowing. G=
iven that the simultaneous sending of many capture encodings is a primary m=
otivator behind CLUE I believe that demultiplexing via payload will not be =
sufficient, and that we would need additional methods of differentiation.

If not all payload types must be unique then the question of under what cir=
cumstances codecs on two m-lines can share the same payload type becomes re=
levant.

Summary:

I believe BUNDLE is a good solution for how we can multiplex media streams =
when doing CLUE, and I can't see any CLUE-specific issues that would need t=
o be addressed in BUNDLE. I think the main decision CLUE would need to make=
 is whether do first do BUNDLE negotiation (including address synchronisati=
on) and then begin adding CLUE m-lines, or whether the CLUE media negotiati=
on process can begin with the BAS, essentially eliminating the extra O/A 'c=
ost' of BUNDLE.

--_000_7594FB04B1934943A5C02806D1A2204B1D271311ESESSMB209erics_
Content-Type: text/html; charset="us-ascii"
Content-Transfer-Encoding: quoted-printable

<html xmlns:v=3D"urn:schemas-microsoft-com:vml" xmlns:o=3D"urn:schemas-micr=
osoft-com:office:office" xmlns:w=3D"urn:schemas-microsoft-com:office:word" =
xmlns:m=3D"http://schemas.microsoft.com/office/2004/12/omml" xmlns=3D"http:=
//www.w3.org/TR/REC-html40">
<head>
<meta http-equiv=3D"Content-Type" content=3D"text/html; charset=3Dus-ascii"=
>
<meta name=3D"Generator" content=3D"Microsoft Word 14 (filtered medium)">
<style><!--
/* Font Definitions */
@font-face
	{font-family:Calibri;
	panose-1:2 15 5 2 2 2 4 3 2 4;}
@font-face
	{font-family:Tahoma;
	panose-1:2 11 6 4 3 5 4 4 2 4;}
/* Style Definitions */
p.MsoNormal, li.MsoNormal, div.MsoNormal
	{margin:0cm;
	margin-bottom:.0001pt;
	font-size:11.0pt;
	font-family:"Calibri","sans-serif";}
a:link, span.MsoHyperlink
	{mso-style-priority:99;
	color:blue;
	text-decoration:underline;}
a:visited, span.MsoHyperlinkFollowed
	{mso-style-priority:99;
	color:purple;
	text-decoration:underline;}
span.EmailStyle17
	{mso-style-type:personal;
	font-family:"Calibri","sans-serif";
	color:windowtext;}
span.EmailStyle18
	{mso-style-type:personal-reply;
	font-family:"Calibri","sans-serif";
	color:#1F497D;}
.MsoChpDefault
	{mso-style-type:export-only;
	font-size:10.0pt;}
@page WordSection1
	{size:612.0pt 792.0pt;
	margin:72.0pt 72.0pt 72.0pt 72.0pt;}
div.WordSection1
	{page:WordSection1;}
--></style><!--[if gte mso 9]><xml>
<o:shapedefaults v:ext=3D"edit" spidmax=3D"1026" />
</xml><![endif]--><!--[if gte mso 9]><xml>
<o:shapelayout v:ext=3D"edit">
<o:idmap v:ext=3D"edit" data=3D"1" />
</o:shapelayout></xml><![endif]-->
</head>
<body lang=3D"EN-US" link=3D"blue" vlink=3D"purple">
<div class=3D"WordSection1">
<p class=3D"MsoNormal"><span style=3D"color:#1F497D">Hi,<o:p></o:p></span><=
/p>
<p class=3D"MsoNormal"><span style=3D"color:#1F497D"><o:p>&nbsp;</o:p></spa=
n></p>
<p class=3D"MsoNormal"><b><span style=3D"color:#1F497D">Q1</span></b><span =
style=3D"color:#1F497D">: If the remote part has indicated support of BUNDL=
E in the first answer, then I see no problem in adding the BUNDLE m- lines =
in the BAS offer. But, of course you can
 also first send a BAS offer containing the initial m- lines, and after tha=
t send a third offer with the BUNDLE m- lines. I am not sure whether we wou=
ld need to mandate one way or another, but it would probably be good to des=
cribe both.<o:p></o:p></span></p>
<p class=3D"MsoNormal"><span style=3D"color:#1F497D"><o:p>&nbsp;</o:p></spa=
n></p>
<p class=3D"MsoNormal"><b><span style=3D"color:#1F497D">Q2</span></b><span =
style=3D"color:#1F497D">: Regarding your example, I hope we&#8217;ll use a =
media feature tag instead of group:CLUE to indicate support of CLUE. At lea=
st nobody has objected to it :)<o:p></o:p></span></p>
<p class=3D"MsoNormal"><span style=3D"color:#1F497D"><o:p>&nbsp;</o:p></spa=
n></p>
<p class=3D"MsoNormal"><b><span style=3D"color:#1F497D">Q3</span></b><span =
style=3D"color:#1F497D">: Regarding your text on Demultiplexing of RTP, the=
 status has changed.</span><span lang=3D"EN-GB" style=3D"color:#1F497D"><o:=
p></o:p></span></p>
<p class=3D"MsoNormal"><span lang=3D"EN-GB" style=3D"color:#1F497D"><o:p>&n=
bsp;</o:p></span></p>
<p class=3D"MsoNormal"><span style=3D"color:#1F497D">You CAN re-use the sam=
e PT value in multiple m- lines (due to the limited number range), as long =
as the codec configuration associated with the PT value is the same. The pl=
an is that BUNDLE will specify usage
 of an identifier, associated with an m- line, which is inserted into the R=
TP packet by the sender, so that the receiver knows to which m- line the RT=
P packet belongs. This will reflected in bundle-06, which will be submitted=
 in the near future.<o:p></o:p></span></p>
<p class=3D"MsoNormal"><span style=3D"color:#1F497D"><o:p>&nbsp;</o:p></spa=
n></p>
<p class=3D"MsoNormal"><b><span style=3D"color:#1F497D">Q4</span></b><span =
style=3D"color:#1F497D">: Regarding the directionality and number of BUNDLE=
 groups, using a BUNDLE group per direction was one option. Another could b=
e to use a BUNDLE group per media type.
 Etc. There could e.g. be QoS related reasons why you want to use different=
 BUNDLE groups (read: different media 5-tuples). I guess the question is wh=
ether it would be useful for CLUE entities to be able to indicate which &#8=
220;BUNDLE configurations&#8221; they support.
 &#8220;Everything in one BUNDLE group&#8221; could be default.<o:p></o:p><=
/span></p>
<p class=3D"MsoNormal"><span style=3D"color:#1F497D"><o:p>&nbsp;</o:p></spa=
n></p>
<p class=3D"MsoNormal"><span style=3D"color:#1F497D">Regards,<o:p></o:p></s=
pan></p>
<p class=3D"MsoNormal"><span style=3D"color:#1F497D"><o:p>&nbsp;</o:p></spa=
n></p>
<p class=3D"MsoNormal"><span style=3D"color:#1F497D">Christer<o:p></o:p></s=
pan></p>
<p class=3D"MsoNormal"><span style=3D"color:#1F497D"><o:p>&nbsp;</o:p></spa=
n></p>
<div>
<div style=3D"border:none;border-top:solid #B5C4DF 1.0pt;padding:3.0pt 0cm =
0cm 0cm">
<p class=3D"MsoNormal"><b><span style=3D"font-size:10.0pt;font-family:&quot=
;Tahoma&quot;,&quot;sans-serif&quot;">From:</span></b><span style=3D"font-s=
ize:10.0pt;font-family:&quot;Tahoma&quot;,&quot;sans-serif&quot;"> clue [ma=
ilto:clue-bounces@ietf.org]
<b>On Behalf Of </b>Robert Hansen (rohanse2)<br>
<b>Sent:</b> 1. huhtikuuta 2014 13:07<br>
<b>To:</b> clue@ietf.org<br>
<b>Subject:</b> [clue] Using BUNDLE with CLUE<o:p></o:p></span></p>
</div>
</div>
<p class=3D"MsoNormal"><o:p>&nbsp;</o:p></p>
<p class=3D"MsoNormal"><span lang=3D"EN-GB">Introduction:<o:p></o:p></span>=
</p>
<p class=3D"MsoNormal"><span lang=3D"EN-GB"><o:p>&nbsp;</o:p></span></p>
<p class=3D"MsoNormal"><span lang=3D"EN-GB">A fundamental aspect of CLUE is=
 the sending of multiple streams of media. We are using conventional SDP to=
 specify the encodings we support, and hence each media stream necessitates=
 a separate m-line. For most use-cases,
 however, using a separate port per m-line is suboptimal: it means opening =
more ports, more NAT work, more resources consumed for ICE, etc. As such, a=
 method to allow multiple m-lines to share the same 5-tuple address is high=
ly desirable.<o:p></o:p></span></p>
<p class=3D"MsoNormal"><span lang=3D"EN-GB"><o:p>&nbsp;</o:p></span></p>
<p class=3D"MsoNormal"><span lang=3D"EN-GB">Until now the signalling work h=
as focused on the demultiplexed case. As such, this is an attempt to evalua=
te the use of BUNDLE with CLUE - those with a much better understanding of =
BUNDLE than I will be able to correct
 the mistakes I'll inevitably make.<o:p></o:p></span></p>
<p class=3D"MsoNormal"><span lang=3D"EN-GB"><o:p>&nbsp;</o:p></span></p>
<p class=3D"MsoNormal"><span lang=3D"EN-GB">We have generally figured that =
CLUE would not raise any particular issues with BUNDLE, given that CLUE now=
 uses separate m-lines for encodings, but we need to go through and make su=
re, as well as providing guidance in
 the CLUE documentation on how BUNDLE interacts with it.<o:p></o:p></span><=
/p>
<p class=3D"MsoNormal"><span lang=3D"EN-GB"><o:p>&nbsp;</o:p></span></p>
<p class=3D"MsoNormal"><span lang=3D"EN-GB">Obviously I'm not going to dupl=
icate the BUNDLE draft here, instead I'll just make reference to the draft.=
<o:p></o:p></span></p>
<p class=3D"MsoNormal"><span lang=3D"EN-GB"><o:p>&nbsp;</o:p></span></p>
<p class=3D"MsoNormal"><span lang=3D"EN-GB">Initial offer/answer:<o:p></o:p=
></span></p>
<p class=3D"MsoNormal"><span lang=3D"EN-GB"><o:p>&nbsp;</o:p></span></p>
<p class=3D"MsoNormal"><span lang=3D"EN-GB">One of the principle concerns o=
f BUNDLE is to ensure that a receiver does not receive an SDP it considers =
invalid due to its lack of support for bundling m-lines. The fact that we r=
ecommend that CLUE-controlled m-lines
 are not included in the initial O/A means that I think that we can combine=
 the extra O/A that BUNDLE normally adds over an unBUNDLEd call with one of=
 the O/As required by CLUE...<o:p></o:p></span></p>
<p class=3D"MsoNormal"><span lang=3D"EN-GB"><o:p>&nbsp;</o:p></span></p>
<p class=3D"MsoNormal"><span lang=3D"EN-GB">Adding CLUE-controlled m-lines =
during the Bundle Address Synchronization (BAS) offer:<o:p></o:p></span></p=
>
<p class=3D"MsoNormal"><span lang=3D"EN-GB"><o:p>&nbsp;</o:p></span></p>
<p class=3D"MsoNormal"><span lang=3D"EN-GB">Having completed the initial O/=
A and established BUNDLE support the initial offerer needs to send a new of=
fer to synchronise the BUNDLE addresses and make any intermediary devices a=
ware of the addresses in use.<o:p></o:p></span></p>
<p class=3D"MsoNormal"><span lang=3D"EN-GB"><o:p>&nbsp;</o:p></span></p>
<p class=3D"MsoNormal"><span lang=3D"EN-GB">In CLUE, the subsequent offer i=
s also when we would like to start adding CLUE-controlled m-lines. However,=
 section 6.4.3. of the BUNDLE specification warns that it important that th=
e BAS offer is accepted, and while it
 makes clear that the offerer MAY change the SDP, it warns to avoid changes=
 that could cause the answerer to reject the new offer.<o:p></o:p></span></=
p>
<p class=3D"MsoNormal"><span lang=3D"EN-GB"><o:p>&nbsp;</o:p></span></p>
<p class=3D"MsoNormal"><span lang=3D"EN-GB">CLUE definitely needs to provid=
e guidance here. My belief is that adding the CLUE-controlled m-lines at th=
is stage should not increase the chance of the offer as a whole being rejec=
ted, so long as they share the same
 media types and attributes as the existing media lines. This would also he=
lp resolve the glare issue at the start of a CLUE call: since a BAS is mand=
atory in BUNDLE this provides an obvious 'who should reINVITE first' case f=
or CLUE, where the initial offerer
 has to send an new offer even if they don't have any CLUE encodings to add=
.<o:p></o:p></span></p>
<p class=3D"MsoNormal"><span lang=3D"EN-GB"><o:p>&nbsp;</o:p></span></p>
<p class=3D"MsoNormal"><span lang=3D"EN-GB">If we did feel that adding new =
m-lines in the BAS offer was too high a risk then BUNDLE and CLUE become se=
quential: BAS should be done first, and then CLUE-controlled m-lines should=
 be added in a subsequent INVITE. In
 this case BUNDLE would include one extra O/A compared to the unBUNDLEd cas=
e.<o:p></o:p></span></p>
<p class=3D"MsoNormal"><span lang=3D"EN-GB"><o:p>&nbsp;</o:p></span></p>
<p class=3D"MsoNormal"><span lang=3D"EN-GB">I've included an example below,=
 modifying the example from the BUNDLE draft to show the initial offerer us=
ing the BAS offer to also add CLUE-controlled media.<o:p></o:p></span></p>
<p class=3D"MsoNormal"><span lang=3D"EN-GB"><o:p>&nbsp;</o:p></span></p>
<p class=3D"MsoNormal"><span lang=3D"EN-GB">Initial offer example - the off=
er includes an audio and video line with unique ports, both in the same BUN=
DLE group (indicating that the offerer supports BUNDLE and wants to multipl=
ex these media flows). There is also
 a data channel that will be used for CLUE, which is included in the CLUE g=
roup.<o:p></o:p></span></p>
<p class=3D"MsoNormal"><span lang=3D"EN-GB"><o:p>&nbsp;</o:p></span></p>
<p class=3D"MsoNormal"><span lang=3D"EN-GB">a=3Dgroup:BUNDLE foo bar<o:p></=
o:p></span></p>
<p class=3D"MsoNormal"><span lang=3D"EN-GB">a=3Dgroup:CLUE zen<o:p></o:p></=
span></p>
<p class=3D"MsoNormal"><span lang=3D"EN-GB">m=3Daudio 10000 RTP/AVP 0 8 106=
<o:p></o:p></span></p>
<p class=3D"MsoNormal"><span lang=3D"EN-GB">...<o:p></o:p></span></p>
<p class=3D"MsoNormal"><span lang=3D"EN-GB">a=3Dmid:foo<o:p></o:p></span></=
p>
<p class=3D"MsoNormal"><span lang=3D"EN-GB">m=3Dvideo 10002 RTP/AVP 96 97<o=
:p></o:p></span></p>
<p class=3D"MsoNormal"><span lang=3D"EN-GB">...<o:p></o:p></span></p>
<p class=3D"MsoNormal"><span lang=3D"EN-GB">a=3Dmid:bar<o:p></o:p></span></=
p>
<p class=3D"MsoNormal"><span lang=3D"EN-GB">m=3Dapplication 10004 SCTP/DTLS=
 10004<o:p></o:p></span></p>
<p class=3D"MsoNormal"><span lang=3D"EN-GB">...<o:p></o:p></span></p>
<p class=3D"MsoNormal"><span lang=3D"EN-GB">a=3Dmid:zen<o:p></o:p></span></=
p>
<p class=3D"MsoNormal"><span lang=3D"EN-GB"><o:p>&nbsp;</o:p></span></p>
<p class=3D"MsoNormal"><span lang=3D"EN-GB">Initial answer example - the an=
swer picks a local BUNDLE address and (via ordering in the group attribute)=
 selects an address for the offerer.<o:p></o:p></span></p>
<p class=3D"MsoNormal"><span lang=3D"EN-GB"><o:p>&nbsp;</o:p></span></p>
<p class=3D"MsoNormal"><span lang=3D"EN-GB">a=3Dgroup:BUNDLE foo bar<o:p></=
o:p></span></p>
<p class=3D"MsoNormal"><span lang=3D"EN-GB">a=3Dgroup:CLUE zen<o:p></o:p></=
span></p>
<p class=3D"MsoNormal"><span lang=3D"EN-GB">m=3Daudio 20000 RTP/AVP 106<o:p=
></o:p></span></p>
<p class=3D"MsoNormal"><span lang=3D"EN-GB">...<o:p></o:p></span></p>
<p class=3D"MsoNormal"><span lang=3D"EN-GB">a=3Dmid:foo<o:p></o:p></span></=
p>
<p class=3D"MsoNormal"><span lang=3D"EN-GB">m=3Dvideo 20000 RTP/AVP 96<o:p>=
</o:p></span></p>
<p class=3D"MsoNormal"><span lang=3D"EN-GB">...<o:p></o:p></span></p>
<p class=3D"MsoNormal"><span lang=3D"EN-GB">a=3Dmid:bar<o:p></o:p></span></=
p>
<p class=3D"MsoNormal"><span lang=3D"EN-GB">m=3Dapplication 20002 SCTP/DTLS=
 20002<o:p></o:p></span></p>
<p class=3D"MsoNormal"><span lang=3D"EN-GB">...<o:p></o:p></span></p>
<p class=3D"MsoNormal"><span lang=3D"EN-GB">a=3Dmid:zen<o:p></o:p></span></=
p>
<p class=3D"MsoNormal"><span lang=3D"EN-GB"><o:p>&nbsp;</o:p></span></p>
<p class=3D"MsoNormal"><span lang=3D"EN-GB">Subsequent offer example - the =
offerer synchronises addresses and adds CLUE-controlled media lines using t=
he same address; these are included in both the BUNDLE and CLUE groups.<o:p=
></o:p></span></p>
<p class=3D"MsoNormal"><span lang=3D"EN-GB"><o:p>&nbsp;</o:p></span></p>
<p class=3D"MsoNormal"><span lang=3D"EN-GB">a=3Dgroup:BUNDLE foo bar enc1 e=
nc2 enc3<o:p></o:p></span></p>
<p class=3D"MsoNormal"><span lang=3D"EN-GB">a=3Dgroup:CLUE zen enc1 enc2 en=
c3<o:p></o:p></span></p>
<p class=3D"MsoNormal"><span lang=3D"EN-GB">m=3Daudio 10000 RTP/AVP 0 8 106=
<o:p></o:p></span></p>
<p class=3D"MsoNormal"><span lang=3D"EN-GB">...<o:p></o:p></span></p>
<p class=3D"MsoNormal"><span lang=3D"EN-GB">a=3Dmid:foo<o:p></o:p></span></=
p>
<p class=3D"MsoNormal"><span lang=3D"EN-GB">m=3Dvideo 10000 RTP/AVP 96 97<o=
:p></o:p></span></p>
<p class=3D"MsoNormal"><span lang=3D"EN-GB">...<o:p></o:p></span></p>
<p class=3D"MsoNormal"><span lang=3D"EN-GB">a=3Dmid:bar<o:p></o:p></span></=
p>
<p class=3D"MsoNormal"><span lang=3D"EN-GB">m=3Dapplication 10004 SCTP/DTLS=
 10004<o:p></o:p></span></p>
<p class=3D"MsoNormal"><span lang=3D"EN-GB">...<o:p></o:p></span></p>
<p class=3D"MsoNormal"><span lang=3D"EN-GB">a=3Dmid:zen<o:p></o:p></span></=
p>
<p class=3D"MsoNormal"><span lang=3D"EN-GB">m=3Dvideo 10000 RTP/AVP 96 97<o=
:p></o:p></span></p>
<p class=3D"MsoNormal"><span lang=3D"EN-GB">...<o:p></o:p></span></p>
<p class=3D"MsoNormal"><span lang=3D"EN-GB">a=3Dmid:enc1<o:p></o:p></span><=
/p>
<p class=3D"MsoNormal"><span lang=3D"EN-GB">a=3Dlabel:1<o:p></o:p></span></=
p>
<p class=3D"MsoNormal"><span lang=3D"EN-GB">m=3Dvideo 10000 RTP/AVP 96 97<o=
:p></o:p></span></p>
<p class=3D"MsoNormal"><span lang=3D"EN-GB">...<o:p></o:p></span></p>
<p class=3D"MsoNormal"><span lang=3D"EN-GB">a=3Dmid:enc2<o:p></o:p></span><=
/p>
<p class=3D"MsoNormal"><span lang=3D"EN-GB">a=3Dlabel:2<o:p></o:p></span></=
p>
<p class=3D"MsoNormal"><span lang=3D"EN-GB">m=3Dvideo 10000 RTP/AVP 96 97<o=
:p></o:p></span></p>
<p class=3D"MsoNormal"><span lang=3D"EN-GB">...<o:p></o:p></span></p>
<p class=3D"MsoNormal"><span lang=3D"EN-GB">a=3Dmid:enc3<o:p></o:p></span><=
/p>
<p class=3D"MsoNormal"><span lang=3D"EN-GB">a=3Dlabel:3<o:p></o:p></span></=
p>
<p class=3D"MsoNormal"><span lang=3D"EN-GB"><o:p>&nbsp;</o:p></span></p>
<p class=3D"MsoNormal"><span lang=3D"EN-GB">Answerer rejects BUNDLE:<o:p></=
o:p></span></p>
<p class=3D"MsoNormal"><span lang=3D"EN-GB"><o:p>&nbsp;</o:p></span></p>
<p class=3D"MsoNormal"><span lang=3D"EN-GB">If the answer does not support =
BUNDLE then the offerer continues as in the disaggregated case, though it m=
ay decide to offer fewer streams, or even not do CLUE at all (in the latter=
 case it should reINVITE and remove
 the CLUE group and data channel).<o:p></o:p></span></p>
<p class=3D"MsoNormal"><span lang=3D"EN-GB"><o:p>&nbsp;</o:p></span></p>
<p class=3D"MsoNormal"><span lang=3D"EN-GB">Effects of multiplexing on CLUE=
-relevant SDP attributes:<o:p></o:p></span></p>
<p class=3D"MsoNormal"><span lang=3D"EN-GB"><o:p>&nbsp;</o:p></span></p>
<p class=3D"MsoNormal"><span lang=3D"EN-GB">The BUNDLE draft specifies how =
certain SDP attributes are affected by multiplexing the m-lines, and draft-=
ietf-mmusic-sdp-mux-attributes describes how multiplexing affects many othe=
r SDP attributes. I can't see any CLUE-specific
 issues here: the 'label' attribute is unaffected by multiplexing, nor is t=
he directionality of the m-lines, and the 'mid' attribute needed by CLUE is=
 also needed by BUNDLE.<o:p></o:p></span></p>
<p class=3D"MsoNormal"><span lang=3D"EN-GB"><o:p>&nbsp;</o:p></span></p>
<p class=3D"MsoNormal"><span lang=3D"EN-GB">Directionality of CLUE-controll=
ed media:<o:p></o:p></span></p>
<p class=3D"MsoNormal"><span lang=3D"EN-GB"><o:p>&nbsp;</o:p></span></p>
<p class=3D"MsoNormal"><span lang=3D"EN-GB">CLUE-controlled m-lines are cur=
rently unidirectional. I don't believe this raises any specific BUNDLE issu=
es. Christer suggests that potentially the encodings on each side could be =
in separate BUNDLE groups - I suggest
 we use the same group for both sides, which would also make it easier to m=
ove to bidirectional streams if we ever wanted to do that.<o:p></o:p></span=
></p>
<p class=3D"MsoNormal"><span lang=3D"EN-GB"><o:p>&nbsp;</o:p></span></p>
<p class=3D"MsoNormal"><span lang=3D"EN-GB">BUNDLE support in CLUE:<o:p></o=
:p></span></p>
<p class=3D"MsoNormal"><span lang=3D"EN-GB"><o:p>&nbsp;</o:p></span></p>
<p class=3D"MsoNormal"><span lang=3D"EN-GB">While in most cases multiplexin=
g m-lines onto a single 5-tuple will be preferable we do have use cases inv=
olving disaggregated media. Because of this, and because my understanding o=
f BUNDLE is that it does not support
 the aggregated to disaggregated use case (where one device send/receives m=
edia associated with multiple m-lines on a single IP/port, while the other =
sends/receives the media associated with multiple m-lines on multiple IP/po=
rts) I don't see a reason to mandate
 BUNDLE support for CLUE devices that only operate in disaggregated mode. A=
s such, while CLUE should document how to use BUNDLE to multiplex media str=
eams, it shouldn't be mandatory for CLUE devices to support it.<o:p></o:p><=
/span></p>
<p class=3D"MsoNormal"><span lang=3D"EN-GB"><o:p>&nbsp;</o:p></span></p>
<p class=3D"MsoNormal"><span lang=3D"EN-GB">Demultiplexing of RTP:<o:p></o:=
p></span></p>
<p class=3D"MsoNormal"><span lang=3D"EN-GB"><o:p>&nbsp;</o:p></span></p>
<p class=3D"MsoNormal"><span lang=3D"EN-GB">Media flows from the same BUNDL=
E group all belong to the same RTP session; hence a receiver needs a method=
 is needed to separate the various capture encodings. The method currently =
defined in BUNDLE is very straightforward:
 all of the dynamic payload types in each m-line in a BUNDLE group must be =
unique; as such, any received RTP packet can be mapped to a specific m-line=
 and codec.<o:p></o:p></span></p>
<p class=3D"MsoNormal"><span lang=3D"EN-GB"><o:p>&nbsp;</o:p></span></p>
<p class=3D"MsoNormal"><span lang=3D"EN-GB">This method, however, does suff=
er from the fact that the dynamic payload type range is limited. In the cas=
e of many m-lines, of audio (with potentially many codecs) and video bundle=
d together, and/or the increasing use
 of scalable video codecs with multiple layers and the use of FEC repair fl=
ows, this space may be insufficient.<o:p></o:p></span></p>
<p class=3D"MsoNormal"><span lang=3D"EN-GB"><o:p>&nbsp;</o:p></span></p>
<p class=3D"MsoNormal"><span lang=3D"EN-GB">There are other methods that co=
uld be used for this demultiplexing. For implementation with static sources=
 or RTP mixers, streams will have a consistent SSRC, and hence the a=3Dssrc=
 attribute can be supplied by the sender
 to allow the receiver to differentiate flows via their SSRC. For implement=
ations where the streams being sent do not have a static SSRC, such as sour=
ce projection mixers, a specific stream-correlator can be used, such as an =
AppID token specified by the sender
 in SDP and in an RTP header extension.<o:p></o:p></span></p>
<p class=3D"MsoNormal"><span lang=3D"EN-GB"><o:p>&nbsp;</o:p></span></p>
<p class=3D"MsoNormal"><span lang=3D"EN-GB">How stream demultiplexing is pe=
rformed with BUNDLE is not a problem specific to CLUE, and is not a problem=
 that should be solved in a CLUE-specific fashion. However, it is one that =
we need to ensure is solved in a way
 consistent with our use-cases, and the number of encodings we envisage all=
owing. Given that the simultaneous sending of many capture encodings is a p=
rimary motivator behind CLUE I believe that demultiplexing via payload will=
 not be sufficient, and that we
 would need additional methods of differentiation.<o:p></o:p></span></p>
<p class=3D"MsoNormal"><span lang=3D"EN-GB"><o:p>&nbsp;</o:p></span></p>
<p class=3D"MsoNormal"><span lang=3D"EN-GB">If not all payload types must b=
e unique then the question of under what circumstances codecs on two m-line=
s can share the same payload type becomes relevant.<o:p></o:p></span></p>
<p class=3D"MsoNormal"><span lang=3D"EN-GB"><o:p>&nbsp;</o:p></span></p>
<p class=3D"MsoNormal"><span lang=3D"EN-GB">Summary:<o:p></o:p></span></p>
<p class=3D"MsoNormal"><span lang=3D"EN-GB"><o:p>&nbsp;</o:p></span></p>
<p class=3D"MsoNormal"><span lang=3D"EN-GB">I believe BUNDLE is a good solu=
tion for how we can multiplex media streams when doing CLUE, and I can't se=
e any CLUE-specific issues that would need to be addressed in BUNDLE. I thi=
nk the main decision CLUE would need
 to make is whether do first do BUNDLE negotiation (including address synch=
ronisation) and then begin adding CLUE m-lines, or whether the CLUE media n=
egotiation process can begin with the BAS, essentially eliminating the extr=
a O/A 'cost' of BUNDLE.<o:p></o:p></span></p>
</div>
</body>
</html>

--_000_7594FB04B1934943A5C02806D1A2204B1D271311ESESSMB209erics_--


From nobody Tue Apr  1 03:44:31 2014
Return-Path: <roberta.presta@unina.it>
X-Original-To: clue@ietfa.amsl.com
Delivered-To: clue@ietfa.amsl.com
Received: from localhost (ietfa.amsl.com [127.0.0.1]) by ietfa.amsl.com (Postfix) with ESMTP id 69CFC1A8034 for <clue@ietfa.amsl.com>; Tue,  1 Apr 2014 03:44:30 -0700 (PDT)
X-Virus-Scanned: amavisd-new at amsl.com
X-Spam-Flag: NO
X-Spam-Score: -0.031
X-Spam-Level: 
X-Spam-Status: No, score=-0.031 tagged_above=-999 required=5 tests=[BAYES_00=-1.9, HELO_EQ_IT=0.635, HOST_EQ_IT=1.245, SPF_PASS=-0.001, T_RP_MATCHES_RCVD=-0.01] autolearn=no
Received: from mail.ietf.org ([4.31.198.44]) by localhost (ietfa.amsl.com [127.0.0.1]) (amavisd-new, port 10024) with ESMTP id Tx1GZTVnhhi4 for <clue@ietfa.amsl.com>; Tue,  1 Apr 2014 03:44:29 -0700 (PDT)
Received: from smtp1.unina.it (smtp1.unina.it [192.132.34.61]) by ietfa.amsl.com (Postfix) with ESMTP id A6EE61A86DD for <clue@ietf.org>; Tue,  1 Apr 2014 03:44:22 -0700 (PDT)
Received: from [127.0.0.1] ([143.225.229.123]) (authenticated bits=0) by smtp1.unina.it (8.14.4/8.14.4) with ESMTP id s31AiIP7029004 (version=TLSv1/SSLv3 cipher=DHE-RSA-AES128-SHA bits=128 verify=NO) for <clue@ietf.org>; Tue, 1 Apr 2014 12:44:18 +0200
Message-ID: <533A9882.1090006@unina.it>
Date: Tue, 01 Apr 2014 12:44:18 +0200
From: Roberta Presta <roberta.presta@unina.it>
User-Agent: Mozilla/5.0 (Windows NT 6.1; WOW64; rv:24.0) Gecko/20100101 Thunderbird/24.4.0
MIME-Version: 1.0
To: clue@ietf.org
References: <49E45C59CA48264997FEBFB29B6BC2D617AC536624@CRPMBOXPRD07.polycom.com> <53329F6C.3070600@unina.it> <49E45C59CA48264997FEBFB29B6BC2D617ACF1AE9A@CRPMBOXPRD07.polycom.com> <53335996.5010608@nteczone.com>
In-Reply-To: <53335996.5010608@nteczone.com>
Content-Type: text/plain; charset=windows-1252; format=flowed
Content-Transfer-Encoding: 8bit
Archived-At: http://mailarchive.ietf.org/arch/msg/clue/yjMMi7pl7gZ0JKOz2g_p_jpQt34
Subject: Re: [clue] <globalCaptureEntry> in data model
X-BeenThere: clue@ietf.org
X-Mailman-Version: 2.1.15
Precedence: list
List-Id: CLUE - ControLling mUltiple streams for TElepresence <clue.ietf.org>
List-Unsubscribe: <https://www.ietf.org/mailman/options/clue>, <mailto:clue-request@ietf.org?subject=unsubscribe>
List-Archive: <http://www.ietf.org/mail-archive/web/clue/>
List-Post: <mailto:clue@ietf.org>
List-Help: <mailto:clue-request@ietf.org?subject=help>
List-Subscribe: <https://www.ietf.org/mailman/listinfo/clue>, <mailto:clue-request@ietf.org?subject=subscribe>
X-List-Received-Date: Tue, 01 Apr 2014 10:44:30 -0000

Christian and Mark,

thank you for your clarification.
I will correct the draft accordingly.

Roberta



Il 26/03/2014 23:49, Christian Groves ha scritto:
> Hello Roberta,
>
> I think Mark has captured my concern about individual captures being 
> used in a GCSE.
>
> Regards, Christian
>
> On 27/03/2014 6:26 AM, Duckworth, Mark wrote:
>>
>> Hi Roberta,
>>
>> Initially, I was thinking the same thing as you, and that is how I 
>> presented it on the list and at London. Then at London I changed my 
>> mind as a few people agreed that using only CSEs better matches the 
>> meaning of the global list. I think Christian first mentioned in 
>> email on the list 
>> <http://www.ietf.org/mail-archive/web/clue/current/msg03541.html> why 
>> it would be better to use only CSEs within the global CSE list rather 
>> than also including individual captures:
>>
>> “I'm also not too sure about GCEs listing individual captures. What 
>> does an individual capture mean at that level? The CSEs indicate what 
>> captures represent a scene. Why would a provider indicate that you 
>> need a set of captures and then contradict itself by suggesting 
>> another set of captures.”
>>
>> The semantic meaning of an item in the global CSE list is that it 
>> represents the entire advertisement. So we thought it didn’t make 
>> sense to specify individual captures which themselves weren’t even a 
>> representation of a complete scene. So by including only CSEs, and 
>> not individual captures, it helps enforce this meaning.
>>
>> Mark
>>
>> *From:*Roberta Presta [mailto:roberta.presta@unina.it]
>> *Sent:* Wednesday, March 26, 2014 5:36 AM
>> *To:* Duckworth, Mark; clue@ietf.org
>> *Subject:* Re: <globalCaptureEntry> in data model
>>
>> Hi Mark,
>>
>> there is of course a typo in the definition, sorry.
>> In my presentation in London I tried to highlight commonalities when 
>> dealing with "content lists" in simultaneous sets, multiple content 
>> captures and global capture entries.
>> I proposed to use for global capture entries the same content format 
>> of simultaneous sets and multiple content captures:
>>
>> "The content can be specified in terms of:
>> 1.A list of (homogeneous) media capture identifiers
>> 2.A list of (homogeneous) capture Scene Entry identifiers (used as 
>> shortcuts)"
>>
>> As far as I remember, there were no objections to that.
>> It does not seem to be in contrast with discussion on the mailing list.
>> Is there any reason why we should not include references to media 
>> captures within GCEs?
>>
>> Roberta
>>
>>
>> Il 25/03/2014 20:47, Duckworth, Mark ha scritto:
>>
>>     I think in London we agreed that a single item in the new Global
>>     CSE List would consist of a set of one or more CSEs. We don’t want
>>     to include direct references to Media Captures. If I’m remembering
>>     correctly, then in section 18 we should remove this part:
>>
>>     <xs:element name="captureIDREF" type="xs:IDREF"
>>
>>     minOccurs="0" maxOccurs="unbounded"/>
>>
>>     I think there is a typo in the first sentence:
>>
>>     “<globalCaptureEntry> represents a set of captures of the same
>>     media time...”
>>
>>     should be:
>>
>>     “<globalCaptureEntry> represents a set of captures of the same
>>     media *type*...”
>>
>>     and then I suggest change to clarify it is built from CSEs:
>>
>>     “<globalCaptureEntry> represents a set of Capture Scene Entries of
>>     the same media *type*...”
>>
>>     I’m in process of adding “Global CSE List” to the framework 
>> document.
>>
>>     Regards,
>>
>>     Mark
>>
>>
>>
>> _______________________________________________
>> clue mailing list
>> clue@ietf.org
>> https://www.ietf.org/mailman/listinfo/clue
>
> _______________________________________________
> clue mailing list
> clue@ietf.org
> https://www.ietf.org/mailman/listinfo/clue
>


From nobody Tue Apr  1 03:54:01 2014
Return-Path: <roberta.presta@unina.it>
X-Original-To: clue@ietfa.amsl.com
Delivered-To: clue@ietfa.amsl.com
Received: from localhost (ietfa.amsl.com [127.0.0.1]) by ietfa.amsl.com (Postfix) with ESMTP id B47551A8035 for <clue@ietfa.amsl.com>; Tue,  1 Apr 2014 03:54:00 -0700 (PDT)
X-Virus-Scanned: amavisd-new at amsl.com
X-Spam-Flag: NO
X-Spam-Score: -0.03
X-Spam-Level: 
X-Spam-Status: No, score=-0.03 tagged_above=-999 required=5 tests=[BAYES_00=-1.9, HELO_EQ_IT=0.635, HOST_EQ_IT=1.245, HTML_MESSAGE=0.001, SPF_PASS=-0.001, T_RP_MATCHES_RCVD=-0.01] autolearn=no
Received: from mail.ietf.org ([4.31.198.44]) by localhost (ietfa.amsl.com [127.0.0.1]) (amavisd-new, port 10024) with ESMTP id cQf68X-IEHmy for <clue@ietfa.amsl.com>; Tue,  1 Apr 2014 03:53:59 -0700 (PDT)
Received: from smtp1.unina.it (smtp1.unina.it [192.132.34.61]) by ietfa.amsl.com (Postfix) with ESMTP id 15F111A86DD for <clue@ietf.org>; Tue,  1 Apr 2014 03:53:58 -0700 (PDT)
Received: from [127.0.0.1] ([143.225.229.123]) (authenticated bits=0) by smtp1.unina.it (8.14.4/8.14.4) with ESMTP id s31ArrAV030485 (version=TLSv1/SSLv3 cipher=DHE-RSA-AES128-SHA bits=128 verify=NO); Tue, 1 Apr 2014 12:53:54 +0200
Message-ID: <533A9AC1.5060706@unina.it>
Date: Tue, 01 Apr 2014 12:53:53 +0200
From: Roberta Presta <roberta.presta@unina.it>
User-Agent: Mozilla/5.0 (Windows NT 6.1; WOW64; rv:24.0) Gecko/20100101 Thunderbird/24.4.0
MIME-Version: 1.0
To: "Duckworth, Mark" <Mark.Duckworth@polycom.com>, "clue@ietf.org" <clue@ietf.org>
References: <49E45C59CA48264997FEBFB29B6BC2D617ACF1B213@CRPMBOXPRD07.polycom.com>
In-Reply-To: <49E45C59CA48264997FEBFB29B6BC2D617ACF1B213@CRPMBOXPRD07.polycom.com>
Content-Type: multipart/alternative; boundary="------------000202030806010905000908"
Archived-At: http://mailarchive.ietf.org/arch/msg/clue/L-4o3UufNvtte0M0im09p_6XsQA
Subject: Re: [clue] Global CSE List in framework
X-BeenThere: clue@ietf.org
X-Mailman-Version: 2.1.15
Precedence: list
List-Id: CLUE - ControLling mUltiple streams for TElepresence <clue.ietf.org>
List-Unsubscribe: <https://www.ietf.org/mailman/options/clue>, <mailto:clue-request@ietf.org?subject=unsubscribe>
List-Archive: <http://www.ietf.org/mail-archive/web/clue/>
List-Post: <mailto:clue@ietf.org>
List-Help: <mailto:clue-request@ietf.org?subject=help>
List-Subscribe: <https://www.ietf.org/mailman/listinfo/clue>, <mailto:clue-request@ietf.org?subject=subscribe>
X-List-Received-Date: Tue, 01 Apr 2014 10:54:00 -0000

This is a multi-part message in MIME format.
--------------000202030806010905000908
Content-Type: text/plain; charset=ISO-8859-1; format=flowed
Content-Transfer-Encoding: 7bit

Hello Mark,

just a minor question about your definition:


Il 27/03/2014 18:10, Duckworth, Mark ha scritto:
> An Advertisement can include an optional global Capture Scene Entry 
> list, for each media type.  Each item in this list is a set of one or 
> more Capture Scene Entries.

Do you mean that we have a global capture scene entry list for video and 
a global capture scene entry list for audio?
Otherwise, what about "...global Capture Scene Entry list. Each item in 
the list is a set of one or more Capture Scene Entries of the same media 
type."? That last approach is more similar to the one adopted for the 
content of capture scenes.

Regards,

Roberta

--------------000202030806010905000908
Content-Type: text/html; charset=ISO-8859-1
Content-Transfer-Encoding: 7bit

<html>
  <head>
    <meta content="text/html; charset=ISO-8859-1"
      http-equiv="Content-Type">
  </head>
  <body text="#000000" bgcolor="#FFFFFF">
    <div class="moz-cite-prefix">Hello Mark, <br>
      <br>
      just a minor question about your definition:<br>
      <br>
      <br>
      Il 27/03/2014 18:10, Duckworth, Mark ha scritto:<br>
    </div>
    <blockquote
cite="mid:49E45C59CA48264997FEBFB29B6BC2D617ACF1B213@CRPMBOXPRD07.polycom.com"
      type="cite">
      <meta http-equiv="Content-Type" content="text/html;
        charset=ISO-8859-1">
      <meta name="Generator" content="Microsoft Word 14 (filtered
        medium)">
      <style><!--
/* Font Definitions */
@font-face
	{font-family:Calibri;
	panose-1:2 15 5 2 2 2 4 3 2 4;}
/* Style Definitions */
p.MsoNormal, li.MsoNormal, div.MsoNormal
	{margin:0in;
	margin-bottom:.0001pt;
	font-size:11.0pt;
	font-family:"Calibri","sans-serif";}
a:link, span.MsoHyperlink
	{mso-style-priority:99;
	color:blue;
	text-decoration:underline;}
a:visited, span.MsoHyperlinkFollowed
	{mso-style-priority:99;
	color:purple;
	text-decoration:underline;}
span.EmailStyle17
	{mso-style-type:personal-compose;
	font-family:"Calibri","sans-serif";
	color:windowtext;}
.MsoChpDefault
	{mso-style-type:export-only;
	font-family:"Calibri","sans-serif";}
@page WordSection1
	{size:8.5in 11.0in;
	margin:1.0in 1.0in 1.0in 1.0in;}
div.WordSection1
	{page:WordSection1;}
--></style><!--[if gte mso 9]><xml>
<o:shapedefaults v:ext="edit" spidmax="1026" />
</xml><![endif]--><!--[if gte mso 9]><xml>
<o:shapelayout v:ext="edit">
<o:idmap v:ext="edit" data="1" />
</o:shapelayout></xml><![endif]-->
      <div class="WordSection1"><o:p>&nbsp;</o:p>An Advertisement can include
        an optional global Capture Scene Entry list, for each media
        type.&nbsp; Each item in this list is a set of one or more Capture
        Scene Entries.&nbsp; <o:p></o:p></div>
    </blockquote>
    <br>
    Do you mean that we have a global capture scene entry list for video
    and a global capture scene entry list for audio?<br>
    Otherwise, what about "...global Capture Scene Entry list. Each item
    in the list is a set of one or more Capture Scene Entries of the
    same media type."? That last approach is more similar to the one
    adopted for the content of capture scenes. <br>
    <br>
    Regards,<br>
    <br>
    Roberta<br>
  </body>
</html>

--------------000202030806010905000908--


From nobody Tue Apr  1 07:45:55 2014
Return-Path: <mary.ietf.barnes@gmail.com>
X-Original-To: clue@ietfa.amsl.com
Delivered-To: clue@ietfa.amsl.com
Received: from localhost (ietfa.amsl.com [127.0.0.1]) by ietfa.amsl.com (Postfix) with ESMTP id 87DC51A0874 for <clue@ietfa.amsl.com>; Tue,  1 Apr 2014 07:45:53 -0700 (PDT)
X-Virus-Scanned: amavisd-new at amsl.com
X-Spam-Flag: NO
X-Spam-Score: -1.999
X-Spam-Level: 
X-Spam-Status: No, score=-1.999 tagged_above=-999 required=5 tests=[BAYES_00=-1.9, DKIM_SIGNED=0.1, DKIM_VALID=-0.1, DKIM_VALID_AU=-0.1, FREEMAIL_FROM=0.001, HTML_MESSAGE=0.001, SPF_PASS=-0.001] autolearn=ham
Received: from mail.ietf.org ([4.31.198.44]) by localhost (ietfa.amsl.com [127.0.0.1]) (amavisd-new, port 10024) with ESMTP id hx8oXC8zRgXU for <clue@ietfa.amsl.com>; Tue,  1 Apr 2014 07:45:51 -0700 (PDT)
Received: from mail-wi0-x22b.google.com (mail-wi0-x22b.google.com [IPv6:2a00:1450:400c:c05::22b]) by ietfa.amsl.com (Postfix) with ESMTP id A5F651A0871 for <clue@ietf.org>; Tue,  1 Apr 2014 07:45:47 -0700 (PDT)
Received: by mail-wi0-f171.google.com with SMTP id q5so5312586wiv.16 for <clue@ietf.org>; Tue, 01 Apr 2014 07:45:43 -0700 (PDT)
DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20120113;  h=mime-version:in-reply-to:references:date:message-id:subject:from:to :cc:content-type; bh=E0NyOwva3i5jb5kzzF9cHUJ0bBoK15C7p2ImW07cMhE=; b=ZLg7w9U3KZdwl7//WLHkOxmWTiFZaVclGCY8zoZwP2mIhPUZnSejFPCfL8N6pyXJ7u qgBk08tYNwKdcRCxsMChwSd335a45+XlyrSUMUYJ9R8KKFfUbw7Gpex+V3R05WIAMBvB WnS93Q9SEXzH5Pkciv5kI3Ojhl4zrfyci84ZWoMS7Q7hOxYbr3y6AzXqjE5my2UKIZAd +gbk9KNsik2B4Qye0dsZNnQwQdYH4g/4GS4hdXaBw2etOv0pHhbLsIU73UKQqvE/YOuK oqcYfA/tc/RY1mKgDDBvF91ARizkFR/Q2U/eb/IC+NajY/24Jd09oCHIhKYk31NKX9T3 PN3A==
MIME-Version: 1.0
X-Received: by 10.181.13.15 with SMTP id eu15mr20727047wid.38.1396363543430; Tue, 01 Apr 2014 07:45:43 -0700 (PDT)
Received: by 10.216.10.6 with HTTP; Tue, 1 Apr 2014 07:45:43 -0700 (PDT)
In-Reply-To: <CAHBDyN7cL6BQ67MNkOzuMK8N37Htwck+-uPW8jz1Ut79mU2TFw@mail.gmail.com>
References: <1073194987.31242.1395857956833.POLL_ADMIN_PARTICIPATELINK.doodle@worker1> <CAHBDyN7cL6BQ67MNkOzuMK8N37Htwck+-uPW8jz1Ut79mU2TFw@mail.gmail.com>
Date: Tue, 1 Apr 2014 09:45:43 -0500
Message-ID: <CAHBDyN4aryQCf=9zBBAjiw2B=L9RsweuNRtp1C5MMh7cK_NjNA@mail.gmail.com>
From: Mary Barnes <mary.ietf.barnes@gmail.com>
To: CLUE <clue@ietf.org>
Content-Type: multipart/alternative; boundary=f46d04388ec3fd17ea04f5fc38bb
Archived-At: http://mailarchive.ietf.org/arch/msg/clue/N05oBFK8nHPMNvgFYCj2Zd3FGXU
Subject: Re: [clue] Doodle: Link for poll "CLUE Virtual Interim"
X-BeenThere: clue@ietf.org
X-Mailman-Version: 2.1.15
Precedence: list
List-Id: CLUE - ControLling mUltiple streams for TElepresence <clue.ietf.org>
List-Unsubscribe: <https://www.ietf.org/mailman/options/clue>, <mailto:clue-request@ietf.org?subject=unsubscribe>
List-Archive: <http://www.ietf.org/mail-archive/web/clue/>
List-Post: <mailto:clue@ietf.org>
List-Help: <mailto:clue-request@ietf.org?subject=help>
List-Subscribe: <https://www.ietf.org/mailman/listinfo/clue>, <mailto:clue-request@ietf.org?subject=subscribe>
X-List-Received-Date: Tue, 01 Apr 2014 14:45:53 -0000

--f46d04388ec3fd17ea04f5fc38bb
Content-Type: text/plain; charset=ISO-8859-1

HI all,

So, the final time that was the optimal for the majority is 3pm Central on
Thursday, May 29th.  I know that's late for you all in Europe & Israel, but
hopefully manageable this one time.  I'll shortly send the Webex details
and post a preliminary agenda on the wiki.

As noted on the design team call, the expectation is that updates to
documents be completed at least a week before to allow people time to
carefully review so we can have an effective meeting.  We will continue
with design team meetings up until the interim (and likely after) and we'll
try to get the topics identified shortly on the wiki.

Thanks,
Mary.


On Wed, Mar 26, 2014 at 1:43 PM, Mary Barnes <mary.ietf.barnes@gmail.com>wrote:

> HI all,
>
> Unfortunately, we've had a number of people that are not able to attend a
> f2f meeting due to lack of travel funding, time commitments and personal
> conflicts.  I do want to thank Simon et al for their willingness to host a
> f2f meeting and I still intend to get to Naples somehow in the near future
> ;)
>
> So, Paul and I have chatted and we are proposing a virtual interim meeting
> instead.  As we discussed one of the benefits of a f2f is getting everyone
> in one place since we do have a timezone split across the group that makes
> it difficult to get everyone on a call at the same time.  So, we'd really
> like to get 100% participation from our key contributors (i.e., document
> authors and reviewers).  As a compromise, we are proposing a 3pm Eastern
> meeting start for a 2 hour meeting.  I know this is quite late for Europe &
> Israel and quite early for folks in Asia and Australia.  But, hopefully,
> folks can make the sacrifice this one time - I realize that's very easy for
> me say given my timezone ;)
>
> So, here's the doodle:
> http://doodle.com/b5qtvb3ar5t25xad
>
> Please consider using the yellow if it's not an optimal time but if you're
> willing to make the sacrifice for that time.
>
> Please respond to the doodle no later than Wednesday, April 2nd at noon
> Eastern, so that we can make a decision.
>
> Thanks,
> Mary
>
>

--f46d04388ec3fd17ea04f5fc38bb
Content-Type: text/html; charset=ISO-8859-1
Content-Transfer-Encoding: quoted-printable

<div dir=3D"ltr">HI all,<div><br></div><div>So, the final time that was the=
 optimal for the majority is 3pm Central on Thursday, May 29th. =A0I know t=
hat&#39;s late for you all in Europe &amp; Israel, but hopefully manageable=
 this one time. =A0I&#39;ll shortly send the Webex details and post a preli=
minary agenda on the wiki.</div>
<div><br></div><div>As noted on the design team call, the expectation is th=
at updates to documents be completed at least a week before to allow people=
 time to carefully review so we can have an effective meeting. =A0We will c=
ontinue with design team meetings up until the interim (and likely after) a=
nd we&#39;ll try to get the topics identified shortly on the wiki.=A0</div>
<div><br></div><div>Thanks,</div><div>Mary.=A0</div></div><div class=3D"gma=
il_extra"><br><br><div class=3D"gmail_quote">On Wed, Mar 26, 2014 at 1:43 P=
M, Mary Barnes <span dir=3D"ltr">&lt;<a href=3D"mailto:mary.ietf.barnes@gma=
il.com" target=3D"_blank">mary.ietf.barnes@gmail.com</a>&gt;</span> wrote:<=
br>
<blockquote class=3D"gmail_quote" style=3D"margin:0 0 0 .8ex;border-left:1p=
x #ccc solid;padding-left:1ex"><div dir=3D"ltr">HI all,<div><br></div><div>=
Unfortunately, we&#39;ve had a number of people that are not able to attend=
 a f2f meeting due to lack of travel funding, time commitments and personal=
 conflicts. =A0I do want to thank Simon et al for their willingness to host=
 a f2f meeting and I still intend to get to Naples somehow in the near futu=
re ;)</div>

<div><br></div><div>So, Paul and I have chatted and we are proposing a virt=
ual interim meeting instead. =A0As we discussed one of the benefits of a f2=
f is getting everyone in one place since we do have a timezone split across=
 the group that makes it difficult to get everyone on a call at the same ti=
me. =A0So, we&#39;d really like to get 100% participation from our key cont=
ributors (i.e., document authors and reviewers). =A0As a compromise, we are=
 proposing a 3pm Eastern meeting start for a 2 hour meeting. =A0I know this=
 is quite late for Europe &amp; Israel and quite early for folks in Asia an=
d Australia. =A0But, hopefully, folks can make the sacrifice this one time =
- I realize that&#39;s very easy for me say given my timezone ;) =A0</div>

<div><br></div><div>So, here&#39;s the doodle: =A0</div><div><a href=3D"htt=
p://doodle.com/b5qtvb3ar5t25xad" style=3D"color:rgb(17,85,204)" target=3D"_=
blank">http://doodle.com/b5qtvb3ar5t25xad</a><br></div><div><br></div><div>=
Please consider using the yellow if it&#39;s not an optimal time but if you=
&#39;re willing to make the sacrifice for that time. =A0=A0</div>

<div><br></div><div>Please respond to the doodle no later than Wednesday, A=
pril 2nd at noon Eastern, so that we can make a decision.</div><div><br></d=
iv><div>Thanks,</div><div>Mary</div><div>=A0<br></div></div>
</blockquote></div><br></div>

--f46d04388ec3fd17ea04f5fc38bb--


From nobody Tue Apr  1 10:07:15 2014
Return-Path: <pkyzivat@alum.mit.edu>
X-Original-To: clue@ietfa.amsl.com
Delivered-To: clue@ietfa.amsl.com
Received: from localhost (ietfa.amsl.com [127.0.0.1]) by ietfa.amsl.com (Postfix) with ESMTP id 934071A0996 for <clue@ietfa.amsl.com>; Tue,  1 Apr 2014 10:07:14 -0700 (PDT)
X-Virus-Scanned: amavisd-new at amsl.com
X-Spam-Flag: NO
X-Spam-Score: -0.735
X-Spam-Level: 
X-Spam-Status: No, score=-0.735 tagged_above=-999 required=5 tests=[BAYES_00=-1.9, DKIM_SIGNED=0.1, DKIM_VALID=-0.1, MIME_BAD_LINEBREAK=0.5, SPF_SOFTFAIL=0.665] autolearn=no
Received: from mail.ietf.org ([4.31.198.44]) by localhost (ietfa.amsl.com [127.0.0.1]) (amavisd-new, port 10024) with ESMTP id 6R2vcpgCsqKa for <clue@ietfa.amsl.com>; Tue,  1 Apr 2014 10:07:13 -0700 (PDT)
Received: from qmta01.westchester.pa.mail.comcast.net (qmta01.westchester.pa.mail.comcast.net [IPv6:2001:558:fe14:43:76:96:62:16]) by ietfa.amsl.com (Postfix) with ESMTP id B42681A0564 for <clue@ietf.org>; Tue,  1 Apr 2014 10:07:12 -0700 (PDT)
Received: from omta21.westchester.pa.mail.comcast.net ([76.96.62.72]) by qmta01.westchester.pa.mail.comcast.net with comcast id kggv1n0061ZXKqc51h78PD; Tue, 01 Apr 2014 17:07:08 +0000
Received: from Paul-Kyzivats-MacBook-Pro.local ([50.138.229.164]) by omta21.westchester.pa.mail.comcast.net with comcast id kh761n0083ZTu2S3hh78Vr; Tue, 01 Apr 2014 17:07:08 +0000
Message-ID: <533AF239.4040600@alum.mit.edu>
Date: Tue, 01 Apr 2014 13:07:05 -0400
From: Paul Kyzivat <pkyzivat@alum.mit.edu>
User-Agent: Mozilla/5.0 (Macintosh; Intel Mac OS X 10.7; rv:24.0) Gecko/20100101 Thunderbird/24.4.0
MIME-Version: 1.0
To: CLUE <clue@ietf.org>
Content-Type: multipart/mixed; boundary="------------000902080100050200080309"
DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=comcast.net; s=q20140121; t=1396372028; bh=M731OIoscG1U3aRmMj2Cz7jlj1OuUkjeDz83+4S8r7M=; h=Received:Received:Message-ID:Date:From:MIME-Version:To:Subject: Content-Type; b=AAfYwXdYsF0c/+hFhd8agZHuMjGjF+jADmecGqzf7h21Npe2cY4oXFeElkPQAcmHr gt4gBjuoV5/ehiSV3Jq9HANMTTFwht9pBf4LXz8laeqdaQnXATG+pQ2n1Gn1xiCc26 kgFakDIlpK9ELlIxvpw04jaUdUz3SVUz9BWDSQbKGMXpw737KOKGXJgTg9QBJPNUXf Rd6wTqn+xjxKSaJ2ub+DyImYOvjemQC40Rr97X3AWxU/pwopYdyfVs5OnNTR/t5Ezf 0FUMUf+79pUDl3OYpIpkrwB+6sKiEkf639M3KEmL5yAl7d2yQEwEvjxA2VLTSh+wQk lsAPUfqrRX0Kg==
Archived-At: http://mailarchive.ietf.org/arch/msg/clue/TFFx03GuyO0mS84XwM-L6vP4eIw
Subject: [clue] Design team meeting today
X-BeenThere: clue@ietf.org
X-Mailman-Version: 2.1.15
Precedence: list
List-Id: CLUE - ControLling mUltiple streams for TElepresence <clue.ietf.org>
List-Unsubscribe: <https://www.ietf.org/mailman/options/clue>, <mailto:clue-request@ietf.org?subject=unsubscribe>
List-Archive: <http://www.ietf.org/mail-archive/web/clue/>
List-Post: <mailto:clue@ietf.org>
List-Help: <mailto:clue-request@ietf.org?subject=help>
List-Subscribe: <https://www.ietf.org/mailman/listinfo/clue>, <mailto:clue-request@ietf.org?subject=subscribe>
X-List-Received-Date: Tue, 01 Apr 2014 17:07:14 -0000

This is a multi-part message in MIME format.
--------------000902080100050200080309
Content-Type: text/plain; charset=ISO-8859-1; format=flowed
Content-Transfer-Encoding: 7bit

The notes from today's design team meeting are now posted on the clue 
wiki. (I've also attached them to this.)

	Thanks,
	Paul

--------------000902080100050200080309
Content-Type: text/plain; charset=UTF-8;
 name="minutes-140401.txt"
Content-Transfer-Encoding: base64
Content-Disposition: attachment;
 filename="minutes-140401.txt"

Q0xVRSBXb3JraW5nIEdyb3VwIERlc2lnbiB0ZWFtIG1lZXRpbmcgKEFwcmlsIDEsIDIwMTQp
Cj09PT09PT09PT09PT09PT09PT09PT09PT09PT09PT09PT09PT09PT09PT09PT09PT09PT09
PT09PT0KClN1YmplY3Q6IEFzc29ydGVkCgpBdHRlbmRlZXM6IFBhdWwgS3l6aXZhdCwgSm9o
biBMZXNsaWUsIE1hcnkgQmFybmVzLCBSb2JlcnRhIFByZXN0YSwgUm9uaSBFdmVuLCBSb2Ig
SGFuc2VuLCBDaHJpc3RlciBIb2xtYmVyZywgSm9uYXRoYW4gTGVubm94CgpOb3RlIHRha2Vy
czogIFBhdWwsIEpvaG4KClJlY29yZGluZzogaHR0cHM6Ly9pZXRmLndlYmV4LmNvbS9pZXRm
L2xkci5waHA/UkNJRD0xOGFmNzAzNmQxZWRjMDg5NDM1MzZkYTIzYzFmMmZlZgoKU3VtbWFy
eToKPT09PT09PT0KCkFjdGlvbnM6CgpDaGFpcnM6IE9wZW4gbmV3IHRpY2tldCBmb3IgY2xh
cmlmeWluZyBhdWRpbyBzcGF0aWFsIGluZm9ybWF0aW9uLgoKSm9obiBMZXNsaWU6IFB1c2gg
dGhlIGltcHJvdmVtZW50IG9mIGF1ZGlvIHNwZWNpZmljYXRpb24gb24gdGhlIGxpc3QuCgpS
b2IgSGFuc2VuOiBSZXN1Ym1pdCBzaWduYWxpbmcgZHJhZnQgYXMgV0cgZHJhZnQuCgpSb2Ig
SGFuc2VuOiBQdWJsaXNoIHN0cmF3bWFuIG9mIHVwZGF0ZXMgdG8gc2lnbmFsaW5nIChhcyBk
ZXNjcmliZWQgYmVsb3MpIGJ5IDEwLUFwci4KCj09PT09PT09PT09PT09PT09PT09PT09PT09
PT09PT09PT09PT09PT09PT09PT09PT09PT09PT09PT0KTm90ZXMgYnkgUGF1bCBLeXppdmF0
OgoKQXVkaW86Ci0tLS0tLQoKSm9obiBMZXNsaWU6IHdlIG5lZWQgdG8gZG8gbW9yZSBhYm91
dCBhdWRpby4gCgpUaGVyZSB3YXMgZGlzY3Vzc2lvbiBhYm91dCB0aGlzLCBhbmQgd2hhdCBp
dCBtZWFudC4KV2UgY3VycmVudGx5IGRvbuKAmXQgaGF2ZSBhbnl0aGluZyB0byBoZWxwIGRl
Y2lkZSB3aGljaCBhdWRpbyBnb2VzIHdpdGggd2hpY2ggdmlkZW8uIAoKSm9objogZG9uJ3Qg
dGhpbmsgd2Ugc2hvdWxkIHRyeSB0byBkbyB0aGF0LCBleGNlcHQgdG8gYXNzb2NpYXRlIG1l
ZGlhIHdpdGhpbiBhIHNjZW5lLiAKUm9iOiBwZW9wbGUgd2FudCB0byBhc3NvY2lhdGUg4oCT
IGxlZnQgdmlkZW8gd2l0aCBsZWZ0IGF1ZGlvLCBldGMuIAoKUm9uaTogdGhhdCBpcyBvayAt
IGp1c3QgbmVlZCBleGFtcGxlcy4KClBhdWw6IHdpbGwgb3BlbiBhIHRpY2tldCB0byB0cmFj
ayB0aGlzIGlzc3VlLiBCdXQgSm9obiBpcyBvdXIgZXhwZXJ0IG9uIHRoaXMsIGFuZCBuZWVk
cyB0byBwdXNoIHRoZSBpc3N1ZS4gT3RoZXJzIHdpbGwgaGVscCBpbiBmaWd1cmluZyBvdXQg
d2hhdCBkb2N1bWVudHMgbmVlZCB0byBhZGRyZXNzIHRoZSBpc3N1ZS4gCgpUaGUgaW1wYWN0
ZWQgZG9jdW1lbnRzIG1lbnRpb25lZCB3ZXJlOiBmcmFtZXdvcmsgYW5kIGRhdGEgbW9kZWwu
IEluIHBhcnRpY3VsYXIsIHNwYXRpYWwgaW5mb3JtYXRpb24uIEFncmVlZCB0aGF0IHBvaW50
IG9mIGNhcHR1cmUgYW5kIGF4aXMgb2YgY2FwdHVyZSBhcmUgcmVsZXZhbnQuCgpSb25pOiBu
ZWVkIG1vcmUgc3BlY2lmaWNhdGlvbiBvZiBob3cgdGhpcyBpbmZvcm1hdGlvbiB3aWxsIGJl
IHVzZWQgYnkgdGhlIHJlY2VpdmVyLgoKSm9objogQ2FuJ3Qgc2F5IGFueXRoaW5nIG5vcm1h
dGl2ZSBhYm91dCB0aGlzLgoKU2lnbmFsaW5nOgotLS0tLS0tLS0tCgpDaHJpc3RlcjogbmVl
ZCB0byBkaXNjdXNzIHRyYW5zaXRpb24gZnJvbSBjbHVlIG1vZGUgdG8gbGVnYWN5IG1vZGUu
IEhlIHRoaW5rcyB0aGlzIGlzIG1vcmUgaW1wb3J0YW50IHRoYW4gc3BlbGxpbmcgb3V0IGJ1
bmRsZS4gCg1Sb2I6IG9wZW4gaXNzdWVzOiAKMSkgaG93IHRvIHNpZ25hbCBjbHVlIOKAkyBm
ZWF0dXJlIHRhZyBvciBncm91cD8gCjIpIHdoYXQgaWYgY2x1ZSBjaGFubmVsIGVuZHMuIAoz
KSBob3cgdG8gZG8gYnVuZGxlLiAKNCkgZG8gd2Ugd2FudCB0byBjb25zaWRlciBiaS1kaXJl
Y3Rpb25hbCBmbG93cz8gCgpSb2Igd2lsbCBwcm9wb3NlIGEgc3RyYXdtYW4gZm9yIDEtMyBh
Ym92ZS4gKERvZXNuJ3QgaGF2ZSBhIHByb3Bvc2FsIGluIG1pbmQgZm9yIDQuKSBDb21taXRz
IHRvIGRvIGJ5IHRoZSBtb3JuaW5nIG9mIEFwcmlsIDEwIQoKUGF1bCAmIE1hcnk6IGl0IGlz
IG9rIChkZXNpcmFibGUpIGZvciB0aGVzZSBwcm9wb3NhbHMgdG8gYmUgbWFkZSBpbiBhIG5l
dyB2ZXJzaW9uIG9mIHRoZSBkb2N1bWVudC4gU2hvdWxkIGluZGljYXRlIHdoZW4gdGhpbmdz
IGFyZSBwcm9wb3NhbHMgcmF0aGVyIHRoYW4gYWdyZWVkLgoKTWFyeTogd2UgaGF2ZSBjb25z
ZW5zdXMgdG8gYWRvcHQgYXMgV0cgZHJhZnQuCgpQYXVsOiBSb2Igc2hvdWxkIHN1Ym1pdCBh
IHJlbmFtZWQgdmVyc2lvbiBhbmQgY2hhaXJzIHdpbGwgYWxsb3cgaXQuCgpDYWxsIEZsb3cg
RG9jdW1lbnQ6Ci0tLS0tLS0tLS0tLS0tLS0tLS0KCk1hcnk6IFRoaXMgc3RpbGwgcmVtYWlu
cyB0byBiZSBkb25lLiBMb3JlbnpvIGhhZCB2b2x1bnRlZXJlZC4KClRoZXJlIHdhcyBkaXNj
dXNzaW9uIG9mIGVhc3kvZGlmZmljdWx0eSBvZiBnZXR0aW5nIHRoaXMgZG9uZS4KClZpcnR1
YWwgSW50ZXJpbToKLS0tLS0tLS0tLS0tLS0tLQoKQ2hyaXN0ZXI6IGRvIHdlIHJlYWxseSBu
ZWVkIG9uZT8gQXJlbid0IHRoZSBkZXNpZ24gdGVhbSBjYWxscyBnb29kIGVub3VnaD8KClRo
ZXJlIHdhcyBkaXNjdXNzaW9uIG9mIHRoZSBiZW5lZml0cyBvZiBhIHZpcnR1YWwgaW50ZXJp
bS4gTWFyeSBtZW50aW9uZWQgdGhlIHByZXNlbmNlIG9mIGFuIEFELCBhbmQgb3RoZXJzIHdo
byBkb24ndCBjb21lIHRvIG91ciBpbnRlcmltcy4KCkNocmlzdGVyIHdpdGhkcmV3IGhpcyBv
YmplY3Rpb24uCgpGdXR1cmUgRGVzaWduIFRlYW0gbWVldGluZ3M6Ci0tLS0tLS0tLS0tLS0t
LS0tLS0tLS0tLS0tLS0KClBhdWw6IGN1cnJlbnRseSB3ZSBoYXZlIG5vIHRvcGljcyBhc3Np
Z25lZCBmb3IgZnV0dXJlIG1lZXRpbmdzLiBBc2tlZCBpZiBhbnlib2R5IGhhZCBzb21ldGhp
bmcgZm9yIG5leHQgd2Vlay4gKE5vYm9keSBkaWQuKSBPSywgd2lsbCBrZWVwIHRoZSBtZWV0
aW5nIG9wZW4gYW5kIGRlY2lkZSBiZWZvcmUgdGhlbi4KCk1hcnk6IHdlIHNob3VsZCBzdGFy
dCB1c2luZyB0aGUgbWVldGluZ3MgZm9yIGRvY3VtZW50IHJldmlld3MgYXMgcGVvcGxlIGdl
dCB0aGVpciBkb2N1bWVudHMgdXBkYXRlZC4KCj09PT09PT09PT09PT09PT09PT09PT09PT09
PT09PT09PT09PT09PT09PT09PT09PT09PT09PT09PT0KTm90ZXMgYnkgSm9obiBMZXNsaWU6
CgowOTM0IFBhdWwsIE1hcnksIEpvaG4sIFJvYmVydGEsIFJvbmksIFJvYiwgSm9uYXRoYW4g
YWJvdXQgMDk0MCwgQ2hyaXN0ZXIgYWJvdXQgMDk1NQpQYXVsIHdpbGwgdGFrZSBub3RlcwoK
UGF1bDogd2hhdCBhcmUgd2UgZ29pbmcgdG8gdGFsayBhYm91dD8gQXVkaW8/CkpvaG46IApN
YXJ5OiBub3QgdGhlIHJpZ2h0IHBlb3BsZSBvbiB0aGUgY2FsbCwgU3RlcGhlbiBCb3NjbyAt
LSBoZSB3b24ndCBiZSBhdmFpbGFibGUgYXMgYSByZXNvdXJjZQpSb25pOiB3YXMgZGlzY3Vz
c2VkIG9uLWxpc3QsIG5vIHJlYWwgZGlzY3Vzc2lvbiB0aGVuClBhdWw6IHRhbGtlZCBhYm91
dCBpdCwgbm8gcmVzb2x1dGlvbiwgaXQgZHJvcHBlZApSb25pOiBxdWVzdGlvbiBob3cgaXQg
d2lsbCBiZSB1c2VkIGJ5IHJlY2VpdmVyIC0tIHdoYXQgaXQgbWVhbnMsIGhvdyB0byBtYWtl
IHVzZSBvZiBpdC4uLgpQYXVsOiBKb2huIHBsZWFzZSBwdXNoIG9uLWxpc3QsIHNob3VsZCBz
aG93IHVwIGluIGZyYW1ld29yayBhbmQgZGF0YSBtb2RlbApSb2I6IHJlcXVpcmVzIG5ldyB0
ZXh0IGluIHByb3RvY29sClBhdWw6IHRhbGtlZCBvZiBtaWNyb3Bob25lIHBhdHRlcm5zLi4u
IG5vdyB3ZSBoYXZlIHNwYXRpYWwgbW9kZWwgaW50ZW5kZWQgZm9yIHZpZGVvLCBkb24ndCBr
bm93IGlmIGl0IG1ha2VzIHNlbnNlIGZvciBhdWRpbwpSb2I6IHBlb3BsZSBkbyB3YW50IHRv
IHBsYXkgb3V0IGF1ZGlvIHRvIG1hdGNoIHZpZGVvIGxvY2F0aW9uClJvbmk6IGlmIEkgaGF2
ZSB0aHJlZSB2aWRlbyBpbWFnZXMsIEkgd2FudCB0byBrbm93IHdoaWNoIGF1ZGlvIG1hdGNo
ZXMgc3BlYWtlciB3aXRoaW4gb25lIHJvb20KUm9iOiB3aGVyZSB0aGUgbWljcm9waG9uZSBp
czogbW9zdCBpbXBvcnRhbnQgdG8gbWUKUGF1bDogcG9pbnQgb2YgY2FwdHVyZSwgYXhpcyBv
ZiBjYXB0dXJlIC0tIGhvdyB0byBjaG9vc2UgdGhlIGF1ZGlvLi4uIGxldCdzIHRha2UgaXQg
dG8gdGhlIGxpc3QKClBhdWw6IGluIGdlbmVyYWwsIGhvdyB3ZSdyZSBkb2luZyBpbiB0ZXJt
cyBvZiBnZXR0aW5nIGRvbmUKTWFyeTogSSdsbCB1cGRhdGUgdGhlIHNjaGVkdWxlLCBwZW9w
bGUgY2FuIHJlYWN0IG9uLWxpc3QsIGxldCdzIG1vdmUgdG8gc2lnbmFsaW5nIG5vdyB0aGF0
IENocmlzdGVyIGhhcyBqb2luZWQKQ2hyaXN0ZXI6IHN3aXRjaGluZyBmcm9tIGxlZ2FjeSB0
byBDTFVFIG1vZGUsIHdoYXQgaGFwcGVucyBpZiBkYXRhIGNoYW5uZWwgZ29lcyBhd2F5PyBt
YWluIGlzc3VlIHNob3VsZCBiZSBDTFVFIGVzdGFibGlzbWVudCBhbmQgdGVybWluYXRpb24K
UGF1bDogZm9yIHNpZ25hbGluZyBkb2MsIHdoYXQncyBibG9ja2luZz8gZGVjaXNpb25zIHRo
YXQgbmVlZCB0byBiZSBtYWRlLCBvciBqdXN0IHRleHQKUm9iOiBob3cgZG8geW91IGlkZW50
aWZ5IHRoaXMgYXMgYSBDTFVFIGNhbGwuLi4gbWVkaWEtZmVhdHVyZSB0YWc/Li4uIHdoYXQg
aGFwcGVucyB3aGVuIHlvdSBsb3NlIGRhdGEgY2hhbm5lbDsgaG93IHlvdSBkbyBCVU5ETEU/
OyB3aWxsIHdlIGhhbmRsZSBiaS1kaXJlY3Rpb25hbCBmbG93cwpQYXVsOiBkbyB5b3Ugd2Fu
dCB1cyB0byBvcGVuIHRpY2tldHMgb24gdGhlc2U/ClJvYjogSSdtIGhhcHB5IHRvIHByb3Bv
c2UgYSBzdHJhdy1tYW4gZm9yIGVhY2gKUGF1bDogYW55dGhpbmcgbW9yZSB3ZSBuZWVkIHRv
IHRhbGsgYWJvdXQgb24gc2lnbmFsaW5nPwpSb2I6IG5vIHBvaW50IGluIGEgbmV3IGRyYWZ0
CkNocmlzdGVyOiBJIHByZWZlciB0byBwb3N0IGlzc3VlcyB0byBsaXN0Ck1hcnk6IHRvIGtl
ZXAgdHJhY2sgb2YgaXQsIGl0J3MgYmVzdCB0byBoYXZlIGl0IGluIGEgZG9jdW1lbnQKUm9i
OiBob3cgeW91IHN0YXJ0IGEgQ0xVRSBjaGFubmVsLCBob3cgeW91IG1ha2UgaXQgZ28gYXdh
eSwgaG93IHRvIHVzZSBCVU5ETEUsIG1heWJlIGJpLWRpcmVjdGlvbmFsClBhdWw6IEknbSBj
b25jZXJuZWQgYWJvdXQgYmktZGlyZWN0aW9uYWwsIGJ1dCBJIGhhdmUgbm8gc3RhcnRpbmcg
cG9pbnQ7IHdoZW4gaXQncyBhIHN0cmF3LWhvcnNlIGluIHRoZSBkb2N1bWVudCwgaXQncyBn
b29kIHRvIG1hcmsgaXQgY2xlYXJseSBhcyBzdWNoCk1hcnk6IGluIHRoZSBkb2N1bWVudCwg
aXQncyBlYXNpZXIgdG8gdHJhY2sgY2hhbmdlcwpDaHJpc3RlcjogdGhlIGltcG9ydGFudCB0
aGluZyBpcyB0byBnZXQgdGV4dApSb2I6IEknbGwgd3JpdGUgc29tZXRoaW5nCk1hcnk6IGVk
aXRvcidzIG5vdGVzIHdoZXJlIHlvdSBoYXZlIHF1ZXN0aW9ucwpQYXVsOiB3aGVuIGNhbiB5
b3UgZ2V0IHRoYXQgZG9uZT8KUm9iOiBtb3JuaW5nIG9mIDEwIEFwcmlsCkNocmlzdGVyOiBh
bnkgY29uY2VybnMgb24gQlVORExFLCBjb250YWN0IG1lClBhdWw6IEknbGwgYWx3YXlzIGhh
dmUgb3BpbmlvbnMgb24gQlVORExFLi4uClBhdWw6IFJvYiwgeW91IGNhbiBzdWJtaXQgeW91
ciBkcmFmdCBhcyBhIFdHIGRyYWZ0IC0tIGp1c3QgY2hhbmdlIHRoZSBuYW1lIC0tIHN0dWZm
IHdlIHRhbGtlZCBhYm91dCBjYW4gZ28gaW4gbGF0ZXIKCjEwMTIgUGF1bDogd2hhdCBlbHNl
IGlzIGtlZXBpbmcgdXMgZnJvbSBiZWluZyBkb25lCk1hcnk6IGNyaXRpY2FsIHRoaW5nIGlz
IHNpZ25hbGluZywgd2l0aCBhbGwgZGV0YWlsczsgdGhlbiBhIGNhbGwtZmxvdyBkb2N1bWVu
dDsgSSdsbCBmb2xsb3ctdXAgd2l0aCBMb3JlbnpvCk1hcnk6IGRhdGEtY2hhbm5lbCBzdHVm
ZiBpcyBwcmV0dHkgc3RyYWlnaHRmb3J3YXJkLCBjcml0aWNhbCB0aGluZyBpcyBzaWduYWxp
bmcKCjEwMTYgTWFyeTogQ2hyaXN0aWFuIHJlcXVlc3RlZCBhbiBlYXJsaWVyIHRpbWUgb24t
bGlzdCwgSSByZWNhbGwgdGhhdCB3b3VsZG4ndCB3b3JrIGZvciBDaHJpc3RlcgpDaHJpc3Rl
cjogOTAgbWludXRlcyBlYXJsaWVyIHdvdWxkIGJlIGZpbmUgZm9yIG1lCk1hcnk6IEknbGwg
Zm9sbG93LXVwIG9uLWxpc3QKSm9uYXRoYW46IG11bWJsZS4uLiA5MCBtaW51dGVzIGVhcmxp
ZXIgd291bGQgYmUgaGFyZCBmb3IgbWUsIGJ1dCBpdCBzZWVtcyBhIHJlYXNvbmFibGUgdHJh
ZGVvZmYgdG8gZ2V0IHNvbWVvbmUgd3JpdGluZyBhIGRvY3VtZW50LgpDaHJpc3Rlcjogd2h5
IGRvIHdlIG5lZWQgdmlydHVhbCBpbnRlcmltCk1hcnk6IHBvc3NpYmxlIHdlJ2QgZ2V0IFJU
Q1dFQiBwZW9wbGUKUGF1bDogbWFpbiB0aGluZyBpcyBob3cgd2UgdHJlYXQgaXQgLS0gZGVh
ZGxpbmVzLCBzbGlkZXMgaW4gYWR2YW5jZQpNYXJ5OiB3ZSBob3BlIHRvIGdldCBvdXIgbmV3
IEFyZWEgRGlyZWN0b3IgaW52b2x2ZWQgLS0gd2UgYW5ub3VuY2VkIHRvIGNvbW11bml0eSwg
b3RoZXIgV0cgbWF5IGF0dGVuZCwgd2UnbGwgaGF2ZSBBRApQYXVsOiBtYWtlIHN1cmUgd2Un
cmUgc2VyaW91cyBhYm91dCBpdCAtLSBpZiB3ZSB0cmVhdCBpdCBhcyBqdXN0IGFub3RoZXIg
ZGVzaWduLXRlYW0sIGl0IHdvbid0IGJlIGFzIGVmZmVjdGl2ZQpNYXJ5OiB3ZSB3YW50IGRv
Y3VtZW50cyBhIHdlZWsgYmVmb3JlIGludGVyaW0sIHNvIHBlb3BsZSBoYXZlIHRpbWUgdG8g
cmV2aWV3Li4uIG9uZSBkYXkgaW50ZXJpbSB3aWxsIGhhdmUgdG8gd29yayAoY2FuJ3Qgc2No
ZWR1bGUgdHdvIGRheXMpOyB3ZSdsbCBzY2hlZHVsZSB0aGluZ3Mgd2UgY2FuJ3Qgc2V0dGxl
IGZvciBkZXNpZ24gdGVhbSBtZWV0aW5ncwpNYXJ5OiBJJ2xsIGdldCBpc3N1ZXMgaW4gdGhl
IHRyYWNrZXIsIGRvY3VtZW50IGVkaXRvcnMgY2FuIGJlIGFsbG93ZWQgdG8gZW50ZXIgdGhl
aXIgb3duIGlzc3VlcwoKMTAyNyBQYXVsOiBhbnkgc3ViamVjdHMgdGhhdCB3aWxsIGJlIHJl
YWR5IHRvIHRhbGsgYWJvdXQgbmV4dCB3ZWVrClJvbmk6IHdlIGRvbid0IGhhdmUgTWFyYyBv
biB0aGUgY2FsbApNYXJ5OiBoZSBkaWRuJ3Qgc2VlIGFueSBpc3N1ZXMgbmVlZGluZyByZWFs
LXRpbWUgaW50ZXJhY3Rpb24KCjEwMzAgUGF1bDogZW5kIG9mIG1lZXRpbmcgdGltZTsgcGxl
YXNlIGhvbGQgdGhlIHRpbWUgbmV4dCB3ZWVrOyBhZGpvdXJuZWQKCi0tCkpvaG4gTGVzbGll
IDxqb2huQGpsYy5uZXQ+CgoKDQ==
--------------000902080100050200080309--


From nobody Tue Apr  1 10:11:53 2014
Return-Path: <pkyzivat@alum.mit.edu>
X-Original-To: clue@ietfa.amsl.com
Delivered-To: clue@ietfa.amsl.com
Received: from localhost (ietfa.amsl.com [127.0.0.1]) by ietfa.amsl.com (Postfix) with ESMTP id 181621A088C for <clue@ietfa.amsl.com>; Tue,  1 Apr 2014 10:11:52 -0700 (PDT)
X-Virus-Scanned: amavisd-new at amsl.com
X-Spam-Flag: NO
X-Spam-Score: -1.235
X-Spam-Level: 
X-Spam-Status: No, score=-1.235 tagged_above=-999 required=5 tests=[BAYES_00=-1.9, DKIM_SIGNED=0.1, DKIM_VALID=-0.1, SPF_SOFTFAIL=0.665] autolearn=no
Received: from mail.ietf.org ([4.31.198.44]) by localhost (ietfa.amsl.com [127.0.0.1]) (amavisd-new, port 10024) with ESMTP id SZ21YAKsa7k9 for <clue@ietfa.amsl.com>; Tue,  1 Apr 2014 10:11:50 -0700 (PDT)
Received: from qmta13.westchester.pa.mail.comcast.net (qmta13.westchester.pa.mail.comcast.net [IPv6:2001:558:fe14:44:76:96:59:243]) by ietfa.amsl.com (Postfix) with ESMTP id 08C3B1A0564 for <clue@ietf.org>; Tue,  1 Apr 2014 10:11:49 -0700 (PDT)
Received: from omta17.westchester.pa.mail.comcast.net ([76.96.62.89]) by qmta13.westchester.pa.mail.comcast.net with comcast id kda21n0061vXlb85DhBmVl; Tue, 01 Apr 2014 17:11:46 +0000
Received: from Paul-Kyzivats-MacBook-Pro.local ([50.138.229.164]) by omta17.westchester.pa.mail.comcast.net with comcast id khBl1n00Q3ZTu2S3dhBmUo; Tue, 01 Apr 2014 17:11:46 +0000
Message-ID: <533AF351.9050201@alum.mit.edu>
Date: Tue, 01 Apr 2014 13:11:45 -0400
From: Paul Kyzivat <pkyzivat@alum.mit.edu>
User-Agent: Mozilla/5.0 (Macintosh; Intel Mac OS X 10.7; rv:24.0) Gecko/20100101 Thunderbird/24.4.0
MIME-Version: 1.0
To: CLUE <clue@ietf.org>
Content-Type: text/plain; charset=ISO-8859-1; format=flowed
Content-Transfer-Encoding: 7bit
DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=comcast.net; s=q20140121; t=1396372306; bh=/ebBBj9U6FPe1QPGpQcz4FwT3hHsLoh0kfff/gs4Koo=; h=Received:Received:Message-ID:Date:From:MIME-Version:To:Subject: Content-Type; b=ef14NG4HWOQWatpH2veXOtVSelnv1rLaSdzs/vCu362/UTsu5Qa8SXt47NU3KNvEp A2XWhuPNo1liqn60XexNT8Gvs/iXm87rxZOaFVjR5IQK9wiyNbnggE+AECa1X8ysUn GPAy3mlG/eEZ3/g1Tm34APFSbq+1McVbvcfrj7H/8Q10dvYVr+8tf0G7L620RWDyfV YaZyQvSa/EUVPaUWct35/zpNyzviJamCqdswsRwGZqqohLZToYgEQrkmz1qT5/nHwe O06Mr5Y4tPb+/JMPcFK92agtj4YgFWOdfp+oJ8lHtr/RusHaKYXDPyb5fg/Wb6z/te 8YNGvaKcPfe1g==
Archived-At: http://mailarchive.ietf.org/arch/msg/clue/V2y5Ypqg5ltI6xZK_Slw2BTjCyk
Subject: [clue] Improving treatment of audio
X-BeenThere: clue@ietf.org
X-Mailman-Version: 2.1.15
Precedence: list
List-Id: CLUE - ControLling mUltiple streams for TElepresence <clue.ietf.org>
List-Unsubscribe: <https://www.ietf.org/mailman/options/clue>, <mailto:clue-request@ietf.org?subject=unsubscribe>
List-Archive: <http://www.ietf.org/mail-archive/web/clue/>
List-Post: <mailto:clue@ietf.org>
List-Help: <mailto:clue-request@ietf.org?subject=help>
List-Subscribe: <https://www.ietf.org/mailman/listinfo/clue>, <mailto:clue-request@ietf.org?subject=subscribe>
X-List-Received-Date: Tue, 01 Apr 2014 17:11:52 -0000

In today's design team meeting we had a discussion of what is lacking 
about our treatment of audio, and how to fix it.

I want to open a ticket on this topic, but I need some help to properly 
describe the task. IMO it has to do with what sort of spatial 
information should be provided for audio captures, how it can be used to 
correlate audio captures with video captures, and how it can be used to 
choose which audio captures to configure.

Can somebody (John?) make a *concise* statement of what is needed?

	Thanks,
	Paul


From nobody Tue Apr  1 16:13:17 2014
Return-Path: <Christian.Groves@nteczone.com>
X-Original-To: clue@ietfa.amsl.com
Delivered-To: clue@ietfa.amsl.com
Received: from localhost (ietfa.amsl.com [127.0.0.1]) by ietfa.amsl.com (Postfix) with ESMTP id B14391A0009 for <clue@ietfa.amsl.com>; Tue,  1 Apr 2014 16:13:15 -0700 (PDT)
X-Virus-Scanned: amavisd-new at amsl.com
X-Spam-Flag: NO
X-Spam-Score: -1.9
X-Spam-Level: 
X-Spam-Status: No, score=-1.9 tagged_above=-999 required=5 tests=[BAYES_00=-1.9] autolearn=ham
Received: from mail.ietf.org ([4.31.198.44]) by localhost (ietfa.amsl.com [127.0.0.1]) (amavisd-new, port 10024) with ESMTP id bdGlAHSe75tb for <clue@ietfa.amsl.com>; Tue,  1 Apr 2014 16:13:13 -0700 (PDT)
Received: from cserver5.myshophosting.com (cserver5.myshophosting.com [175.107.161.1]) by ietfa.amsl.com (Postfix) with ESMTP id 9393B1A0004 for <clue@ietf.org>; Tue,  1 Apr 2014 16:13:12 -0700 (PDT)
Received: from ppp118-209-153-126.lns20.mel6.internode.on.net ([118.209.153.126]:51479 helo=[127.0.0.1]) by cserver5.myshophosting.com with esmtpsa (TLSv1:DHE-RSA-AES128-SHA:128) (Exim 4.82) (envelope-from <Christian.Groves@nteczone.com>) id 1WV7rc-0001OU-Pm; Wed, 02 Apr 2014 10:13:04 +1100
Message-ID: <533B47FF.5020407@nteczone.com>
Date: Wed, 02 Apr 2014 10:13:03 +1100
From: Christian Groves <Christian.Groves@nteczone.com>
User-Agent: Mozilla/5.0 (Windows NT 6.1; WOW64; rv:24.0) Gecko/20100101 Thunderbird/24.4.0
MIME-Version: 1.0
To: Roberta Presta <roberta.presta@unina.it>, "clue@ietf.org" <clue@ietf.org>
References: <5318809E.2030204@nteczone.com> <5318AE5E.4050404@alum.mit.edu> <5318B0C9.1050603@nteczone.com> <5318B48E.3090300@alum.mit.edu> <53290C9B.6090106@nteczone.com> <5329F3FF.7090606@alum.mit.edu> <533A16BA.40707@nteczone.com> <533A91BF.7040404@unina.it>
In-Reply-To: <533A91BF.7040404@unina.it>
Content-Type: text/plain; charset=ISO-8859-1; format=flowed
Content-Transfer-Encoding: 7bit
X-AntiAbuse: This header was added to track abuse, please include it with any abuse report
X-AntiAbuse: Primary Hostname - cserver5.myshophosting.com
X-AntiAbuse: Original Domain - ietf.org
X-AntiAbuse: Originator/Caller UID/GID - [47 12] / [47 12]
X-AntiAbuse: Sender Address Domain - nteczone.com
X-Get-Message-Sender-Via: cserver5.myshophosting.com: authenticated_id: christian.groves@nteczone.com
X-Source: 
X-Source-Args: 
X-Source-Dir: 
Archived-At: http://mailarchive.ietf.org/arch/msg/clue/JPDVI21dnNgHKUiO0ei7UFv2YZQ
Subject: Re: [clue] Participant info/type followup
X-BeenThere: clue@ietf.org
X-Mailman-Version: 2.1.15
Precedence: list
List-Id: CLUE - ControLling mUltiple streams for TElepresence <clue.ietf.org>
List-Unsubscribe: <https://www.ietf.org/mailman/options/clue>, <mailto:clue-request@ietf.org?subject=unsubscribe>
List-Archive: <http://www.ietf.org/mail-archive/web/clue/>
List-Post: <mailto:clue@ietf.org>
List-Help: <mailto:clue-request@ietf.org?subject=help>
List-Subscribe: <https://www.ietf.org/mailman/listinfo/clue>, <mailto:clue-request@ietf.org?subject=subscribe>
X-List-Received-Date: Tue, 01 Apr 2014 23:13:16 -0000

Hello Roberta,

 From a data model perspective I agree that it makes sense to have the 
participant metadata "referenced" to minimize duplication in the 
messages. So I support your solution in the data model to the redundancy 
problem.

In general I think its good to maintain consistency between the 
framework and the data model. However in London I thought this was more 
of a syntax shortcut (like the captureID wildcarding) rather than 
something we'd need to formalize in the framework. I guess we could add 
it at a higher level in the framework (it would functionally be the same 
thing), it would just be more work for the editor...

I could go either way on this.

Regards, Christian

On 1/04/2014 9:15 PM, Roberta Presta wrote:
> Hi Christian,
>
> In section 7.1.1.X we are dealing with media capture attributes and in 
> section 7.1.1.11 we want to define an attribute conveying information 
> about participants.
> I would say that here we can find both information (*references*, in 
> the data model) about "who is represented in the capture" ("captured 
> participants"?) and "who is the owner of the generating device" 
> ("owner"?).
>
> Maybe participant metadata, such as Participant Information and 
> Participant Type as they are currently defined, should be treated in a 
> separate section of the same level of capture scenes or media captures 
> as.
> Indeed we showed in London that repeating the vcard and the role of 
> the participants in each capture, as if they were capture attributes,  
> causes redundancy.
>
> I know that I have a data model definition perspective, but I would 
> propose to make a change that is more coherent with what we will 
> describe formally.
>
> Cheers,
>
> Roberta
>
>
>
>
>
> Il 01/04/2014 03:30, Christian Groves ha scritto:
>> Hello Paul, all,
>>
>> If we follow the approach that there is a specific indicating of 
>> whether the participant information is based on an explicit 
>> indication then I would suggest the following text for the framework:
>>
>> Clause 7.1.11 Participant information
>> (Under the 1st paragraph)
>>
>> The participant information contains an explicit indication of 
>> whether it relates to a participant contained in the capture, from a 
>> participants capture device or both. For example a video camera may 
>> capture an image containing the participant, or a participant may 
>> send a video capture with a presentation that does not depict the 
>> participant.
>>
>> Something similar would be needed under participant type.
>>
>> Thoughts?
>>
>> Regards, Christian
>>
>> On 20/03/2014 6:46 AM, Paul Kyzivat wrote:
>>> On 3/18/14 11:18 PM, Christian Groves wrote:
>>>> Hello Paul,
>>>>
>>>> "How" they differ is given by the example bullets below the 
>>>> sentence. If
>>>> you want something more normative we could remove the "For example".
>>>
>>> Yeah, I don't believe in specification by example. :-)
>>>
>>> IMO it is a bit dicey to base this distinction on the type of capture.
>>>
>>> I'm more comfortable with an explicit syntactic indication of the 
>>> distinction, such as proposed by Roberta.
>>>
>>>     Thanks,
>>>     Paul
>>>
>>>> Regards, Christian
>>>>
>>>> On 7/03/2014 4:46 AM, Paul Kyzivat wrote:
>>>>> On 3/6/14 5:30 PM, Christian Groves wrote:
>>>>>> Hello Paul,
>>>>>>
>>>>>> The text says media type and presentation attribute. Is that the
>>>>>> relationship you're talking about?
>>>>>
>>>>> "How the generated content relates to the entity described in the
>>>>> participant info is dependent on media type and and the presentation
>>>>> attribute."
>>>>>
>>>>> I take that to mean that the relationship may be different for
>>>>> presentation streams than non-presentation streams. But it doesn't 
>>>>> say
>>>>> *how* they differ.
>>>>>
>>>>>     Thanks,
>>>>>     Paul
>>>>>
>>>>>> Regards, Christian
>>>>>>
>>>>>> On 7/03/2014 4:20 AM, Paul Kyzivat wrote:
>>>>>>> Christian,
>>>>>>>
>>>>>>> I've read the quoted text several times, and I cannot figure out 
>>>>>>> how
>>>>>>> to *derive* your example conclusions from it. The text says the
>>>>>>> relationship is dependent on the presentation attribute, but not 
>>>>>>> how.
>>>>>>>
>>>>>>> AFAICT I could make a new definition where the a participant 
>>>>>>> attached
>>>>>>> to a presentation capture means that the participant is shown in 
>>>>>>> the
>>>>>>> presentation, and that would be equally compatible with the text.
>>>>>>>
>>>>>>> ISTM that more text is required to actually specify the 
>>>>>>> relationships.
>>>>>>>
>>>>>>>     Thanks,
>>>>>>>     Paul
>>>>>>>
>>>>>>> On 3/6/14 2:05 PM, Christian Groves wrote:
>>>>>>>> Hello all,
>>>>>>>>
>>>>>>>> To follow up on Jonathon's comments on participant info/type 
>>>>>>>> and the
>>>>>>>> semantics and particularly how it relates to a presentation. 
>>>>>>>> Here's a
>>>>>>>> first stab at some text to stimulate some discussions.
>>>>>>>>
>>>>>>>>
>>>>>>>> "The participant info attribute allows a provider to associate
>>>>>>>> participant information with the capture source. When used in an
>>>>>>>> individual capture it indicates that the captured content (e.g.
>>>>>>>> video/audio/text etc.) as opposed to the actual media streams is
>>>>>>>> generated from the entity described. How the generated content 
>>>>>>>> relates
>>>>>>>> to the entity described in the participant info is dependent on 
>>>>>>>> media
>>>>>>>> type and and the presentation attribute.
>>>>>>>>
>>>>>>>> For example:
>>>>>>>> - a video capture with participant info would indicate that the 
>>>>>>>> video
>>>>>>>> contains a picture of the entity associated with the information
>>>>>>>> provided.
>>>>>>>> - a video capture with participant info and the presentation 
>>>>>>>> attribute
>>>>>>>> would indicate that the presentation video is associated with the
>>>>>>>> participant but could contain any video content.
>>>>>>>> - a text capture with participant info would indicate that the 
>>>>>>>> text is
>>>>>>>> generated from the actual participant.
>>>>>>>> - a text capture with participant info and the presentation 
>>>>>>>> attribute
>>>>>>>> would indicate that the text is associated with the participant 
>>>>>>>> but
>>>>>>>> could contain any text content."
>>>>>>>>
>>>>>>>> Comments?
>>>>>>>>
>>>>>>>>
>>>>>>>> Regards, Christian
>>>>>>>>
>>>>>>>> _______________________________________________
>>>>>>>> clue mailing list
>>>>>>>> clue@ietf.org
>>>>>>>> https://www.ietf.org/mailman/listinfo/clue
>>>>>>>>
>>>>>>>
>>>>>>> _______________________________________________
>>>>>>> clue mailing list
>>>>>>> clue@ietf.org
>>>>>>> https://www.ietf.org/mailman/listinfo/clue
>>>>>>>
>>>>>>
>>>>>> _______________________________________________
>>>>>> clue mailing list
>>>>>> clue@ietf.org
>>>>>> https://www.ietf.org/mailman/listinfo/clue
>>>>>>
>>>>>
>>>>>
>>>>
>>>>
>>>
>>>
>>
>> _______________________________________________
>> clue mailing list
>> clue@ietf.org
>> https://www.ietf.org/mailman/listinfo/clue
>>
>
>


From nobody Tue Apr  1 21:14:00 2014
Return-Path: <Christian.Groves@nteczone.com>
X-Original-To: clue@ietfa.amsl.com
Delivered-To: clue@ietfa.amsl.com
Received: from localhost (ietfa.amsl.com [127.0.0.1]) by ietfa.amsl.com (Postfix) with ESMTP id 216561A00F3 for <clue@ietfa.amsl.com>; Tue,  1 Apr 2014 21:13:56 -0700 (PDT)
X-Virus-Scanned: amavisd-new at amsl.com
X-Spam-Flag: NO
X-Spam-Score: 0.5
X-Spam-Level: 
X-Spam-Status: No, score=0.5 tagged_above=-999 required=5 tests=[BAYES_00=-1.9, J_CHICKENPOX_111=0.6, J_CHICKENPOX_14=0.6, J_CHICKENPOX_15=0.6, J_CHICKENPOX_54=0.6] autolearn=no
Received: from mail.ietf.org ([4.31.198.44]) by localhost (ietfa.amsl.com [127.0.0.1]) (amavisd-new, port 10024) with ESMTP id ramB2m87SjW1 for <clue@ietfa.amsl.com>; Tue,  1 Apr 2014 21:13:52 -0700 (PDT)
Received: from cserver5.myshophosting.com (cserver5.myshophosting.com [175.107.161.1]) by ietfa.amsl.com (Postfix) with ESMTP id 686641A0107 for <clue@ietf.org>; Tue,  1 Apr 2014 21:13:52 -0700 (PDT)
Received: from ppp118-209-153-126.lns20.mel6.internode.on.net ([118.209.153.126]:51405 helo=[127.0.0.1]) by cserver5.myshophosting.com with esmtpsa (TLSv1:DHE-RSA-AES128-SHA:128) (Exim 4.82) (envelope-from <Christian.Groves@nteczone.com>) id 1WVCYa-0005Af-2p for clue@ietf.org; Wed, 02 Apr 2014 15:13:44 +1100
Message-ID: <533B8E76.2090009@nteczone.com>
Date: Wed, 02 Apr 2014 15:13:42 +1100
From: Christian Groves <Christian.Groves@nteczone.com>
User-Agent: Mozilla/5.0 (Windows NT 6.1; WOW64; rv:24.0) Gecko/20100101 Thunderbird/24.4.0
MIME-Version: 1.0
To: clue@ietf.org
References: <C6252EA94E00E44EADC3A2FEB59D4402020DF93F@xmb-aln-x07.cisco.com> <7594FB04B1934943A5C02806D1A2204B1D271311@ESESSMB209.ericsson.se>
In-Reply-To: <7594FB04B1934943A5C02806D1A2204B1D271311@ESESSMB209.ericsson.se>
Content-Type: text/plain; charset=windows-1252; format=flowed
Content-Transfer-Encoding: 8bit
X-AntiAbuse: This header was added to track abuse, please include it with any abuse report
X-AntiAbuse: Primary Hostname - cserver5.myshophosting.com
X-AntiAbuse: Original Domain - ietf.org
X-AntiAbuse: Originator/Caller UID/GID - [47 12] / [47 12]
X-AntiAbuse: Sender Address Domain - nteczone.com
X-Get-Message-Sender-Via: cserver5.myshophosting.com: authenticated_id: christian.groves@nteczone.com
X-Source: 
X-Source-Args: 
X-Source-Dir: 
Archived-At: http://mailarchive.ietf.org/arch/msg/clue/gIFJZeVgAguTo9KJ5LcuDVtUYkA
Subject: Re: [clue] Using BUNDLE with CLUE
X-BeenThere: clue@ietf.org
X-Mailman-Version: 2.1.15
Precedence: list
List-Id: CLUE - ControLling mUltiple streams for TElepresence <clue.ietf.org>
List-Unsubscribe: <https://www.ietf.org/mailman/options/clue>, <mailto:clue-request@ietf.org?subject=unsubscribe>
List-Archive: <http://www.ietf.org/mail-archive/web/clue/>
List-Post: <mailto:clue@ietf.org>
List-Help: <mailto:clue-request@ietf.org?subject=help>
List-Subscribe: <https://www.ietf.org/mailman/listinfo/clue>, <mailto:clue-request@ietf.org?subject=subscribe>
X-List-Received-Date: Wed, 02 Apr 2014 04:13:56 -0000

Hello Christer,

With respect to Q2, I didn't think the media feature tag was instead of 
group:CLUE?

What happens if we have two m=application SCTP/DTLS lines in an SDP? If 
we only use the media feature tag how do we indicate which m= would be 
specific to CLUE via SDP?

Regards, Christian

On 1/04/2014 9:27 PM, Christer Holmberg wrote:
>
> Hi,
>
> *Q1*: If the remote part has indicated support of BUNDLE in the first 
> answer, then I see no problem in adding the BUNDLE m- lines in the BAS 
> offer. But, of course you can also first send a BAS offer containing 
> the initial m- lines, and after that send a third offer with the 
> BUNDLE m- lines. I am not sure whether we would need to mandate one 
> way or another, but it would probably be good to describe both.
>
> *Q2*: Regarding your example, I hope we’ll use a media feature tag 
> instead of group:CLUE to indicate support of CLUE. At least nobody has 
> objected to it :)
>
> *Q3*: Regarding your text on Demultiplexing of RTP, the status has 
> changed.
>
> You CAN re-use the same PT value in multiple m- lines (due to the 
> limited number range), as long as the codec configuration associated 
> with the PT value is the same. The plan is that BUNDLE will specify 
> usage of an identifier, associated with an m- line, which is inserted 
> into the RTP packet by the sender, so that the receiver knows to which 
> m- line the RTP packet belongs. This will reflected in bundle-06, 
> which will be submitted in the near future.
>
> *Q4*: Regarding the directionality and number of BUNDLE groups, using 
> a BUNDLE group per direction was one option. Another could be to use a 
> BUNDLE group per media type. Etc. There could e.g. be QoS related 
> reasons why you want to use different BUNDLE groups (read: different 
> media 5-tuples). I guess the question is whether it would be useful 
> for CLUE entities to be able to indicate which “BUNDLE configurations” 
> they support. “Everything in one BUNDLE group” could be default.
>
> Regards,
>
> Christer
>
> *From:*clue [mailto:clue-bounces@ietf.org] *On Behalf Of *Robert 
> Hansen (rohanse2)
> *Sent:* 1. huhtikuuta 2014 13:07
> *To:* clue@ietf.org
> *Subject:* [clue] Using BUNDLE with CLUE
>
> Introduction:
>
> A fundamental aspect of CLUE is the sending of multiple streams of 
> media. We are using conventional SDP to specify the encodings we 
> support, and hence each media stream necessitates a separate m-line. 
> For most use-cases, however, using a separate port per m-line is 
> suboptimal: it means opening more ports, more NAT work, more resources 
> consumed for ICE, etc. As such, a method to allow multiple m-lines to 
> share the same 5-tuple address is highly desirable.
>
> Until now the signalling work has focused on the demultiplexed case. 
> As such, this is an attempt to evaluate the use of BUNDLE with CLUE - 
> those with a much better understanding of BUNDLE than I will be able 
> to correct the mistakes I'll inevitably make.
>
> We have generally figured that CLUE would not raise any particular 
> issues with BUNDLE, given that CLUE now uses separate m-lines for 
> encodings, but we need to go through and make sure, as well as 
> providing guidance in the CLUE documentation on how BUNDLE interacts 
> with it.
>
> Obviously I'm not going to duplicate the BUNDLE draft here, instead 
> I'll just make reference to the draft.
>
> Initial offer/answer:
>
> One of the principle concerns of BUNDLE is to ensure that a receiver 
> does not receive an SDP it considers invalid due to its lack of 
> support for bundling m-lines. The fact that we recommend that 
> CLUE-controlled m-lines are not included in the initial O/A means that 
> I think that we can combine the extra O/A that BUNDLE normally adds 
> over an unBUNDLEd call with one of the O/As required by CLUE...
>
> Adding CLUE-controlled m-lines during the Bundle Address 
> Synchronization (BAS) offer:
>
> Having completed the initial O/A and established BUNDLE support the 
> initial offerer needs to send a new offer to synchronise the BUNDLE 
> addresses and make any intermediary devices aware of the addresses in use.
>
> In CLUE, the subsequent offer is also when we would like to start 
> adding CLUE-controlled m-lines. However, section 6.4.3. of the BUNDLE 
> specification warns that it important that the BAS offer is accepted, 
> and while it makes clear that the offerer MAY change the SDP, it warns 
> to avoid changes that could cause the answerer to reject the new offer.
>
> CLUE definitely needs to provide guidance here. My belief is that 
> adding the CLUE-controlled m-lines at this stage should not increase 
> the chance of the offer as a whole being rejected, so long as they 
> share the same media types and attributes as the existing media lines. 
> This would also help resolve the glare issue at the start of a CLUE 
> call: since a BAS is mandatory in BUNDLE this provides an obvious 'who 
> should reINVITE first' case for CLUE, where the initial offerer has to 
> send an new offer even if they don't have any CLUE encodings to add.
>
> If we did feel that adding new m-lines in the BAS offer was too high a 
> risk then BUNDLE and CLUE become sequential: BAS should be done first, 
> and then CLUE-controlled m-lines should be added in a subsequent 
> INVITE. In this case BUNDLE would include one extra O/A compared to 
> the unBUNDLEd case.
>
> I've included an example below, modifying the example from the BUNDLE 
> draft to show the initial offerer using the BAS offer to also add 
> CLUE-controlled media.
>
> Initial offer example - the offer includes an audio and video line 
> with unique ports, both in the same BUNDLE group (indicating that the 
> offerer supports BUNDLE and wants to multiplex these media flows). 
> There is also a data channel that will be used for CLUE, which is 
> included in the CLUE group.
>
> a=group:BUNDLE foo bar
>
> a=group:CLUE zen
>
> m=audio 10000 RTP/AVP 0 8 106
>
> ...
>
> a=mid:foo
>
> m=video 10002 RTP/AVP 96 97
>
> ...
>
> a=mid:bar
>
> m=application 10004 SCTP/DTLS 10004
>
> ...
>
> a=mid:zen
>
> Initial answer example - the answer picks a local BUNDLE address and 
> (via ordering in the group attribute) selects an address for the offerer.
>
> a=group:BUNDLE foo bar
>
> a=group:CLUE zen
>
> m=audio 20000 RTP/AVP 106
>
> ...
>
> a=mid:foo
>
> m=video 20000 RTP/AVP 96
>
> ...
>
> a=mid:bar
>
> m=application 20002 SCTP/DTLS 20002
>
> ...
>
> a=mid:zen
>
> Subsequent offer example - the offerer synchronises addresses and adds 
> CLUE-controlled media lines using the same address; these are included 
> in both the BUNDLE and CLUE groups.
>
> a=group:BUNDLE foo bar enc1 enc2 enc3
>
> a=group:CLUE zen enc1 enc2 enc3
>
> m=audio 10000 RTP/AVP 0 8 106
>
> ...
>
> a=mid:foo
>
> m=video 10000 RTP/AVP 96 97
>
> ...
>
> a=mid:bar
>
> m=application 10004 SCTP/DTLS 10004
>
> ...
>
> a=mid:zen
>
> m=video 10000 RTP/AVP 96 97
>
> ...
>
> a=mid:enc1
>
> a=label:1
>
> m=video 10000 RTP/AVP 96 97
>
> ...
>
> a=mid:enc2
>
> a=label:2
>
> m=video 10000 RTP/AVP 96 97
>
> ...
>
> a=mid:enc3
>
> a=label:3
>
> Answerer rejects BUNDLE:
>
> If the answer does not support BUNDLE then the offerer continues as in 
> the disaggregated case, though it may decide to offer fewer streams, 
> or even not do CLUE at all (in the latter case it should reINVITE and 
> remove the CLUE group and data channel).
>
> Effects of multiplexing on CLUE-relevant SDP attributes:
>
> The BUNDLE draft specifies how certain SDP attributes are affected by 
> multiplexing the m-lines, and draft-ietf-mmusic-sdp-mux-attributes 
> describes how multiplexing affects many other SDP attributes. I can't 
> see any CLUE-specific issues here: the 'label' attribute is unaffected 
> by multiplexing, nor is the directionality of the m-lines, and the 
> 'mid' attribute needed by CLUE is also needed by BUNDLE.
>
> Directionality of CLUE-controlled media:
>
> CLUE-controlled m-lines are currently unidirectional. I don't believe 
> this raises any specific BUNDLE issues. Christer suggests that 
> potentially the encodings on each side could be in separate BUNDLE 
> groups - I suggest we use the same group for both sides, which would 
> also make it easier to move to bidirectional streams if we ever wanted 
> to do that.
>
> BUNDLE support in CLUE:
>
> While in most cases multiplexing m-lines onto a single 5-tuple will be 
> preferable we do have use cases involving disaggregated media. Because 
> of this, and because my understanding of BUNDLE is that it does not 
> support the aggregated to disaggregated use case (where one device 
> send/receives media associated with multiple m-lines on a single 
> IP/port, while the other sends/receives the media associated with 
> multiple m-lines on multiple IP/ports) I don't see a reason to mandate 
> BUNDLE support for CLUE devices that only operate in disaggregated 
> mode. As such, while CLUE should document how to use BUNDLE to 
> multiplex media streams, it shouldn't be mandatory for CLUE devices to 
> support it.
>
> Demultiplexing of RTP:
>
> Media flows from the same BUNDLE group all belong to the same RTP 
> session; hence a receiver needs a method is needed to separate the 
> various capture encodings. The method currently defined in BUNDLE is 
> very straightforward: all of the dynamic payload types in each m-line 
> in a BUNDLE group must be unique; as such, any received RTP packet can 
> be mapped to a specific m-line and codec.
>
> This method, however, does suffer from the fact that the dynamic 
> payload type range is limited. In the case of many m-lines, of audio 
> (with potentially many codecs) and video bundled together, and/or the 
> increasing use of scalable video codecs with multiple layers and the 
> use of FEC repair flows, this space may be insufficient.
>
> There are other methods that could be used for this demultiplexing. 
> For implementation with static sources or RTP mixers, streams will 
> have a consistent SSRC, and hence the a=ssrc attribute can be supplied 
> by the sender to allow the receiver to differentiate flows via their 
> SSRC. For implementations where the streams being sent do not have a 
> static SSRC, such as source projection mixers, a specific 
> stream-correlator can be used, such as an AppID token specified by the 
> sender in SDP and in an RTP header extension.
>
> How stream demultiplexing is performed with BUNDLE is not a problem 
> specific to CLUE, and is not a problem that should be solved in a 
> CLUE-specific fashion. However, it is one that we need to ensure is 
> solved in a way consistent with our use-cases, and the number of 
> encodings we envisage allowing. Given that the simultaneous sending of 
> many capture encodings is a primary motivator behind CLUE I believe 
> that demultiplexing via payload will not be sufficient, and that we 
> would need additional methods of differentiation.
>
> If not all payload types must be unique then the question of under 
> what circumstances codecs on two m-lines can share the same payload 
> type becomes relevant.
>
> Summary:
>
> I believe BUNDLE is a good solution for how we can multiplex media 
> streams when doing CLUE, and I can't see any CLUE-specific issues that 
> would need to be addressed in BUNDLE. I think the main decision CLUE 
> would need to make is whether do first do BUNDLE negotiation 
> (including address synchronisation) and then begin adding CLUE 
> m-lines, or whether the CLUE media negotiation process can begin with 
> the BAS, essentially eliminating the extra O/A 'cost' of BUNDLE.
>
>
>
> _______________________________________________
> clue mailing list
> clue@ietf.org
> https://www.ietf.org/mailman/listinfo/clue


From nobody Tue Apr  1 22:47:23 2014
Return-Path: <christer.holmberg@ericsson.com>
X-Original-To: clue@ietfa.amsl.com
Delivered-To: clue@ietfa.amsl.com
Received: from localhost (ietfa.amsl.com [127.0.0.1]) by ietfa.amsl.com (Postfix) with ESMTP id E80011A0137 for <clue@ietfa.amsl.com>; Tue,  1 Apr 2014 22:47:19 -0700 (PDT)
X-Virus-Scanned: amavisd-new at amsl.com
X-Spam-Flag: NO
X-Spam-Score: -1.451
X-Spam-Level: 
X-Spam-Status: No, score=-1.451 tagged_above=-999 required=5 tests=[BAYES_00=-1.9, HELO_EQ_SE=0.35, J_CHICKENPOX_111=0.6, J_CHICKENPOX_14=0.6, J_CHICKENPOX_15=0.6, J_CHICKENPOX_54=0.6, RCVD_IN_DNSWL_MED=-2.3, SPF_PASS=-0.001] autolearn=no
Received: from mail.ietf.org ([4.31.198.44]) by localhost (ietfa.amsl.com [127.0.0.1]) (amavisd-new, port 10024) with ESMTP id NBAI7Wj-ib3s for <clue@ietfa.amsl.com>; Tue,  1 Apr 2014 22:47:16 -0700 (PDT)
Received: from mailgw2.ericsson.se (mailgw2.ericsson.se [193.180.251.37]) by ietfa.amsl.com (Postfix) with ESMTP id 325311A0135 for <clue@ietf.org>; Tue,  1 Apr 2014 22:47:15 -0700 (PDT)
X-AuditID: c1b4fb25-b7f3b8e0000006f1-72-533ba45eeb25
Received: from ESESSHC022.ericsson.se (Unknown_Domain [153.88.253.124]) by mailgw2.ericsson.se (Symantec Mail Security) with SMTP id 15.B8.01777.E54AB335; Wed,  2 Apr 2014 07:47:10 +0200 (CEST)
Received: from ESESSMB209.ericsson.se ([169.254.9.213]) by ESESSHC022.ericsson.se ([153.88.183.84]) with mapi id 14.03.0174.001; Wed, 2 Apr 2014 07:47:09 +0200
From: Christer Holmberg <christer.holmberg@ericsson.com>
To: Christian Groves <Christian.Groves@nteczone.com>, "clue@ietf.org" <clue@ietf.org>
Thread-Topic: [clue] Using BUNDLE with CLUE
Thread-Index: Ac9Nkg5IUPNx7ldUTDK8kHk510O6AQAAOGwQACGOwwAAB1PM0A==
Date: Wed, 2 Apr 2014 05:47:09 +0000
Message-ID: <7594FB04B1934943A5C02806D1A2204B1D272932@ESESSMB209.ericsson.se>
References: <C6252EA94E00E44EADC3A2FEB59D4402020DF93F@xmb-aln-x07.cisco.com> <7594FB04B1934943A5C02806D1A2204B1D271311@ESESSMB209.ericsson.se> <533B8E76.2090009@nteczone.com>
In-Reply-To: <533B8E76.2090009@nteczone.com>
Accept-Language: en-US
Content-Language: en-US
X-MS-Has-Attach: 
X-MS-TNEF-Correlator: 
x-originating-ip: [153.88.183.19]
Content-Type: text/plain; charset="us-ascii"
Content-Transfer-Encoding: quoted-printable
MIME-Version: 1.0
X-Brightmail-Tracker: H4sIAAAAAAAAA+NgFtrHLMWRmVeSWpSXmKPExsUyM+JvjW7cEutgg8ldahZf3jeyWOw/dZnZ gcljyZKfTB4rzs9kCWCK4rJJSc3JLEst0rdL4Mo4v3wpa8HclIo55yewNjBuCuhi5OSQEDCR 2HB+NwuELSZx4d56NhBbSOAwo8Tr9SVdjFxA9mJGiTMz/zB2MXJwsAlYSHT/0wapEREIl+jY doURxBYW0JJYff0PK0RcW2LF/R4mkHIRASeJrWtzQMIsAioSSzdsACvnFfCVeLJ3HxvE+E2M EpdenWQBqecU0JHoPgxWzwh0zvdTa5hAbGYBcYlbT+YzQZwpILFkz3lmCFtU4uXjf6wQtqJE +9MGRoh6HYkFuz+xQdjaEssWvmaG2CsocXLmE5YJjKKzkIydhaRlFpKWWUhaFjCyrGJkz03M zEkvN9rECIyDg1t+q+5gvHNO5BCjNAeLkjjvh7fOQUIC6YklqdmpqQWpRfFFpTmpxYcYmTg4 pRoY10rLft+wV5V3vlrFHvE7m62dIlIdPvjuesByUTnU/DkjT8+H3W8cmHMm31uadfOtme4T bjObBWK63/W27b/MxSHPdUfG/W9twNFCro5fGyoSmfUZWGfdiqprSa6fv4kr7WTD7xeJk67V 7I95U/wjO6889UdPorLw8/jW4kcZZjKyin/2vDmlxFKckWioxVxUnAgAVfSFBFECAAA=
Archived-At: http://mailarchive.ietf.org/arch/msg/clue/LA9V6U2LtyG_Ki1ptuaNsbnacyk
Subject: Re: [clue] Using BUNDLE with CLUE
X-BeenThere: clue@ietf.org
X-Mailman-Version: 2.1.15
Precedence: list
List-Id: CLUE - ControLling mUltiple streams for TElepresence <clue.ietf.org>
List-Unsubscribe: <https://www.ietf.org/mailman/options/clue>, <mailto:clue-request@ietf.org?subject=unsubscribe>
List-Archive: <http://www.ietf.org/mail-archive/web/clue/>
List-Post: <mailto:clue@ietf.org>
List-Help: <mailto:clue-request@ietf.org?subject=help>
List-Subscribe: <https://www.ietf.org/mailman/listinfo/clue>, <mailto:clue-request@ietf.org?subject=subscribe>
X-List-Received-Date: Wed, 02 Apr 2014 05:47:20 -0000

Hi,

>With respect to Q2, I didn't think the media feature tag was instead of gr=
oup:CLUE?

We would still use group:CLUE to group m=3D lines controlled by CLUE, but w=
e would not use group:CLUE as a generic CLUE support indicator. I.e. you wo=
uld not include group:CLUE unless you are actually grouping CLUE m=3D lines=
.

>What happens if we have two m=3Dapplication SCTP/DTLS lines in an SDP? If =
we only use the media feature tag how do we indicate which m=3D would be sp=
ecific to CLUE via SDP?

The idea has been that, within a CLUE group, we only have ONE m=3Dapplicati=
on SCTP/DTLS line.

Regards,

Christer


Regards, Christian

On 1/04/2014 9:27 PM, Christer Holmberg wrote:
>
> Hi,
>
> *Q1*: If the remote part has indicated support of BUNDLE in the first=20
> answer, then I see no problem in adding the BUNDLE m- lines in the BAS=20
> offer. But, of course you can also first send a BAS offer containing=20
> the initial m- lines, and after that send a third offer with the=20
> BUNDLE m- lines. I am not sure whether we would need to mandate one=20
> way or another, but it would probably be good to describe both.
>
> *Q2*: Regarding your example, I hope we'll use a media feature tag=20
> instead of group:CLUE to indicate support of CLUE. At least nobody has=20
> objected to it :)
>
> *Q3*: Regarding your text on Demultiplexing of RTP, the status has=20
> changed.
>
> You CAN re-use the same PT value in multiple m- lines (due to the=20
> limited number range), as long as the codec configuration associated=20
> with the PT value is the same. The plan is that BUNDLE will specify=20
> usage of an identifier, associated with an m- line, which is inserted=20
> into the RTP packet by the sender, so that the receiver knows to which
> m- line the RTP packet belongs. This will reflected in bundle-06,=20
> which will be submitted in the near future.
>
> *Q4*: Regarding the directionality and number of BUNDLE groups, using=20
> a BUNDLE group per direction was one option. Another could be to use a=20
> BUNDLE group per media type. Etc. There could e.g. be QoS related=20
> reasons why you want to use different BUNDLE groups (read: different=20
> media 5-tuples). I guess the question is whether it would be useful=20
> for CLUE entities to be able to indicate which "BUNDLE configurations"
> they support. "Everything in one BUNDLE group" could be default.
>
> Regards,
>
> Christer
>
> *From:*clue [mailto:clue-bounces@ietf.org] *On Behalf Of *Robert=20
> Hansen (rohanse2)
> *Sent:* 1. huhtikuuta 2014 13:07
> *To:* clue@ietf.org
> *Subject:* [clue] Using BUNDLE with CLUE
>
> Introduction:
>
> A fundamental aspect of CLUE is the sending of multiple streams of=20
> media. We are using conventional SDP to specify the encodings we=20
> support, and hence each media stream necessitates a separate m-line.
> For most use-cases, however, using a separate port per m-line is
> suboptimal: it means opening more ports, more NAT work, more resources=20
> consumed for ICE, etc. As such, a method to allow multiple m-lines to=20
> share the same 5-tuple address is highly desirable.
>
> Until now the signalling work has focused on the demultiplexed case.=20
> As such, this is an attempt to evaluate the use of BUNDLE with CLUE -=20
> those with a much better understanding of BUNDLE than I will be able=20
> to correct the mistakes I'll inevitably make.
>
> We have generally figured that CLUE would not raise any particular=20
> issues with BUNDLE, given that CLUE now uses separate m-lines for=20
> encodings, but we need to go through and make sure, as well as=20
> providing guidance in the CLUE documentation on how BUNDLE interacts=20
> with it.
>
> Obviously I'm not going to duplicate the BUNDLE draft here, instead=20
> I'll just make reference to the draft.
>
> Initial offer/answer:
>
> One of the principle concerns of BUNDLE is to ensure that a receiver=20
> does not receive an SDP it considers invalid due to its lack of=20
> support for bundling m-lines. The fact that we recommend that=20
> CLUE-controlled m-lines are not included in the initial O/A means that=20
> I think that we can combine the extra O/A that BUNDLE normally adds=20
> over an unBUNDLEd call with one of the O/As required by CLUE...
>
> Adding CLUE-controlled m-lines during the Bundle Address=20
> Synchronization (BAS) offer:
>
> Having completed the initial O/A and established BUNDLE support the=20
> initial offerer needs to send a new offer to synchronise the BUNDLE=20
> addresses and make any intermediary devices aware of the addresses in use=
.
>
> In CLUE, the subsequent offer is also when we would like to start=20
> adding CLUE-controlled m-lines. However, section 6.4.3. of the BUNDLE=20
> specification warns that it important that the BAS offer is accepted,=20
> and while it makes clear that the offerer MAY change the SDP, it warns=20
> to avoid changes that could cause the answerer to reject the new offer.
>
> CLUE definitely needs to provide guidance here. My belief is that=20
> adding the CLUE-controlled m-lines at this stage should not increase=20
> the chance of the offer as a whole being rejected, so long as they=20
> share the same media types and attributes as the existing media lines.
> This would also help resolve the glare issue at the start of a CLUE
> call: since a BAS is mandatory in BUNDLE this provides an obvious 'who=20
> should reINVITE first' case for CLUE, where the initial offerer has to=20
> send an new offer even if they don't have any CLUE encodings to add.
>
> If we did feel that adding new m-lines in the BAS offer was too high a=20
> risk then BUNDLE and CLUE become sequential: BAS should be done first,=20
> and then CLUE-controlled m-lines should be added in a subsequent=20
> INVITE. In this case BUNDLE would include one extra O/A compared to=20
> the unBUNDLEd case.
>
> I've included an example below, modifying the example from the BUNDLE=20
> draft to show the initial offerer using the BAS offer to also add=20
> CLUE-controlled media.
>
> Initial offer example - the offer includes an audio and video line=20
> with unique ports, both in the same BUNDLE group (indicating that the=20
> offerer supports BUNDLE and wants to multiplex these media flows).
> There is also a data channel that will be used for CLUE, which is=20
> included in the CLUE group.
>
> a=3Dgroup:BUNDLE foo bar
>
> a=3Dgroup:CLUE zen
>
> m=3Daudio 10000 RTP/AVP 0 8 106
>
> ...
>
> a=3Dmid:foo
>
> m=3Dvideo 10002 RTP/AVP 96 97
>
> ...
>
> a=3Dmid:bar
>
> m=3Dapplication 10004 SCTP/DTLS 10004
>
> ...
>
> a=3Dmid:zen
>
> Initial answer example - the answer picks a local BUNDLE address and=20
> (via ordering in the group attribute) selects an address for the offerer.
>
> a=3Dgroup:BUNDLE foo bar
>
> a=3Dgroup:CLUE zen
>
> m=3Daudio 20000 RTP/AVP 106
>
> ...
>
> a=3Dmid:foo
>
> m=3Dvideo 20000 RTP/AVP 96
>
> ...
>
> a=3Dmid:bar
>
> m=3Dapplication 20002 SCTP/DTLS 20002
>
> ...
>
> a=3Dmid:zen
>
> Subsequent offer example - the offerer synchronises addresses and adds=20
> CLUE-controlled media lines using the same address; these are included=20
> in both the BUNDLE and CLUE groups.
>
> a=3Dgroup:BUNDLE foo bar enc1 enc2 enc3
>
> a=3Dgroup:CLUE zen enc1 enc2 enc3
>
> m=3Daudio 10000 RTP/AVP 0 8 106
>
> ...
>
> a=3Dmid:foo
>
> m=3Dvideo 10000 RTP/AVP 96 97
>
> ...
>
> a=3Dmid:bar
>
> m=3Dapplication 10004 SCTP/DTLS 10004
>
> ...
>
> a=3Dmid:zen
>
> m=3Dvideo 10000 RTP/AVP 96 97
>
> ...
>
> a=3Dmid:enc1
>
> a=3Dlabel:1
>
> m=3Dvideo 10000 RTP/AVP 96 97
>
> ...
>
> a=3Dmid:enc2
>
> a=3Dlabel:2
>
> m=3Dvideo 10000 RTP/AVP 96 97
>
> ...
>
> a=3Dmid:enc3
>
> a=3Dlabel:3
>
> Answerer rejects BUNDLE:
>
> If the answer does not support BUNDLE then the offerer continues as in=20
> the disaggregated case, though it may decide to offer fewer streams,=20
> or even not do CLUE at all (in the latter case it should reINVITE and=20
> remove the CLUE group and data channel).
>
> Effects of multiplexing on CLUE-relevant SDP attributes:
>
> The BUNDLE draft specifies how certain SDP attributes are affected by=20
> multiplexing the m-lines, and draft-ietf-mmusic-sdp-mux-attributes
> describes how multiplexing affects many other SDP attributes. I can't=20
> see any CLUE-specific issues here: the 'label' attribute is unaffected=20
> by multiplexing, nor is the directionality of the m-lines, and the=20
> 'mid' attribute needed by CLUE is also needed by BUNDLE.
>
> Directionality of CLUE-controlled media:
>
> CLUE-controlled m-lines are currently unidirectional. I don't believe=20
> this raises any specific BUNDLE issues. Christer suggests that=20
> potentially the encodings on each side could be in separate BUNDLE=20
> groups - I suggest we use the same group for both sides, which would=20
> also make it easier to move to bidirectional streams if we ever wanted=20
> to do that.
>
> BUNDLE support in CLUE:
>
> While in most cases multiplexing m-lines onto a single 5-tuple will be=20
> preferable we do have use cases involving disaggregated media. Because=20
> of this, and because my understanding of BUNDLE is that it does not=20
> support the aggregated to disaggregated use case (where one device=20
> send/receives media associated with multiple m-lines on a single=20
> IP/port, while the other sends/receives the media associated with=20
> multiple m-lines on multiple IP/ports) I don't see a reason to mandate=20
> BUNDLE support for CLUE devices that only operate in disaggregated=20
> mode. As such, while CLUE should document how to use BUNDLE to=20
> multiplex media streams, it shouldn't be mandatory for CLUE devices to=20
> support it.
>
> Demultiplexing of RTP:
>
> Media flows from the same BUNDLE group all belong to the same RTP=20
> session; hence a receiver needs a method is needed to separate the=20
> various capture encodings. The method currently defined in BUNDLE is=20
> very straightforward: all of the dynamic payload types in each m-line=20
> in a BUNDLE group must be unique; as such, any received RTP packet can=20
> be mapped to a specific m-line and codec.
>
> This method, however, does suffer from the fact that the dynamic=20
> payload type range is limited. In the case of many m-lines, of audio=20
> (with potentially many codecs) and video bundled together, and/or the=20
> increasing use of scalable video codecs with multiple layers and the=20
> use of FEC repair flows, this space may be insufficient.
>
> There are other methods that could be used for this demultiplexing.=20
> For implementation with static sources or RTP mixers, streams will=20
> have a consistent SSRC, and hence the a=3Dssrc attribute can be supplied=
=20
> by the sender to allow the receiver to differentiate flows via their=20
> SSRC. For implementations where the streams being sent do not have a=20
> static SSRC, such as source projection mixers, a specific=20
> stream-correlator can be used, such as an AppID token specified by the=20
> sender in SDP and in an RTP header extension.
>
> How stream demultiplexing is performed with BUNDLE is not a problem=20
> specific to CLUE, and is not a problem that should be solved in a=20
> CLUE-specific fashion. However, it is one that we need to ensure is=20
> solved in a way consistent with our use-cases, and the number of=20
> encodings we envisage allowing. Given that the simultaneous sending of=20
> many capture encodings is a primary motivator behind CLUE I believe=20
> that demultiplexing via payload will not be sufficient, and that we=20
> would need additional methods of differentiation.
>
> If not all payload types must be unique then the question of under=20
> what circumstances codecs on two m-lines can share the same payload=20
> type becomes relevant.
>
> Summary:
>
> I believe BUNDLE is a good solution for how we can multiplex media=20
> streams when doing CLUE, and I can't see any CLUE-specific issues that=20
> would need to be addressed in BUNDLE. I think the main decision CLUE=20
> would need to make is whether do first do BUNDLE negotiation=20
> (including address synchronisation) and then begin adding CLUE=20
> m-lines, or whether the CLUE media negotiation process can begin with=20
> the BAS, essentially eliminating the extra O/A 'cost' of BUNDLE.
>
>
>
> _______________________________________________
> clue mailing list
> clue@ietf.org
> https://www.ietf.org/mailman/listinfo/clue

_______________________________________________
clue mailing list
clue@ietf.org
https://www.ietf.org/mailman/listinfo/clue


From nobody Tue Apr  1 23:38:19 2014
Return-Path: <Christian.Groves@nteczone.com>
X-Original-To: clue@ietfa.amsl.com
Delivered-To: clue@ietfa.amsl.com
Received: from localhost (ietfa.amsl.com [127.0.0.1]) by ietfa.amsl.com (Postfix) with ESMTP id 9C5561A013A for <clue@ietfa.amsl.com>; Tue,  1 Apr 2014 23:38:17 -0700 (PDT)
X-Virus-Scanned: amavisd-new at amsl.com
X-Spam-Flag: NO
X-Spam-Score: -0.7
X-Spam-Level: 
X-Spam-Status: No, score=-0.7 tagged_above=-999 required=5 tests=[BAYES_00=-1.9, J_CHICKENPOX_111=0.6, J_CHICKENPOX_54=0.6] autolearn=no
Received: from mail.ietf.org ([4.31.198.44]) by localhost (ietfa.amsl.com [127.0.0.1]) (amavisd-new, port 10024) with ESMTP id MqAYBIQEV0Xb for <clue@ietfa.amsl.com>; Tue,  1 Apr 2014 23:38:13 -0700 (PDT)
Received: from cserver5.myshophosting.com (cserver5.myshophosting.com [175.107.161.1]) by ietfa.amsl.com (Postfix) with ESMTP id 3A4CF1A00DE for <clue@ietf.org>; Tue,  1 Apr 2014 23:38:13 -0700 (PDT)
Received: from ppp118-209-153-126.lns20.mel6.internode.on.net ([118.209.153.126]:53178 helo=[127.0.0.1]) by cserver5.myshophosting.com with esmtpsa (TLSv1:DHE-RSA-AES128-SHA:128) (Exim 4.82) (envelope-from <Christian.Groves@nteczone.com>) id 1WVEoH-0005sp-CU; Wed, 02 Apr 2014 17:38:05 +1100
Message-ID: <533BB04B.50405@nteczone.com>
Date: Wed, 02 Apr 2014 17:38:03 +1100
From: Christian Groves <Christian.Groves@nteczone.com>
User-Agent: Mozilla/5.0 (Windows NT 6.1; WOW64; rv:24.0) Gecko/20100101 Thunderbird/24.4.0
MIME-Version: 1.0
To: Christer Holmberg <christer.holmberg@ericsson.com>,  "clue@ietf.org" <clue@ietf.org>
References: <C6252EA94E00E44EADC3A2FEB59D4402020DF93F@xmb-aln-x07.cisco.com> <7594FB04B1934943A5C02806D1A2204B1D271311@ESESSMB209.ericsson.se> <533B8E76.2090009@nteczone.com> <7594FB04B1934943A5C02806D1A2204B1D272932@ESESSMB209.ericsson.se>
In-Reply-To: <7594FB04B1934943A5C02806D1A2204B1D272932@ESESSMB209.ericsson.se>
Content-Type: text/plain; charset=ISO-8859-1; format=flowed
Content-Transfer-Encoding: 7bit
X-AntiAbuse: This header was added to track abuse, please include it with any abuse report
X-AntiAbuse: Primary Hostname - cserver5.myshophosting.com
X-AntiAbuse: Original Domain - ietf.org
X-AntiAbuse: Originator/Caller UID/GID - [47 12] / [47 12]
X-AntiAbuse: Sender Address Domain - nteczone.com
X-Get-Message-Sender-Via: cserver5.myshophosting.com: authenticated_id: christian.groves@nteczone.com
X-Source: 
X-Source-Args: 
X-Source-Dir: 
Archived-At: http://mailarchive.ietf.org/arch/msg/clue/cmHWLTAj9VHzunlG4i7-I5JRJz4
Subject: Re: [clue] Using BUNDLE with CLUE
X-BeenThere: clue@ietf.org
X-Mailman-Version: 2.1.15
Precedence: list
List-Id: CLUE - ControLling mUltiple streams for TElepresence <clue.ietf.org>
List-Unsubscribe: <https://www.ietf.org/mailman/options/clue>, <mailto:clue-request@ietf.org?subject=unsubscribe>
List-Archive: <http://www.ietf.org/mail-archive/web/clue/>
List-Post: <mailto:clue@ietf.org>
List-Help: <mailto:clue-request@ietf.org?subject=help>
List-Subscribe: <https://www.ietf.org/mailman/listinfo/clue>, <mailto:clue-request@ietf.org?subject=subscribe>
X-List-Received-Date: Wed, 02 Apr 2014 06:38:17 -0000

Hello Christer,


On 2/04/2014 4:47 PM, Christer Holmberg wrote:
> Hi,
>
>> With respect to Q2, I didn't think the media feature tag was instead of group:CLUE?
> We would still use group:CLUE to group m= lines controlled by CLUE, but we would not use group:CLUE as a generic CLUE support indicator. I.e. you would not include group:CLUE unless you are actually grouping CLUE m= lines.
>
>> What happens if we have two m=application SCTP/DTLS lines in an SDP? If we only use the media feature tag how do we indicate which m= would be specific to CLUE via SDP?
> The idea has been that, within a CLUE group, we only have ONE m=application SCTP/DTLS line.
[CNG] I know, I was a little confused when you were talking about using 
a media tag "instead" of groupLCLUE. If you're still planning on using 
group:CLUE then its not an issue. So in Rob's example you'd still need a 
media tag CLUE AND group:CLUE in the first offer if you're advertising a 
m=application SCTP/DTLS line.
Regards, Christian


From nobody Wed Apr  2 00:01:52 2014
Return-Path: <christer.holmberg@ericsson.com>
X-Original-To: clue@ietfa.amsl.com
Delivered-To: clue@ietfa.amsl.com
Received: from localhost (ietfa.amsl.com [127.0.0.1]) by ietfa.amsl.com (Postfix) with ESMTP id 11D8A1A010D for <clue@ietfa.amsl.com>; Wed,  2 Apr 2014 00:01:50 -0700 (PDT)
X-Virus-Scanned: amavisd-new at amsl.com
X-Spam-Flag: NO
X-Spam-Score: -0.701
X-Spam-Level: 
X-Spam-Status: No, score=-0.701 tagged_above=-999 required=5 tests=[BAYES_00=-1.9, J_CHICKENPOX_111=0.6, J_CHICKENPOX_54=0.6, SPF_PASS=-0.001] autolearn=no
Received: from mail.ietf.org ([4.31.198.44]) by localhost (ietfa.amsl.com [127.0.0.1]) (amavisd-new, port 10024) with ESMTP id 9IfFnCuskWHJ for <clue@ietfa.amsl.com>; Wed,  2 Apr 2014 00:01:46 -0700 (PDT)
Received: from sesbmg20.ericsson.net (sesbmg20.ericsson.net [193.180.251.56]) by ietfa.amsl.com (Postfix) with ESMTP id 554931A00BE for <clue@ietf.org>; Wed,  2 Apr 2014 00:01:44 -0700 (PDT)
X-AuditID: c1b4fb38-b7f518e000000889-0f-533bb5d49ede
Received: from ESESSHC017.ericsson.se (Unknown_Domain [153.88.253.124]) by sesbmg20.ericsson.net (Symantec Mail Security) with SMTP id CD.71.02185.4D5BB335; Wed,  2 Apr 2014 09:01:40 +0200 (CEST)
Received: from ESESSMB209.ericsson.se ([169.254.9.213]) by ESESSHC017.ericsson.se ([153.88.183.69]) with mapi id 14.03.0174.001; Wed, 2 Apr 2014 09:01:38 +0200
From: Christer Holmberg <christer.holmberg@ericsson.com>
To: Christian Groves <Christian.Groves@nteczone.com>, "clue@ietf.org" <clue@ietf.org>
Thread-Topic: [clue] Using BUNDLE with CLUE
Thread-Index: Ac9Nkg5IUPNx7ldUTDK8kHk510O6AQAAOGwQACGOwwAAB1PM0P//7baA///ZsyA=
Date: Wed, 2 Apr 2014 07:01:38 +0000
Message-ID: <7594FB04B1934943A5C02806D1A2204B1D272B45@ESESSMB209.ericsson.se>
References: <C6252EA94E00E44EADC3A2FEB59D4402020DF93F@xmb-aln-x07.cisco.com> <7594FB04B1934943A5C02806D1A2204B1D271311@ESESSMB209.ericsson.se> <533B8E76.2090009@nteczone.com> <7594FB04B1934943A5C02806D1A2204B1D272932@ESESSMB209.ericsson.se> <533BB04B.50405@nteczone.com>
In-Reply-To: <533BB04B.50405@nteczone.com>
Accept-Language: en-US
Content-Language: en-US
X-MS-Has-Attach: 
X-MS-TNEF-Correlator: 
x-originating-ip: [153.88.183.19]
Content-Type: text/plain; charset="us-ascii"
Content-Transfer-Encoding: quoted-printable
MIME-Version: 1.0
X-Brightmail-Tracker: H4sIAAAAAAAAA+NgFjrILMWRmVeSWpSXmKPExsUyM+Jvje6VrdbBBguPaFl8ed/IYrH/1GVm ByaPJUt+MnmsOD+TJYApissmJTUnsyy1SN8ugSujq9W54D97xaFJ/9kbGFezdTFyckgImEgc eLySBcIWk7hwbz1QnItDSOAoo8TTT7NYIJzFjBJtf/cCZTg42AQsJLr/aYM0iAiES3Rsu8II YgsLaEmsvv6HFSKuLbHifg8ThO0ncf/vHDCbRUBF4tO5dcwgNq+Ar8SX59eYIeZPY5J4+us+ O0iCE2jQ64NdYEMZgS76fmoNWDOzgLjErSfzmSAuFZBYsuc8M4QtKvHy8T9WCFtRov1pAyNE vY7Egt2f2CBsbYllC19DLRaUODnzCcsERtFZSMbOQtIyC0nLLCQtCxhZVjFyFKcWJ+WmGxls YgRGw8Etvy12MF7+a3OIUZqDRUmc9+Nb5yAhgfTEktTs1NSC1KL4otKc1OJDjEwcnFINjGVz L50v8W9z/noj567chFcuPmZfH07bv6OheKuiV4zepl29Vyqz/x/rn7l568s6k1AmB7enMrW/ r/alTer0esdg1hy/+pS35/Q+/x+KLfy3pt3dxS2TbucqpPec83n7y7Xyt46vne54Qnx71Kd4 y/jVfo8df33jnte6dqFA0yehBIvVS353aSixFGckGmoxFxUnAgD4HqyqVAIAAA==
Archived-At: http://mailarchive.ietf.org/arch/msg/clue/gJ16OewXV2iOoYF7s0NY9QDICss
Subject: Re: [clue] Using BUNDLE with CLUE
X-BeenThere: clue@ietf.org
X-Mailman-Version: 2.1.15
Precedence: list
List-Id: CLUE - ControLling mUltiple streams for TElepresence <clue.ietf.org>
List-Unsubscribe: <https://www.ietf.org/mailman/options/clue>, <mailto:clue-request@ietf.org?subject=unsubscribe>
List-Archive: <http://www.ietf.org/mail-archive/web/clue/>
List-Post: <mailto:clue@ietf.org>
List-Help: <mailto:clue-request@ietf.org?subject=help>
List-Subscribe: <https://www.ietf.org/mailman/listinfo/clue>, <mailto:clue-request@ietf.org?subject=subscribe>
X-List-Received-Date: Wed, 02 Apr 2014 07:01:50 -0000

Hi,

>>>> With respect to Q2, I didn't think the media feature tag was instead o=
f group:CLUE?
>>> We would still use group:CLUE to group m=3D lines controlled by CLUE, b=
ut we would not use group:CLUE as a=20
>>> generic CLUE support indicator. I.e. you would not include group:CLUE u=
nless you are actually grouping CLUE=20
>>> m=3D lines.
>>
>>> What happens if we have two m=3Dapplication SCTP/DTLS lines in an SDP? =
If we only use the media feature tag how do we indicate which m=3D would be=
 specific to CLUE via SDP?
>> The idea has been that, within a CLUE group, we only have ONE m=3Dapplic=
ation SCTP/DTLS line.
> [CNG] I know, I was a little confused when you were talking about using a=
 media tag "instead" of groupLCLUE. If you're still planning on=20
> using group:CLUE then its not an issue. So in Rob's example you'd still n=
eed a media tag CLUE AND group:CLUE in the first offer if you're=20
> advertising a m=3Dapplication SCTP/DTLS line.

Correct.

Regards,

Christer


From nobody Wed Apr  2 07:12:17 2014
Return-Path: <mary.ietf.barnes@gmail.com>
X-Original-To: clue@ietfa.amsl.com
Delivered-To: clue@ietfa.amsl.com
Received: from localhost (ietfa.amsl.com [127.0.0.1]) by ietfa.amsl.com (Postfix) with ESMTP id 5232D1A0228 for <clue@ietfa.amsl.com>; Wed,  2 Apr 2014 07:12:15 -0700 (PDT)
X-Virus-Scanned: amavisd-new at amsl.com
X-Spam-Flag: NO
X-Spam-Score: -1.999
X-Spam-Level: 
X-Spam-Status: No, score=-1.999 tagged_above=-999 required=5 tests=[BAYES_00=-1.9, DKIM_SIGNED=0.1, DKIM_VALID=-0.1, DKIM_VALID_AU=-0.1, FREEMAIL_FROM=0.001, HTML_MESSAGE=0.001, SPF_PASS=-0.001] autolearn=ham
Received: from mail.ietf.org ([4.31.198.44]) by localhost (ietfa.amsl.com [127.0.0.1]) (amavisd-new, port 10024) with ESMTP id rBM9HehkjzFN for <clue@ietfa.amsl.com>; Wed,  2 Apr 2014 07:12:12 -0700 (PDT)
Received: from mail-wi0-x22c.google.com (mail-wi0-x22c.google.com [IPv6:2a00:1450:400c:c05::22c]) by ietfa.amsl.com (Postfix) with ESMTP id 05A711A0220 for <clue@ietf.org>; Wed,  2 Apr 2014 07:12:02 -0700 (PDT)
Received: by mail-wi0-f172.google.com with SMTP id hi2so5786004wib.5 for <clue@ietf.org>; Wed, 02 Apr 2014 07:11:58 -0700 (PDT)
DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20120113;  h=mime-version:in-reply-to:references:date:message-id:subject:from:to :cc:content-type; bh=hhgwyRHr+rHCFx1mQbW7LM7l8ezpfsx7j4UCCqNh4gM=; b=IdTWu9dR0g3ShoX6qREXCNcxOM70EP3LxX4Z2rIVNd/8Gc2wpscHdhi6UK11qWqkJD 6hEzc0Os0uYfJKFhlcxnlHZYEviroWz5cPNbsW6dv/Z6j8JiWJyf2MrYZTgy37z/deSY xVFqQPyO9uVHmckosbHBnQtqaXeq8DVuKxu1bYLTjdysA/IPmEsYoRIz3o3DcPoKIW8R Z6TZhH4qz8QSZ7TVDxybeETFb7OhNaHFCr1TeOWxmxUfTYRpNuu+oXrblHE1cD7pjz4w fg/I0NiBmGxrhvj5M4gtZ2KtKiQSXXqVFJLm6mULvU5xOj5Z1AWhsy4AihIUjn7LzUKv RT5g==
MIME-Version: 1.0
X-Received: by 10.194.203.170 with SMTP id kr10mr1118297wjc.19.1396447900556;  Wed, 02 Apr 2014 07:11:40 -0700 (PDT)
Received: by 10.216.10.6 with HTTP; Wed, 2 Apr 2014 07:11:40 -0700 (PDT)
In-Reply-To: <1502356412.3765966.1396447486992.POLL_ADMIN_PARTICIPATELINK.doodle@worker1>
References: <1502356412.3765966.1396447486992.POLL_ADMIN_PARTICIPATELINK.doodle@worker1>
Date: Wed, 2 Apr 2014 09:11:40 -0500
Message-ID: <CAHBDyN4Sit0MbG84oE0kDc-jRKCWQSX6Bf-Q2XaE15zF5CHj_g@mail.gmail.com>
From: Mary Barnes <mary.ietf.barnes@gmail.com>
To: CLUE <clue@ietf.org>
Content-Type: multipart/alternative; boundary=047d7b6d87de0f641104f60fdd0f
Archived-At: http://mailarchive.ietf.org/arch/msg/clue/FUAzGOOwwqI3UbHMv1GkvJhlIoo
Subject: [clue] Fwd: Doodle: Link for poll "CLUE DT meeting"
X-BeenThere: clue@ietf.org
X-Mailman-Version: 2.1.15
Precedence: list
List-Id: CLUE - ControLling mUltiple streams for TElepresence <clue.ietf.org>
List-Unsubscribe: <https://www.ietf.org/mailman/options/clue>, <mailto:clue-request@ietf.org?subject=unsubscribe>
List-Archive: <http://www.ietf.org/mail-archive/web/clue/>
List-Post: <mailto:clue@ietf.org>
List-Help: <mailto:clue-request@ietf.org?subject=help>
List-Subscribe: <https://www.ietf.org/mailman/listinfo/clue>, <mailto:clue-request@ietf.org?subject=subscribe>
X-List-Received-Date: Wed, 02 Apr 2014 14:12:15 -0000

--047d7b6d87de0f641104f60fdd0f
Content-Type: text/plain; charset=ISO-8859-1

As discussed on the call yesterday, there has been a suggestion to see if
we can change the time for the weekly design team meetings. I've created a
doodle with some options (including the current).   This will help us plan
topics if you will fill out for each week - e.g., if you won't make the
call on a specific week then we'll know that we shouldn't schedule
discussion of your document.

I've left next week's the same, however, if folks can please respond no
later than noon Central US on Monday, April 7th, we'll confirm by Tuesday.

Thanks,
Mary.

---------- Forwarded message ----------
From: Doodle <mailer@doodle.com>
Date: Wed, Apr 2, 2014 at 9:04 AM
Subject: Doodle: Link for poll "CLUE DT meeting"
To: Mary Barnes <mary.ietf.barnes@gmail.com>


You have initiated a poll "CLUE DT meeting" at Doodle. The link to your
poll is:

http://doodle.com/4xa7ndpfx42rgfr2

Share this link with all those who should cast their votes. Do not forget
to cast your vote, too.
(If you did not initiate this poll, somebody must accidentally have used
your e-mail address; simply ignore this e-mail, please.)

--047d7b6d87de0f641104f60fdd0f
Content-Type: text/html; charset=ISO-8859-1
Content-Transfer-Encoding: quoted-printable

<div dir=3D"ltr">As discussed on the call yesterday, there has been a sugge=
stion to see if we can change the time for the weekly design team meetings.=
 I&#39;ve created a doodle with some options (including the current). =A0 T=
his will help us plan topics if you will fill out for each week - e.g., if =
you won&#39;t make the call on a specific week then we&#39;ll know that we =
shouldn&#39;t schedule discussion of your document.=A0<div>
<br></div><div>I&#39;ve left next week&#39;s the same, however, if folks ca=
n please respond no later than noon Central US on Monday, April 7th, we&#39=
;ll confirm by Tuesday.=A0</div><div><br></div><div>Thanks,</div><div>Mary.=
<br>
<br><div class=3D"gmail_quote">---------- Forwarded message ----------<br>F=
rom: <b class=3D"gmail_sendername">Doodle</b> <span dir=3D"ltr">&lt;<a href=
=3D"mailto:mailer@doodle.com">mailer@doodle.com</a>&gt;</span><br>Date: Wed=
, Apr 2, 2014 at 9:04 AM<br>
Subject: Doodle: Link for poll &quot;CLUE DT meeting&quot;<br>To: Mary Barn=
es &lt;<a href=3D"mailto:mary.ietf.barnes@gmail.com">mary.ietf.barnes@gmail=
.com</a>&gt;<br><br><br>You have initiated a poll &quot;CLUE DT meeting&quo=
t; at Doodle. The link to your poll is:<br>

<br>
<a href=3D"http://doodle.com/4xa7ndpfx42rgfr2" target=3D"_blank">http://doo=
dle.com/4xa7ndpfx42rgfr2</a><br>
<br>
Share this link with all those who should cast their votes. Do not forget t=
o cast your vote, too.<br>
(If you did not initiate this poll, somebody must accidentally have used yo=
ur e-mail address; simply ignore this e-mail, please.)<br>
</div><br></div></div>

--047d7b6d87de0f641104f60fdd0f--


From nobody Wed Apr  2 14:56:51 2014
Return-Path: <Mark.Duckworth@polycom.com>
X-Original-To: clue@ietfa.amsl.com
Delivered-To: clue@ietfa.amsl.com
Received: from localhost (ietfa.amsl.com [127.0.0.1]) by ietfa.amsl.com (Postfix) with ESMTP id 3524A1A03F9 for <clue@ietfa.amsl.com>; Wed,  2 Apr 2014 14:56:50 -0700 (PDT)
X-Virus-Scanned: amavisd-new at amsl.com
X-Spam-Flag: NO
X-Spam-Score: -1.82
X-Spam-Level: 
X-Spam-Status: No, score=-1.82 tagged_above=-999 required=5 tests=[BAYES_00=-1.9, RCVD_IN_DNSWL_LOW=-0.7, SPF_NEUTRAL=0.779, UNPARSEABLE_RELAY=0.001] autolearn=ham
Received: from mail.ietf.org ([4.31.198.44]) by localhost (ietfa.amsl.com [127.0.0.1]) (amavisd-new, port 10024) with ESMTP id 57QfzEj_5jOx for <clue@ietfa.amsl.com>; Wed,  2 Apr 2014 14:56:46 -0700 (PDT)
Received: from mail1.bemta7.messagelabs.com (mail1.bemta7.messagelabs.com [216.82.254.106]) by ietfa.amsl.com (Postfix) with ESMTP id 143891A03FB for <clue@ietf.org>; Wed,  2 Apr 2014 14:56:44 -0700 (PDT)
Received: from [216.82.254.20:13636] by server-10.bemta-7.messagelabs.com id AE/E0-11882-7978C335; Wed, 02 Apr 2014 21:56:39 +0000
X-Env-Sender: Mark.Duckworth@polycom.com
X-Msg-Ref: server-12.tower-47.messagelabs.com!1396475799!4837505!1
X-Originating-IP: [140.242.64.158]
X-StarScan-Received: 
X-StarScan-Version: 6.11.1; banners=-,-,-
X-VirusChecked: Checked
Received: (qmail 20943 invoked from network); 2 Apr 2014 21:56:39 -0000
Received: from crpehubprd01.polycom.com (HELO crpehubprd02.polycom.com) (140.242.64.158) by server-12.tower-47.messagelabs.com with AES128-SHA encrypted SMTP; 2 Apr 2014 21:56:39 -0000
Received: from CRPMBOXPRD07.polycom.com ([fe80::8113:9ad1:f9be:53f1]) by crpehubprd02.polycom.com ([fe80::5efe:10.236.0.154%12]) with mapi; Wed, 2 Apr 2014 14:55:53 -0700
From: "Duckworth, Mark" <Mark.Duckworth@polycom.com>
To: Christian Groves <Christian.Groves@nteczone.com>, "clue@ietf.org" <clue@ietf.org>
Date: Wed, 2 Apr 2014 14:55:50 -0700
Thread-Topic: [clue] Global CSE List in framework
Thread-Index: Ac9NSMBYfWfdYYkHQreXi/z+PPijOABcka1w
Message-ID: <49E45C59CA48264997FEBFB29B6BC2D62151A17223@CRPMBOXPRD07.polycom.com>
References: <49E45C59CA48264997FEBFB29B6BC2D617ACF1B213@CRPMBOXPRD07.polycom.com> <533A14A0.6030807@nteczone.com>
In-Reply-To: <533A14A0.6030807@nteczone.com>
Accept-Language: en-US
Content-Language: en-US
X-MS-Has-Attach: 
X-MS-TNEF-Correlator: 
acceptlanguage: en-US
Content-Type: text/plain; charset="us-ascii"
Content-Transfer-Encoding: quoted-printable
MIME-Version: 1.0
Archived-At: http://mailarchive.ietf.org/arch/msg/clue/8K7c7U3fGUKXWOZqcb7KHj1bVI8
Subject: Re: [clue] Global CSE List in framework
X-BeenThere: clue@ietf.org
X-Mailman-Version: 2.1.15
Precedence: list
List-Id: CLUE - ControLling mUltiple streams for TElepresence <clue.ietf.org>
List-Unsubscribe: <https://www.ietf.org/mailman/options/clue>, <mailto:clue-request@ietf.org?subject=unsubscribe>
List-Archive: <http://www.ietf.org/mail-archive/web/clue/>
List-Post: <mailto:clue@ietf.org>
List-Help: <mailto:clue-request@ietf.org?subject=help>
List-Subscribe: <https://www.ietf.org/mailman/listinfo/clue>, <mailto:clue-request@ietf.org?subject=subscribe>
X-List-Received-Date: Wed, 02 Apr 2014 21:56:50 -0000

Hi Christian, thanks for the suggestions.  Comments below.
Mark

> -----Original Message-----
> From: clue [mailto:clue-bounces@ietf.org] On Behalf Of Christian Groves
> Sent: Monday, March 31, 2014 9:22 PM
> To: clue@ietf.org
> Subject: Re: [clue] Global CSE List in framework
>=20
> Hello Mark,
>=20
> Thanks for providing a start for the text. Some comments below.
>=20
> Regards, Christian
>=20
> On 28/03/2014 4:10 AM, Duckworth, Mark wrote:
> >
> > Here is a new section I'm adding to the framework. Please send any
> > comments or suggestions.
> >
> > =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=
=3D=3D=3D=3D=3D=3D=3D
> >
> > 7.3.3. Global Capture Scene Entry List
> >
> > An Advertisement can include an optional global Capture Scene Entry
> > list, for each media type. Each item in this list is a set of one or
> > more Capture Scene Entries. Each set of CSEs in the list is a
> > suggestion from the Provider to the Consumer for which CSEs represent
> > the entire advertisement, across multiple scenes.
> >
> [CNG] I'm not sure about the wording about "..for which CSEs represent th=
e
> entire advertisement..." For me the "advertisement" is the CLUE message.
> So it seems strange to talk about CSEs representing the advertisement.
> Perhaps something like "... for which CSEs provide a complete representat=
ion
> of the simultaneous captures provided by the provider, across multiple
> scenes."?
[Duckworth, Mark] I agree

> > The Provider can include multiple sets, to accommodate different
> > Consumers with varying capability to receive multiple encodings. This
> > is very similar to how each CSE represents a particular scene.
> >
> [CNG] Is the GCSE about providing sets for multiple "encodings"? You can
> provide multiple encodings without GSEs. Isn't it about providing differe=
nt
> sets of captures so that a receiver can best determine what the set of
> captures it wants based on the service it wants to provide.
> I.e. a consumer may only support three media streams (3 encodings) but
> these may be used in a number of ways.
[Duckworth, Mark] Are you suggesting just replace "encodings" with "capture=
s"?  That's fine with me, maybe "captures" better expresses the meaning.

> > As an example, suppose an advertisement has three scenes, and each
> > scene has three CSEs, ranging from one to three video captures in each
> > CSE. The provider is advertising a total of nine video captures across
> > three scenes. The provider can use the Global CSE list to suggest
> > alternatives for consumers that can't receive all nine video captures.
> >
> [CNG] Maybe we should add at the end "... all nine video captures <as
> separate media streams>."?
[Duckworth, Mark] I agree

> > For accommodating a consumer that wants to receive three video
> > captures, a provider might suggest a single CSE with three captures
> > and nothing from the other two scenes. Or a provider might suggest
> > three different CSEs, one from each scene, with a single video capture
> > in each. The choice of how to make these suggestions in the Global CSE
> > list for what represents the entire advertisement is up to the provider=
.
> >
> > Some additional rules:
> >
> [CNG] Perhaps also? "A CSE may be used in multiple sets." Otherwise its n=
ot
> clear whether this is allowed or not. Also I take it there's at most one =
CSE
> instance per SET.
[Duckworth, Mark] I agree

> > * The ordering of items (sets of CSEs) in the global CSE list is not
> > important.
> >
> > * The ordering of CSEs within each set is not important.
> >
> > * The Provider must be capable of encoding and sending all Captures
> > within the CSEs of a given set simultaneously.
> >
> [CNG] Do we need to add some text to the simultaneous set section of the
> framework indicating this? We should probably say something about the
> interaction. i.e. If a GCSE set is included should the provider actually =
need to
> provide a STS? Sendings CSEs in a STS indicates the those CSE/captures ma=
y
> be used simultaneously. There appears to be alot of overlap with the GCSE=
 in
> this respect.
[Duckworth, Mark] I think I see what you mean.  We have this paragraph alre=
ady:
"If an Advertisement does not include Simultaneous Transmission Sets, then =
the Provider MUST be able to provide all Capture Scenes simultaneously.  If=
 multiple capture Scene Entries are in a Capture Scene then the Consumer ch=
ooses at most one Capture Scene Entry per Capture Scene for each media type=
."

[Duckworth, Mark] I suggest we also add "If there is no STS and there is a =
global CSE list, then the Consumer chooses at most one set of CSEs of each =
media type, from the global CSE list."

>=20
>=20
> > =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=
=3D=3D=3D=3D=3D=3D=3D
> >
> > I also added a statement in 7.3 saying a CSE has an advertisement
> > unique identity.
> >
> > I also plan to add an example in 12.3 for Global CSE List, taken from
> > my IETF89 slides.
> >
> [CNG] OK
> >
> > Mark
> >
> >
> >
> > _______________________________________________
> > clue mailing list
> > clue@ietf.org
> > https://www.ietf.org/mailman/listinfo/clue
>=20
> _______________________________________________
> clue mailing list
> clue@ietf.org
> https://www.ietf.org/mailman/listinfo/clue


From nobody Wed Apr  2 14:58:30 2014
Return-Path: <Mark.Duckworth@polycom.com>
X-Original-To: clue@ietfa.amsl.com
Delivered-To: clue@ietfa.amsl.com
Received: from localhost (ietfa.amsl.com [127.0.0.1]) by ietfa.amsl.com (Postfix) with ESMTP id 0F9C31A03FB for <clue@ietfa.amsl.com>; Wed,  2 Apr 2014 14:58:29 -0700 (PDT)
X-Virus-Scanned: amavisd-new at amsl.com
X-Spam-Flag: NO
X-Spam-Score: -1.819
X-Spam-Level: 
X-Spam-Status: No, score=-1.819 tagged_above=-999 required=5 tests=[BAYES_00=-1.9, HTML_MESSAGE=0.001, RCVD_IN_DNSWL_LOW=-0.7, SPF_NEUTRAL=0.779, UNPARSEABLE_RELAY=0.001] autolearn=ham
Received: from mail.ietf.org ([4.31.198.44]) by localhost (ietfa.amsl.com [127.0.0.1]) (amavisd-new, port 10024) with ESMTP id mZuBA48-ow8h for <clue@ietfa.amsl.com>; Wed,  2 Apr 2014 14:58:24 -0700 (PDT)
Received: from mail1.bemta7.messagelabs.com (mail1.bemta7.messagelabs.com [216.82.254.106]) by ietfa.amsl.com (Postfix) with ESMTP id 5094E1A03FA for <clue@ietf.org>; Wed,  2 Apr 2014 14:58:24 -0700 (PDT)
Received: from [216.82.254.20:23223] by server-10.bemta-7.messagelabs.com id A5/A3-11882-CF78C335; Wed, 02 Apr 2014 21:58:20 +0000
X-Env-Sender: Mark.Duckworth@polycom.com
X-Msg-Ref: server-2.tower-47.messagelabs.com!1396475899!4104346!1
X-Originating-IP: [140.242.64.158]
X-StarScan-Received: 
X-StarScan-Version: 6.11.1; banners=-,-,-
X-VirusChecked: Checked
Received: (qmail 2530 invoked from network); 2 Apr 2014 21:58:19 -0000
Received: from crpehubprd01.polycom.com (HELO Crpehubprd01.polycom.com) (140.242.64.158) by server-2.tower-47.messagelabs.com with AES128-SHA encrypted SMTP; 2 Apr 2014 21:58:19 -0000
Received: from PWEHUB01.polycom.com (10.236.2.221) by Crpehubprd01.polycom.com (10.236.0.158) with Microsoft SMTP Server (TLS) id 8.3.192.1; Wed, 2 Apr 2014 14:58:19 -0700
Received: from CRPMBOXPRD07.polycom.com ([fe80::8113:9ad1:f9be:53f1]) by PWEHUB01.polycom.com ([fe80::99a8:f785:3f0c:2bb6%17]) with mapi; Wed, 2 Apr 2014 14:58:18 -0700
From: "Duckworth, Mark" <Mark.Duckworth@polycom.com>
To: Roberta Presta <roberta.presta@unina.it>, "clue@ietf.org" <clue@ietf.org>
Date: Wed, 2 Apr 2014 14:58:18 -0700
Thread-Topic: [clue] Global CSE List in framework
Thread-Index: Ac9NmLB7/SnwzYGaQMi0vcFcarJiWwBJbSgA
Message-ID: <49E45C59CA48264997FEBFB29B6BC2D62151A1722A@CRPMBOXPRD07.polycom.com>
References: <49E45C59CA48264997FEBFB29B6BC2D617ACF1B213@CRPMBOXPRD07.polycom.com> <533A9AC1.5060706@unina.it>
In-Reply-To: <533A9AC1.5060706@unina.it>
Accept-Language: en-US
Content-Language: en-US
X-MS-Has-Attach: 
X-MS-TNEF-Correlator: 
acceptlanguage: en-US
Content-Type: multipart/alternative; boundary="_000_49E45C59CA48264997FEBFB29B6BC2D62151A1722ACRPMBOXPRD07p_"
MIME-Version: 1.0
Archived-At: http://mailarchive.ietf.org/arch/msg/clue/y29W4V_YzBS5LRoCp4jMx_JJStE
Subject: Re: [clue] Global CSE List in framework
X-BeenThere: clue@ietf.org
X-Mailman-Version: 2.1.15
Precedence: list
List-Id: CLUE - ControLling mUltiple streams for TElepresence <clue.ietf.org>
List-Unsubscribe: <https://www.ietf.org/mailman/options/clue>, <mailto:clue-request@ietf.org?subject=unsubscribe>
List-Archive: <http://www.ietf.org/mail-archive/web/clue/>
List-Post: <mailto:clue@ietf.org>
List-Help: <mailto:clue-request@ietf.org?subject=help>
List-Subscribe: <https://www.ietf.org/mailman/listinfo/clue>, <mailto:clue-request@ietf.org?subject=subscribe>
X-List-Received-Date: Wed, 02 Apr 2014 21:58:29 -0000

--_000_49E45C59CA48264997FEBFB29B6BC2D62151A1722ACRPMBOXPRD07p_
Content-Type: text/plain; charset="us-ascii"
Content-Transfer-Encoding: quoted-printable

Hi Roberta,

Thanks for the suggestion, it makes sense.  I think the meaning is the same=
 either way, but keeping it consistent with the different media types in a =
CSE list sounds good.  I'll change the text in the framework.

Mark

From: Roberta Presta [mailto:roberta.presta@unina.it]
Sent: Tuesday, April 01, 2014 6:54 AM
To: Duckworth, Mark; clue@ietf.org
Subject: Re: [clue] Global CSE List in framework

Hello Mark,

just a minor question about your definition:


Il 27/03/2014 18:10, Duckworth, Mark ha scritto:
 An Advertisement can include an optional global Capture Scene Entry list, =
for each media type.  Each item in this list is a set of one or more Captur=
e Scene Entries.

Do you mean that we have a global capture scene entry list for video and a =
global capture scene entry list for audio?
Otherwise, what about "...global Capture Scene Entry list. Each item in the=
 list is a set of one or more Capture Scene Entries of the same media type.=
"? That last approach is more similar to the one adopted for the content of=
 capture scenes.

Regards,

Roberta

--_000_49E45C59CA48264997FEBFB29B6BC2D62151A1722ACRPMBOXPRD07p_
Content-Type: text/html; charset="us-ascii"
Content-Transfer-Encoding: quoted-printable

<html xmlns:v=3D"urn:schemas-microsoft-com:vml" xmlns:o=3D"urn:schemas-micr=
osoft-com:office:office" xmlns:w=3D"urn:schemas-microsoft-com:office:word" =
xmlns:m=3D"http://schemas.microsoft.com/office/2004/12/omml" xmlns=3D"http:=
//www.w3.org/TR/REC-html40"><head><META HTTP-EQUIV=3D"Content-Type" CONTENT=
=3D"text/html; charset=3Dus-ascii"><meta name=3DGenerator content=3D"Micros=
oft Word 14 (filtered medium)"><style><!--
/* Font Definitions */
@font-face
	{font-family:Calibri;
	panose-1:2 15 5 2 2 2 4 3 2 4;}
@font-face
	{font-family:Tahoma;
	panose-1:2 11 6 4 3 5 4 4 2 4;}
/* Style Definitions */
p.MsoNormal, li.MsoNormal, div.MsoNormal
	{margin:0in;
	margin-bottom:.0001pt;
	font-size:11.0pt;
	font-family:"Calibri","sans-serif";
	color:black;}
a:link, span.MsoHyperlink
	{mso-style-priority:99;
	color:blue;
	text-decoration:underline;}
a:visited, span.MsoHyperlinkFollowed
	{mso-style-priority:99;
	color:purple;
	text-decoration:underline;}
span.EmailStyle17
	{mso-style-type:personal;
	font-family:"Calibri","sans-serif";
	color:windowtext;}
span.EmailStyle18
	{mso-style-type:personal-reply;
	font-family:"Calibri","sans-serif";
	color:#1F497D;}
.MsoChpDefault
	{mso-style-type:export-only;
	font-size:10.0pt;}
@page WordSection1
	{size:8.5in 11.0in;
	margin:1.0in 1.0in 1.0in 1.0in;}
div.WordSection1
	{page:WordSection1;}
--></style><!--[if gte mso 9]><xml>
<o:shapedefaults v:ext=3D"edit" spidmax=3D"1026" />
</xml><![endif]--><!--[if gte mso 9]><xml>
<o:shapelayout v:ext=3D"edit">
<o:idmap v:ext=3D"edit" data=3D"1" />
</o:shapelayout></xml><![endif]--></head><body bgcolor=3Dwhite lang=3DEN-US=
 link=3Dblue vlink=3Dpurple><div class=3DWordSection1><p class=3DMsoNormal>=
<span style=3D'color:#1F497D'>Hi Roberta,<o:p></o:p></span></p><p class=3DM=
soNormal><span style=3D'color:#1F497D'><o:p>&nbsp;</o:p></span></p><p class=
=3DMsoNormal><span style=3D'color:#1F497D'>Thanks for the suggestion, it ma=
kes sense.&nbsp; I think the meaning is the same either way, but keeping it=
 consistent with the different media types in a CSE list sounds good.&nbsp;=
 I&#8217;ll change the text in the framework.<o:p></o:p></span></p><p class=
=3DMsoNormal><span style=3D'color:#1F497D'><o:p>&nbsp;</o:p></span></p><p c=
lass=3DMsoNormal><span style=3D'color:#1F497D'>Mark<o:p></o:p></span></p><p=
 class=3DMsoNormal><span style=3D'color:#1F497D'><o:p>&nbsp;</o:p></span></=
p><div style=3D'border:none;border-left:solid blue 1.5pt;padding:0in 0in 0i=
n 4.0pt'><div><div style=3D'border:none;border-top:solid #B5C4DF 1.0pt;padd=
ing:3.0pt 0in 0in 0in'><p class=3DMsoNormal><b><span style=3D'font-size:10.=
0pt;font-family:"Tahoma","sans-serif";color:windowtext'>From:</span></b><sp=
an style=3D'font-size:10.0pt;font-family:"Tahoma","sans-serif";color:window=
text'> Roberta Presta [mailto:roberta.presta@unina.it] <br><b>Sent:</b> Tue=
sday, April 01, 2014 6:54 AM<br><b>To:</b> Duckworth, Mark; clue@ietf.org<b=
r><b>Subject:</b> Re: [clue] Global CSE List in framework<o:p></o:p></span>=
</p></div></div><p class=3DMsoNormal><o:p>&nbsp;</o:p></p><div><p class=3DM=
soNormal>Hello Mark, <br><br>just a minor question about your definition:<b=
r><br><br>Il 27/03/2014 18:10, Duckworth, Mark ha scritto:<o:p></o:p></p></=
div><blockquote style=3D'margin-top:5.0pt;margin-bottom:5.0pt'><p class=3DM=
soNormal><span style=3D'font-size:12.0pt;font-family:"Times New Roman","ser=
if"'>&nbsp;An Advertisement can include an optional global Capture Scene En=
try list, for each media type.&nbsp; Each item in this list is a set of one=
 or more Capture Scene Entries.&nbsp; <o:p></o:p></span></p></blockquote><p=
 class=3DMsoNormal><span style=3D'font-size:12.0pt;font-family:"Times New R=
oman","serif"'><br>Do you mean that we have a global capture scene entry li=
st for video and a global capture scene entry list for audio?<br>Otherwise,=
 what about &quot;...global Capture Scene Entry list. Each item in the list=
 is a set of one or more Capture Scene Entries of the same media type.&quot=
;? That last approach is more similar to the one adopted for the content of=
 capture scenes. <br><br>Regards,<br><br>Roberta<o:p></o:p></span></p></div=
></div></body></html>=

--_000_49E45C59CA48264997FEBFB29B6BC2D62151A1722ACRPMBOXPRD07p_--


From nobody Thu Apr  3 08:16:12 2014
Return-Path: <pkyzivat@alum.mit.edu>
X-Original-To: clue@ietfa.amsl.com
Delivered-To: clue@ietfa.amsl.com
Received: from localhost (ietfa.amsl.com [127.0.0.1]) by ietfa.amsl.com (Postfix) with ESMTP id EE00E1A0207 for <clue@ietfa.amsl.com>; Thu,  3 Apr 2014 08:16:08 -0700 (PDT)
X-Virus-Scanned: amavisd-new at amsl.com
X-Spam-Flag: NO
X-Spam-Score: -1.235
X-Spam-Level: 
X-Spam-Status: No, score=-1.235 tagged_above=-999 required=5 tests=[BAYES_00=-1.9, DKIM_SIGNED=0.1, DKIM_VALID=-0.1, SPF_SOFTFAIL=0.665] autolearn=no
Received: from mail.ietf.org ([4.31.198.44]) by localhost (ietfa.amsl.com [127.0.0.1]) (amavisd-new, port 10024) with ESMTP id Ahf68gPlfHGy for <clue@ietfa.amsl.com>; Thu,  3 Apr 2014 08:16:05 -0700 (PDT)
Received: from qmta15.westchester.pa.mail.comcast.net (qmta15.westchester.pa.mail.comcast.net [IPv6:2001:558:fe14:44:76:96:59:228]) by ietfa.amsl.com (Postfix) with ESMTP id 9882E1A0214 for <clue@ietf.org>; Thu,  3 Apr 2014 08:16:04 -0700 (PDT)
Received: from omta16.westchester.pa.mail.comcast.net ([76.96.62.88]) by qmta15.westchester.pa.mail.comcast.net with comcast id lSPN1n0051uE5Es5FTG0EL; Thu, 03 Apr 2014 15:16:00 +0000
Received: from Paul-Kyzivats-MacBook-Pro.local ([IPv6:2002:328a:e5a4:0:f9e7:d4b2:db8d:afda]) by omta16.westchester.pa.mail.comcast.net with comcast id lTFy1n00C4wFaYb3cTFyJS; Thu, 03 Apr 2014 15:16:00 +0000
Message-ID: <533D7B2E.5020609@alum.mit.edu>
Date: Thu, 03 Apr 2014 11:15:58 -0400
From: Paul Kyzivat <pkyzivat@alum.mit.edu>
User-Agent: Mozilla/5.0 (Macintosh; Intel Mac OS X 10.7; rv:24.0) Gecko/20100101 Thunderbird/24.4.0
MIME-Version: 1.0
To: clue@ietf.org
References: <49E45C59CA48264997FEBFB29B6BC2D617AC5365DF@CRPMBOXPRD07.polycom.com> <20140326001139.GD99797@verdi> <53329AE2.4050302@unina.it> <49E45C59CA48264997FEBFB29B6BC2D617ACF1ACE8@CRPMBOXPRD07.polycom.com> <53330844.9040705@unina.it> <49E45C59CA48264997FEBFB29B6BC2D617ACF1ADBF@CRPMBOXPRD07.polycom.com> <20140326183424.GF99797@verdi> <53359141.9000404@unina.it> <20140328155515.GM99797@verdi> <49E45C59CA48264997FEBFB29B6BC2D617ACF1B5BB@CRPMBOXPRD07.polycom.com> <20140328174257.GN99797@verdi> <533A0DE1.2050806@nteczone.com>
In-Reply-To: <533A0DE1.2050806@nteczone.com>
Content-Type: text/plain; charset=ISO-8859-1; format=flowed
Content-Transfer-Encoding: 7bit
DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=comcast.net; s=q20140121; t=1396538160; bh=5FuBXimuFeMhyEhf1VJBuclrS89RKT98ifc+/O8rIrI=; h=Received:Received:Message-ID:Date:From:MIME-Version:To:Subject: Content-Type; b=h2w8wb/jaF43BUzkx7s4kKNy7LeOIhY7vX2Up7A9EHyiXOU6WMLrZiCtTqmM91QeP deViLV2omGraaF9U1Hp49gzdTULquKi5SwLGu4Qg2wMKZxobuO3/WLfUQwKnoA1AC4 +apZNKDUwV7fKQqLj0CcY22MFFisH8+fxQOi+OUctU3Nb13R9PZnSXzxuPvSGbCTu1 bfGCFi8+XFRtcvwBRcBCJw7UphwcZxjFiGKYXjCUjVnJCH1waAtcPktjejxmJywdCn 98MdF1LNjBR9L1cmF6DqYzdJw++q3veFBdPpAAV4EhSdBAKy9Fqwjz3WxftJc3hA3u e/M7Q3jAQVgkg==
Archived-At: http://mailarchive.ietf.org/arch/msg/clue/46AEW9R-_MwX9i_iSG4Am3fu4f8
Subject: Re: [clue] <capturePoint> should be optional
X-BeenThere: clue@ietf.org
X-Mailman-Version: 2.1.15
Precedence: list
List-Id: CLUE - ControLling mUltiple streams for TElepresence <clue.ietf.org>
List-Unsubscribe: <https://www.ietf.org/mailman/options/clue>, <mailto:clue-request@ietf.org?subject=unsubscribe>
List-Archive: <http://www.ietf.org/mail-archive/web/clue/>
List-Post: <mailto:clue@ietf.org>
List-Help: <mailto:clue-request@ietf.org?subject=help>
List-Subscribe: <https://www.ietf.org/mailman/listinfo/clue>, <mailto:clue-request@ietf.org?subject=subscribe>
X-List-Received-Date: Thu, 03 Apr 2014 15:16:09 -0000

On 3/31/14 8:52 PM, Christian Groves wrote:
> Hello
>
> In some respects I agree with John. The discussions around the use of
> spatial information for audio captures just seemed to die out without
> any real conclusion. As John has alluded to the use of the spatial
> attributes for audio is problematic. We don't seem to have any text in
> the framework discussing this or forbidding the use of the spatial
> information for audio.

Certainly there is nothing forbidding the use of the existing spatial 
information for audio. But neither is there any guidance in how it 
could/should be used. John has repeatedly made the point that the 
existing info doesn't work for audio. But we have never actually *done* 
anything about that.

I have been of the assumption that people had in mind some way they 
would use the existing spatial info for audio. But without some clarity 
about how that is to be done I suspect that there could be major 
interoperability problems.

> With regards to the video related spatial information I didn't realize
> that there was deep seated confusion over the use of the attributes.

The lack of apparent confusion doesn't mean there is consistent 
understanding. :-)

>  From a previous email:
>
> where the "to be rendered" information are represented by means of the
> <capturedArea> within the <spatialInformation> tag associated to the MCCs.
>
>     <captureArea> has always bothered me, because it is a two-dimensional
> quantity, while we interact in three-dimensional space. I suppose I should
> have objected to including it at all in <spatialInformation>; but until
> now I thought we understood it to be defining space in a plane within
> our three-dimensional interaction space (and therefore a three-dimensional
> space from <capturePoint> to the <captureArea>).
>
> I thought of the <captureArea> being a quadrilateral on a plane in 3D
> space, i.e. a 2D area. The <capturePoint> and <lineOfCapturePoint> were
> points in this 3D space giving the relationship between the capture
> device and the <captureArea>. This was in order to allow the Renderer to
> perform what ever geometric corrections it needed to do in order to
> provide an optimal rendering of the image on a screen i.e. quadrilateral
> on a plane.

*My* understanding is that the capture area and the point of capture 
define a "capture volume" as a sort of pyramid, with the capture area 
being in focus.

In the case of captures by cameras, the content captured may occur 
anywhere in that volume. It clearly doesn't all come from the 2d capture 
area. But in most cases we would imagine that it will be displayed in a 
comparable 2d area.

But clearly this model doesn't work well for audio.

ISTM that, as for video, the goal here must be to provide the 
information needed for the consumer to:
- decide which audio captures to configure;
- decide how to render those on the equipment available to the receiver.

> In the case where the image has been through a device that performs the
> geometric corrections etc. (e.g. an MCU) the resultant image still
> represents an quadrilateral on a plane (albeit with Y=0). In terms of
> the <capturePoint> and <lineOfCapturePoint> there would be "virtual"
> values for these as a result of any transformations applied to the
> image. I guess the assumption is that is would be meaningless to send
> these as no further transformation requiring these points would be
> required.

The extreme here is for synthetic captures, such as presentations. They 
may truly not have a point of capture, and the volume of capture is 
truly just the 2d area of capture.

	Thanks,
	Paul

> Regards, Christian
>
>
> On 29/03/2014 4:42 AM, John Leslie wrote:
>> Duckworth, Mark <Mark.Duckworth@polycom.com> wrote:
>>> The group spent a lot of time discussing and finalizing the spatial
>>> information as currently defined.
>>     Indeed!
>>
>>> It has been very stable for a long time now.
>>     Umm, Seems more "dormant" to me...
>>
>>> I'd rather not re-open that discussion now, unless there is strong
>>> consensus in the group to do so.
>>     Hmm...
>>
>>> My view is that both types of usages described below are real spatial
>>> information, but only one is directly derived from physical space.
>>     An interesting view...
>>
>>     ISTM that one describes physical space (three-dimensional) and the
>> other describes a mathematical concept, instantiated on a visual
>> display. Mixing them is inherently confusing.
>>
>>     (Arguments that the mathematical concept is a "space" do not move
>> me -- everything in math is sometimes described as a "space".)
>>
>>     I seriously fear that we've only scratched the surface of the
>> confusion we've uncovered.
>>
>>     I do understand that bikeshedding can prevent progress until it's
>> too late to consider the important issues. But keeping physical space
>> distinct from mathematical "space" strikes me as an important issue.
>>
>>> They are distinguished by the "scale" attribute of <captureScene>.
>>     I doubt this. I remember only "millimeters", "unknown", and
>> "noscale".
>> I can imagine folks using "noscale" for three-dimensional space.
>>
>>     Also, I find "scale" a rather strange way to disambiguate physical
>> from mathematical space...
>>
>> --
>> John Leslie <john@jlc.net>
>>
>> _______________________________________________
>> clue mailing list
>> clue@ietf.org
>> https://www.ietf.org/mailman/listinfo/clue
>>
>
> _______________________________________________
> clue mailing list
> clue@ietf.org
> https://www.ietf.org/mailman/listinfo/clue
>


From nobody Thu Apr  3 08:31:14 2014
Return-Path: <pkyzivat@alum.mit.edu>
X-Original-To: clue@ietfa.amsl.com
Delivered-To: clue@ietfa.amsl.com
Received: from localhost (ietfa.amsl.com [127.0.0.1]) by ietfa.amsl.com (Postfix) with ESMTP id 500231A021B for <clue@ietfa.amsl.com>; Thu,  3 Apr 2014 08:31:07 -0700 (PDT)
X-Virus-Scanned: amavisd-new at amsl.com
X-Spam-Flag: NO
X-Spam-Score: -1.235
X-Spam-Level: 
X-Spam-Status: No, score=-1.235 tagged_above=-999 required=5 tests=[BAYES_00=-1.9, DKIM_SIGNED=0.1, DKIM_VALID=-0.1, SPF_SOFTFAIL=0.665] autolearn=no
Received: from mail.ietf.org ([4.31.198.44]) by localhost (ietfa.amsl.com [127.0.0.1]) (amavisd-new, port 10024) with ESMTP id X4559_x0yBnp for <clue@ietfa.amsl.com>; Thu,  3 Apr 2014 08:31:03 -0700 (PDT)
Received: from qmta12.westchester.pa.mail.comcast.net (qmta12.westchester.pa.mail.comcast.net [IPv6:2001:558:fe14:44:76:96:59:227]) by ietfa.amsl.com (Postfix) with ESMTP id 0D77F1A0207 for <clue@ietf.org>; Thu,  3 Apr 2014 08:31:02 -0700 (PDT)
Received: from omta03.westchester.pa.mail.comcast.net ([76.96.62.27]) by qmta12.westchester.pa.mail.comcast.net with comcast id lPvl1n0040bG4ec5CTWymJ; Thu, 03 Apr 2014 15:30:58 +0000
Received: from Paul-Kyzivats-MacBook-Pro.local ([IPv6:2002:328a:e5a4:0:f9e7:d4b2:db8d:afda]) by omta03.westchester.pa.mail.comcast.net with comcast id lTWS1n00Y4wFaYb3PTWS3p; Thu, 03 Apr 2014 15:30:58 +0000
Message-ID: <533D7E8F.4000303@alum.mit.edu>
Date: Thu, 03 Apr 2014 11:30:23 -0400
From: Paul Kyzivat <pkyzivat@alum.mit.edu>
User-Agent: Mozilla/5.0 (Macintosh; Intel Mac OS X 10.7; rv:24.0) Gecko/20100101 Thunderbird/24.4.0
MIME-Version: 1.0
To: clue@ietf.org
References: <5318809E.2030204@nteczone.com> <5318AE5E.4050404@alum.mit.edu> <5318B0C9.1050603@nteczone.com> <5318B48E.3090300@alum.mit.edu> <53290C9B.6090106@nteczone.com> <5329F3FF.7090606@alum.mit.edu> <533A16BA.40707@nteczone.com> <533A91BF.7040404@unina.it> <533B47FF.5020407@nteczone.com>
In-Reply-To: <533B47FF.5020407@nteczone.com>
Content-Type: text/plain; charset=ISO-8859-1; format=flowed
Content-Transfer-Encoding: 7bit
DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=comcast.net; s=q20140121; t=1396539058; bh=aafB0/LAEWGd80mWtpfZobgKFEqWQ3dWtZoPIxJdTF0=; h=Received:Received:Message-ID:Date:From:MIME-Version:To:Subject: Content-Type; b=RjxSKLLOoEnNwjnKGvjoZBqgJ0BsLPp27kuPm33wznb6inwGKsmWn9XTh+2w2OQXd 1FyUoLTZzjRUBVHL31RwZMrgakt2jhYkdJQ+LvOjYAFyj6anquFyfNR+KOylRaEZOD eYuVM/6G4gvSUhRXwNL43sE2o9//mhZ0JcT6N2vRH1g/ZV33sWbhykTZBl6xrm2Zz2 sW/TcgcuUDxsRSbI5bej1wc5ynvh6WBj8GttVII3HzGtrYRqHzhew9zo1GOmDHAaUK A031aEQ3h93rLPBHlFPeNsa6Y9pt7A9ZjKUMwHn6GKoExHbIiNM6V/YZPhjdQEqFJx 2iIim1FCJBaIg==
Archived-At: http://mailarchive.ietf.org/arch/msg/clue/-ySYie6zVFDyehyArLaQKzribRo
Subject: Re: [clue] Participant info/type followup
X-BeenThere: clue@ietf.org
X-Mailman-Version: 2.1.15
Precedence: list
List-Id: CLUE - ControLling mUltiple streams for TElepresence <clue.ietf.org>
List-Unsubscribe: <https://www.ietf.org/mailman/options/clue>, <mailto:clue-request@ietf.org?subject=unsubscribe>
List-Archive: <http://www.ietf.org/mail-archive/web/clue/>
List-Post: <mailto:clue@ietf.org>
List-Help: <mailto:clue-request@ietf.org?subject=help>
List-Subscribe: <https://www.ietf.org/mailman/listinfo/clue>, <mailto:clue-request@ietf.org?subject=subscribe>
X-List-Received-Date: Thu, 03 Apr 2014 15:31:07 -0000

ISTM we have both a semantic issue and a syntactic/terminological issue:

- we have the semantic distinction between who is represented within
   the capture and who is responsible for sending the capture

- we have a terminology issue regarding what to call these two things

I think we now agree that the semantic distinction is real and should be 
identified in the advertisement.

Regarding terminology, IMO we should use two different terms to identify 
these. At most one of them should be called "participant". Given the 
conflict with xcon, perhaps *neither* of them should be called participant.

This is mostly a data model issue. Some of it might need to peek through 
into the framework. And what does should be consistent.

	Thanks,
	Paul

On 4/1/14 7:13 PM, Christian Groves wrote:
> Hello Roberta,
>
>  From a data model perspective I agree that it makes sense to have the
> participant metadata "referenced" to minimize duplication in the
> messages. So I support your solution in the data model to the redundancy
> problem.
>
> In general I think its good to maintain consistency between the
> framework and the data model. However in London I thought this was more
> of a syntax shortcut (like the captureID wildcarding) rather than
> something we'd need to formalize in the framework. I guess we could add
> it at a higher level in the framework (it would functionally be the same
> thing), it would just be more work for the editor...
>
> I could go either way on this.
>
> Regards, Christian
>
> On 1/04/2014 9:15 PM, Roberta Presta wrote:
>> Hi Christian,
>>
>> In section 7.1.1.X we are dealing with media capture attributes and in
>> section 7.1.1.11 we want to define an attribute conveying information
>> about participants.
>> I would say that here we can find both information (*references*, in
>> the data model) about "who is represented in the capture" ("captured
>> participants"?) and "who is the owner of the generating device"
>> ("owner"?).
>>
>> Maybe participant metadata, such as Participant Information and
>> Participant Type as they are currently defined, should be treated in a
>> separate section of the same level of capture scenes or media captures
>> as.
>> Indeed we showed in London that repeating the vcard and the role of
>> the participants in each capture, as if they were capture attributes,
>> causes redundancy.
>>
>> I know that I have a data model definition perspective, but I would
>> propose to make a change that is more coherent with what we will
>> describe formally.
>>
>> Cheers,
>>
>> Roberta
>>
>>
>>
>>
>>
>> Il 01/04/2014 03:30, Christian Groves ha scritto:
>>> Hello Paul, all,
>>>
>>> If we follow the approach that there is a specific indicating of
>>> whether the participant information is based on an explicit
>>> indication then I would suggest the following text for the framework:
>>>
>>> Clause 7.1.11 Participant information
>>> (Under the 1st paragraph)
>>>
>>> The participant information contains an explicit indication of
>>> whether it relates to a participant contained in the capture, from a
>>> participants capture device or both. For example a video camera may
>>> capture an image containing the participant, or a participant may
>>> send a video capture with a presentation that does not depict the
>>> participant.
>>>
>>> Something similar would be needed under participant type.
>>>
>>> Thoughts?
>>>
>>> Regards, Christian
>>>
>>> On 20/03/2014 6:46 AM, Paul Kyzivat wrote:
>>>> On 3/18/14 11:18 PM, Christian Groves wrote:
>>>>> Hello Paul,
>>>>>
>>>>> "How" they differ is given by the example bullets below the
>>>>> sentence. If
>>>>> you want something more normative we could remove the "For example".
>>>>
>>>> Yeah, I don't believe in specification by example. :-)
>>>>
>>>> IMO it is a bit dicey to base this distinction on the type of capture.
>>>>
>>>> I'm more comfortable with an explicit syntactic indication of the
>>>> distinction, such as proposed by Roberta.
>>>>
>>>>     Thanks,
>>>>     Paul
>>>>
>>>>> Regards, Christian
>>>>>
>>>>> On 7/03/2014 4:46 AM, Paul Kyzivat wrote:
>>>>>> On 3/6/14 5:30 PM, Christian Groves wrote:
>>>>>>> Hello Paul,
>>>>>>>
>>>>>>> The text says media type and presentation attribute. Is that the
>>>>>>> relationship you're talking about?
>>>>>>
>>>>>> "How the generated content relates to the entity described in the
>>>>>> participant info is dependent on media type and and the presentation
>>>>>> attribute."
>>>>>>
>>>>>> I take that to mean that the relationship may be different for
>>>>>> presentation streams than non-presentation streams. But it doesn't
>>>>>> say
>>>>>> *how* they differ.
>>>>>>
>>>>>>     Thanks,
>>>>>>     Paul
>>>>>>
>>>>>>> Regards, Christian
>>>>>>>
>>>>>>> On 7/03/2014 4:20 AM, Paul Kyzivat wrote:
>>>>>>>> Christian,
>>>>>>>>
>>>>>>>> I've read the quoted text several times, and I cannot figure out
>>>>>>>> how
>>>>>>>> to *derive* your example conclusions from it. The text says the
>>>>>>>> relationship is dependent on the presentation attribute, but not
>>>>>>>> how.
>>>>>>>>
>>>>>>>> AFAICT I could make a new definition where the a participant
>>>>>>>> attached
>>>>>>>> to a presentation capture means that the participant is shown in
>>>>>>>> the
>>>>>>>> presentation, and that would be equally compatible with the text.
>>>>>>>>
>>>>>>>> ISTM that more text is required to actually specify the
>>>>>>>> relationships.
>>>>>>>>
>>>>>>>>     Thanks,
>>>>>>>>     Paul
>>>>>>>>
>>>>>>>> On 3/6/14 2:05 PM, Christian Groves wrote:
>>>>>>>>> Hello all,
>>>>>>>>>
>>>>>>>>> To follow up on Jonathon's comments on participant info/type
>>>>>>>>> and the
>>>>>>>>> semantics and particularly how it relates to a presentation.
>>>>>>>>> Here's a
>>>>>>>>> first stab at some text to stimulate some discussions.
>>>>>>>>>
>>>>>>>>>
>>>>>>>>> "The participant info attribute allows a provider to associate
>>>>>>>>> participant information with the capture source. When used in an
>>>>>>>>> individual capture it indicates that the captured content (e.g.
>>>>>>>>> video/audio/text etc.) as opposed to the actual media streams is
>>>>>>>>> generated from the entity described. How the generated content
>>>>>>>>> relates
>>>>>>>>> to the entity described in the participant info is dependent on
>>>>>>>>> media
>>>>>>>>> type and and the presentation attribute.
>>>>>>>>>
>>>>>>>>> For example:
>>>>>>>>> - a video capture with participant info would indicate that the
>>>>>>>>> video
>>>>>>>>> contains a picture of the entity associated with the information
>>>>>>>>> provided.
>>>>>>>>> - a video capture with participant info and the presentation
>>>>>>>>> attribute
>>>>>>>>> would indicate that the presentation video is associated with the
>>>>>>>>> participant but could contain any video content.
>>>>>>>>> - a text capture with participant info would indicate that the
>>>>>>>>> text is
>>>>>>>>> generated from the actual participant.
>>>>>>>>> - a text capture with participant info and the presentation
>>>>>>>>> attribute
>>>>>>>>> would indicate that the text is associated with the participant
>>>>>>>>> but
>>>>>>>>> could contain any text content."
>>>>>>>>>
>>>>>>>>> Comments?
>>>>>>>>>
>>>>>>>>>
>>>>>>>>> Regards, Christian
>>>>>>>>>
>>>>>>>>> _______________________________________________
>>>>>>>>> clue mailing list
>>>>>>>>> clue@ietf.org
>>>>>>>>> https://www.ietf.org/mailman/listinfo/clue
>>>>>>>>>
>>>>>>>>
>>>>>>>> _______________________________________________
>>>>>>>> clue mailing list
>>>>>>>> clue@ietf.org
>>>>>>>> https://www.ietf.org/mailman/listinfo/clue
>>>>>>>>
>>>>>>>
>>>>>>> _______________________________________________
>>>>>>> clue mailing list
>>>>>>> clue@ietf.org
>>>>>>> https://www.ietf.org/mailman/listinfo/clue
>>>>>>>
>>>>>>
>>>>>>
>>>>>
>>>>>
>>>>
>>>>
>>>
>>> _______________________________________________
>>> clue mailing list
>>> clue@ietf.org
>>> https://www.ietf.org/mailman/listinfo/clue
>>>
>>
>>
>
> _______________________________________________
> clue mailing list
> clue@ietf.org
> https://www.ietf.org/mailman/listinfo/clue
>


From nobody Thu Apr  3 16:27:02 2014
Return-Path: <Christian.Groves@nteczone.com>
X-Original-To: clue@ietfa.amsl.com
Delivered-To: clue@ietfa.amsl.com
Received: from localhost (ietfa.amsl.com [127.0.0.1]) by ietfa.amsl.com (Postfix) with ESMTP id 738F51A02DF for <clue@ietfa.amsl.com>; Thu,  3 Apr 2014 16:27:01 -0700 (PDT)
X-Virus-Scanned: amavisd-new at amsl.com
X-Spam-Flag: NO
X-Spam-Score: 0.8
X-Spam-Level: 
X-Spam-Status: No, score=0.8 tagged_above=-999 required=5 tests=[BAYES_50=0.8] autolearn=ham
Received: from mail.ietf.org ([4.31.198.44]) by localhost (ietfa.amsl.com [127.0.0.1]) (amavisd-new, port 10024) with ESMTP id l2smyBV_XTpi for <clue@ietfa.amsl.com>; Thu,  3 Apr 2014 16:26:56 -0700 (PDT)
Received: from cserver5.myshophosting.com (cserver5.myshophosting.com [175.107.161.1]) by ietfa.amsl.com (Postfix) with ESMTP id 8AE141A02A6 for <clue@ietf.org>; Thu,  3 Apr 2014 16:26:56 -0700 (PDT)
Received: from ppp118-209-196-251.lns20.mel6.internode.on.net ([118.209.196.251]:52682 helo=[127.0.0.1]) by cserver5.myshophosting.com with esmtpsa (TLSv1:DHE-RSA-AES128-SHA:128) (Exim 4.82) (envelope-from <Christian.Groves@nteczone.com>) id 1WVr1z-0005aq-Cg; Fri, 04 Apr 2014 10:26:47 +1100
Message-ID: <533DEE34.7070809@nteczone.com>
Date: Fri, 04 Apr 2014 10:26:44 +1100
From: Christian Groves <Christian.Groves@nteczone.com>
User-Agent: Mozilla/5.0 (Windows NT 6.1; WOW64; rv:24.0) Gecko/20100101 Thunderbird/24.4.0
MIME-Version: 1.0
To: "Duckworth, Mark" <Mark.Duckworth@polycom.com>,  "clue@ietf.org" <clue@ietf.org>
References: <49E45C59CA48264997FEBFB29B6BC2D617ACF1B213@CRPMBOXPRD07.polycom.com> <533A14A0.6030807@nteczone.com> <49E45C59CA48264997FEBFB29B6BC2D62151A17223@CRPMBOXPRD07.polycom.com>
In-Reply-To: <49E45C59CA48264997FEBFB29B6BC2D62151A17223@CRPMBOXPRD07.polycom.com>
Content-Type: text/plain; charset=ISO-8859-1; format=flowed
Content-Transfer-Encoding: 7bit
X-AntiAbuse: This header was added to track abuse, please include it with any abuse report
X-AntiAbuse: Primary Hostname - cserver5.myshophosting.com
X-AntiAbuse: Original Domain - ietf.org
X-AntiAbuse: Originator/Caller UID/GID - [47 12] / [47 12]
X-AntiAbuse: Sender Address Domain - nteczone.com
X-Get-Message-Sender-Via: cserver5.myshophosting.com: authenticated_id: christian.groves@nteczone.com
X-Source: 
X-Source-Args: 
X-Source-Dir: 
Archived-At: http://mailarchive.ietf.org/arch/msg/clue/0jQb17JLmo6q7TRKx0AryZvhdwU
Subject: Re: [clue] Global CSE List in framework
X-BeenThere: clue@ietf.org
X-Mailman-Version: 2.1.15
Precedence: list
List-Id: CLUE - ControLling mUltiple streams for TElepresence <clue.ietf.org>
List-Unsubscribe: <https://www.ietf.org/mailman/options/clue>, <mailto:clue-request@ietf.org?subject=unsubscribe>
List-Archive: <http://www.ietf.org/mail-archive/web/clue/>
List-Post: <mailto:clue@ietf.org>
List-Help: <mailto:clue-request@ietf.org?subject=help>
List-Subscribe: <https://www.ietf.org/mailman/listinfo/clue>, <mailto:clue-request@ietf.org?subject=subscribe>
X-List-Received-Date: Thu, 03 Apr 2014 23:27:01 -0000

Hello Mark,

Please see below.

Regards, Christian

..snip..
>>> The Provider can include multiple sets, to accommodate different
>>> Consumers with varying capability to receive multiple encodings. This
>>> is very similar to how each CSE represents a particular scene.
>>>
>> [CNG] Is the GCSE about providing sets for multiple "encodings"? You can
>> provide multiple encodings without GSEs. Isn't it about providing different
>> sets of captures so that a receiver can best determine what the set of
>> captures it wants based on the service it wants to provide.
>> I.e. a consumer may only support three media streams (3 encodings) but
>> these may be used in a number of ways.
> [Duckworth, Mark] Are you suggesting just replace "encodings" with "captures"?  That's fine with me, maybe "captures" better expresses the meaning.
[CNG] Whilst "captures" is probably better than "encodings" I'm not sure 
that "receive multiple captures" is the point. How about "The Provider 
can include multiple sets, to allow a consumer to choose sets of 
captures appropriate to its capabilities or application..."

> ...snip ...
>>> * The ordering of items (sets of CSEs) in the global CSE list is not
>>> important.
>>>
>>> * The ordering of CSEs within each set is not important.
>>>
>>> * The Provider must be capable of encoding and sending all Captures
>>> within the CSEs of a given set simultaneously.
>>>
>> [CNG] Do we need to add some text to the simultaneous set section of the
>> framework indicating this? We should probably say something about the
>> interaction. i.e. If a GCSE set is included should the provider actually need to
>> provide a STS? Sendings CSEs in a STS indicates the those CSE/captures may
>> be used simultaneously. There appears to be alot of overlap with the GCSE in
>> this respect.
> [Duckworth, Mark] I think I see what you mean.  We have this paragraph already:
> "If an Advertisement does not include Simultaneous Transmission Sets, then the Provider MUST be able to provide all Capture Scenes simultaneously.  If multiple capture Scene Entries are in a Capture Scene then the Consumer chooses at most one Capture Scene Entry per Capture Scene for each media type."
>
> [Duckworth, Mark] I suggest we also add "If there is no STS and there is a global CSE list, then the Consumer chooses at most one set of CSEs of each media type, from the global CSE list."
[CNG] I'm OK with the above sentence. I think in addition we need to add 
some text regarding what happens when an STS AND GCSE is included. 
There's some overlap between them.

E.g.
Scene1(CSE1(VC1,VC2,VC3)
Scene2(CSE2(VC4,VC5,VC6)
Scene3(CSE3(VC7,VC8,VC9)
GCSE([CSE1,CSE3],[CSE2,CSE3])
STS([CSE1,CSE3],[CSE2,CSE3])

If the provider includes a GCSE as above and wants to include a STS does 
it have to contain the exact same CSE sets?, i.e. because according to 
the above bullet the CSE must capable of being send simultaneously. Or 
is the interaction that a GCSE set indicates that at least one capture 
from each CSE can be sent simultaneously, or?




From nobody Thu Apr  3 16:56:01 2014
Return-Path: <Christian.Groves@nteczone.com>
X-Original-To: clue@ietfa.amsl.com
Delivered-To: clue@ietfa.amsl.com
Received: from localhost (ietfa.amsl.com [127.0.0.1]) by ietfa.amsl.com (Postfix) with ESMTP id 3DD0B1A03CB for <clue@ietfa.amsl.com>; Thu,  3 Apr 2014 16:56:00 -0700 (PDT)
X-Virus-Scanned: amavisd-new at amsl.com
X-Spam-Flag: NO
X-Spam-Score: -1.9
X-Spam-Level: 
X-Spam-Status: No, score=-1.9 tagged_above=-999 required=5 tests=[BAYES_00=-1.9] autolearn=ham
Received: from mail.ietf.org ([4.31.198.44]) by localhost (ietfa.amsl.com [127.0.0.1]) (amavisd-new, port 10024) with ESMTP id oma0xIw3y1wo for <clue@ietfa.amsl.com>; Thu,  3 Apr 2014 16:55:55 -0700 (PDT)
Received: from cserver5.myshophosting.com (cserver5.myshophosting.com [175.107.161.1]) by ietfa.amsl.com (Postfix) with ESMTP id 54CB91A03C6 for <clue@ietf.org>; Thu,  3 Apr 2014 16:55:55 -0700 (PDT)
Received: from ppp118-209-196-251.lns20.mel6.internode.on.net ([118.209.196.251]:53047 helo=[127.0.0.1]) by cserver5.myshophosting.com with esmtpsa (TLSv1:DHE-RSA-AES128-SHA:128) (Exim 4.82) (envelope-from <Christian.Groves@nteczone.com>) id 1WVrU2-0000rh-QX for clue@ietf.org; Fri, 04 Apr 2014 10:55:47 +1100
Message-ID: <533DF500.1080403@nteczone.com>
Date: Fri, 04 Apr 2014 10:55:44 +1100
From: Christian Groves <Christian.Groves@nteczone.com>
User-Agent: Mozilla/5.0 (Windows NT 6.1; WOW64; rv:24.0) Gecko/20100101 Thunderbird/24.4.0
MIME-Version: 1.0
To: clue@ietf.org
References: <5318809E.2030204@nteczone.com> <5318AE5E.4050404@alum.mit.edu> <5318B0C9.1050603@nteczone.com> <5318B48E.3090300@alum.mit.edu> <53290C9B.6090106@nteczone.com> <5329F3FF.7090606@alum.mit.edu> <533A16BA.40707@nteczone.com> <533A91BF.7040404@unina.it> <533B47FF.5020407@nteczone.com> <533D7E8F.4000303@alum.mit.edu>
In-Reply-To: <533D7E8F.4000303@alum.mit.edu>
Content-Type: text/plain; charset=ISO-8859-1; format=flowed
Content-Transfer-Encoding: 7bit
X-AntiAbuse: This header was added to track abuse, please include it with any abuse report
X-AntiAbuse: Primary Hostname - cserver5.myshophosting.com
X-AntiAbuse: Original Domain - ietf.org
X-AntiAbuse: Originator/Caller UID/GID - [47 12] / [47 12]
X-AntiAbuse: Sender Address Domain - nteczone.com
X-Get-Message-Sender-Via: cserver5.myshophosting.com: authenticated_id: christian.groves@nteczone.com
X-Source: 
X-Source-Args: 
X-Source-Dir: 
Archived-At: http://mailarchive.ietf.org/arch/msg/clue/cr7hkqlDUVU8Tt21i3F7Cv0jB3U
Subject: Re: [clue] Participant info/type followup
X-BeenThere: clue@ietf.org
X-Mailman-Version: 2.1.15
Precedence: list
List-Id: CLUE - ControLling mUltiple streams for TElepresence <clue.ietf.org>
List-Unsubscribe: <https://www.ietf.org/mailman/options/clue>, <mailto:clue-request@ietf.org?subject=unsubscribe>
List-Archive: <http://www.ietf.org/mail-archive/web/clue/>
List-Post: <mailto:clue@ietf.org>
List-Help: <mailto:clue-request@ietf.org?subject=help>
List-Subscribe: <https://www.ietf.org/mailman/listinfo/clue>, <mailto:clue-request@ietf.org?subject=subscribe>
X-List-Received-Date: Thu, 03 Apr 2014 23:56:00 -0000

Hello Paul,

I think its a framework issue as much as a data model issue. The data 
model issue is really just the syntactic issue of minimizing the 
repeating of the participant information through a reference.

In the framework I think there are two options:
1. A participant information attribute where the applicability of this 
attribute (e.g. a participant is depicted by the capture, and/or a 
capture is from a participant) is indicated.

2. Two separate attributes: "Captured Participant/s" and "Capture 
originating from participant"

I think either will work. From a data model perspective where a 
reference will be used and given the XML structure probably the 2nd 
option would allow most alignment between the two.

Regards, Christian

On 4/04/2014 2:30 AM, Paul Kyzivat wrote:
> ISTM we have both a semantic issue and a syntactic/terminological issue:
>
> - we have the semantic distinction between who is represented within
>   the capture and who is responsible for sending the capture
>
> - we have a terminology issue regarding what to call these two things
>
> I think we now agree that the semantic distinction is real and should 
> be identified in the advertisement.
>
> Regarding terminology, IMO we should use two different terms to 
> identify these. At most one of them should be called "participant". 
> Given the conflict with xcon, perhaps *neither* of them should be 
> called participant.
>
> This is mostly a data model issue. Some of it might need to peek 
> through into the framework. And what does should be consistent.
>
>     Thanks,
>     Paul
>
> On 4/1/14 7:13 PM, Christian Groves wrote:
>> Hello Roberta,
>>
>>  From a data model perspective I agree that it makes sense to have the
>> participant metadata "referenced" to minimize duplication in the
>> messages. So I support your solution in the data model to the redundancy
>> problem.
>>
>> In general I think its good to maintain consistency between the
>> framework and the data model. However in London I thought this was more
>> of a syntax shortcut (like the captureID wildcarding) rather than
>> something we'd need to formalize in the framework. I guess we could add
>> it at a higher level in the framework (it would functionally be the same
>> thing), it would just be more work for the editor...
>>
>> I could go either way on this.
>>
>> Regards, Christian
>>
>> On 1/04/2014 9:15 PM, Roberta Presta wrote:
>>> Hi Christian,
>>>
>>> In section 7.1.1.X we are dealing with media capture attributes and in
>>> section 7.1.1.11 we want to define an attribute conveying information
>>> about participants.
>>> I would say that here we can find both information (*references*, in
>>> the data model) about "who is represented in the capture" ("captured
>>> participants"?) and "who is the owner of the generating device"
>>> ("owner"?).
>>>
>>> Maybe participant metadata, such as Participant Information and
>>> Participant Type as they are currently defined, should be treated in a
>>> separate section of the same level of capture scenes or media captures
>>> as.
>>> Indeed we showed in London that repeating the vcard and the role of
>>> the participants in each capture, as if they were capture attributes,
>>> causes redundancy.
>>>
>>> I know that I have a data model definition perspective, but I would
>>> propose to make a change that is more coherent with what we will
>>> describe formally.
>>>
>>> Cheers,
>>>
>>> Roberta
>>>
>>>
>>>
>>>
>>>
>>> Il 01/04/2014 03:30, Christian Groves ha scritto:
>>>> Hello Paul, all,
>>>>
>>>> If we follow the approach that there is a specific indicating of
>>>> whether the participant information is based on an explicit
>>>> indication then I would suggest the following text for the framework:
>>>>
>>>> Clause 7.1.11 Participant information
>>>> (Under the 1st paragraph)
>>>>
>>>> The participant information contains an explicit indication of
>>>> whether it relates to a participant contained in the capture, from a
>>>> participants capture device or both. For example a video camera may
>>>> capture an image containing the participant, or a participant may
>>>> send a video capture with a presentation that does not depict the
>>>> participant.
>>>>
>>>> Something similar would be needed under participant type.
>>>>
>>>> Thoughts?
>>>>
>>>> Regards, Christian
>>>>
>>>> On 20/03/2014 6:46 AM, Paul Kyzivat wrote:
>>>>> On 3/18/14 11:18 PM, Christian Groves wrote:
>>>>>> Hello Paul,
>>>>>>
>>>>>> "How" they differ is given by the example bullets below the
>>>>>> sentence. If
>>>>>> you want something more normative we could remove the "For example".
>>>>>
>>>>> Yeah, I don't believe in specification by example. :-)
>>>>>
>>>>> IMO it is a bit dicey to base this distinction on the type of 
>>>>> capture.
>>>>>
>>>>> I'm more comfortable with an explicit syntactic indication of the
>>>>> distinction, such as proposed by Roberta.
>>>>>
>>>>>     Thanks,
>>>>>     Paul
>>>>>
>>>>>> Regards, Christian
>>>>>>
>>>>>> On 7/03/2014 4:46 AM, Paul Kyzivat wrote:
>>>>>>> On 3/6/14 5:30 PM, Christian Groves wrote:
>>>>>>>> Hello Paul,
>>>>>>>>
>>>>>>>> The text says media type and presentation attribute. Is that the
>>>>>>>> relationship you're talking about?
>>>>>>>
>>>>>>> "How the generated content relates to the entity described in the
>>>>>>> participant info is dependent on media type and and the 
>>>>>>> presentation
>>>>>>> attribute."
>>>>>>>
>>>>>>> I take that to mean that the relationship may be different for
>>>>>>> presentation streams than non-presentation streams. But it doesn't
>>>>>>> say
>>>>>>> *how* they differ.
>>>>>>>
>>>>>>>     Thanks,
>>>>>>>     Paul
>>>>>>>
>>>>>>>> Regards, Christian
>>>>>>>>
>>>>>>>> On 7/03/2014 4:20 AM, Paul Kyzivat wrote:
>>>>>>>>> Christian,
>>>>>>>>>
>>>>>>>>> I've read the quoted text several times, and I cannot figure out
>>>>>>>>> how
>>>>>>>>> to *derive* your example conclusions from it. The text says the
>>>>>>>>> relationship is dependent on the presentation attribute, but not
>>>>>>>>> how.
>>>>>>>>>
>>>>>>>>> AFAICT I could make a new definition where the a participant
>>>>>>>>> attached
>>>>>>>>> to a presentation capture means that the participant is shown in
>>>>>>>>> the
>>>>>>>>> presentation, and that would be equally compatible with the text.
>>>>>>>>>
>>>>>>>>> ISTM that more text is required to actually specify the
>>>>>>>>> relationships.
>>>>>>>>>
>>>>>>>>>     Thanks,
>>>>>>>>>     Paul
>>>>>>>>>
>>>>>>>>> On 3/6/14 2:05 PM, Christian Groves wrote:
>>>>>>>>>> Hello all,
>>>>>>>>>>
>>>>>>>>>> To follow up on Jonathon's comments on participant info/type
>>>>>>>>>> and the
>>>>>>>>>> semantics and particularly how it relates to a presentation.
>>>>>>>>>> Here's a
>>>>>>>>>> first stab at some text to stimulate some discussions.
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>> "The participant info attribute allows a provider to associate
>>>>>>>>>> participant information with the capture source. When used in an
>>>>>>>>>> individual capture it indicates that the captured content (e.g.
>>>>>>>>>> video/audio/text etc.) as opposed to the actual media streams is
>>>>>>>>>> generated from the entity described. How the generated content
>>>>>>>>>> relates
>>>>>>>>>> to the entity described in the participant info is dependent on
>>>>>>>>>> media
>>>>>>>>>> type and and the presentation attribute.
>>>>>>>>>>
>>>>>>>>>> For example:
>>>>>>>>>> - a video capture with participant info would indicate that the
>>>>>>>>>> video
>>>>>>>>>> contains a picture of the entity associated with the information
>>>>>>>>>> provided.
>>>>>>>>>> - a video capture with participant info and the presentation
>>>>>>>>>> attribute
>>>>>>>>>> would indicate that the presentation video is associated with 
>>>>>>>>>> the
>>>>>>>>>> participant but could contain any video content.
>>>>>>>>>> - a text capture with participant info would indicate that the
>>>>>>>>>> text is
>>>>>>>>>> generated from the actual participant.
>>>>>>>>>> - a text capture with participant info and the presentation
>>>>>>>>>> attribute
>>>>>>>>>> would indicate that the text is associated with the participant
>>>>>>>>>> but
>>>>>>>>>> could contain any text content."
>>>>>>>>>>
>>>>>>>>>> Comments?
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>> Regards, Christian
>>>>>>>>>>
>>>>>>>>>> _______________________________________________
>>>>>>>>>> clue mailing list
>>>>>>>>>> clue@ietf.org
>>>>>>>>>> https://www.ietf.org/mailman/listinfo/clue
>>>>>>>>>>
>>>>>>>>>
>>>>>>>>> _______________________________________________
>>>>>>>>> clue mailing list
>>>>>>>>> clue@ietf.org
>>>>>>>>> https://www.ietf.org/mailman/listinfo/clue
>>>>>>>>>
>>>>>>>>
>>>>>>>> _______________________________________________
>>>>>>>> clue mailing list
>>>>>>>> clue@ietf.org
>>>>>>>> https://www.ietf.org/mailman/listinfo/clue
>>>>>>>>
>>>>>>>
>>>>>>>
>>>>>>
>>>>>>
>>>>>
>>>>>
>>>>
>>>> _______________________________________________
>>>> clue mailing list
>>>> clue@ietf.org
>>>> https://www.ietf.org/mailman/listinfo/clue
>>>>
>>>
>>>
>>
>> _______________________________________________
>> clue mailing list
>> clue@ietf.org
>> https://www.ietf.org/mailman/listinfo/clue
>>
>
> _______________________________________________
> clue mailing list
> clue@ietf.org
> https://www.ietf.org/mailman/listinfo/clue
>


From nobody Fri Apr  4 07:25:53 2014
Return-Path: <Mark.Duckworth@polycom.com>
X-Original-To: clue@ietfa.amsl.com
Delivered-To: clue@ietfa.amsl.com
Received: from localhost (ietfa.amsl.com [127.0.0.1]) by ietfa.amsl.com (Postfix) with ESMTP id 8C0031A01E3 for <clue@ietfa.amsl.com>; Fri,  4 Apr 2014 07:25:51 -0700 (PDT)
X-Virus-Scanned: amavisd-new at amsl.com
X-Spam-Flag: NO
X-Spam-Score: -1.82
X-Spam-Level: 
X-Spam-Status: No, score=-1.82 tagged_above=-999 required=5 tests=[BAYES_00=-1.9, RCVD_IN_DNSWL_LOW=-0.7, SPF_NEUTRAL=0.779, UNPARSEABLE_RELAY=0.001] autolearn=ham
Received: from mail.ietf.org ([4.31.198.44]) by localhost (ietfa.amsl.com [127.0.0.1]) (amavisd-new, port 10024) with ESMTP id y-TNOA7_Fjk6 for <clue@ietfa.amsl.com>; Fri,  4 Apr 2014 07:25:47 -0700 (PDT)
Received: from mail1.bemta7.messagelabs.com (mail1.bemta7.messagelabs.com [216.82.254.106]) by ietfa.amsl.com (Postfix) with ESMTP id 20D6F1A01FB for <clue@ietf.org>; Fri,  4 Apr 2014 07:25:47 -0700 (PDT)
Received: from [216.82.254.19:7088] by server-10.bemta-7.messagelabs.com id F2/DD-11882-6E0CE335; Fri, 04 Apr 2014 14:25:42 +0000
X-Env-Sender: Mark.Duckworth@polycom.com
X-Msg-Ref: server-3.tower-96.messagelabs.com!1396621537!4842507!11
X-Originating-IP: [140.242.64.158]
X-StarScan-Received: 
X-StarScan-Version: 6.11.1; banners=-,-,-
X-VirusChecked: Checked
Received: (qmail 32098 invoked from network); 4 Apr 2014 14:25:41 -0000
Received: from crpehubprd01.polycom.com (HELO crpehubprd02.polycom.com) (140.242.64.158) by server-3.tower-96.messagelabs.com with AES128-SHA encrypted SMTP; 4 Apr 2014 14:25:41 -0000
Received: from CRPMBOXPRD07.polycom.com ([fe80::8113:9ad1:f9be:53f1]) by crpehubprd02.polycom.com ([fe80::5efe:10.236.0.154%12]) with mapi; Fri, 4 Apr 2014 07:25:23 -0700
From: "Duckworth, Mark" <Mark.Duckworth@polycom.com>
To: Christian Groves <Christian.Groves@nteczone.com>, "clue@ietf.org" <clue@ietf.org>
Date: Fri, 4 Apr 2014 07:25:18 -0700
Thread-Topic: [clue] Global CSE List in framework
Thread-Index: Ac9PlDnOHO/fCt+hSMWZVZ6WbC5NVwAeOvSQ
Message-ID: <49E45C59CA48264997FEBFB29B6BC2D62151CDC8AC@CRPMBOXPRD07.polycom.com>
References: <49E45C59CA48264997FEBFB29B6BC2D617ACF1B213@CRPMBOXPRD07.polycom.com> <533A14A0.6030807@nteczone.com> <49E45C59CA48264997FEBFB29B6BC2D62151A17223@CRPMBOXPRD07.polycom.com> <533DEE34.7070809@nteczone.com>
In-Reply-To: <533DEE34.7070809@nteczone.com>
Accept-Language: en-US
Content-Language: en-US
X-MS-Has-Attach: 
X-MS-TNEF-Correlator: 
acceptlanguage: en-US
Content-Type: text/plain; charset="us-ascii"
Content-Transfer-Encoding: quoted-printable
MIME-Version: 1.0
Archived-At: http://mailarchive.ietf.org/arch/msg/clue/OkgbIrzechNyc7lfFvC9iWn5rEY
Subject: Re: [clue] Global CSE List in framework
X-BeenThere: clue@ietf.org
X-Mailman-Version: 2.1.15
Precedence: list
List-Id: CLUE - ControLling mUltiple streams for TElepresence <clue.ietf.org>
List-Unsubscribe: <https://www.ietf.org/mailman/options/clue>, <mailto:clue-request@ietf.org?subject=unsubscribe>
List-Archive: <http://www.ietf.org/mail-archive/web/clue/>
List-Post: <mailto:clue@ietf.org>
List-Help: <mailto:clue-request@ietf.org?subject=help>
List-Subscribe: <https://www.ietf.org/mailman/listinfo/clue>, <mailto:clue-request@ietf.org?subject=subscribe>
X-List-Received-Date: Fri, 04 Apr 2014 14:25:51 -0000

Hi Christian,
Thanks again.  See below.
Mark

> -----Original Message-----
> From: Christian Groves [mailto:Christian.Groves@nteczone.com]
> Sent: Thursday, April 03, 2014 7:27 PM
> To: Duckworth, Mark; clue@ietf.org
> Subject: Re: [clue] Global CSE List in framework
>=20
> Hello Mark,
>=20
> Please see below.
>=20
> Regards, Christian
>=20
> ..snip..
> >>> The Provider can include multiple sets, to accommodate different
> >>> Consumers with varying capability to receive multiple encodings.
> >>> This is very similar to how each CSE represents a particular scene.
> >>>
> >> [CNG] Is the GCSE about providing sets for multiple "encodings"? You
> >> can provide multiple encodings without GSEs. Isn't it about providing
> >> different sets of captures so that a receiver can best determine what
> >> the set of captures it wants based on the service it wants to provide.
> >> I.e. a consumer may only support three media streams (3 encodings)
> >> but these may be used in a number of ways.
> > [Duckworth, Mark] Are you suggesting just replace "encodings" with
> "captures"?  That's fine with me, maybe "captures" better expresses the
> meaning.
> [CNG] Whilst "captures" is probably better than "encodings" I'm not sure =
that
> "receive multiple captures" is the point. How about "The Provider can inc=
lude
> multiple sets, to allow a consumer to choose sets of captures appropriate=
 to
> its capabilities or application..."
[Duckworth, Mark] Thanks, I think that says it better.

> > ...snip ...
> >>> * The ordering of items (sets of CSEs) in the global CSE list is not
> >>> important.
> >>>
> >>> * The ordering of CSEs within each set is not important.
> >>>
> >>> * The Provider must be capable of encoding and sending all Captures
> >>> within the CSEs of a given set simultaneously.
> >>>
> >> [CNG] Do we need to add some text to the simultaneous set section of
> >> the framework indicating this? We should probably say something about
> >> the interaction. i.e. If a GCSE set is included should the provider
> >> actually need to provide a STS? Sendings CSEs in a STS indicates the
> >> those CSE/captures may be used simultaneously. There appears to be
> >> alot of overlap with the GCSE in this respect.
> > [Duckworth, Mark] I think I see what you mean.  We have this paragraph
> already:
> > "If an Advertisement does not include Simultaneous Transmission Sets,
> then the Provider MUST be able to provide all Capture Scenes
> simultaneously.  If multiple capture Scene Entries are in a Capture Scene=
 then
> the Consumer chooses at most one Capture Scene Entry per Capture Scene
> for each media type."
> >
> > [Duckworth, Mark] I suggest we also add "If there is no STS and there i=
s a
> global CSE list, then the Consumer chooses at most one set of CSEs of eac=
h
> media type, from the global CSE list."
> [CNG] I'm OK with the above sentence. I think in addition we need to add
> some text regarding what happens when an STS AND GCSE is included.
> There's some overlap between them.
>=20
> E.g.
> Scene1(CSE1(VC1,VC2,VC3)
> Scene2(CSE2(VC4,VC5,VC6)
> Scene3(CSE3(VC7,VC8,VC9)
> GCSE([CSE1,CSE3],[CSE2,CSE3])
> STS([CSE1,CSE3],[CSE2,CSE3])
>=20
> If the provider includes a GCSE as above and wants to include a STS does =
it
> have to contain the exact same CSE sets?, i.e. because according to the
> above bullet the CSE must capable of being send simultaneously. Or is the
> interaction that a GCSE set indicates that at least one capture from each=
 CSE
> can be sent simultaneously, or?
>=20
[Duckworth, Mark] No, it shouldn't have to contain the exact same CSE sets =
in both.  The point is that if there are STS in the advertisement, then the=
 STSs must make it possible for: "The Provider must be capable of encoding =
and sending all Captures within the CSEs of a given set (in the global CSE =
list) simultaneously."  In other words, the STSs must not express any restr=
iction that prohibits the consumer from requesting all the captures in all =
the CSEs in a set of CSEs in the global CSE list.  But it can express restr=
ictions that do not violate the global CSE list.
So from your example, the GCSE list is saying the consumer can choose from =
these sets of captures:
    VC1,VC2,VC3,VC7,VC8,VC9
    or VC4,VC5,VC6,VC7,VC8,VC9
So there could still be STSs that don't contradict that, but do express lim=
itations.  For example:
    STS([VC1,VC2,VC3,VC4,VC7,VC8,VC9],[VC1,VC4,VC5,VC6,VC7,VC8,VC9])
So with these STSs, the consumer could choose everything from the first set=
 in the GCSE list (VC1,VC2,VC3,VC7,VC8,VC9) plus also VC4 if it wants.  But=
 not VC5 or VC6.

[Duckworth, Mark] Do you have a suggestion for how to make this more clear =
in the framework?


From nobody Fri Apr  4 08:47:04 2014
Return-Path: <Mark.Duckworth@polycom.com>
X-Original-To: clue@ietfa.amsl.com
Delivered-To: clue@ietfa.amsl.com
Received: from localhost (ietfa.amsl.com [127.0.0.1]) by ietfa.amsl.com (Postfix) with ESMTP id 9D3941A0230 for <clue@ietfa.amsl.com>; Fri,  4 Apr 2014 08:47:01 -0700 (PDT)
X-Virus-Scanned: amavisd-new at amsl.com
X-Spam-Flag: NO
X-Spam-Score: -1.12
X-Spam-Level: 
X-Spam-Status: No, score=-1.12 tagged_above=-999 required=5 tests=[BAYES_00=-1.9, RCVD_IN_DNSWL_NONE=-0.0001, SPF_NEUTRAL=0.779, UNPARSEABLE_RELAY=0.001] autolearn=no
Received: from mail.ietf.org ([4.31.198.44]) by localhost (ietfa.amsl.com [127.0.0.1]) (amavisd-new, port 10024) with ESMTP id Ry_u2RSl2k9X for <clue@ietfa.amsl.com>; Fri,  4 Apr 2014 08:46:57 -0700 (PDT)
Received: from mail1.bemta12.messagelabs.com (mail1.bemta12.messagelabs.com [216.82.251.6]) by ietfa.amsl.com (Postfix) with ESMTP id 9DA221A021E for <clue@ietf.org>; Fri,  4 Apr 2014 08:46:57 -0700 (PDT)
Received: from [216.82.249.212:36220] by server-6.bemta-12.messagelabs.com id 8B/13-08935-DE3DE335; Fri, 04 Apr 2014 15:46:53 +0000
X-Env-Sender: Mark.Duckworth@polycom.com
X-Msg-Ref: server-3.tower-219.messagelabs.com!1396626410!4020269!4
X-Originating-IP: [140.242.64.158]
X-StarScan-Received: 
X-StarScan-Version: 6.11.1; banners=-,-,-
X-VirusChecked: Checked
Received: (qmail 6860 invoked from network); 4 Apr 2014 15:46:52 -0000
Received: from crpehubprd01.polycom.com (HELO crpehubprd02.polycom.com) (140.242.64.158) by server-3.tower-219.messagelabs.com with AES128-SHA encrypted SMTP; 4 Apr 2014 15:46:52 -0000
Received: from CRPMBOXPRD07.polycom.com ([fe80::91fc:8a0f:5258:aff0]) by crpehubprd02.polycom.com ([fe80::5efe:10.236.0.154%12]) with mapi; Fri, 4 Apr 2014 08:46:11 -0700
From: "Duckworth, Mark" <Mark.Duckworth@polycom.com>
To: "clue@ietf.org" <clue@ietf.org>
Date: Fri, 4 Apr 2014 08:46:09 -0700
Thread-Topic: [clue] Participant info/type followup
Thread-Index: Ac9PUdPhqezSWk/zRZGlqD1bHP+nPwAyOqmQ
Message-ID: <49E45C59CA48264997FEBFB29B6BC2D6215264287E@CRPMBOXPRD07.polycom.com>
References: <5318809E.2030204@nteczone.com> <5318AE5E.4050404@alum.mit.edu> <5318B0C9.1050603@nteczone.com> <5318B48E.3090300@alum.mit.edu> <53290C9B.6090106@nteczone.com> <5329F3FF.7090606@alum.mit.edu> <533A16BA.40707@nteczone.com> <533A91BF.7040404@unina.it> <533B47FF.5020407@nteczone.com> <533D7E8F.4000303@alum.mit.edu>
In-Reply-To: <533D7E8F.4000303@alum.mit.edu>
Accept-Language: en-US
Content-Language: en-US
X-MS-Has-Attach: 
X-MS-TNEF-Correlator: 
acceptlanguage: en-US
Content-Type: text/plain; charset="us-ascii"
Content-Transfer-Encoding: quoted-printable
MIME-Version: 1.0
Archived-At: http://mailarchive.ietf.org/arch/msg/clue/XILTwoUwBzUkSAhfCehUqPlu8jI
Subject: Re: [clue] Participant info/type followup
X-BeenThere: clue@ietf.org
X-Mailman-Version: 2.1.15
Precedence: list
List-Id: CLUE - ControLling mUltiple streams for TElepresence <clue.ietf.org>
List-Unsubscribe: <https://www.ietf.org/mailman/options/clue>, <mailto:clue-request@ietf.org?subject=unsubscribe>
List-Archive: <http://www.ietf.org/mail-archive/web/clue/>
List-Post: <mailto:clue@ietf.org>
List-Help: <mailto:clue-request@ietf.org?subject=help>
List-Subscribe: <https://www.ietf.org/mailman/listinfo/clue>, <mailto:clue-request@ietf.org?subject=subscribe>
X-List-Received-Date: Fri, 04 Apr 2014 15:47:01 -0000

> -----Original Message-----
> From: clue [mailto:clue-bounces@ietf.org] On Behalf Of Paul Kyzivat
> Sent: Thursday, April 03, 2014 11:30 AM
> To: clue@ietf.org
> Subject: Re: [clue] Participant info/type followup
>=20
> ISTM we have both a semantic issue and a syntactic/terminological issue:
>=20
> - we have the semantic distinction between who is represented within
>    the capture and who is responsible for sending the capture

[Duckworth, Mark] I agree, but I thought the intent of the participant attr=
ibutes was to give information about "who is represented within the capture=
" and not about "who is responsible for sending the capture".  From the fra=
mework:

    "... information regarding the conference participants *in* a Capture."
    "... indicates the type of participant/s contained *in* the capture..."

[Duckworth, Mark] I think we don't need to add a way to indicate "who is re=
sponsible for sending the capture".  IIRC, this was not a goal of the origi=
nal proposal and discussion about "roles".  But if the group agrees we shou=
ld add this now, I wouldn't object too strongly.

> - we have a terminology issue regarding what to call these two things
>=20
> I think we now agree that the semantic distinction is real and should be
> identified in the advertisement.

[Duckworth, Mark] I'm mildly against adding the ability to describe "who is=
 responsible for sending the capture".

> Regarding terminology, IMO we should use two different terms to identify
> these. At most one of them should be called "participant". Given the conf=
lict
> with xcon, perhaps *neither* of them should be called participant.
>=20
> This is mostly a data model issue. Some of it might need to peek through =
into
> the framework. And what does should be consistent.
>=20
> 	Thanks,
> 	Paul
>=20
> On 4/1/14 7:13 PM, Christian Groves wrote:
> > Hello Roberta,
> >
> >  From a data model perspective I agree that it makes sense to have the
> > participant metadata "referenced" to minimize duplication in the
> > messages. So I support your solution in the data model to the
> > redundancy problem.

[Duckworth, Mark] I agree too.

> > In general I think its good to maintain consistency between the
> > framework and the data model. However in London I thought this was
> > more of a syntax shortcut (like the captureID wildcarding) rather than
> > something we'd need to formalize in the framework. I guess we could
> > add it at a higher level in the framework (it would functionally be
> > the same thing), it would just be more work for the editor...
> >
> > I could go either way on this.
> >
> > Regards, Christian
snip


From nobody Fri Apr  4 13:49:34 2014
Return-Path: <pkyzivat@alum.mit.edu>
X-Original-To: clue@ietfa.amsl.com
Delivered-To: clue@ietfa.amsl.com
Received: from localhost (ietfa.amsl.com [127.0.0.1]) by ietfa.amsl.com (Postfix) with ESMTP id 94DE31A0392 for <clue@ietfa.amsl.com>; Fri,  4 Apr 2014 13:49:31 -0700 (PDT)
X-Virus-Scanned: amavisd-new at amsl.com
X-Spam-Flag: NO
X-Spam-Score: -1.235
X-Spam-Level: 
X-Spam-Status: No, score=-1.235 tagged_above=-999 required=5 tests=[BAYES_00=-1.9, DKIM_SIGNED=0.1, DKIM_VALID=-0.1, SPF_SOFTFAIL=0.665] autolearn=no
Received: from mail.ietf.org ([4.31.198.44]) by localhost (ietfa.amsl.com [127.0.0.1]) (amavisd-new, port 10024) with ESMTP id WKCjYLsxtdXg for <clue@ietfa.amsl.com>; Fri,  4 Apr 2014 13:49:27 -0700 (PDT)
Received: from qmta10.westchester.pa.mail.comcast.net (qmta10.westchester.pa.mail.comcast.net [IPv6:2001:558:fe14:43:76:96:62:17]) by ietfa.amsl.com (Postfix) with ESMTP id 60C751A0116 for <clue@ietf.org>; Fri,  4 Apr 2014 13:49:26 -0700 (PDT)
Received: from omta20.westchester.pa.mail.comcast.net ([76.96.62.71]) by qmta10.westchester.pa.mail.comcast.net with comcast id lnPL1n0011YDfWL5AwpMoG; Fri, 04 Apr 2014 20:49:21 +0000
Received: from Paul-Kyzivats-MacBook-Pro.local ([50.138.229.164]) by omta20.westchester.pa.mail.comcast.net with comcast id lwpM1n00G3ZTu2S3gwpMt6; Fri, 04 Apr 2014 20:49:21 +0000
Message-ID: <533F1AD1.7060507@alum.mit.edu>
Date: Fri, 04 Apr 2014 16:49:21 -0400
From: Paul Kyzivat <pkyzivat@alum.mit.edu>
User-Agent: Mozilla/5.0 (Macintosh; Intel Mac OS X 10.7; rv:24.0) Gecko/20100101 Thunderbird/24.4.0
MIME-Version: 1.0
To: Roni Even <ron.even.tlv@gmail.com>, Jonathan Lennox <jonathan@vidyo.com>
References: <533E7A50.5040909@ericsson.com>
In-Reply-To: <533E7A50.5040909@ericsson.com>
Content-Type: text/plain; charset=ISO-8859-1; format=flowed
Content-Transfer-Encoding: 8bit
DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=comcast.net; s=q20140121; t=1396644561; bh=sCrAxzK0TA/T9JTfnPzUgffTfjz0f7EeP3YYETo9k5o=; h=Received:Received:Message-ID:Date:From:MIME-Version:To:Subject: Content-Type; b=OKWyVKTMmfFt5aHs+OGhnFw2CJZAO+Q5rXwnWZQ+bqjzKWfwzIbcfxa9LGzFYUImA Gb6JXcH8IfTAdfrM85E70/2t/ydkCgagvdWFJXsaUXsW1T0Gq1K3wOV1gvaHnzgZGA ovjwQLuWME/QXGJ3RoqK0kDyA8BcIVPxKOyj1spPUIis9rq3jhbmmKPZNn0UFDI6Zb fDqYCgCM6v6X8aNV3EDjxO2FX407sQm+UNd/nfyi3tQhGoFTOyN6Wy6qZat40XfEi4 g6Y4Bvjos7EgSl6yAJhJ8S9+x4ZzNXu4TrRp9/N5jkuHVvRIjSjNc+2HvrJcIsViAT muKkJ7uB6pOTw==
Archived-At: http://mailarchive.ietf.org/arch/msg/clue/MybTKWiXw9BlHfcHGCTGkJXVacY
Cc: CLUE <clue@ietf.org>
Subject: Re: [clue] [rtcweb] Draft proposal for updating Multiparty topologies in draft-ietf-rtcweb-rtp-usage
X-BeenThere: clue@ietf.org
X-Mailman-Version: 2.1.15
Precedence: list
List-Id: CLUE - ControLling mUltiple streams for TElepresence <clue.ietf.org>
List-Unsubscribe: <https://www.ietf.org/mailman/options/clue>, <mailto:clue-request@ietf.org?subject=unsubscribe>
List-Archive: <http://www.ietf.org/mail-archive/web/clue/>
List-Post: <mailto:clue@ietf.org>
List-Help: <mailto:clue-request@ietf.org?subject=help>
List-Subscribe: <https://www.ietf.org/mailman/listinfo/clue>, <mailto:clue-request@ietf.org?subject=subscribe>
X-List-Received-Date: Fri, 04 Apr 2014 20:49:31 -0000

Roni & Jonathan,

Can one of you comment on whether any of this has an impact on CLUE interop?

	Thanks,
	Paul

On 4/4/14 5:24 AM, Magnus Westerlund wrote:
> WG,
> (As Author)
>
> Colin and I have been working on resolving terminology usage and in
> general giving the draft a polish over before the WG last call. One
> section that has gotten quite some attention from us is the below one.
> The changes are significantly from an editorial stand point. However,
> the intended content has not been intended to be changed, although
> clarified and better motivated. So please review it, we intended to
> include this in the upcoming draft update. Feedback appreciated.
>
> 5.1.  Conferencing Extensions and Topologies
>
>     RTP is a protocol that inherently supports group communication.
>     Groups can be implemented by having each endpoint sending its RTP
>     packet streams to an RTP middlebox that redistributes the traffic, by
>     using a mesh of unicast RTP packet streams between endpoints, or by
>     using an IP multicast group to distribute the RTP packet streams.
>     These topologies can be implemented in a number of ways as discussed
>     in [I-D.ietf-avtcore-rtp-topologies-update].
>
>     While the use of IP multicast groups is popular in IPTV systems, the
>     topologies based on RTP middleboxes are dominant in interactive video
>     conferencing environments.  Topologies based on a mesh of unicast
>     transport-layer flows to create a common RTP session have not seen
>     widespread deployment.  Accordingly, WebRTC implementations are not
>     expected to support topologies based on IP multicast groups.  WebRTC
>     implementations are also not expected to support mesh-based
>     topologies, such as a point-to-multipoint mesh configured as a single
>     RTP session (Topo-Mesh in the terminology of
>     [I-D.ietf-avtcore-rtp-topologies-update]).  However, a point-to-
>     multipoint mesh constructed using several RTP sessions, in the WebRTC
>     context using independent RTCPeerConnections can be expected to be
>     utilised by WebRTC applications.
>
>     WebRTC implementations of RTP endpoints implemented according to this
>     memo are expected to support all the topologies described in
>     [I-D.ietf-avtcore-rtp-topologies-update] where the RTP endpoints send
>     and receive unicast RTP packet streams to some peer device, provided
>     that peer can participate in performing congestion control on the RTP
>     packet streams.  The peer device could be another RTP endpoint, or it
>     could be an RTP middlebox that redistributes the RTP packet streams
>     to other RTP endpoints.  This limitation means that some of the RTP
>     middlebox-based topologies are not suitable for use in the WebRTC
>     environment.  Specifically:
>
>     o  Video switching MCUs (Topo-Video-switch-MCU) SHOULD NOT be used,
>        since they make the use of RTCP for congestion control and quality
>        of service reports problematic (see Section 3.6.2 of
>        [I-D.ietf-avtcore-rtp-topologies-update]).
>
>     o  Content modifying MCUs with RTCP termination (Topo-RTCP-
>        terminating-MCU) SHOULD NOT be used since they break RTP loop
>        detection, and prevent receivers from identifying active senders
>        (see section 3.8 of [I-D.ietf-avtcore-rtp-topologies-update]).
>
>     o  The Relay-Transport Translator (Topo-PtM-Trn-Translator) topology
>        SHOULD NOT be used because its safe use requires a point to
>        multipoint congestion control algorithm or RTP circuit breaker,
>        which has not yet been standardised.
>
>     The RTP extensions described in Section 5.1.1 to Section 5.1.6 are
>     designed to be used with centralised conferencing, where an RTP
>     middlebox (e.g., a conference bridge) receives a participant's RTP
>     packet streams and distributes them to the other participants.  These
>     extensions are not necessary for interoperability; an RTP end-point
>     that does not implement these extensions will work correctly, but
>     might offer poor performance.  Support for the listed extensions will
>     greatly improve the quality of experience and, to provide a
>     reasonable baseline quality, some of these extensions are mandatory
>     to be supported by WebRTC end-points.
>
>     The RTCP conferencing extensions are defined in Extended RTP Profile
>     for Real-time Transport Control Protocol (RTCP)-Based Feedback (RTP/
>     AVPF) [RFC4585] and the "Codec Control Messages in the RTP Audio-
>     Visual Profile with Feedback (AVPF)" (CCM) [RFC5104] and are fully
>     usable by the Secure variant of this profile (RTP/SAVPF) [RFC5124].
>
>
> Cheers
>
> Magnus Westerlund
>
> ----------------------------------------------------------------------
> Services, Media and Network features, Ericsson Research EAB/TXM
> ----------------------------------------------------------------------
> Ericsson AB                 | Phone  +46 10 7148287
> Färögatan 6                 | Mobile +46 73 0949079
> SE-164 80 Stockholm, Sweden | mailto: magnus.westerlund@ericsson.com
> ----------------------------------------------------------------------
>
> _______________________________________________
> rtcweb mailing list
> rtcweb@ietf.org
> https://www.ietf.org/mailman/listinfo/rtcweb
>


From nobody Fri Apr  4 14:28:50 2014
Return-Path: <mary.ietf.barnes@gmail.com>
X-Original-To: clue@ietfa.amsl.com
Delivered-To: clue@ietfa.amsl.com
Received: from localhost (ietfa.amsl.com [127.0.0.1]) by ietfa.amsl.com (Postfix) with ESMTP id B96941A0104 for <clue@ietfa.amsl.com>; Fri,  4 Apr 2014 14:28:48 -0700 (PDT)
X-Virus-Scanned: amavisd-new at amsl.com
X-Spam-Flag: NO
X-Spam-Score: -1.999
X-Spam-Level: 
X-Spam-Status: No, score=-1.999 tagged_above=-999 required=5 tests=[BAYES_00=-1.9, DKIM_SIGNED=0.1, DKIM_VALID=-0.1, DKIM_VALID_AU=-0.1, FREEMAIL_FROM=0.001, HTML_MESSAGE=0.001, SPF_PASS=-0.001] autolearn=ham
Received: from mail.ietf.org ([4.31.198.44]) by localhost (ietfa.amsl.com [127.0.0.1]) (amavisd-new, port 10024) with ESMTP id eEjuSygHjQ9R for <clue@ietfa.amsl.com>; Fri,  4 Apr 2014 14:28:44 -0700 (PDT)
Received: from mail-we0-x230.google.com (mail-we0-x230.google.com [IPv6:2a00:1450:400c:c03::230]) by ietfa.amsl.com (Postfix) with ESMTP id B14701A00E5 for <clue@ietf.org>; Fri,  4 Apr 2014 14:28:43 -0700 (PDT)
Received: by mail-we0-f176.google.com with SMTP id x48so3917371wes.7 for <clue@ietf.org>; Fri, 04 Apr 2014 14:28:38 -0700 (PDT)
DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20120113;  h=mime-version:in-reply-to:references:date:message-id:subject:from:to :cc:content-type; bh=c4BHpWF8BeDNKsKfyMYIE2tfCroTrLEa2iMxoTe2zRo=; b=Z/GmRHIfBDf8Rb0Ss9lPVYKvS2ECftQsKw6aqRH597mgpIJMyFhjOvWnnftDDMK9VG qaw9doTM58XQ5qM2C1+eQBBmIlprotkYM/gcP/c/P9Jn+sc6ps9F9+fAQPml29UPWMa7 g4L0jne9eXCgRVBYp5yDNeso0+zytqs9frRQGvQsyHv3zlvDp8/E48R0eCC6Qgix2eHy pELT83BKTZqbz1KJMvowP4C0o7l2oEatP2rDAWRohf6HN+orT28BIOte2ELWNudn0hif Ujgcv1szsccqJnTR/fP1z6DtxCAW99yfWlyO9ne+FQHD/6FBk9D+Jb2fq3JlfVkHtQvX qOfQ==
MIME-Version: 1.0
X-Received: by 10.180.19.69 with SMTP id c5mr7505814wie.7.1396646918620; Fri, 04 Apr 2014 14:28:38 -0700 (PDT)
Received: by 10.216.10.6 with HTTP; Fri, 4 Apr 2014 14:28:38 -0700 (PDT)
In-Reply-To: <533F1AD1.7060507@alum.mit.edu>
References: <533E7A50.5040909@ericsson.com> <533F1AD1.7060507@alum.mit.edu>
Date: Fri, 4 Apr 2014 16:28:38 -0500
Message-ID: <CAHBDyN7duDtr0b+uHGn8im5Ny2vCUN5EvSurYYt_M5YXYx0KAA@mail.gmail.com>
From: Mary Barnes <mary.ietf.barnes@gmail.com>
To: Paul Kyzivat <pkyzivat@alum.mit.edu>
Content-Type: multipart/alternative; boundary=bcaec53d5a8b760bd404f63e332b
Archived-At: http://mailarchive.ietf.org/arch/msg/clue/xSRyb3VGGuQ-0MDRH4ScmqVVwaM
Cc: Jonathan Lennox <jonathan@vidyo.com>, CLUE <clue@ietf.org>
Subject: Re: [clue] [rtcweb] Draft proposal for updating Multiparty topologies in draft-ietf-rtcweb-rtp-usage
X-BeenThere: clue@ietf.org
X-Mailman-Version: 2.1.15
Precedence: list
List-Id: CLUE - ControLling mUltiple streams for TElepresence <clue.ietf.org>
List-Unsubscribe: <https://www.ietf.org/mailman/options/clue>, <mailto:clue-request@ietf.org?subject=unsubscribe>
List-Archive: <http://www.ietf.org/mail-archive/web/clue/>
List-Post: <mailto:clue@ietf.org>
List-Help: <mailto:clue-request@ietf.org?subject=help>
List-Subscribe: <https://www.ietf.org/mailman/listinfo/clue>, <mailto:clue-request@ietf.org?subject=subscribe>
X-List-Received-Date: Fri, 04 Apr 2014 21:28:49 -0000

--bcaec53d5a8b760bd404f63e332b
Content-Type: text/plain; charset=ISO-8859-1
Content-Transfer-Encoding: quoted-printable

Note, this is also a good reminder, that we should do the same scrubbing of
CLUE documents in terms of terminology (per the action from IETF-89).

Mary.


On Fri, Apr 4, 2014 at 3:49 PM, Paul Kyzivat <pkyzivat@alum.mit.edu> wrote:

> Roni & Jonathan,
>
> Can one of you comment on whether any of this has an impact on CLUE
> interop?
>
>         Thanks,
>         Paul
>
>
> On 4/4/14 5:24 AM, Magnus Westerlund wrote:
>
>> WG,
>> (As Author)
>>
>> Colin and I have been working on resolving terminology usage and in
>> general giving the draft a polish over before the WG last call. One
>> section that has gotten quite some attention from us is the below one.
>> The changes are significantly from an editorial stand point. However,
>> the intended content has not been intended to be changed, although
>> clarified and better motivated. So please review it, we intended to
>> include this in the upcoming draft update. Feedback appreciated.
>>
>> 5.1.  Conferencing Extensions and Topologies
>>
>>     RTP is a protocol that inherently supports group communication.
>>     Groups can be implemented by having each endpoint sending its RTP
>>     packet streams to an RTP middlebox that redistributes the traffic, b=
y
>>     using a mesh of unicast RTP packet streams between endpoints, or by
>>     using an IP multicast group to distribute the RTP packet streams.
>>     These topologies can be implemented in a number of ways as discussed
>>     in [I-D.ietf-avtcore-rtp-topologies-update].
>>
>>     While the use of IP multicast groups is popular in IPTV systems, the
>>     topologies based on RTP middleboxes are dominant in interactive vide=
o
>>     conferencing environments.  Topologies based on a mesh of unicast
>>     transport-layer flows to create a common RTP session have not seen
>>     widespread deployment.  Accordingly, WebRTC implementations are not
>>     expected to support topologies based on IP multicast groups.  WebRTC
>>     implementations are also not expected to support mesh-based
>>     topologies, such as a point-to-multipoint mesh configured as a singl=
e
>>     RTP session (Topo-Mesh in the terminology of
>>     [I-D.ietf-avtcore-rtp-topologies-update]).  However, a point-to-
>>     multipoint mesh constructed using several RTP sessions, in the WebRT=
C
>>     context using independent RTCPeerConnections can be expected to be
>>     utilised by WebRTC applications.
>>
>>     WebRTC implementations of RTP endpoints implemented according to thi=
s
>>     memo are expected to support all the topologies described in
>>     [I-D.ietf-avtcore-rtp-topologies-update] where the RTP endpoints sen=
d
>>     and receive unicast RTP packet streams to some peer device, provided
>>     that peer can participate in performing congestion control on the RT=
P
>>     packet streams.  The peer device could be another RTP endpoint, or i=
t
>>     could be an RTP middlebox that redistributes the RTP packet streams
>>     to other RTP endpoints.  This limitation means that some of the RTP
>>     middlebox-based topologies are not suitable for use in the WebRTC
>>     environment.  Specifically:
>>
>>     o  Video switching MCUs (Topo-Video-switch-MCU) SHOULD NOT be used,
>>        since they make the use of RTCP for congestion control and qualit=
y
>>        of service reports problematic (see Section 3.6.2 of
>>        [I-D.ietf-avtcore-rtp-topologies-update]).
>>
>>     o  Content modifying MCUs with RTCP termination (Topo-RTCP-
>>        terminating-MCU) SHOULD NOT be used since they break RTP loop
>>        detection, and prevent receivers from identifying active senders
>>        (see section 3.8 of [I-D.ietf-avtcore-rtp-topologies-update]).
>>
>>     o  The Relay-Transport Translator (Topo-PtM-Trn-Translator) topology
>>        SHOULD NOT be used because its safe use requires a point to
>>        multipoint congestion control algorithm or RTP circuit breaker,
>>        which has not yet been standardised.
>>
>>     The RTP extensions described in Section 5.1.1 to Section 5.1.6 are
>>     designed to be used with centralised conferencing, where an RTP
>>     middlebox (e.g., a conference bridge) receives a participant's RTP
>>     packet streams and distributes them to the other participants.  Thes=
e
>>     extensions are not necessary for interoperability; an RTP end-point
>>     that does not implement these extensions will work correctly, but
>>     might offer poor performance.  Support for the listed extensions wil=
l
>>     greatly improve the quality of experience and, to provide a
>>     reasonable baseline quality, some of these extensions are mandatory
>>     to be supported by WebRTC end-points.
>>
>>     The RTCP conferencing extensions are defined in Extended RTP Profile
>>     for Real-time Transport Control Protocol (RTCP)-Based Feedback (RTP/
>>     AVPF) [RFC4585] and the "Codec Control Messages in the RTP Audio-
>>     Visual Profile with Feedback (AVPF)" (CCM) [RFC5104] and are fully
>>     usable by the Secure variant of this profile (RTP/SAVPF) [RFC5124].
>>
>>
>> Cheers
>>
>> Magnus Westerlund
>>
>> ----------------------------------------------------------------------
>> Services, Media and Network features, Ericsson Research EAB/TXM
>> ----------------------------------------------------------------------
>> Ericsson AB                 | Phone  +46 10 7148287
>> F=E4r=F6gatan 6                 | Mobile +46 73 0949079
>> SE-164 80 Stockholm, Sweden | mailto: magnus.westerlund@ericsson.com
>> ----------------------------------------------------------------------
>>
>> _______________________________________________
>> rtcweb mailing list
>> rtcweb@ietf.org
>> https://www.ietf.org/mailman/listinfo/rtcweb
>>
>>
> _______________________________________________
> clue mailing list
> clue@ietf.org
> https://www.ietf.org/mailman/listinfo/clue
>

--bcaec53d5a8b760bd404f63e332b
Content-Type: text/html; charset=ISO-8859-1
Content-Transfer-Encoding: quoted-printable

<div dir=3D"ltr">Note, this is also a good reminder, that we should do the =
same scrubbing of CLUE documents in terms of terminology (per the action fr=
om IETF-89).<div><br></div><div>Mary.</div></div><div class=3D"gmail_extra"=
>
<br><br><div class=3D"gmail_quote">On Fri, Apr 4, 2014 at 3:49 PM, Paul Kyz=
ivat <span dir=3D"ltr">&lt;<a href=3D"mailto:pkyzivat@alum.mit.edu" target=
=3D"_blank">pkyzivat@alum.mit.edu</a>&gt;</span> wrote:<br><blockquote clas=
s=3D"gmail_quote" style=3D"margin:0 0 0 .8ex;border-left:1px #ccc solid;pad=
ding-left:1ex">
Roni &amp; Jonathan,<br>
<br>
Can one of you comment on whether any of this has an impact on CLUE interop=
?<br>
<br>
=A0 =A0 =A0 =A0 Thanks,<br>
=A0 =A0 =A0 =A0 Paul<div><div class=3D"h5"><br>
<br>
On 4/4/14 5:24 AM, Magnus Westerlund wrote:<br>
<blockquote class=3D"gmail_quote" style=3D"margin:0 0 0 .8ex;border-left:1p=
x #ccc solid;padding-left:1ex">
WG,<br>
(As Author)<br>
<br>
Colin and I have been working on resolving terminology usage and in<br>
general giving the draft a polish over before the WG last call. One<br>
section that has gotten quite some attention from us is the below one.<br>
The changes are significantly from an editorial stand point. However,<br>
the intended content has not been intended to be changed, although<br>
clarified and better motivated. So please review it, we intended to<br>
include this in the upcoming draft update. Feedback appreciated.<br>
<br>
5.1. =A0Conferencing Extensions and Topologies<br>
<br>
=A0 =A0 RTP is a protocol that inherently supports group communication.<br>
=A0 =A0 Groups can be implemented by having each endpoint sending its RTP<b=
r>
=A0 =A0 packet streams to an RTP middlebox that redistributes the traffic, =
by<br>
=A0 =A0 using a mesh of unicast RTP packet streams between endpoints, or by=
<br>
=A0 =A0 using an IP multicast group to distribute the RTP packet streams.<b=
r>
=A0 =A0 These topologies can be implemented in a number of ways as discusse=
d<br>
=A0 =A0 in [I-D.ietf-avtcore-rtp-<u></u>topologies-update].<br>
<br>
=A0 =A0 While the use of IP multicast groups is popular in IPTV systems, th=
e<br>
=A0 =A0 topologies based on RTP middleboxes are dominant in interactive vid=
eo<br>
=A0 =A0 conferencing environments. =A0Topologies based on a mesh of unicast=
<br>
=A0 =A0 transport-layer flows to create a common RTP session have not seen<=
br>
=A0 =A0 widespread deployment. =A0Accordingly, WebRTC implementations are n=
ot<br>
=A0 =A0 expected to support topologies based on IP multicast groups. =A0Web=
RTC<br>
=A0 =A0 implementations are also not expected to support mesh-based<br>
=A0 =A0 topologies, such as a point-to-multipoint mesh configured as a sing=
le<br>
=A0 =A0 RTP session (Topo-Mesh in the terminology of<br>
=A0 =A0 [I-D.ietf-avtcore-rtp-<u></u>topologies-update]). =A0However, a poi=
nt-to-<br>
=A0 =A0 multipoint mesh constructed using several RTP sessions, in the WebR=
TC<br>
=A0 =A0 context using independent RTCPeerConnections can be expected to be<=
br>
=A0 =A0 utilised by WebRTC applications.<br>
<br>
=A0 =A0 WebRTC implementations of RTP endpoints implemented according to th=
is<br>
=A0 =A0 memo are expected to support all the topologies described in<br>
=A0 =A0 [I-D.ietf-avtcore-rtp-<u></u>topologies-update] where the RTP endpo=
ints send<br>
=A0 =A0 and receive unicast RTP packet streams to some peer device, provide=
d<br>
=A0 =A0 that peer can participate in performing congestion control on the R=
TP<br>
=A0 =A0 packet streams. =A0The peer device could be another RTP endpoint, o=
r it<br>
=A0 =A0 could be an RTP middlebox that redistributes the RTP packet streams=
<br>
=A0 =A0 to other RTP endpoints. =A0This limitation means that some of the R=
TP<br>
=A0 =A0 middlebox-based topologies are not suitable for use in the WebRTC<b=
r>
=A0 =A0 environment. =A0Specifically:<br>
<br>
=A0 =A0 o =A0Video switching MCUs (Topo-Video-switch-MCU) SHOULD NOT be use=
d,<br>
=A0 =A0 =A0 =A0since they make the use of RTCP for congestion control and q=
uality<br>
=A0 =A0 =A0 =A0of service reports problematic (see Section 3.6.2 of<br>
=A0 =A0 =A0 =A0[I-D.ietf-avtcore-rtp-<u></u>topologies-update]).<br>
<br>
=A0 =A0 o =A0Content modifying MCUs with RTCP termination (Topo-RTCP-<br>
=A0 =A0 =A0 =A0terminating-MCU) SHOULD NOT be used since they break RTP loo=
p<br>
=A0 =A0 =A0 =A0detection, and prevent receivers from identifying active sen=
ders<br>
=A0 =A0 =A0 =A0(see section 3.8 of [I-D.ietf-avtcore-rtp-<u></u>topologies-=
update]).<br>
<br>
=A0 =A0 o =A0The Relay-Transport Translator (Topo-PtM-Trn-Translator) topol=
ogy<br>
=A0 =A0 =A0 =A0SHOULD NOT be used because its safe use requires a point to<=
br>
=A0 =A0 =A0 =A0multipoint congestion control algorithm or RTP circuit break=
er,<br>
=A0 =A0 =A0 =A0which has not yet been standardised.<br>
<br>
=A0 =A0 The RTP extensions described in Section 5.1.1 to Section 5.1.6 are<=
br>
=A0 =A0 designed to be used with centralised conferencing, where an RTP<br>
=A0 =A0 middlebox (e.g., a conference bridge) receives a participant&#39;s =
RTP<br>
=A0 =A0 packet streams and distributes them to the other participants. =A0T=
hese<br>
=A0 =A0 extensions are not necessary for interoperability; an RTP end-point=
<br>
=A0 =A0 that does not implement these extensions will work correctly, but<b=
r>
=A0 =A0 might offer poor performance. =A0Support for the listed extensions =
will<br>
=A0 =A0 greatly improve the quality of experience and, to provide a<br>
=A0 =A0 reasonable baseline quality, some of these extensions are mandatory=
<br>
=A0 =A0 to be supported by WebRTC end-points.<br>
<br>
=A0 =A0 The RTCP conferencing extensions are defined in Extended RTP Profil=
e<br>
=A0 =A0 for Real-time Transport Control Protocol (RTCP)-Based Feedback (RTP=
/<br>
=A0 =A0 AVPF) [RFC4585] and the &quot;Codec Control Messages in the RTP Aud=
io-<br>
=A0 =A0 Visual Profile with Feedback (AVPF)&quot; (CCM) [RFC5104] and are f=
ully<br>
=A0 =A0 usable by the Secure variant of this profile (RTP/SAVPF) [RFC5124].=
<br>
<br>
<br>
Cheers<br>
<br>
Magnus Westerlund<br>
<br>
------------------------------<u></u>------------------------------<u></u>-=
---------<br>
Services, Media and Network features, Ericsson Research EAB/TXM<br>
------------------------------<u></u>------------------------------<u></u>-=
---------<br>
Ericsson AB =A0 =A0 =A0 =A0 =A0 =A0 =A0 =A0 | Phone =A0<a href=3D"tel:%2B46=
%2010%207148287" value=3D"+46107148287" target=3D"_blank">+46 10 7148287</a=
><br>
F=E4r=F6gatan 6 =A0 =A0 =A0 =A0 =A0 =A0 =A0 =A0 | Mobile <a href=3D"tel:%2B=
46%2073%200949079" value=3D"+46730949079" target=3D"_blank">+46 73 0949079<=
/a><br>
SE-164 80 Stockholm, Sweden | mailto: <a href=3D"mailto:magnus.westerlund@e=
ricsson.com" target=3D"_blank">magnus.westerlund@ericsson.com</a><br>
------------------------------<u></u>------------------------------<u></u>-=
---------<br>
<br>
______________________________<u></u>_________________<br>
rtcweb mailing list<br>
<a href=3D"mailto:rtcweb@ietf.org" target=3D"_blank">rtcweb@ietf.org</a><br=
>
<a href=3D"https://www.ietf.org/mailman/listinfo/rtcweb" target=3D"_blank">=
https://www.ietf.org/mailman/<u></u>listinfo/rtcweb</a><br>
<br>
</blockquote>
<br></div></div>
______________________________<u></u>_________________<br>
clue mailing list<br>
<a href=3D"mailto:clue@ietf.org" target=3D"_blank">clue@ietf.org</a><br>
<a href=3D"https://www.ietf.org/mailman/listinfo/clue" target=3D"_blank">ht=
tps://www.ietf.org/mailman/<u></u>listinfo/clue</a><br>
</blockquote></div><br></div>

--bcaec53d5a8b760bd404f63e332b--


From nobody Sun Apr  6 19:17:44 2014
Return-Path: <Christian.Groves@nteczone.com>
X-Original-To: clue@ietfa.amsl.com
Delivered-To: clue@ietfa.amsl.com
Received: from localhost (ietfa.amsl.com [127.0.0.1]) by ietfa.amsl.com (Postfix) with ESMTP id 4A6981A064F for <clue@ietfa.amsl.com>; Sun,  6 Apr 2014 19:17:43 -0700 (PDT)
X-Virus-Scanned: amavisd-new at amsl.com
X-Spam-Flag: NO
X-Spam-Score: 0.8
X-Spam-Level: 
X-Spam-Status: No, score=0.8 tagged_above=-999 required=5 tests=[BAYES_50=0.8] autolearn=ham
Received: from mail.ietf.org ([4.31.198.44]) by localhost (ietfa.amsl.com [127.0.0.1]) (amavisd-new, port 10024) with ESMTP id H3jorqzSvVzG for <clue@ietfa.amsl.com>; Sun,  6 Apr 2014 19:17:38 -0700 (PDT)
Received: from cserver5.myshophosting.com (cserver5.myshophosting.com [175.107.161.1]) by ietfa.amsl.com (Postfix) with ESMTP id F295C1A0197 for <clue@ietf.org>; Sun,  6 Apr 2014 19:17:36 -0700 (PDT)
Received: from ppp118-209-187-61.lns20.mel6.internode.on.net ([118.209.187.61]:63417 helo=[127.0.0.1]) by cserver5.myshophosting.com with esmtpsa (TLSv1:DHE-RSA-AES128-SHA:128) (Exim 4.82) (envelope-from <Christian.Groves@nteczone.com>) id 1WWz7l-0007hM-JW; Mon, 07 Apr 2014 12:17:25 +1000
Message-ID: <53420AB8.1000106@nteczone.com>
Date: Mon, 07 Apr 2014 12:17:28 +1000
From: Christian Groves <Christian.Groves@nteczone.com>
User-Agent: Mozilla/5.0 (Windows NT 6.1; WOW64; rv:24.0) Gecko/20100101 Thunderbird/24.4.0
MIME-Version: 1.0
To: "Duckworth, Mark" <Mark.Duckworth@polycom.com>,  "clue@ietf.org" <clue@ietf.org>
References: <49E45C59CA48264997FEBFB29B6BC2D617ACF1B213@CRPMBOXPRD07.polycom.com> <533A14A0.6030807@nteczone.com> <49E45C59CA48264997FEBFB29B6BC2D62151A17223@CRPMBOXPRD07.polycom.com> <533DEE34.7070809@nteczone.com> <49E45C59CA48264997FEBFB29B6BC2D62151CDC8AC@CRPMBOXPRD07.polycom.com>
In-Reply-To: <49E45C59CA48264997FEBFB29B6BC2D62151CDC8AC@CRPMBOXPRD07.polycom.com>
Content-Type: text/plain; charset=ISO-8859-1; format=flowed
Content-Transfer-Encoding: 7bit
X-AntiAbuse: This header was added to track abuse, please include it with any abuse report
X-AntiAbuse: Primary Hostname - cserver5.myshophosting.com
X-AntiAbuse: Original Domain - ietf.org
X-AntiAbuse: Originator/Caller UID/GID - [47 12] / [47 12]
X-AntiAbuse: Sender Address Domain - nteczone.com
X-Get-Message-Sender-Via: cserver5.myshophosting.com: authenticated_id: christian.groves@nteczone.com
X-Source: 
X-Source-Args: 
X-Source-Dir: 
Archived-At: http://mailarchive.ietf.org/arch/msg/clue/ZEDc9gh-Tr1gnC10ajasCaojOzg
Subject: Re: [clue] Global CSE List in framework
X-BeenThere: clue@ietf.org
X-Mailman-Version: 2.1.15
Precedence: list
List-Id: CLUE - ControLling mUltiple streams for TElepresence <clue.ietf.org>
List-Unsubscribe: <https://www.ietf.org/mailman/options/clue>, <mailto:clue-request@ietf.org?subject=unsubscribe>
List-Archive: <http://www.ietf.org/mail-archive/web/clue/>
List-Post: <mailto:clue@ietf.org>
List-Help: <mailto:clue-request@ietf.org?subject=help>
List-Subscribe: <https://www.ietf.org/mailman/listinfo/clue>, <mailto:clue-request@ietf.org?subject=subscribe>
X-List-Received-Date: Mon, 07 Apr 2014 02:17:43 -0000

Hello Mark,

Please see below.

Regards, Christian

..snip..
>>> ...snip ...
>>>>> * The ordering of items (sets of CSEs) in the global CSE list is not
>>>>> important.
>>>>>
>>>>> * The ordering of CSEs within each set is not important.
>>>>>
>>>>> * The Provider must be capable of encoding and sending all Captures
>>>>> within the CSEs of a given set simultaneously.
>>>>>
>>>> [CNG] Do we need to add some text to the simultaneous set section of
>>>> the framework indicating this? We should probably say something about
>>>> the interaction. i.e. If a GCSE set is included should the provider
>>>> actually need to provide a STS? Sendings CSEs in a STS indicates the
>>>> those CSE/captures may be used simultaneously. There appears to be
>>>> alot of overlap with the GCSE in this respect.
>>> [Duckworth, Mark] I think I see what you mean.  We have this paragraph
>> already:
>>> "If an Advertisement does not include Simultaneous Transmission Sets,
>> then the Provider MUST be able to provide all Capture Scenes
>> simultaneously.  If multiple capture Scene Entries are in a Capture Scene then
>> the Consumer chooses at most one Capture Scene Entry per Capture Scene
>> for each media type."
>>> [Duckworth, Mark] I suggest we also add "If there is no STS and there is a
>> global CSE list, then the Consumer chooses at most one set of CSEs of each
>> media type, from the global CSE list."
>> [CNG] I'm OK with the above sentence. I think in addition we need to add
>> some text regarding what happens when an STS AND GCSE is included.
>> There's some overlap between them.
>>
>> E.g.
>> Scene1(CSE1(VC1,VC2,VC3)
>> Scene2(CSE2(VC4,VC5,VC6)
>> Scene3(CSE3(VC7,VC8,VC9)
>> GCSE([CSE1,CSE3],[CSE2,CSE3])
>> STS([CSE1,CSE3],[CSE2,CSE3])
>>
>> If the provider includes a GCSE as above and wants to include a STS does it
>> have to contain the exact same CSE sets?, i.e. because according to the
>> above bullet the CSE must capable of being send simultaneously. Or is the
>> interaction that a GCSE set indicates that at least one capture from each CSE
>> can be sent simultaneously, or?
>>
> [Duckworth, Mark] No, it shouldn't have to contain the exact same CSE sets in both.  The point is that if there are STS in the advertisement, then the STSs must make it possible for: "The Provider must be capable of encoding and sending all Captures within the CSEs of a given set (in the global CSE list) simultaneously."  In other words, the STSs must not express any restriction that prohibits the consumer from requesting all the captures in all the CSEs in a set of CSEs in the global CSE list.  But it can express restrictions that do not violate the global CSE list.
> So from your example, the GCSE list is saying the consumer can choose from these sets of captures:
>      VC1,VC2,VC3,VC7,VC8,VC9
>      or VC4,VC5,VC6,VC7,VC8,VC9
> So there could still be STSs that don't contradict that, but do express limitations.  For example:
>      STS([VC1,VC2,VC3,VC4,VC7,VC8,VC9],[VC1,VC4,VC5,VC6,VC7,VC8,VC9])
> So with these STSs, the consumer could choose everything from the first set in the GCSE list (VC1,VC2,VC3,VC7,VC8,VC9) plus also VC4 if it wants.  But not VC5 or VC6.

>
> [Duckworth, Mark] Do you have a suggestion for how to make this more clear in the framework?
[CNG] I think we agree. Using "exact" probably wasn't the best word, I 
probably should have used "at least".
I think what we're trying to say is: "If an STS is used then it a 
minimum MUST contain captures that reflect the simultaneity expressed by 
any global CSE sets."

Not for the framework, but we should probably make a mental note that we 
include an error code/case for if the STSs, CSEs and GCSEs don't match. 
I would assume that the consumer would return an error, unless people 
can think of another handling?
>
>


From nobody Sun Apr  6 19:21:24 2014
Return-Path: <Christian.Groves@nteczone.com>
X-Original-To: clue@ietfa.amsl.com
Delivered-To: clue@ietfa.amsl.com
Received: from localhost (ietfa.amsl.com [127.0.0.1]) by ietfa.amsl.com (Postfix) with ESMTP id 12C071A064F for <clue@ietfa.amsl.com>; Sun,  6 Apr 2014 19:21:23 -0700 (PDT)
X-Virus-Scanned: amavisd-new at amsl.com
X-Spam-Flag: NO
X-Spam-Score: -0.001
X-Spam-Level: 
X-Spam-Status: No, score=-0.001 tagged_above=-999 required=5 tests=[BAYES_20=-0.001] autolearn=ham
Received: from mail.ietf.org ([4.31.198.44]) by localhost (ietfa.amsl.com [127.0.0.1]) (amavisd-new, port 10024) with ESMTP id V84cOOP6hxj9 for <clue@ietfa.amsl.com>; Sun,  6 Apr 2014 19:21:18 -0700 (PDT)
Received: from cserver5.myshophosting.com (cserver5.myshophosting.com [175.107.161.1]) by ietfa.amsl.com (Postfix) with ESMTP id 903511A0329 for <clue@ietf.org>; Sun,  6 Apr 2014 19:21:18 -0700 (PDT)
Received: from ppp118-209-187-61.lns20.mel6.internode.on.net ([118.209.187.61]:63453 helo=[127.0.0.1]) by cserver5.myshophosting.com with esmtpsa (TLSv1:DHE-RSA-AES128-SHA:128) (Exim 4.82) (envelope-from <Christian.Groves@nteczone.com>) id 1WWzBM-0008IA-4Q for clue@ietf.org; Mon, 07 Apr 2014 12:21:08 +1000
Message-ID: <53420B96.3010603@nteczone.com>
Date: Mon, 07 Apr 2014 12:21:10 +1000
From: Christian Groves <Christian.Groves@nteczone.com>
User-Agent: Mozilla/5.0 (Windows NT 6.1; WOW64; rv:24.0) Gecko/20100101 Thunderbird/24.4.0
MIME-Version: 1.0
To: clue@ietf.org
References: <5318809E.2030204@nteczone.com> <5318AE5E.4050404@alum.mit.edu> <5318B0C9.1050603@nteczone.com> <5318B48E.3090300@alum.mit.edu> <53290C9B.6090106@nteczone.com> <5329F3FF.7090606@alum.mit.edu> <533A16BA.40707@nteczone.com> <533A91BF.7040404@unina.it> <533B47FF.5020407@nteczone.com> <533D7E8F.4000303@alum.mit.edu> <49E45C59CA48264997FEBFB29B6BC2D6215264287E@CRPMBOXPRD07.polycom.com>
In-Reply-To: <49E45C59CA48264997FEBFB29B6BC2D6215264287E@CRPMBOXPRD07.polycom.com>
Content-Type: text/plain; charset=ISO-8859-1; format=flowed
Content-Transfer-Encoding: 7bit
X-AntiAbuse: This header was added to track abuse, please include it with any abuse report
X-AntiAbuse: Primary Hostname - cserver5.myshophosting.com
X-AntiAbuse: Original Domain - ietf.org
X-AntiAbuse: Originator/Caller UID/GID - [47 12] / [47 12]
X-AntiAbuse: Sender Address Domain - nteczone.com
X-Get-Message-Sender-Via: cserver5.myshophosting.com: authenticated_id: christian.groves@nteczone.com
X-Source: 
X-Source-Args: 
X-Source-Dir: 
Archived-At: http://mailarchive.ietf.org/arch/msg/clue/viJfwtukAJWg4Rlmg2naSRzmTXM
Subject: Re: [clue] Participant info/type followup
X-BeenThere: clue@ietf.org
X-Mailman-Version: 2.1.15
Precedence: list
List-Id: CLUE - ControLling mUltiple streams for TElepresence <clue.ietf.org>
List-Unsubscribe: <https://www.ietf.org/mailman/options/clue>, <mailto:clue-request@ietf.org?subject=unsubscribe>
List-Archive: <http://www.ietf.org/mail-archive/web/clue/>
List-Post: <mailto:clue@ietf.org>
List-Help: <mailto:clue-request@ietf.org?subject=help>
List-Subscribe: <https://www.ietf.org/mailman/listinfo/clue>, <mailto:clue-request@ietf.org?subject=subscribe>
X-List-Received-Date: Mon, 07 Apr 2014 02:21:23 -0000

Hello Mark,

Please see below.

Regards, Christian

On 5/04/2014 2:46 AM, Duckworth, Mark wrote:
>> -----Original Message-----
>> From: clue [mailto:clue-bounces@ietf.org] On Behalf Of Paul Kyzivat
>> Sent: Thursday, April 03, 2014 11:30 AM
>> To: clue@ietf.org
>> Subject: Re: [clue] Participant info/type followup
>>
>> ISTM we have both a semantic issue and a syntactic/terminological issue:
>>
>> - we have the semantic distinction between who is represented within
>>     the capture and who is responsible for sending the capture
> [Duckworth, Mark] I agree, but I thought the intent of the participant attributes was to give information about "who is represented within the capture" and not about "who is responsible for sending the capture".  From the framework:
>
>      "... information regarding the conference participants *in* a Capture."
>      "... indicates the type of participant/s contained *in* the capture..."
>
> [Duckworth, Mark] I think we don't need to add a way to indicate "who is responsible for sending the capture".  IIRC, this was not a goal of the original proposal and discussion about "roles".  But if the group agrees we should add this now, I wouldn't object too strongly.
[CNG] I agree when I proposed it I was thinking about "who was in the 
capture" rather than who was sending the capture. The issue only came to 
life when Jonathon asked what it meant for a presentation capture. I'd 
be OK to drop "who sends the capture" but we would still need to clarify 
the presentation capture case.



From nobody Mon Apr  7 06:34:50 2014
Return-Path: <Mark.Duckworth@polycom.com>
X-Original-To: clue@ietfa.amsl.com
Delivered-To: clue@ietfa.amsl.com
Received: from localhost (ietfa.amsl.com [127.0.0.1]) by ietfa.amsl.com (Postfix) with ESMTP id BE3F41A0789 for <clue@ietfa.amsl.com>; Mon,  7 Apr 2014 06:34:46 -0700 (PDT)
X-Virus-Scanned: amavisd-new at amsl.com
X-Spam-Flag: NO
X-Spam-Score: -1.12
X-Spam-Level: 
X-Spam-Status: No, score=-1.12 tagged_above=-999 required=5 tests=[BAYES_00=-1.9, RCVD_IN_DNSWL_NONE=-0.0001, SPF_NEUTRAL=0.779, UNPARSEABLE_RELAY=0.001] autolearn=no
Received: from mail.ietf.org ([4.31.198.44]) by localhost (ietfa.amsl.com [127.0.0.1]) (amavisd-new, port 10024) with ESMTP id DRFQoEW8Y3BQ for <clue@ietfa.amsl.com>; Mon,  7 Apr 2014 06:34:42 -0700 (PDT)
Received: from mail1.bemta8.messagelabs.com (mail1.bemta8.messagelabs.com [216.82.243.202]) by ietfa.amsl.com (Postfix) with ESMTP id 492381A0782 for <clue@ietf.org>; Mon,  7 Apr 2014 06:32:49 -0700 (PDT)
Received: from [216.82.241.100:23545] by server-10.bemta-8.messagelabs.com id 59/04-27324-BF8A2435; Mon, 07 Apr 2014 13:32:43 +0000
X-Env-Sender: Mark.Duckworth@polycom.com
X-Msg-Ref: server-8.tower-220.messagelabs.com!1396877559!4967161!1
X-Originating-IP: [140.242.64.158]
X-StarScan-Received: 
X-StarScan-Version: 6.11.1; banners=-,-,-
X-VirusChecked: Checked
Received: (qmail 27988 invoked from network); 7 Apr 2014 13:32:40 -0000
Received: from crpehubprd01.polycom.com (HELO Crpehubprd01.polycom.com) (140.242.64.158) by server-8.tower-220.messagelabs.com with AES128-SHA encrypted SMTP; 7 Apr 2014 13:32:40 -0000
Received: from CRPMBOXPRD07.polycom.com ([fe80::91fc:8a0f:5258:aff0]) by Crpehubprd01.polycom.com ([fe80::5efe:10.236.0.158%14]) with mapi; Mon, 7 Apr 2014 06:32:38 -0700
From: "Duckworth, Mark" <Mark.Duckworth@polycom.com>
To: Christian Groves <Christian.Groves@nteczone.com>, "clue@ietf.org" <clue@ietf.org>
Date: Mon, 7 Apr 2014 06:32:36 -0700
Thread-Topic: [clue] Global CSE List in framework
Thread-Index: Ac9SB4oGZDX3wh19RP+V9yzDY1N5bQAXiJTQ
Message-ID: <49E45C59CA48264997FEBFB29B6BC2D62152642AE9@CRPMBOXPRD07.polycom.com>
References: <49E45C59CA48264997FEBFB29B6BC2D617ACF1B213@CRPMBOXPRD07.polycom.com> <533A14A0.6030807@nteczone.com> <49E45C59CA48264997FEBFB29B6BC2D62151A17223@CRPMBOXPRD07.polycom.com> <533DEE34.7070809@nteczone.com> <49E45C59CA48264997FEBFB29B6BC2D62151CDC8AC@CRPMBOXPRD07.polycom.com> <53420AB8.1000106@nteczone.com>
In-Reply-To: <53420AB8.1000106@nteczone.com>
Accept-Language: en-US
Content-Language: en-US
X-MS-Has-Attach: 
X-MS-TNEF-Correlator: 
acceptlanguage: en-US
Content-Type: text/plain; charset="us-ascii"
Content-Transfer-Encoding: quoted-printable
MIME-Version: 1.0
Archived-At: http://mailarchive.ietf.org/arch/msg/clue/K3wLyqf5RSnb7ncpYHgkxvaN8qE
Subject: Re: [clue] Global CSE List in framework
X-BeenThere: clue@ietf.org
X-Mailman-Version: 2.1.15
Precedence: list
List-Id: CLUE - ControLling mUltiple streams for TElepresence <clue.ietf.org>
List-Unsubscribe: <https://www.ietf.org/mailman/options/clue>, <mailto:clue-request@ietf.org?subject=unsubscribe>
List-Archive: <http://www.ietf.org/mail-archive/web/clue/>
List-Post: <mailto:clue@ietf.org>
List-Help: <mailto:clue-request@ietf.org?subject=help>
List-Subscribe: <https://www.ietf.org/mailman/listinfo/clue>, <mailto:clue-request@ietf.org?subject=subscribe>
X-List-Received-Date: Mon, 07 Apr 2014 13:34:47 -0000

Thanks Christian.  I will collect all the clarifications from this conversa=
tion and add them to the next version of the framework.
Mark

> -----Original Message-----
> From: Christian Groves [mailto:Christian.Groves@nteczone.com]
> Sent: Sunday, April 06, 2014 10:17 PM
> To: Duckworth, Mark; clue@ietf.org
> Subject: Re: [clue] Global CSE List in framework
>=20
> Hello Mark,
>=20
> Please see below.
>=20
> Regards, Christian
>=20
> ..snip..
> >>> ...snip ...
> >>>>> * The ordering of items (sets of CSEs) in the global CSE list is
> >>>>> not important.
> >>>>>
> >>>>> * The ordering of CSEs within each set is not important.
> >>>>>
> >>>>> * The Provider must be capable of encoding and sending all
> >>>>> Captures within the CSEs of a given set simultaneously.
> >>>>>
> >>>> [CNG] Do we need to add some text to the simultaneous set section
> >>>> of the framework indicating this? We should probably say something
> >>>> about the interaction. i.e. If a GCSE set is included should the
> >>>> provider actually need to provide a STS? Sendings CSEs in a STS
> >>>> indicates the those CSE/captures may be used simultaneously. There
> >>>> appears to be alot of overlap with the GCSE in this respect.
> >>> [Duckworth, Mark] I think I see what you mean.  We have this
> >>> paragraph
> >> already:
> >>> "If an Advertisement does not include Simultaneous Transmission
> >>> Sets,
> >> then the Provider MUST be able to provide all Capture Scenes
> >> simultaneously.  If multiple capture Scene Entries are in a Capture
> >> Scene then the Consumer chooses at most one Capture Scene Entry per
> >> Capture Scene for each media type."
> >>> [Duckworth, Mark] I suggest we also add "If there is no STS and
> >>> there is a
> >> global CSE list, then the Consumer chooses at most one set of CSEs of
> >> each media type, from the global CSE list."
> >> [CNG] I'm OK with the above sentence. I think in addition we need to
> >> add some text regarding what happens when an STS AND GCSE is
> included.
> >> There's some overlap between them.
> >>
> >> E.g.
> >> Scene1(CSE1(VC1,VC2,VC3)
> >> Scene2(CSE2(VC4,VC5,VC6)
> >> Scene3(CSE3(VC7,VC8,VC9)
> >> GCSE([CSE1,CSE3],[CSE2,CSE3])
> >> STS([CSE1,CSE3],[CSE2,CSE3])
> >>
> >> If the provider includes a GCSE as above and wants to include a STS
> >> does it have to contain the exact same CSE sets?, i.e. because
> >> according to the above bullet the CSE must capable of being send
> >> simultaneously. Or is the interaction that a GCSE set indicates that
> >> at least one capture from each CSE can be sent simultaneously, or?
> >>
> > [Duckworth, Mark] No, it shouldn't have to contain the exact same CSE s=
ets
> in both.  The point is that if there are STS in the advertisement, then t=
he STSs
> must make it possible for: "The Provider must be capable of encoding and
> sending all Captures within the CSEs of a given set (in the global CSE li=
st)
> simultaneously."  In other words, the STSs must not express any restricti=
on
> that prohibits the consumer from requesting all the captures in all the C=
SEs in
> a set of CSEs in the global CSE list.  But it can express restrictions th=
at do not
> violate the global CSE list.
> > So from your example, the GCSE list is saying the consumer can choose
> from these sets of captures:
> >      VC1,VC2,VC3,VC7,VC8,VC9
> >      or VC4,VC5,VC6,VC7,VC8,VC9
> > So there could still be STSs that don't contradict that, but do express
> limitations.  For example:
> >      STS([VC1,VC2,VC3,VC4,VC7,VC8,VC9],[VC1,VC4,VC5,VC6,VC7,VC8,VC9])
> > So with these STSs, the consumer could choose everything from the first
> set in the GCSE list (VC1,VC2,VC3,VC7,VC8,VC9) plus also VC4 if it wants.=
  But
> not VC5 or VC6.
>=20
> >
> > [Duckworth, Mark] Do you have a suggestion for how to make this more
> clear in the framework?
> [CNG] I think we agree. Using "exact" probably wasn't the best word, I
> probably should have used "at least".
> I think what we're trying to say is: "If an STS is used then it a minimum=
 MUST
> contain captures that reflect the simultaneity expressed by any global CS=
E
> sets."
>=20
> Not for the framework, but we should probably make a mental note that we
> include an error code/case for if the STSs, CSEs and GCSEs don't match.
> I would assume that the consumer would return an error, unless people can
> think of another handling?
> >
> >


From nobody Tue Apr  8 08:06:24 2014
Return-Path: <ron.even.tlv@gmail.com>
X-Original-To: clue@ietfa.amsl.com
Delivered-To: clue@ietfa.amsl.com
Received: from localhost (ietfa.amsl.com [127.0.0.1]) by ietfa.amsl.com (Postfix) with ESMTP id A05131A044B for <clue@ietfa.amsl.com>; Tue,  8 Apr 2014 08:06:22 -0700 (PDT)
X-Virus-Scanned: amavisd-new at amsl.com
X-Spam-Flag: NO
X-Spam-Score: -2
X-Spam-Level: 
X-Spam-Status: No, score=-2 tagged_above=-999 required=5 tests=[BAYES_00=-1.9,  DKIM_SIGNED=0.1, DKIM_VALID=-0.1, DKIM_VALID_AU=-0.1, FREEMAIL_FROM=0.001, SPF_PASS=-0.001] autolearn=ham
Received: from mail.ietf.org ([4.31.198.44]) by localhost (ietfa.amsl.com [127.0.0.1]) (amavisd-new, port 10024) with ESMTP id ms8QekqJ9m-9 for <clue@ietfa.amsl.com>; Tue,  8 Apr 2014 08:06:18 -0700 (PDT)
Received: from mail-ee0-x229.google.com (mail-ee0-x229.google.com [IPv6:2a00:1450:4013:c00::229]) by ietfa.amsl.com (Postfix) with ESMTP id 7AA681A043F for <clue@ietf.org>; Tue,  8 Apr 2014 08:06:17 -0700 (PDT)
Received: by mail-ee0-f41.google.com with SMTP id t10so813022eei.28 for <clue@ietf.org>; Tue, 08 Apr 2014 08:06:16 -0700 (PDT)
DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20120113;  h=from:to:cc:references:in-reply-to:subject:date:message-id :mime-version:content-type:content-transfer-encoding:thread-index :content-language; bh=gHdjEcIFHzKN0WTe49SVa+n5R+ZSw53bJkoHC7yXKK0=; b=rmWOK8X3ZzD4jVcIQrFG4pguyu5uPaV/PXphuuUj+SjwqqmM5vxq1AFNduh5P3b9QH MYYhgJv4wmQ0xsnLbvjggKtBu53TEFlOp2XqXJo3Lv4TV+SQSBEUGPgMIj0FAueTSKRe go/z3Rnt01WCaNpA+DbUleqSE6UixRnVo25GktoIciowx6pQR4ud1jUKmQIRv+3StwHR o0uPxBXFRj0mO/KdPKPsQyM86fg+FcKfuqmAWG91n4qaNJX2j17IbpDIzQ2C48X5bHbS E2fsZf95/vyg6L3T2W5S0kwRMAvVPgIAdQ9nF5LOjGEW50b2pRkGB+4CimkcI5HxPlz0 56nw==
X-Received: by 10.14.184.66 with SMTP id r42mr2800936eem.84.1396969576712; Tue, 08 Apr 2014 08:06:16 -0700 (PDT)
Received: from RoniE (bzq-79-183-174-212.red.bezeqint.net. [79.183.174.212]) by mx.google.com with ESMTPSA id q41sm5161204eez.7.2014.04.08.08.06.14 for <multiple recipients> (version=TLSv1 cipher=ECDHE-RSA-AES128-SHA bits=128/128); Tue, 08 Apr 2014 08:06:16 -0700 (PDT)
From: "Roni Even" <ron.even.tlv@gmail.com>
To: "'Paul Kyzivat'" <pkyzivat@alum.mit.edu>, "'Jonathan Lennox'" <jonathan@vidyo.com>
References: <533E7A50.5040909@ericsson.com> <533F1AD1.7060507@alum.mit.edu>
In-Reply-To: <533F1AD1.7060507@alum.mit.edu>
Date: Tue, 8 Apr 2014 18:06:12 +0300
Message-ID: <010401cf533c$161b6130$42522390$@gmail.com>
MIME-Version: 1.0
Content-Type: text/plain; charset="ISO-8859-1"
Content-Transfer-Encoding: quoted-printable
X-Mailer: Microsoft Outlook 14.0
Thread-Index: AQKAdRtyny0O/9Ki8/9GVI+Yi+W4VAJvCvqgmZIVb3A=
Content-Language: en-us
Archived-At: http://mailarchive.ietf.org/arch/msg/clue/hy4DAT3gGlN0BuDJhlUvyMK-Uos
Cc: 'CLUE' <clue@ietf.org>
Subject: Re: [clue] [rtcweb] Draft proposal for updating Multiparty topologies in draft-ietf-rtcweb-rtp-usage
X-BeenThere: clue@ietf.org
X-Mailman-Version: 2.1.15
Precedence: list
List-Id: CLUE - ControLling mUltiple streams for TElepresence <clue.ietf.org>
List-Unsubscribe: <https://www.ietf.org/mailman/options/clue>, <mailto:clue-request@ietf.org?subject=unsubscribe>
List-Archive: <http://www.ietf.org/mail-archive/web/clue/>
List-Post: <mailto:clue@ietf.org>
List-Help: <mailto:clue-request@ietf.org?subject=help>
List-Subscribe: <https://www.ietf.org/mailman/listinfo/clue>, <mailto:clue-request@ietf.org?subject=subscribe>
X-List-Received-Date: Tue, 08 Apr 2014 15:06:23 -0000

Hi Paul,
I responded on the RTCweb mailing list. I agree with Harald comment in
RTCweb.
Roni

> -----Original Message-----
> From: Paul Kyzivat [mailto:pkyzivat@alum.mit.edu]
> Sent: 04 April, 2014 11:49 PM
> To: Roni Even; Jonathan Lennox
> Cc: CLUE
> Subject: Re: [rtcweb] Draft proposal for updating Multiparty =
topologies in
draft-
> ietf-rtcweb-rtp-usage
>=20
> Roni & Jonathan,
>=20
> Can one of you comment on whether any of this has an impact on CLUE
> interop?
>=20
> 	Thanks,
> 	Paul
>=20
> On 4/4/14 5:24 AM, Magnus Westerlund wrote:
> > WG,
> > (As Author)
> >
> > Colin and I have been working on resolving terminology usage and in
> > general giving the draft a polish over before the WG last call. One
> > section that has gotten quite some attention from us is the below =
one.
> > The changes are significantly from an editorial stand point. =
However,
> > the intended content has not been intended to be changed, although
> > clarified and better motivated. So please review it, we intended to
> > include this in the upcoming draft update. Feedback appreciated.
> >
> > 5.1.  Conferencing Extensions and Topologies
> >
> >     RTP is a protocol that inherently supports group communication.
> >     Groups can be implemented by having each endpoint sending its =
RTP
> >     packet streams to an RTP middlebox that redistributes the =
traffic,
by
> >     using a mesh of unicast RTP packet streams between endpoints, or =
by
> >     using an IP multicast group to distribute the RTP packet =
streams.
> >     These topologies can be implemented in a number of ways as =
discussed
> >     in [I-D.ietf-avtcore-rtp-topologies-update].
> >
> >     While the use of IP multicast groups is popular in IPTV systems, =
the
> >     topologies based on RTP middleboxes are dominant in interactive
video
> >     conferencing environments.  Topologies based on a mesh of =
unicast
> >     transport-layer flows to create a common RTP session have not =
seen
> >     widespread deployment.  Accordingly, WebRTC implementations are =
not
> >     expected to support topologies based on IP multicast groups.  =
WebRTC
> >     implementations are also not expected to support mesh-based
> >     topologies, such as a point-to-multipoint mesh configured as a
single
> >     RTP session (Topo-Mesh in the terminology of
> >     [I-D.ietf-avtcore-rtp-topologies-update]).  However, a point-to-
> >     multipoint mesh constructed using several RTP sessions, in the
WebRTC
> >     context using independent RTCPeerConnections can be expected to =
be
> >     utilised by WebRTC applications.
> >
> >     WebRTC implementations of RTP endpoints implemented according to
this
> >     memo are expected to support all the topologies described in
> >     [I-D.ietf-avtcore-rtp-topologies-update] where the RTP endpoints
send
> >     and receive unicast RTP packet streams to some peer device, =
provided
> >     that peer can participate in performing congestion control on =
the
RTP
> >     packet streams.  The peer device could be another RTP endpoint, =
or
it
> >     could be an RTP middlebox that redistributes the RTP packet =
streams
> >     to other RTP endpoints.  This limitation means that some of the =
RTP
> >     middlebox-based topologies are not suitable for use in the =
WebRTC
> >     environment.  Specifically:
> >
> >     o  Video switching MCUs (Topo-Video-switch-MCU) SHOULD NOT be =
used,
> >        since they make the use of RTCP for congestion control and
quality
> >        of service reports problematic (see Section 3.6.2 of
> >        [I-D.ietf-avtcore-rtp-topologies-update]).
> >
> >     o  Content modifying MCUs with RTCP termination (Topo-RTCP-
> >        terminating-MCU) SHOULD NOT be used since they break RTP loop
> >        detection, and prevent receivers from identifying active =
senders
> >        (see section 3.8 of =
[I-D.ietf-avtcore-rtp-topologies-update]).
> >
> >     o  The Relay-Transport Translator (Topo-PtM-Trn-Translator) =
topology
> >        SHOULD NOT be used because its safe use requires a point to
> >        multipoint congestion control algorithm or RTP circuit =
breaker,
> >        which has not yet been standardised.
> >
> >     The RTP extensions described in Section 5.1.1 to Section 5.1.6 =
are
> >     designed to be used with centralised conferencing, where an RTP
> >     middlebox (e.g., a conference bridge) receives a participant's =
RTP
> >     packet streams and distributes them to the other participants.
These
> >     extensions are not necessary for interoperability; an RTP =
end-point
> >     that does not implement these extensions will work correctly, =
but
> >     might offer poor performance.  Support for the listed extensions
will
> >     greatly improve the quality of experience and, to provide a
> >     reasonable baseline quality, some of these extensions are =
mandatory
> >     to be supported by WebRTC end-points.
> >
> >     The RTCP conferencing extensions are defined in Extended RTP =
Profile
> >     for Real-time Transport Control Protocol (RTCP)-Based Feedback =
(RTP/
> >     AVPF) [RFC4585] and the "Codec Control Messages in the RTP =
Audio-
> >     Visual Profile with Feedback (AVPF)" (CCM) [RFC5104] and are =
fully
> >     usable by the Secure variant of this profile (RTP/SAVPF) =
[RFC5124].
> >
> >
> > Cheers
> >
> > Magnus Westerlund
> >
> > =
----------------------------------------------------------------------
> > Services, Media and Network features, Ericsson Research EAB/TXM
> > =
----------------------------------------------------------------------
> > Ericsson AB                 | Phone  +46 10 7148287
> > F=E4r=F6gatan 6                 | Mobile +46 73 0949079
> > SE-164 80 Stockholm, Sweden | mailto: magnus.westerlund@ericsson.com
> > =
----------------------------------------------------------------------
> >
> > _______________________________________________
> > rtcweb mailing list
> > rtcweb@ietf.org
> > https://www.ietf.org/mailman/listinfo/rtcweb
> >


From nobody Tue Apr  8 08:27:40 2014
Return-Path: <pkyzivat@alum.mit.edu>
X-Original-To: clue@ietfa.amsl.com
Delivered-To: clue@ietfa.amsl.com
Received: from localhost (ietfa.amsl.com [127.0.0.1]) by ietfa.amsl.com (Postfix) with ESMTP id 3BA4C1A043A for <clue@ietfa.amsl.com>; Tue,  8 Apr 2014 08:27:38 -0700 (PDT)
X-Virus-Scanned: amavisd-new at amsl.com
X-Spam-Flag: NO
X-Spam-Score: 2.665
X-Spam-Level: **
X-Spam-Status: No, score=2.665 tagged_above=-999 required=5 tests=[BAYES_50=0.8, DKIM_SIGNED=0.1, DKIM_VALID=-0.1, J_CHICKENPOX_110=0.6, J_CHICKENPOX_15=0.6, SPF_SOFTFAIL=0.665] autolearn=no
Received: from mail.ietf.org ([4.31.198.44]) by localhost (ietfa.amsl.com [127.0.0.1]) (amavisd-new, port 10024) with ESMTP id TWBzcS17JElM for <clue@ietfa.amsl.com>; Tue,  8 Apr 2014 08:27:37 -0700 (PDT)
Received: from qmta05.westchester.pa.mail.comcast.net (qmta05.westchester.pa.mail.comcast.net [IPv6:2001:558:fe14:43:76:96:62:48]) by ietfa.amsl.com (Postfix) with ESMTP id 42C281A0433 for <clue@ietf.org>; Tue,  8 Apr 2014 08:27:36 -0700 (PDT)
Received: from omta08.westchester.pa.mail.comcast.net ([76.96.62.12]) by qmta05.westchester.pa.mail.comcast.net with comcast id nPBk1n0020Fqzac55TTcdF; Tue, 08 Apr 2014 15:27:36 +0000
Received: from Paul-Kyzivats-MacBook-Pro.local ([50.138.229.164]) by omta08.westchester.pa.mail.comcast.net with comcast id nTTc1n00Z3ZTu2S3UTTczq; Tue, 08 Apr 2014 15:27:36 +0000
Message-ID: <53441568.6090704@alum.mit.edu>
Date: Tue, 08 Apr 2014 11:27:36 -0400
From: Paul Kyzivat <pkyzivat@alum.mit.edu>
User-Agent: Mozilla/5.0 (Macintosh; Intel Mac OS X 10.7; rv:24.0) Gecko/20100101 Thunderbird/24.4.0
MIME-Version: 1.0
To: CLUE <clue@ietf.org>
Content-Type: text/plain; charset=ISO-8859-1; format=flowed
Content-Transfer-Encoding: 7bit
DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=comcast.net; s=q20140121; t=1396970856; bh=RcA2FdMwuB7WcT/slQEml4DiL1vKuS0P+JjR6S7/baY=; h=Received:Received:Message-ID:Date:From:MIME-Version:To:Subject: Content-Type; b=Ft8jphp7GIHXr9voKDoRrQ66G8reETwWkFEMKFPgk8gjnTqZTSrqf0oFQYj8gHU+4 VeZ5odFwiP0CqiVkiukuPBnkPaLKtIy0wYVPdiArnki1TObx0MEE1Hzshp9l+PWPMs DTCO3fG1eRz2Q28GyFoUabpIH051AT3TntZE9T1xi7hLxDz+dmu6HC8Gu7vd/zMNEz AWBDEU6mPloW+Mt1J9tjAbvMC8m1ztp5jztXZUBDrnJl0E/Rkpw7gwl/bZKRhyJtjG gwUAnwLdn3LIRtb2jAFXTqYoQkYOY8PcBNtOijO7cSpu3fabkfVr9hfeC3a/6wUrbl KEhix528hONKA==
Archived-At: http://mailarchive.ietf.org/arch/msg/clue/sQdrUMIkY3UHwHAvAYxUzo7o1bI
Subject: [clue] decisions and issues re establishment of clue channel
X-BeenThere: clue@ietf.org
X-Mailman-Version: 2.1.15
Precedence: list
List-Id: CLUE - ControLling mUltiple streams for TElepresence <clue.ietf.org>
List-Unsubscribe: <https://www.ietf.org/mailman/options/clue>, <mailto:clue-request@ietf.org?subject=unsubscribe>
List-Archive: <http://www.ietf.org/mail-archive/web/clue/>
List-Post: <mailto:clue@ietf.org>
List-Help: <mailto:clue-request@ietf.org?subject=help>
List-Subscribe: <https://www.ietf.org/mailman/listinfo/clue>, <mailto:clue-request@ietf.org?subject=subscribe>
X-List-Received-Date: Tue, 08 Apr 2014 15:27:38 -0000

In the design team call today we got into a discussion about how to 
specify which end of the sip session is responsible for opening the clue 
channel. The following is a summary of what I think we agreed:

- There was a proposal that it is an error to have a clue group that
   doesn't include an sctp m-line. I think there was general agreement
   with this.

- The end that creates the sctp association - the end that is
   *active* according to RFC 4572/4145 - will be responsible for
   the initial open, and for reopening after a reset as long as
   association is working and is part of the clue group.

- If the association is reestablished (signaled by a=connection:new)
   then the end that is responsible is redefined at that time by the
   end that is active for establishing the new association.

- we didn't discuss a case where the sctp association is first
   established but is not part of a clue group, and is later added
   to the clue group. IMO, in this case the responsibility should
   still be the end that was active for establishment of the association.

- another we didn't discuss is what happens if there is clue channel
   open and then the sctp association is removed from the clue group.
   (This means that there should be no clue group, so indicates a
   change to legacy mode.) It is possible that the sctp m-line is
   still in the SDP, just not in the clue group, and is still functional.
   I think, in this case that what happens to the clue channel is
   undefined. But if it is reset, then the active end of the association
   is *not* responsible for opening it again.

Based on this, we have some issues to bring to MMUSIC for clarification:

- draft-ietf-mmusic-sdp-mux-attributes specifies that a=setup
   and a=connection must be identical across all bundled m-lines.

- draft-ietf-mmusic-sctp-sdp specifies use of a=setup and
   a=connection as specified in rfc4145.

- RFC 5763 (section 5) forbids use of a=connection for
   UDP/TLS/RTP/SAVP.

This is a conflict. It appears that draft-ietf-mmusic-sdp-mux-attributes 
needs to be revised. Perhaps it can say that a=connection must be 
identical across all bundled m-lines *that permit the use of a=connection.*

And then we have a general concern over the progress of 
draft-ietf-mmusic-sctp-sdp.

	Thanks,
	Paul


From nobody Wed Apr  9 09:10:08 2014
Return-Path: <rohanse2@cisco.com>
X-Original-To: clue@ietfa.amsl.com
Delivered-To: clue@ietfa.amsl.com
Received: from localhost (ietfa.amsl.com [127.0.0.1]) by ietfa.amsl.com (Postfix) with ESMTP id C61F31A03BE for <clue@ietfa.amsl.com>; Wed,  9 Apr 2014 09:10:04 -0700 (PDT)
X-Virus-Scanned: amavisd-new at amsl.com
X-Spam-Flag: NO
X-Spam-Score: -14.772
X-Spam-Level: 
X-Spam-Status: No, score=-14.772 tagged_above=-999 required=5 tests=[BAYES_00=-1.9, DKIM_SIGNED=0.1, DKIM_VALID=-0.1, DKIM_VALID_AU=-0.1, HTML_MESSAGE=0.001, RCVD_IN_DNSWL_HI=-5, RP_MATCHES_RCVD=-0.272, SPF_PASS=-0.001, USER_IN_DEF_DKIM_WL=-7.5] autolearn=ham
Received: from mail.ietf.org ([4.31.198.44]) by localhost (ietfa.amsl.com [127.0.0.1]) (amavisd-new, port 10024) with ESMTP id MJRJXacjpXgg for <clue@ietfa.amsl.com>; Wed,  9 Apr 2014 09:10:03 -0700 (PDT)
Received: from rcdn-iport-2.cisco.com (rcdn-iport-2.cisco.com [173.37.86.73]) by ietfa.amsl.com (Postfix) with ESMTP id ECD291A03BD for <clue@ietf.org>; Wed,  9 Apr 2014 09:10:02 -0700 (PDT)
DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/simple; d=cisco.com; i=@cisco.com; l=2660; q=dns/txt; s=iport; t=1397059802; x=1398269402; h=from:to:subject:date:message-id:mime-version; bh=MxB32FROA0JjNdn5YxhnvR/DMd63xjaDlwLIp+y0uJU=; b=nHH/ZlLTs65XqpDqIEKpBHC6OX03Fxm35jaHsggjW6P3IG76DxpqXx6U 6kd/ok55D2RvhbS7kMP2CDQ96nU5JRV0tyoCDZsycVxIUSPnvyN6MHu4O SkArUYGhB9D/E4KJ4N5bX6JedfhXufy2dZe3gNFjVsbIsRS2vRTHBFtFU Q=;
X-IronPort-Anti-Spam-Filtered: true
X-IronPort-Anti-Spam-Result: AhwFALxvRVOtJV2b/2dsb2JhbABZgkJEO1fEDYEgFnSCJwEELV4BDB5WJgEEG4d0myqxMheOChEBH4NcgRQEqyCDMIFyOQ
X-IronPort-AV: E=Sophos;i="4.97,827,1389744000";  d="scan'208,217";a="316393397"
Received: from rcdn-core-4.cisco.com ([173.37.93.155]) by rcdn-iport-2.cisco.com with ESMTP; 09 Apr 2014 16:10:02 +0000
Received: from xhc-aln-x08.cisco.com (xhc-aln-x08.cisco.com [173.36.12.82]) by rcdn-core-4.cisco.com (8.14.5/8.14.5) with ESMTP id s39GA1Lw026506 (version=TLSv1/SSLv3 cipher=AES128-SHA bits=128 verify=FAIL) for <clue@ietf.org>; Wed, 9 Apr 2014 16:10:01 GMT
Received: from xmb-aln-x07.cisco.com ([169.254.2.162]) by xhc-aln-x08.cisco.com ([173.36.12.82]) with mapi id 14.03.0123.003; Wed, 9 Apr 2014 11:10:01 -0500
From: "Robert Hansen (rohanse2)" <rohanse2@cisco.com>
To: "clue@ietf.org" <clue@ietf.org>
Thread-Topic: Signalling updates
Thread-Index: Ac9UDigOJS8cB6DNQLCgxfxNf8/mPA==
Date: Wed, 9 Apr 2014 16:10:01 +0000
Message-ID: <C6252EA94E00E44EADC3A2FEB59D4402020F54D5@xmb-aln-x07.cisco.com>
Accept-Language: en-GB, en-US
Content-Language: en-US
X-MS-Has-Attach: 
X-MS-TNEF-Correlator: 
x-originating-ip: [10.21.124.108]
Content-Type: multipart/alternative; boundary="_000_C6252EA94E00E44EADC3A2FEB59D4402020F54D5xmbalnx07ciscoc_"
MIME-Version: 1.0
Archived-At: http://mailarchive.ietf.org/arch/msg/clue/bVbhJvF90WbidybX92CqfU4yPlQ
Subject: [clue] Signalling updates
X-BeenThere: clue@ietf.org
X-Mailman-Version: 2.1.15
Precedence: list
List-Id: CLUE - ControLling mUltiple streams for TElepresence <clue.ietf.org>
List-Unsubscribe: <https://www.ietf.org/mailman/options/clue>, <mailto:clue-request@ietf.org?subject=unsubscribe>
List-Archive: <http://www.ietf.org/mail-archive/web/clue/>
List-Post: <mailto:clue@ietf.org>
List-Help: <mailto:clue-request@ietf.org?subject=help>
List-Subscribe: <https://www.ietf.org/mailman/listinfo/clue>, <mailto:clue-request@ietf.org?subject=subscribe>
X-List-Received-Date: Wed, 09 Apr 2014 16:10:04 -0000

--_000_C6252EA94E00E44EADC3A2FEB59D4402020F54D5xmbalnx07ciscoc_
Content-Type: text/plain; charset="us-ascii"
Content-Transfer-Encoding: quoted-printable

Hi all,

I decided to update the draft directly rather than put the proposals down i=
n an email, and as a result it's taken longer than I expected; I'll have it=
 done by tomorrow morning.

Rob

--_000_C6252EA94E00E44EADC3A2FEB59D4402020F54D5xmbalnx07ciscoc_
Content-Type: text/html; charset="us-ascii"
Content-Transfer-Encoding: quoted-printable

<html xmlns:v=3D"urn:schemas-microsoft-com:vml" xmlns:o=3D"urn:schemas-micr=
osoft-com:office:office" xmlns:w=3D"urn:schemas-microsoft-com:office:word" =
xmlns:m=3D"http://schemas.microsoft.com/office/2004/12/omml" xmlns=3D"http:=
//www.w3.org/TR/REC-html40">
<head>
<meta http-equiv=3D"Content-Type" content=3D"text/html; charset=3Dus-ascii"=
>
<meta name=3D"Generator" content=3D"Microsoft Word 14 (filtered medium)">
<style><!--
/* Font Definitions */
@font-face
	{font-family:Calibri;
	panose-1:2 15 5 2 2 2 4 3 2 4;}
/* Style Definitions */
p.MsoNormal, li.MsoNormal, div.MsoNormal
	{margin:0cm;
	margin-bottom:.0001pt;
	font-size:11.0pt;
	font-family:"Calibri","sans-serif";
	mso-fareast-language:EN-US;}
a:link, span.MsoHyperlink
	{mso-style-priority:99;
	color:blue;
	text-decoration:underline;}
a:visited, span.MsoHyperlinkFollowed
	{mso-style-priority:99;
	color:purple;
	text-decoration:underline;}
span.EmailStyle17
	{mso-style-type:personal-compose;
	font-family:"Calibri","sans-serif";
	color:windowtext;}
.MsoChpDefault
	{mso-style-type:export-only;
	font-family:"Calibri","sans-serif";
	mso-fareast-language:EN-US;}
@page WordSection1
	{size:612.0pt 792.0pt;
	margin:72.0pt 72.0pt 72.0pt 72.0pt;}
div.WordSection1
	{page:WordSection1;}
--></style><!--[if gte mso 9]><xml>
<o:shapedefaults v:ext=3D"edit" spidmax=3D"1026" />
</xml><![endif]--><!--[if gte mso 9]><xml>
<o:shapelayout v:ext=3D"edit">
<o:idmap v:ext=3D"edit" data=3D"1" />
</o:shapelayout></xml><![endif]-->
</head>
<body lang=3D"EN-GB" link=3D"blue" vlink=3D"purple">
<div class=3D"WordSection1">
<p class=3D"MsoNormal">Hi all,<o:p></o:p></p>
<p class=3D"MsoNormal"><o:p>&nbsp;</o:p></p>
<p class=3D"MsoNormal">I decided to update the draft directly rather than p=
ut the proposals down in an email, and as a result it&#8217;s taken longer =
than I expected; I&#8217;ll have it done by tomorrow morning.<o:p></o:p></p=
>
<p class=3D"MsoNormal"><o:p>&nbsp;</o:p></p>
<p class=3D"MsoNormal">Rob<o:p></o:p></p>
</div>
</body>
</html>

--_000_C6252EA94E00E44EADC3A2FEB59D4402020F54D5xmbalnx07ciscoc_--


From nobody Thu Apr 10 15:27:12 2014
Return-Path: <john@jlc.net>
X-Original-To: clue@ietfa.amsl.com
Delivered-To: clue@ietfa.amsl.com
Received: from localhost (ietfa.amsl.com [127.0.0.1]) by ietfa.amsl.com (Postfix) with ESMTP id 935B91A0259 for <clue@ietfa.amsl.com>; Thu, 10 Apr 2014 15:27:11 -0700 (PDT)
X-Virus-Scanned: amavisd-new at amsl.com
X-Spam-Flag: NO
X-Spam-Score: -1.772
X-Spam-Level: 
X-Spam-Status: No, score=-1.772 tagged_above=-999 required=5 tests=[BAYES_50=0.8, RCVD_IN_DNSWL_MED=-2.3, RP_MATCHES_RCVD=-0.272] autolearn=ham
Received: from mail.ietf.org ([4.31.198.44]) by localhost (ietfa.amsl.com [127.0.0.1]) (amavisd-new, port 10024) with ESMTP id kSTxRrdSKttx for <clue@ietfa.amsl.com>; Thu, 10 Apr 2014 15:27:09 -0700 (PDT)
Received: from mailhost.jlc.net (mailhost.jlc.net [199.201.159.4]) by ietfa.amsl.com (Postfix) with ESMTP id C59681A024F for <clue@ietf.org>; Thu, 10 Apr 2014 15:27:08 -0700 (PDT)
Received: by mailhost.jlc.net (Postfix, from userid 104) id 070E0C94A9; Thu, 10 Apr 2014 18:27:05 -0400 (EDT)
Date: Thu, 10 Apr 2014 18:27:05 -0400
From: John Leslie <john@jlc.net>
To: Paul Kyzivat <pkyzivat@alum.mit.edu>
Message-ID: <20140410222705.GW39240@verdi>
References: <533AF351.9050201@alum.mit.edu>
Mime-Version: 1.0
Content-Type: text/plain; charset=us-ascii
Content-Disposition: inline
In-Reply-To: <533AF351.9050201@alum.mit.edu>
User-Agent: Mutt/1.4.1i
Archived-At: http://mailarchive.ietf.org/arch/msg/clue/n_pYMZ8KZl8WEKIvhiwv93IHljw
Cc: CLUE <clue@ietf.org>
Subject: Re: [clue] Improving treatment of audio
X-BeenThere: clue@ietf.org
X-Mailman-Version: 2.1.15
Precedence: list
List-Id: CLUE - ControLling mUltiple streams for TElepresence <clue.ietf.org>
List-Unsubscribe: <https://www.ietf.org/mailman/options/clue>, <mailto:clue-request@ietf.org?subject=unsubscribe>
List-Archive: <http://www.ietf.org/mail-archive/web/clue/>
List-Post: <mailto:clue@ietf.org>
List-Help: <mailto:clue-request@ietf.org?subject=help>
List-Subscribe: <https://www.ietf.org/mailman/listinfo/clue>, <mailto:clue-request@ietf.org?subject=subscribe>
X-List-Received-Date: Thu, 10 Apr 2014 22:27:11 -0000

Paul Kyzivat <pkyzivat@alum.mit.edu> wrote:
> Date: Tue, 01 Apr 2014 13:11:45 -0400
> From: Paul Kyzivat <pkyzivat@alum.mit.edu>
> To: CLUE <clue@ietf.org>
> Subject: [clue] Improving treatment of audio
> 
> In today's design team meeting (01 April) we had a discussion of what is
> lacking about our treatment of audio, and how to fix it.
> 
> I want to open a ticket on this topic, but I need some help to properly 
> describe the task. IMO it has to do with what sort of spatial 
> information should be provided for audio captures, how it can be used to 
> correlate audio captures with video captures, and how it can be used to 
> choose which audio captures to configure.
> 
> Can somebody (John?) make a *concise* statement of what is needed?

   No.

   I don't believe we understand the problem space well enough to agree
on a _concise_ statement.

   I've been playinng around with an "Audio 101" presentation about
the problem space. I'm not particularly happy with it, but it's high
time I posted it and moved on.

   (Hopefully, it's light-hearted enough that reading it won't be too
painful...)

====

Audio 101

Q: What is sound?
A: Sound is periodic localized air-pressure differences in the
   audio-frequency range.

Q: What do you mean by "periodic"?
A: Periodic refers to some sort of repetition enabling reinforcement
   by combining with a time-delayed signal from the past.

Q: What is the audio-frequency range?
A: Human hearing is generally accepted to cover the range from
   20 to 20,000 Hertz (repetitions per second).

Q: How does sound propagate?
A: Nature abhors a vacuum. Air molecules move towards lower pressure.
   Sound can also travel in solid objects, such as a conference tables
   when the CEO is pounding his fist on it.

Q: What is a microphone?
A: It's a device with a membrane which responds to air-pressure
   differences between one side and the other, converting those
   differences to electrical signals.

Q: What is the output of a microphone?
A: Roughly speaking, it's a amplitude modulated voltage which tracks
   the air-pressure difference, within limits.

Q: What is an audio mixer?
A: It's a device that combines multiple input signals at adjustable
   levels into one or more output signals. Incoming signals may be
   muted, reduced, or amplified depending on the desired result.

Q: How well does that work, really?
A: Better than you'd expect! The human ear expects any original sound
   to be combined with "reverberation"

Q: What is "reverberation"?
A: Every room and even outdoor spaces have some level of reflection
   of sounds. A microphone or human ear can only hear the combination
   of original sound and reflections. The combination of reflections
   is called reverberation.

Q: Isn't that confusing?
A: Actually rather un-nerving to find yourself in a space designed
   to prevent any reverberation: with no reverberation your voice
   "disappears". Your ears have long since learned how to "extract"
   the original sound from the reverberation.

Q: With so many reflections, doesn't reverberation overwhelm the
   original sound?
A: Yes, but no real-world reflection is 100%, and in most cases the
   sound continues to dissipate with distance. The ear can make sense
   of it in "ordinary" rooms even when the original sound becomes
   inaudible.

Q: What's "dissipation with distance"?
A: In an object-free space, a localized air-pressure difference
   (sound) will spread in _every_ direction; thus there's less
   pressure difference as you get farther from the source. Thus,
   at twice the distance there's only one-quarter the pressure
   difference.

Q: Is that why we try to put microphones near the person speaking?
A: Yes: If the microphone is close enough, the person speaking
   dominates all other sound sources.

Q: What if there are more speakers than microphones?
A: We tend to put microphones within arm's reach and encourage
   speakers to grab the nearest one. (There is some equipment
   that can cope with multiple speakers at arms length: it depends
   on hardware and configuration.)

Q: Does that work well?
A: No.

Q: Why not?
A: First, because often they don't (grab). People assume if they can
   hear themselves, so can the other people.

Q: So, should we have a microphone in front of every speaker?
A: It would help, but it may not be close enough. A lapel mike isi
   perhaps 6 inches from the mouth. A tabletop may be 3 feet.

Q: How would you manage sound mixing for lapel mikes for every
   participant?
A: Usually, you can simply send each signal into a mixer, more or
   less equalize the levels, and leave it alone.

Q: How would you manage sound mixing for a cardioid mike in front
   of each person:
A: Poorly!

Q: Why can't it be the same as for lapel mikes?
A: If you feed an in-room speaker with the mix, you get feedback; if
   you feed it only to remote rooms, you get echo.

Q: So what do sound engineers do?
A: They "ride gain" by keeping most levels low and bringing up the
   person speaking.

Q: Can that be automated?
A: Easily.

Q: Why can't the same trick be used for one mike per three persons?
A: It can; but it will sound as if they're speaking from three different
   rooms. And, you'll probably get feedback.

Q: Aren't there devices to "eliminate" feedback?
A: Indeed there are: when the problem is mild they work pretty well;
   but if the problem is severe, instead of deafening feedback, it
   sounds like the mike is malfunctioning.

Q: What do good sound engineers do faced with that difficult situation?
A: Order another beer. Alcohol improves the sound.

Q: What is "feedback"?
A: When a microphone "hears" the signal it sent to the mixer returning
   from a loudspeaker in the room, the delay between the microphone
   picking up a sound and the sound returning from the loudspeaker
   will be just right to reinforce certain frequencies: if that
   reinforcing signal is not sufficiently damped, that "feedback"
   frequency will overwhelm normal speech.

Q: Are there ways to automatically damp such "feedback"?
A: Many, but that's outside the scope of this tutorial.

Q: How does "feedback" differ from "echo"?
A: Mainly it's the amount of delay. With loudspeakers being unable to
   reproduce sounds of less than 20 Hertz, the sound no longer gets
   loud enough to overwhelm, but the ear is confused by an acoustic
   very unlike any room it knows.

Q: So "echo cancellation" is needed?
A: Yes, although "echo avoidance" is better. Echo can be avoided if
   everyone wears earbuds, for example.

Q: Is that why telephones work so well?
A: Basically yes. However, telephones fake a slight echo by feeding
   back (electrically) a signal from the mouthpiece to the earpiece
   so that the person can "believe" the other end of the call is
   hearing him/her.

Q: Can we do that trick in CLUE?
A: To some degree, yes. It works quite well for headsets, but only if
   the actual echo from the room(s) at the other end of the call can
   be kept quiet enough.

Q: What about room to room, with no headsets?
A: Almost always it's better to depend on the room acoustics; but
   the trick can still be used to simulate echo from the other room(s)
   at an "appropriate" delay (coming from the loudspeakers).

Q: How is echo cancellation done?
A: It's a black art! Fundamentally the equipment guesses at the actual
   echo delay and "subtracts" a predicted echo from the actual
   returned echo. The details go way beyond this tutorial.

Q: Is that a worse problem for CLUE than for telephony?
A: Much worse! The actual delay changes unpredictably, especially if
   anything about the audio is "switched".

Q: What do you recommend that CLUE does about it?
A: To fix it, we'd need to synchronize time at both rooms and find
   the loudspeaker-to-microphone delays. That's too hard right now.
   To some degree, the problem can be hidden with artificial
   reverberation.

Q: What is artificial reverberation?
A: That's generating an artificial signal to simulate a larger room
   than what you're in, usually by applying an "impulse response"
   to the original signal.

Q: What's an "impulse response"?
A: Conceptually, it's what you would hear in a real room where the
   only sound input is an "infinitely short" step function. In practice
   it's often done by recording a (blank) pistol shot.

Q: Will that magic actually fool the human ear?
A: Close enough. It will sound "almost natural" and the brain will
   suspend its disbelief. :^)

====

   If I succeeded in outlining the scope, perhaps somebody else will take
a stab at the wording Paul wants...

--
John Leslie <john@jlc.net>


From nobody Thu Apr 10 16:17:57 2014
Return-Path: <coverdale@sympatico.ca>
X-Original-To: clue@ietfa.amsl.com
Delivered-To: clue@ietfa.amsl.com
Received: from localhost (ietfa.amsl.com [127.0.0.1]) by ietfa.amsl.com (Postfix) with ESMTP id D86DE1A030F for <clue@ietfa.amsl.com>; Thu, 10 Apr 2014 16:17:54 -0700 (PDT)
X-Virus-Scanned: amavisd-new at amsl.com
X-Spam-Flag: NO
X-Spam-Score: 0.8
X-Spam-Level: 
X-Spam-Status: No, score=0.8 tagged_above=-999 required=5 tests=[BAYES_50=0.8,  MSGID_FROM_MTA_HEADER=0.001, RCVD_IN_DNSWL_NONE=-0.0001, SPF_PASS=-0.001] autolearn=ham
Received: from mail.ietf.org ([4.31.198.44]) by localhost (ietfa.amsl.com [127.0.0.1]) (amavisd-new, port 10024) with ESMTP id G1VoTvcM5qJg for <clue@ietfa.amsl.com>; Thu, 10 Apr 2014 16:17:50 -0700 (PDT)
Received: from blu0-omc1-s21.blu0.hotmail.com (blu0-omc1-s21.blu0.hotmail.com [65.55.116.32]) by ietfa.amsl.com (Postfix) with ESMTP id 6A0B21A026E for <clue@ietf.org>; Thu, 10 Apr 2014 16:17:50 -0700 (PDT)
Received: from BLU0-SMTP69 ([65.55.116.7]) by blu0-omc1-s21.blu0.hotmail.com with Microsoft SMTPSVC(6.0.3790.4675);  Thu, 10 Apr 2014 16:17:49 -0700
X-TMN: [MkTs4ID8x/ywd889aGGE+tYZPMsK4lzZ]
X-Originating-Email: [coverdale@sympatico.ca]
Message-ID: <BLU0-SMTP6910DA20C9A9287137B585D0550@phx.gbl>
Received: from PaulNewPC ([74.15.60.251]) by BLU0-SMTP69.phx.gbl over TLS secured channel with Microsoft SMTPSVC(6.0.3790.4675);  Thu, 10 Apr 2014 16:17:48 -0700
From: Paul Coverdale <coverdale@sympatico.ca>
To: "'John Leslie'" <john@jlc.net>, "'Paul Kyzivat'" <pkyzivat@alum.mit.edu>
References: <533AF351.9050201@alum.mit.edu> <20140410222705.GW39240@verdi>
In-Reply-To: <20140410222705.GW39240@verdi>
Date: Thu, 10 Apr 2014 19:17:38 -0400
MIME-Version: 1.0
Content-Type: text/plain; charset="us-ascii"
Content-Transfer-Encoding: 7bit
X-Mailer: Microsoft Office Outlook 12.0
Thread-Index: Ac9VDAR4qOgLYHOmQSqhuTA1U8M5OQABUU5Q
Content-Language: en-us
X-OriginalArrivalTime: 10 Apr 2014 23:17:49.0000 (UTC) FILETIME=[16A9CC80:01CF5513]
Archived-At: http://mailarchive.ietf.org/arch/msg/clue/cxLJ3ypPk_VTAyhbzvw5pxanHFc
Cc: 'CLUE' <clue@ietf.org>
Subject: Re: [clue] Improving treatment of audio
X-BeenThere: clue@ietf.org
X-Mailman-Version: 2.1.15
Precedence: list
List-Id: CLUE - ControLling mUltiple streams for TElepresence <clue.ietf.org>
List-Unsubscribe: <https://www.ietf.org/mailman/options/clue>, <mailto:clue-request@ietf.org?subject=unsubscribe>
List-Archive: <http://www.ietf.org/mail-archive/web/clue/>
List-Post: <mailto:clue@ietf.org>
List-Help: <mailto:clue-request@ietf.org?subject=help>
List-Subscribe: <https://www.ietf.org/mailman/listinfo/clue>, <mailto:clue-request@ietf.org?subject=subscribe>
X-List-Received-Date: Thu, 10 Apr 2014 23:17:55 -0000

I don't disagree with John's "Audio 101" (well, most of it anyway). It takes
me back to my old days in the Audio & Acoustics group at BNR/Nortel when we
were trying to figure out what to do about acoustic echo and
de-reverberation. But those problems have pretty well been solved for audio
conferencing systems. So it's not clear to me what extra we need to do from
a CLUE perspective. What is the problem?

...Paul

>-----Original Message-----
>From: clue [mailto:clue-bounces@ietf.org] On Behalf Of John Leslie
>Sent: Thursday, April 10, 2014 6:27 PM
>To: Paul Kyzivat
>Cc: CLUE
>Subject: Re: [clue] Improving treatment of audio
>
>Paul Kyzivat <pkyzivat@alum.mit.edu> wrote:
>> Date: Tue, 01 Apr 2014 13:11:45 -0400
>> From: Paul Kyzivat <pkyzivat@alum.mit.edu>
>> To: CLUE <clue@ietf.org>
>> Subject: [clue] Improving treatment of audio
>>
>> In today's design team meeting (01 April) we had a discussion of what
>> is lacking about our treatment of audio, and how to fix it.
>>
>> I want to open a ticket on this topic, but I need some help to
>> properly describe the task. IMO it has to do with what sort of spatial
>> information should be provided for audio captures, how it can be used
>> to correlate audio captures with video captures, and how it can be
>> used to choose which audio captures to configure.
>>
>> Can somebody (John?) make a *concise* statement of what is needed?
>
>   No.
>
>   I don't believe we understand the problem space well enough to agree
>on a _concise_ statement.
>
>   I've been playinng around with an "Audio 101" presentation about the
>problem space. I'm not particularly happy with it, but it's high time I
>posted it and moved on.
>
>   (Hopefully, it's light-hearted enough that reading it won't be too
>painful...)
>
>====
>
>Audio 101
>
>Q: What is sound?
>A: Sound is periodic localized air-pressure differences in the
>   audio-frequency range.
>
>Q: What do you mean by "periodic"?
>A: Periodic refers to some sort of repetition enabling reinforcement
>   by combining with a time-delayed signal from the past.
>
>Q: What is the audio-frequency range?
>A: Human hearing is generally accepted to cover the range from
>   20 to 20,000 Hertz (repetitions per second).
>
>Q: How does sound propagate?
>A: Nature abhors a vacuum. Air molecules move towards lower pressure.
>   Sound can also travel in solid objects, such as a conference tables
>   when the CEO is pounding his fist on it.
>
>Q: What is a microphone?
>A: It's a device with a membrane which responds to air-pressure
>   differences between one side and the other, converting those
>   differences to electrical signals.
>
>Q: What is the output of a microphone?
>A: Roughly speaking, it's a amplitude modulated voltage which tracks
>   the air-pressure difference, within limits.
>
>Q: What is an audio mixer?
>A: It's a device that combines multiple input signals at adjustable
>   levels into one or more output signals. Incoming signals may be
>   muted, reduced, or amplified depending on the desired result.
>
>Q: How well does that work, really?
>A: Better than you'd expect! The human ear expects any original sound
>   to be combined with "reverberation"
>
>Q: What is "reverberation"?
>A: Every room and even outdoor spaces have some level of reflection
>   of sounds. A microphone or human ear can only hear the combination
>   of original sound and reflections. The combination of reflections
>   is called reverberation.
>
>Q: Isn't that confusing?
>A: Actually rather un-nerving to find yourself in a space designed
>   to prevent any reverberation: with no reverberation your voice
>   "disappears". Your ears have long since learned how to "extract"
>   the original sound from the reverberation.
>
>Q: With so many reflections, doesn't reverberation overwhelm the
>   original sound?
>A: Yes, but no real-world reflection is 100%, and in most cases the
>   sound continues to dissipate with distance. The ear can make sense
>   of it in "ordinary" rooms even when the original sound becomes
>   inaudible.
>
>Q: What's "dissipation with distance"?
>A: In an object-free space, a localized air-pressure difference
>   (sound) will spread in _every_ direction; thus there's less
>   pressure difference as you get farther from the source. Thus,
>   at twice the distance there's only one-quarter the pressure
>   difference.
>
>Q: Is that why we try to put microphones near the person speaking?
>A: Yes: If the microphone is close enough, the person speaking
>   dominates all other sound sources.
>
>Q: What if there are more speakers than microphones?
>A: We tend to put microphones within arm's reach and encourage
>   speakers to grab the nearest one. (There is some equipment
>   that can cope with multiple speakers at arms length: it depends
>   on hardware and configuration.)
>
>Q: Does that work well?
>A: No.
>
>Q: Why not?
>A: First, because often they don't (grab). People assume if they can
>   hear themselves, so can the other people.
>
>Q: So, should we have a microphone in front of every speaker?
>A: It would help, but it may not be close enough. A lapel mike isi
>   perhaps 6 inches from the mouth. A tabletop may be 3 feet.
>
>Q: How would you manage sound mixing for lapel mikes for every
>   participant?
>A: Usually, you can simply send each signal into a mixer, more or
>   less equalize the levels, and leave it alone.
>
>Q: How would you manage sound mixing for a cardioid mike in front
>   of each person:
>A: Poorly!
>
>Q: Why can't it be the same as for lapel mikes?
>A: If you feed an in-room speaker with the mix, you get feedback; if
>   you feed it only to remote rooms, you get echo.
>
>Q: So what do sound engineers do?
>A: They "ride gain" by keeping most levels low and bringing up the
>   person speaking.
>
>Q: Can that be automated?
>A: Easily.
>
>Q: Why can't the same trick be used for one mike per three persons?
>A: It can; but it will sound as if they're speaking from three different
>   rooms. And, you'll probably get feedback.
>
>Q: Aren't there devices to "eliminate" feedback?
>A: Indeed there are: when the problem is mild they work pretty well;
>   but if the problem is severe, instead of deafening feedback, it
>   sounds like the mike is malfunctioning.
>
>Q: What do good sound engineers do faced with that difficult situation?
>A: Order another beer. Alcohol improves the sound.
>
>Q: What is "feedback"?
>A: When a microphone "hears" the signal it sent to the mixer returning
>   from a loudspeaker in the room, the delay between the microphone
>   picking up a sound and the sound returning from the loudspeaker
>   will be just right to reinforce certain frequencies: if that
>   reinforcing signal is not sufficiently damped, that "feedback"
>   frequency will overwhelm normal speech.
>
>Q: Are there ways to automatically damp such "feedback"?
>A: Many, but that's outside the scope of this tutorial.
>
>Q: How does "feedback" differ from "echo"?
>A: Mainly it's the amount of delay. With loudspeakers being unable to
>   reproduce sounds of less than 20 Hertz, the sound no longer gets
>   loud enough to overwhelm, but the ear is confused by an acoustic
>   very unlike any room it knows.
>
>Q: So "echo cancellation" is needed?
>A: Yes, although "echo avoidance" is better. Echo can be avoided if
>   everyone wears earbuds, for example.
>
>Q: Is that why telephones work so well?
>A: Basically yes. However, telephones fake a slight echo by feeding
>   back (electrically) a signal from the mouthpiece to the earpiece
>   so that the person can "believe" the other end of the call is
>   hearing him/her.
>
>Q: Can we do that trick in CLUE?
>A: To some degree, yes. It works quite well for headsets, but only if
>   the actual echo from the room(s) at the other end of the call can
>   be kept quiet enough.
>
>Q: What about room to room, with no headsets?
>A: Almost always it's better to depend on the room acoustics; but
>   the trick can still be used to simulate echo from the other room(s)
>   at an "appropriate" delay (coming from the loudspeakers).
>
>Q: How is echo cancellation done?
>A: It's a black art! Fundamentally the equipment guesses at the actual
>   echo delay and "subtracts" a predicted echo from the actual
>   returned echo. The details go way beyond this tutorial.
>
>Q: Is that a worse problem for CLUE than for telephony?
>A: Much worse! The actual delay changes unpredictably, especially if
>   anything about the audio is "switched".
>
>Q: What do you recommend that CLUE does about it?
>A: To fix it, we'd need to synchronize time at both rooms and find
>   the loudspeaker-to-microphone delays. That's too hard right now.
>   To some degree, the problem can be hidden with artificial
>   reverberation.
>
>Q: What is artificial reverberation?
>A: That's generating an artificial signal to simulate a larger room
>   than what you're in, usually by applying an "impulse response"
>   to the original signal.
>
>Q: What's an "impulse response"?
>A: Conceptually, it's what you would hear in a real room where the
>   only sound input is an "infinitely short" step function. In practice
>   it's often done by recording a (blank) pistol shot.
>
>Q: Will that magic actually fool the human ear?
>A: Close enough. It will sound "almost natural" and the brain will
>   suspend its disbelief. :^)
>
>====
>
>   If I succeeded in outlining the scope, perhaps somebody else will
>take a stab at the wording Paul wants...
>
>--
>John Leslie <john@jlc.net>
>
>_______________________________________________
>clue mailing list
>clue@ietf.org
>https://www.ietf.org/mailman/listinfo/clue


From nobody Thu Apr 10 19:26:32 2014
Return-Path: <pkyzivat@alum.mit.edu>
X-Original-To: clue@ietfa.amsl.com
Delivered-To: clue@ietfa.amsl.com
Received: from localhost (ietfa.amsl.com [127.0.0.1]) by ietfa.amsl.com (Postfix) with ESMTP id 115FE1A0002 for <clue@ietfa.amsl.com>; Thu, 10 Apr 2014 19:26:30 -0700 (PDT)
X-Virus-Scanned: amavisd-new at amsl.com
X-Spam-Flag: NO
X-Spam-Score: 0.664
X-Spam-Level: 
X-Spam-Status: No, score=0.664 tagged_above=-999 required=5 tests=[BAYES_20=-0.001, DKIM_SIGNED=0.1, DKIM_VALID=-0.1, SPF_SOFTFAIL=0.665] autolearn=no
Received: from mail.ietf.org ([4.31.198.44]) by localhost (ietfa.amsl.com [127.0.0.1]) (amavisd-new, port 10024) with ESMTP id weC-m_Lpfzvg for <clue@ietfa.amsl.com>; Thu, 10 Apr 2014 19:26:25 -0700 (PDT)
Received: from qmta13.westchester.pa.mail.comcast.net (qmta13.westchester.pa.mail.comcast.net [IPv6:2001:558:fe14:44:76:96:59:243]) by ietfa.amsl.com (Postfix) with ESMTP id C4A871A03F9 for <clue@ietf.org>; Thu, 10 Apr 2014 19:26:24 -0700 (PDT)
Received: from omta04.westchester.pa.mail.comcast.net ([76.96.62.35]) by qmta13.westchester.pa.mail.comcast.net with comcast id oS441n0040ldTLk5DSSP3v; Fri, 11 Apr 2014 02:26:23 +0000
Received: from Paul-Kyzivats-MacBook-Pro.local ([50.138.229.164]) by omta04.westchester.pa.mail.comcast.net with comcast id oSSN1n0113ZTu2S01SSNU8; Fri, 11 Apr 2014 02:26:23 +0000
Message-ID: <534752CE.9010507@alum.mit.edu>
Date: Thu, 10 Apr 2014 22:26:22 -0400
From: Paul Kyzivat <pkyzivat@alum.mit.edu>
User-Agent: Mozilla/5.0 (Macintosh; Intel Mac OS X 10.7; rv:24.0) Gecko/20100101 Thunderbird/24.4.0
MIME-Version: 1.0
To: John Leslie <john@jlc.net>
References: <533AF351.9050201@alum.mit.edu> <20140410222705.GW39240@verdi>
In-Reply-To: <20140410222705.GW39240@verdi>
Content-Type: text/plain; charset=ISO-8859-1; format=flowed
Content-Transfer-Encoding: 7bit
DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=comcast.net; s=q20140121; t=1397183183; bh=1GSdEUOBOJD2tJrBAJ4IFgV0SpVeI0bN9w2jV6+YI/A=; h=Received:Received:Message-ID:Date:From:MIME-Version:To:Subject: Content-Type; b=vtZOLqTXWIcpZoBXrIw1n8qxALSGAnZC3wQISl0+hj2LgRuwf1FuFkwtacO6Pssma cPbfj+3MQZO+w673wVA5oZg7sBkjM9uESxez8j9Tt3Od1A2gbhYpxK3WmRcqfVg7A2 geLCD6MDaZQbtsXe/fr8N5FtQoq5uDqwWwNKVCN94OuL04hIDL09omYKsN0O+N+pcL 57LCvMBQPMp8xQsvfcNjmmtsociZZuJTpw2WGq0JNxcT9J95mSVHsWJnPnb1812lRJ qHeJ0WQonpvlw9AGNLobIxBZzEWYWhRVOfyJOsF6RPD8uhxs0xkjrUVWgBGeTdg3M/ cN7L8KDATYr/w==
Archived-At: http://mailarchive.ietf.org/arch/msg/clue/A8RTaGl_feGZzeCTCO2uTzrVA3E
Cc: CLUE <clue@ietf.org>
Subject: Re: [clue] Improving treatment of audio
X-BeenThere: clue@ietf.org
X-Mailman-Version: 2.1.15
Precedence: list
List-Id: CLUE - ControLling mUltiple streams for TElepresence <clue.ietf.org>
List-Unsubscribe: <https://www.ietf.org/mailman/options/clue>, <mailto:clue-request@ietf.org?subject=unsubscribe>
List-Archive: <http://www.ietf.org/mail-archive/web/clue/>
List-Post: <mailto:clue@ietf.org>
List-Help: <mailto:clue-request@ietf.org?subject=help>
List-Subscribe: <https://www.ietf.org/mailman/listinfo/clue>, <mailto:clue-request@ietf.org?subject=subscribe>
X-List-Received-Date: Fri, 11 Apr 2014 02:26:30 -0000

John,

While interesting, I can definitely say that what you gave is *not* what 
I was looking for. :-)

The concise statement I am looking for is not a description of the 
solution we are seeking. I am only (right now) looking for a description 
of the *issue* that needs to be solved.

E.g. "Define what sort of spatial information should be provided with 
the description of audio captures in order to support consumer selection 
of captures and proper playback."

*Then* the next step will be to actually do it, in the framework and the 
data model. So the issue should be phrased in a way that is achievable.

	Thanks,
	Paul


On 4/10/14 6:27 PM, John Leslie wrote:
> Paul Kyzivat <pkyzivat@alum.mit.edu> wrote:
>> Date: Tue, 01 Apr 2014 13:11:45 -0400
>> From: Paul Kyzivat <pkyzivat@alum.mit.edu>
>> To: CLUE <clue@ietf.org>
>> Subject: [clue] Improving treatment of audio
>>
>> In today's design team meeting (01 April) we had a discussion of what is
>> lacking about our treatment of audio, and how to fix it.
>>
>> I want to open a ticket on this topic, but I need some help to properly
>> describe the task. IMO it has to do with what sort of spatial
>> information should be provided for audio captures, how it can be used to
>> correlate audio captures with video captures, and how it can be used to
>> choose which audio captures to configure.
>>
>> Can somebody (John?) make a *concise* statement of what is needed?
>
>     No.
>
>     I don't believe we understand the problem space well enough to agree
> on a _concise_ statement.
>
>     I've been playinng around with an "Audio 101" presentation about
> the problem space. I'm not particularly happy with it, but it's high
> time I posted it and moved on.
>
>     (Hopefully, it's light-hearted enough that reading it won't be too
> painful...)
>
> ====
>
> Audio 101
>
> Q: What is sound?
> A: Sound is periodic localized air-pressure differences in the
>     audio-frequency range.
>
> Q: What do you mean by "periodic"?
> A: Periodic refers to some sort of repetition enabling reinforcement
>     by combining with a time-delayed signal from the past.
>
> Q: What is the audio-frequency range?
> A: Human hearing is generally accepted to cover the range from
>     20 to 20,000 Hertz (repetitions per second).
>
> Q: How does sound propagate?
> A: Nature abhors a vacuum. Air molecules move towards lower pressure.
>     Sound can also travel in solid objects, such as a conference tables
>     when the CEO is pounding his fist on it.
>
> Q: What is a microphone?
> A: It's a device with a membrane which responds to air-pressure
>     differences between one side and the other, converting those
>     differences to electrical signals.
>
> Q: What is the output of a microphone?
> A: Roughly speaking, it's a amplitude modulated voltage which tracks
>     the air-pressure difference, within limits.
>
> Q: What is an audio mixer?
> A: It's a device that combines multiple input signals at adjustable
>     levels into one or more output signals. Incoming signals may be
>     muted, reduced, or amplified depending on the desired result.
>
> Q: How well does that work, really?
> A: Better than you'd expect! The human ear expects any original sound
>     to be combined with "reverberation"
>
> Q: What is "reverberation"?
> A: Every room and even outdoor spaces have some level of reflection
>     of sounds. A microphone or human ear can only hear the combination
>     of original sound and reflections. The combination of reflections
>     is called reverberation.
>
> Q: Isn't that confusing?
> A: Actually rather un-nerving to find yourself in a space designed
>     to prevent any reverberation: with no reverberation your voice
>     "disappears". Your ears have long since learned how to "extract"
>     the original sound from the reverberation.
>
> Q: With so many reflections, doesn't reverberation overwhelm the
>     original sound?
> A: Yes, but no real-world reflection is 100%, and in most cases the
>     sound continues to dissipate with distance. The ear can make sense
>     of it in "ordinary" rooms even when the original sound becomes
>     inaudible.
>
> Q: What's "dissipation with distance"?
> A: In an object-free space, a localized air-pressure difference
>     (sound) will spread in _every_ direction; thus there's less
>     pressure difference as you get farther from the source. Thus,
>     at twice the distance there's only one-quarter the pressure
>     difference.
>
> Q: Is that why we try to put microphones near the person speaking?
> A: Yes: If the microphone is close enough, the person speaking
>     dominates all other sound sources.
>
> Q: What if there are more speakers than microphones?
> A: We tend to put microphones within arm's reach and encourage
>     speakers to grab the nearest one. (There is some equipment
>     that can cope with multiple speakers at arms length: it depends
>     on hardware and configuration.)
>
> Q: Does that work well?
> A: No.
>
> Q: Why not?
> A: First, because often they don't (grab). People assume if they can
>     hear themselves, so can the other people.
>
> Q: So, should we have a microphone in front of every speaker?
> A: It would help, but it may not be close enough. A lapel mike isi
>     perhaps 6 inches from the mouth. A tabletop may be 3 feet.
>
> Q: How would you manage sound mixing for lapel mikes for every
>     participant?
> A: Usually, you can simply send each signal into a mixer, more or
>     less equalize the levels, and leave it alone.
>
> Q: How would you manage sound mixing for a cardioid mike in front
>     of each person:
> A: Poorly!
>
> Q: Why can't it be the same as for lapel mikes?
> A: If you feed an in-room speaker with the mix, you get feedback; if
>     you feed it only to remote rooms, you get echo.
>
> Q: So what do sound engineers do?
> A: They "ride gain" by keeping most levels low and bringing up the
>     person speaking.
>
> Q: Can that be automated?
> A: Easily.
>
> Q: Why can't the same trick be used for one mike per three persons?
> A: It can; but it will sound as if they're speaking from three different
>     rooms. And, you'll probably get feedback.
>
> Q: Aren't there devices to "eliminate" feedback?
> A: Indeed there are: when the problem is mild they work pretty well;
>     but if the problem is severe, instead of deafening feedback, it
>     sounds like the mike is malfunctioning.
>
> Q: What do good sound engineers do faced with that difficult situation?
> A: Order another beer. Alcohol improves the sound.
>
> Q: What is "feedback"?
> A: When a microphone "hears" the signal it sent to the mixer returning
>     from a loudspeaker in the room, the delay between the microphone
>     picking up a sound and the sound returning from the loudspeaker
>     will be just right to reinforce certain frequencies: if that
>     reinforcing signal is not sufficiently damped, that "feedback"
>     frequency will overwhelm normal speech.
>
> Q: Are there ways to automatically damp such "feedback"?
> A: Many, but that's outside the scope of this tutorial.
>
> Q: How does "feedback" differ from "echo"?
> A: Mainly it's the amount of delay. With loudspeakers being unable to
>     reproduce sounds of less than 20 Hertz, the sound no longer gets
>     loud enough to overwhelm, but the ear is confused by an acoustic
>     very unlike any room it knows.
>
> Q: So "echo cancellation" is needed?
> A: Yes, although "echo avoidance" is better. Echo can be avoided if
>     everyone wears earbuds, for example.
>
> Q: Is that why telephones work so well?
> A: Basically yes. However, telephones fake a slight echo by feeding
>     back (electrically) a signal from the mouthpiece to the earpiece
>     so that the person can "believe" the other end of the call is
>     hearing him/her.
>
> Q: Can we do that trick in CLUE?
> A: To some degree, yes. It works quite well for headsets, but only if
>     the actual echo from the room(s) at the other end of the call can
>     be kept quiet enough.
>
> Q: What about room to room, with no headsets?
> A: Almost always it's better to depend on the room acoustics; but
>     the trick can still be used to simulate echo from the other room(s)
>     at an "appropriate" delay (coming from the loudspeakers).
>
> Q: How is echo cancellation done?
> A: It's a black art! Fundamentally the equipment guesses at the actual
>     echo delay and "subtracts" a predicted echo from the actual
>     returned echo. The details go way beyond this tutorial.
>
> Q: Is that a worse problem for CLUE than for telephony?
> A: Much worse! The actual delay changes unpredictably, especially if
>     anything about the audio is "switched".
>
> Q: What do you recommend that CLUE does about it?
> A: To fix it, we'd need to synchronize time at both rooms and find
>     the loudspeaker-to-microphone delays. That's too hard right now.
>     To some degree, the problem can be hidden with artificial
>     reverberation.
>
> Q: What is artificial reverberation?
> A: That's generating an artificial signal to simulate a larger room
>     than what you're in, usually by applying an "impulse response"
>     to the original signal.
>
> Q: What's an "impulse response"?
> A: Conceptually, it's what you would hear in a real room where the
>     only sound input is an "infinitely short" step function. In practice
>     it's often done by recording a (blank) pistol shot.
>
> Q: Will that magic actually fool the human ear?
> A: Close enough. It will sound "almost natural" and the brain will
>     suspend its disbelief. :^)
>
> ====
>
>     If I succeeded in outlining the scope, perhaps somebody else will take
> a stab at the wording Paul wants...
>
> --
> John Leslie <john@jlc.net>
>


From nobody Thu Apr 10 19:38:02 2014
Return-Path: <Christian.Groves@nteczone.com>
X-Original-To: clue@ietfa.amsl.com
Delivered-To: clue@ietfa.amsl.com
Received: from localhost (ietfa.amsl.com [127.0.0.1]) by ietfa.amsl.com (Postfix) with ESMTP id D8EE41A000A for <clue@ietfa.amsl.com>; Thu, 10 Apr 2014 19:37:59 -0700 (PDT)
X-Virus-Scanned: amavisd-new at amsl.com
X-Spam-Flag: NO
X-Spam-Score: -0.001
X-Spam-Level: 
X-Spam-Status: No, score=-0.001 tagged_above=-999 required=5 tests=[BAYES_40=-0.001] autolearn=ham
Received: from mail.ietf.org ([4.31.198.44]) by localhost (ietfa.amsl.com [127.0.0.1]) (amavisd-new, port 10024) with ESMTP id bsbKHRysW9d7 for <clue@ietfa.amsl.com>; Thu, 10 Apr 2014 19:37:53 -0700 (PDT)
Received: from cserver5.myshophosting.com (cserver5.myshophosting.com [175.107.161.1]) by ietfa.amsl.com (Postfix) with ESMTP id 3B3531A0391 for <clue@ietf.org>; Thu, 10 Apr 2014 19:37:43 -0700 (PDT)
Received: from ppp118-209-176-12.lns20.mel6.internode.on.net ([118.209.176.12]:51591 helo=[127.0.0.1]) by cserver5.myshophosting.com with esmtpsa (TLSv1:DHE-RSA-AES128-SHA:128) (Exim 4.82) (envelope-from <Christian.Groves@nteczone.com>) id 1WYRLX-0007mP-G4 for clue@ietf.org; Fri, 11 Apr 2014 12:37:39 +1000
Message-ID: <53475572.4040807@nteczone.com>
Date: Fri, 11 Apr 2014 12:37:38 +1000
From: Christian Groves <Christian.Groves@nteczone.com>
User-Agent: Mozilla/5.0 (Windows NT 6.1; WOW64; rv:24.0) Gecko/20100101 Thunderbird/24.4.0
MIME-Version: 1.0
To: clue@ietf.org
References: <533AF351.9050201@alum.mit.edu> <20140410222705.GW39240@verdi> <BLU0-SMTP6910DA20C9A9287137B585D0550@phx.gbl>
In-Reply-To: <BLU0-SMTP6910DA20C9A9287137B585D0550@phx.gbl>
Content-Type: text/plain; charset=ISO-8859-1; format=flowed
Content-Transfer-Encoding: 7bit
X-AntiAbuse: This header was added to track abuse, please include it with any abuse report
X-AntiAbuse: Primary Hostname - cserver5.myshophosting.com
X-AntiAbuse: Original Domain - ietf.org
X-AntiAbuse: Originator/Caller UID/GID - [47 12] / [47 12]
X-AntiAbuse: Sender Address Domain - nteczone.com
X-Get-Message-Sender-Via: cserver5.myshophosting.com: authenticated_id: christian.groves@nteczone.com
X-Source: 
X-Source-Args: 
X-Source-Dir: 
Archived-At: http://mailarchive.ietf.org/arch/msg/clue/PCvQkmr0Ox8AXGrmoVlkWJC03vw
Subject: Re: [clue] Improving treatment of audio
X-BeenThere: clue@ietf.org
X-Mailman-Version: 2.1.15
Precedence: list
List-Id: CLUE - ControLling mUltiple streams for TElepresence <clue.ietf.org>
List-Unsubscribe: <https://www.ietf.org/mailman/options/clue>, <mailto:clue-request@ietf.org?subject=unsubscribe>
List-Archive: <http://www.ietf.org/mail-archive/web/clue/>
List-Post: <mailto:clue@ietf.org>
List-Help: <mailto:clue-request@ietf.org?subject=help>
List-Subscribe: <https://www.ietf.org/mailman/listinfo/clue>, <mailto:clue-request@ietf.org?subject=subscribe>
X-List-Received-Date: Fri, 11 Apr 2014 02:38:00 -0000

"So it's not clear to me what extra we need to do from a CLUE 
perspective. What is the problem?"
I guess this is the pertinent point.

 From John's Audio 101 it seems that anything to do with spatial 
information that would relate to reverberation is in the too hard 
basket. So it seems "area of capture" for audio could be marked "not 
applicable" in the framework. There doesn't seem to be any driver for 
having the "point of capture" apply to an audio capture either.

 From the Audio101 there does seem to be a dependency on where the 
microphone is located with respect to the person speaking (i.e. lapel 
mic, desk mic, room mic) to how it is handled at mixing/playout. Perhaps 
this is useful to signal via CLUE? If it is possible to signal this them 
perhaps tying a particular audio capture to a video capture makes sense? 
i.e. a talker giving a presentation using a lapel mic captured by a 
particular video.

Regards, Christian


On 11/04/2014 9:17 AM, Paul Coverdale wrote:
> I don't disagree with John's "Audio 101" (well, most of it anyway). It takes
> me back to my old days in the Audio & Acoustics group at BNR/Nortel when we
> were trying to figure out what to do about acoustic echo and
> de-reverberation. But those problems have pretty well been solved for audio
> conferencing systems. So it's not clear to me what extra we need to do from
> a CLUE perspective. What is the problem?
>
> ...Paul
>
>> -----Original Message-----
>> From: clue [mailto:clue-bounces@ietf.org] On Behalf Of John Leslie
>> Sent: Thursday, April 10, 2014 6:27 PM
>> To: Paul Kyzivat
>> Cc: CLUE
>> Subject: Re: [clue] Improving treatment of audio
>>
>> Paul Kyzivat <pkyzivat@alum.mit.edu> wrote:
>>> Date: Tue, 01 Apr 2014 13:11:45 -0400
>>> From: Paul Kyzivat <pkyzivat@alum.mit.edu>
>>> To: CLUE <clue@ietf.org>
>>> Subject: [clue] Improving treatment of audio
>>>
>>> In today's design team meeting (01 April) we had a discussion of what
>>> is lacking about our treatment of audio, and how to fix it.
>>>
>>> I want to open a ticket on this topic, but I need some help to
>>> properly describe the task. IMO it has to do with what sort of spatial
>>> information should be provided for audio captures, how it can be used
>>> to correlate audio captures with video captures, and how it can be
>>> used to choose which audio captures to configure.
>>>
>>> Can somebody (John?) make a *concise* statement of what is needed?
>>    No.
>>
>>    I don't believe we understand the problem space well enough to agree
>> on a _concise_ statement.
>>
>>    I've been playinng around with an "Audio 101" presentation about the
>> problem space. I'm not particularly happy with it, but it's high time I
>> posted it and moved on.
>>
>>    (Hopefully, it's light-hearted enough that reading it won't be too
>> painful...)
>>
>> ====
>>
>> Audio 101
>>
>> Q: What is sound?
>> A: Sound is periodic localized air-pressure differences in the
>>    audio-frequency range.
>>
>> Q: What do you mean by "periodic"?
>> A: Periodic refers to some sort of repetition enabling reinforcement
>>    by combining with a time-delayed signal from the past.
>>
>> Q: What is the audio-frequency range?
>> A: Human hearing is generally accepted to cover the range from
>>    20 to 20,000 Hertz (repetitions per second).
>>
>> Q: How does sound propagate?
>> A: Nature abhors a vacuum. Air molecules move towards lower pressure.
>>    Sound can also travel in solid objects, such as a conference tables
>>    when the CEO is pounding his fist on it.
>>
>> Q: What is a microphone?
>> A: It's a device with a membrane which responds to air-pressure
>>    differences between one side and the other, converting those
>>    differences to electrical signals.
>>
>> Q: What is the output of a microphone?
>> A: Roughly speaking, it's a amplitude modulated voltage which tracks
>>    the air-pressure difference, within limits.
>>
>> Q: What is an audio mixer?
>> A: It's a device that combines multiple input signals at adjustable
>>    levels into one or more output signals. Incoming signals may be
>>    muted, reduced, or amplified depending on the desired result.
>>
>> Q: How well does that work, really?
>> A: Better than you'd expect! The human ear expects any original sound
>>    to be combined with "reverberation"
>>
>> Q: What is "reverberation"?
>> A: Every room and even outdoor spaces have some level of reflection
>>    of sounds. A microphone or human ear can only hear the combination
>>    of original sound and reflections. The combination of reflections
>>    is called reverberation.
>>
>> Q: Isn't that confusing?
>> A: Actually rather un-nerving to find yourself in a space designed
>>    to prevent any reverberation: with no reverberation your voice
>>    "disappears". Your ears have long since learned how to "extract"
>>    the original sound from the reverberation.
>>
>> Q: With so many reflections, doesn't reverberation overwhelm the
>>    original sound?
>> A: Yes, but no real-world reflection is 100%, and in most cases the
>>    sound continues to dissipate with distance. The ear can make sense
>>    of it in "ordinary" rooms even when the original sound becomes
>>    inaudible.
>>
>> Q: What's "dissipation with distance"?
>> A: In an object-free space, a localized air-pressure difference
>>    (sound) will spread in _every_ direction; thus there's less
>>    pressure difference as you get farther from the source. Thus,
>>    at twice the distance there's only one-quarter the pressure
>>    difference.
>>
>> Q: Is that why we try to put microphones near the person speaking?
>> A: Yes: If the microphone is close enough, the person speaking
>>    dominates all other sound sources.
>>
>> Q: What if there are more speakers than microphones?
>> A: We tend to put microphones within arm's reach and encourage
>>    speakers to grab the nearest one. (There is some equipment
>>    that can cope with multiple speakers at arms length: it depends
>>    on hardware and configuration.)
>>
>> Q: Does that work well?
>> A: No.
>>
>> Q: Why not?
>> A: First, because often they don't (grab). People assume if they can
>>    hear themselves, so can the other people.
>>
>> Q: So, should we have a microphone in front of every speaker?
>> A: It would help, but it may not be close enough. A lapel mike isi
>>    perhaps 6 inches from the mouth. A tabletop may be 3 feet.
>>
>> Q: How would you manage sound mixing for lapel mikes for every
>>    participant?
>> A: Usually, you can simply send each signal into a mixer, more or
>>    less equalize the levels, and leave it alone.
>>
>> Q: How would you manage sound mixing for a cardioid mike in front
>>    of each person:
>> A: Poorly!
>>
>> Q: Why can't it be the same as for lapel mikes?
>> A: If you feed an in-room speaker with the mix, you get feedback; if
>>    you feed it only to remote rooms, you get echo.
>>
>> Q: So what do sound engineers do?
>> A: They "ride gain" by keeping most levels low and bringing up the
>>    person speaking.
>>
>> Q: Can that be automated?
>> A: Easily.
>>
>> Q: Why can't the same trick be used for one mike per three persons?
>> A: It can; but it will sound as if they're speaking from three different
>>    rooms. And, you'll probably get feedback.
>>
>> Q: Aren't there devices to "eliminate" feedback?
>> A: Indeed there are: when the problem is mild they work pretty well;
>>    but if the problem is severe, instead of deafening feedback, it
>>    sounds like the mike is malfunctioning.
>>
>> Q: What do good sound engineers do faced with that difficult situation?
>> A: Order another beer. Alcohol improves the sound.
>>
>> Q: What is "feedback"?
>> A: When a microphone "hears" the signal it sent to the mixer returning
>>    from a loudspeaker in the room, the delay between the microphone
>>    picking up a sound and the sound returning from the loudspeaker
>>    will be just right to reinforce certain frequencies: if that
>>    reinforcing signal is not sufficiently damped, that "feedback"
>>    frequency will overwhelm normal speech.
>>
>> Q: Are there ways to automatically damp such "feedback"?
>> A: Many, but that's outside the scope of this tutorial.
>>
>> Q: How does "feedback" differ from "echo"?
>> A: Mainly it's the amount of delay. With loudspeakers being unable to
>>    reproduce sounds of less than 20 Hertz, the sound no longer gets
>>    loud enough to overwhelm, but the ear is confused by an acoustic
>>    very unlike any room it knows.
>>
>> Q: So "echo cancellation" is needed?
>> A: Yes, although "echo avoidance" is better. Echo can be avoided if
>>    everyone wears earbuds, for example.
>>
>> Q: Is that why telephones work so well?
>> A: Basically yes. However, telephones fake a slight echo by feeding
>>    back (electrically) a signal from the mouthpiece to the earpiece
>>    so that the person can "believe" the other end of the call is
>>    hearing him/her.
>>
>> Q: Can we do that trick in CLUE?
>> A: To some degree, yes. It works quite well for headsets, but only if
>>    the actual echo from the room(s) at the other end of the call can
>>    be kept quiet enough.
>>
>> Q: What about room to room, with no headsets?
>> A: Almost always it's better to depend on the room acoustics; but
>>    the trick can still be used to simulate echo from the other room(s)
>>    at an "appropriate" delay (coming from the loudspeakers).
>>
>> Q: How is echo cancellation done?
>> A: It's a black art! Fundamentally the equipment guesses at the actual
>>    echo delay and "subtracts" a predicted echo from the actual
>>    returned echo. The details go way beyond this tutorial.
>>
>> Q: Is that a worse problem for CLUE than for telephony?
>> A: Much worse! The actual delay changes unpredictably, especially if
>>    anything about the audio is "switched".
>>
>> Q: What do you recommend that CLUE does about it?
>> A: To fix it, we'd need to synchronize time at both rooms and find
>>    the loudspeaker-to-microphone delays. That's too hard right now.
>>    To some degree, the problem can be hidden with artificial
>>    reverberation.
>>
>> Q: What is artificial reverberation?
>> A: That's generating an artificial signal to simulate a larger room
>>    than what you're in, usually by applying an "impulse response"
>>    to the original signal.
>>
>> Q: What's an "impulse response"?
>> A: Conceptually, it's what you would hear in a real room where the
>>    only sound input is an "infinitely short" step function. In practice
>>    it's often done by recording a (blank) pistol shot.
>>
>> Q: Will that magic actually fool the human ear?
>> A: Close enough. It will sound "almost natural" and the brain will
>>    suspend its disbelief. :^)
>>
>> ====
>>
>>    If I succeeded in outlining the scope, perhaps somebody else will
>> take a stab at the wording Paul wants...
>>
>> --
>> John Leslie <john@jlc.net>
>>
>> _______________________________________________
>> clue mailing list
>> clue@ietf.org
>> https://www.ietf.org/mailman/listinfo/clue
> _______________________________________________
> clue mailing list
> clue@ietf.org
> https://www.ietf.org/mailman/listinfo/clue
>


From nobody Thu Apr 10 20:00:15 2014
Return-Path: <rohanse2@cisco.com>
X-Original-To: clue@ietfa.amsl.com
Delivered-To: clue@ietfa.amsl.com
Received: from localhost (ietfa.amsl.com [127.0.0.1]) by ietfa.amsl.com (Postfix) with ESMTP id 87D221A0407 for <clue@ietfa.amsl.com>; Thu, 10 Apr 2014 20:00:00 -0700 (PDT)
X-Virus-Scanned: amavisd-new at amsl.com
X-Spam-Flag: NO
X-Spam-Score: -9.773
X-Spam-Level: 
X-Spam-Status: No, score=-9.773 tagged_above=-999 required=5 tests=[BAYES_00=-1.9, DKIM_SIGNED=0.1, DKIM_VALID=-0.1, DKIM_VALID_AU=-0.1, RP_MATCHES_RCVD=-0.272, SPF_PASS=-0.001, USER_IN_DEF_DKIM_WL=-7.5] autolearn=ham
Received: from mail.ietf.org ([4.31.198.44]) by localhost (ietfa.amsl.com [127.0.0.1]) (amavisd-new, port 10024) with ESMTP id y0rfKrIztgQj for <clue@ietfa.amsl.com>; Thu, 10 Apr 2014 19:59:40 -0700 (PDT)
Received: from alln-iport-1.cisco.com (alln-iport-1.cisco.com [173.37.142.88]) by ietfa.amsl.com (Postfix) with ESMTP id 177C01A001D for <clue@ietf.org>; Thu, 10 Apr 2014 19:59:38 -0700 (PDT)
DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/simple; d=cisco.com; i=@cisco.com; l=3486; q=dns/txt; s=iport; t=1397185177; x=1398394777; h=from:to:subject:date:message-id:references:in-reply-to: content-transfer-encoding:mime-version; bh=ydaR7QAUEOQTjGJ5WKFTNWez0DsLXoQUoux0jQByeSc=; b=I3q6Ih9K+y0jVjdY9tnynvXSv6Viisjv/3sC+ndjrtLl5E1dDTJQWoMh DvwU/LP0j4nuuZGXncTRxyR3XaLUo7H71JPuZKOwYOx9aav27ij4yMjvH bXbJyZI+BGf+Lr3TvV/LprxD62qq8vyqA2FgELiaTcR+G7UdQxOHBH8Mf M=;
X-IronPort-Anti-Spam-Filtered: true
X-IronPort-Anti-Spam-Result: AgEGAHhaR1OtJV2b/2dsb2JhbABagwY7UQaDDsEpGYEFFnSCJQEBAQQjEUMOBAIBCBEEAQEDAgYdAwICAjAUAQYBAQUDAgQTCAGHcwgFqiOiWReBKY0SOAaCaTWBFASaE5ENgzCCKw
X-IronPort-AV: E=Sophos;i="4.97,838,1389744000"; d="scan'208";a="34829426"
Received: from rcdn-core-4.cisco.com ([173.37.93.155]) by alln-iport-1.cisco.com with ESMTP; 11 Apr 2014 02:59:36 +0000
Received: from xhc-rcd-x13.cisco.com (xhc-rcd-x13.cisco.com [173.37.183.87]) by rcdn-core-4.cisco.com (8.14.5/8.14.5) with ESMTP id s3B2xabv031325 (version=TLSv1/SSLv3 cipher=AES128-SHA bits=128 verify=FAIL) for <clue@ietf.org>; Fri, 11 Apr 2014 02:59:36 GMT
Received: from xmb-aln-x07.cisco.com ([169.254.2.162]) by xhc-rcd-x13.cisco.com ([173.37.183.87]) with mapi id 14.03.0123.003; Thu, 10 Apr 2014 21:59:36 -0500
From: "Robert Hansen (rohanse2)" <rohanse2@cisco.com>
To: "clue@ietf.org" <clue@ietf.org>
Thread-Topic: New Version Notification for draft-kyzivat-clue-signaling-08.txt
Thread-Index: AQHPVTD2urOS1UDktke+aMmvDx+QDJsLt9IA
Date: Fri, 11 Apr 2014 02:59:35 +0000
Message-ID: <C6252EA94E00E44EADC3A2FEB59D440202125EA3@xmb-aln-x07.cisco.com>
References: <20140411025128.30125.21894.idtracker@ietfa.amsl.com>
In-Reply-To: <20140411025128.30125.21894.idtracker@ietfa.amsl.com>
Accept-Language: en-GB, en-US
Content-Language: en-US
X-MS-Has-Attach: 
X-MS-TNEF-Correlator: 
x-originating-ip: [10.61.105.32]
Content-Type: text/plain; charset="utf-8"
Content-Transfer-Encoding: base64
MIME-Version: 1.0
Archived-At: http://mailarchive.ietf.org/arch/msg/clue/s4BwTwD1ZIBrEsSjCRolQaa7kTA
Subject: [clue] FW: New Version Notification for draft-kyzivat-clue-signaling-08.txt
X-BeenThere: clue@ietf.org
X-Mailman-Version: 2.1.15
Precedence: list
List-Id: CLUE - ControLling mUltiple streams for TElepresence <clue.ietf.org>
List-Unsubscribe: <https://www.ietf.org/mailman/options/clue>, <mailto:clue-request@ietf.org?subject=unsubscribe>
List-Archive: <http://www.ietf.org/mail-archive/web/clue/>
List-Post: <mailto:clue@ietf.org>
List-Help: <mailto:clue-request@ietf.org?subject=help>
List-Subscribe: <https://www.ietf.org/mailman/listinfo/clue>, <mailto:clue-request@ietf.org?subject=subscribe>
X-List-Received-Date: Fri, 11 Apr 2014 03:00:01 -0000

QXBvbG9naWVzIC0gZGlkbid0IGVuZCB1cCBnZXR0aW5nIGl0IGRvbmUgdW50aWwgYSBkYXkgYW5k
IGEgaGFsZiBhZnRlciBJIHNhaWQgSSB3b3VsZC4NCg0KSSd2ZSBpbmNsdWRlZCBwcm9wb3NhbHMg
Zm9yIGEgbnVtYmVyIG9mIG9wZW4gaXNzdWVzLCB3aXRoIHRoZSBleGNlcHRpb24gb2YgdGFja2xp
bmcgdW5pZGlyZWN0aW9uYWwgc3RyZWFtcyBvciByZS11c2Ugb2YgaW5pdGlhbCBhdWRpby92aWRl
byBtIGxpbmVzIGFzIENMVUUtY29udHJvbGxlZCBsaW5lcyAocGFybHkgYmVjYXVzZSB0aGV5J3Jl
IG1vcmUgb3B0aW1pc2F0aW9ucyB0aGFuIGFueXRoaW5nIHdlIG5lZWQgdG8gbWFrZSBDTFVFIHdv
cmssIGFuZCBwYXJ0bHkgYmVjYXVzZSBJIGRvbuKAmXQgaGF2ZSBzdHJvbmcgb3BpbmlvbnMgYWJv
dXQgdGhlbSkuDQoNCkkndmUgYWxzbyByZWZvcm1hdHRlZCB0aGUgU0RQIHJlbGF0ZWQgc2VjdGlv
biBvZiB0aGUgZG9jdW1lbnQgdG8gbWFrZSBpdCBtb3JlIGNlbnRyZWQgYXJvdW5kIHRoZSBncm91
cGluZyBmcmFtZXdvcmsgYW5kIGl0cyB1c2UgLSB0aGFua3MgdG8gQ2hyaXN0ZXIgZm9yIHN1Z2dl
c3Rpb25zIG9uIGhvdyB0byBkbyB0aGF0Lg0KDQpPaCwgYW5kIHJhdGhlciB0aGFuICdzaXAuY2x1
ZScgZm9yIHRoZSBtZWRpYSB0YWcgYW5kICdDTFVFJyBmb3IgdGhlIGdyb3VwaW5nIHNlbWFudGlj
IEkgd2VudCB3aXRoICdzaXAudGVsZXByZXNlbmNlJyBhbmQgJ1RFTEVQUkVTRU5DRScsIGFzIEkg
dGhpbmsgdGhlIGxhdHRlciB3aWxsIG1ha2UgaXQgbW9yZSBvYnZpb3VzIHRvIHBlb3BsZSB3aGF0
IHRoZSBmdW5jdGlvbmFsaXR5IHRoYXQgaXMgYmVpbmcgaW52b2tlZCBpcyBmb3IuIEkga25vdyBw
ZW9wbGUgbWF5IGhhdmUgc3Ryb25nIG9waW5pb25zIHRvIHRoZSBjb250cmFyeSwgYnV0IEkgZmln
dXJlZCBJJ2QgcmFpc2UgdGhlIGlzc3VlIG5vdyBiZWZvcmUgd2UgZ290IHRvbyBjb21mb3J0YWJs
ZSB3aXRoIHRoZSBleGlzdGluZyBzeW50YXguDQoNClJvYg0KDQotLS0tLU9yaWdpbmFsIE1lc3Nh
Z2UtLS0tLQ0KRnJvbTogaW50ZXJuZXQtZHJhZnRzQGlldGYub3JnIFttYWlsdG86aW50ZXJuZXQt
ZHJhZnRzQGlldGYub3JnXSANClNlbnQ6IDExIEFwcmlsIDIwMTQgMDM6NTENClRvOiBMZW5uYXJk
IFhpYW87IFBhdWwgS3l6aXZhdDsgTGVubmFyZCBYaWFvOyBSb2JlcnQgSGFuc2VuIChyb2hhbnNl
Mik7IFJvYmVydCBIYW5zZW4gKHJvaGFuc2UyKTsgQ2hyaXN0aWFuIEdyb3ZlczsgQ2hyaXN0aWFu
IEdyb3ZlczsgUGF1bCBLeXppdmF0DQpTdWJqZWN0OiBOZXcgVmVyc2lvbiBOb3RpZmljYXRpb24g
Zm9yIGRyYWZ0LWt5eml2YXQtY2x1ZS1zaWduYWxpbmctMDgudHh0DQoNCg0KQSBuZXcgdmVyc2lv
biBvZiBJLUQsIGRyYWZ0LWt5eml2YXQtY2x1ZS1zaWduYWxpbmctMDgudHh0DQpoYXMgYmVlbiBz
dWNjZXNzZnVsbHkgc3VibWl0dGVkIGJ5IFJvYmVydCBIYW5zZW4gYW5kIHBvc3RlZCB0byB0aGUg
SUVURiByZXBvc2l0b3J5Lg0KDQpOYW1lOgkJZHJhZnQta3l6aXZhdC1jbHVlLXNpZ25hbGluZw0K
UmV2aXNpb246CTA4DQpUaXRsZToJCUNMVUUgU2lnbmFsaW5nDQpEb2N1bWVudCBkYXRlOgkyMDE0
LTA0LTExDQpHcm91cDoJCUluZGl2aWR1YWwgU3VibWlzc2lvbg0KUGFnZXM6CQk0MQ0KVVJMOiAg
ICAgICAgICAgIGh0dHA6Ly93d3cuaWV0Zi5vcmcvaW50ZXJuZXQtZHJhZnRzL2RyYWZ0LWt5eml2
YXQtY2x1ZS1zaWduYWxpbmctMDgudHh0DQpTdGF0dXM6ICAgICAgICAgaHR0cHM6Ly9kYXRhdHJh
Y2tlci5pZXRmLm9yZy9kb2MvZHJhZnQta3l6aXZhdC1jbHVlLXNpZ25hbGluZy8NCkh0bWxpemVk
OiAgICAgICBodHRwOi8vdG9vbHMuaWV0Zi5vcmcvaHRtbC9kcmFmdC1reXppdmF0LWNsdWUtc2ln
bmFsaW5nLTA4DQpEaWZmOiAgICAgICAgICAgaHR0cDovL3d3dy5pZXRmLm9yZy9yZmNkaWZmP3Vy
bDI9ZHJhZnQta3l6aXZhdC1jbHVlLXNpZ25hbGluZy0wOA0KDQpBYnN0cmFjdDoNCiAgIFRoaXMg
ZG9jdW1lbnQgc3BlY2lmaWVzIGhvdyBDTFVFLXNwZWNpZmljIHNpZ25hbGluZyBzdWNoIGFzIHRo
ZSBDTFVFDQogICBwcm90b2NvbCBbSS1ELnByZXN0YS1jbHVlLXByb3RvY29sXSBhbmQgdGhlIENM
VUUgZGF0YSBjaGFubmVsDQogICBbSS1ELmlldGYtY2x1ZS1kYXRhY2hhbm5lbF0gYXJlIHVzZWQg
d2l0aCBlYWNoIG90aGVyIGFuZCB3aXRoDQogICBleGlzdGluZyBzaWduYWxpbmcgbWVjaGFuaXNt
cyBzdWNoIGFzIFNJUCBhbmQgU0RQIHRvIHByb2R1Y2UgYQ0KICAgdGVsZXByZXNlbmNlIGNhbGwu
DQoNCiAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAg
ICAgICAgICAgICAgICAgICAgICAgICAgICAgICANCg0KDQpQbGVhc2Ugbm90ZSB0aGF0IGl0IG1h
eSB0YWtlIGEgY291cGxlIG9mIG1pbnV0ZXMgZnJvbSB0aGUgdGltZSBvZiBzdWJtaXNzaW9uIHVu
dGlsIHRoZSBodG1saXplZCB2ZXJzaW9uIGFuZCBkaWZmIGFyZSBhdmFpbGFibGUgYXQgdG9vbHMu
aWV0Zi5vcmcuDQoNClRoZSBJRVRGIFNlY3JldGFyaWF0DQoNCg==


From nobody Fri Apr 11 06:52:52 2014
Return-Path: <coverdale@sympatico.ca>
X-Original-To: clue@ietfa.amsl.com
Delivered-To: clue@ietfa.amsl.com
Received: from localhost (ietfa.amsl.com [127.0.0.1]) by ietfa.amsl.com (Postfix) with ESMTP id 33D6D1A0695 for <clue@ietfa.amsl.com>; Fri, 11 Apr 2014 06:52:49 -0700 (PDT)
X-Virus-Scanned: amavisd-new at amsl.com
X-Spam-Flag: NO
X-Spam-Score: 0.8
X-Spam-Level: 
X-Spam-Status: No, score=0.8 tagged_above=-999 required=5 tests=[BAYES_50=0.8,  MSGID_FROM_MTA_HEADER=0.001, RCVD_IN_DNSWL_NONE=-0.0001, SPF_PASS=-0.001] autolearn=ham
Received: from mail.ietf.org ([4.31.198.44]) by localhost (ietfa.amsl.com [127.0.0.1]) (amavisd-new, port 10024) with ESMTP id 08at1DDXRyJB for <clue@ietfa.amsl.com>; Fri, 11 Apr 2014 06:52:47 -0700 (PDT)
Received: from blu0-omc1-s27.blu0.hotmail.com (blu0-omc1-s27.blu0.hotmail.com [65.55.116.38]) by ietfa.amsl.com (Postfix) with ESMTP id EDAD91A0694 for <clue@ietf.org>; Fri, 11 Apr 2014 06:52:46 -0700 (PDT)
Received: from BLU0-SMTP62 ([65.55.116.9]) by blu0-omc1-s27.blu0.hotmail.com with Microsoft SMTPSVC(6.0.3790.4675);  Fri, 11 Apr 2014 06:52:45 -0700
X-TMN: [uUaNmAzfK3aT7cOfr2OG9Kjq/jMUdqwu]
X-Originating-Email: [coverdale@sympatico.ca]
Message-ID: <BLU0-SMTP625327C11EEA9F47C744A1D0540@phx.gbl>
Received: from PaulNewPC ([74.15.60.251]) by BLU0-SMTP62.phx.gbl over TLS secured channel with Microsoft SMTPSVC(6.0.3790.4675);  Fri, 11 Apr 2014 06:52:45 -0700
From: Paul Coverdale <coverdale@sympatico.ca>
To: "'John Leslie'" <john@jlc.net>, <clue@ietf.org>
References: <533AF351.9050201@alum.mit.edu> <20140410222705.GW39240@verdi> <BLU0-SMTP6910DA20C9A9287137B585D0550@phx.gbl> <20140410234249.GX39240@verdi>
In-Reply-To: <20140410234249.GX39240@verdi>
Date: Fri, 11 Apr 2014 09:52:34 -0400
MIME-Version: 1.0
Content-Type: text/plain; charset="us-ascii"
Content-Transfer-Encoding: 7bit
X-Mailer: Microsoft Office Outlook 12.0
Thread-Index: Ac9VFpaFCkqO54LvRUCePVPgSfv4BAAa/uGA
Content-Language: en-us
X-OriginalArrivalTime: 11 Apr 2014 13:52:45.0157 (UTC) FILETIME=[50CFA950:01CF558D]
Archived-At: http://mailarchive.ietf.org/arch/msg/clue/-JZxBrxXGwsGqIaSD9f-k37ngQ0
Subject: Re: [clue] Improving treatment of audio
X-BeenThere: clue@ietf.org
X-Mailman-Version: 2.1.15
Precedence: list
List-Id: CLUE - ControLling mUltiple streams for TElepresence <clue.ietf.org>
List-Unsubscribe: <https://www.ietf.org/mailman/options/clue>, <mailto:clue-request@ietf.org?subject=unsubscribe>
List-Archive: <http://www.ietf.org/mail-archive/web/clue/>
List-Post: <mailto:clue@ietf.org>
List-Help: <mailto:clue-request@ietf.org?subject=help>
List-Subscribe: <https://www.ietf.org/mailman/listinfo/clue>, <mailto:clue-request@ietf.org?subject=subscribe>
X-List-Received-Date: Fri, 11 Apr 2014 13:52:49 -0000

>> It takes me back to my old days in the Audio & Acoustics group at
>> BNR/Nortel when we were trying to figure out what to do about acoustic
>> echo and de-reverberation. But those problems have pretty well been
>> solved for audio conferencing systems.
>
>   They have indeed been "solved" for POTS. I am less convinced they
>have been solved for audio-conferencing systems -- though I'd be happy
>to be convinced...

Well, I think that there are some quite sophisticated acoustic echo
cancellers and even de-reverberation algorithms on the market nowadays, so I
believe that the potential is certainly there to give a high-quality audio
experience within Telepresence. Having said that, the implementers of these
systems need to have some skill in setting them up to work correctly.
However, I think this goes beyond the scope of CLUE.


>   But IMHO, prior solutions can't simply be tacked on -- we need to
>recognize where these problems come from; then we need to evaluate what
>new difficulties arise in our immersion-teleconferencing setting.
>
>   I don't mean to _solve_ these problems here; but I do want to make
>sufficient tools available to the folks that will have to solve them:
>preferably early in CLUE's rollout, but at worst "eventually".
>
>   Recall that we will have to deal with wildly varying delays and
>jitter, and that the human ear's expectations become more discriminating
>when trying to match the imagined room. Et cetera...
>
Don't disagree, but it's still not clear to me what specifically needs to be
done in a CLUE context. 


>> So it's not clear to me what extra we need to do from a CLUE
>> perspective.
>
>   That's the question, isn't it. I expect we will need to better inform
>echo-cancellation about the varying delays, for one thing.
>
Well, it's not going to be easy to figure out the end to end delay, since it
will depend on the terminal equipment and the network transport. But even if
we had such a figure, I'm not sure what the acoustic echo-canceller would do
with it. Basically, the canceller plus non-linear processor tries to reduce
the level of any echo to below the level of audibility, regardless of the
end to end delay. The only delay it needs to worry about is in the
tail-path, which is at the local end anyway.





From nobody Fri Apr 11 07:25:07 2014
Return-Path: <coverdale@sympatico.ca>
X-Original-To: clue@ietfa.amsl.com
Delivered-To: clue@ietfa.amsl.com
Received: from localhost (ietfa.amsl.com [127.0.0.1]) by ietfa.amsl.com (Postfix) with ESMTP id C947C1A06C8 for <clue@ietfa.amsl.com>; Fri, 11 Apr 2014 07:25:04 -0700 (PDT)
X-Virus-Scanned: amavisd-new at amsl.com
X-Spam-Flag: NO
X-Spam-Score: -0.001
X-Spam-Level: 
X-Spam-Status: No, score=-0.001 tagged_above=-999 required=5 tests=[BAYES_20=-0.001, MSGID_FROM_MTA_HEADER=0.001, RCVD_IN_DNSWL_NONE=-0.0001, SPF_PASS=-0.001] autolearn=ham
Received: from mail.ietf.org ([4.31.198.44]) by localhost (ietfa.amsl.com [127.0.0.1]) (amavisd-new, port 10024) with ESMTP id daeQQAg-1PZf for <clue@ietfa.amsl.com>; Fri, 11 Apr 2014 07:25:03 -0700 (PDT)
Received: from blu0-omc1-s7.blu0.hotmail.com (blu0-omc1-s7.blu0.hotmail.com [65.55.116.18]) by ietfa.amsl.com (Postfix) with ESMTP id 1570A1A06B2 for <clue@ietf.org>; Fri, 11 Apr 2014 07:25:02 -0700 (PDT)
Received: from BLU0-SMTP72 ([65.55.116.8]) by blu0-omc1-s7.blu0.hotmail.com with Microsoft SMTPSVC(6.0.3790.4675);  Fri, 11 Apr 2014 07:25:01 -0700
X-TMN: [uJquoFrXMKlw0I5A/+9/Jspejux3TXih]
X-Originating-Email: [coverdale@sympatico.ca]
Message-ID: <BLU0-SMTP72E8824027ED6662C24095D0540@phx.gbl>
Received: from PaulNewPC ([74.15.60.251]) by BLU0-SMTP72.phx.gbl over TLS secured channel with Microsoft SMTPSVC(6.0.3790.4675);  Fri, 11 Apr 2014 07:25:01 -0700
From: Paul Coverdale <coverdale@sympatico.ca>
To: "'Christian Groves'" <Christian.Groves@nteczone.com>, <clue@ietf.org>
References: <533AF351.9050201@alum.mit.edu> <20140410222705.GW39240@verdi> <BLU0-SMTP6910DA20C9A9287137B585D0550@phx.gbl> <53475572.4040807@nteczone.com>
In-Reply-To: <53475572.4040807@nteczone.com>
Date: Fri, 11 Apr 2014 10:24:51 -0400
MIME-Version: 1.0
Content-Type: text/plain; charset="us-ascii"
Content-Transfer-Encoding: 7bit
X-Mailer: Microsoft Office Outlook 12.0
Thread-Index: Ac9VLw74hQF7b1cCS2W3F7psbhpyMwAVIZqQ
Content-Language: en-us
X-OriginalArrivalTime: 11 Apr 2014 14:25:01.0197 (UTC) FILETIME=[D2C7EBD0:01CF5591]
Archived-At: http://mailarchive.ietf.org/arch/msg/clue/T7K7voiiOLC9czzqIr-y3Sd32RU
Subject: Re: [clue] Improving treatment of audio
X-BeenThere: clue@ietf.org
X-Mailman-Version: 2.1.15
Precedence: list
List-Id: CLUE - ControLling mUltiple streams for TElepresence <clue.ietf.org>
List-Unsubscribe: <https://www.ietf.org/mailman/options/clue>, <mailto:clue-request@ietf.org?subject=unsubscribe>
List-Archive: <http://www.ietf.org/mail-archive/web/clue/>
List-Post: <mailto:clue@ietf.org>
List-Help: <mailto:clue-request@ietf.org?subject=help>
List-Subscribe: <https://www.ietf.org/mailman/listinfo/clue>, <mailto:clue-request@ietf.org?subject=subscribe>
X-List-Received-Date: Fri, 11 Apr 2014 14:25:04 -0000

>-----Original Message-----
>From: clue [mailto:clue-bounces@ietf.org] On Behalf Of Christian Groves
>Sent: Thursday, April 10, 2014 10:38 PM
>To: clue@ietf.org
>Subject: Re: [clue] Improving treatment of audio
>
>"So it's not clear to me what extra we need to do from a CLUE
>perspective. What is the problem?"
>I guess this is the pertinent point.
>
> From John's Audio 101 it seems that anything to do with spatial
>information that would relate to reverberation is in the too hard
>basket. So it seems "area of capture" for audio could be marked "not
>applicable" in the framework. There doesn't seem to be any driver for
>having the "point of capture" apply to an audio capture either.
>
> From the Audio101 there does seem to be a dependency on where the
>microphone is located with respect to the person speaking (i.e. lapel
>mic, desk mic, room mic) to how it is handled at mixing/playout. Perhaps
>this is useful to signal via CLUE? If it is possible to signal this them
>perhaps tying a particular audio capture to a video capture makes sense?
>i.e. a talker giving a presentation using a lapel mic captured by a
>particular video.
>


Yes, I think we will be opening a can of worms if we try to get too fancy
about signalling spatial audio information. However, I think that as a
minimum we need to signal the position of each microphone, its pattern (and
axis of directivity if directional), and its relation to the audio capture
and video capture.


...Paul






From nobody Fri Apr 11 08:01:05 2014
Return-Path: <john@jlc.net>
X-Original-To: clue@ietfa.amsl.com
Delivered-To: clue@ietfa.amsl.com
Received: from localhost (ietfa.amsl.com [127.0.0.1]) by ietfa.amsl.com (Postfix) with ESMTP id BB2551A026B for <clue@ietfa.amsl.com>; Fri, 11 Apr 2014 08:01:03 -0700 (PDT)
X-Virus-Scanned: amavisd-new at amsl.com
X-Spam-Flag: NO
X-Spam-Score: -1.772
X-Spam-Level: 
X-Spam-Status: No, score=-1.772 tagged_above=-999 required=5 tests=[BAYES_50=0.8, RCVD_IN_DNSWL_MED=-2.3, RP_MATCHES_RCVD=-0.272] autolearn=ham
Received: from mail.ietf.org ([4.31.198.44]) by localhost (ietfa.amsl.com [127.0.0.1]) (amavisd-new, port 10024) with ESMTP id 0igqjvFjrjCA for <clue@ietfa.amsl.com>; Fri, 11 Apr 2014 08:00:59 -0700 (PDT)
Received: from mailhost.jlc.net (mailhost.jlc.net [199.201.159.4]) by ietfa.amsl.com (Postfix) with ESMTP id B33A21A03D3 for <clue@ietf.org>; Fri, 11 Apr 2014 08:00:58 -0700 (PDT)
Received: by mailhost.jlc.net (Postfix, from userid 104) id 8010AC94BE; Fri, 11 Apr 2014 11:00:55 -0400 (EDT)
Date: Fri, 11 Apr 2014 11:00:55 -0400
From: John Leslie <john@jlc.net>
To: Christian Groves <Christian.Groves@nteczone.com>
Message-ID: <20140411150055.GE60844@verdi>
References: <533AF351.9050201@alum.mit.edu> <20140410222705.GW39240@verdi> <BLU0-SMTP6910DA20C9A9287137B585D0550@phx.gbl> <53475572.4040807@nteczone.com>
Mime-Version: 1.0
Content-Type: text/plain; charset=us-ascii
Content-Disposition: inline
In-Reply-To: <53475572.4040807@nteczone.com>
User-Agent: Mutt/1.4.1i
Archived-At: http://mailarchive.ietf.org/arch/msg/clue/THtnxxF15YBg6cfPDzeOaaonXZI
Cc: clue@ietf.org
Subject: Re: [clue] Improving treatment of audio
X-BeenThere: clue@ietf.org
X-Mailman-Version: 2.1.15
Precedence: list
List-Id: CLUE - ControLling mUltiple streams for TElepresence <clue.ietf.org>
List-Unsubscribe: <https://www.ietf.org/mailman/options/clue>, <mailto:clue-request@ietf.org?subject=unsubscribe>
List-Archive: <http://www.ietf.org/mail-archive/web/clue/>
List-Post: <mailto:clue@ietf.org>
List-Help: <mailto:clue-request@ietf.org?subject=help>
List-Subscribe: <https://www.ietf.org/mailman/listinfo/clue>, <mailto:clue-request@ietf.org?subject=subscribe>
X-List-Received-Date: Fri, 11 Apr 2014 15:01:03 -0000

Christian Groves <Christian.Groves@nteczone.com> wrote:
> 
> "So it's not clear to me what extra we need to do from a CLUE 
> perspective. What is the problem?"
> I guess this is the pertinent point.
> 
> From John's Audio 101 it seems that anything to do with spatial 
> information that would relate to reverberation is in the too hard 
> basket.

   What precisely does "too hard basket" mean?

   I haven't even reached the point of suggesting what metrics to
define the _ability_ to send. To me, "too hard" merely means that
an individual site could choose not to send them or to ignore them
on receipt.

   But it sounds as if you're suggesting reverberation is "to hard
to understand" and thus we should have no metrics about it.

   I _hope_ that's not what you mean.

> So it seems "area of capture" for audio could be marked "not 
> applicable" in the framework.

   I hope so.

> There doesn't seem to be any driver for having the "point of capture"
> apply to an audio capture either.

   I don't understand this. Point of capture for a microphone may be
"too hard" to track (today) for a microphone which moves; nonetheless
it seems to me the most fundamental metric there could be.

> From the Audio101 there does seem to be a dependency on where the 
> microphone is located with respect to the person speaking (i.e. lapel 
> mic, desk mic, room mic) to how it is handled at mixing/playout.

   I don't think I really got that far...

   IMHO, it's more flexible to have mixing be the responsibily of
the receiver, but I haven't tried to specify that and I'm not at all
sure I want to specify that.

   I expect the actual sound systems in different rooms to vary wildly,
from one monaural speaker to stereo to full surround-sound. Mixing
for these without knowing which is the actual target seems hard; but
I'm sure there will be sites which prefer to do so. A question which
will arise, IMHO, is how to specify the _intent_ of a mix generated
in one room to be fed to other rooms. (I'd prefer not to go there yet.)

> Perhaps this is useful to signal via CLUE? If it is possible to signal
> this then perhaps tying a particular audio capture to a video capture
> makes sense?

   I'm not thinking along those lines. (That doesn't mean we shouldn't
think along those lines...) I'm thinking in terms of providing several
audio streams per room, associated with position information about
the position of the source of those sounds, and allowing the receiver
to choose how to mix them and how to present the mix in his/her room.

> i.e. a talker giving a presentation using a lapel mic captured by a 
> particular video.

   In fact, there's only limited tendency for humans to strictly
attach the sound they hear to the video they see. Clearly, during a
presentation we want to _hear_ the presenter talking, but having the
sound move back and forth as the speaker walks can be confusing.

   At the same time, we will want to hear the questions to which the
presenter may respond. It will be far easier on the listener if these
do _not_ seem to be coming from the same physical position, especially
when the questioner _is_ in the same room as the presenter.

   Hope this helps...

--
John Leslie <john@jlc.net>


From nobody Fri Apr 11 08:14:50 2014
Return-Path: <john@jlc.net>
X-Original-To: clue@ietfa.amsl.com
Delivered-To: clue@ietfa.amsl.com
Received: from localhost (ietfa.amsl.com [127.0.0.1]) by ietfa.amsl.com (Postfix) with ESMTP id F1B391A02C7 for <clue@ietfa.amsl.com>; Fri, 11 Apr 2014 08:14:46 -0700 (PDT)
X-Virus-Scanned: amavisd-new at amsl.com
X-Spam-Flag: NO
X-Spam-Score: -2.573
X-Spam-Level: 
X-Spam-Status: No, score=-2.573 tagged_above=-999 required=5 tests=[BAYES_20=-0.001, RCVD_IN_DNSWL_MED=-2.3, RP_MATCHES_RCVD=-0.272] autolearn=ham
Received: from mail.ietf.org ([4.31.198.44]) by localhost (ietfa.amsl.com [127.0.0.1]) (amavisd-new, port 10024) with ESMTP id mDhFqbnoEbdj for <clue@ietfa.amsl.com>; Fri, 11 Apr 2014 08:14:38 -0700 (PDT)
Received: from mailhost.jlc.net (mailhost.jlc.net [199.201.159.4]) by ietfa.amsl.com (Postfix) with ESMTP id B403B1A025A for <clue@ietf.org>; Fri, 11 Apr 2014 08:14:08 -0700 (PDT)
Received: by mailhost.jlc.net (Postfix, from userid 104) id BB641C94EB; Fri, 11 Apr 2014 11:14:05 -0400 (EDT)
Date: Fri, 11 Apr 2014 11:14:05 -0400
From: John Leslie <john@jlc.net>
To: Paul Coverdale <coverdale@sympatico.ca>
Message-ID: <20140411151405.GF60844@verdi>
References: <533AF351.9050201@alum.mit.edu> <20140410222705.GW39240@verdi> <BLU0-SMTP6910DA20C9A9287137B585D0550@phx.gbl> <53475572.4040807@nteczone.com> <BLU0-SMTP72E8824027ED6662C24095D0540@phx.gbl>
Mime-Version: 1.0
Content-Type: text/plain; charset=us-ascii
Content-Disposition: inline
In-Reply-To: <BLU0-SMTP72E8824027ED6662C24095D0540@phx.gbl>
User-Agent: Mutt/1.4.1i
Archived-At: http://mailarchive.ietf.org/arch/msg/clue/10aH-hAcjnVbCZ7ryDxFqUehq04
Cc: clue@ietf.org
Subject: Re: [clue] Improving treatment of audio
X-BeenThere: clue@ietf.org
X-Mailman-Version: 2.1.15
Precedence: list
List-Id: CLUE - ControLling mUltiple streams for TElepresence <clue.ietf.org>
List-Unsubscribe: <https://www.ietf.org/mailman/options/clue>, <mailto:clue-request@ietf.org?subject=unsubscribe>
List-Archive: <http://www.ietf.org/mail-archive/web/clue/>
List-Post: <mailto:clue@ietf.org>
List-Help: <mailto:clue-request@ietf.org?subject=help>
List-Subscribe: <https://www.ietf.org/mailman/listinfo/clue>, <mailto:clue-request@ietf.org?subject=subscribe>
X-List-Received-Date: Fri, 11 Apr 2014 15:14:47 -0000

Paul Coverdale <coverdale@sympatico.ca> wrote:
> 
> Yes, I think we will be opening a can of worms if we try to get too
> fancy about signalling spatial audio information.

   There are worm-cans strewn about, I agree...

   But, even if we tend to be minimalist for the initial spec, I believe
we will need to be able to expand the limits of what we _can_ specify.

> However, I think that as a minimum we need to signal the position of
> each microphone,

   Yes, but there will be microphones which move (especially lapel
microphones of presenters). There is a real question how to specify
such motion (which I'm not in a hurry to get to yet).

> its pattern (and axis of directivity if directional),

   Pattern, to tell truth, always is related to axis of capture for
physically-achievable microphones. (I suspect, however, that sometimes
the patter will simply be "unknown". This isn't necessarily fatal.)

> and its relation to the audio capture and video capture.

   I wouldn't know how to specify "relation to the audio capture"
(or even what a "the audio capture" might be). "Relation to the video
capture" seems not worth the trouble, at first blush...

   When a sender _chooses_ to send a mix rather than individual audio
sources, I haven't given a lot of thought to how to describe the mix.
My initial guess would be to reference the sources (which hopefully
have a physical position) and the relative gain settings.

   To me, a mix doesn't really _have_ a physical position.

   Hope this helps...

--
John Leslie <john@jlc.net>


From nobody Fri Apr 11 10:56:43 2014
Return-Path: <coverdale@sympatico.ca>
X-Original-To: clue@ietfa.amsl.com
Delivered-To: clue@ietfa.amsl.com
Received: from localhost (ietfa.amsl.com [127.0.0.1]) by ietfa.amsl.com (Postfix) with ESMTP id 1B9C11A0761 for <clue@ietfa.amsl.com>; Fri, 11 Apr 2014 10:56:41 -0700 (PDT)
X-Virus-Scanned: amavisd-new at amsl.com
X-Spam-Flag: NO
X-Spam-Score: -0.001
X-Spam-Level: 
X-Spam-Status: No, score=-0.001 tagged_above=-999 required=5 tests=[BAYES_20=-0.001, MSGID_FROM_MTA_HEADER=0.001, RCVD_IN_DNSWL_NONE=-0.0001, SPF_PASS=-0.001] autolearn=ham
Received: from mail.ietf.org ([4.31.198.44]) by localhost (ietfa.amsl.com [127.0.0.1]) (amavisd-new, port 10024) with ESMTP id Wu3apo_S6TEi for <clue@ietfa.amsl.com>; Fri, 11 Apr 2014 10:56:39 -0700 (PDT)
Received: from blu0-omc1-s23.blu0.hotmail.com (blu0-omc1-s23.blu0.hotmail.com [65.55.116.34]) by ietfa.amsl.com (Postfix) with ESMTP id 07AB51A0758 for <clue@ietf.org>; Fri, 11 Apr 2014 10:56:37 -0700 (PDT)
Received: from BLU0-SMTP76 ([65.55.116.8]) by blu0-omc1-s23.blu0.hotmail.com with Microsoft SMTPSVC(6.0.3790.4675);  Fri, 11 Apr 2014 10:56:36 -0700
X-TMN: [GNev0O1h9tQ9LHaei4YX6c+cD2aJUDF5]
X-Originating-Email: [coverdale@sympatico.ca]
Message-ID: <BLU0-SMTP7680549C7179BB61E4F216D0540@phx.gbl>
Received: from PaulNewPC ([74.15.60.251]) by BLU0-SMTP76.phx.gbl over TLS secured channel with Microsoft SMTPSVC(6.0.3790.4675);  Fri, 11 Apr 2014 10:56:35 -0700
From: Paul Coverdale <coverdale@sympatico.ca>
To: "'John Leslie'" <john@jlc.net>
References: <533AF351.9050201@alum.mit.edu> <20140410222705.GW39240@verdi> <BLU0-SMTP6910DA20C9A9287137B585D0550@phx.gbl> <53475572.4040807@nteczone.com> <BLU0-SMTP72E8824027ED6662C24095D0540@phx.gbl> <20140411151405.GF60844@verdi>
In-Reply-To: <20140411151405.GF60844@verdi>
Date: Fri, 11 Apr 2014 13:56:25 -0400
MIME-Version: 1.0
Content-Type: text/plain; charset="us-ascii"
Content-Transfer-Encoding: 7bit
X-Mailer: Microsoft Office Outlook 12.0
Thread-Index: Ac9VmK6do3LspRN6TH6QdxLf+g/yoAABnw6Q
Content-Language: en-us
X-OriginalArrivalTime: 11 Apr 2014 17:56:35.0825 (UTC) FILETIME=[615E6E10:01CF55AF]
Archived-At: http://mailarchive.ietf.org/arch/msg/clue/-TC8OQU0PmvVRExdw_n4TWTVm8M
Cc: clue@ietf.org
Subject: Re: [clue] Improving treatment of audio
X-BeenThere: clue@ietf.org
X-Mailman-Version: 2.1.15
Precedence: list
List-Id: CLUE - ControLling mUltiple streams for TElepresence <clue.ietf.org>
List-Unsubscribe: <https://www.ietf.org/mailman/options/clue>, <mailto:clue-request@ietf.org?subject=unsubscribe>
List-Archive: <http://www.ietf.org/mail-archive/web/clue/>
List-Post: <mailto:clue@ietf.org>
List-Help: <mailto:clue-request@ietf.org?subject=help>
List-Subscribe: <https://www.ietf.org/mailman/listinfo/clue>, <mailto:clue-request@ietf.org?subject=subscribe>
X-List-Received-Date: Fri, 11 Apr 2014 17:56:41 -0000

>-----Original Message-----
>From: John Leslie [mailto:john@jlc.net]
>Sent: Friday, April 11, 2014 11:14 AM
>To: Paul Coverdale
>Cc: clue@ietf.org
>Subject: Re: [clue] Improving treatment of audio
>
>> However, I think that as a minimum we need to signal the position of
>> each microphone,
>
>   Yes, but there will be microphones which move (especially lapel
>microphones of presenters). There is a real question how to specify such
>motion (which I'm not in a hurry to get to yet).

Yes, it would be difficult to track the motion, but I don't think we need
to. The lapel microphone is likely to be the best option for picking up
audio, giving (hopefully) a good signal to noise ratio regardless of where
the user is. All we need to do is associate it with the appropriate video
capture.

>> its pattern (and axis of directivity if directional),
>
>   Pattern, to tell truth, always is related to axis of capture for
>physically-achievable microphones. (I suspect, however, that sometimes
>the patter will simply be "unknown". This isn't necessarily fatal.)

Maybe I didn't get the terminology right, but I was thinking of describing
the microphone position in terms of Point of Capture (7.1.1.1 in Framework).
By pattern, I meant the inherent microphone pattern (in free-space), eg
omni, cardioid. In the case of a directional microphone, the direction of
the major lobe could be described using Point of Capture and Point on Line
of Capture (7.1.1.2). I agree that Area of Capture has little meaning for
audio. As you say, in some cases the pattern may be unknown. In this case,
the default would be omni.


>> and its relation to the audio capture and video capture.
>
>   I wouldn't know how to specify "relation to the audio capture"
>(or even what a "the audio capture" might be). "Relation to the video
>capture" seems not worth the trouble, at first blush...

Again, maybe I didn't use the right terminology. What I meant was that a
given microphone needs to be associated with a particular audio capture, and
this audio capture is normally associated with a specific video capture eg,
you hear what the guy in the picture is saying. But I suppose there could be
situations where the audio capture has no associated video eg, the guy
doesn't want anyone to see the wart on his nose. And there may be situations
where the same audio capture is associated with more than one video capture,
ie several subjects are sharing the same microphone.

>   When a sender _chooses_ to send a mix rather than individual audio
>sources, I haven't given a lot of thought to how to describe the mix.
>My initial guess would be to reference the sources (which hopefully have
>a physical position) and the relative gain settings.
>
>   To me, a mix doesn't really _have_ a physical position.

Agreed, in the case of a mix from more than microphone, just reference the
individual microphone locations and the relative gain settings.


>   Hope this helps...

Yes, it does.



From nobody Fri Apr 11 12:24:37 2014
Return-Path: <Mark.Duckworth@polycom.com>
X-Original-To: clue@ietfa.amsl.com
Delivered-To: clue@ietfa.amsl.com
Received: from localhost (ietfa.amsl.com [127.0.0.1]) by ietfa.amsl.com (Postfix) with ESMTP id 482971A0395 for <clue@ietfa.amsl.com>; Fri, 11 Apr 2014 12:24:36 -0700 (PDT)
X-Virus-Scanned: amavisd-new at amsl.com
X-Spam-Flag: NO
X-Spam-Score: 0.28
X-Spam-Level: 
X-Spam-Status: No, score=0.28 tagged_above=-999 required=5 tests=[BAYES_05=-0.5, RCVD_IN_DNSWL_NONE=-0.0001, SPF_NEUTRAL=0.779, UNPARSEABLE_RELAY=0.001] autolearn=no
Received: from mail.ietf.org ([4.31.198.44]) by localhost (ietfa.amsl.com [127.0.0.1]) (amavisd-new, port 10024) with ESMTP id cdKXPsgpmC2I for <clue@ietfa.amsl.com>; Fri, 11 Apr 2014 12:24:32 -0700 (PDT)
Received: from mail1.bemta8.messagelabs.com (mail1.bemta8.messagelabs.com [216.82.243.205]) by ietfa.amsl.com (Postfix) with ESMTP id 1C4F41A034A for <clue@ietf.org>; Fri, 11 Apr 2014 12:24:31 -0700 (PDT)
Received: from [216.82.241.100:11137] by server-13.bemta-8.messagelabs.com id 8F/45-14006-D6148435; Fri, 11 Apr 2014 19:24:29 +0000
X-Env-Sender: Mark.Duckworth@polycom.com
X-Msg-Ref: server-10.tower-220.messagelabs.com!1397244262!6257195!7
X-Originating-IP: [140.242.64.158]
X-StarScan-Received: 
X-StarScan-Version: 6.11.1; banners=-,-,-
X-VirusChecked: Checked
Received: (qmail 31480 invoked from network); 11 Apr 2014 19:24:28 -0000
Received: from crpehubprd01.polycom.com (HELO crpehubprd02.polycom.com) (140.242.64.158) by server-10.tower-220.messagelabs.com with AES128-SHA encrypted SMTP; 11 Apr 2014 19:24:28 -0000
Received: from CRPMBOXPRD07.polycom.com ([fe80::91fc:8a0f:5258:aff0]) by crpehubprd02.polycom.com ([::1]) with mapi; Fri, 11 Apr 2014 12:23:56 -0700
From: "Duckworth, Mark" <Mark.Duckworth@polycom.com>
To: John Leslie <john@jlc.net>, Christian Groves <Christian.Groves@nteczone.com>
Date: Fri, 11 Apr 2014 12:23:55 -0700
Thread-Topic: [clue] Improving treatment of audio
Thread-Index: Ac9Vlt1dlqHYS7oTRTqUSFX4G036yQAJD77Q
Message-ID: <49E45C59CA48264997FEBFB29B6BC2D6215295A8AE@CRPMBOXPRD07.polycom.com>
References: <533AF351.9050201@alum.mit.edu> <20140410222705.GW39240@verdi> <BLU0-SMTP6910DA20C9A9287137B585D0550@phx.gbl> <53475572.4040807@nteczone.com> <20140411150055.GE60844@verdi>
In-Reply-To: <20140411150055.GE60844@verdi>
Accept-Language: en-US
Content-Language: en-US
X-MS-Has-Attach: 
X-MS-TNEF-Correlator: 
acceptlanguage: en-US
Content-Type: text/plain; charset="us-ascii"
Content-Transfer-Encoding: quoted-printable
MIME-Version: 1.0
Archived-At: http://mailarchive.ietf.org/arch/msg/clue/VjYilFL-FmR7Zvnf9o1jYjBYc8o
Cc: "clue@ietf.org" <clue@ietf.org>
Subject: Re: [clue] Improving treatment of audio
X-BeenThere: clue@ietf.org
X-Mailman-Version: 2.1.15
Precedence: list
List-Id: CLUE - ControLling mUltiple streams for TElepresence <clue.ietf.org>
List-Unsubscribe: <https://www.ietf.org/mailman/options/clue>, <mailto:clue-request@ietf.org?subject=unsubscribe>
List-Archive: <http://www.ietf.org/mail-archive/web/clue/>
List-Post: <mailto:clue@ietf.org>
List-Help: <mailto:clue-request@ietf.org?subject=help>
List-Subscribe: <https://www.ietf.org/mailman/listinfo/clue>, <mailto:clue-request@ietf.org?subject=subscribe>
X-List-Received-Date: Fri, 11 Apr 2014 19:24:36 -0000

Hi all,
The original intent of using "area of capture" for all media types, includi=
ng audio and video, was to provide a way to associate audio with video.  So=
 the consumer can choose captures that go together and render them together=
.  An example of this is in the framework document, section 12.1.1.  I stil=
l think this works fine for this purpose.

In that example, the consumer can choose VC0, VC1, VC2, AC0, AC1, and AC2. =
 AC0 and VC0 have the same area of capture, indicating they are capturing t=
he same area of the scene.  If the consumer wants to render AC0 from a loud=
speaker close to the display for VC0 it can do so.  Similarly, AC3 has an a=
rea of capture covering the whole scene (the full extent of the areas of VC=
0, VC1, VC2) so the consumer knows AC3 includes audio associated with all o=
f VC0, VC1, and VC2.

So I'm puzzled why you are proposing we remove the ability to use area of c=
apture for audio for this purpose.

Regards,
Mark

> -----Original Message-----
> From: clue [mailto:clue-bounces@ietf.org] On Behalf Of John Leslie
> Sent: Friday, April 11, 2014 11:01 AM
> To: Christian Groves
> Cc: clue@ietf.org
> Subject: Re: [clue] Improving treatment of audio
>=20
> Christian Groves <Christian.Groves@nteczone.com> wrote:
> >
> > "So it's not clear to me what extra we need to do from a CLUE
> > perspective. What is the problem?"
> > I guess this is the pertinent point.
> >
> > From John's Audio 101 it seems that anything to do with spatial
> > information that would relate to reverberation is in the too hard
> > basket.
>=20
>    What precisely does "too hard basket" mean?
>=20
>    I haven't even reached the point of suggesting what metrics to define =
the
> _ability_ to send. To me, "too hard" merely means that an individual site
> could choose not to send them or to ignore them on receipt.
>=20
>    But it sounds as if you're suggesting reverberation is "to hard to
> understand" and thus we should have no metrics about it.
>=20
>    I _hope_ that's not what you mean.
>=20
> > So it seems "area of capture" for audio could be marked "not
> > applicable" in the framework.
>=20
>    I hope so.
>=20
> > There doesn't seem to be any driver for having the "point of capture"
> > apply to an audio capture either.
>=20
>    I don't understand this. Point of capture for a microphone may be "too
> hard" to track (today) for a microphone which moves; nonetheless it seems
> to me the most fundamental metric there could be.
>=20
> > From the Audio101 there does seem to be a dependency on where the
> > microphone is located with respect to the person speaking (i.e. lapel
> > mic, desk mic, room mic) to how it is handled at mixing/playout.
>=20
>    I don't think I really got that far...
>=20
>    IMHO, it's more flexible to have mixing be the responsibily of the rec=
eiver,
> but I haven't tried to specify that and I'm not at all sure I want to spe=
cify that.
>=20
>    I expect the actual sound systems in different rooms to vary wildly, f=
rom
> one monaural speaker to stereo to full surround-sound. Mixing for these
> without knowing which is the actual target seems hard; but I'm sure there
> will be sites which prefer to do so. A question which will arise, IMHO, i=
s how
> to specify the _intent_ of a mix generated in one room to be fed to other
> rooms. (I'd prefer not to go there yet.)
>=20
> > Perhaps this is useful to signal via CLUE? If it is possible to signal
> > this then perhaps tying a particular audio capture to a video capture
> > makes sense?
>=20
>    I'm not thinking along those lines. (That doesn't mean we shouldn't th=
ink
> along those lines...) I'm thinking in terms of providing several audio st=
reams
> per room, associated with position information about the position of the
> source of those sounds, and allowing the receiver to choose how to mix
> them and how to present the mix in his/her room.
>=20
> > i.e. a talker giving a presentation using a lapel mic captured by a
> > particular video.
>=20
>    In fact, there's only limited tendency for humans to strictly attach t=
he sound
> they hear to the video they see. Clearly, during a presentation we want t=
o
> _hear_ the presenter talking, but having the sound move back and forth as
> the speaker walks can be confusing.
>=20
>    At the same time, we will want to hear the questions to which the
> presenter may respond. It will be far easier on the listener if these do =
_not_
> seem to be coming from the same physical position, especially when the
> questioner _is_ in the same room as the presenter.
>=20
>    Hope this helps...
>=20
> --
> John Leslie <john@jlc.net>
>=20
> _______________________________________________
> clue mailing list
> clue@ietf.org
> https://www.ietf.org/mailman/listinfo/clue


From nobody Fri Apr 11 13:48:44 2014
Return-Path: <coverdale@sympatico.ca>
X-Original-To: clue@ietfa.amsl.com
Delivered-To: clue@ietfa.amsl.com
Received: from localhost (ietfa.amsl.com [127.0.0.1]) by ietfa.amsl.com (Postfix) with ESMTP id 04E971A0392 for <clue@ietfa.amsl.com>; Fri, 11 Apr 2014 13:48:44 -0700 (PDT)
X-Virus-Scanned: amavisd-new at amsl.com
X-Spam-Flag: NO
X-Spam-Score: -0.001
X-Spam-Level: 
X-Spam-Status: No, score=-0.001 tagged_above=-999 required=5 tests=[BAYES_40=-0.001, MSGID_FROM_MTA_HEADER=0.001, RCVD_IN_DNSWL_NONE=-0.0001, SPF_PASS=-0.001] autolearn=ham
Received: from mail.ietf.org ([4.31.198.44]) by localhost (ietfa.amsl.com [127.0.0.1]) (amavisd-new, port 10024) with ESMTP id KPHpe7VQbeMZ for <clue@ietfa.amsl.com>; Fri, 11 Apr 2014 13:48:40 -0700 (PDT)
Received: from blu0-omc1-s16.blu0.hotmail.com (blu0-omc1-s16.blu0.hotmail.com [65.55.116.27]) by ietfa.amsl.com (Postfix) with ESMTP id DB3041A076B for <clue@ietf.org>; Fri, 11 Apr 2014 13:48:39 -0700 (PDT)
Received: from BLU0-SMTP18 ([65.55.116.8]) by blu0-omc1-s16.blu0.hotmail.com with Microsoft SMTPSVC(6.0.3790.4675);  Fri, 11 Apr 2014 13:48:38 -0700
X-TMN: [tATgEhqhdigXqi50SNh+FXceKFTKjH8n]
X-Originating-Email: [coverdale@sympatico.ca]
Message-ID: <BLU0-SMTP1804A70E3E0F666EAEAE84D0540@phx.gbl>
Received: from PaulNewPC ([74.15.60.251]) by BLU0-SMTP18.phx.gbl over TLS secured channel with Microsoft SMTPSVC(6.0.3790.4675);  Fri, 11 Apr 2014 13:48:37 -0700
From: Paul Coverdale <coverdale@sympatico.ca>
To: "'Duckworth, Mark'" <Mark.Duckworth@polycom.com>, "'John Leslie'" <john@jlc.net>, "'Christian Groves'" <Christian.Groves@nteczone.com>
References: <533AF351.9050201@alum.mit.edu> <20140410222705.GW39240@verdi> <BLU0-SMTP6910DA20C9A9287137B585D0550@phx.gbl> <53475572.4040807@nteczone.com> <20140411150055.GE60844@verdi> <49E45C59CA48264997FEBFB29B6BC2D6215295A8AE@CRPMBOXPRD07.polycom.com>
In-Reply-To: <49E45C59CA48264997FEBFB29B6BC2D6215295A8AE@CRPMBOXPRD07.polycom.com>
Date: Fri, 11 Apr 2014 16:48:27 -0400
MIME-Version: 1.0
Content-Type: text/plain; charset="us-ascii"
Content-Transfer-Encoding: 7bit
X-Mailer: Microsoft Office Outlook 12.0
Thread-Index: Ac9Vlt1dlqHYS7oTRTqUSFX4G036yQAJD77QAABTnFA=
Content-Language: en-us
X-OriginalArrivalTime: 11 Apr 2014 20:48:37.0791 (UTC) FILETIME=[69BD72F0:01CF55C7]
Archived-At: http://mailarchive.ietf.org/arch/msg/clue/qdOuV5GFlIYOpWKYet_62iJCe5U
Cc: clue@ietf.org
Subject: Re: [clue] Improving treatment of audio
X-BeenThere: clue@ietf.org
X-Mailman-Version: 2.1.15
Precedence: list
List-Id: CLUE - ControLling mUltiple streams for TElepresence <clue.ietf.org>
List-Unsubscribe: <https://www.ietf.org/mailman/options/clue>, <mailto:clue-request@ietf.org?subject=unsubscribe>
List-Archive: <http://www.ietf.org/mail-archive/web/clue/>
List-Post: <mailto:clue@ietf.org>
List-Help: <mailto:clue-request@ietf.org?subject=help>
List-Subscribe: <https://www.ietf.org/mailman/listinfo/clue>, <mailto:clue-request@ietf.org?subject=subscribe>
X-List-Received-Date: Fri, 11 Apr 2014 20:48:44 -0000

Hi Mark,

I can see what you're trying to do in the example given in Framework section
12.1.1. The problem I have is that, because of the huge difference in
wavelength between light waves and audio waves, you can't really define an
area of capture for audio in the same way that you can for video. A camera
can focus on a scene defined by 4 co-planar X,Y,Z coordinates. It will
capture the video inside that quadrilateral, and nothing outside. A
microphone can capture audio coming from inside of the same quadrilateral,
but it will also capture a lot of audio from outside. What this means is
that we can never define a unique association of an audio capture with a
video capture based on pure geometrical considerations, but we can still
simply define that ACx is associated with VCx, but maybe with VCy and VCz
too.

...Paul

>-----Original Message-----
>From: clue [mailto:clue-bounces@ietf.org] On Behalf Of Duckworth, Mark
>Sent: Friday, April 11, 2014 3:24 PM
>To: John Leslie; Christian Groves
>Cc: clue@ietf.org
>Subject: Re: [clue] Improving treatment of audio
>
>Hi all,
>The original intent of using "area of capture" for all media types,
>including audio and video, was to provide a way to associate audio with
>video.  So the consumer can choose captures that go together and render
>them together.  An example of this is in the framework document, section
>12.1.1.  I still think this works fine for this purpose.
>
>In that example, the consumer can choose VC0, VC1, VC2, AC0, AC1, and
>AC2.  AC0 and VC0 have the same area of capture, indicating they are
>capturing the same area of the scene.  If the consumer wants to render
>AC0 from a loudspeaker close to the display for VC0 it can do so.
>Similarly, AC3 has an area of capture covering the whole scene (the full
>extent of the areas of VC0, VC1, VC2) so the consumer knows AC3 includes
>audio associated with all of VC0, VC1, and VC2.
>
>So I'm puzzled why you are proposing we remove the ability to use area
>of capture for audio for this purpose.
>
>Regards,
>Mark
>
>> -----Original Message-----
>> From: clue [mailto:clue-bounces@ietf.org] On Behalf Of John Leslie
>> Sent: Friday, April 11, 2014 11:01 AM
>> To: Christian Groves
>> Cc: clue@ietf.org
>> Subject: Re: [clue] Improving treatment of audio
>>
>> Christian Groves <Christian.Groves@nteczone.com> wrote:
>> >
>> > "So it's not clear to me what extra we need to do from a CLUE
>> > perspective. What is the problem?"
>> > I guess this is the pertinent point.
>> >
>> > From John's Audio 101 it seems that anything to do with spatial
>> > information that would relate to reverberation is in the too hard
>> > basket.
>>
>>    What precisely does "too hard basket" mean?
>>
>>    I haven't even reached the point of suggesting what metrics to
>> define the _ability_ to send. To me, "too hard" merely means that an
>> individual site could choose not to send them or to ignore them on
>receipt.
>>
>>    But it sounds as if you're suggesting reverberation is "to hard to
>> understand" and thus we should have no metrics about it.
>>
>>    I _hope_ that's not what you mean.
>>
>> > So it seems "area of capture" for audio could be marked "not
>> > applicable" in the framework.
>>
>>    I hope so.
>>
>> > There doesn't seem to be any driver for having the "point of
>capture"
>> > apply to an audio capture either.
>>
>>    I don't understand this. Point of capture for a microphone may be
>> "too hard" to track (today) for a microphone which moves; nonetheless
>> it seems to me the most fundamental metric there could be.
>>
>> > From the Audio101 there does seem to be a dependency on where the
>> > microphone is located with respect to the person speaking (i.e.
>> > lapel mic, desk mic, room mic) to how it is handled at
>mixing/playout.
>>
>>    I don't think I really got that far...
>>
>>    IMHO, it's more flexible to have mixing be the responsibily of the
>> receiver, but I haven't tried to specify that and I'm not at all sure
>I want to specify that.
>>
>>    I expect the actual sound systems in different rooms to vary
>> wildly, from one monaural speaker to stereo to full surround-sound.
>> Mixing for these without knowing which is the actual target seems
>> hard; but I'm sure there will be sites which prefer to do so. A
>> question which will arise, IMHO, is how to specify the _intent_ of a
>> mix generated in one room to be fed to other rooms. (I'd prefer not to
>> go there yet.)
>>
>> > Perhaps this is useful to signal via CLUE? If it is possible to
>> > signal this then perhaps tying a particular audio capture to a video
>> > capture makes sense?
>>
>>    I'm not thinking along those lines. (That doesn't mean we shouldn't
>> think along those lines...) I'm thinking in terms of providing several
>> audio streams per room, associated with position information about the
>> position of the source of those sounds, and allowing the receiver to
>> choose how to mix them and how to present the mix in his/her room.
>>
>> > i.e. a talker giving a presentation using a lapel mic captured by a
>> > particular video.
>>
>>    In fact, there's only limited tendency for humans to strictly
>> attach the sound they hear to the video they see. Clearly, during a
>> presentation we want to _hear_ the presenter talking, but having the
>> sound move back and forth as the speaker walks can be confusing.
>>
>>    At the same time, we will want to hear the questions to which the
>> presenter may respond. It will be far easier on the listener if these
>> do _not_ seem to be coming from the same physical position, especially
>> when the questioner _is_ in the same room as the presenter.
>>
>>    Hope this helps...
>>
>> --
>> John Leslie <john@jlc.net>
>>
>> _______________________________________________
>> clue mailing list
>> clue@ietf.org
>> https://www.ietf.org/mailman/listinfo/clue
>
>_______________________________________________
>clue mailing list
>clue@ietf.org
>https://www.ietf.org/mailman/listinfo/clue


From nobody Fri Apr 11 14:42:42 2014
Return-Path: <Mark.Duckworth@polycom.com>
X-Original-To: clue@ietfa.amsl.com
Delivered-To: clue@ietfa.amsl.com
Received: from localhost (ietfa.amsl.com [127.0.0.1]) by ietfa.amsl.com (Postfix) with ESMTP id BFD741A01DF for <clue@ietfa.amsl.com>; Fri, 11 Apr 2014 14:42:40 -0700 (PDT)
X-Virus-Scanned: amavisd-new at amsl.com
X-Spam-Flag: NO
X-Spam-Score: -1.12
X-Spam-Level: 
X-Spam-Status: No, score=-1.12 tagged_above=-999 required=5 tests=[BAYES_00=-1.9, RCVD_IN_DNSWL_NONE=-0.0001, SPF_NEUTRAL=0.779, UNPARSEABLE_RELAY=0.001] autolearn=no
Received: from mail.ietf.org ([4.31.198.44]) by localhost (ietfa.amsl.com [127.0.0.1]) (amavisd-new, port 10024) with ESMTP id gM6VPya1mDdN for <clue@ietfa.amsl.com>; Fri, 11 Apr 2014 14:42:37 -0700 (PDT)
Received: from mail1.bemta8.messagelabs.com (mail1.bemta8.messagelabs.com [216.82.243.204]) by ietfa.amsl.com (Postfix) with ESMTP id 39E421A02AF for <clue@ietf.org>; Fri, 11 Apr 2014 14:42:28 -0700 (PDT)
Received: from [216.82.241.100:36049] by server-12.bemta-8.messagelabs.com id 4B/A8-31743-2C168435; Fri, 11 Apr 2014 21:42:26 +0000
X-Env-Sender: Mark.Duckworth@polycom.com
X-Msg-Ref: server-11.tower-220.messagelabs.com!1397252544!6293765!2
X-Originating-IP: [140.242.64.158]
X-StarScan-Received: 
X-StarScan-Version: 6.11.1; banners=-,-,-
X-VirusChecked: Checked
Received: (qmail 11063 invoked from network); 11 Apr 2014 21:42:26 -0000
Received: from crpehubprd01.polycom.com (HELO crpehubprd02.polycom.com) (140.242.64.158) by server-11.tower-220.messagelabs.com with AES128-SHA encrypted SMTP; 11 Apr 2014 21:42:26 -0000
Received: from CRPMBOXPRD07.polycom.com ([fe80::8113:9ad1:f9be:53f1]) by crpehubprd02.polycom.com ([::1]) with mapi; Fri, 11 Apr 2014 14:42:05 -0700
From: "Duckworth, Mark" <Mark.Duckworth@polycom.com>
To: Paul Coverdale <coverdale@sympatico.ca>
Date: Fri, 11 Apr 2014 14:42:04 -0700
Thread-Topic: [clue] Improving treatment of audio
Thread-Index: Ac9Vlt1dlqHYS7oTRTqUSFX4G036yQAJD77QAABTnFAABIolsA==
Message-ID: <49E45C59CA48264997FEBFB29B6BC2D62409219636@CRPMBOXPRD07.polycom.com>
References: <533AF351.9050201@alum.mit.edu> <20140410222705.GW39240@verdi> <BLU0-SMTP6910DA20C9A9287137B585D0550@phx.gbl> <53475572.4040807@nteczone.com> <20140411150055.GE60844@verdi> <49E45C59CA48264997FEBFB29B6BC2D6215295A8AE@CRPMBOXPRD07.polycom.com> <BLU0-SMTP1804A70E3E0F666EAEAE84D0540@phx.gbl>
In-Reply-To: <BLU0-SMTP1804A70E3E0F666EAEAE84D0540@phx.gbl>
Accept-Language: en-US
Content-Language: en-US
X-MS-Has-Attach: 
X-MS-TNEF-Correlator: 
acceptlanguage: en-US
Content-Type: text/plain; charset="us-ascii"
Content-Transfer-Encoding: quoted-printable
MIME-Version: 1.0
Archived-At: http://mailarchive.ietf.org/arch/msg/clue/HBgJ96mOqzSKAe7Gze9B3RdIVJo
Cc: "clue@ietf.org" <clue@ietf.org>
Subject: Re: [clue] Improving treatment of audio
X-BeenThere: clue@ietf.org
X-Mailman-Version: 2.1.15
Precedence: list
List-Id: CLUE - ControLling mUltiple streams for TElepresence <clue.ietf.org>
List-Unsubscribe: <https://www.ietf.org/mailman/options/clue>, <mailto:clue-request@ietf.org?subject=unsubscribe>
List-Archive: <http://www.ietf.org/mail-archive/web/clue/>
List-Post: <mailto:clue@ietf.org>
List-Help: <mailto:clue-request@ietf.org?subject=help>
List-Subscribe: <https://www.ietf.org/mailman/listinfo/clue>, <mailto:clue-request@ietf.org?subject=subscribe>
X-List-Received-Date: Fri, 11 Apr 2014 21:42:40 -0000

Hi Paul,
Yes, I understand the area of capture for audio can't be nearly as precise =
as it can be for video.  But for this usage I think it doesn't matter.
Mark

> -----Original Message-----
> From: Paul Coverdale [mailto:coverdale@sympatico.ca]
> Sent: Friday, April 11, 2014 4:48 PM
> To: Duckworth, Mark; 'John Leslie'; 'Christian Groves'
> Cc: clue@ietf.org
> Subject: RE: [clue] Improving treatment of audio
>=20
> Hi Mark,
>=20
> I can see what you're trying to do in the example given in Framework sect=
ion
> 12.1.1. The problem I have is that, because of the huge difference in
> wavelength between light waves and audio waves, you can't really define a=
n
> area of capture for audio in the same way that you can for video. A camer=
a
> can focus on a scene defined by 4 co-planar X,Y,Z coordinates. It will ca=
pture
> the video inside that quadrilateral, and nothing outside. A microphone ca=
n
> capture audio coming from inside of the same quadrilateral, but it will a=
lso
> capture a lot of audio from outside. What this means is that we can never
> define a unique association of an audio capture with a video capture base=
d
> on pure geometrical considerations, but we can still simply define that A=
Cx is
> associated with VCx, but maybe with VCy and VCz too.
>=20
> ...Paul
>=20
> >-----Original Message-----
> >From: clue [mailto:clue-bounces@ietf.org] On Behalf Of Duckworth, Mark
> >Sent: Friday, April 11, 2014 3:24 PM
> >To: John Leslie; Christian Groves
> >Cc: clue@ietf.org
> >Subject: Re: [clue] Improving treatment of audio
> >
> >Hi all,
> >The original intent of using "area of capture" for all media types,
> >including audio and video, was to provide a way to associate audio with
> >video.  So the consumer can choose captures that go together and render
> >them together.  An example of this is in the framework document,
> >section 12.1.1.  I still think this works fine for this purpose.
> >
> >In that example, the consumer can choose VC0, VC1, VC2, AC0, AC1, and
> >AC2.  AC0 and VC0 have the same area of capture, indicating they are
> >capturing the same area of the scene.  If the consumer wants to render
> >AC0 from a loudspeaker close to the display for VC0 it can do so.
> >Similarly, AC3 has an area of capture covering the whole scene (the
> >full extent of the areas of VC0, VC1, VC2) so the consumer knows AC3
> >includes audio associated with all of VC0, VC1, and VC2.
> >
> >So I'm puzzled why you are proposing we remove the ability to use area
> >of capture for audio for this purpose.
> >
> >Regards,
> >Mark
> >
> >> -----Original Message-----
> >> From: clue [mailto:clue-bounces@ietf.org] On Behalf Of John Leslie
> >> Sent: Friday, April 11, 2014 11:01 AM
> >> To: Christian Groves
> >> Cc: clue@ietf.org
> >> Subject: Re: [clue] Improving treatment of audio
> >>
> >> Christian Groves <Christian.Groves@nteczone.com> wrote:
> >> >
> >> > "So it's not clear to me what extra we need to do from a CLUE
> >> > perspective. What is the problem?"
> >> > I guess this is the pertinent point.
> >> >
> >> > From John's Audio 101 it seems that anything to do with spatial
> >> > information that would relate to reverberation is in the too hard
> >> > basket.
> >>
> >>    What precisely does "too hard basket" mean?
> >>
> >>    I haven't even reached the point of suggesting what metrics to
> >> define the _ability_ to send. To me, "too hard" merely means that an
> >> individual site could choose not to send them or to ignore them on
> >receipt.
> >>
> >>    But it sounds as if you're suggesting reverberation is "to hard to
> >> understand" and thus we should have no metrics about it.
> >>
> >>    I _hope_ that's not what you mean.
> >>
> >> > So it seems "area of capture" for audio could be marked "not
> >> > applicable" in the framework.
> >>
> >>    I hope so.
> >>
> >> > There doesn't seem to be any driver for having the "point of
> >capture"
> >> > apply to an audio capture either.
> >>
> >>    I don't understand this. Point of capture for a microphone may be
> >> "too hard" to track (today) for a microphone which moves; nonetheless
> >> it seems to me the most fundamental metric there could be.
> >>
> >> > From the Audio101 there does seem to be a dependency on where the
> >> > microphone is located with respect to the person speaking (i.e.
> >> > lapel mic, desk mic, room mic) to how it is handled at
> >mixing/playout.
> >>
> >>    I don't think I really got that far...
> >>
> >>    IMHO, it's more flexible to have mixing be the responsibily of the
> >> receiver, but I haven't tried to specify that and I'm not at all sure
> >I want to specify that.
> >>
> >>    I expect the actual sound systems in different rooms to vary
> >> wildly, from one monaural speaker to stereo to full surround-sound.
> >> Mixing for these without knowing which is the actual target seems
> >> hard; but I'm sure there will be sites which prefer to do so. A
> >> question which will arise, IMHO, is how to specify the _intent_ of a
> >> mix generated in one room to be fed to other rooms. (I'd prefer not
> >> to go there yet.)
> >>
> >> > Perhaps this is useful to signal via CLUE? If it is possible to
> >> > signal this then perhaps tying a particular audio capture to a
> >> > video capture makes sense?
> >>
> >>    I'm not thinking along those lines. (That doesn't mean we
> >> shouldn't think along those lines...) I'm thinking in terms of
> >> providing several audio streams per room, associated with position
> >> information about the position of the source of those sounds, and
> >> allowing the receiver to choose how to mix them and how to present the
> mix in his/her room.
> >>
> >> > i.e. a talker giving a presentation using a lapel mic captured by a
> >> > particular video.
> >>
> >>    In fact, there's only limited tendency for humans to strictly
> >> attach the sound they hear to the video they see. Clearly, during a
> >> presentation we want to _hear_ the presenter talking, but having the
> >> sound move back and forth as the speaker walks can be confusing.
> >>
> >>    At the same time, we will want to hear the questions to which the
> >> presenter may respond. It will be far easier on the listener if these
> >> do _not_ seem to be coming from the same physical position,
> >> especially when the questioner _is_ in the same room as the presenter.
> >>
> >>    Hope this helps...
> >>
> >> --
> >> John Leslie <john@jlc.net>
> >>
> >> _______________________________________________
> >> clue mailing list
> >> clue@ietf.org
> >> https://www.ietf.org/mailman/listinfo/clue
> >
> >_______________________________________________
> >clue mailing list
> >clue@ietf.org
> >https://www.ietf.org/mailman/listinfo/clue


From nobody Fri Apr 11 17:13:15 2014
Return-Path: <coverdale@sympatico.ca>
X-Original-To: clue@ietfa.amsl.com
Delivered-To: clue@ietfa.amsl.com
Received: from localhost (ietfa.amsl.com [127.0.0.1]) by ietfa.amsl.com (Postfix) with ESMTP id 4A2301A02DE for <clue@ietfa.amsl.com>; Fri, 11 Apr 2014 17:13:10 -0700 (PDT)
X-Virus-Scanned: amavisd-new at amsl.com
X-Spam-Flag: NO
X-Spam-Score: -1.9
X-Spam-Level: 
X-Spam-Status: No, score=-1.9 tagged_above=-999 required=5 tests=[BAYES_00=-1.9, MSGID_FROM_MTA_HEADER=0.001, RCVD_IN_DNSWL_NONE=-0.0001, SPF_PASS=-0.001] autolearn=ham
Received: from mail.ietf.org ([4.31.198.44]) by localhost (ietfa.amsl.com [127.0.0.1]) (amavisd-new, port 10024) with ESMTP id axngnCrSFuyw for <clue@ietfa.amsl.com>; Fri, 11 Apr 2014 17:13:07 -0700 (PDT)
Received: from blu0-omc1-s7.blu0.hotmail.com (blu0-omc1-s7.blu0.hotmail.com [65.55.116.18]) by ietfa.amsl.com (Postfix) with ESMTP id 94B391A02D9 for <clue@ietf.org>; Fri, 11 Apr 2014 17:13:07 -0700 (PDT)
Received: from BLU0-SMTP12 ([65.55.116.9]) by blu0-omc1-s7.blu0.hotmail.com with Microsoft SMTPSVC(6.0.3790.4675);  Fri, 11 Apr 2014 17:13:05 -0700
X-TMN: [5Q2NpFAcuJXkoi4XzMq8UPewO2QayiAA]
X-Originating-Email: [coverdale@sympatico.ca]
Message-ID: <BLU0-SMTP128D704AF84C6BFEE6CC53D0570@phx.gbl>
Received: from PaulNewPC ([74.15.60.251]) by BLU0-SMTP12.phx.gbl over TLS secured channel with Microsoft SMTPSVC(6.0.3790.4675);  Fri, 11 Apr 2014 17:13:05 -0700
From: Paul Coverdale <coverdale@sympatico.ca>
To: "'Duckworth, Mark'" <Mark.Duckworth@polycom.com>
References: <533AF351.9050201@alum.mit.edu> <20140410222705.GW39240@verdi> <BLU0-SMTP6910DA20C9A9287137B585D0550@phx.gbl> <53475572.4040807@nteczone.com> <20140411150055.GE60844@verdi> <49E45C59CA48264997FEBFB29B6BC2D6215295A8AE@CRPMBOXPRD07.polycom.com> <BLU0-SMTP1804A70E3E0F666EAEAE84D0540@phx.gbl> <49E45C59CA48264997FEBFB29B6BC2D62409219636@CRPMBOXPRD07.polycom.com>
In-Reply-To: <49E45C59CA48264997FEBFB29B6BC2D62409219636@CRPMBOXPRD07.polycom.com>
Date: Fri, 11 Apr 2014 20:12:54 -0400
MIME-Version: 1.0
Content-Type: text/plain; charset="us-ascii"
Content-Transfer-Encoding: 7bit
X-Mailer: Microsoft Office Outlook 12.0
Thread-Index: Ac9Vlt1dlqHYS7oTRTqUSFX4G036yQAJD77QAABTnFAABIolsAADvfOg
Content-Language: en-us
X-OriginalArrivalTime: 12 Apr 2014 00:13:05.0310 (UTC) FILETIME=[F9C04BE0:01CF55E3]
Archived-At: http://mailarchive.ietf.org/arch/msg/clue/CJptrB-m1xM3ysaXIjkksixRTAg
Cc: clue@ietf.org
Subject: Re: [clue] Improving treatment of audio
X-BeenThere: clue@ietf.org
X-Mailman-Version: 2.1.15
Precedence: list
List-Id: CLUE - ControLling mUltiple streams for TElepresence <clue.ietf.org>
List-Unsubscribe: <https://www.ietf.org/mailman/options/clue>, <mailto:clue-request@ietf.org?subject=unsubscribe>
List-Archive: <http://www.ietf.org/mail-archive/web/clue/>
List-Post: <mailto:clue@ietf.org>
List-Help: <mailto:clue-request@ietf.org?subject=help>
List-Subscribe: <https://www.ietf.org/mailman/listinfo/clue>, <mailto:clue-request@ietf.org?subject=subscribe>
X-List-Received-Date: Sat, 12 Apr 2014 00:13:10 -0000

Hi Mark,

Maybe I mis-understood, but I had the impression that the capture area would
provide the association between audio and video captures, without any other
explicit assignment. In other words, if capture areas for two different
media define exactly the same space, they must be associated with each
other. This is the part that I have some concerns about, for the reasons
already given for audio capture area. But I have no objection to defining an
association on the basis of the video capture area and an audio capture that
most closely intersects that area.


...Paul

>-----Original Message-----
>From: Duckworth, Mark [mailto:Mark.Duckworth@polycom.com]
>Sent: Friday, April 11, 2014 5:42 PM
>To: Paul Coverdale
>Cc: clue@ietf.org
>Subject: RE: [clue] Improving treatment of audio
>
>Hi Paul,
>Yes, I understand the area of capture for audio can't be nearly as
>precise as it can be for video.  But for this usage I think it doesn't
>matter.
>Mark
>
>> -----Original Message-----
>> From: Paul Coverdale [mailto:coverdale@sympatico.ca]
>> Sent: Friday, April 11, 2014 4:48 PM
>> To: Duckworth, Mark; 'John Leslie'; 'Christian Groves'
>> Cc: clue@ietf.org
>> Subject: RE: [clue] Improving treatment of audio
>>
>> Hi Mark,
>>
>> I can see what you're trying to do in the example given in Framework
>> section 12.1.1. The problem I have is that, because of the huge
>> difference in wavelength between light waves and audio waves, you
>> can't really define an area of capture for audio in the same way that
>> you can for video. A camera can focus on a scene defined by 4
>> co-planar X,Y,Z coordinates. It will capture the video inside that
>> quadrilateral, and nothing outside. A microphone can capture audio
>> coming from inside of the same quadrilateral, but it will also capture
>> a lot of audio from outside. What this means is that we can never
>> define a unique association of an audio capture with a video capture
>> based on pure geometrical considerations, but we can still simply
>define that ACx is associated with VCx, but maybe with VCy and VCz too.
>>
>> ...Paul
>>
>> >-----Original Message-----
>> >From: clue [mailto:clue-bounces@ietf.org] On Behalf Of Duckworth,
>> >Mark
>> >Sent: Friday, April 11, 2014 3:24 PM
>> >To: John Leslie; Christian Groves
>> >Cc: clue@ietf.org
>> >Subject: Re: [clue] Improving treatment of audio
>> >
>> >Hi all,
>> >The original intent of using "area of capture" for all media types,
>> >including audio and video, was to provide a way to associate audio
>> >with video.  So the consumer can choose captures that go together and
>> >render them together.  An example of this is in the framework
>> >document, section 12.1.1.  I still think this works fine for this
>purpose.
>> >
>> >In that example, the consumer can choose VC0, VC1, VC2, AC0, AC1, and
>> >AC2.  AC0 and VC0 have the same area of capture, indicating they are
>> >capturing the same area of the scene.  If the consumer wants to
>> >render
>> >AC0 from a loudspeaker close to the display for VC0 it can do so.
>> >Similarly, AC3 has an area of capture covering the whole scene (the
>> >full extent of the areas of VC0, VC1, VC2) so the consumer knows AC3
>> >includes audio associated with all of VC0, VC1, and VC2.
>> >
>> >So I'm puzzled why you are proposing we remove the ability to use
>> >area of capture for audio for this purpose.
>> >
>> >Regards,
>> >Mark
>> >
>> >> -----Original Message-----
>> >> From: clue [mailto:clue-bounces@ietf.org] On Behalf Of John Leslie
>> >> Sent: Friday, April 11, 2014 11:01 AM
>> >> To: Christian Groves
>> >> Cc: clue@ietf.org
>> >> Subject: Re: [clue] Improving treatment of audio
>> >>
>> >> Christian Groves <Christian.Groves@nteczone.com> wrote:
>> >> >
>> >> > "So it's not clear to me what extra we need to do from a CLUE
>> >> > perspective. What is the problem?"
>> >> > I guess this is the pertinent point.
>> >> >
>> >> > From John's Audio 101 it seems that anything to do with spatial
>> >> > information that would relate to reverberation is in the too hard
>> >> > basket.
>> >>
>> >>    What precisely does "too hard basket" mean?
>> >>
>> >>    I haven't even reached the point of suggesting what metrics to
>> >> define the _ability_ to send. To me, "too hard" merely means that
>> >> an individual site could choose not to send them or to ignore them
>> >> on
>> >receipt.
>> >>
>> >>    But it sounds as if you're suggesting reverberation is "to hard
>> >> to understand" and thus we should have no metrics about it.
>> >>
>> >>    I _hope_ that's not what you mean.
>> >>
>> >> > So it seems "area of capture" for audio could be marked "not
>> >> > applicable" in the framework.
>> >>
>> >>    I hope so.
>> >>
>> >> > There doesn't seem to be any driver for having the "point of
>> >capture"
>> >> > apply to an audio capture either.
>> >>
>> >>    I don't understand this. Point of capture for a microphone may
>> >> be "too hard" to track (today) for a microphone which moves;
>> >> nonetheless it seems to me the most fundamental metric there could
>be.
>> >>
>> >> > From the Audio101 there does seem to be a dependency on where the
>> >> > microphone is located with respect to the person speaking (i.e.
>> >> > lapel mic, desk mic, room mic) to how it is handled at
>> >mixing/playout.
>> >>
>> >>    I don't think I really got that far...
>> >>
>> >>    IMHO, it's more flexible to have mixing be the responsibily of
>> >> the receiver, but I haven't tried to specify that and I'm not at
>> >> all sure
>> >I want to specify that.
>> >>
>> >>    I expect the actual sound systems in different rooms to vary
>> >> wildly, from one monaural speaker to stereo to full surround-sound.
>> >> Mixing for these without knowing which is the actual target seems
>> >> hard; but I'm sure there will be sites which prefer to do so. A
>> >> question which will arise, IMHO, is how to specify the _intent_ of
>> >> a mix generated in one room to be fed to other rooms. (I'd prefer
>> >> not to go there yet.)
>> >>
>> >> > Perhaps this is useful to signal via CLUE? If it is possible to
>> >> > signal this then perhaps tying a particular audio capture to a
>> >> > video capture makes sense?
>> >>
>> >>    I'm not thinking along those lines. (That doesn't mean we
>> >> shouldn't think along those lines...) I'm thinking in terms of
>> >> providing several audio streams per room, associated with position
>> >> information about the position of the source of those sounds, and
>> >> allowing the receiver to choose how to mix them and how to present
>> >> the
>> mix in his/her room.
>> >>
>> >> > i.e. a talker giving a presentation using a lapel mic captured by
>> >> > a particular video.
>> >>
>> >>    In fact, there's only limited tendency for humans to strictly
>> >> attach the sound they hear to the video they see. Clearly, during a
>> >> presentation we want to _hear_ the presenter talking, but having
>> >> the sound move back and forth as the speaker walks can be
>confusing.
>> >>
>> >>    At the same time, we will want to hear the questions to which
>> >> the presenter may respond. It will be far easier on the listener if
>> >> these do _not_ seem to be coming from the same physical position,
>> >> especially when the questioner _is_ in the same room as the
>presenter.
>> >>
>> >>    Hope this helps...
>> >>
>> >> --
>> >> John Leslie <john@jlc.net>
>> >>
>> >> _______________________________________________
>> >> clue mailing list
>> >> clue@ietf.org
>> >> https://www.ietf.org/mailman/listinfo/clue
>> >
>> >_______________________________________________
>> >clue mailing list
>> >clue@ietf.org
>> >https://www.ietf.org/mailman/listinfo/clue



From nobody Sun Apr 13 12:59:17 2014
Return-Path: <coverdale@sympatico.ca>
X-Original-To: clue@ietfa.amsl.com
Delivered-To: clue@ietfa.amsl.com
Received: from localhost (ietfa.amsl.com [127.0.0.1]) by ietfa.amsl.com (Postfix) with ESMTP id 149201A0221 for <clue@ietfa.amsl.com>; Sun, 13 Apr 2014 12:59:16 -0700 (PDT)
X-Virus-Scanned: amavisd-new at amsl.com
X-Spam-Flag: NO
X-Spam-Score: 0.8
X-Spam-Level: 
X-Spam-Status: No, score=0.8 tagged_above=-999 required=5 tests=[BAYES_50=0.8,  MSGID_FROM_MTA_HEADER=0.001, RCVD_IN_DNSWL_NONE=-0.0001, SPF_PASS=-0.001] autolearn=ham
Received: from mail.ietf.org ([4.31.198.44]) by localhost (ietfa.amsl.com [127.0.0.1]) (amavisd-new, port 10024) with ESMTP id uNT379ZnUmnN for <clue@ietfa.amsl.com>; Sun, 13 Apr 2014 12:59:14 -0700 (PDT)
Received: from blu0-omc1-s29.blu0.hotmail.com (blu0-omc1-s29.blu0.hotmail.com [65.55.116.40]) by ietfa.amsl.com (Postfix) with ESMTP id 2E72D1A021D for <clue@ietf.org>; Sun, 13 Apr 2014 12:59:14 -0700 (PDT)
Received: from BLU0-SMTP94 ([65.55.116.8]) by blu0-omc1-s29.blu0.hotmail.com with Microsoft SMTPSVC(6.0.3790.4675);  Sun, 13 Apr 2014 12:59:11 -0700
X-TMN: [cIgQSk1N/82IUjvhnhkoVJHNV/Q8JXqY]
X-Originating-Email: [coverdale@sympatico.ca]
Message-ID: <BLU0-SMTP942A03E86972E7B8A212C6D0560@phx.gbl>
Received: from PaulNewPC ([74.15.60.251]) by BLU0-SMTP94.phx.gbl over TLS secured channel with Microsoft SMTPSVC(6.0.3790.4675);  Sun, 13 Apr 2014 12:59:11 -0700
From: Paul Coverdale <coverdale@sympatico.ca>
To: "'CLUE'" <clue@ietf.org>
References: <533AF351.9050201@alum.mit.edu>
In-Reply-To: <533AF351.9050201@alum.mit.edu>
Date: Sun, 13 Apr 2014 15:59:10 -0400
MIME-Version: 1.0
Content-Type: text/plain; charset="us-ascii"
Content-Transfer-Encoding: 7bit
X-Mailer: Microsoft Office Outlook 12.0
Thread-Index: Ac9NzXhcFrxH9DijQqmnx4bwUJgguAJbRPTw
Content-Language: en-us
X-OriginalArrivalTime: 13 Apr 2014 19:59:11.0478 (UTC) FILETIME=[D6818160:01CF5752]
Archived-At: http://mailarchive.ietf.org/arch/msg/clue/9IUKOjreD6LTqXW8OoFBkNocPzw
Subject: Re: [clue] Improving treatment of audio
X-BeenThere: clue@ietf.org
X-Mailman-Version: 2.1.15
Precedence: list
List-Id: CLUE - ControLling mUltiple streams for TElepresence <clue.ietf.org>
List-Unsubscribe: <https://www.ietf.org/mailman/options/clue>, <mailto:clue-request@ietf.org?subject=unsubscribe>
List-Archive: <http://www.ietf.org/mail-archive/web/clue/>
List-Post: <mailto:clue@ietf.org>
List-Help: <mailto:clue-request@ietf.org?subject=help>
List-Subscribe: <https://www.ietf.org/mailman/listinfo/clue>, <mailto:clue-request@ietf.org?subject=subscribe>
X-List-Received-Date: Sun, 13 Apr 2014 19:59:16 -0000

So, going back to Paul K's request for a concise statement of what needs to
be done with respect to the treatment of audio in CLUE, I'm not sure what
else is really required apart from satisfying REQMT-2 and REQMT-3, with the
possible exception of including the directional characteristics of the
microphone with the audio capture.

...Paul

>-----Original Message-----
>From: clue [mailto:clue-bounces@ietf.org] On Behalf Of Paul Kyzivat
>Sent: Tuesday, April 01, 2014 1:12 PM
>To: CLUE
>Subject: [clue] Improving treatment of audio
>
>In today's design team meeting we had a discussion of what is lacking
>about our treatment of audio, and how to fix it.
>
>I want to open a ticket on this topic, but I need some help to properly
>describe the task. IMO it has to do with what sort of spatial
>information should be provided for audio captures, how it can be used to
>correlate audio captures with video captures, and how it can be used to
>choose which audio captures to configure.
>
>Can somebody (John?) make a *concise* statement of what is needed?
>
>	Thanks,
>	Paul
>
>_______________________________________________
>clue mailing list
>clue@ietf.org
>https://www.ietf.org/mailman/listinfo/clue


From nobody Sun Apr 13 16:32:04 2014
Return-Path: <john@jlc.net>
X-Original-To: clue@ietfa.amsl.com
Delivered-To: clue@ietfa.amsl.com
Received: from localhost (ietfa.amsl.com [127.0.0.1]) by ietfa.amsl.com (Postfix) with ESMTP id 42A211A0253 for <clue@ietfa.amsl.com>; Sun, 13 Apr 2014 16:32:02 -0700 (PDT)
X-Virus-Scanned: amavisd-new at amsl.com
X-Spam-Flag: NO
X-Spam-Score: -1.772
X-Spam-Level: 
X-Spam-Status: No, score=-1.772 tagged_above=-999 required=5 tests=[BAYES_50=0.8, RCVD_IN_DNSWL_MED=-2.3, RP_MATCHES_RCVD=-0.272] autolearn=ham
Received: from mail.ietf.org ([4.31.198.44]) by localhost (ietfa.amsl.com [127.0.0.1]) (amavisd-new, port 10024) with ESMTP id TKYwBKocEj5K for <clue@ietfa.amsl.com>; Sun, 13 Apr 2014 16:31:59 -0700 (PDT)
Received: from mailhost.jlc.net (mailhost.jlc.net [199.201.159.4]) by ietfa.amsl.com (Postfix) with ESMTP id 78AB31A02F3 for <clue@ietf.org>; Sun, 13 Apr 2014 16:31:59 -0700 (PDT)
Received: by mailhost.jlc.net (Postfix, from userid 104) id 2C258C94C1; Sun, 13 Apr 2014 19:31:54 -0400 (EDT)
Date: Sun, 13 Apr 2014 19:31:54 -0400
From: John Leslie <john@jlc.net>
To: Paul Coverdale <coverdale@sympatico.ca>
Message-ID: <20140413233154.GL60844@verdi>
References: <533AF351.9050201@alum.mit.edu> <BLU0-SMTP942A03E86972E7B8A212C6D0560@phx.gbl>
Mime-Version: 1.0
Content-Type: text/plain; charset=us-ascii
Content-Disposition: inline
In-Reply-To: <BLU0-SMTP942A03E86972E7B8A212C6D0560@phx.gbl>
User-Agent: Mutt/1.4.1i
Archived-At: http://mailarchive.ietf.org/arch/msg/clue/fep_E8tss-CFoWZrcIXoyraQIFI
Cc: 'CLUE' <clue@ietf.org>
Subject: Re: [clue] Improving treatment of audio
X-BeenThere: clue@ietf.org
X-Mailman-Version: 2.1.15
Precedence: list
List-Id: CLUE - ControLling mUltiple streams for TElepresence <clue.ietf.org>
List-Unsubscribe: <https://www.ietf.org/mailman/options/clue>, <mailto:clue-request@ietf.org?subject=unsubscribe>
List-Archive: <http://www.ietf.org/mail-archive/web/clue/>
List-Post: <mailto:clue@ietf.org>
List-Help: <mailto:clue-request@ietf.org?subject=help>
List-Subscribe: <https://www.ietf.org/mailman/listinfo/clue>, <mailto:clue-request@ietf.org?subject=subscribe>
X-List-Received-Date: Sun, 13 Apr 2014 23:32:02 -0000

Paul Coverdale <coverdale@sympatico.ca> wrote:
> 
> So, going back to Paul K's request for a concise statement of what needs to
> be done with respect to the treatment of audio in CLUE, I'm not sure what
> else is really required apart from satisfying REQMT-2 and REQMT-3,

   I have to agree with Paul-C here: clue-telepresence-requirements is
pretty much set in stone, now that it has entered the RFCed queue. But
I don't think it fits Paul-K's criteria of "concise:
] 
] REQMT-2: The solution MUST support a description of the spatial
]          arrangement of captured source audio sent in audio streams
]          which enables a satisfactory reproduction at the receiver
]          in a spatially correct manner.  This applies to each site
]          in a point to point or a multipoint meeting and refers to
]          the spatial ordering within a site, not the ordering of
]          channels between sites.
]
]          Use case point to point symmetric, and all use cases,
]          especially heterogeneous.
]
]  REQMT-2a: The solution MUST support a means of preserving
]            the spatial order of audio in the captured
]            scene.  For example, if John sounds as if he is
]            at Susan's right in the captured audio, John
]            voice is also placed at Susan's right in the
]            rendered image.
]
]  REQMT-2b: The solution MUST support a means to identify
]            the number and spatial arrangement of audio
]            channels including monaural, stereophonic
]            (2.0), and 3.0 (left, center, right) audio
]            channels.
]
]  REQMT-2c: The solution MUST support a means to identify
]            the point of capture of individual audio
]            captures in three dimensions.
]
]  REQMT-2d: The solution MUST support a means to identify
]            the area of coverage of individual audio
]            captures in three dimensions.
]
] REQMT-3: The solution MUST enable individual audio streams to be
]          associated with one or more video image captures, and
]          individual video image captures to be associated with one
]          or more audio captures, for the purpose of rendering
]          proper position.
]
]          Use case is point to point symmetric, and all use cases.

   Worse, I'm not confident that all of this is even attainable. :^(

   Mea culpa, of course -- I should have said that in WGLC. :^( :^(

> with the possible exception of including the directional
> characteristics of the microphone with the audio capture.

   While I'd be happy to see that signaled, I'm really not pushing it
for our first spec -- merely hoping we'll get there someday.

   There is, alas, an elephant in each room which can render our
audio treatment useless -- it's called "echo". :^(

====
   Despite being too late, I feel I ought to comment on these REQMTs.
Feel free to ignore what follows:

] REQMT-2: The solution MUST support a description of the spatial
]          arrangement of captured source audio sent in audio streams
]          which enables a satisfactory reproduction at the receiver
]          in a spatially correct manner.

   Other than wondering what "spatially correct means, I view this as
attainable, since the responsibily must belong to the receiver, and
that only if the sender chooses to specify enough audio sources.

]  REQMT-2a: The solution MUST support a means of preserving
]            the spatial order of audio in the captured
]            scene.  For example, if John sounds as if he is
]            at Susan's right in the captured audio, John
]            voice is also placed at Susan's right in the
]            rendered image.

   Likewise, this responsibility must belong to the receiver.

]  REQMT-2b: The solution MUST support a means to identify
]            the number and spatial arrangement of audio
]            channels including monaural, stereophonic
]            (2.0), and 3.0 (left, center, right) audio
]            channels.

   This, IMHO, isn't worth doing (although perhaps I don't understand
what REQMT-2b means).

   The actual layout of loudspeakers should belong entirely to the
receiving room. Further, "stereo" doesn't actually define the relationship
of the two channels; and "3.0" suffers a similar problem.

   If a sender _actually_ sends "stereo" or "3.0" the receiver will have
to punt. I can only hope that our actual standard doesn't encourage this.

]  REQMT-2c: The solution MUST support a means to identify
]            the point of capture of individual audio
]            captures in three dimensions.

   Well stated. Adding point on line of capture helps slightly, but only
if we have a useful approximation of sensitivity pattern (which I don't
believe we can expect to get).

]  REQMT-2d: The solution MUST support a means to identify
]            the area of coverage of individual audio
]            captures in three dimensions.

   Poorly stated. :^( _Every_ microphone in a room can cover _every_
source of sound in a room. Thus, the "area of coverage" must necessarily
be the entire room...

   (It gets worse... much worse!)

   Inevitably, some yahoo is going to put one or more loudspeakers in
the room. And those will expand the "area of coverage" to include the
other rooms those loudspeakers attempt to render.

   With delay!!!

] REQMT-3: The solution MUST enable individual audio streams to be
]          associated with one or more video image captures, and
]          individual video image captures to be associated with one
]          or more audio captures, for the purpose of rendering
]          proper position.
]
]          Use case is point to point symmetric, and all use cases.

   I accept there are many folks who will insist on doing this. I don't
expect to stop them.

   But audio streams are intrinsically linked to rooms, not video captures.
And the most useful audio streams are linked to microphones "close" to
individual people, while video captures will _very_ often try to cover
more than one person -- thus video capures will very typically be associated
with several audio captures.

====

   It will help if more of us pay attention to the relationship of audio
in television in movies. I find that they _don't_ track each other too
closely.

   The audio provides the continuity; the video tantalizes the senses.
When the camera switches between two people talking to each other, the
audio doesn't switch -- it tries to make you believe you're in the room.

   Thus I most sincerely hope that folks won't choose to switch their
audio based on who's speaking -- but I won't try to stop them from doing
this. (And I must admit that sometimes I _shouldn't_ blame them, because
they find themselves sabotaged by audio sources contaminated by echo.)

====

   A word on "echo suppression" might help here: this is a well-solved
in plain-old-telephony where end-to-end delay can be held constand. We
won't be able to do that over the Internet.

   An individual room is probably able to cancel the echo in what it
_sends_ but a receiving room won't be able to cancel echo in what it
receives. :^(

====

   That's as much as I'm willing to sqeeze into this email.

   Hope _something_ in this helps...

--
John Leslie <john@jlc.net>


From nobody Sun Apr 13 18:49:39 2014
Return-Path: <Christian.Groves@nteczone.com>
X-Original-To: clue@ietfa.amsl.com
Delivered-To: clue@ietfa.amsl.com
Received: from localhost (ietfa.amsl.com [127.0.0.1]) by ietfa.amsl.com (Postfix) with ESMTP id EE8CC1A02E5 for <clue@ietfa.amsl.com>; Sun, 13 Apr 2014 18:49:38 -0700 (PDT)
X-Virus-Scanned: amavisd-new at amsl.com
X-Spam-Flag: NO
X-Spam-Score: 0.8
X-Spam-Level: 
X-Spam-Status: No, score=0.8 tagged_above=-999 required=5 tests=[BAYES_50=0.8] autolearn=ham
Received: from mail.ietf.org ([4.31.198.44]) by localhost (ietfa.amsl.com [127.0.0.1]) (amavisd-new, port 10024) with ESMTP id YBSsLh_BEkSk for <clue@ietfa.amsl.com>; Sun, 13 Apr 2014 18:49:37 -0700 (PDT)
Received: from cserver5.myshophosting.com (cserver5.myshophosting.com [175.107.161.1]) by ietfa.amsl.com (Postfix) with ESMTP id 2C3BD1A02E3 for <clue@ietf.org>; Sun, 13 Apr 2014 18:49:37 -0700 (PDT)
Received: from ppp118-209-221-158.lns20.mel6.internode.on.net ([118.209.221.158]:53099 helo=[127.0.0.1]) by cserver5.myshophosting.com with esmtpsa (TLSv1:DHE-RSA-AES128-SHA:128) (Exim 4.82) (envelope-from <Christian.Groves@nteczone.com>) id 1WZW1W-0000x1-TN; Mon, 14 Apr 2014 11:49:27 +1000
Message-ID: <534B3EA8.4060406@nteczone.com>
Date: Mon, 14 Apr 2014 11:49:28 +1000
From: Christian Groves <Christian.Groves@nteczone.com>
User-Agent: Mozilla/5.0 (Windows NT 6.1; WOW64; rv:24.0) Gecko/20100101 Thunderbird/24.4.0
MIME-Version: 1.0
To: John Leslie <john@jlc.net>
References: <533AF351.9050201@alum.mit.edu> <20140410222705.GW39240@verdi> <BLU0-SMTP6910DA20C9A9287137B585D0550@phx.gbl> <53475572.4040807@nteczone.com> <20140411150055.GE60844@verdi>
In-Reply-To: <20140411150055.GE60844@verdi>
Content-Type: text/plain; charset=ISO-8859-1; format=flowed
Content-Transfer-Encoding: 7bit
X-AntiAbuse: This header was added to track abuse, please include it with any abuse report
X-AntiAbuse: Primary Hostname - cserver5.myshophosting.com
X-AntiAbuse: Original Domain - ietf.org
X-AntiAbuse: Originator/Caller UID/GID - [47 12] / [47 12]
X-AntiAbuse: Sender Address Domain - nteczone.com
X-Get-Message-Sender-Via: cserver5.myshophosting.com: authenticated_id: christian.groves@nteczone.com
X-Source: 
X-Source-Args: 
X-Source-Dir: 
Archived-At: http://mailarchive.ietf.org/arch/msg/clue/ZnuElhYaJClsNvN43PI5Mzws5-Y
Cc: clue@ietf.org
Subject: Re: [clue] Improving treatment of audio
X-BeenThere: clue@ietf.org
X-Mailman-Version: 2.1.15
Precedence: list
List-Id: CLUE - ControLling mUltiple streams for TElepresence <clue.ietf.org>
List-Unsubscribe: <https://www.ietf.org/mailman/options/clue>, <mailto:clue-request@ietf.org?subject=unsubscribe>
List-Archive: <http://www.ietf.org/mail-archive/web/clue/>
List-Post: <mailto:clue@ietf.org>
List-Help: <mailto:clue-request@ietf.org?subject=help>
List-Subscribe: <https://www.ietf.org/mailman/listinfo/clue>, <mailto:clue-request@ietf.org?subject=subscribe>
X-List-Received-Date: Mon, 14 Apr 2014 01:49:39 -0000

Hello John,

Thanks for the reply. Please see my responses below.

Regards, Christian

On 12/04/2014 1:00 AM, John Leslie wrote:
> Christian Groves <Christian.Groves@nteczone.com> wrote:
>> "So it's not clear to me what extra we need to do from a CLUE
>> perspective. What is the problem?"
>> I guess this is the pertinent point.
>>
>>  From John's Audio 101 it seems that anything to do with spatial
>> information that would relate to reverberation is in the too hard
>> basket.
>     What precisely does "too hard basket" mean?
[CNG] Well what did you mean by "That's too hard right now" in your 
audio 101? :-) The tone of the 101 (and previous emails) seemed to imply 
that using the current spatial information is not appropriate and using 
other sorts of physical information is going to be hard.
>
>     I haven't even reached the point of suggesting what metrics to
> define the _ability_ to send. To me, "too hard" merely means that
> an individual site could choose not to send them or to ignore them
> on receipt.
[CNG] I didn't understand your usage of "too hard" to imply optionality 
in sending or receipt.
>
>     But it sounds as if you're suggesting reverberation is "to hard
> to understand" and thus we should have no metrics about it.
[CNG] I'm not saying its too hard to understand, more questioning about 
the amount of effort that would be required at this point in time of the 
development of CLUE.
>
>     I _hope_ that's not what you mean.
>
>> So it seems "area of capture" for audio could be marked "not
>> applicable" in the framework.
>     I hope so.
>
>> There doesn't seem to be any driver for having the "point of capture"
>> apply to an audio capture either.
>     I don't understand this. Point of capture for a microphone may be
> "too hard" to track (today) for a microphone which moves; nonetheless
> it seems to me the most fundamental metric there could be.
[CNG] I didn't understand the 101 wanting the current "point of capture" 
metric as a physical point. I agree it is a fundamental metric. However 
rather than a physical point in space is it enough to say the microphone 
is associated with a certain video capture or a location (e.g. lapel 
mic.)? I guess this comes down to what people actually want to do with it.
>
>>  From the Audio101 there does seem to be a dependency on where the
>> microphone is located with respect to the person speaking (i.e. lapel
>> mic, desk mic, room mic) to how it is handled at mixing/playout.
>     I don't think I really got that far...
[CNG] You did mention different possibilities of handling sound from 
different microphone types and where they are placed. That's what I was 
picking up on.
>
>     IMHO, it's more flexible to have mixing be the responsibily of
> the receiver, but I haven't tried to specify that and I'm not at all
> sure I want to specify that.
>
>     I expect the actual sound systems in different rooms to vary wildly,
> from one monaural speaker to stereo to full surround-sound. Mixing
> for these without knowing which is the actual target seems hard; but
> I'm sure there will be sites which prefer to do so. A question which
> will arise, IMHO, is how to specify the _intent_ of a mix generated
> in one room to be fed to other rooms. (I'd prefer not to go there yet.)
>
>> Perhaps this is useful to signal via CLUE? If it is possible to signal
>> this then perhaps tying a particular audio capture to a video capture
>> makes sense?
>     I'm not thinking along those lines. (That doesn't mean we shouldn't
> think along those lines...) I'm thinking in terms of providing several
> audio streams per room, associated with position information about
> the position of the source of those sounds, and allowing the receiver
> to choose how to mix them and how to present the mix in his/her room.
[CNG] I think we're thinking in a similar vain, the issue is how to 
specify the "position". Whether its a physical point or more 
representional, i.e. lapel mic.
>
>> i.e. a talker giving a presentation using a lapel mic captured by a
>> particular video.
>     In fact, there's only limited tendency for humans to strictly
> attach the sound they hear to the video they see. Clearly, during a
> presentation we want to _hear_ the presenter talking, but having the
> sound move back and forth as the speaker walks can be confusing.
>
>     At the same time, we will want to hear the questions to which the
> presenter may respond. It will be far easier on the listener if these
> do _not_ seem to be coming from the same physical position, especially
> when the questioner _is_ in the same room as the presenter.
>
>     Hope this helps...
>
> --
> John Leslie <john@jlc.net>
>


From nobody Sun Apr 13 19:08:18 2014
Return-Path: <Christian.Groves@nteczone.com>
X-Original-To: clue@ietfa.amsl.com
Delivered-To: clue@ietfa.amsl.com
Received: from localhost (ietfa.amsl.com [127.0.0.1]) by ietfa.amsl.com (Postfix) with ESMTP id D491A1A02F9 for <clue@ietfa.amsl.com>; Sun, 13 Apr 2014 19:08:13 -0700 (PDT)
X-Virus-Scanned: amavisd-new at amsl.com
X-Spam-Flag: NO
X-Spam-Score: -0.5
X-Spam-Level: 
X-Spam-Status: No, score=-0.5 tagged_above=-999 required=5 tests=[BAYES_05=-0.5] autolearn=ham
Received: from mail.ietf.org ([4.31.198.44]) by localhost (ietfa.amsl.com [127.0.0.1]) (amavisd-new, port 10024) with ESMTP id Bj0sEvO0xBd2 for <clue@ietfa.amsl.com>; Sun, 13 Apr 2014 19:08:10 -0700 (PDT)
Received: from cserver5.myshophosting.com (cserver5.myshophosting.com [175.107.161.1]) by ietfa.amsl.com (Postfix) with ESMTP id 726AC1A02EC for <clue@ietf.org>; Sun, 13 Apr 2014 19:08:10 -0700 (PDT)
Received: from ppp118-209-221-158.lns20.mel6.internode.on.net ([118.209.221.158]:53184 helo=[127.0.0.1]) by cserver5.myshophosting.com with esmtpsa (TLSv1:DHE-RSA-AES128-SHA:128) (Exim 4.82) (envelope-from <Christian.Groves@nteczone.com>) id 1WZWJW-0003md-Op for clue@ietf.org; Mon, 14 Apr 2014 12:08:02 +1000
Message-ID: <534B4303.7060707@nteczone.com>
Date: Mon, 14 Apr 2014 12:08:03 +1000
From: Christian Groves <Christian.Groves@nteczone.com>
User-Agent: Mozilla/5.0 (Windows NT 6.1; WOW64; rv:24.0) Gecko/20100101 Thunderbird/24.4.0
MIME-Version: 1.0
To: clue@ietf.org
References: <533AF351.9050201@alum.mit.edu> <20140410222705.GW39240@verdi> <BLU0-SMTP6910DA20C9A9287137B585D0550@phx.gbl> <53475572.4040807@nteczone.com> <20140411150055.GE60844@verdi> <49E45C59CA48264997FEBFB29B6BC2D6215295A8AE@CRPMBOXPRD07.polycom.com> <BLU0-SMTP1804A70E3E0F666EAEAE84D0540@phx.gbl> <49E45C59CA48264997FEBFB29B6BC2D62409219636@CRPMBOXPRD07.polycom.com>
In-Reply-To: <49E45C59CA48264997FEBFB29B6BC2D62409219636@CRPMBOXPRD07.polycom.com>
Content-Type: text/plain; charset=ISO-8859-1; format=flowed
Content-Transfer-Encoding: 7bit
X-AntiAbuse: This header was added to track abuse, please include it with any abuse report
X-AntiAbuse: Primary Hostname - cserver5.myshophosting.com
X-AntiAbuse: Original Domain - ietf.org
X-AntiAbuse: Originator/Caller UID/GID - [47 12] / [47 12]
X-AntiAbuse: Sender Address Domain - nteczone.com
X-Get-Message-Sender-Via: cserver5.myshophosting.com: authenticated_id: christian.groves@nteczone.com
X-Source: 
X-Source-Args: 
X-Source-Dir: 
Archived-At: http://mailarchive.ietf.org/arch/msg/clue/zI2sRHsLBz5nexVLlhOO29RYGUE
Subject: Re: [clue] Improving treatment of audio
X-BeenThere: clue@ietf.org
X-Mailman-Version: 2.1.15
Precedence: list
List-Id: CLUE - ControLling mUltiple streams for TElepresence <clue.ietf.org>
List-Unsubscribe: <https://www.ietf.org/mailman/options/clue>, <mailto:clue-request@ietf.org?subject=unsubscribe>
List-Archive: <http://www.ietf.org/mail-archive/web/clue/>
List-Post: <mailto:clue@ietf.org>
List-Help: <mailto:clue-request@ietf.org?subject=help>
List-Subscribe: <https://www.ietf.org/mailman/listinfo/clue>, <mailto:clue-request@ietf.org?subject=subscribe>
X-List-Received-Date: Mon, 14 Apr 2014 02:08:14 -0000

Hello all,

Perhaps the confusion is that some see Audio Capture area as a means to 
associate an Audio capture with a video capture as you've shown Mark's 
example. e.g. The audio spatial information isn't really used for any 
audio transformation other than associating it with a particular video 
stream.

Whereas others were more thinking the audio spatial information as an 
input to mixing and more complicated audio processing.

If this is the case perhaps rather than linking ACs and VCs through the 
physical or virtual co-ordinates of the area of capture information, we 
simplify things and in each audio capture we say what VC it relates to? 
e.g.using Mark's example AC0(VC0), AC3(V0,V1,V2)


Regards, Christian

On 12/04/2014 7:42 AM, Duckworth, Mark wrote:
> Hi Paul,
> Yes, I understand the area of capture for audio can't be nearly as precise as it can be for video.  But for this usage I think it doesn't matter.
> Mark
>
>> -----Original Message-----
>> From: Paul Coverdale [mailto:coverdale@sympatico.ca]
>> Sent: Friday, April 11, 2014 4:48 PM
>> To: Duckworth, Mark; 'John Leslie'; 'Christian Groves'
>> Cc: clue@ietf.org
>> Subject: RE: [clue] Improving treatment of audio
>>
>> Hi Mark,
>>
>> I can see what you're trying to do in the example given in Framework section
>> 12.1.1. The problem I have is that, because of the huge difference in
>> wavelength between light waves and audio waves, you can't really define an
>> area of capture for audio in the same way that you can for video. A camera
>> can focus on a scene defined by 4 co-planar X,Y,Z coordinates. It will capture
>> the video inside that quadrilateral, and nothing outside. A microphone can
>> capture audio coming from inside of the same quadrilateral, but it will also
>> capture a lot of audio from outside. What this means is that we can never
>> define a unique association of an audio capture with a video capture based
>> on pure geometrical considerations, but we can still simply define that ACx is
>> associated with VCx, but maybe with VCy and VCz too.
>>
>> ...Paul
>>
>>> -----Original Message-----
>>> From: clue [mailto:clue-bounces@ietf.org] On Behalf Of Duckworth, Mark
>>> Sent: Friday, April 11, 2014 3:24 PM
>>> To: John Leslie; Christian Groves
>>> Cc: clue@ietf.org
>>> Subject: Re: [clue] Improving treatment of audio
>>>
>>> Hi all,
>>> The original intent of using "area of capture" for all media types,
>>> including audio and video, was to provide a way to associate audio with
>>> video.  So the consumer can choose captures that go together and render
>>> them together.  An example of this is in the framework document,
>>> section 12.1.1.  I still think this works fine for this purpose.
>>>
>>> In that example, the consumer can choose VC0, VC1, VC2, AC0, AC1, and
>>> AC2.  AC0 and VC0 have the same area of capture, indicating they are
>>> capturing the same area of the scene.  If the consumer wants to render
>>> AC0 from a loudspeaker close to the display for VC0 it can do so.
>>> Similarly, AC3 has an area of capture covering the whole scene (the
>>> full extent of the areas of VC0, VC1, VC2) so the consumer knows AC3
>>> includes audio associated with all of VC0, VC1, and VC2.
>>>
>>> So I'm puzzled why you are proposing we remove the ability to use area
>>> of capture for audio for this purpose.
>>>
>>> Regards,
>>> Mark
>>>
>>>> -----Original Message-----
>>>> From: clue [mailto:clue-bounces@ietf.org] On Behalf Of John Leslie
>>>> Sent: Friday, April 11, 2014 11:01 AM
>>>> To: Christian Groves
>>>> Cc: clue@ietf.org
>>>> Subject: Re: [clue] Improving treatment of audio
>>>>
>>>> Christian Groves <Christian.Groves@nteczone.com> wrote:
>>>>> "So it's not clear to me what extra we need to do from a CLUE
>>>>> perspective. What is the problem?"
>>>>> I guess this is the pertinent point.
>>>>>
>>>>>  From John's Audio 101 it seems that anything to do with spatial
>>>>> information that would relate to reverberation is in the too hard
>>>>> basket.
>>>>     What precisely does "too hard basket" mean?
>>>>
>>>>     I haven't even reached the point of suggesting what metrics to
>>>> define the _ability_ to send. To me, "too hard" merely means that an
>>>> individual site could choose not to send them or to ignore them on
>>> receipt.
>>>>     But it sounds as if you're suggesting reverberation is "to hard to
>>>> understand" and thus we should have no metrics about it.
>>>>
>>>>     I _hope_ that's not what you mean.
>>>>
>>>>> So it seems "area of capture" for audio could be marked "not
>>>>> applicable" in the framework.
>>>>     I hope so.
>>>>
>>>>> There doesn't seem to be any driver for having the "point of
>>> capture"
>>>>> apply to an audio capture either.
>>>>     I don't understand this. Point of capture for a microphone may be
>>>> "too hard" to track (today) for a microphone which moves; nonetheless
>>>> it seems to me the most fundamental metric there could be.
>>>>
>>>>>  From the Audio101 there does seem to be a dependency on where the
>>>>> microphone is located with respect to the person speaking (i.e.
>>>>> lapel mic, desk mic, room mic) to how it is handled at
>>> mixing/playout.
>>>>     I don't think I really got that far...
>>>>
>>>>     IMHO, it's more flexible to have mixing be the responsibily of the
>>>> receiver, but I haven't tried to specify that and I'm not at all sure
>>> I want to specify that.
>>>>     I expect the actual sound systems in different rooms to vary
>>>> wildly, from one monaural speaker to stereo to full surround-sound.
>>>> Mixing for these without knowing which is the actual target seems
>>>> hard; but I'm sure there will be sites which prefer to do so. A
>>>> question which will arise, IMHO, is how to specify the _intent_ of a
>>>> mix generated in one room to be fed to other rooms. (I'd prefer not
>>>> to go there yet.)
>>>>
>>>>> Perhaps this is useful to signal via CLUE? If it is possible to
>>>>> signal this then perhaps tying a particular audio capture to a
>>>>> video capture makes sense?
>>>>     I'm not thinking along those lines. (That doesn't mean we
>>>> shouldn't think along those lines...) I'm thinking in terms of
>>>> providing several audio streams per room, associated with position
>>>> information about the position of the source of those sounds, and
>>>> allowing the receiver to choose how to mix them and how to present the
>> mix in his/her room.
>>>>> i.e. a talker giving a presentation using a lapel mic captured by a
>>>>> particular video.
>>>>     In fact, there's only limited tendency for humans to strictly
>>>> attach the sound they hear to the video they see. Clearly, during a
>>>> presentation we want to _hear_ the presenter talking, but having the
>>>> sound move back and forth as the speaker walks can be confusing.
>>>>
>>>>     At the same time, we will want to hear the questions to which the
>>>> presenter may respond. It will be far easier on the listener if these
>>>> do _not_ seem to be coming from the same physical position,
>>>> especially when the questioner _is_ in the same room as the presenter.
>>>>
>>>>     Hope this helps...
>>>>
>>>> --
>>>> John Leslie <john@jlc.net>
>>>>
>>>> _______________________________________________
>>>> clue mailing list
>>>> clue@ietf.org
>>>> https://www.ietf.org/mailman/listinfo/clue
>>> _______________________________________________
>>> clue mailing list
>>> clue@ietf.org
>>> https://www.ietf.org/mailman/listinfo/clue
> _______________________________________________
> clue mailing list
> clue@ietf.org
> https://www.ietf.org/mailman/listinfo/clue
>


From nobody Sun Apr 13 19:26:51 2014
Return-Path: <Mark.Duckworth@polycom.com>
X-Original-To: clue@ietfa.amsl.com
Delivered-To: clue@ietfa.amsl.com
Received: from localhost (ietfa.amsl.com [127.0.0.1]) by ietfa.amsl.com (Postfix) with ESMTP id 5712C1A0309 for <clue@ietfa.amsl.com>; Sun, 13 Apr 2014 19:26:48 -0700 (PDT)
X-Virus-Scanned: amavisd-new at amsl.com
X-Spam-Flag: NO
X-Spam-Score: -1.82
X-Spam-Level: 
X-Spam-Status: No, score=-1.82 tagged_above=-999 required=5 tests=[BAYES_00=-1.9, RCVD_IN_DNSWL_LOW=-0.7, SPF_NEUTRAL=0.779, UNPARSEABLE_RELAY=0.001] autolearn=ham
Received: from mail.ietf.org ([4.31.198.44]) by localhost (ietfa.amsl.com [127.0.0.1]) (amavisd-new, port 10024) with ESMTP id eA4gdCiSNO64 for <clue@ietfa.amsl.com>; Sun, 13 Apr 2014 19:26:45 -0700 (PDT)
Received: from mail1.bemta7.messagelabs.com (mail1.bemta7.messagelabs.com [216.82.254.100]) by ietfa.amsl.com (Postfix) with ESMTP id AC4371A02F9 for <clue@ietf.org>; Sun, 13 Apr 2014 19:26:45 -0700 (PDT)
Received: from [216.82.254.20:14096] by server-4.bemta-7.messagelabs.com id 82/9C-05677-3674B435; Mon, 14 Apr 2014 02:26:43 +0000
X-Env-Sender: Mark.Duckworth@polycom.com
X-Msg-Ref: server-16.tower-47.messagelabs.com!1397442402!7196776!1
X-Originating-IP: [140.242.64.158]
X-StarScan-Received: 
X-StarScan-Version: 6.11.1; banners=-,-,-
X-VirusChecked: Checked
Received: (qmail 14855 invoked from network); 14 Apr 2014 02:26:42 -0000
Received: from crpehubprd01.polycom.com (HELO crpehubprd02.polycom.com) (140.242.64.158) by server-16.tower-47.messagelabs.com with AES128-SHA encrypted SMTP; 14 Apr 2014 02:26:42 -0000
Received: from CRPMBOXPRD08.polycom.com ([169.254.1.94]) by crpehubprd02.polycom.com ([::1]) with mapi; Sun, 13 Apr 2014 19:26:19 -0700
From: "Duckworth, Mark" <Mark.Duckworth@polycom.com>
To: Christian Groves <Christian.Groves@nteczone.com>, "clue@ietf.org" <clue@ietf.org>
Date: Sun, 13 Apr 2014 19:26:18 -0700
Thread-Topic: [clue] Improving treatment of audio
Thread-Index: Ac9XhmgttdowIAHITNquLh18jhFG6wAAaLbQ
Message-ID: <5C4AC54BFF7A0842A6A11F554D6FB52F07152B@CRPMBOXPRD08.polycom.com>
References: <533AF351.9050201@alum.mit.edu> <20140410222705.GW39240@verdi> <BLU0-SMTP6910DA20C9A9287137B585D0550@phx.gbl> <53475572.4040807@nteczone.com> <20140411150055.GE60844@verdi> <49E45C59CA48264997FEBFB29B6BC2D6215295A8AE@CRPMBOXPRD07.polycom.com> <BLU0-SMTP1804A70E3E0F666EAEAE84D0540@phx.gbl> <49E45C59CA48264997FEBFB29B6BC2D62409219636@CRPMBOXPRD07.polycom.com> <534B4303.7060707@nteczone.com>
In-Reply-To: <534B4303.7060707@nteczone.com>
Accept-Language: en-US
Content-Language: en-US
X-MS-Has-Attach: 
X-MS-TNEF-Correlator: 
acceptlanguage: en-US
Content-Type: text/plain; charset="us-ascii"
Content-Transfer-Encoding: quoted-printable
MIME-Version: 1.0
Archived-At: http://mailarchive.ietf.org/arch/msg/clue/uZRFauhc2j_UaR04vDEmafji8_k
Subject: Re: [clue] Improving treatment of audio
X-BeenThere: clue@ietf.org
X-Mailman-Version: 2.1.15
Precedence: list
List-Id: CLUE - ControLling mUltiple streams for TElepresence <clue.ietf.org>
List-Unsubscribe: <https://www.ietf.org/mailman/options/clue>, <mailto:clue-request@ietf.org?subject=unsubscribe>
List-Archive: <http://www.ietf.org/mail-archive/web/clue/>
List-Post: <mailto:clue@ietf.org>
List-Help: <mailto:clue-request@ietf.org?subject=help>
List-Subscribe: <https://www.ietf.org/mailman/listinfo/clue>, <mailto:clue-request@ietf.org?subject=subscribe>
X-List-Received-Date: Mon, 14 Apr 2014 02:26:48 -0000

Hello Christian,
please see below.
Mark

> -----Original Message-----
> From: clue [mailto:clue-bounces@ietf.org] On Behalf Of Christian Groves
> Sent: Sunday, April 13, 2014 10:08 PM
> To: clue@ietf.org
> Subject: Re: [clue] Improving treatment of audio
>=20
> Hello all,
>=20
> Perhaps the confusion is that some see Audio Capture area as a means to
> associate an Audio capture with a video capture as you've shown Mark's
> example. e.g. The audio spatial information isn't really used for any aud=
io
> transformation other than associating it with a particular video stream.
>=20
> Whereas others were more thinking the audio spatial information as an inp=
ut
> to mixing and more complicated audio processing.

[Duckworth, Mark] You could be right about this being a source of confusion=
.

> If this is the case perhaps rather than linking ACs and VCs through the
> physical or virtual co-ordinates of the area of capture information, we
> simplify things and in each audio capture we say what VC it relates to?
> e.g.using Mark's example AC0(VC0), AC3(V0,V1,V2)

[Duckworth, Mark] I'm not sure how this would really work.  Because even in=
 this simple example, we also have AC0 relates to VC3 (sometimes), VC4, and=
 VC5 (at least part of it), and so on.  AC3 relates to all of VC1 through V=
C5.  I think using area of capture works better than trying to do something=
 like this.

Regards,
Mark

> Regards, Christian
>=20
> On 12/04/2014 7:42 AM, Duckworth, Mark wrote:
> > Hi Paul,
> > Yes, I understand the area of capture for audio can't be nearly as prec=
ise as
> it can be for video.  But for this usage I think it doesn't matter.
> > Mark
> >
> >> -----Original Message-----
> >> From: Paul Coverdale [mailto:coverdale@sympatico.ca]
> >> Sent: Friday, April 11, 2014 4:48 PM
> >> To: Duckworth, Mark; 'John Leslie'; 'Christian Groves'
> >> Cc: clue@ietf.org
> >> Subject: RE: [clue] Improving treatment of audio
> >>
> >> Hi Mark,
> >>
> >> I can see what you're trying to do in the example given in Framework
> >> section 12.1.1. The problem I have is that, because of the huge
> >> difference in wavelength between light waves and audio waves, you
> >> can't really define an area of capture for audio in the same way that
> >> you can for video. A camera can focus on a scene defined by 4
> >> co-planar X,Y,Z coordinates. It will capture the video inside that
> >> quadrilateral, and nothing outside. A microphone can capture audio
> >> coming from inside of the same quadrilateral, but it will also
> >> capture a lot of audio from outside. What this means is that we can
> >> never define a unique association of an audio capture with a video
> >> capture based on pure geometrical considerations, but we can still sim=
ply
> define that ACx is associated with VCx, but maybe with VCy and VCz too.
> >>
> >> ...Paul
> >>
> >>> -----Original Message-----
> >>> From: clue [mailto:clue-bounces@ietf.org] On Behalf Of Duckworth,
> >>> Mark
> >>> Sent: Friday, April 11, 2014 3:24 PM
> >>> To: John Leslie; Christian Groves
> >>> Cc: clue@ietf.org
> >>> Subject: Re: [clue] Improving treatment of audio
> >>>
> >>> Hi all,
> >>> The original intent of using "area of capture" for all media types,
> >>> including audio and video, was to provide a way to associate audio
> >>> with video.  So the consumer can choose captures that go together
> >>> and render them together.  An example of this is in the framework
> >>> document, section 12.1.1.  I still think this works fine for this pur=
pose.
> >>>
> >>> In that example, the consumer can choose VC0, VC1, VC2, AC0, AC1,
> >>> and AC2.  AC0 and VC0 have the same area of capture, indicating they
> >>> are capturing the same area of the scene.  If the consumer wants to
> >>> render
> >>> AC0 from a loudspeaker close to the display for VC0 it can do so.
> >>> Similarly, AC3 has an area of capture covering the whole scene (the
> >>> full extent of the areas of VC0, VC1, VC2) so the consumer knows AC3
> >>> includes audio associated with all of VC0, VC1, and VC2.
> >>>
> >>> So I'm puzzled why you are proposing we remove the ability to use
> >>> area of capture for audio for this purpose.
> >>>
> >>> Regards,
> >>> Mark
> >>>
> >>>> -----Original Message-----
> >>>> From: clue [mailto:clue-bounces@ietf.org] On Behalf Of John Leslie
> >>>> Sent: Friday, April 11, 2014 11:01 AM
> >>>> To: Christian Groves
> >>>> Cc: clue@ietf.org
> >>>> Subject: Re: [clue] Improving treatment of audio
> >>>>
> >>>> Christian Groves <Christian.Groves@nteczone.com> wrote:
> >>>>> "So it's not clear to me what extra we need to do from a CLUE
> >>>>> perspective. What is the problem?"
> >>>>> I guess this is the pertinent point.
> >>>>>
> >>>>>  From John's Audio 101 it seems that anything to do with spatial
> >>>>> information that would relate to reverberation is in the too hard
> >>>>> basket.
> >>>>     What precisely does "too hard basket" mean?
> >>>>
> >>>>     I haven't even reached the point of suggesting what metrics to
> >>>> define the _ability_ to send. To me, "too hard" merely means that
> >>>> an individual site could choose not to send them or to ignore them
> >>>> on
> >>> receipt.
> >>>>     But it sounds as if you're suggesting reverberation is "to hard
> >>>> to understand" and thus we should have no metrics about it.
> >>>>
> >>>>     I _hope_ that's not what you mean.
> >>>>
> >>>>> So it seems "area of capture" for audio could be marked "not
> >>>>> applicable" in the framework.
> >>>>     I hope so.
> >>>>
> >>>>> There doesn't seem to be any driver for having the "point of
> >>> capture"
> >>>>> apply to an audio capture either.
> >>>>     I don't understand this. Point of capture for a microphone may
> >>>> be "too hard" to track (today) for a microphone which moves;
> >>>> nonetheless it seems to me the most fundamental metric there could
> be.
> >>>>
> >>>>>  From the Audio101 there does seem to be a dependency on where
> the
> >>>>> microphone is located with respect to the person speaking (i.e.
> >>>>> lapel mic, desk mic, room mic) to how it is handled at
> >>> mixing/playout.
> >>>>     I don't think I really got that far...
> >>>>
> >>>>     IMHO, it's more flexible to have mixing be the responsibily of
> >>>> the receiver, but I haven't tried to specify that and I'm not at
> >>>> all sure
> >>> I want to specify that.
> >>>>     I expect the actual sound systems in different rooms to vary
> >>>> wildly, from one monaural speaker to stereo to full surround-sound.
> >>>> Mixing for these without knowing which is the actual target seems
> >>>> hard; but I'm sure there will be sites which prefer to do so. A
> >>>> question which will arise, IMHO, is how to specify the _intent_ of
> >>>> a mix generated in one room to be fed to other rooms. (I'd prefer
> >>>> not to go there yet.)
> >>>>
> >>>>> Perhaps this is useful to signal via CLUE? If it is possible to
> >>>>> signal this then perhaps tying a particular audio capture to a
> >>>>> video capture makes sense?
> >>>>     I'm not thinking along those lines. (That doesn't mean we
> >>>> shouldn't think along those lines...) I'm thinking in terms of
> >>>> providing several audio streams per room, associated with position
> >>>> information about the position of the source of those sounds, and
> >>>> allowing the receiver to choose how to mix them and how to present
> >>>> the
> >> mix in his/her room.
> >>>>> i.e. a talker giving a presentation using a lapel mic captured by
> >>>>> a particular video.
> >>>>     In fact, there's only limited tendency for humans to strictly
> >>>> attach the sound they hear to the video they see. Clearly, during a
> >>>> presentation we want to _hear_ the presenter talking, but having
> >>>> the sound move back and forth as the speaker walks can be confusing.
> >>>>
> >>>>     At the same time, we will want to hear the questions to which
> >>>> the presenter may respond. It will be far easier on the listener if
> >>>> these do _not_ seem to be coming from the same physical position,
> >>>> especially when the questioner _is_ in the same room as the
> presenter.
> >>>>
> >>>>     Hope this helps...
> >>>>
> >>>> --
> >>>> John Leslie <john@jlc.net>
> >>>>
> >>>> _______________________________________________
> >>>> clue mailing list
> >>>> clue@ietf.org
> >>>> https://www.ietf.org/mailman/listinfo/clue
> >>> _______________________________________________
> >>> clue mailing list
> >>> clue@ietf.org
> >>> https://www.ietf.org/mailman/listinfo/clue
> > _______________________________________________
> > clue mailing list
> > clue@ietf.org
> > https://www.ietf.org/mailman/listinfo/clue
> >
>=20
> _______________________________________________
> clue mailing list
> clue@ietf.org
> https://www.ietf.org/mailman/listinfo/clue


From nobody Sun Apr 13 20:18:13 2014
Return-Path: <Christian.Groves@nteczone.com>
X-Original-To: clue@ietfa.amsl.com
Delivered-To: clue@ietfa.amsl.com
Received: from localhost (ietfa.amsl.com [127.0.0.1]) by ietfa.amsl.com (Postfix) with ESMTP id 3D5E41A0324 for <clue@ietfa.amsl.com>; Sun, 13 Apr 2014 20:18:11 -0700 (PDT)
X-Virus-Scanned: amavisd-new at amsl.com
X-Spam-Flag: NO
X-Spam-Score: -1.9
X-Spam-Level: 
X-Spam-Status: No, score=-1.9 tagged_above=-999 required=5 tests=[BAYES_00=-1.9] autolearn=ham
Received: from mail.ietf.org ([4.31.198.44]) by localhost (ietfa.amsl.com [127.0.0.1]) (amavisd-new, port 10024) with ESMTP id AKoAD0JH8YT5 for <clue@ietfa.amsl.com>; Sun, 13 Apr 2014 20:18:08 -0700 (PDT)
Received: from cserver5.myshophosting.com (cserver5.myshophosting.com [175.107.161.1]) by ietfa.amsl.com (Postfix) with ESMTP id 23F7C1A031C for <clue@ietf.org>; Sun, 13 Apr 2014 20:18:08 -0700 (PDT)
Received: from ppp118-209-221-158.lns20.mel6.internode.on.net ([118.209.221.158]:53884 helo=[127.0.0.1]) by cserver5.myshophosting.com with esmtpsa (TLSv1:DHE-RSA-AES128-SHA:128) (Exim 4.82) (envelope-from <Christian.Groves@nteczone.com>) id 1WZXPG-0005LQ-OB; Mon, 14 Apr 2014 13:18:02 +1000
Message-ID: <534B536B.8000205@nteczone.com>
Date: Mon, 14 Apr 2014 13:18:03 +1000
From: Christian Groves <Christian.Groves@nteczone.com>
User-Agent: Mozilla/5.0 (Windows NT 6.1; WOW64; rv:24.0) Gecko/20100101 Thunderbird/24.4.0
MIME-Version: 1.0
To: "Duckworth, Mark" <Mark.Duckworth@polycom.com>,  "clue@ietf.org" <clue@ietf.org>
References: <533AF351.9050201@alum.mit.edu> <20140410222705.GW39240@verdi> <BLU0-SMTP6910DA20C9A9287137B585D0550@phx.gbl> <53475572.4040807@nteczone.com> <20140411150055.GE60844@verdi> <49E45C59CA48264997FEBFB29B6BC2D6215295A8AE@CRPMBOXPRD07.polycom.com> <BLU0-SMTP1804A70E3E0F666EAEAE84D0540@phx.gbl> <49E45C59CA48264997FEBFB29B6BC2D62409219636@CRPMBOXPRD07.polycom.com> <534B4303.7060707@nteczone.com> <5C4AC54BFF7A0842A6A11F554D6FB52F07152B@CRPMBOXPRD08.polycom.com>
In-Reply-To: <5C4AC54BFF7A0842A6A11F554D6FB52F07152B@CRPMBOXPRD08.polycom.com>
Content-Type: text/plain; charset=ISO-8859-1; format=flowed
Content-Transfer-Encoding: 7bit
X-AntiAbuse: This header was added to track abuse, please include it with any abuse report
X-AntiAbuse: Primary Hostname - cserver5.myshophosting.com
X-AntiAbuse: Original Domain - ietf.org
X-AntiAbuse: Originator/Caller UID/GID - [47 12] / [47 12]
X-AntiAbuse: Sender Address Domain - nteczone.com
X-Get-Message-Sender-Via: cserver5.myshophosting.com: authenticated_id: christian.groves@nteczone.com
X-Source: 
X-Source-Args: 
X-Source-Dir: 
Archived-At: http://mailarchive.ietf.org/arch/msg/clue/mtoHsWZjIlWpTQZxR-yH53O5M80
Subject: Re: [clue] Improving treatment of audio
X-BeenThere: clue@ietf.org
X-Mailman-Version: 2.1.15
Precedence: list
List-Id: CLUE - ControLling mUltiple streams for TElepresence <clue.ietf.org>
List-Unsubscribe: <https://www.ietf.org/mailman/options/clue>, <mailto:clue-request@ietf.org?subject=unsubscribe>
List-Archive: <http://www.ietf.org/mail-archive/web/clue/>
List-Post: <mailto:clue@ietf.org>
List-Help: <mailto:clue-request@ietf.org?subject=help>
List-Subscribe: <https://www.ietf.org/mailman/listinfo/clue>, <mailto:clue-request@ietf.org?subject=subscribe>
X-List-Received-Date: Mon, 14 Apr 2014 03:18:11 -0000

Hello Mark,

Sorry I didn't consider the entire 12.1.1. if I do that, according to 
that example:

Video areas of capture:

        bottom left    bottom right  top left         top right
    VC0 (-2011,2850,0) (-673,3000,0) (-2011,2850,757) (-673,3000,757)
    VC1 ( -673,3000,0) ( 673,3000,0) ( -673,3000,757) ( 673,3000,757)
    VC2 (  673,3000,0) (2011,2850,0) (  673,3000,757) (2011,3000,757)
    VC3 (-2011,2850,0) (2011,2850,0) (-2011,2850,757) (2011,3000,757)
    VC4 (-2011,2850,0) (2011,2850,0) (-2011,2850,757) (2011,3000,757)
    VC5 (-2011,2850,0) (2011,2850,0) (-2011,2850,757) (2011,3000,757)
    VC6 none

Areas of capture for audio (from 12.1.1):

        bottom left    bottom right  top left         top right

    AC0 (-2011,2850,0) (-673,3000,0) (-2011,2850,757) (-673,3000,757)
    AC1 (  673,3000,0) (2011,2850,0) (  673,3000,757) (2011,3000,757)
    AC2 ( -673,3000,0) ( 673,3000,0) ( -673,3000,757) ( 673,3000,757)
    AC3 (-2011,2850,0) (2011,2850,0) (-2011,2850,757) (2011,3000,757)
    AC4 none

Using a reference rather than area of capture:
     AC0 (VC0)
     AC1 (VC2)
     AC2 (VC1)
     AC3 (VC3,VC4,VC5)
I would think they are basically conveying the same information???

Regards, Christian

On 14/04/2014 12:26 PM, Duckworth, Mark wrote:
> Hello Christian,
> please see below.
> Mark
>
>> -----Original Message-----
>> From: clue [mailto:clue-bounces@ietf.org] On Behalf Of Christian Groves
>> Sent: Sunday, April 13, 2014 10:08 PM
>> To: clue@ietf.org
>> Subject: Re: [clue] Improving treatment of audio
>>
>> Hello all,
>>
>> Perhaps the confusion is that some see Audio Capture area as a means to
>> associate an Audio capture with a video capture as you've shown Mark's
>> example. e.g. The audio spatial information isn't really used for any audio
>> transformation other than associating it with a particular video stream.
>>
>> Whereas others were more thinking the audio spatial information as an input
>> to mixing and more complicated audio processing.
> [Duckworth, Mark] You could be right about this being a source of confusion.
>
>> If this is the case perhaps rather than linking ACs and VCs through the
>> physical or virtual co-ordinates of the area of capture information, we
>> simplify things and in each audio capture we say what VC it relates to?
>> e.g.using Mark's example AC0(VC0), AC3(V0,V1,V2)
> [Duckworth, Mark] I'm not sure how this would really work.  Because even in this simple example, we also have AC0 relates to VC3 (sometimes), VC4, and VC5 (at least part of it), and so on.  AC3 relates to all of VC1 through VC5.  I think using area of capture works better than trying to do something like this.
>
> Regards,
> Mark
>
>> Regards, Christian
>>
>> On 12/04/2014 7:42 AM, Duckworth, Mark wrote:
>>> Hi Paul,
>>> Yes, I understand the area of capture for audio can't be nearly as precise as
>> it can be for video.  But for this usage I think it doesn't matter.
>>> Mark
>>>
>>>> -----Original Message-----
>>>> From: Paul Coverdale [mailto:coverdale@sympatico.ca]
>>>> Sent: Friday, April 11, 2014 4:48 PM
>>>> To: Duckworth, Mark; 'John Leslie'; 'Christian Groves'
>>>> Cc: clue@ietf.org
>>>> Subject: RE: [clue] Improving treatment of audio
>>>>
>>>> Hi Mark,
>>>>
>>>> I can see what you're trying to do in the example given in Framework
>>>> section 12.1.1. The problem I have is that, because of the huge
>>>> difference in wavelength between light waves and audio waves, you
>>>> can't really define an area of capture for audio in the same way that
>>>> you can for video. A camera can focus on a scene defined by 4
>>>> co-planar X,Y,Z coordinates. It will capture the video inside that
>>>> quadrilateral, and nothing outside. A microphone can capture audio
>>>> coming from inside of the same quadrilateral, but it will also
>>>> capture a lot of audio from outside. What this means is that we can
>>>> never define a unique association of an audio capture with a video
>>>> capture based on pure geometrical considerations, but we can still simply
>> define that ACx is associated with VCx, but maybe with VCy and VCz too.
>>>> ...Paul
>>>>
>>>>> -----Original Message-----
>>>>> From: clue [mailto:clue-bounces@ietf.org] On Behalf Of Duckworth,
>>>>> Mark
>>>>> Sent: Friday, April 11, 2014 3:24 PM
>>>>> To: John Leslie; Christian Groves
>>>>> Cc: clue@ietf.org
>>>>> Subject: Re: [clue] Improving treatment of audio
>>>>>
>>>>> Hi all,
>>>>> The original intent of using "area of capture" for all media types,
>>>>> including audio and video, was to provide a way to associate audio
>>>>> with video.  So the consumer can choose captures that go together
>>>>> and render them together.  An example of this is in the framework
>>>>> document, section 12.1.1.  I still think this works fine for this purpose.
>>>>>
>>>>> In that example, the consumer can choose VC0, VC1, VC2, AC0, AC1,
>>>>> and AC2.  AC0 and VC0 have the same area of capture, indicating they
>>>>> are capturing the same area of the scene.  If the consumer wants to
>>>>> render
>>>>> AC0 from a loudspeaker close to the display for VC0 it can do so.
>>>>> Similarly, AC3 has an area of capture covering the whole scene (the
>>>>> full extent of the areas of VC0, VC1, VC2) so the consumer knows AC3
>>>>> includes audio associated with all of VC0, VC1, and VC2.
>>>>>
>>>>> So I'm puzzled why you are proposing we remove the ability to use
>>>>> area of capture for audio for this purpose.
>>>>>
>>>>> Regards,
>>>>> Mark
>>>>>
>>>>>> -----Original Message-----
>>>>>> From: clue [mailto:clue-bounces@ietf.org] On Behalf Of John Leslie
>>>>>> Sent: Friday, April 11, 2014 11:01 AM
>>>>>> To: Christian Groves
>>>>>> Cc: clue@ietf.org
>>>>>> Subject: Re: [clue] Improving treatment of audio
>>>>>>
>>>>>> Christian Groves <Christian.Groves@nteczone.com> wrote:
>>>>>>> "So it's not clear to me what extra we need to do from a CLUE
>>>>>>> perspective. What is the problem?"
>>>>>>> I guess this is the pertinent point.
>>>>>>>
>>>>>>>   From John's Audio 101 it seems that anything to do with spatial
>>>>>>> information that would relate to reverberation is in the too hard
>>>>>>> basket.
>>>>>>      What precisely does "too hard basket" mean?
>>>>>>
>>>>>>      I haven't even reached the point of suggesting what metrics to
>>>>>> define the _ability_ to send. To me, "too hard" merely means that
>>>>>> an individual site could choose not to send them or to ignore them
>>>>>> on
>>>>> receipt.
>>>>>>      But it sounds as if you're suggesting reverberation is "to hard
>>>>>> to understand" and thus we should have no metrics about it.
>>>>>>
>>>>>>      I _hope_ that's not what you mean.
>>>>>>
>>>>>>> So it seems "area of capture" for audio could be marked "not
>>>>>>> applicable" in the framework.
>>>>>>      I hope so.
>>>>>>
>>>>>>> There doesn't seem to be any driver for having the "point of
>>>>> capture"
>>>>>>> apply to an audio capture either.
>>>>>>      I don't understand this. Point of capture for a microphone may
>>>>>> be "too hard" to track (today) for a microphone which moves;
>>>>>> nonetheless it seems to me the most fundamental metric there could
>> be.
>>>>>>>   From the Audio101 there does seem to be a dependency on where
>> the
>>>>>>> microphone is located with respect to the person speaking (i.e.
>>>>>>> lapel mic, desk mic, room mic) to how it is handled at
>>>>> mixing/playout.
>>>>>>      I don't think I really got that far...
>>>>>>
>>>>>>      IMHO, it's more flexible to have mixing be the responsibily of
>>>>>> the receiver, but I haven't tried to specify that and I'm not at
>>>>>> all sure
>>>>> I want to specify that.
>>>>>>      I expect the actual sound systems in different rooms to vary
>>>>>> wildly, from one monaural speaker to stereo to full surround-sound.
>>>>>> Mixing for these without knowing which is the actual target seems
>>>>>> hard; but I'm sure there will be sites which prefer to do so. A
>>>>>> question which will arise, IMHO, is how to specify the _intent_ of
>>>>>> a mix generated in one room to be fed to other rooms. (I'd prefer
>>>>>> not to go there yet.)
>>>>>>
>>>>>>> Perhaps this is useful to signal via CLUE? If it is possible to
>>>>>>> signal this then perhaps tying a particular audio capture to a
>>>>>>> video capture makes sense?
>>>>>>      I'm not thinking along those lines. (That doesn't mean we
>>>>>> shouldn't think along those lines...) I'm thinking in terms of
>>>>>> providing several audio streams per room, associated with position
>>>>>> information about the position of the source of those sounds, and
>>>>>> allowing the receiver to choose how to mix them and how to present
>>>>>> the
>>>> mix in his/her room.
>>>>>>> i.e. a talker giving a presentation using a lapel mic captured by
>>>>>>> a particular video.
>>>>>>      In fact, there's only limited tendency for humans to strictly
>>>>>> attach the sound they hear to the video they see. Clearly, during a
>>>>>> presentation we want to _hear_ the presenter talking, but having
>>>>>> the sound move back and forth as the speaker walks can be confusing.
>>>>>>
>>>>>>      At the same time, we will want to hear the questions to which
>>>>>> the presenter may respond. It will be far easier on the listener if
>>>>>> these do _not_ seem to be coming from the same physical position,
>>>>>> especially when the questioner _is_ in the same room as the
>> presenter.
>>>>>>      Hope this helps...
>>>>>>
>>>>>> --
>>>>>> John Leslie <john@jlc.net>
>>>>>>
>>>>>> _______________________________________________
>>>>>> clue mailing list
>>>>>> clue@ietf.org
>>>>>> https://www.ietf.org/mailman/listinfo/clue
>>>>> _______________________________________________
>>>>> clue mailing list
>>>>> clue@ietf.org
>>>>> https://www.ietf.org/mailman/listinfo/clue
>>> _______________________________________________
>>> clue mailing list
>>> clue@ietf.org
>>> https://www.ietf.org/mailman/listinfo/clue
>>>
>> _______________________________________________
>> clue mailing list
>> clue@ietf.org
>> https://www.ietf.org/mailman/listinfo/clue


From nobody Mon Apr 14 05:31:52 2014
Return-Path: <scarlett.liuyan@huawei.com>
X-Original-To: clue@ietfa.amsl.com
Delivered-To: clue@ietfa.amsl.com
Received: from localhost (ietfa.amsl.com [127.0.0.1]) by ietfa.amsl.com (Postfix) with ESMTP id CCE5F1A01E6 for <clue@ietfa.amsl.com>; Mon, 14 Apr 2014 05:31:48 -0700 (PDT)
X-Virus-Scanned: amavisd-new at amsl.com
X-Spam-Flag: NO
X-Spam-Score: -4.473
X-Spam-Level: 
X-Spam-Status: No, score=-4.473 tagged_above=-999 required=5 tests=[BAYES_00=-1.9, RCVD_IN_DNSWL_MED=-2.3, RP_MATCHES_RCVD=-0.272, SPF_PASS=-0.001] autolearn=ham
Received: from mail.ietf.org ([4.31.198.44]) by localhost (ietfa.amsl.com [127.0.0.1]) (amavisd-new, port 10024) with ESMTP id U02_9SJFFEjA for <clue@ietfa.amsl.com>; Mon, 14 Apr 2014 05:31:46 -0700 (PDT)
Received: from lhrrgout.huawei.com (lhrrgout.huawei.com [194.213.3.17]) by ietfa.amsl.com (Postfix) with ESMTP id 0544C1A01CE for <clue@ietf.org>; Mon, 14 Apr 2014 05:31:45 -0700 (PDT)
Received: from 172.18.7.190 (EHLO lhreml204-edg.china.huawei.com) ([172.18.7.190]) by lhrrg01-dlp.huawei.com (MOS 4.3.7-GA FastPath queued) with ESMTP id BFP64906; Mon, 14 Apr 2014 12:31:42 +0000 (GMT)
Received: from LHREML405-HUB.china.huawei.com (10.201.5.242) by lhreml204-edg.china.huawei.com (172.18.7.223) with Microsoft SMTP Server (TLS) id 14.3.158.1; Mon, 14 Apr 2014 13:30:24 +0100
Received: from SZXEMA408-HUB.china.huawei.com (10.82.72.40) by lhreml405-hub.china.huawei.com (10.201.5.242) with Microsoft SMTP Server (TLS) id 14.3.158.1; Mon, 14 Apr 2014 13:31:40 +0100
Received: from SZXEMA503-MBX.china.huawei.com ([169.254.5.101]) by SZXEMA408-HUB.china.huawei.com ([10.82.72.40]) with mapi id 14.03.0158.001; Mon, 14 Apr 2014 20:31:37 +0800
From: "Liuyan (Scarlett)" <scarlett.liuyan@huawei.com>
To: "clue@ietf.org" <clue@ietf.org>
Thread-Topic: [clue] I-D Action: draft-ietf-clue-data-model-schema-04.txt
Thread-Index: AQHPRShhZNm1KN9smkaiRMdYFcxj5psRLFHQ
Date: Mon, 14 Apr 2014 12:31:37 +0000
Message-ID: <E97C33E5D67C2843AFA535DE2E5B80C1573A6D70@SZXEMA503-MBX.china.huawei.com>
References: <20140321170900.10202.65398.idtracker@ietfa.amsl.com>
In-Reply-To: <20140321170900.10202.65398.idtracker@ietfa.amsl.com>
Accept-Language: zh-CN, en-US
Content-Language: zh-CN
X-MS-Has-Attach: 
X-MS-TNEF-Correlator: 
x-originating-ip: [10.66.137.62]
Content-Type: text/plain; charset="us-ascii"
Content-Transfer-Encoding: quoted-printable
MIME-Version: 1.0
X-CFilter-Loop: Reflected
Archived-At: http://mailarchive.ietf.org/arch/msg/clue/_4PXRgJfGNpgnPjUvwvI_8zRusU
Subject: Re: [clue] I-D Action: draft-ietf-clue-data-model-schema-04.txt
X-BeenThere: clue@ietf.org
X-Mailman-Version: 2.1.15
Precedence: list
List-Id: CLUE - ControLling mUltiple streams for TElepresence <clue.ietf.org>
List-Unsubscribe: <https://www.ietf.org/mailman/options/clue>, <mailto:clue-request@ietf.org?subject=unsubscribe>
List-Archive: <http://www.ietf.org/mail-archive/web/clue/>
List-Post: <mailto:clue@ietf.org>
List-Help: <mailto:clue-request@ietf.org?subject=help>
List-Subscribe: <https://www.ietf.org/mailman/listinfo/clue>, <mailto:clue-request@ietf.org?subject=subscribe>
X-List-Received-Date: Mon, 14 Apr 2014 12:31:49 -0000

Hi Roberta,

   There is a typo in the third line:
   "
      <sceneEntry mediaType=3D"audio" sceneEntryID=3D"SE4">
                    <mediaCaptureIDs>
                        <captureIDREF>VC4</captureIDREF>
                    </mediaCaptureIDs>
                </sceneEntry>
   "=20
   Should be
   "
      <sceneEntry mediaType=3D"audio" sceneEntryID=3D"SE4">
                    <mediaCaptureIDs>
                        <captureIDREF>AC0</captureIDREF>
                    </mediaCaptureIDs>
                </sceneEntry>
   "
 =20

Best Regards,
Scarlett
  =20
-----Original Message-----
From: clue [mailto:clue-bounces@ietf.org] On Behalf Of internet-drafts@ietf=
.org
Sent: Saturday, March 22, 2014 1:09 AM
To: i-d-announce@ietf.org
Cc: clue@ietf.org
Subject: [clue] I-D Action: draft-ietf-clue-data-model-schema-04.txt


A New Internet-Draft is available from the on-line Internet-Drafts director=
ies.
 This draft is a work item of the ControLling mUltiple streams for tElepres=
ence Working Group of the IETF.

        Title           : An XML Schema for the CLUE data model
        Authors         : Roberta Presta
                          Simon Pietro Romano
	Filename        : draft-ietf-clue-data-model-schema-04.txt
	Pages           : 54
	Date            : 2014-03-21

Abstract:
   This document provides an XML schema file for the definition of CLUE
   data model types.


The IETF datatracker status page for this draft is:
https://datatracker.ietf.org/doc/draft-ietf-clue-data-model-schema/

There's also a htmlized version available at:
http://tools.ietf.org/html/draft-ietf-clue-data-model-schema-04

A diff from the previous version is available at:
http://www.ietf.org/rfcdiff?url2=3Ddraft-ietf-clue-data-model-schema-04


Please note that it may take a couple of minutes from the time of submissio=
n until the htmlized version and diff are available at tools.ietf.org.

Internet-Drafts are also available by anonymous FTP at:
ftp://ftp.ietf.org/internet-drafts/

_______________________________________________
clue mailing list
clue@ietf.org
https://www.ietf.org/mailman/listinfo/clue


From nobody Mon Apr 14 07:51:39 2014
Return-Path: <pkyzivat@alum.mit.edu>
X-Original-To: clue@ietfa.amsl.com
Delivered-To: clue@ietfa.amsl.com
Received: from localhost (ietfa.amsl.com [127.0.0.1]) by ietfa.amsl.com (Postfix) with ESMTP id 46E5D1A047E for <clue@ietfa.amsl.com>; Mon, 14 Apr 2014 07:51:38 -0700 (PDT)
X-Virus-Scanned: amavisd-new at amsl.com
X-Spam-Flag: NO
X-Spam-Score: 0.664
X-Spam-Level: 
X-Spam-Status: No, score=0.664 tagged_above=-999 required=5 tests=[BAYES_40=-0.001, DKIM_SIGNED=0.1, DKIM_VALID=-0.1, SPF_SOFTFAIL=0.665] autolearn=no
Received: from mail.ietf.org ([4.31.198.44]) by localhost (ietfa.amsl.com [127.0.0.1]) (amavisd-new, port 10024) with ESMTP id ffultUGe-1NB for <clue@ietfa.amsl.com>; Mon, 14 Apr 2014 07:51:37 -0700 (PDT)
Received: from qmta15.westchester.pa.mail.comcast.net (qmta15.westchester.pa.mail.comcast.net [IPv6:2001:558:fe14:44:76:96:59:228]) by ietfa.amsl.com (Postfix) with ESMTP id 1D63D1A0468 for <clue@ietf.org>; Mon, 14 Apr 2014 07:51:30 -0700 (PDT)
Received: from omta02.westchester.pa.mail.comcast.net ([76.96.62.19]) by qmta15.westchester.pa.mail.comcast.net with comcast id povh1n0030QuhwU5FqrTHW; Mon, 14 Apr 2014 14:51:27 +0000
Received: from Paul-Kyzivats-MacBook-Pro.local ([50.138.229.164]) by omta02.westchester.pa.mail.comcast.net with comcast id pqrT1n00D3ZTu2S3NqrTBF; Mon, 14 Apr 2014 14:51:27 +0000
Message-ID: <534BF5EF.6010601@alum.mit.edu>
Date: Mon, 14 Apr 2014 10:51:27 -0400
From: Paul Kyzivat <pkyzivat@alum.mit.edu>
User-Agent: Mozilla/5.0 (Macintosh; Intel Mac OS X 10.7; rv:24.0) Gecko/20100101 Thunderbird/24.4.0
MIME-Version: 1.0
To: clue@ietf.org
References: <20140411025128.30125.21894.idtracker@ietfa.amsl.com> <C6252EA94E00E44EADC3A2FEB59D440202125EA3@xmb-aln-x07.cisco.com>
In-Reply-To: <C6252EA94E00E44EADC3A2FEB59D440202125EA3@xmb-aln-x07.cisco.com>
Content-Type: text/plain; charset=UTF-8; format=flowed
Content-Transfer-Encoding: 7bit
DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=comcast.net; s=q20140121; t=1397487087; bh=UEqXsWFrchI58z8lP2+qvRcHSS8xAsIMLD2qw445b4Y=; h=Received:Received:Message-ID:Date:From:MIME-Version:To:Subject: Content-Type; b=HTqQJklbxvXy+0paLGtZRKbhF8V9VSvvjeDi+Zq+VxQkAAW+0XvC4BdeVLV+OY0vs Hn5IV6ltLWHnSkvUmwpktQES7GC05fXMupq/8Tfcw68mijblyXlVALAb+gjjs69iy/ 4JZ0JO6zHVRyeVPFjh+whJIUh64MRULtfFVLZI8/sdwXI5KGoOrIkYfdPBXK9Hd5kj 4o6wLlv/x1vrJUwuDANxX9xdmOQabLgUFMyc6ZC1se+SpuVIvx4DwCTtSrrusn6abN PJiB8lswivT7Rqu5qpOTLuvKSPTNag89Bf2fpVmXKRC9M2AhOLGe8oaoAaWMaHiqtA gbAiMEm8RLGeg==
Archived-At: http://mailarchive.ietf.org/arch/msg/clue/pIJQNufy11EDmHX0x96w7Xeh6iw
Subject: Re: [clue] FW: New Version Notification for draft-kyzivat-clue-signaling-08.txt
X-BeenThere: clue@ietf.org
X-Mailman-Version: 2.1.15
Precedence: list
List-Id: CLUE - ControLling mUltiple streams for TElepresence <clue.ietf.org>
List-Unsubscribe: <https://www.ietf.org/mailman/options/clue>, <mailto:clue-request@ietf.org?subject=unsubscribe>
List-Archive: <http://www.ietf.org/mail-archive/web/clue/>
List-Post: <mailto:clue@ietf.org>
List-Help: <mailto:clue-request@ietf.org?subject=help>
List-Subscribe: <https://www.ietf.org/mailman/listinfo/clue>, <mailto:clue-request@ietf.org?subject=subscribe>
X-List-Received-Date: Mon, 14 Apr 2014 14:51:38 -0000

Rob,

Thanks for this update. It is much improved.
Following are some comments on this new version.

	Thanks,
	Paul

Section 4.2:

There is some ambiguity in the use of the term "data channel" in this 
section. In some places it refers (properly) to an individual channel 
over the SCTP association. In other places it refers to the m-line of 
the SCTP association that carries the clue data channel. (E.g., "mid"s 
for more than one data channel.)

Perhaps the last paragraph of this section could say a bit more. The 
feature tag indicates that CLUE is *supported* but does not indicate a 
desire to use it. The telepresence group indicates a desire to *use* 
clue - is CLUE-enabled. So "SHOULD" isn't quite right - it is impossible 
to be CLUE-enabled without including the group. And if the group is 
included, then I think the SCTP m-line MUST be included.

Section 4.3:

    CLUE-controlled media is controlled by the CLUE protocol as
    negotiated on the CLUE data channel with an "mid" included in the
    TELEPRESENCE group.  If no data channel is included in the group the
    other "m" lines in the group are still considered CLUE-controlled and
    under all the restrictions of CLUE-controlled media specified in this
    document.

This allows a telepresence group that doesn't have an SCTP association. 
Do we want to allow that?

Section 4.4:

General observation: this section is now structured into subsections 
according to the usual pattern for describing SDP O/A. When following 
this pattern we need to keep in mind what the subsections mean. The 
"Initial Offer" refers to the first offer of the SDP session, which 
should also be the first offer of the SIP session.

"Modifying the Session" applies to any subsequent offer within the SIP 
and SDP session.

I *think* we have been assuming that the first O/A will contain "legacy" 
media, while sorting out whether CLUE is supported by both ends. Then 
once that is known, a subsequent offer will include the clue-controlled 
media.

We have been working on the assumption that the CLUE channel and/or the 
SCTP association and clue group would be in the initial offer as an 
indicator of clue support. Now that we are proposing to use a feature 
tag to indicate clue support, the initial offer doesn't need to include 
either a tp group or the SCTP m-line.

So, I think that *all* of the SDP that is specific to clue (the 
telepresence group, the SCTP association in the telepresence group, and 
the encodings) should all fall under the "Modifying the Session" section 
of the document. Within that section there will then probably need to be 
section on "Enabling CLUE" and then "Modifying the CLUE session".

Section 4.4.1.3:

If you agree with my point on 4.4, then this section doesn't need a 
special case for the SCTP association.

Section 4.4.2.2:

Regarding RTCP - IIUC this gets complicated once we introduce bundling, 
because there isn't separate RTCP per m-line in the bundle. I don't know 
if that affects what is said here or not. (Let's hear from the RTP experts.)

Section 4.4.1.2 talked about preemptively including recvonly m-lines in 
an offer. But I find nothing talking about generating the answer in that 
case. I think some discussion of that belongs in this section.

Section 4.4.4.1:

This seems to say that the call is clue enabled if *one* side adds the 
SDP association to the telepresence group. IIUC, one side doing this is 
a request to enable clue, and it takes both sides to *enable* clue.

Also note that it would be possible to negotiate an SCTP association 
m-line that is not in the telepresence group, and then at some later 
point simply add it to the telepresence group. That too would enable 
clue. (I don't think the wording needs to change to cover this. But I 
want to make sure people have considered it.)

Section 6.2.1:

If my earlier comments are followed, it may be that the SCTP association 
is first offered concurrently with the first encodings. That might mean 
that no bundle is negotiated until then.

We might want to discuss how to avoid an extra O/A, or extra ports, in 
this case.

One possibility is to create a bundle group around the legacy media. But 
that might not be desired.

Another would be to include the SCTP association as a bundle group with 
only one m-line in the initial offer even though it isn't otherwise 
needed, so that when the encodings are added they can be added to the 
bundle without separate ports.

Another would be to make an offer with the SCTP association in a bundle 
group, and all of the encodings as bundle-only. (That runs the risk that 
the other side doesn't support bundle. If it doesn't, then another offer 
could be made to offer them on separate ports.)


From nobody Mon Apr 14 08:06:04 2014
Return-Path: <john@jlc.net>
X-Original-To: clue@ietfa.amsl.com
Delivered-To: clue@ietfa.amsl.com
Received: from localhost (ietfa.amsl.com [127.0.0.1]) by ietfa.amsl.com (Postfix) with ESMTP id A59F01A04B6 for <clue@ietfa.amsl.com>; Mon, 14 Apr 2014 08:06:02 -0700 (PDT)
X-Virus-Scanned: amavisd-new at amsl.com
X-Spam-Flag: NO
X-Spam-Score: -4.472
X-Spam-Level: 
X-Spam-Status: No, score=-4.472 tagged_above=-999 required=5 tests=[BAYES_00=-1.9, RCVD_IN_DNSWL_MED=-2.3, RP_MATCHES_RCVD=-0.272] autolearn=ham
Received: from mail.ietf.org ([4.31.198.44]) by localhost (ietfa.amsl.com [127.0.0.1]) (amavisd-new, port 10024) with ESMTP id 0X6_hky-G02X for <clue@ietfa.amsl.com>; Mon, 14 Apr 2014 08:05:58 -0700 (PDT)
Received: from mailhost.jlc.net (mailhost.jlc.net [199.201.159.4]) by ietfa.amsl.com (Postfix) with ESMTP id C40C31A02CB for <clue@ietf.org>; Mon, 14 Apr 2014 08:05:53 -0700 (PDT)
Received: by mailhost.jlc.net (Postfix, from userid 104) id 957F5C94C3; Mon, 14 Apr 2014 11:05:49 -0400 (EDT)
Date: Mon, 14 Apr 2014 11:05:49 -0400
From: John Leslie <john@jlc.net>
To: "Duckworth, Mark" <Mark.Duckworth@polycom.com>
Message-ID: <20140414150549.GM60844@verdi>
References: <533AF351.9050201@alum.mit.edu> <20140410222705.GW39240@verdi> <BLU0-SMTP6910DA20C9A9287137B585D0550@phx.gbl> <53475572.4040807@nteczone.com> <20140411150055.GE60844@verdi> <49E45C59CA48264997FEBFB29B6BC2D6215295A8AE@CRPMBOXPRD07.polycom.com> <BLU0-SMTP1804A70E3E0F666EAEAE84D0540@phx.gbl> <49E45C59CA48264997FEBFB29B6BC2D62409219636@CRPMBOXPRD07.polycom.com> <534B4303.7060707@nteczone.com> <5C4AC54BFF7A0842A6A11F554D6FB52F07152B@CRPMBOXPRD08.polycom.com>
Mime-Version: 1.0
Content-Type: text/plain; charset=us-ascii
Content-Disposition: inline
In-Reply-To: <5C4AC54BFF7A0842A6A11F554D6FB52F07152B@CRPMBOXPRD08.polycom.com>
User-Agent: Mutt/1.4.1i
Archived-At: http://mailarchive.ietf.org/arch/msg/clue/aynbcuq2z1sUBjH01S0-uZ___dQ
Cc: "clue@ietf.org" <clue@ietf.org>
Subject: Re: [clue] Improving treatment of audio
X-BeenThere: clue@ietf.org
X-Mailman-Version: 2.1.15
Precedence: list
List-Id: CLUE - ControLling mUltiple streams for TElepresence <clue.ietf.org>
List-Unsubscribe: <https://www.ietf.org/mailman/options/clue>, <mailto:clue-request@ietf.org?subject=unsubscribe>
List-Archive: <http://www.ietf.org/mail-archive/web/clue/>
List-Post: <mailto:clue@ietf.org>
List-Help: <mailto:clue-request@ietf.org?subject=help>
List-Subscribe: <https://www.ietf.org/mailman/listinfo/clue>, <mailto:clue-request@ietf.org?subject=subscribe>
X-List-Received-Date: Mon, 14 Apr 2014 15:06:03 -0000

Duckworth, Mark <Mark.Duckworth@polycom.com> wrote:
> From: clue [mailto:clue-bounces@ietf.org] On Behalf Of Christian Groves
> 
>> If this is the case perhaps rather than linking ACs and VCs through the
>> physical or virtual co-ordinates of the area of capture information, we
>> simplify things and in each audio capture we say what VC it relates to?
>> e.g.using Mark's example AC0(VC0), AC3(V0,V1,V2)
> 
> I'm not sure how this would really work.  Because even in this simple
> example, we also have AC0 relates to VC3 (sometimes), VC4, and VC5
> (at least part of it), and so on. AC3 relates to all of VC1 through VC5.

   Agreed.

> I think using area of capture works better than trying to do something
> like this.

   I need to admit I'm pretty clueless what Mark means here. :^(

   Area-of-capture _is_ four points in 3-space, hopefully co-planar.
I find the concept of relating area-of-capture for a video source to
area-of-capture for an audio source (where the two area-of-captures
will not be coplanar) _extremely_ confusing -- and I really don't
understand how future operators of telepresence systems are going to
find it less confusing.

   (Sorry, Mark...)

--
John Leslie <john@jlc.net>


From nobody Mon Apr 14 08:46:13 2014
Return-Path: <pkyzivat@alum.mit.edu>
X-Original-To: clue@ietfa.amsl.com
Delivered-To: clue@ietfa.amsl.com
Received: from localhost (ietfa.amsl.com [127.0.0.1]) by ietfa.amsl.com (Postfix) with ESMTP id C98AA1A048A for <clue@ietfa.amsl.com>; Mon, 14 Apr 2014 08:46:11 -0700 (PDT)
X-Virus-Scanned: amavisd-new at amsl.com
X-Spam-Flag: NO
X-Spam-Score: 1.465
X-Spam-Level: *
X-Spam-Status: No, score=1.465 tagged_above=-999 required=5 tests=[BAYES_50=0.8, DKIM_SIGNED=0.1, DKIM_VALID=-0.1, SPF_SOFTFAIL=0.665] autolearn=no
Received: from mail.ietf.org ([4.31.198.44]) by localhost (ietfa.amsl.com [127.0.0.1]) (amavisd-new, port 10024) with ESMTP id bZr4OYefHQFd for <clue@ietfa.amsl.com>; Mon, 14 Apr 2014 08:46:08 -0700 (PDT)
Received: from qmta03.westchester.pa.mail.comcast.net (qmta03.westchester.pa.mail.comcast.net [IPv6:2001:558:fe14:43:76:96:62:32]) by ietfa.amsl.com (Postfix) with ESMTP id D8ACF1A0479 for <clue@ietf.org>; Mon, 14 Apr 2014 08:46:07 -0700 (PDT)
Received: from omta24.westchester.pa.mail.comcast.net ([76.96.62.76]) by qmta03.westchester.pa.mail.comcast.net with comcast id prZS1n0071ei1Bg53rm51t; Mon, 14 Apr 2014 15:46:05 +0000
Received: from Paul-Kyzivats-MacBook-Pro.local ([50.138.229.164]) by omta24.westchester.pa.mail.comcast.net with comcast id prm41n00t3ZTu2S3krm5L1; Mon, 14 Apr 2014 15:46:05 +0000
Message-ID: <534C02BC.6000705@alum.mit.edu>
Date: Mon, 14 Apr 2014 11:46:04 -0400
From: Paul Kyzivat <pkyzivat@alum.mit.edu>
User-Agent: Mozilla/5.0 (Macintosh; Intel Mac OS X 10.7; rv:24.0) Gecko/20100101 Thunderbird/24.4.0
MIME-Version: 1.0
To: clue@ietf.org
References: <533AF351.9050201@alum.mit.edu> <BLU0-SMTP942A03E86972E7B8A212C6D0560@phx.gbl> <20140413233154.GL60844@verdi>
In-Reply-To: <20140413233154.GL60844@verdi>
Content-Type: text/plain; charset=ISO-8859-1; format=flowed
Content-Transfer-Encoding: 7bit
DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=comcast.net; s=q20140121; t=1397490365; bh=a9yBDkqISq+4P7qdlMt3BVxrkH4CKFvnVxyPAZXZlTI=; h=Received:Received:Message-ID:Date:From:MIME-Version:To:Subject: Content-Type; b=gDkHY+PnVgSuaV3wjt2UBXDc97fVJIgQarS6dCHJMZhQL50jyXUqvTYShAxcMsnmj RB03GV7/bvdPKkyadENt1oSwOdwm5/aoTMOo3UO6lIAT/SoS4dfAmWYEZ+Lqt68zPm FLyv8KAdaOFQpm6XPeT4EGkT+tIhnVyrvuFrLX6AcJkf/0OpzoG++yoOS39oSJrgzm rzrKuGPBIWbnTQC3qGuLyFxU5YUHaD4Km188trzRZFC6zuwzjNrsiRNA5YD5J146x2 +95o+FSDJPdkV33c5unT6Bi5j52d02F6ztIdBk5YLoK6ih9xu+9kySFxWH9rBlv9rG Qc6WbbTon37FQ==
Archived-At: http://mailarchive.ietf.org/arch/msg/clue/iO3GGo2FWop4O-LHqqEUjOPM1Eo
Subject: Re: [clue] Improving treatment of audio
X-BeenThere: clue@ietf.org
X-Mailman-Version: 2.1.15
Precedence: list
List-Id: CLUE - ControLling mUltiple streams for TElepresence <clue.ietf.org>
List-Unsubscribe: <https://www.ietf.org/mailman/options/clue>, <mailto:clue-request@ietf.org?subject=unsubscribe>
List-Archive: <http://www.ietf.org/mail-archive/web/clue/>
List-Post: <mailto:clue@ietf.org>
List-Help: <mailto:clue-request@ietf.org?subject=help>
List-Subscribe: <https://www.ietf.org/mailman/listinfo/clue>, <mailto:clue-request@ietf.org?subject=subscribe>
X-List-Received-Date: Mon, 14 Apr 2014 15:46:11 -0000

On 4/13/14 7:31 PM, John Leslie wrote:
> Paul Coverdale <coverdale@sympatico.ca> wrote:
>>
>> So, going back to Paul K's request for a concise statement of what needs to
>> be done with respect to the treatment of audio in CLUE, I'm not sure what
>> else is really required apart from satisfying REQMT-2 and REQMT-3,
>
>     I have to agree with Paul-C here: clue-telepresence-requirements is
> pretty much set in stone, now that it has entered the RFCed queue. But
> I don't think it fits Paul-K's criteria of "concise:
> ]
> ] REQMT-2: The solution MUST support a description of the spatial
> ]          arrangement of captured source audio sent in audio streams
> ]          which enables a satisfactory reproduction at the receiver
> ]          in a spatially correct manner.  This applies to each site
> ]          in a point to point or a multipoint meeting and refers to
> ]          the spatial ordering within a site, not the ordering of
> ]          channels between sites.
> ]
> ]          Use case point to point symmetric, and all use cases,
> ]          especially heterogeneous.
> ]
> ]  REQMT-2a: The solution MUST support a means of preserving
> ]            the spatial order of audio in the captured
> ]            scene.  For example, if John sounds as if he is
> ]            at Susan's right in the captured audio, John
> ]            voice is also placed at Susan's right in the
> ]            rendered image.
> ]
> ]  REQMT-2b: The solution MUST support a means to identify
> ]            the number and spatial arrangement of audio
> ]            channels including monaural, stereophonic
> ]            (2.0), and 3.0 (left, center, right) audio
> ]            channels.
> ]
> ]  REQMT-2c: The solution MUST support a means to identify
> ]            the point of capture of individual audio
> ]            captures in three dimensions.
> ]
> ]  REQMT-2d: The solution MUST support a means to identify
> ]            the area of coverage of individual audio
> ]            captures in three dimensions.
> ]
> ] REQMT-3: The solution MUST enable individual audio streams to be
> ]          associated with one or more video image captures, and
> ]          individual video image captures to be associated with one
> ]          or more audio captures, for the purpose of rendering
> ]          proper position.
> ]
> ]          Use case is point to point symmetric, and all use cases.
>
>     Worse, I'm not confident that all of this is even attainable. :^(

I don't know if they are unattainable. But IIUC they cannot be attained 
with the information we currently have in the fw and data model.

If we want to give up on attaining some of them, then let's do so, and 
explicitly note that we have. And then ensure that we have sufficient 
information to attain the rest.

>     Mea culpa, of course -- I should have said that in WGLC. :^( :^(
>
>> with the possible exception of including the directional
>> characteristics of the microphone with the audio capture.
>
>     While I'd be happy to see that signaled, I'm really not pushing it
> for our first spec -- merely hoping we'll get there someday.
>
>     There is, alas, an elephant in each room which can render our
> audio treatment useless -- it's called "echo". :^(

IMO, if we don't have a solution that can prevent this then we don't 
have a solution worth publishing.

> ====
>     Despite being too late, I feel I ought to comment on these REQMTs.
> Feel free to ignore what follows:

As I noted above, I think it is important to confront these issues.

> ] REQMT-2: The solution MUST support a description of the spatial
> ]          arrangement of captured source audio sent in audio streams
> ]          which enables a satisfactory reproduction at the receiver
> ]          in a spatially correct manner.
>
>     Other than wondering what "spatially correct means, I view this as
> attainable, since the responsibily must belong to the receiver, and
> that only if the sender chooses to specify enough audio sources.

What do you mean by "attainable"? Can this be achieved using the data 
model we have now? (I don't think so.)

ISTM that the responsibility is *shared*:
- the advertiser must provide sufficient sources, *and* sufficient
   description of those sources to permit the receiver to do the right
   thing.

- the receiver must utilize the available information in order to
   reproduce the audio in a spatially correct manner, within the limits
   of its available equipment.

- *we* are responsible for providing sufficient expressiveness in
   the advertisement and RTP so that the advertiser and receiver can
   do the above. *That* is the part we need to work on now.

> ]  REQMT-2a: The solution MUST support a means of preserving
> ]            the spatial order of audio in the captured
> ]            scene.  For example, if John sounds as if he is
> ]            at Susan's right in the captured audio, John
> ]            voice is also placed at Susan's right in the
> ]            rendered image.
>
>     Likewise, this responsibility must belong to the receiver.
>
> ]  REQMT-2b: The solution MUST support a means to identify
> ]            the number and spatial arrangement of audio
> ]            channels including monaural, stereophonic
> ]            (2.0), and 3.0 (left, center, right) audio
> ]            channels.
>
>     This, IMHO, isn't worth doing (although perhaps I don't understand
> what REQMT-2b means).
>
>     The actual layout of loudspeakers should belong entirely to the
> receiving room. Further, "stereo" doesn't actually define the relationship
> of the two channels; and "3.0" suffers a similar problem.
>
>     If a sender _actually_ sends "stereo" or "3.0" the receiver will have
> to punt. I can only hope that our actual standard doesn't encourage this.

Sounds like something we need to talk about. Where?

> ]  REQMT-2c: The solution MUST support a means to identify
> ]            the point of capture of individual audio
> ]            captures in three dimensions.
>
>     Well stated. Adding point on line of capture helps slightly, but only
> if we have a useful approximation of sensitivity pattern (which I don't
> believe we can expect to get).
>
> ]  REQMT-2d: The solution MUST support a means to identify
> ]            the area of coverage of individual audio
> ]            captures in three dimensions.
>
>     Poorly stated. :^( _Every_ microphone in a room can cover _every_
> source of sound in a room. Thus, the "area of coverage" must necessarily
> be the entire room...
>
>     (It gets worse... much worse!)
>
>     Inevitably, some yahoo is going to put one or more loudspeakers in
> the room. And those will expand the "area of coverage" to include the
> other rooms those loudspeakers attempt to render.
>
>     With delay!!!

OK. But surely the point of capture isn't sufficient by itself. 
*Something* more is required. Can you suggest something that is stated 
better, that would be attainable and useful for our purposes?

> ] REQMT-3: The solution MUST enable individual audio streams to be
> ]          associated with one or more video image captures, and
> ]          individual video image captures to be associated with one
> ]          or more audio captures, for the purpose of rendering
> ]          proper position.
> ]
> ]          Use case is point to point symmetric, and all use cases.
>
>     I accept there are many folks who will insist on doing this. I don't
> expect to stop them.
>
>     But audio streams are intrinsically linked to rooms, not video captures.
> And the most useful audio streams are linked to microphones "close" to
> individual people, while video captures will _very_ often try to cover
> more than one person -- thus video capures will very typically be associated
> with several audio captures.

Are you saying that within a scene the choice of audio captures is 
unrelated to the choice of video captures? (Other than if you configure 
*some* video then you probably also want to configure *some* audio.)

If so, then does it ever make sense to choose a subset of the advertised 
audio? And if that is true, then what criteria make sense?

> ====
>
>     It will help if more of us pay attention to the relationship of audio
> in television in movies. I find that they _don't_ track each other too
> closely.
>
>     The audio provides the continuity; the video tantalizes the senses.
> When the camera switches between two people talking to each other, the
> audio doesn't switch -- it tries to make you believe you're in the room.
>
>     Thus I most sincerely hope that folks won't choose to switch their
> audio based on who's speaking -- but I won't try to stop them from doing
> this. (And I must admit that sometimes I _shouldn't_ blame them, because
> they find themselves sabotaged by audio sources contaminated by echo.)

I assume you mean when the speakers are in the same room, right?

Isn't the situation different when we are switching among speakers in 
different rooms?

> ====
>
>     A word on "echo suppression" might help here: this is a well-solved
> in plain-old-telephony where end-to-end delay can be held constand. We
> won't be able to do that over the Internet.
>
>     An individual room is probably able to cancel the echo in what it
> _sends_ but a receiving room won't be able to cancel echo in what it
> receives. :^(
>
> ====
>
>     That's as much as I'm willing to sqeeze into this email.
>
>     Hope _something_ in this helps...

Some! This is going to take more work.

	Thanks,
	Paul

> --
> John Leslie <john@jlc.net>
>
> _______________________________________________
> clue mailing list
> clue@ietf.org
> https://www.ietf.org/mailman/listinfo/clue
>


From nobody Mon Apr 14 08:52:07 2014
Return-Path: <pkyzivat@alum.mit.edu>
X-Original-To: clue@ietfa.amsl.com
Delivered-To: clue@ietfa.amsl.com
Received: from localhost (ietfa.amsl.com [127.0.0.1]) by ietfa.amsl.com (Postfix) with ESMTP id 8AED31A03D7 for <clue@ietfa.amsl.com>; Mon, 14 Apr 2014 08:52:05 -0700 (PDT)
X-Virus-Scanned: amavisd-new at amsl.com
X-Spam-Flag: NO
X-Spam-Score: -1.235
X-Spam-Level: 
X-Spam-Status: No, score=-1.235 tagged_above=-999 required=5 tests=[BAYES_00=-1.9, DKIM_SIGNED=0.1, DKIM_VALID=-0.1, SPF_SOFTFAIL=0.665] autolearn=no
Received: from mail.ietf.org ([4.31.198.44]) by localhost (ietfa.amsl.com [127.0.0.1]) (amavisd-new, port 10024) with ESMTP id Dbzud_OQ7xtP for <clue@ietfa.amsl.com>; Mon, 14 Apr 2014 08:52:01 -0700 (PDT)
Received: from qmta05.westchester.pa.mail.comcast.net (qmta05.westchester.pa.mail.comcast.net [IPv6:2001:558:fe14:43:76:96:62:48]) by ietfa.amsl.com (Postfix) with ESMTP id 91AB51A0505 for <clue@ietf.org>; Mon, 14 Apr 2014 08:52:00 -0700 (PDT)
Received: from omta18.westchester.pa.mail.comcast.net ([76.96.62.90]) by qmta05.westchester.pa.mail.comcast.net with comcast id pp2g1n00H1wpRvQ55rry6W; Mon, 14 Apr 2014 15:51:58 +0000
Received: from Paul-Kyzivats-MacBook-Pro.local ([50.138.229.164]) by omta18.westchester.pa.mail.comcast.net with comcast id prrx1n00k3ZTu2S3errxWW; Mon, 14 Apr 2014 15:51:58 +0000
Message-ID: <534C041D.5050903@alum.mit.edu>
Date: Mon, 14 Apr 2014 11:51:57 -0400
From: Paul Kyzivat <pkyzivat@alum.mit.edu>
User-Agent: Mozilla/5.0 (Macintosh; Intel Mac OS X 10.7; rv:24.0) Gecko/20100101 Thunderbird/24.4.0
MIME-Version: 1.0
To: clue@ietf.org
References: <533AF351.9050201@alum.mit.edu> <20140410222705.GW39240@verdi> <BLU0-SMTP6910DA20C9A9287137B585D0550@phx.gbl> <53475572.4040807@nteczone.com> <20140411150055.GE60844@verdi> <49E45C59CA48264997FEBFB29B6BC2D6215295A8AE@CRPMBOXPRD07.polycom.com> <BLU0-SMTP1804A70E3E0F666EAEAE84D0540@phx.gbl> <49E45C59CA48264997FEBFB29B6BC2D62409219636@CRPMBOXPRD07.polycom.com> <534B4303.7060707@nteczone.com>
In-Reply-To: <534B4303.7060707@nteczone.com>
Content-Type: text/plain; charset=ISO-8859-1; format=flowed
Content-Transfer-Encoding: 7bit
DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=comcast.net; s=q20140121; t=1397490718; bh=cOCUn494FVGommCIkxwYbzweCbp6UygHgk0vJbwmTIM=; h=Received:Received:Message-ID:Date:From:MIME-Version:To:Subject: Content-Type; b=lgMD2RfkOW/xC5jWSHzVc+e87bbvY3UhwJL/z2MhZS9Q2WFLs4ZhJjikom9u3mSwd 1JpqqefcYIpUhsgXNKG1X6mrR8oeZTrL93mrv0v7oXllKaK1pIQmgkABYFd6Vjt77D qRlZtG4acPr09W4pGWg05rYhuSFnyhlOoWTXZl7zJQaV046BJrgQNicgHGF4Y7D6z5 niRt+FFGIUBnsAXKzMcU6ekFHjix+K/u4CISAPQ9x0gOtB7vRBfMcMBFu/C+oA98Ky G/W7SWlArMAU7zMn3F9cC2uxxD0m/v5sNRy0iiAUHxlRzBi+Hiy/gNqiaKH16YKPvM /29h864YMjV9g==
Archived-At: http://mailarchive.ietf.org/arch/msg/clue/2TRGQx4PJYY5AV43a_2H5wmN2Qc
Subject: Re: [clue] Improving treatment of audio
X-BeenThere: clue@ietf.org
X-Mailman-Version: 2.1.15
Precedence: list
List-Id: CLUE - ControLling mUltiple streams for TElepresence <clue.ietf.org>
List-Unsubscribe: <https://www.ietf.org/mailman/options/clue>, <mailto:clue-request@ietf.org?subject=unsubscribe>
List-Archive: <http://www.ietf.org/mail-archive/web/clue/>
List-Post: <mailto:clue@ietf.org>
List-Help: <mailto:clue-request@ietf.org?subject=help>
List-Subscribe: <https://www.ietf.org/mailman/listinfo/clue>, <mailto:clue-request@ietf.org?subject=subscribe>
X-List-Received-Date: Mon, 14 Apr 2014 15:52:05 -0000

On 4/13/14 10:08 PM, Christian Groves wrote:
> Hello all,
>
> Perhaps the confusion is that some see Audio Capture area as a means to
> associate an Audio capture with a video capture as you've shown Mark's
> example. e.g. The audio spatial information isn't really used for any
> audio transformation other than associating it with a particular video
> stream.
>
> Whereas others were more thinking the audio spatial information as an
> input to mixing and more complicated audio processing.

Yes.

> If this is the case perhaps rather than linking ACs and VCs through the
> physical or virtual co-ordinates of the area of capture information, we
> simplify things and in each audio capture we say what VC it relates to?
> e.g.using Mark's example AC0(VC0), AC3(V0,V1,V2)

Based on John's comments, I wonder if this is possible at all, for a 
single scene.

Perhaps within a scene it makes no sense to choose audio based on which 
video is chosen. Instead maybe we should simply be trying to satisfy the 
need for mixing and audio processing.

	Thanks,
	Paul

> Regards, Christian
>
> On 12/04/2014 7:42 AM, Duckworth, Mark wrote:
>> Hi Paul,
>> Yes, I understand the area of capture for audio can't be nearly as
>> precise as it can be for video.  But for this usage I think it doesn't
>> matter.
>> Mark
>>
>>> -----Original Message-----
>>> From: Paul Coverdale [mailto:coverdale@sympatico.ca]
>>> Sent: Friday, April 11, 2014 4:48 PM
>>> To: Duckworth, Mark; 'John Leslie'; 'Christian Groves'
>>> Cc: clue@ietf.org
>>> Subject: RE: [clue] Improving treatment of audio
>>>
>>> Hi Mark,
>>>
>>> I can see what you're trying to do in the example given in Framework
>>> section
>>> 12.1.1. The problem I have is that, because of the huge difference in
>>> wavelength between light waves and audio waves, you can't really
>>> define an
>>> area of capture for audio in the same way that you can for video. A
>>> camera
>>> can focus on a scene defined by 4 co-planar X,Y,Z coordinates. It
>>> will capture
>>> the video inside that quadrilateral, and nothing outside. A
>>> microphone can
>>> capture audio coming from inside of the same quadrilateral, but it
>>> will also
>>> capture a lot of audio from outside. What this means is that we can
>>> never
>>> define a unique association of an audio capture with a video capture
>>> based
>>> on pure geometrical considerations, but we can still simply define
>>> that ACx is
>>> associated with VCx, but maybe with VCy and VCz too.
>>>
>>> ...Paul
>>>
>>>> -----Original Message-----
>>>> From: clue [mailto:clue-bounces@ietf.org] On Behalf Of Duckworth, Mark
>>>> Sent: Friday, April 11, 2014 3:24 PM
>>>> To: John Leslie; Christian Groves
>>>> Cc: clue@ietf.org
>>>> Subject: Re: [clue] Improving treatment of audio
>>>>
>>>> Hi all,
>>>> The original intent of using "area of capture" for all media types,
>>>> including audio and video, was to provide a way to associate audio with
>>>> video.  So the consumer can choose captures that go together and render
>>>> them together.  An example of this is in the framework document,
>>>> section 12.1.1.  I still think this works fine for this purpose.
>>>>
>>>> In that example, the consumer can choose VC0, VC1, VC2, AC0, AC1, and
>>>> AC2.  AC0 and VC0 have the same area of capture, indicating they are
>>>> capturing the same area of the scene.  If the consumer wants to render
>>>> AC0 from a loudspeaker close to the display for VC0 it can do so.
>>>> Similarly, AC3 has an area of capture covering the whole scene (the
>>>> full extent of the areas of VC0, VC1, VC2) so the consumer knows AC3
>>>> includes audio associated with all of VC0, VC1, and VC2.
>>>>
>>>> So I'm puzzled why you are proposing we remove the ability to use area
>>>> of capture for audio for this purpose.
>>>>
>>>> Regards,
>>>> Mark
>>>>
>>>>> -----Original Message-----
>>>>> From: clue [mailto:clue-bounces@ietf.org] On Behalf Of John Leslie
>>>>> Sent: Friday, April 11, 2014 11:01 AM
>>>>> To: Christian Groves
>>>>> Cc: clue@ietf.org
>>>>> Subject: Re: [clue] Improving treatment of audio
>>>>>
>>>>> Christian Groves <Christian.Groves@nteczone.com> wrote:
>>>>>> "So it's not clear to me what extra we need to do from a CLUE
>>>>>> perspective. What is the problem?"
>>>>>> I guess this is the pertinent point.
>>>>>>
>>>>>>  From John's Audio 101 it seems that anything to do with spatial
>>>>>> information that would relate to reverberation is in the too hard
>>>>>> basket.
>>>>>     What precisely does "too hard basket" mean?
>>>>>
>>>>>     I haven't even reached the point of suggesting what metrics to
>>>>> define the _ability_ to send. To me, "too hard" merely means that an
>>>>> individual site could choose not to send them or to ignore them on
>>>> receipt.
>>>>>     But it sounds as if you're suggesting reverberation is "to hard to
>>>>> understand" and thus we should have no metrics about it.
>>>>>
>>>>>     I _hope_ that's not what you mean.
>>>>>
>>>>>> So it seems "area of capture" for audio could be marked "not
>>>>>> applicable" in the framework.
>>>>>     I hope so.
>>>>>
>>>>>> There doesn't seem to be any driver for having the "point of
>>>> capture"
>>>>>> apply to an audio capture either.
>>>>>     I don't understand this. Point of capture for a microphone may be
>>>>> "too hard" to track (today) for a microphone which moves; nonetheless
>>>>> it seems to me the most fundamental metric there could be.
>>>>>
>>>>>>  From the Audio101 there does seem to be a dependency on where the
>>>>>> microphone is located with respect to the person speaking (i.e.
>>>>>> lapel mic, desk mic, room mic) to how it is handled at
>>>> mixing/playout.
>>>>>     I don't think I really got that far...
>>>>>
>>>>>     IMHO, it's more flexible to have mixing be the responsibily of the
>>>>> receiver, but I haven't tried to specify that and I'm not at all sure
>>>> I want to specify that.
>>>>>     I expect the actual sound systems in different rooms to vary
>>>>> wildly, from one monaural speaker to stereo to full surround-sound.
>>>>> Mixing for these without knowing which is the actual target seems
>>>>> hard; but I'm sure there will be sites which prefer to do so. A
>>>>> question which will arise, IMHO, is how to specify the _intent_ of a
>>>>> mix generated in one room to be fed to other rooms. (I'd prefer not
>>>>> to go there yet.)
>>>>>
>>>>>> Perhaps this is useful to signal via CLUE? If it is possible to
>>>>>> signal this then perhaps tying a particular audio capture to a
>>>>>> video capture makes sense?
>>>>>     I'm not thinking along those lines. (That doesn't mean we
>>>>> shouldn't think along those lines...) I'm thinking in terms of
>>>>> providing several audio streams per room, associated with position
>>>>> information about the position of the source of those sounds, and
>>>>> allowing the receiver to choose how to mix them and how to present the
>>> mix in his/her room.
>>>>>> i.e. a talker giving a presentation using a lapel mic captured by a
>>>>>> particular video.
>>>>>     In fact, there's only limited tendency for humans to strictly
>>>>> attach the sound they hear to the video they see. Clearly, during a
>>>>> presentation we want to _hear_ the presenter talking, but having the
>>>>> sound move back and forth as the speaker walks can be confusing.
>>>>>
>>>>>     At the same time, we will want to hear the questions to which the
>>>>> presenter may respond. It will be far easier on the listener if these
>>>>> do _not_ seem to be coming from the same physical position,
>>>>> especially when the questioner _is_ in the same room as the presenter.
>>>>>
>>>>>     Hope this helps...
>>>>>
>>>>> --
>>>>> John Leslie <john@jlc.net>
>>>>>
>>>>> _______________________________________________
>>>>> clue mailing list
>>>>> clue@ietf.org
>>>>> https://www.ietf.org/mailman/listinfo/clue
>>>> _______________________________________________
>>>> clue mailing list
>>>> clue@ietf.org
>>>> https://www.ietf.org/mailman/listinfo/clue
>> _______________________________________________
>> clue mailing list
>> clue@ietf.org
>> https://www.ietf.org/mailman/listinfo/clue
>>
>
> _______________________________________________
> clue mailing list
> clue@ietf.org
> https://www.ietf.org/mailman/listinfo/clue
>


From nobody Mon Apr 14 11:11:56 2014
Return-Path: <coverdale@sympatico.ca>
X-Original-To: clue@ietfa.amsl.com
Delivered-To: clue@ietfa.amsl.com
Received: from localhost (ietfa.amsl.com [127.0.0.1]) by ietfa.amsl.com (Postfix) with ESMTP id A6DFA1A0471 for <clue@ietfa.amsl.com>; Mon, 14 Apr 2014 11:11:55 -0700 (PDT)
X-Virus-Scanned: amavisd-new at amsl.com
X-Spam-Flag: NO
X-Spam-Score: 0.8
X-Spam-Level: 
X-Spam-Status: No, score=0.8 tagged_above=-999 required=5 tests=[BAYES_50=0.8,  MSGID_FROM_MTA_HEADER=0.001, RCVD_IN_DNSWL_NONE=-0.0001, SPF_PASS=-0.001] autolearn=ham
Received: from mail.ietf.org ([4.31.198.44]) by localhost (ietfa.amsl.com [127.0.0.1]) (amavisd-new, port 10024) with ESMTP id BWuj8UDrMuXy for <clue@ietfa.amsl.com>; Mon, 14 Apr 2014 11:11:54 -0700 (PDT)
Received: from blu0-omc1-s11.blu0.hotmail.com (blu0-omc1-s11.blu0.hotmail.com [65.55.116.22]) by ietfa.amsl.com (Postfix) with ESMTP id 194E71A032E for <clue@ietf.org>; Mon, 14 Apr 2014 11:11:53 -0700 (PDT)
Received: from BLU0-SMTP26 ([65.55.116.9]) by blu0-omc1-s11.blu0.hotmail.com with Microsoft SMTPSVC(6.0.3790.4675);  Mon, 14 Apr 2014 11:11:51 -0700
X-TMN: [MY/7ly2wRVDWBEfwIa4rfW5PBkhzvIM3]
X-Originating-Email: [coverdale@sympatico.ca]
Message-ID: <BLU0-SMTP26739EFCE85636668D142CD0510@phx.gbl>
Received: from PaulNewPC ([74.15.60.251]) by BLU0-SMTP26.phx.gbl over TLS secured channel with Microsoft SMTPSVC(6.0.3790.4675);  Mon, 14 Apr 2014 11:11:50 -0700
From: Paul Coverdale <coverdale@sympatico.ca>
To: "'Paul Kyzivat'" <pkyzivat@alum.mit.edu>, <clue@ietf.org>
References: <533AF351.9050201@alum.mit.edu> <BLU0-SMTP942A03E86972E7B8A212C6D0560@phx.gbl> <20140413233154.GL60844@verdi> <534C02BC.6000705@alum.mit.edu>
In-Reply-To: <534C02BC.6000705@alum.mit.edu>
Date: Mon, 14 Apr 2014 14:11:48 -0400
MIME-Version: 1.0
Content-Type: text/plain; charset="us-ascii"
Content-Transfer-Encoding: 7bit
X-Mailer: Microsoft Office Outlook 12.0
Thread-Index: Ac9X+Kim/PDqrki0TBCBa1CEWCaUOwABfLDA
Content-Language: en-us
X-OriginalArrivalTime: 14 Apr 2014 18:11:50.0692 (UTC) FILETIME=[01E94E40:01CF580D]
Archived-At: http://mailarchive.ietf.org/arch/msg/clue/cxX7j35f45WpVhdI0UkuCcA4BC8
Subject: Re: [clue] Improving treatment of audio
X-BeenThere: clue@ietf.org
X-Mailman-Version: 2.1.15
Precedence: list
List-Id: CLUE - ControLling mUltiple streams for TElepresence <clue.ietf.org>
List-Unsubscribe: <https://www.ietf.org/mailman/options/clue>, <mailto:clue-request@ietf.org?subject=unsubscribe>
List-Archive: <http://www.ietf.org/mail-archive/web/clue/>
List-Post: <mailto:clue@ietf.org>
List-Help: <mailto:clue-request@ietf.org?subject=help>
List-Subscribe: <https://www.ietf.org/mailman/listinfo/clue>, <mailto:clue-request@ietf.org?subject=subscribe>
X-List-Received-Date: Mon, 14 Apr 2014 18:11:55 -0000

>>     There is, alas, an elephant in each room which can render our
>> audio treatment useless -- it's called "echo". :^(
>
>IMO, if we don't have a solution that can prevent this then we don't
>have a solution worth publishing.
>

Of course echo is important, but I wouldn't get too pessimistic about it.
While it's always possible to screw it up, I suspect that any serious
Telepresence vendor understands acoustic echo control quite well and will
have taken steps to control it. I don't know that there's much else we can
usefully do from a CLUE perspective. 

...Paul


From nobody Mon Apr 14 11:16:46 2014
Return-Path: <pkyzivat@alum.mit.edu>
X-Original-To: clue@ietfa.amsl.com
Delivered-To: clue@ietfa.amsl.com
Received: from localhost (ietfa.amsl.com [127.0.0.1]) by ietfa.amsl.com (Postfix) with ESMTP id 92FA21A0692 for <clue@ietfa.amsl.com>; Mon, 14 Apr 2014 11:16:44 -0700 (PDT)
X-Virus-Scanned: amavisd-new at amsl.com
X-Spam-Flag: NO
X-Spam-Score: -1.235
X-Spam-Level: 
X-Spam-Status: No, score=-1.235 tagged_above=-999 required=5 tests=[BAYES_00=-1.9, DKIM_SIGNED=0.1, DKIM_VALID=-0.1, SPF_SOFTFAIL=0.665] autolearn=no
Received: from mail.ietf.org ([4.31.198.44]) by localhost (ietfa.amsl.com [127.0.0.1]) (amavisd-new, port 10024) with ESMTP id rZ_rF1vEI5r8 for <clue@ietfa.amsl.com>; Mon, 14 Apr 2014 11:16:43 -0700 (PDT)
Received: from qmta01.westchester.pa.mail.comcast.net (qmta01.westchester.pa.mail.comcast.net [IPv6:2001:558:fe14:43:76:96:62:16]) by ietfa.amsl.com (Postfix) with ESMTP id 67FFE1A0673 for <clue@ietf.org>; Mon, 14 Apr 2014 11:16:43 -0700 (PDT)
Received: from omta18.westchester.pa.mail.comcast.net ([76.96.62.90]) by qmta01.westchester.pa.mail.comcast.net with comcast id ppBm1n0011wpRvQ51uGgPk; Mon, 14 Apr 2014 18:16:40 +0000
Received: from Paul-Kyzivats-MacBook-Pro.local ([50.138.229.164]) by omta18.westchester.pa.mail.comcast.net with comcast id puGg1n00R3ZTu2S3euGg2p; Mon, 14 Apr 2014 18:16:40 +0000
Message-ID: <534C2608.9040506@alum.mit.edu>
Date: Mon, 14 Apr 2014 14:16:40 -0400
From: Paul Kyzivat <pkyzivat@alum.mit.edu>
User-Agent: Mozilla/5.0 (Macintosh; Intel Mac OS X 10.7; rv:24.0) Gecko/20100101 Thunderbird/24.4.0
MIME-Version: 1.0
To: Paul Coverdale <coverdale@sympatico.ca>, clue@ietf.org
References: <533AF351.9050201@alum.mit.edu> <BLU0-SMTP942A03E86972E7B8A212C6D0560@phx.gbl> <20140413233154.GL60844@verdi> <534C02BC.6000705@alum.mit.edu> <BLU0-SMTP26739EFCE85636668D142CD0510@phx.gbl>
In-Reply-To: <BLU0-SMTP26739EFCE85636668D142CD0510@phx.gbl>
Content-Type: text/plain; charset=ISO-8859-1; format=flowed
Content-Transfer-Encoding: 7bit
DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=comcast.net; s=q20140121; t=1397499400; bh=cBAz7L71SZoyYrgW6Mlc3eLTdGbgXOXfH9yuvXnCYE0=; h=Received:Received:Message-ID:Date:From:MIME-Version:To:Subject: Content-Type; b=C0plQoDhroDYPuAFANENhkMDrkHHus8IVZVnGPiBlw7bdCkr5tmTDkV1HbCnQDie9 EVs8gOusa/BRC/V9LevPqndC6GHqROj8KZirFBhevRNwbx37bPeb/HowAFDeclhtB2 c/wM0+Kq2q2mLM1KaPaDfFgV/0T2/Dfh7fJNw4vMyY2G8yEjbK5Ls3vaAHU2vXxTwD 543R/1QRBwvJqDtbR51SYJfv+OxqGLdfwfCWyxUoNyjOdBQTs2tTqDhI5P2PlY7c66 eUmb6LIJLUcBgxcnUrNwnq6vi1TTMseobclMNgQkA/UUn+21UpyceJPAig2OtEUCwx q7eGVpH+vvfHA==
Archived-At: http://mailarchive.ietf.org/arch/msg/clue/VSEeCbcJHXVypRL-uycRnU7wFxQ
Subject: Re: [clue] Improving treatment of audio
X-BeenThere: clue@ietf.org
X-Mailman-Version: 2.1.15
Precedence: list
List-Id: CLUE - ControLling mUltiple streams for TElepresence <clue.ietf.org>
List-Unsubscribe: <https://www.ietf.org/mailman/options/clue>, <mailto:clue-request@ietf.org?subject=unsubscribe>
List-Archive: <http://www.ietf.org/mail-archive/web/clue/>
List-Post: <mailto:clue@ietf.org>
List-Help: <mailto:clue-request@ietf.org?subject=help>
List-Subscribe: <https://www.ietf.org/mailman/listinfo/clue>, <mailto:clue-request@ietf.org?subject=subscribe>
X-List-Received-Date: Mon, 14 Apr 2014 18:16:44 -0000

On 4/14/14 2:11 PM, Paul Coverdale wrote:
>>>      There is, alas, an elephant in each room which can render our
>>> audio treatment useless -- it's called "echo". :^(
>>
>> IMO, if we don't have a solution that can prevent this then we don't
>> have a solution worth publishing.
>>
>
> Of course echo is important, but I wouldn't get too pessimistic about it.
> While it's always possible to screw it up, I suspect that any serious
> Telepresence vendor understands acoustic echo control quite well and will
> have taken steps to control it. I don't know that there's much else we can
> usefully do from a CLUE perspective.

I will happily accept the judgement of those who know about this, as 
long as they actively think it is ok. But its bad if there is risk and 
people haven't thought about it.

	Thanks,
	Paul


From nobody Mon Apr 14 12:20:28 2014
Return-Path: <coverdale@sympatico.ca>
X-Original-To: clue@ietfa.amsl.com
Delivered-To: clue@ietfa.amsl.com
Received: from localhost (ietfa.amsl.com [127.0.0.1]) by ietfa.amsl.com (Postfix) with ESMTP id C563E1A06E4 for <clue@ietfa.amsl.com>; Mon, 14 Apr 2014 12:20:21 -0700 (PDT)
X-Virus-Scanned: amavisd-new at amsl.com
X-Spam-Flag: NO
X-Spam-Score: 2.147
X-Spam-Level: **
X-Spam-Status: No, score=2.147 tagged_above=-999 required=5 tests=[BAYES_50=0.8, MSGID_FROM_MTA_HEADER=0.001, RCVD_IN_BL_SPAMCOP_NET=1.347, RCVD_IN_DNSWL_NONE=-0.0001, SPF_PASS=-0.001] autolearn=no
Received: from mail.ietf.org ([4.31.198.44]) by localhost (ietfa.amsl.com [127.0.0.1]) (amavisd-new, port 10024) with ESMTP id bJaa1qk5jaSw for <clue@ietfa.amsl.com>; Mon, 14 Apr 2014 12:20:21 -0700 (PDT)
Received: from blu0-omc1-s12.blu0.hotmail.com (blu0-omc1-s12.blu0.hotmail.com [65.55.116.23]) by ietfa.amsl.com (Postfix) with ESMTP id D43CE1A06BD for <clue@ietf.org>; Mon, 14 Apr 2014 12:20:20 -0700 (PDT)
Received: from BLU0-SMTP50 ([65.55.116.8]) by blu0-omc1-s12.blu0.hotmail.com with Microsoft SMTPSVC(6.0.3790.4675);  Mon, 14 Apr 2014 12:20:18 -0700
X-TMN: [2zSPt05vwyNawH9WysF4WJtVK42mbThC]
X-Originating-Email: [coverdale@sympatico.ca]
Message-ID: <BLU0-SMTP502386A531D7333075CF76D0510@phx.gbl>
Received: from PaulNewPC ([74.15.60.251]) by BLU0-SMTP50.phx.gbl over TLS secured channel with Microsoft SMTPSVC(6.0.3790.4675);  Mon, 14 Apr 2014 12:20:17 -0700
From: Paul Coverdale <coverdale@sympatico.ca>
To: "'Paul Kyzivat'" <pkyzivat@alum.mit.edu>, <clue@ietf.org>
References: <533AF351.9050201@alum.mit.edu> <BLU0-SMTP942A03E86972E7B8A212C6D0560@phx.gbl> <20140413233154.GL60844@verdi> <534C02BC.6000705@alum.mit.edu> <BLU0-SMTP26739EFCE85636668D142CD0510@phx.gbl> <534C2608.9040506@alum.mit.edu>
In-Reply-To: <534C2608.9040506@alum.mit.edu>
Date: Mon, 14 Apr 2014 15:20:15 -0400
MIME-Version: 1.0
Content-Type: text/plain; charset="us-ascii"
Content-Transfer-Encoding: 7bit
X-Mailer: Microsoft Office Outlook 12.0
Thread-Index: Ac9YDa5adurbKSKWTMiuZ7E0CIa/MQABbNaA
Content-Language: en-us
X-OriginalArrivalTime: 14 Apr 2014 19:20:17.0755 (UTC) FILETIME=[91E962B0:01CF5816]
Archived-At: http://mailarchive.ietf.org/arch/msg/clue/1IZ6XFA9h7-1HbADMuFhYH7CYlM
Subject: Re: [clue] Improving treatment of audio
X-BeenThere: clue@ietf.org
X-Mailman-Version: 2.1.15
Precedence: list
List-Id: CLUE - ControLling mUltiple streams for TElepresence <clue.ietf.org>
List-Unsubscribe: <https://www.ietf.org/mailman/options/clue>, <mailto:clue-request@ietf.org?subject=unsubscribe>
List-Archive: <http://www.ietf.org/mail-archive/web/clue/>
List-Post: <mailto:clue@ietf.org>
List-Help: <mailto:clue-request@ietf.org?subject=help>
List-Subscribe: <https://www.ietf.org/mailman/listinfo/clue>, <mailto:clue-request@ietf.org?subject=subscribe>
X-List-Received-Date: Mon, 14 Apr 2014 19:20:22 -0000

>>>> There is, alas, an elephant in each room which can render our
>>>> audio treatment useless -- it's called "echo". :^(
>>>
>>> IMO, if we don't have a solution that can prevent this then we don't
>>> have a solution worth publishing.
>>
>> Of course echo is important, but I wouldn't get too pessimistic about it.
>> While it's always possible to screw it up, I suspect that any serious
>> Telepresence vendor understands acoustic echo control quite well and
>> will have taken steps to control it. I don't know that there's much
>> else we can usefully do from a CLUE perspective.
>
>I will happily accept the judgement of those who know about this, as
>long as they actively think it is ok. But its bad if there is risk and
>people haven't thought about it.

[PVC]: Well, if we want to be really bullet-proof about echo we would need
to write detailed specifications and implementation guidelines for acoustic
echo cancellers. But it seems to me that this goes way beyond the current
CLUE charter.

...Paul


From nobody Mon Apr 14 12:52:54 2014
Return-Path: <pkyzivat@alum.mit.edu>
X-Original-To: clue@ietfa.amsl.com
Delivered-To: clue@ietfa.amsl.com
Received: from localhost (ietfa.amsl.com [127.0.0.1]) by ietfa.amsl.com (Postfix) with ESMTP id C99F81A0734 for <clue@ietfa.amsl.com>; Mon, 14 Apr 2014 12:52:52 -0700 (PDT)
X-Virus-Scanned: amavisd-new at amsl.com
X-Spam-Flag: NO
X-Spam-Score: 1.465
X-Spam-Level: *
X-Spam-Status: No, score=1.465 tagged_above=-999 required=5 tests=[BAYES_50=0.8, DKIM_SIGNED=0.1, DKIM_VALID=-0.1, SPF_SOFTFAIL=0.665] autolearn=no
Received: from mail.ietf.org ([4.31.198.44]) by localhost (ietfa.amsl.com [127.0.0.1]) (amavisd-new, port 10024) with ESMTP id COZt0PidtCuV for <clue@ietfa.amsl.com>; Mon, 14 Apr 2014 12:52:51 -0700 (PDT)
Received: from qmta01.westchester.pa.mail.comcast.net (qmta01.westchester.pa.mail.comcast.net [IPv6:2001:558:fe14:43:76:96:62:16]) by ietfa.amsl.com (Postfix) with ESMTP id 7EF391A072A for <clue@ietf.org>; Mon, 14 Apr 2014 12:52:51 -0700 (PDT)
Received: from omta14.westchester.pa.mail.comcast.net ([76.96.62.60]) by qmta01.westchester.pa.mail.comcast.net with comcast id puMe1n00E1HzFnQ51vsobn; Mon, 14 Apr 2014 19:52:48 +0000
Received: from Paul-Kyzivats-MacBook-Pro.local ([50.138.229.164]) by omta14.westchester.pa.mail.comcast.net with comcast id pvso1n00X3ZTu2S3avsoq6; Mon, 14 Apr 2014 19:52:48 +0000
Message-ID: <534C3C90.50106@alum.mit.edu>
Date: Mon, 14 Apr 2014 15:52:48 -0400
From: Paul Kyzivat <pkyzivat@alum.mit.edu>
User-Agent: Mozilla/5.0 (Macintosh; Intel Mac OS X 10.7; rv:24.0) Gecko/20100101 Thunderbird/24.4.0
MIME-Version: 1.0
To: Paul Coverdale <coverdale@sympatico.ca>, clue@ietf.org
References: <533AF351.9050201@alum.mit.edu> <BLU0-SMTP942A03E86972E7B8A212C6D0560@phx.gbl> <20140413233154.GL60844@verdi> <534C02BC.6000705@alum.mit.edu> <BLU0-SMTP26739EFCE85636668D142CD0510@phx.gbl> <534C2608.9040506@alum.mit.edu> <BLU0-SMTP502386A531D7333075CF76D0510@phx.gbl>
In-Reply-To: <BLU0-SMTP502386A531D7333075CF76D0510@phx.gbl>
Content-Type: text/plain; charset=ISO-8859-1; format=flowed
Content-Transfer-Encoding: 7bit
DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=comcast.net; s=q20140121; t=1397505168; bh=5Jmc6fg44WpjmHF+Rwl/8hYVEVTWAbESfF5LTlI6JKs=; h=Received:Received:Message-ID:Date:From:MIME-Version:To:Subject: Content-Type; b=DUM2uMs/gFOIpPjU/AejMkr8RxLAWGl7MAdbm2/Ja4mXmNfv3ux6O+RTuLSFyelmh cv9eDYDEoviIYw/LB4ucKQjQpLHFo8vzUhZ5qQHwEV+nNKmVCRRo9CSRTAI1hxRLLt E4JYgZEmc41vQD2YzAl9dRehwBoRFvs4RfN/Gt82FktHCxxmIlOTwjeUo5ehZS07Ls knO4gfsq+HR7EAt2fzlDfrPPtEuPoxmEoUJ0gffqUxw+Kj3aVKzUVeCFMxbsarq+c+ Tm9x9LJTQ/Ti4AR8mc5eABiTOOM7Dybt3NfKuMnQKwoy6TXVJI4baBEBCxyycA2JAe CgMlQ90t31WfA==
Archived-At: http://mailarchive.ietf.org/arch/msg/clue/QNn-t-wfDX3kFlPfN_m_l0vc3zw
Subject: Re: [clue] Improving treatment of audio
X-BeenThere: clue@ietf.org
X-Mailman-Version: 2.1.15
Precedence: list
List-Id: CLUE - ControLling mUltiple streams for TElepresence <clue.ietf.org>
List-Unsubscribe: <https://www.ietf.org/mailman/options/clue>, <mailto:clue-request@ietf.org?subject=unsubscribe>
List-Archive: <http://www.ietf.org/mail-archive/web/clue/>
List-Post: <mailto:clue@ietf.org>
List-Help: <mailto:clue-request@ietf.org?subject=help>
List-Subscribe: <https://www.ietf.org/mailman/listinfo/clue>, <mailto:clue-request@ietf.org?subject=subscribe>
X-List-Received-Date: Mon, 14 Apr 2014 19:52:53 -0000

On 4/14/14 3:20 PM, Paul Coverdale wrote:
>>>>> There is, alas, an elephant in each room which can render our
>>>>> audio treatment useless -- it's called "echo". :^(
>>>>
>>>> IMO, if we don't have a solution that can prevent this then we don't
>>>> have a solution worth publishing.
>>>
>>> Of course echo is important, but I wouldn't get too pessimistic about it.
>>> While it's always possible to screw it up, I suspect that any serious
>>> Telepresence vendor understands acoustic echo control quite well and
>>> will have taken steps to control it. I don't know that there's much
>>> else we can usefully do from a CLUE perspective.
>>
>> I will happily accept the judgement of those who know about this, as
>> long as they actively think it is ok. But its bad if there is risk and
>> people haven't thought about it.
>
> [PVC]: Well, if we want to be really bullet-proof about echo we would need
> to write detailed specifications and implementation guidelines for acoustic
> echo cancellers. But it seems to me that this goes way beyond the current
> CLUE charter.

I'll reiterate: we only need to ensure that there is enough information 
so that the endpoints can do this. We don't have to tell them how.

Based on John's comments, it sounds like we might need to specify that 
the advertiser is responsible for removing echo from the captures it 
supplies, rather than thinking that the receiver can do it. (But I don't 
really know if what I said makes any sense.)

	Thanks,
	Paul


From nobody Mon Apr 14 13:41:26 2014
Return-Path: <Mark.Duckworth@polycom.com>
X-Original-To: clue@ietfa.amsl.com
Delivered-To: clue@ietfa.amsl.com
Received: from localhost (ietfa.amsl.com [127.0.0.1]) by ietfa.amsl.com (Postfix) with ESMTP id 8AAE41A06F8 for <clue@ietfa.amsl.com>; Mon, 14 Apr 2014 13:41:24 -0700 (PDT)
X-Virus-Scanned: amavisd-new at amsl.com
X-Spam-Flag: NO
X-Spam-Score: -1.82
X-Spam-Level: 
X-Spam-Status: No, score=-1.82 tagged_above=-999 required=5 tests=[BAYES_00=-1.9, RCVD_IN_DNSWL_LOW=-0.7, SPF_NEUTRAL=0.779, UNPARSEABLE_RELAY=0.001] autolearn=ham
Received: from mail.ietf.org ([4.31.198.44]) by localhost (ietfa.amsl.com [127.0.0.1]) (amavisd-new, port 10024) with ESMTP id V3Z6PFqdiNSY for <clue@ietfa.amsl.com>; Mon, 14 Apr 2014 13:41:23 -0700 (PDT)
Received: from mail1.bemta7.messagelabs.com (mail1.bemta7.messagelabs.com [216.82.254.100]) by ietfa.amsl.com (Postfix) with ESMTP id 2484A1A021F for <clue@ietf.org>; Mon, 14 Apr 2014 13:41:23 -0700 (PDT)
Received: from [216.82.254.19:57543] by server-4.bemta-7.messagelabs.com id F3/B6-05677-0F74C435; Mon, 14 Apr 2014 20:41:20 +0000
X-Env-Sender: Mark.Duckworth@polycom.com
X-Msg-Ref: server-12.tower-96.messagelabs.com!1397508079!7071724!1
X-Originating-IP: [140.242.64.158]
X-StarScan-Received: 
X-StarScan-Version: 6.11.1; banners=-,-,-
X-VirusChecked: Checked
Received: (qmail 1069 invoked from network); 14 Apr 2014 20:41:20 -0000
Received: from crpehubprd01.polycom.com (HELO Crpehubprd01.polycom.com) (140.242.64.158) by server-12.tower-96.messagelabs.com with AES128-SHA encrypted SMTP; 14 Apr 2014 20:41:20 -0000
Received: from CRPMBOXPRD08.polycom.com ([169.254.1.94]) by Crpehubprd01.polycom.com ([fe80::5efe:10.236.0.158%14]) with mapi; Mon, 14 Apr 2014 13:41:12 -0700
From: "Duckworth, Mark" <Mark.Duckworth@polycom.com>
To: Paul Coverdale <coverdale@sympatico.ca>, 'Paul Kyzivat' <pkyzivat@alum.mit.edu>, "clue@ietf.org" <clue@ietf.org>
Date: Mon, 14 Apr 2014 13:41:12 -0700
Thread-Topic: [clue] Improving treatment of audio
Thread-Index: Ac9X+Kim/PDqrki0TBCBa1CEWCaUOwABfLDAAAit9RA=
Message-ID: <5C4AC54BFF7A0842A6A11F554D6FB52F071638@CRPMBOXPRD08.polycom.com>
References: <533AF351.9050201@alum.mit.edu> <BLU0-SMTP942A03E86972E7B8A212C6D0560@phx.gbl> <20140413233154.GL60844@verdi> <534C02BC.6000705@alum.mit.edu> <BLU0-SMTP26739EFCE85636668D142CD0510@phx.gbl>
In-Reply-To: <BLU0-SMTP26739EFCE85636668D142CD0510@phx.gbl>
Accept-Language: en-US
Content-Language: en-US
X-MS-Has-Attach: 
X-MS-TNEF-Correlator: 
acceptlanguage: en-US
Content-Type: text/plain; charset="us-ascii"
Content-Transfer-Encoding: quoted-printable
MIME-Version: 1.0
Archived-At: http://mailarchive.ietf.org/arch/msg/clue/DmHamtHfsu1FXNf6mck2l39dv5I
Subject: Re: [clue] Improving treatment of audio
X-BeenThere: clue@ietf.org
X-Mailman-Version: 2.1.15
Precedence: list
List-Id: CLUE - ControLling mUltiple streams for TElepresence <clue.ietf.org>
List-Unsubscribe: <https://www.ietf.org/mailman/options/clue>, <mailto:clue-request@ietf.org?subject=unsubscribe>
List-Archive: <http://www.ietf.org/mail-archive/web/clue/>
List-Post: <mailto:clue@ietf.org>
List-Help: <mailto:clue-request@ietf.org?subject=help>
List-Subscribe: <https://www.ietf.org/mailman/listinfo/clue>, <mailto:clue-request@ietf.org?subject=subscribe>
X-List-Received-Date: Mon, 14 Apr 2014 20:41:24 -0000

I agree echo cancellation is outside the scope of CLUE.  Echo cancellation =
for audio/video conferencing is nothing new related to CLUE.  CLUE doesn't =
make echo cancellation any easier or harder than conferencing without CLUE.

Mark

> -----Original Message-----
> From: clue [mailto:clue-bounces@ietf.org] On Behalf Of Paul Coverdale
> Sent: Monday, April 14, 2014 2:12 PM
> To: 'Paul Kyzivat'; clue@ietf.org
> Subject: Re: [clue] Improving treatment of audio
>=20
> >>     There is, alas, an elephant in each room which can render our
> >> audio treatment useless -- it's called "echo". :^(
> >
> >IMO, if we don't have a solution that can prevent this then we don't
> >have a solution worth publishing.
> >
>=20
> Of course echo is important, but I wouldn't get too pessimistic about it.
> While it's always possible to screw it up, I suspect that any serious
> Telepresence vendor understands acoustic echo control quite well and will
> have taken steps to control it. I don't know that there's much else we ca=
n
> usefully do from a CLUE perspective.
>=20
> ...Paul
>=20
> _______________________________________________
> clue mailing list
> clue@ietf.org
> https://www.ietf.org/mailman/listinfo/clue


From nobody Mon Apr 14 14:36:47 2014
Return-Path: <coverdale@sympatico.ca>
X-Original-To: clue@ietfa.amsl.com
Delivered-To: clue@ietfa.amsl.com
Received: from localhost (ietfa.amsl.com [127.0.0.1]) by ietfa.amsl.com (Postfix) with ESMTP id 86AEF1A075C for <clue@ietfa.amsl.com>; Mon, 14 Apr 2014 14:36:46 -0700 (PDT)
X-Virus-Scanned: amavisd-new at amsl.com
X-Spam-Flag: NO
X-Spam-Score: 0.8
X-Spam-Level: 
X-Spam-Status: No, score=0.8 tagged_above=-999 required=5 tests=[BAYES_50=0.8,  MSGID_FROM_MTA_HEADER=0.001, RCVD_IN_DNSWL_NONE=-0.0001, SPF_PASS=-0.001] autolearn=ham
Received: from mail.ietf.org ([4.31.198.44]) by localhost (ietfa.amsl.com [127.0.0.1]) (amavisd-new, port 10024) with ESMTP id frew77Vg4GTI for <clue@ietfa.amsl.com>; Mon, 14 Apr 2014 14:36:45 -0700 (PDT)
Received: from blu0-omc1-s3.blu0.hotmail.com (blu0-omc1-s3.blu0.hotmail.com [65.55.116.14]) by ietfa.amsl.com (Postfix) with ESMTP id 244AA1A0231 for <clue@ietf.org>; Mon, 14 Apr 2014 14:36:45 -0700 (PDT)
Received: from BLU0-SMTP34 ([65.55.116.7]) by blu0-omc1-s3.blu0.hotmail.com with Microsoft SMTPSVC(6.0.3790.4675);  Mon, 14 Apr 2014 14:36:42 -0700
X-TMN: [TXUyKgX59loqd3TeVY0n2zro+AwS8PCo]
X-Originating-Email: [coverdale@sympatico.ca]
Message-ID: <BLU0-SMTP345CA8C2F79C69A3190232D0510@phx.gbl>
Received: from PaulNewPC ([74.15.60.251]) by BLU0-SMTP34.phx.gbl over TLS secured channel with Microsoft SMTPSVC(6.0.3790.4675);  Mon, 14 Apr 2014 14:36:42 -0700
From: Paul Coverdale <coverdale@sympatico.ca>
To: "'Paul Kyzivat'" <pkyzivat@alum.mit.edu>, <clue@ietf.org>
References: <533AF351.9050201@alum.mit.edu> <BLU0-SMTP942A03E86972E7B8A212C6D0560@phx.gbl> <20140413233154.GL60844@verdi> <534C02BC.6000705@alum.mit.edu> <BLU0-SMTP26739EFCE85636668D142CD0510@phx.gbl> <534C2608.9040506@alum.mit.edu> <BLU0-SMTP502386A531D7333075CF76D0510@phx.gbl> <534C3C90.50106@alum.mit.edu>
In-Reply-To: <534C3C90.50106@alum.mit.edu>
Date: Mon, 14 Apr 2014 17:36:39 -0400
MIME-Version: 1.0
Content-Type: text/plain; charset="us-ascii"
Content-Transfer-Encoding: 7bit
X-Mailer: Microsoft Office Outlook 12.0
Thread-Index: Ac9YGxzx9Bt2cz3JSamZkxrK6Ri7qQAAeF7g
Content-Language: en-us
X-OriginalArrivalTime: 14 Apr 2014 21:36:42.0415 (UTC) FILETIME=[A05963F0:01CF5829]
Archived-At: http://mailarchive.ietf.org/arch/msg/clue/81_guW68TeE_k1lC58_kzhONjBM
Subject: Re: [clue] Improving treatment of audio
X-BeenThere: clue@ietf.org
X-Mailman-Version: 2.1.15
Precedence: list
List-Id: CLUE - ControLling mUltiple streams for TElepresence <clue.ietf.org>
List-Unsubscribe: <https://www.ietf.org/mailman/options/clue>, <mailto:clue-request@ietf.org?subject=unsubscribe>
List-Archive: <http://www.ietf.org/mail-archive/web/clue/>
List-Post: <mailto:clue@ietf.org>
List-Help: <mailto:clue-request@ietf.org?subject=help>
List-Subscribe: <https://www.ietf.org/mailman/listinfo/clue>, <mailto:clue-request@ietf.org?subject=subscribe>
X-List-Received-Date: Mon, 14 Apr 2014 21:36:46 -0000

>-----Original Message-----
>From: Paul Kyzivat [mailto:pkyzivat@alum.mit.edu]
>Sent: Monday, April 14, 2014 3:53 PM
>To: Paul Coverdale; clue@ietf.org
>Subject: Re: [clue] Improving treatment of audio
>
>>>>>> There is, alas, an elephant in each room which can render our
>>>>>> audio treatment useless -- it's called "echo". :^(
>>>>>
>>>>> IMO, if we don't have a solution that can prevent this then we
>>>>> don't have a solution worth publishing.
>>>>
>>>> Of course echo is important, but I wouldn't get too pessimistic
>about it.
>>>> While it's always possible to screw it up, I suspect that any
>>>> serious Telepresence vendor understands acoustic echo control quite
>>>> well and will have taken steps to control it. I don't know that
>>>> there's much else we can usefully do from a CLUE perspective.
>>>
>>> I will happily accept the judgement of those who know about this, as
>>> long as they actively think it is ok. But its bad if there is risk
>>> and people haven't thought about it.
>>
>> [PVC]: Well, if we want to be really bullet-proof about echo we would
>> need to write detailed specifications and implementation guidelines
>> for acoustic echo cancellers. But it seems to me that this goes way
>> beyond the current CLUE charter.
>
>I'll reiterate: we only need to ensure that there is enough information
>so that the endpoints can do this. We don't have to tell them how.

[PVC]: The endpoints should already know what to do. It's not clear to me
what CLUE information needs to be signaled between endpoints.
>
>Based on John's comments, it sounds like we might need to specify that
>the advertiser is responsible for removing echo from the captures it
>supplies, rather than thinking that the receiver can do it. (But I don't
>really know if what I said makes any sense.)

[PVC]: I'm not sure what this means. It is the far-end that is always
responsible for removing echo. The near-end can do nothing. Unless by echo
you mean reverberance (barrel effect). The best way to get rid of
reverberance is acoustically (microphone closer to talker, or more
directional microphone), but de-reverberation algorithms exist which can
work on the electrical signal. These algorithms could be applied at the
near-end or the far-end.


...Paul


>




From nobody Tue Apr 15 00:02:48 2014
Return-Path: <Christian.Groves@nteczone.com>
X-Original-To: clue@ietfa.amsl.com
Delivered-To: clue@ietfa.amsl.com
Received: from localhost (ietfa.amsl.com [127.0.0.1]) by ietfa.amsl.com (Postfix) with ESMTP id 27E561A0368 for <clue@ietfa.amsl.com>; Tue, 15 Apr 2014 00:02:44 -0700 (PDT)
X-Virus-Scanned: amavisd-new at amsl.com
X-Spam-Flag: NO
X-Spam-Score: 0.8
X-Spam-Level: 
X-Spam-Status: No, score=0.8 tagged_above=-999 required=5 tests=[BAYES_50=0.8] autolearn=ham
Received: from mail.ietf.org ([4.31.198.44]) by localhost (ietfa.amsl.com [127.0.0.1]) (amavisd-new, port 10024) with ESMTP id FpvzDQD3FmK3 for <clue@ietfa.amsl.com>; Tue, 15 Apr 2014 00:02:41 -0700 (PDT)
Received: from cserver5.myshophosting.com (cserver5.myshophosting.com [175.107.161.1]) by ietfa.amsl.com (Postfix) with ESMTP id C7AAA1A022D for <clue@ietf.org>; Tue, 15 Apr 2014 00:02:40 -0700 (PDT)
Received: from ppp118-209-199-135.lns20.mel6.internode.on.net ([118.209.199.135]:59497 helo=[127.0.0.1]) by cserver5.myshophosting.com with esmtpsa (TLSv1:DHE-RSA-AES128-SHA:128) (Exim 4.82) (envelope-from <Christian.Groves@nteczone.com>) id 1WZxO7-0004DN-1S for clue@ietf.org; Tue, 15 Apr 2014 17:02:35 +1000
Message-ID: <534CD98A.2090605@nteczone.com>
Date: Tue, 15 Apr 2014 17:02:34 +1000
From: Christian Groves <Christian.Groves@nteczone.com>
User-Agent: Mozilla/5.0 (Windows NT 6.1; WOW64; rv:24.0) Gecko/20100101 Thunderbird/24.4.0
MIME-Version: 1.0
To: "clue@ietf.org" <clue@ietf.org>
Content-Type: text/plain; charset=ISO-8859-1; format=flowed
Content-Transfer-Encoding: 7bit
X-AntiAbuse: This header was added to track abuse, please include it with any abuse report
X-AntiAbuse: Primary Hostname - cserver5.myshophosting.com
X-AntiAbuse: Original Domain - ietf.org
X-AntiAbuse: Originator/Caller UID/GID - [47 12] / [47 12]
X-AntiAbuse: Sender Address Domain - nteczone.com
X-Get-Message-Sender-Via: cserver5.myshophosting.com: authenticated_id: christian.groves@nteczone.com
X-Source: 
X-Source-Args: 
X-Source-Dir: 
Archived-At: http://mailarchive.ietf.org/arch/msg/clue/8riEaZVguSuqzUqiLhbl0b9z_08
Subject: [clue] Data model GCSE syntax
X-BeenThere: clue@ietf.org
X-Mailman-Version: 2.1.15
Precedence: list
List-Id: CLUE - ControLling mUltiple streams for TElepresence <clue.ietf.org>
List-Unsubscribe: <https://www.ietf.org/mailman/options/clue>, <mailto:clue-request@ietf.org?subject=unsubscribe>
List-Archive: <http://www.ietf.org/mail-archive/web/clue/>
List-Post: <mailto:clue@ietf.org>
List-Help: <mailto:clue-request@ietf.org?subject=help>
List-Subscribe: <https://www.ietf.org/mailman/listinfo/clue>, <mailto:clue-request@ietf.org?subject=subscribe>
X-List-Received-Date: Tue, 15 Apr 2014 07:02:44 -0000

Hello Roberta,

http://tools.ietf.org/html/draft-ietf-clue-data-model-schema-04#section-18
indicates that a media type is associated with GCSE.

Is this really needed?

If we look at the <simultaneousSet> definition 
http://tools.ietf.org/html/draft-ietf-clue-data-model-schema-04#section-17 
there's no media type here. The media type would be given by the 
underlying capture description.

It seems that the GCSE syntax is basically identical to the 
simultaneousSet apart from the fact that we decided that a GSE can only 
specify CSEs, e.g.

    <!-- GLOBAL CAPTURE ENTRY TYPE -->
    <xs:complexType name="globalCaptureEntryType">
     <xs:sequence>
       <xs:element name="sceneEntryIDREF" type="xs:IDREF"
       minOccurs="0" maxOccurs="unbounded"/>
     </xs:sequence>
    </xs:complexType>

Btw I also noticed a typo in the 1st paragraph of 18, "...of the same 
media time..." should be "... of the same media type...".

Regards, Christian


From nobody Tue Apr 15 06:58:53 2014
Return-Path: <john@jlc.net>
X-Original-To: clue@ietfa.amsl.com
Delivered-To: clue@ietfa.amsl.com
Received: from localhost (ietfa.amsl.com [127.0.0.1]) by ietfa.amsl.com (Postfix) with ESMTP id 5D88C1A0658 for <clue@ietfa.amsl.com>; Tue, 15 Apr 2014 06:58:50 -0700 (PDT)
X-Virus-Scanned: amavisd-new at amsl.com
X-Spam-Flag: NO
X-Spam-Score: -1.772
X-Spam-Level: 
X-Spam-Status: No, score=-1.772 tagged_above=-999 required=5 tests=[BAYES_50=0.8, RCVD_IN_DNSWL_MED=-2.3, RP_MATCHES_RCVD=-0.272] autolearn=ham
Received: from mail.ietf.org ([4.31.198.44]) by localhost (ietfa.amsl.com [127.0.0.1]) (amavisd-new, port 10024) with ESMTP id 7pjMvIe-E4tN for <clue@ietfa.amsl.com>; Tue, 15 Apr 2014 06:58:45 -0700 (PDT)
Received: from mailhost.jlc.net (mailhost.jlc.net [199.201.159.4]) by ietfa.amsl.com (Postfix) with ESMTP id 9C6021A0213 for <clue@ietf.org>; Tue, 15 Apr 2014 06:58:44 -0700 (PDT)
Received: by mailhost.jlc.net (Postfix, from userid 104) id 348A0C94C3; Tue, 15 Apr 2014 09:58:40 -0400 (EDT)
Date: Tue, 15 Apr 2014 09:58:40 -0400
From: John Leslie <john@jlc.net>
To: Paul Kyzivat <pkyzivat@alum.mit.edu>
Message-ID: <20140415135840.GA89388@verdi>
References: <533AF351.9050201@alum.mit.edu> <BLU0-SMTP942A03E86972E7B8A212C6D0560@phx.gbl> <20140413233154.GL60844@verdi> <534C02BC.6000705@alum.mit.edu>
Mime-Version: 1.0
Content-Type: text/plain; charset=us-ascii
Content-Disposition: inline
In-Reply-To: <534C02BC.6000705@alum.mit.edu>
User-Agent: Mutt/1.4.1i
Archived-At: http://mailarchive.ietf.org/arch/msg/clue/_BUg84dNvLMD6BE8EkbQ4J82n-U
Cc: clue@ietf.org
Subject: Re: [clue] Improving treatment of audio
X-BeenThere: clue@ietf.org
X-Mailman-Version: 2.1.15
Precedence: list
List-Id: CLUE - ControLling mUltiple streams for TElepresence <clue.ietf.org>
List-Unsubscribe: <https://www.ietf.org/mailman/options/clue>, <mailto:clue-request@ietf.org?subject=unsubscribe>
List-Archive: <http://www.ietf.org/mail-archive/web/clue/>
List-Post: <mailto:clue@ietf.org>
List-Help: <mailto:clue-request@ietf.org?subject=help>
List-Subscribe: <https://www.ietf.org/mailman/listinfo/clue>, <mailto:clue-request@ietf.org?subject=subscribe>
X-List-Received-Date: Tue, 15 Apr 2014 13:58:50 -0000

Paul Kyzivat <pkyzivat@alum.mit.edu> wrote:
> On 4/13/14 7:31 PM, John Leslie wrote:
>> 
>>] REQMT-2:...
>>
>> Worse, I'm not confident that all of this is even attainable. :^(
> 
> I don't know if they are unattainable. But IIUC they cannot be attained 
> with the information we currently have in the fw and data model.
> 
> If we want to give up on attaining some of them, then let's do so, and 
> explicitly note that we have. And then ensure that we have sufficient 
> information to attain the rest.

   ISTM we're not ready to give up on any of them, but I'm willing to
continue the conversation...

>> There is, alas, an elephant in each room which can render our
>> audio treatment useless -- it's called "echo". :^(
> 
> IMO, if we don't have a solution that can prevent this then we don't 
> have a solution worth publishing.

   I think this deserves a separate thread...

>>] REQMT-2: The solution MUST support a description of the spatial
>>]          arrangement of captured source audio sent in audio streams
>>]          which enables a satisfactory reproduction at the receiver
>>]          in a spatially correct manner.
>>
>>    Other than wondering what "spatially correct means, I view this as
>> attainable, since the responsibily must belong to the receiver, and
>> that only if the sender chooses to specify enough audio sources.
> 
> What do you mean by "attainable"? Can this be achieved using the data 
> model we have now? (I don't think so.)

   Perhaps REQMT-2 means different things to different folk... :^(

   (Unquestionably, we have different ideas of _how_ to "enable a
satisfactory reproduction" -- but this doesn't bother me: I expect
end-users to adapt to the reproduction they pay for.)

   (BTW, I don't think we need to establish a single meaning for
"spatially correct" -- I only meant to stipulate that I was trying to
avoid that issue.)

   IMHO, a "satisfactory reproduction" can be achieved given separate
audio feeds with a clear specification of the location of each (so long
as you don't have to deal with echo of audio you sent yourself).

   Others, no doubt, prefer a different paradigm where each room sends
a surround-sound representation of their own room. It _is_ possible
to process that into a _different_ surround-sound representation which
will be satisfactory. (Of course, if the received surround-sound has
echo of what _you_ sent, my mind starts to boggle about cancelling echo.)

> ISTM that the responsibility is *shared*:
> - the advertiser must provide sufficient sources, *and* sufficient
>   description of those sources to permit the receiver to do the right
>   thing.
> 
> - the receiver must utilize the available information in order to
>   reproduce the audio in a spatially correct manner, within the limits
>   of its available equipment.

   Yes.

> - *we* are responsible for providing sufficient expressiveness in
>   the advertisement and RTP so that the advertiser and receiver can
>   do the above. *That* is the part we need to work on now.

   Speaking for myself, I'll settle for multiple channels of audio
with point-of-capture for each mike.

   (I know I won't always get that -- but I don't blame the CLUE spec
for that.)

>> If a sender _actually_ sends "stereo" or "3.0" the receiver will have
>> to punt. I can only hope that our actual standard doesn't encourage
>> this.
> 
> Sounds like something we need to talk about. Where?

   I have no suggestions...

>> Inevitably, some yahoo is going to put one or more loudspeakers in
>> the room. And those will expand the "area of coverage" to include the
>> other rooms those loudspeakers attempt to render.
>>
>> With delay!!!
> 
> OK. But surely the point of capture isn't sufficient by itself. 
> *Something* more is required. Can you suggest something that is stated 
> better, that would be attainable and useful for our purposes?

   Distance from the speaker's mouth would help. If it's "close enough"
the receiver needn't worry -- when an individual want's to mute, that's
the sender's problem.

>> ... the most useful audio streams are linked to microphones "close" to
>> individual people, while video captures will _very_ often try to cover
>> more than one person -- thus video capures will very typically be 
>> associated with several audio captures.
> 
> Are you saying that within a scene the choice of audio captures is 
> unrelated to the choice of video captures?

   Mostly, yes.

> If so, then does it ever make sense to choose a subset of the advertised 
> audio? And if that is true, then what criteria make sense?

   Clearly, there will be cases where an individual in one room is
particularly interested in some limited area within another room. IMHO
this is _in_addition_to_ the general room audio, so I see no reason
to worry about it in our specification.

>> Thus I most sincerely hope that folks won't choose to switch their
>> audio based on who's speaking -- but I won't try to stop them from doing
>> this. (And I must admit that sometimes I _shouldn't_ blame them, because
>> they find themselves sabotaged by audio sources contaminated by echo.)
> 
> I assume you mean when the speakers are in the same room, right?

   What I meant was audio feeds they receive coming in with so much echo
that they find themselves abliged to exclude some of their own mikes
from what they send.

   This, to tell truth, is a somewhat weird corner case. They should be
able to exclude the _received_ echo from their outgoing feed by normal
echo-cancellation algorithms. If I had thought this through while
writing the above text, I would have erased it before sending.

> Isn't the situation different when we are switching among speakers in 
> different rooms?

   I don't understand the question...

--
John Leslie <john@jlc.net>


From nobody Tue Apr 15 09:55:02 2014
Return-Path: <pkyzivat@alum.mit.edu>
X-Original-To: clue@ietfa.amsl.com
Delivered-To: clue@ietfa.amsl.com
Received: from localhost (ietfa.amsl.com [127.0.0.1]) by ietfa.amsl.com (Postfix) with ESMTP id 1CFFB1A02E0 for <clue@ietfa.amsl.com>; Tue, 15 Apr 2014 09:54:56 -0700 (PDT)
X-Virus-Scanned: amavisd-new at amsl.com
X-Spam-Flag: NO
X-Spam-Score: 0.664
X-Spam-Level: 
X-Spam-Status: No, score=0.664 tagged_above=-999 required=5 tests=[BAYES_40=-0.001, DKIM_SIGNED=0.1, DKIM_VALID=-0.1, SPF_SOFTFAIL=0.665] autolearn=no
Received: from mail.ietf.org ([4.31.198.44]) by localhost (ietfa.amsl.com [127.0.0.1]) (amavisd-new, port 10024) with ESMTP id OcoPnF5jUrfU for <clue@ietfa.amsl.com>; Tue, 15 Apr 2014 09:54:54 -0700 (PDT)
Received: from qmta03.westchester.pa.mail.comcast.net (qmta03.westchester.pa.mail.comcast.net [IPv6:2001:558:fe14:43:76:96:62:32]) by ietfa.amsl.com (Postfix) with ESMTP id 2EEB91A0499 for <clue@ietf.org>; Tue, 15 Apr 2014 09:54:53 -0700 (PDT)
Received: from omta11.westchester.pa.mail.comcast.net ([76.96.62.36]) by qmta03.westchester.pa.mail.comcast.net with comcast id qD7y1n0030mv7h053Gurkz; Tue, 15 Apr 2014 16:54:51 +0000
Received: from Paul-Kyzivats-MacBook-Pro.local ([50.138.229.164]) by omta11.westchester.pa.mail.comcast.net with comcast id qGuq1n0173ZTu2S3XGuqzk; Tue, 15 Apr 2014 16:54:51 +0000
Message-ID: <534D645A.9010703@alum.mit.edu>
Date: Tue, 15 Apr 2014 12:54:50 -0400
From: Paul Kyzivat <pkyzivat@alum.mit.edu>
User-Agent: Mozilla/5.0 (Macintosh; Intel Mac OS X 10.7; rv:24.0) Gecko/20100101 Thunderbird/24.4.0
MIME-Version: 1.0
To: John Leslie <john@jlc.net>
References: <533AF351.9050201@alum.mit.edu> <BLU0-SMTP942A03E86972E7B8A212C6D0560@phx.gbl> <20140413233154.GL60844@verdi> <534C02BC.6000705@alum.mit.edu> <20140415135840.GA89388@verdi>
In-Reply-To: <20140415135840.GA89388@verdi>
Content-Type: text/plain; charset=ISO-8859-1; format=flowed
Content-Transfer-Encoding: 7bit
DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=comcast.net; s=q20140121; t=1397580891; bh=caxRLIo+f3hi4vP1/eoqrS8LPKMQExGkBv60a8xb1Nw=; h=Received:Received:Message-ID:Date:From:MIME-Version:To:Subject: Content-Type; b=pQBoJOJHRRaFRQTnjXTauuVlb2lD3lo0ba46S6qh29FXcRxF9tQD/jz1oQ7fKm7mX jFLVFTvtkTXRRSw2uROKfNm7k/38Ej+uxkt2+R/LaOhj8ICdszqQ67roztu1dvLkjX lD8u3s5QixFbU5GhQVU4xVRJVOIz5MS+iJCJBnkPuaYHq0lUZox5MxNkAd2zV1ZkOX wfZvedHOVN5trXtU7X8PYAcaZs7Gz9ZX4scnGOL11kPtb1IyBOpg85F4qS1dJDgvxU 2ZsmIF4dWGK+VBqFnT+omNKC65Lx6XcIF/hAunS34HLknDObq636i7ARQn97zIRG3o zierxQH4kAV/w==
Archived-At: http://mailarchive.ietf.org/arch/msg/clue/qYjI8RwNBoIgXgVAuS-hDYR470Y
Cc: clue@ietf.org
Subject: Re: [clue] Improving treatment of audio
X-BeenThere: clue@ietf.org
X-Mailman-Version: 2.1.15
Precedence: list
List-Id: CLUE - ControLling mUltiple streams for TElepresence <clue.ietf.org>
List-Unsubscribe: <https://www.ietf.org/mailman/options/clue>, <mailto:clue-request@ietf.org?subject=unsubscribe>
List-Archive: <http://www.ietf.org/mail-archive/web/clue/>
List-Post: <mailto:clue@ietf.org>
List-Help: <mailto:clue-request@ietf.org?subject=help>
List-Subscribe: <https://www.ietf.org/mailman/listinfo/clue>, <mailto:clue-request@ietf.org?subject=subscribe>
X-List-Received-Date: Tue, 15 Apr 2014 16:54:56 -0000

On 4/15/14 9:58 AM, John Leslie wrote:
> Paul Kyzivat <pkyzivat@alum.mit.edu> wrote:
>> On 4/13/14 7:31 PM, John Leslie wrote:
>>>
>>> ] REQMT-2:...
>>>
>>> Worse, I'm not confident that all of this is even attainable. :^(
>>
>> I don't know if they are unattainable. But IIUC they cannot be attained
>> with the information we currently have in the fw and data model.
>>
>> If we want to give up on attaining some of them, then let's do so, and
>> explicitly note that we have. And then ensure that we have sufficient
>> information to attain the rest.
>
>     ISTM we're not ready to give up on any of them, but I'm willing to
> continue the conversation...

We may have a conflict here between continuing the conversation and 
getting something done. We've already been working on this stuff for 
years. If we haven't figured it out by now, then maybe we need to simplify.

So please continue the conversation, maybe for as much as a few days. 
But not for a few more months!

>>> There is, alas, an elephant in each room which can render our
>>> audio treatment useless -- it's called "echo". :^(
>>
>> IMO, if we don't have a solution that can prevent this then we don't
>> have a solution worth publishing.
>
>     I think this deserves a separate thread...

Fine. I don't understand. I get the impression that there is something 
that must be said about echo. I leave it to you.

>>> ] REQMT-2: The solution MUST support a description of the spatial
>>> ]          arrangement of captured source audio sent in audio streams
>>> ]          which enables a satisfactory reproduction at the receiver
>>> ]          in a spatially correct manner.
>>>
>>>     Other than wondering what "spatially correct means, I view this as
>>> attainable, since the responsibily must belong to the receiver, and
>>> that only if the sender chooses to specify enough audio sources.
>>
>> What do you mean by "attainable"? Can this be achieved using the data
>> model we have now? (I don't think so.)
>
>     Perhaps REQMT-2 means different things to different folk... :^(
>
>     (Unquestionably, we have different ideas of _how_ to "enable a
> satisfactory reproduction" -- but this doesn't bother me: I expect
> end-users to adapt to the reproduction they pay for.)

IMO it is fine that some things are left to "quality of implementation", 
as long as those things are entirely within the scope of a single endpoint.

>     (BTW, I don't think we need to establish a single meaning for
> "spatially correct" -- I only meant to stipulate that I was trying to
> avoid that issue.)
>
>     IMHO, a "satisfactory reproduction" can be achieved given separate
> audio feeds with a clear specification of the location of each (so long
> as you don't have to deal with echo of audio you sent yourself).
>
>     Others, no doubt, prefer a different paradigm where each room sends
> a surround-sound representation of their own room. It _is_ possible
> to process that into a _different_ surround-sound representation which
> will be satisfactory. (Of course, if the received surround-sound has
> echo of what _you_ sent, my mind starts to boggle about cancelling echo.)

Your last two paragraphs above leave me thinking that there is something 
more that must be said in the framework. But I have no idea what it is.

>> ISTM that the responsibility is *shared*:
>> - the advertiser must provide sufficient sources, *and* sufficient
>>    description of those sources to permit the receiver to do the right
>>    thing.
>>
>> - the receiver must utilize the available information in order to
>>    reproduce the audio in a spatially correct manner, within the limits
>>    of its available equipment.
>
>     Yes.
>
>> - *we* are responsible for providing sufficient expressiveness in
>>    the advertisement and RTP so that the advertiser and receiver can
>>    do the above. *That* is the part we need to work on now.
>
>     Speaking for myself, I'll settle for multiple channels of audio
> with point-of-capture for each mike.
>
>     (I know I won't always get that -- but I don't blame the CLUE spec
> for that.)

Well, if there is some minimal set of information that will always be 
required to get a workable result, then we should make that mandatory.

E.g., suppose the advertisement has a scene with multiple audio captures 
and *no* spatial information about them. Can a receiver to anything 
useful with that? If not, then we should require more.

>>> If a sender _actually_ sends "stereo" or "3.0" the receiver will have
>>> to punt. I can only hope that our actual standard doesn't encourage
>>> this.
>>
>> Sounds like something we need to talk about. Where?
>
>     I have no suggestions...

Then maybe we want to declare it "out of scope".

>>> Inevitably, some yahoo is going to put one or more loudspeakers in
>>> the room. And those will expand the "area of coverage" to include the
>>> other rooms those loudspeakers attempt to render.
>>>
>>> With delay!!!
>>
>> OK. But surely the point of capture isn't sufficient by itself.
>> *Something* more is required. Can you suggest something that is stated
>> better, that would be attainable and useful for our purposes?
>
>     Distance from the speaker's mouth would help. If it's "close enough"
> the receiver needn't worry -- when an individual want's to mute, that's
> the sender's problem.

I want those of you who know something about this to decide what we 
ought to make provision for in the advertisement.

>>> ... the most useful audio streams are linked to microphones "close" to
>>> individual people, while video captures will _very_ often try to cover
>>> more than one person -- thus video capures will very typically be
>>> associated with several audio captures.
>>
>> Are you saying that within a scene the choice of audio captures is
>> unrelated to the choice of video captures?
>
>     Mostly, yes.

That would be a really good thing to say in the framework!

>> If so, then does it ever make sense to choose a subset of the advertised
>> audio? And if that is true, then what criteria make sense?
>
>     Clearly, there will be cases where an individual in one room is
> particularly interested in some limited area within another room. IMHO
> this is _in_addition_to_ the general room audio, so I see no reason
> to worry about it in our specification.

What about a case where the receiver is limited in the number of audio 
captures it can receive and/or process? Can the advertisement help the 
receiver decide which one(s) it should get?

(We do have the priority attribute if nothing else.)

>>> Thus I most sincerely hope that folks won't choose to switch their
>>> audio based on who's speaking -- but I won't try to stop them from doing
>>> this. (And I must admit that sometimes I _shouldn't_ blame them, because
>>> they find themselves sabotaged by audio sources contaminated by echo.)

Maybe I misunderstood you above. By "folks" did you mean providers or 
consumers? I was thinking you meant providers.

>> I assume you mean when the speakers are in the same room, right?
>
>     What I meant was audio feeds they receive coming in with so much echo
> that they find themselves abliged to exclude some of their own mikes
> from what they send.
>
>     This, to tell truth, is a somewhat weird corner case. They should be
> able to exclude the _received_ echo from their outgoing feed by normal
> echo-cancellation algorithms. If I had thought this through while
> writing the above text, I would have erased it before sending.

Again, I don't fully understand, and probably don't need to. But it 
sounds like something more is needed in the fw to discuss expectations.

>> Isn't the situation different when we are switching among speakers in
>> different rooms?
>
>     I don't understand the question...

I was thinking of an MCU advertising, among other things, switched audio 
and video captures, where the alternatives come from different rooms. 
Then the whole point is to switch the audio and video together based on 
who is speaking.

	Thanks,
	Paul


From nobody Wed Apr 16 11:09:08 2014
Return-Path: <Mark.Duckworth@polycom.com>
X-Original-To: clue@ietfa.amsl.com
Delivered-To: clue@ietfa.amsl.com
Received: from localhost (ietfa.amsl.com [127.0.0.1]) by ietfa.amsl.com (Postfix) with ESMTP id 96C491A0227 for <clue@ietfa.amsl.com>; Wed, 16 Apr 2014 11:09:05 -0700 (PDT)
X-Virus-Scanned: amavisd-new at amsl.com
X-Spam-Flag: NO
X-Spam-Score: 0.88
X-Spam-Level: 
X-Spam-Status: No, score=0.88 tagged_above=-999 required=5 tests=[BAYES_50=0.8, RCVD_IN_DNSWL_LOW=-0.7, SPF_NEUTRAL=0.779, UNPARSEABLE_RELAY=0.001] autolearn=ham
Received: from mail.ietf.org ([4.31.198.44]) by localhost (ietfa.amsl.com [127.0.0.1]) (amavisd-new, port 10024) with ESMTP id iYA-gKuOZPzZ for <clue@ietfa.amsl.com>; Wed, 16 Apr 2014 11:08:59 -0700 (PDT)
Received: from mail1.bemta7.messagelabs.com (mail1.bemta7.messagelabs.com [216.82.254.107]) by ietfa.amsl.com (Postfix) with ESMTP id 7DBDC1A02A4 for <clue@ietf.org>; Wed, 16 Apr 2014 11:08:59 -0700 (PDT)
Received: from [216.82.254.20:25763] by server-11.bemta-7.messagelabs.com id DD/5D-21432-737CE435; Wed, 16 Apr 2014 18:08:55 +0000
X-Env-Sender: Mark.Duckworth@polycom.com
X-Msg-Ref: server-14.tower-47.messagelabs.com!1397671728!7892549!14
X-Originating-IP: [140.242.64.158]
X-StarScan-Received: 
X-StarScan-Version: 6.11.1; banners=-,-,-
X-VirusChecked: Checked
Received: (qmail 11018 invoked from network); 16 Apr 2014 18:08:55 -0000
Received: from crpehubprd01.polycom.com (HELO crpehubprd02.polycom.com) (140.242.64.158) by server-14.tower-47.messagelabs.com with AES128-SHA encrypted SMTP; 16 Apr 2014 18:08:55 -0000
Received: from CRPMBOXPRD08.polycom.com ([169.254.1.94]) by crpehubprd02.polycom.com ([::1]) with mapi; Wed, 16 Apr 2014 11:08:40 -0700
From: "Duckworth, Mark" <Mark.Duckworth@polycom.com>
To: Paul Kyzivat <pkyzivat@alum.mit.edu>, "clue@ietf.org" <clue@ietf.org>
Date: Wed, 16 Apr 2014 11:08:38 -0700
Thread-Topic: Requirements - RE: [clue] Improving treatment of audio
Thread-Index: Ac9ZnuQMDguaLxNfSK6fy9F3TftPxQ==
Message-ID: <5C4AC54BFF7A0842A6A11F554D6FB52F08C3EA@CRPMBOXPRD08.polycom.com>
Accept-Language: en-US
Content-Language: en-US
X-MS-Has-Attach: 
X-MS-TNEF-Correlator: 
acceptlanguage: en-US
Content-Type: text/plain; charset="us-ascii"
Content-Transfer-Encoding: quoted-printable
MIME-Version: 1.0
Archived-At: http://mailarchive.ietf.org/arch/msg/clue/UCLssaV9b7h0vY9XZBOqWk1CPFs
Subject: [clue] Requirements - RE:  Improving treatment of audio
X-BeenThere: clue@ietf.org
X-Mailman-Version: 2.1.15
Precedence: list
List-Id: CLUE - ControLling mUltiple streams for TElepresence <clue.ietf.org>
List-Unsubscribe: <https://www.ietf.org/mailman/options/clue>, <mailto:clue-request@ietf.org?subject=unsubscribe>
List-Archive: <http://www.ietf.org/mail-archive/web/clue/>
List-Post: <mailto:clue@ietf.org>
List-Help: <mailto:clue-request@ietf.org?subject=help>
List-Subscribe: <https://www.ietf.org/mailman/listinfo/clue>, <mailto:clue-request@ietf.org?subject=subscribe>
X-List-Received-Date: Wed, 16 Apr 2014 18:09:06 -0000

I think we might not have a common understanding about what the requirement=
s mean, and what are the use cases regarding audio.  Maybe we should discus=
s this before getting too deep into the solution.

I included some notes below about my interpretation of the requirements.

Regards,
Mark

> -----Original Message-----
> From: clue [mailto:clue-bounces@ietf.org] On Behalf Of Paul Kyzivat
> Sent: Monday, April 14, 2014 11:46 AM
> To: clue@ietf.org
> Subject: Re: [clue] Improving treatment of audio
>=20
> On 4/13/14 7:31 PM, John Leslie wrote:
> > Paul Coverdale <coverdale@sympatico.ca> wrote:
> >>
> >> So, going back to Paul K's request for a concise statement of what
> >> needs to be done with respect to the treatment of audio in CLUE, I'm
> >> not sure what else is really required apart from satisfying REQMT-2
> >> and REQMT-3,
> >
> >     I have to agree with Paul-C here: clue-telepresence-requirements
> > is pretty much set in stone, now that it has entered the RFCed queue.
> > But I don't think it fits Paul-K's criteria of "concise:
> > ]
> > ] REQMT-2: The solution MUST support a description of the spatial
> > ]          arrangement of captured source audio sent in audio streams
> > ]          which enables a satisfactory reproduction at the receiver
> > ]          in a spatially correct manner.  This applies to each site
> > ]          in a point to point or a multipoint meeting and refers to
> > ]          the spatial ordering within a site, not the ordering of
> > ]          channels between sites.
> > ]
> > ]          Use case point to point symmetric, and all use cases,
> > ]          especially heterogeneous.
> > ]
> > ]  REQMT-2a: The solution MUST support a means of preserving
> > ]            the spatial order of audio in the captured
> > ]            scene.  For example, if John sounds as if he is
> > ]            at Susan's right in the captured audio, John
> > ]            voice is also placed at Susan's right in the
> > ]            rendered image.
> > ]
> > ]  REQMT-2b: The solution MUST support a means to identify
> > ]            the number and spatial arrangement of audio
> > ]            channels including monaural, stereophonic
> > ]            (2.0), and 3.0 (left, center, right) audio
> > ]            channels.
> > ]
> > ]  REQMT-2c: The solution MUST support a means to identify
> > ]            the point of capture of individual audio
> > ]            captures in three dimensions.
> > ]
> > ]  REQMT-2d: The solution MUST support a means to identify
> > ]            the area of coverage of individual audio
> > ]            captures in three dimensions.
> > ]
> > ] REQMT-3: The solution MUST enable individual audio streams to be
> > ]          associated with one or more video image captures, and
> > ]          individual video image captures to be associated with one
> > ]          or more audio captures, for the purpose of rendering
> > ]          proper position.
> > ]
> > ]          Use case is point to point symmetric, and all use cases.
> >
> >     Worse, I'm not confident that all of this is even attainable. :^(
>=20
> I don't know if they are unattainable. But IIUC they cannot be attained w=
ith
> the information we currently have in the fw and data model.
>=20
> If we want to give up on attaining some of them, then let's do so, and
> explicitly note that we have. And then ensure that we have sufficient
> information to attain the rest.
>=20
> >     Mea culpa, of course -- I should have said that in WGLC. :^( :^(
> >
> >> with the possible exception of including the directional
> >> characteristics of the microphone with the audio capture.
> >
> >     While I'd be happy to see that signaled, I'm really not pushing it
> > for our first spec -- merely hoping we'll get there someday.
> >
> >     There is, alas, an elephant in each room which can render our
> > audio treatment useless -- it's called "echo". :^(
>=20
> IMO, if we don't have a solution that can prevent this then we don't have=
 a
> solution worth publishing.
>=20
> > =3D=3D=3D=3D
> >     Despite being too late, I feel I ought to comment on these REQMTs.
> > Feel free to ignore what follows:
>=20
> As I noted above, I think it is important to confront these issues.
>=20
> > ] REQMT-2: The solution MUST support a description of the spatial
> > ]          arrangement of captured source audio sent in audio streams
> > ]          which enables a satisfactory reproduction at the receiver
> > ]          in a spatially correct manner.
> >
> >     Other than wondering what "spatially correct means, I view this as
> > attainable, since the responsibily must belong to the receiver, and
> > that only if the sender chooses to specify enough audio sources.
>=20
> What do you mean by "attainable"? Can this be achieved using the data
> model we have now? (I don't think so.)
>=20
> ISTM that the responsibility is *shared*:
> - the advertiser must provide sufficient sources, *and* sufficient
>    description of those sources to permit the receiver to do the right
>    thing.
>=20
> - the receiver must utilize the available information in order to
>    reproduce the audio in a spatially correct manner, within the limits
>    of its available equipment.
>=20
> - *we* are responsible for providing sufficient expressiveness in
>    the advertisement and RTP so that the advertiser and receiver can
>    do the above. *That* is the part we need to work on now.

[Duckworth, Mark] I think REQMT-2 means only that audio can be rendered spa=
tially consistent with video, within a scene.  To me this is very simple, l=
ike what a DVD player (with home theater receiver) does.  It renders audio =
and video, spatially consistent with each other (person on left of screen c=
an be heard coming from left side of room).  It can render different audio =
formats (e.g. stereo, surround) depending on the physical loudspeaker confi=
guration and the audio rendering capability.  Likewise in CLUE, the sender =
can send one audio capture (could be multiple channels) for a scene, and th=
e receiver can just render it in same manner as the DVD player would.  The =
receiver could choose different encoding of the audio capture depending on =
the receiver's capability (different encoding for mono, stereo, 3-channel, =
for example), and assuming the sender advertises those encodings. (or maybe=
 they are different alternative captures, rather than different encodings o=
f the same capture.  Need to think about this and how it's related to SDP m=
-line for audio).

[Duckworth, Mark] It gets a bit more complicated when the sender wants to s=
end multiple captures for the same scene, for example different captures th=
at represent different parts of the room.  This is how TIP works, with thre=
e separate captures for left, center, and right.  Now we have to identify s=
patially how those multiple audio captures relate to the video.  This is wh=
ere the capture area was intended to be used.  But people are confused abou=
t that, and have reasons why it won't work well.  Maybe this topic needs mo=
re discussion.  TIP is very simple and limited, in that it forces exactly 3=
 video and 3 audio with one to one correspondence with no flexibility.

> > ]  REQMT-2a: The solution MUST support a means of preserving
> > ]            the spatial order of audio in the captured
> > ]            scene.  For example, if John sounds as if he is
> > ]            at Susan's right in the captured audio, John
> > ]            voice is also placed at Susan's right in the
> > ]            rendered image.
> >
> >     Likewise, this responsibility must belong to the receiver.

[Duckworth, Mark] I think this is very simple, as explained above.  The bas=
ic scenario is the media consumer receives one audio capture for room audio=
 (for example stereo) and renders it.  The receiver doesn't have to do anyt=
hing special, other than know how to decode and render stereo audio.  This =
works today in typical videoconferencing, and it can work easily the same w=
hen using CLUE to spatially arrange multiple video displays for a single sc=
ene.

> > ]  REQMT-2b: The solution MUST support a means to identify
> > ]            the number and spatial arrangement of audio
> > ]            channels including monaural, stereophonic
> > ]            (2.0), and 3.0 (left, center, right) audio
> > ]            channels.
> >
> >     This, IMHO, isn't worth doing (although perhaps I don't understand
> > what REQMT-2b means).
> >
> >     The actual layout of loudspeakers should belong entirely to the
> > receiving room. Further, "stereo" doesn't actually define the
> > relationship of the two channels; and "3.0" suffers a similar problem.
> >
> >     If a sender _actually_ sends "stereo" or "3.0" the receiver will
> > have to punt. I can only hope that our actual standard doesn't encourag=
e
> this.
>=20
> Sounds like something we need to talk about. Where?

[Duckworth, Mark] Conferencing systems have been using stereo and 3-channel=
 audio for years, without CLUE.  I don't understand the concern.  Like I sa=
id above, I'm not sure any more about identifying different formats using t=
he "audio channel format" CLUE capture attribute, or by using information t=
hat would already be in the SDP audio m-lines anyway.

> > ]  REQMT-2c: The solution MUST support a means to identify
> > ]            the point of capture of individual audio
> > ]            captures in three dimensions.
> >
> >     Well stated. Adding point on line of capture helps slightly, but
> > only if we have a useful approximation of sensitivity pattern (which I
> > don't believe we can expect to get).

[Duckworth, Mark] We have this now in CLUE as a media capture attribute, bu=
t I don't understand how it can be used by the audio receiver as an enhance=
ment to the rendering process.  I don't see a use case for it.

> > ]  REQMT-2d: The solution MUST support a means to identify
> > ]            the area of coverage of individual audio
> > ]            captures in three dimensions.
> >
> >     Poorly stated. :^( _Every_ microphone in a room can cover _every_
> > source of sound in a room. Thus, the "area of coverage" must
> > necessarily be the entire room...

[Duckworth, Mark] I think this requirement was for the use case where you n=
eed multiple audio captures for covering an entire scene, like the TIP case=
.  The "area of capture" attribute was supposed to provide this information=
, as a generalization of the limited TIP example of just left, center, righ=
t.

> >     (It gets worse... much worse!)
> >
> >     Inevitably, some yahoo is going to put one or more loudspeakers in
> > the room. And those will expand the "area of coverage" to include the
> > other rooms those loudspeakers attempt to render.
> >
> >     With delay!!!

[Duckworth, Mark] I think echo cancellation is already well understood in t=
he industry and outside the scope of CLUE.

> OK. But surely the point of capture isn't sufficient by itself.
> *Something* more is required. Can you suggest something that is stated
> better, that would be attainable and useful for our purposes?
>=20
> > ] REQMT-3: The solution MUST enable individual audio streams to be
> > ]          associated with one or more video image captures, and
> > ]          individual video image captures to be associated with one
> > ]          or more audio captures, for the purpose of rendering
> > ]          proper position.
> > ]
> > ]          Use case is point to point symmetric, and all use cases.
> >
> >     I accept there are many folks who will insist on doing this. I
> > don't expect to stop them.
> >
> >     But audio streams are intrinsically linked to rooms, not video capt=
ures.
> > And the most useful audio streams are linked to microphones "close" to
> > individual people, while video captures will _very_ often try to cover
> > more than one person -- thus video capures will very typically be
> > associated with several audio captures.

[Duckworth, Mark] Since REQMT-3 is about "rendering proper position", I thi=
nk it doesn't really add any additional requirement that isn't already in R=
EQMT-2.  Maybe REQMT-3 is starting to assume something about the solution, =
and putting lower level requirement on it?
=20
> Are you saying that within a scene the choice of audio captures is unrela=
ted
> to the choice of video captures? (Other than if you configure
> *some* video then you probably also want to configure *some* audio.)
>=20
> If so, then does it ever make sense to choose a subset of the advertised
> audio? And if that is true, then what criteria make sense?
>=20
> > =3D=3D=3D=3D
> >
> >     It will help if more of us pay attention to the relationship of
> > audio in television in movies. I find that they _don't_ track each
> > other too closely.
> >
> >     The audio provides the continuity; the video tantalizes the senses.
> > When the camera switches between two people talking to each other, the
> > audio doesn't switch -- it tries to make you believe you're in the room=
.
> >
> >     Thus I most sincerely hope that folks won't choose to switch their
> > audio based on who's speaking -- but I won't try to stop them from
> > doing this. (And I must admit that sometimes I _shouldn't_ blame them,
> > because they find themselves sabotaged by audio sources contaminated
> > by echo.)
>=20
> I assume you mean when the speakers are in the same room, right?
>=20
> Isn't the situation different when we are switching among speakers in
> different rooms?
>=20
> > =3D=3D=3D=3D
> >
> >     A word on "echo suppression" might help here: this is a
> > well-solved in plain-old-telephony where end-to-end delay can be held
> > constand. We won't be able to do that over the Internet.
> >
> >     An individual room is probably able to cancel the echo in what it
> > _sends_ but a receiving room won't be able to cancel echo in what it
> > receives. :^(
> >
> > =3D=3D=3D=3D
> >
> >     That's as much as I'm willing to sqeeze into this email.
> >
> >     Hope _something_ in this helps...
>=20
> Some! This is going to take more work.
>=20
> 	Thanks,
> 	Paul
>=20
> > --
> > John Leslie <john@jlc.net>


From nobody Wed Apr 16 16:00:18 2014
Return-Path: <coverdale@sympatico.ca>
X-Original-To: clue@ietfa.amsl.com
Delivered-To: clue@ietfa.amsl.com
Received: from localhost (ietfa.amsl.com [127.0.0.1]) by ietfa.amsl.com (Postfix) with ESMTP id BFF441A038D for <clue@ietfa.amsl.com>; Wed, 16 Apr 2014 16:00:12 -0700 (PDT)
X-Virus-Scanned: amavisd-new at amsl.com
X-Spam-Flag: NO
X-Spam-Score: -1.9
X-Spam-Level: 
X-Spam-Status: No, score=-1.9 tagged_above=-999 required=5 tests=[BAYES_00=-1.9, MSGID_FROM_MTA_HEADER=0.001, RCVD_IN_DNSWL_NONE=-0.0001, SPF_PASS=-0.001] autolearn=ham
Received: from mail.ietf.org ([4.31.198.44]) by localhost (ietfa.amsl.com [127.0.0.1]) (amavisd-new, port 10024) with ESMTP id JTNknLSWTCma for <clue@ietfa.amsl.com>; Wed, 16 Apr 2014 16:00:06 -0700 (PDT)
Received: from blu0-omc1-s22.blu0.hotmail.com (blu0-omc1-s22.blu0.hotmail.com [65.55.116.33]) by ietfa.amsl.com (Postfix) with ESMTP id 617591A0387 for <clue@ietf.org>; Wed, 16 Apr 2014 16:00:06 -0700 (PDT)
Received: from BLU0-SMTP38 ([65.55.116.8]) by blu0-omc1-s22.blu0.hotmail.com with Microsoft SMTPSVC(6.0.3790.4675);  Wed, 16 Apr 2014 16:00:02 -0700
X-TMN: [fUTxlLrubl5fja32QnpvuMsXwndSf9uk]
X-Originating-Email: [coverdale@sympatico.ca]
Message-ID: <BLU0-SMTP3873D4CD2388D7E5E49059D0530@phx.gbl>
Received: from PaulNewPC ([74.15.60.251]) by BLU0-SMTP38.phx.gbl over TLS secured channel with Microsoft SMTPSVC(6.0.3790.4675);  Wed, 16 Apr 2014 16:00:02 -0700
From: Paul Coverdale <coverdale@sympatico.ca>
To: "'Duckworth, Mark'" <Mark.Duckworth@polycom.com>, <clue@ietf.org>
References: <5C4AC54BFF7A0842A6A11F554D6FB52F08C3EA@CRPMBOXPRD08.polycom.com>
In-Reply-To: <5C4AC54BFF7A0842A6A11F554D6FB52F08C3EA@CRPMBOXPRD08.polycom.com>
Date: Wed, 16 Apr 2014 18:59:58 -0400
MIME-Version: 1.0
Content-Type: text/plain; charset="us-ascii"
Content-Transfer-Encoding: 7bit
X-Mailer: Microsoft Office Outlook 12.0
Thread-Index: Ac9ZnuQMDguaLxNfSK6fy9F3TftPxQAA32aA
Content-Language: en-us
X-OriginalArrivalTime: 16 Apr 2014 23:00:02.0471 (UTC) FILETIME=[9970E370:01CF59C7]
Archived-At: http://mailarchive.ietf.org/arch/msg/clue/Aco9EM8urskL3UVDCAo70_dLo7E
Subject: Re: [clue] Requirements - RE:  Improving treatment of audio
X-BeenThere: clue@ietf.org
X-Mailman-Version: 2.1.15
Precedence: list
List-Id: CLUE - ControLling mUltiple streams for TElepresence <clue.ietf.org>
List-Unsubscribe: <https://www.ietf.org/mailman/options/clue>, <mailto:clue-request@ietf.org?subject=unsubscribe>
List-Archive: <http://www.ietf.org/mail-archive/web/clue/>
List-Post: <mailto:clue@ietf.org>
List-Help: <mailto:clue-request@ietf.org?subject=help>
List-Subscribe: <https://www.ietf.org/mailman/listinfo/clue>, <mailto:clue-request@ietf.org?subject=subscribe>
X-List-Received-Date: Wed, 16 Apr 2014 23:00:14 -0000

I generally agree with Mark's points. See my comments inline.

..Paul

>-----Original Message-----
>From: clue [mailto:clue-bounces@ietf.org] On Behalf Of Duckworth, Mark
>Sent: Wednesday, April 16, 2014 2:09 PM
>To: Paul Kyzivat; clue@ietf.org
>Subject: [clue] Requirements - RE: Improving treatment of audio
>
>I think we might not have a common understanding about what the
>requirements mean, and what are the use cases regarding audio.  Maybe we
>should discuss this before getting too deep into the solution.
>
>I included some notes below about my interpretation of the requirements.
>
>Regards,
>Mark
>
>> -----Original Message-----
>> From: clue [mailto:clue-bounces@ietf.org] On Behalf Of Paul Kyzivat
>> Sent: Monday, April 14, 2014 11:46 AM
>> To: clue@ietf.org
>> Subject: Re: [clue] Improving treatment of audio
>>
>> On 4/13/14 7:31 PM, John Leslie wrote:
>> > Paul Coverdale <coverdale@sympatico.ca> wrote:
>> >>
>> >> So, going back to Paul K's request for a concise statement of what
>> >> needs to be done with respect to the treatment of audio in CLUE,
>> >> I'm not sure what else is really required apart from satisfying
>> >> REQMT-2 and REQMT-3,
>> >
>> >     I have to agree with Paul-C here: clue-telepresence-requirements
>> > is pretty much set in stone, now that it has entered the RFCed
>queue.
>> > But I don't think it fits Paul-K's criteria of "concise:
>> > ]
>> > ] REQMT-2: The solution MUST support a description of the spatial
>> > ]          arrangement of captured source audio sent in audio
>streams
>> > ]          which enables a satisfactory reproduction at the receiver
>> > ]          in a spatially correct manner.  This applies to each site
>> > ]          in a point to point or a multipoint meeting and refers to
>> > ]          the spatial ordering within a site, not the ordering of
>> > ]          channels between sites.
>> > ]
>> > ]          Use case point to point symmetric, and all use cases,
>> > ]          especially heterogeneous.
>> > ]
>> > ]  REQMT-2a: The solution MUST support a means of preserving
>> > ]            the spatial order of audio in the captured
>> > ]            scene.  For example, if John sounds as if he is
>> > ]            at Susan's right in the captured audio, John
>> > ]            voice is also placed at Susan's right in the
>> > ]            rendered image.
>> > ]
>> > ]  REQMT-2b: The solution MUST support a means to identify
>> > ]            the number and spatial arrangement of audio
>> > ]            channels including monaural, stereophonic
>> > ]            (2.0), and 3.0 (left, center, right) audio
>> > ]            channels.
>> > ]
>> > ]  REQMT-2c: The solution MUST support a means to identify
>> > ]            the point of capture of individual audio
>> > ]            captures in three dimensions.
>> > ]
>> > ]  REQMT-2d: The solution MUST support a means to identify
>> > ]            the area of coverage of individual audio
>> > ]            captures in three dimensions.
>> > ]
>> > ] REQMT-3: The solution MUST enable individual audio streams to be
>> > ]          associated with one or more video image captures, and
>> > ]          individual video image captures to be associated with one
>> > ]          or more audio captures, for the purpose of rendering
>> > ]          proper position.
>> > ]
>> > ]          Use case is point to point symmetric, and all use cases.
>> >
>> >     Worse, I'm not confident that all of this is even attainable.
>> > :^(
>>
>> I don't know if they are unattainable. But IIUC they cannot be
>> attained with the information we currently have in the fw and data
>model.
>>
>> If we want to give up on attaining some of them, then let's do so, and
>> explicitly note that we have. And then ensure that we have sufficient
>> information to attain the rest.
>>
>> >     Mea culpa, of course -- I should have said that in WGLC. :^( :^(
>> >
>> >> with the possible exception of including the directional
>> >> characteristics of the microphone with the audio capture.
>> >
>> >     While I'd be happy to see that signaled, I'm really not pushing
>> > it for our first spec -- merely hoping we'll get there someday.
>> >
>> >     There is, alas, an elephant in each room which can render our
>> > audio treatment useless -- it's called "echo". :^(
>>
>> IMO, if we don't have a solution that can prevent this then we don't
>> have a solution worth publishing.
>>
>> > ====
>> >     Despite being too late, I feel I ought to comment on these
>REQMTs.
>> > Feel free to ignore what follows:
>>
>> As I noted above, I think it is important to confront these issues.
>>
>> > ] REQMT-2: The solution MUST support a description of the spatial
>> > ]          arrangement of captured source audio sent in audio
>streams
>> > ]          which enables a satisfactory reproduction at the receiver
>> > ]          in a spatially correct manner.
>> >
>> >     Other than wondering what "spatially correct means, I view this
>> > as attainable, since the responsibily must belong to the receiver,
>> > and that only if the sender chooses to specify enough audio sources.
>>
>> What do you mean by "attainable"? Can this be achieved using the data
>> model we have now? (I don't think so.)
>>
>> ISTM that the responsibility is *shared*:
>> - the advertiser must provide sufficient sources, *and* sufficient
>>    description of those sources to permit the receiver to do the right
>>    thing.
>>
>> - the receiver must utilize the available information in order to
>>    reproduce the audio in a spatially correct manner, within the
>limits
>>    of its available equipment.
>>
>> - *we* are responsible for providing sufficient expressiveness in
>>    the advertisement and RTP so that the advertiser and receiver can
>>    do the above. *That* is the part we need to work on now.
>
>[Duckworth, Mark] I think REQMT-2 means only that audio can be rendered
>spatially consistent with video, within a scene.  To me this is very
>simple, like what a DVD player (with home theater receiver) does.  It
>renders audio and video, spatially consistent with each other (person on
>left of screen can be heard coming from left side of room).  It can
>render different audio formats (e.g. stereo, surround) depending on the
>physical loudspeaker configuration and the audio rendering capability.
>Likewise in CLUE, the sender can send one audio capture (could be
>multiple channels) for a scene, and the receiver can just render it in
>same manner as the DVD player would.  The receiver could choose
>different encoding of the audio capture depending on the receiver's
>capability (different encoding for mono, stereo, 3-channel, for
>example), and assuming the sender advertises those encodings. (or maybe
>they are different alternative captures, rather than different encodings
>of the same capture.  Ne  ed to think about this and how it's related to
>SDP m-line for audio).

[PVC] This is my interpretation of REQMT-2, also.


>[Duckworth, Mark] It gets a bit more complicated when the sender wants
>to send multiple captures for the same scene, for example different
>captures that represent different parts of the room.  This is how TIP
>works, with three separate captures for left, center, and right.  Now we
>have to identify spatially how those multiple audio captures relate to
>the video.  This is where the capture area was intended to be used.  But
>people are confused about that, and have reasons why it won't work well.
>Maybe this topic needs more discussion.  TIP is very simple and limited,
>in that it forces exactly 3 video and 3 audio with one to one
>correspondence with no flexibility.

[PVC] I agree that the TIP example is very simple, but for the more general
case of multiple audio and video captures, why can't the Advertisement from
the Provider provide information on the association between them?


>> > ]  REQMT-2a: The solution MUST support a means of preserving
>> > ]            the spatial order of audio in the captured
>> > ]            scene.  For example, if John sounds as if he is
>> > ]            at Susan's right in the captured audio, John
>> > ]            voice is also placed at Susan's right in the
>> > ]            rendered image.
>> >
>> >     Likewise, this responsibility must belong to the receiver.
>
>[Duckworth, Mark] I think this is very simple, as explained above.  The
>basic scenario is the media consumer receives one audio capture for room
>audio (for example stereo) and renders it.  The receiver doesn't have to
>do anything special, other than know how to decode and render stereo
>audio.  This works today in typical videoconferencing, and it can work
>easily the same when using CLUE to spatially arrange multiple video
>displays for a single scene.

>
>> > ]  REQMT-2b: The solution MUST support a means to identify
>> > ]            the number and spatial arrangement of audio
>> > ]            channels including monaural, stereophonic
>> > ]            (2.0), and 3.0 (left, center, right) audio
>> > ]            channels.
>> >
>> >     This, IMHO, isn't worth doing (although perhaps I don't
>> > understand what REQMT-2b means).
>> >
>> >     The actual layout of loudspeakers should belong entirely to the
>> > receiving room. Further, "stereo" doesn't actually define the
>> > relationship of the two channels; and "3.0" suffers a similar
>problem.
>> >
>> >     If a sender _actually_ sends "stereo" or "3.0" the receiver will
>> > have to punt. I can only hope that our actual standard doesn't
>> > encourage
>> this.
>>
>> Sounds like something we need to talk about. Where?
>
>[Duckworth, Mark] Conferencing systems have been using stereo and 3-
>channel audio for years, without CLUE.  I don't understand the concern.
>Like I said above, I'm not sure any more about identifying different
>formats using the "audio channel format" CLUE capture attribute, or by
>using information that would already be in the SDP audio m-lines anyway.

>
>> > ]  REQMT-2c: The solution MUST support a means to identify
>> > ]            the point of capture of individual audio
>> > ]            captures in three dimensions.
>> >
>> >     Well stated. Adding point on line of capture helps slightly, but
>> > only if we have a useful approximation of sensitivity pattern (which
>> > I don't believe we can expect to get).
>
>[Duckworth, Mark] We have this now in CLUE as a media capture attribute,
>but I don't understand how it can be used by the audio receiver as an
>enhancement to the rendering process.  I don't see a use case for it.

[PVC] This may not necessary for satisfying current use cases, but knowing
the spatial relationships of all microphones at the sending end (and their
sensitivities and directional patterns), could allow the receiving end to do
some fancy spatial audio processing. 


>> > ]  REQMT-2d: The solution MUST support a means to identify
>> > ]            the area of coverage of individual audio
>> > ]            captures in three dimensions.
>> >
>> >     Poorly stated. :^( _Every_ microphone in a room can cover
>> > _every_ source of sound in a room. Thus, the "area of coverage" must
>> > necessarily be the entire room...
>
>[Duckworth, Mark] I think this requirement was for the use case where
>you need multiple audio captures for covering an entire scene, like the
>TIP case.  The "area of capture" attribute was supposed to provide this
>information, as a generalization of the limited TIP example of just
>left, center, right.

[PVC] As a practical matter, this requirement which doesn't much sense.


>> >     (It gets worse... much worse!)
>> >
>> >     Inevitably, some yahoo is going to put one or more loudspeakers
>> > in the room. And those will expand the "area of coverage" to include
>> > the other rooms those loudspeakers attempt to render.
>> >
>> >     With delay!!!
>
>[Duckworth, Mark] I think echo cancellation is already well understood
>in the industry and outside the scope of CLUE.

[PVC] I agree!


>> OK. But surely the point of capture isn't sufficient by itself.
>> *Something* more is required. Can you suggest something that is stated
>> better, that would be attainable and useful for our purposes?
>>
>> > ] REQMT-3: The solution MUST enable individual audio streams to be
>> > ]          associated with one or more video image captures, and
>> > ]          individual video image captures to be associated with one
>> > ]          or more audio captures, for the purpose of rendering
>> > ]          proper position.
>> > ]
>> > ]          Use case is point to point symmetric, and all use cases.
>> >
>> >     I accept there are many folks who will insist on doing this. I
>> > don't expect to stop them.
>> >
>> >     But audio streams are intrinsically linked to rooms, not video
>captures.
>> > And the most useful audio streams are linked to microphones "close"
>> > to individual people, while video captures will _very_ often try to
>> > cover more than one person -- thus video capures will very typically
>> > be associated with several audio captures.
>
>[Duckworth, Mark] Since REQMT-3 is about "rendering proper position", I
>think it doesn't really add any additional requirement that isn't
>already in REQMT-2.  Maybe REQMT-3 is starting to assume something about
>the solution, and putting lower level requirement on it?

[PVC] Yes, REQMT-3 just seems to be re-stating REQT-2.


>> Are you saying that within a scene the choice of audio captures is
>> unrelated to the choice of video captures? (Other than if you
>> configure
>> *some* video then you probably also want to configure *some* audio.)
>>
>> If so, then does it ever make sense to choose a subset of the
>> advertised audio? And if that is true, then what criteria make sense?
>>
>> > ====
>> >
>> >     It will help if more of us pay attention to the relationship of
>> > audio in television in movies. I find that they _don't_ track each
>> > other too closely.
>> >
>> >     The audio provides the continuity; the video tantalizes the
>senses.
>> > When the camera switches between two people talking to each other,
>> > the audio doesn't switch -- it tries to make you believe you're in
>the room.
>> >
>> >     Thus I most sincerely hope that folks won't choose to switch
>> > their audio based on who's speaking -- but I won't try to stop them
>> > from doing this. (And I must admit that sometimes I _shouldn't_
>> > blame them, because they find themselves sabotaged by audio sources
>> > contaminated by echo.)
>>
>> I assume you mean when the speakers are in the same room, right?
>>
>> Isn't the situation different when we are switching among speakers in
>> different rooms?
>>
>> > ====
>> >
>> >     A word on "echo suppression" might help here: this is a
>> > well-solved in plain-old-telephony where end-to-end delay can be
>> > held constand. We won't be able to do that over the Internet.
>> >
>> >     An individual room is probably able to cancel the echo in what
>> > it _sends_ but a receiving room won't be able to cancel echo in what
>> > it receives. :^(
>> >
>> > ====
>> >
>> >     That's as much as I'm willing to sqeeze into this email.
>> >
>> >     Hope _something_ in this helps...
>>
>> Some! This is going to take more work.
>>
>> 	Thanks,
>> 	Paul
>>
>> > --
>> > John Leslie <john@jlc.net>
>
>_______________________________________________
>clue mailing list
>clue@ietf.org
>https://www.ietf.org/mailman/listinfo/clue


From nobody Wed Apr 16 17:23:58 2014
Return-Path: <Mark.Duckworth@polycom.com>
X-Original-To: clue@ietfa.amsl.com
Delivered-To: clue@ietfa.amsl.com
Received: from localhost (ietfa.amsl.com [127.0.0.1]) by ietfa.amsl.com (Postfix) with ESMTP id 3FA611A0059 for <clue@ietfa.amsl.com>; Wed, 16 Apr 2014 17:23:56 -0700 (PDT)
X-Virus-Scanned: amavisd-new at amsl.com
X-Spam-Flag: NO
X-Spam-Score: -1.82
X-Spam-Level: 
X-Spam-Status: No, score=-1.82 tagged_above=-999 required=5 tests=[BAYES_00=-1.9, RCVD_IN_DNSWL_LOW=-0.7, SPF_NEUTRAL=0.779, UNPARSEABLE_RELAY=0.001] autolearn=ham
Received: from mail.ietf.org ([4.31.198.44]) by localhost (ietfa.amsl.com [127.0.0.1]) (amavisd-new, port 10024) with ESMTP id 6Ywubf93HG36 for <clue@ietfa.amsl.com>; Wed, 16 Apr 2014 17:23:53 -0700 (PDT)
Received: from mail1.bemta7.messagelabs.com (mail1.bemta7.messagelabs.com [216.82.254.98]) by ietfa.amsl.com (Postfix) with ESMTP id 41FBF1A0063 for <clue@ietf.org>; Wed, 16 Apr 2014 17:23:50 -0700 (PDT)
Received: from [216.82.254.19:3041] by server-2.bemta-7.messagelabs.com id 19/7B-13907-21F1F435; Thu, 17 Apr 2014 00:23:46 +0000
X-Env-Sender: Mark.Duckworth@polycom.com
X-Msg-Ref: server-15.tower-96.messagelabs.com!1397694224!7667303!5
X-Originating-IP: [140.242.64.158]
X-StarScan-Received: 
X-StarScan-Version: 6.11.1; banners=-,-,-
X-VirusChecked: Checked
Received: (qmail 17342 invoked from network); 17 Apr 2014 00:23:46 -0000
Received: from crpehubprd01.polycom.com (HELO crpehubprd02.polycom.com) (140.242.64.158) by server-15.tower-96.messagelabs.com with AES128-SHA encrypted SMTP; 17 Apr 2014 00:23:46 -0000
Received: from CRPMBOXPRD08.polycom.com ([169.254.1.94]) by crpehubprd02.polycom.com ([::1]) with mapi; Wed, 16 Apr 2014 17:23:43 -0700
From: "Duckworth, Mark" <Mark.Duckworth@polycom.com>
To: Paul Coverdale <coverdale@sympatico.ca>, "clue@ietf.org" <clue@ietf.org>
Date: Wed, 16 Apr 2014 17:23:41 -0700
Thread-Topic: [clue] Requirements - RE:  Improving treatment of audio
Thread-Index: Ac9ZnuQMDguaLxNfSK6fy9F3TftPxQAA32aAAAvqkHA=
Message-ID: <5C4AC54BFF7A0842A6A11F554D6FB52F08C48F@CRPMBOXPRD08.polycom.com>
References: <5C4AC54BFF7A0842A6A11F554D6FB52F08C3EA@CRPMBOXPRD08.polycom.com> <BLU0-SMTP3873D4CD2388D7E5E49059D0530@phx.gbl>
In-Reply-To: <BLU0-SMTP3873D4CD2388D7E5E49059D0530@phx.gbl>
Accept-Language: en-US
Content-Language: en-US
X-MS-Has-Attach: 
X-MS-TNEF-Correlator: 
acceptlanguage: en-US
Content-Type: text/plain; charset="us-ascii"
Content-Transfer-Encoding: quoted-printable
MIME-Version: 1.0
Archived-At: http://mailarchive.ietf.org/arch/msg/clue/CSFYk-DLvqCwD4Dv3rcioNWj6-k
Subject: Re: [clue] Requirements - RE:  Improving treatment of audio
X-BeenThere: clue@ietf.org
X-Mailman-Version: 2.1.15
Precedence: list
List-Id: CLUE - ControLling mUltiple streams for TElepresence <clue.ietf.org>
List-Unsubscribe: <https://www.ietf.org/mailman/options/clue>, <mailto:clue-request@ietf.org?subject=unsubscribe>
List-Archive: <http://www.ietf.org/mail-archive/web/clue/>
List-Post: <mailto:clue@ietf.org>
List-Help: <mailto:clue-request@ietf.org?subject=help>
List-Subscribe: <https://www.ietf.org/mailman/listinfo/clue>, <mailto:clue-request@ietf.org?subject=subscribe>
X-List-Received-Date: Thu, 17 Apr 2014 00:23:56 -0000

Hi Paul,
A couple of responses are inline below.
Mark

> -----Original Message-----
> From: Paul Coverdale [mailto:coverdale@sympatico.ca]
> Sent: Wednesday, April 16, 2014 7:00 PM
> To: Duckworth, Mark; clue@ietf.org
> Subject: RE: [clue] Requirements - RE: Improving treatment of audio
>=20
> I generally agree with Mark's points. See my comments inline.
>=20
> ..Paul
=20
 ...snip...

> >> > ] REQMT-2: The solution MUST support a description of the spatial
> >> > ]          arrangement of captured source audio sent in audio
> >streams
> >> > ]          which enables a satisfactory reproduction at the receiver
> >> > ]          in a spatially correct manner.
> >> >
> >> >     Other than wondering what "spatially correct means, I view this
> >> > as attainable, since the responsibily must belong to the receiver,
> >> > and that only if the sender chooses to specify enough audio sources.
> >>
> >> What do you mean by "attainable"? Can this be achieved using the data
> >> model we have now? (I don't think so.)
> >>
> >> ISTM that the responsibility is *shared*:
> >> - the advertiser must provide sufficient sources, *and* sufficient
> >>    description of those sources to permit the receiver to do the right
> >>    thing.
> >>
> >> - the receiver must utilize the available information in order to
> >>    reproduce the audio in a spatially correct manner, within the
> >limits
> >>    of its available equipment.
> >>
> >> - *we* are responsible for providing sufficient expressiveness in
> >>    the advertisement and RTP so that the advertiser and receiver can
> >>    do the above. *That* is the part we need to work on now.
> >
> >[Duckworth, Mark] I think REQMT-2 means only that audio can be rendered
> >spatially consistent with video, within a scene.  To me this is very
> >simple, like what a DVD player (with home theater receiver) does.  It
> >renders audio and video, spatially consistent with each other (person
> >on left of screen can be heard coming from left side of room).  It can
> >render different audio formats (e.g. stereo, surround) depending on the
> >physical loudspeaker configuration and the audio rendering capability.
> >Likewise in CLUE, the sender can send one audio capture (could be
> >multiple channels) for a scene, and the receiver can just render it in
> >same manner as the DVD player would.  The receiver could choose
> >different encoding of the audio capture depending on the receiver's
> >capability (different encoding for mono, stereo, 3-channel, for
> >example), and assuming the sender advertises those encodings. (or maybe
> >they are different alternative captures, rather than different
> >encodings of the same capture.  Ne  ed to think about this and how it's
> >related to SDP m-line for audio).
>=20
> [PVC] This is my interpretation of REQMT-2, also.
>=20
>=20
> >[Duckworth, Mark] It gets a bit more complicated when the sender wants
> >to send multiple captures for the same scene, for example different
> >captures that represent different parts of the room.  This is how TIP
> >works, with three separate captures for left, center, and right.  Now
> >we have to identify spatially how those multiple audio captures relate
> >to the video.  This is where the capture area was intended to be used.
> >But people are confused about that, and have reasons why it won't work
> well.
> >Maybe this topic needs more discussion.  TIP is very simple and
> >limited, in that it forces exactly 3 video and 3 audio with one to one
> >correspondence with no flexibility.
>=20
> [PVC] I agree that the TIP example is very simple, but for the more gener=
al
> case of multiple audio and video captures, why can't the Advertisement fr=
om
> the Provider provide information on the association between them?

[Duckworth, Mark] The area of capture attribute was intended to provide thi=
s association between audio and video captures.  But now using that mechani=
sm is in question.  If somebody has another proposal, we can discuss it.

...snip...

> >> > ]  REQMT-2c: The solution MUST support a means to identify
> >> > ]            the point of capture of individual audio
> >> > ]            captures in three dimensions.
> >> >
> >> >     Well stated. Adding point on line of capture helps slightly,
> >> > but only if we have a useful approximation of sensitivity pattern
> >> > (which I don't believe we can expect to get).
> >
> >[Duckworth, Mark] We have this now in CLUE as a media capture
> >attribute, but I don't understand how it can be used by the audio
> >receiver as an enhancement to the rendering process.  I don't see a use
> case for it.
>=20
> [PVC] This may not necessary for satisfying current use cases, but knowin=
g
> the spatial relationships of all microphones at the sending end (and thei=
r
> sensitivities and directional patterns), could allow the receiving end to=
 do
> some fancy spatial audio processing.

[Duckworth, Mark] I wasn't suggesting we remove these attributes.  I wouldn=
't oppose adding new attributes for audio capture sensitivities and directi=
onal patterns either, as long as it isn't mandatory for an advertisement to=
 include them.
=20
...snip...=20


From nobody Thu Apr 17 11:26:06 2014
Return-Path: <stephen.botzko@gmail.com>
X-Original-To: clue@ietfa.amsl.com
Delivered-To: clue@ietfa.amsl.com
Received: from localhost (ietfa.amsl.com [127.0.0.1]) by ietfa.amsl.com (Postfix) with ESMTP id 2C9EF1A0149 for <clue@ietfa.amsl.com>; Thu, 17 Apr 2014 11:26:05 -0700 (PDT)
X-Virus-Scanned: amavisd-new at amsl.com
X-Spam-Flag: NO
X-Spam-Score: -0.1
X-Spam-Level: 
X-Spam-Status: No, score=-0.1 tagged_above=-999 required=5 tests=[BAYES_20=-0.001, DKIM_SIGNED=0.1, DKIM_VALID=-0.1, DKIM_VALID_AU=-0.1, FREEMAIL_FROM=0.001, HTML_MESSAGE=0.001, SPF_PASS=-0.001] autolearn=ham
Received: from mail.ietf.org ([4.31.198.44]) by localhost (ietfa.amsl.com [127.0.0.1]) (amavisd-new, port 10024) with ESMTP id OBeIx2x4ywfv for <clue@ietfa.amsl.com>; Thu, 17 Apr 2014 11:26:00 -0700 (PDT)
Received: from mail-ve0-x22e.google.com (mail-ve0-x22e.google.com [IPv6:2607:f8b0:400c:c01::22e]) by ietfa.amsl.com (Postfix) with ESMTP id 1ECCE1A0110 for <clue@ietf.org>; Thu, 17 Apr 2014 11:26:00 -0700 (PDT)
Received: by mail-ve0-f174.google.com with SMTP id oz11so921954veb.33 for <clue@ietf.org>; Thu, 17 Apr 2014 11:25:56 -0700 (PDT)
DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20120113;  h=mime-version:in-reply-to:references:date:message-id:subject:from:to :cc:content-type; bh=+y7ijYbHv3pLQMgfQpRePwQpilnzelqHupYeegBhKX4=; b=JiUrMD8WmYPbUZyJOVYnCggyUaUXZgVsN1sSQynZRwt9b6eesF+XJDv5c8cAq/PTRl hXt9pnHXSj52gTGtm9Phr4gucK8vYOGhSylDvdeLKwvwHurzyRJL19XdEQk33aYNkjZq +e7iMZFvUoMuCDoVlx0AUgDSS1154/Vix5Lz6WcFTTZZluvqibpqQxN30o0JE4o1b6Aq OhIx6CZidcYMMM2oIuO5hmdj3QlmM1AZ/LFfniBzYgZ+zJxthjEG7YB7s8sIMDKoExRq krIuMbOIG7xyldECnXuf8SlT3MOpl+rgGDbGZrB/W1Vyaxb3DTcf9Gj9mRwY1zWy9KQR Ucdw==
MIME-Version: 1.0
X-Received: by 10.220.95.204 with SMTP id e12mr48017vcn.37.1397759156320; Thu, 17 Apr 2014 11:25:56 -0700 (PDT)
Received: by 10.221.40.135 with HTTP; Thu, 17 Apr 2014 11:25:56 -0700 (PDT)
In-Reply-To: <534D645A.9010703@alum.mit.edu>
References: <533AF351.9050201@alum.mit.edu> <BLU0-SMTP942A03E86972E7B8A212C6D0560@phx.gbl> <20140413233154.GL60844@verdi> <534C02BC.6000705@alum.mit.edu> <20140415135840.GA89388@verdi> <534D645A.9010703@alum.mit.edu>
Date: Thu, 17 Apr 2014 14:25:56 -0400
Message-ID: <CAMC7SJ4K4FPAj-XWzwQWK6wie20YcFSkGB8tDrYS=TEa0wTrig@mail.gmail.com>
From: Stephen Botzko <stephen.botzko@gmail.com>
To: Paul Kyzivat <pkyzivat@alum.mit.edu>
Content-Type: multipart/alternative; boundary=001a11c1f3ecfe7edb04f7412954
Archived-At: http://mailarchive.ietf.org/arch/msg/clue/uvcG6_ZISz46diUyRzSbsFw8kcw
Cc: CLUE <clue@ietf.org>
Subject: Re: [clue] Improving treatment of audio
X-BeenThere: clue@ietf.org
X-Mailman-Version: 2.1.15
Precedence: list
List-Id: CLUE - ControLling mUltiple streams for TElepresence <clue.ietf.org>
List-Unsubscribe: <https://www.ietf.org/mailman/options/clue>, <mailto:clue-request@ietf.org?subject=unsubscribe>
List-Archive: <http://www.ietf.org/mail-archive/web/clue/>
List-Post: <mailto:clue@ietf.org>
List-Help: <mailto:clue-request@ietf.org?subject=help>
List-Subscribe: <https://www.ietf.org/mailman/listinfo/clue>, <mailto:clue-request@ietf.org?subject=subscribe>
X-List-Received-Date: Thu, 17 Apr 2014 18:26:05 -0000

--001a11c1f3ecfe7edb04f7412954
Content-Type: text/plain; charset=UTF-8

Just a general comment on echo...

There are telepresence systems in the field already from multiple vendors.
 They use multiple microphones and multiple speakers, rendering
multichannel audio.  Echo cancellation in those systems works, even if they
are disparate systems using TIP.

Even my home non-telepresence system negotiates full duplex fullband stereo
audio, and it properly cancels the echo even when the received sound is
being rendered through 5 speakers via an external surround sound processor.

And if you build a system which can't handle multichannel echo
cancellation, you can still be CLUE-full.  All you need to do is downmix
the receive audio to mono before you render it and use your single
microphone for capture.  That keeps your system from injecting echo into
the conference.

So echo is not a problem for CLUE.  I agree with Mark that it is
out-of-scope.  We don't see the need to talk about echo when we write up
SIP signalling for multichannel payload formats, and I think the same
principle applies here.

BR,
Steve


On Tue, Apr 15, 2014 at 12:54 PM, Paul Kyzivat <pkyzivat@alum.mit.edu>wrote:

> On 4/15/14 9:58 AM, John Leslie wrote:
>
>> Paul Kyzivat <pkyzivat@alum.mit.edu> wrote:
>>
>>> On 4/13/14 7:31 PM, John Leslie wrote:
>>>
>>>>
>>>> ] REQMT-2:...
>>>>
>>>> Worse, I'm not confident that all of this is even attainable. :^(
>>>>
>>>
>>> I don't know if they are unattainable. But IIUC they cannot be attained
>>> with the information we currently have in the fw and data model.
>>>
>>> If we want to give up on attaining some of them, then let's do so, and
>>> explicitly note that we have. And then ensure that we have sufficient
>>> information to attain the rest.
>>>
>>
>>     ISTM we're not ready to give up on any of them, but I'm willing to
>> continue the conversation...
>>
>
> We may have a conflict here between continuing the conversation and
> getting something done. We've already been working on this stuff for years.
> If we haven't figured it out by now, then maybe we need to simplify.
>
> So please continue the conversation, maybe for as much as a few days. But
> not for a few more months!
>
>
>  There is, alas, an elephant in each room which can render our
>>>> audio treatment useless -- it's called "echo". :^(
>>>>
>>>
>>> IMO, if we don't have a solution that can prevent this then we don't
>>> have a solution worth publishing.
>>>
>>
>>     I think this deserves a separate thread...
>>
>
> Fine. I don't understand. I get the impression that there is something
> that must be said about echo. I leave it to you.
>
>
>  ] REQMT-2: The solution MUST support a description of the spatial
>>>> ]          arrangement of captured source audio sent in audio streams
>>>> ]          which enables a satisfactory reproduction at the receiver
>>>> ]          in a spatially correct manner.
>>>>
>>>>     Other than wondering what "spatially correct means, I view this as
>>>> attainable, since the responsibily must belong to the receiver, and
>>>> that only if the sender chooses to specify enough audio sources.
>>>>
>>>
>>> What do you mean by "attainable"? Can this be achieved using the data
>>> model we have now? (I don't think so.)
>>>
>>
>>     Perhaps REQMT-2 means different things to different folk... :^(
>>
>>     (Unquestionably, we have different ideas of _how_ to "enable a
>> satisfactory reproduction" -- but this doesn't bother me: I expect
>> end-users to adapt to the reproduction they pay for.)
>>
>
> IMO it is fine that some things are left to "quality of implementation",
> as long as those things are entirely within the scope of a single endpoint.
>
>
>      (BTW, I don't think we need to establish a single meaning for
>> "spatially correct" -- I only meant to stipulate that I was trying to
>> avoid that issue.)
>>
>>     IMHO, a "satisfactory reproduction" can be achieved given separate
>> audio feeds with a clear specification of the location of each (so long
>> as you don't have to deal with echo of audio you sent yourself).
>>
>>     Others, no doubt, prefer a different paradigm where each room sends
>> a surround-sound representation of their own room. It _is_ possible
>> to process that into a _different_ surround-sound representation which
>> will be satisfactory. (Of course, if the received surround-sound has
>> echo of what _you_ sent, my mind starts to boggle about cancelling echo.)
>>
>
> Your last two paragraphs above leave me thinking that there is something
> more that must be said in the framework. But I have no idea what it is.
>
>
>  ISTM that the responsibility is *shared*:
>>> - the advertiser must provide sufficient sources, *and* sufficient
>>>    description of those sources to permit the receiver to do the right
>>>    thing.
>>>
>>> - the receiver must utilize the available information in order to
>>>    reproduce the audio in a spatially correct manner, within the limits
>>>    of its available equipment.
>>>
>>
>>     Yes.
>>
>>  - *we* are responsible for providing sufficient expressiveness in
>>>    the advertisement and RTP so that the advertiser and receiver can
>>>    do the above. *That* is the part we need to work on now.
>>>
>>
>>     Speaking for myself, I'll settle for multiple channels of audio
>> with point-of-capture for each mike.
>>
>>     (I know I won't always get that -- but I don't blame the CLUE spec
>> for that.)
>>
>
> Well, if there is some minimal set of information that will always be
> required to get a workable result, then we should make that mandatory.
>
> E.g., suppose the advertisement has a scene with multiple audio captures
> and *no* spatial information about them. Can a receiver to anything useful
> with that? If not, then we should require more.
>
>
>  If a sender _actually_ sends "stereo" or "3.0" the receiver will have
>>>> to punt. I can only hope that our actual standard doesn't encourage
>>>> this.
>>>>
>>>
>>> Sounds like something we need to talk about. Where?
>>>
>>
>>     I have no suggestions...
>>
>
> Then maybe we want to declare it "out of scope".
>
>
>  Inevitably, some yahoo is going to put one or more loudspeakers in
>>>> the room. And those will expand the "area of coverage" to include the
>>>> other rooms those loudspeakers attempt to render.
>>>>
>>>> With delay!!!
>>>>
>>>
>>> OK. But surely the point of capture isn't sufficient by itself.
>>> *Something* more is required. Can you suggest something that is stated
>>> better, that would be attainable and useful for our purposes?
>>>
>>
>>     Distance from the speaker's mouth would help. If it's "close enough"
>> the receiver needn't worry -- when an individual want's to mute, that's
>> the sender's problem.
>>
>
> I want those of you who know something about this to decide what we ought
> to make provision for in the advertisement.
>
>
>  ... the most useful audio streams are linked to microphones "close" to
>>>> individual people, while video captures will _very_ often try to cover
>>>> more than one person -- thus video capures will very typically be
>>>> associated with several audio captures.
>>>>
>>>
>>> Are you saying that within a scene the choice of audio captures is
>>> unrelated to the choice of video captures?
>>>
>>
>>     Mostly, yes.
>>
>
> That would be a really good thing to say in the framework!
>
>
>  If so, then does it ever make sense to choose a subset of the advertised
>>> audio? And if that is true, then what criteria make sense?
>>>
>>
>>     Clearly, there will be cases where an individual in one room is
>> particularly interested in some limited area within another room. IMHO
>> this is _in_addition_to_ the general room audio, so I see no reason
>> to worry about it in our specification.
>>
>
> What about a case where the receiver is limited in the number of audio
> captures it can receive and/or process? Can the advertisement help the
> receiver decide which one(s) it should get?
>
> (We do have the priority attribute if nothing else.)
>
>
>  Thus I most sincerely hope that folks won't choose to switch their
>>>> audio based on who's speaking -- but I won't try to stop them from doing
>>>> this. (And I must admit that sometimes I _shouldn't_ blame them, because
>>>> they find themselves sabotaged by audio sources contaminated by echo.)
>>>>
>>>
> Maybe I misunderstood you above. By "folks" did you mean providers or
> consumers? I was thinking you meant providers.
>
>
>  I assume you mean when the speakers are in the same room, right?
>>>
>>
>>     What I meant was audio feeds they receive coming in with so much echo
>> that they find themselves abliged to exclude some of their own mikes
>> from what they send.
>>
>>     This, to tell truth, is a somewhat weird corner case. They should be
>> able to exclude the _received_ echo from their outgoing feed by normal
>> echo-cancellation algorithms. If I had thought this through while
>> writing the above text, I would have erased it before sending.
>>
>
> Again, I don't fully understand, and probably don't need to. But it sounds
> like something more is needed in the fw to discuss expectations.
>
>
>  Isn't the situation different when we are switching among speakers in
>>> different rooms?
>>>
>>
>>     I don't understand the question...
>>
>
> I was thinking of an MCU advertising, among other things, switched audio
> and video captures, where the alternatives come from different rooms. Then
> the whole point is to switch the audio and video together based on who is
> speaking.
>
>         Thanks,
>         Paul
>
>
> _______________________________________________
> clue mailing list
> clue@ietf.org
> https://www.ietf.org/mailman/listinfo/clue
>

--001a11c1f3ecfe7edb04f7412954
Content-Type: text/html; charset=UTF-8
Content-Transfer-Encoding: quoted-printable

<div dir=3D"ltr">Just a general comment on echo...<div><br></div><div>There=
 are telepresence systems in the field already from multiple vendors. =C2=
=A0They use multiple microphones and multiple speakers, rendering multichan=
nel audio. =C2=A0Echo cancellation in those systems works, even if they are=
 disparate systems using TIP.</div>
<div><br></div><div>Even my home non-telepresence system negotiates full du=
plex fullband stereo audio, and it properly cancels the echo even when the =
received sound is being rendered through 5 speakers via an external surroun=
d sound processor.</div>
<div><br></div><div>And if you build a system which can&#39;t handle multic=
hannel echo cancellation, you can still be CLUE-full. =C2=A0All you need to=
 do is downmix the receive audio to mono before you render it and use your =
single microphone for capture. =C2=A0That keeps your system from injecting =
echo into the conference.=C2=A0</div>
<div><br></div><div>So echo is not a problem for CLUE. =C2=A0I agree with M=
ark that it is out-of-scope. =C2=A0We don&#39;t see the need to talk about =
echo when we write up SIP signalling for multichannel payload formats, and =
I think the same principle applies here.=C2=A0</div>
<div><br></div><div>BR,</div><div>Steve</div></div><div class=3D"gmail_extr=
a"><br><br><div class=3D"gmail_quote">On Tue, Apr 15, 2014 at 12:54 PM, Pau=
l Kyzivat <span dir=3D"ltr">&lt;<a href=3D"mailto:pkyzivat@alum.mit.edu" ta=
rget=3D"_blank">pkyzivat@alum.mit.edu</a>&gt;</span> wrote:<br>
<blockquote class=3D"gmail_quote" style=3D"margin:0 0 0 .8ex;border-left:1p=
x #ccc solid;padding-left:1ex"><div class=3D"">On 4/15/14 9:58 AM, John Les=
lie wrote:<br>
<blockquote class=3D"gmail_quote" style=3D"margin:0 0 0 .8ex;border-left:1p=
x #ccc solid;padding-left:1ex">
Paul Kyzivat &lt;<a href=3D"mailto:pkyzivat@alum.mit.edu" target=3D"_blank"=
>pkyzivat@alum.mit.edu</a>&gt; wrote:<br>
<blockquote class=3D"gmail_quote" style=3D"margin:0 0 0 .8ex;border-left:1p=
x #ccc solid;padding-left:1ex">
On 4/13/14 7:31 PM, John Leslie wrote:<br>
<blockquote class=3D"gmail_quote" style=3D"margin:0 0 0 .8ex;border-left:1p=
x #ccc solid;padding-left:1ex">
<br>
] REQMT-2:...<br>
<br>
Worse, I&#39;m not confident that all of this is even attainable. :^(<br>
</blockquote>
<br>
I don&#39;t know if they are unattainable. But IIUC they cannot be attained=
<br>
with the information we currently have in the fw and data model.<br>
<br>
If we want to give up on attaining some of them, then let&#39;s do so, and<=
br>
explicitly note that we have. And then ensure that we have sufficient<br>
information to attain the rest.<br>
</blockquote>
<br>
=C2=A0 =C2=A0 ISTM we&#39;re not ready to give up on any of them, but I&#39=
;m willing to<br>
continue the conversation...<br>
</blockquote>
<br></div>
We may have a conflict here between continuing the conversation and getting=
 something done. We&#39;ve already been working on this stuff for years. If=
 we haven&#39;t figured it out by now, then maybe we need to simplify.<br>

<br>
So please continue the conversation, maybe for as much as a few days. But n=
ot for a few more months!<div class=3D""><br>
<br>
<blockquote class=3D"gmail_quote" style=3D"margin:0 0 0 .8ex;border-left:1p=
x #ccc solid;padding-left:1ex"><blockquote class=3D"gmail_quote" style=3D"m=
argin:0 0 0 .8ex;border-left:1px #ccc solid;padding-left:1ex"><blockquote c=
lass=3D"gmail_quote" style=3D"margin:0 0 0 .8ex;border-left:1px #ccc solid;=
padding-left:1ex">

There is, alas, an elephant in each room which can render our<br>
audio treatment useless -- it&#39;s called &quot;echo&quot;. :^(<br>
</blockquote>
<br>
IMO, if we don&#39;t have a solution that can prevent this then we don&#39;=
t<br>
have a solution worth publishing.<br>
</blockquote>
<br>
=C2=A0 =C2=A0 I think this deserves a separate thread...<br>
</blockquote>
<br></div>
Fine. I don&#39;t understand. I get the impression that there is something =
that must be said about echo. I leave it to you.<div class=3D""><br>
<br>
<blockquote class=3D"gmail_quote" style=3D"margin:0 0 0 .8ex;border-left:1p=
x #ccc solid;padding-left:1ex"><blockquote class=3D"gmail_quote" style=3D"m=
argin:0 0 0 .8ex;border-left:1px #ccc solid;padding-left:1ex"><blockquote c=
lass=3D"gmail_quote" style=3D"margin:0 0 0 .8ex;border-left:1px #ccc solid;=
padding-left:1ex">

] REQMT-2: The solution MUST support a description of the spatial<br>
] =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0arrangement of captured source audio se=
nt in audio streams<br>
] =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0which enables a satisfactory reproducti=
on at the receiver<br>
] =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0in a spatially correct manner.<br>
<br>
=C2=A0 =C2=A0 Other than wondering what &quot;spatially correct means, I vi=
ew this as<br>
attainable, since the responsibily must belong to the receiver, and<br>
that only if the sender chooses to specify enough audio sources.<br>
</blockquote>
<br>
What do you mean by &quot;attainable&quot;? Can this be achieved using the =
data<br>
model we have now? (I don&#39;t think so.)<br>
</blockquote>
<br>
=C2=A0 =C2=A0 Perhaps REQMT-2 means different things to different folk... :=
^(<br>
<br>
=C2=A0 =C2=A0 (Unquestionably, we have different ideas of _how_ to &quot;en=
able a<br>
satisfactory reproduction&quot; -- but this doesn&#39;t bother me: I expect=
<br>
end-users to adapt to the reproduction they pay for.)<br>
</blockquote>
<br></div>
IMO it is fine that some things are left to &quot;quality of implementation=
&quot;, as long as those things are entirely within the scope of a single e=
ndpoint.<div class=3D""><br>
<br>
<blockquote class=3D"gmail_quote" style=3D"margin:0 0 0 .8ex;border-left:1p=
x #ccc solid;padding-left:1ex">
=C2=A0 =C2=A0 (BTW, I don&#39;t think we need to establish a single meaning=
 for<br>
&quot;spatially correct&quot; -- I only meant to stipulate that I was tryin=
g to<br>
avoid that issue.)<br>
<br>
=C2=A0 =C2=A0 IMHO, a &quot;satisfactory reproduction&quot; can be achieved=
 given separate<br>
audio feeds with a clear specification of the location of each (so long<br>
as you don&#39;t have to deal with echo of audio you sent yourself).<br>
<br>
=C2=A0 =C2=A0 Others, no doubt, prefer a different paradigm where each room=
 sends<br>
a surround-sound representation of their own room. It _is_ possible<br>
to process that into a _different_ surround-sound representation which<br>
will be satisfactory. (Of course, if the received surround-sound has<br>
echo of what _you_ sent, my mind starts to boggle about cancelling echo.)<b=
r>
</blockquote>
<br></div>
Your last two paragraphs above leave me thinking that there is something mo=
re that must be said in the framework. But I have no idea what it is.<div c=
lass=3D""><br>
<br>
<blockquote class=3D"gmail_quote" style=3D"margin:0 0 0 .8ex;border-left:1p=
x #ccc solid;padding-left:1ex"><blockquote class=3D"gmail_quote" style=3D"m=
argin:0 0 0 .8ex;border-left:1px #ccc solid;padding-left:1ex">
ISTM that the responsibility is *shared*:<br>
- the advertiser must provide sufficient sources, *and* sufficient<br>
=C2=A0 =C2=A0description of those sources to permit the receiver to do the =
right<br>
=C2=A0 =C2=A0thing.<br>
<br>
- the receiver must utilize the available information in order to<br>
=C2=A0 =C2=A0reproduce the audio in a spatially correct manner, within the =
limits<br>
=C2=A0 =C2=A0of its available equipment.<br>
</blockquote>
<br>
=C2=A0 =C2=A0 Yes.<br>
<br>
<blockquote class=3D"gmail_quote" style=3D"margin:0 0 0 .8ex;border-left:1p=
x #ccc solid;padding-left:1ex">
- *we* are responsible for providing sufficient expressiveness in<br>
=C2=A0 =C2=A0the advertisement and RTP so that the advertiser and receiver =
can<br>
=C2=A0 =C2=A0do the above. *That* is the part we need to work on now.<br>
</blockquote>
<br>
=C2=A0 =C2=A0 Speaking for myself, I&#39;ll settle for multiple channels of=
 audio<br>
with point-of-capture for each mike.<br>
<br>
=C2=A0 =C2=A0 (I know I won&#39;t always get that -- but I don&#39;t blame =
the CLUE spec<br>
for that.)<br>
</blockquote>
<br></div>
Well, if there is some minimal set of information that will always be requi=
red to get a workable result, then we should make that mandatory.<br>
<br>
E.g., suppose the advertisement has a scene with multiple audio captures an=
d *no* spatial information about them. Can a receiver to anything useful wi=
th that? If not, then we should require more.<div class=3D""><br>
<br>
<blockquote class=3D"gmail_quote" style=3D"margin:0 0 0 .8ex;border-left:1p=
x #ccc solid;padding-left:1ex"><blockquote class=3D"gmail_quote" style=3D"m=
argin:0 0 0 .8ex;border-left:1px #ccc solid;padding-left:1ex"><blockquote c=
lass=3D"gmail_quote" style=3D"margin:0 0 0 .8ex;border-left:1px #ccc solid;=
padding-left:1ex">

If a sender _actually_ sends &quot;stereo&quot; or &quot;3.0&quot; the rece=
iver will have<br>
to punt. I can only hope that our actual standard doesn&#39;t encourage<br>
this.<br>
</blockquote>
<br>
Sounds like something we need to talk about. Where?<br>
</blockquote>
<br>
=C2=A0 =C2=A0 I have no suggestions...<br>
</blockquote>
<br></div>
Then maybe we want to declare it &quot;out of scope&quot;.<div class=3D""><=
br>
<br>
<blockquote class=3D"gmail_quote" style=3D"margin:0 0 0 .8ex;border-left:1p=
x #ccc solid;padding-left:1ex"><blockquote class=3D"gmail_quote" style=3D"m=
argin:0 0 0 .8ex;border-left:1px #ccc solid;padding-left:1ex"><blockquote c=
lass=3D"gmail_quote" style=3D"margin:0 0 0 .8ex;border-left:1px #ccc solid;=
padding-left:1ex">

Inevitably, some yahoo is going to put one or more loudspeakers in<br>
the room. And those will expand the &quot;area of coverage&quot; to include=
 the<br>
other rooms those loudspeakers attempt to render.<br>
<br>
With delay!!!<br>
</blockquote>
<br>
OK. But surely the point of capture isn&#39;t sufficient by itself.<br>
*Something* more is required. Can you suggest something that is stated<br>
better, that would be attainable and useful for our purposes?<br>
</blockquote>
<br>
=C2=A0 =C2=A0 Distance from the speaker&#39;s mouth would help. If it&#39;s=
 &quot;close enough&quot;<br>
the receiver needn&#39;t worry -- when an individual want&#39;s to mute, th=
at&#39;s<br>
the sender&#39;s problem.<br>
</blockquote>
<br></div>
I want those of you who know something about this to decide what we ought t=
o make provision for in the advertisement.<div class=3D""><br>
<br>
<blockquote class=3D"gmail_quote" style=3D"margin:0 0 0 .8ex;border-left:1p=
x #ccc solid;padding-left:1ex"><blockquote class=3D"gmail_quote" style=3D"m=
argin:0 0 0 .8ex;border-left:1px #ccc solid;padding-left:1ex"><blockquote c=
lass=3D"gmail_quote" style=3D"margin:0 0 0 .8ex;border-left:1px #ccc solid;=
padding-left:1ex">

... the most useful audio streams are linked to microphones &quot;close&quo=
t; to<br>
individual people, while video captures will _very_ often try to cover<br>
more than one person -- thus video capures will very typically be<br>
associated with several audio captures.<br>
</blockquote>
<br>
Are you saying that within a scene the choice of audio captures is<br>
unrelated to the choice of video captures?<br>
</blockquote>
<br>
=C2=A0 =C2=A0 Mostly, yes.<br>
</blockquote>
<br></div>
That would be a really good thing to say in the framework!<div class=3D""><=
br>
<br>
<blockquote class=3D"gmail_quote" style=3D"margin:0 0 0 .8ex;border-left:1p=
x #ccc solid;padding-left:1ex"><blockquote class=3D"gmail_quote" style=3D"m=
argin:0 0 0 .8ex;border-left:1px #ccc solid;padding-left:1ex">
If so, then does it ever make sense to choose a subset of the advertised<br=
>
audio? And if that is true, then what criteria make sense?<br>
</blockquote>
<br>
=C2=A0 =C2=A0 Clearly, there will be cases where an individual in one room =
is<br>
particularly interested in some limited area within another room. IMHO<br>
this is _in_addition_to_ the general room audio, so I see no reason<br>
to worry about it in our specification.<br>
</blockquote>
<br></div>
What about a case where the receiver is limited in the number of audio capt=
ures it can receive and/or process? Can the advertisement help the receiver=
 decide which one(s) it should get?<br>
<br>
(We do have the priority attribute if nothing else.)<div class=3D""><br>
<br>
<blockquote class=3D"gmail_quote" style=3D"margin:0 0 0 .8ex;border-left:1p=
x #ccc solid;padding-left:1ex"><blockquote class=3D"gmail_quote" style=3D"m=
argin:0 0 0 .8ex;border-left:1px #ccc solid;padding-left:1ex"><blockquote c=
lass=3D"gmail_quote" style=3D"margin:0 0 0 .8ex;border-left:1px #ccc solid;=
padding-left:1ex">

Thus I most sincerely hope that folks won&#39;t choose to switch their<br>
audio based on who&#39;s speaking -- but I won&#39;t try to stop them from =
doing<br>
this. (And I must admit that sometimes I _shouldn&#39;t_ blame them, becaus=
e<br>
they find themselves sabotaged by audio sources contaminated by echo.)<br>
</blockquote></blockquote></blockquote>
<br></div>
Maybe I misunderstood you above. By &quot;folks&quot; did you mean provider=
s or consumers? I was thinking you meant providers.<div class=3D""><br>
<br>
<blockquote class=3D"gmail_quote" style=3D"margin:0 0 0 .8ex;border-left:1p=
x #ccc solid;padding-left:1ex"><blockquote class=3D"gmail_quote" style=3D"m=
argin:0 0 0 .8ex;border-left:1px #ccc solid;padding-left:1ex">
I assume you mean when the speakers are in the same room, right?<br>
</blockquote>
<br>
=C2=A0 =C2=A0 What I meant was audio feeds they receive coming in with so m=
uch echo<br>
that they find themselves abliged to exclude some of their own mikes<br>
from what they send.<br>
<br>
=C2=A0 =C2=A0 This, to tell truth, is a somewhat weird corner case. They sh=
ould be<br>
able to exclude the _received_ echo from their outgoing feed by normal<br>
echo-cancellation algorithms. If I had thought this through while<br>
writing the above text, I would have erased it before sending.<br>
</blockquote>
<br></div>
Again, I don&#39;t fully understand, and probably don&#39;t need to. But it=
 sounds like something more is needed in the fw to discuss expectations.<di=
v class=3D""><br>
<br>
<blockquote class=3D"gmail_quote" style=3D"margin:0 0 0 .8ex;border-left:1p=
x #ccc solid;padding-left:1ex"><blockquote class=3D"gmail_quote" style=3D"m=
argin:0 0 0 .8ex;border-left:1px #ccc solid;padding-left:1ex">
Isn&#39;t the situation different when we are switching among speakers in<b=
r>
different rooms?<br>
</blockquote>
<br>
=C2=A0 =C2=A0 I don&#39;t understand the question...<br>
</blockquote>
<br></div>
I was thinking of an MCU advertising, among other things, switched audio an=
d video captures, where the alternatives come from different rooms. Then th=
e whole point is to switch the audio and video together based on who is spe=
aking.<br>

<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 Thanks,<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 Paul<div class=3D"HOEnZb"><div class=3D"h5"><br=
>
<br>
______________________________<u></u>_________________<br>
clue mailing list<br>
<a href=3D"mailto:clue@ietf.org" target=3D"_blank">clue@ietf.org</a><br>
<a href=3D"https://www.ietf.org/mailman/listinfo/clue" target=3D"_blank">ht=
tps://www.ietf.org/mailman/<u></u>listinfo/clue</a><br>
</div></div></blockquote></div><br></div>

--001a11c1f3ecfe7edb04f7412954--


From nobody Thu Apr 17 16:03:50 2014
Return-Path: <wwwrun@rfc-editor.org>
X-Original-To: clue@ietfa.amsl.com
Delivered-To: clue@ietfa.amsl.com
Received: from localhost (ietfa.amsl.com [127.0.0.1]) by ietfa.amsl.com (Postfix) with ESMTP id 5B6091A0217; Thu, 17 Apr 2014 16:03:44 -0700 (PDT)
X-Virus-Scanned: amavisd-new at amsl.com
X-Spam-Flag: NO
X-Spam-Score: -2.174
X-Spam-Level: 
X-Spam-Status: No, score=-2.174 tagged_above=-999 required=5 tests=[BAYES_00=-1.9, RP_MATCHES_RCVD=-0.272, SPF_HELO_PASS=-0.001, SPF_PASS=-0.001] autolearn=ham
Received: from mail.ietf.org ([4.31.198.44]) by localhost (ietfa.amsl.com [127.0.0.1]) (amavisd-new, port 10024) with ESMTP id KH83EvQo7Tu6; Thu, 17 Apr 2014 16:03:34 -0700 (PDT)
Received: from rfc-editor.org (rfc-editor.org [4.31.198.49]) by ietfa.amsl.com (Postfix) with ESMTP id 992871A0176; Thu, 17 Apr 2014 16:03:34 -0700 (PDT)
Received: by rfc-editor.org (Postfix, from userid 30) id 5294A18000C; Thu, 17 Apr 2014 16:03:01 -0700 (PDT)
To: ietf-announce@ietf.org, rfc-dist@rfc-editor.org
X-PHP-Originating-Script: 6000:ams_util_lib.php
From: rfc-editor@rfc-editor.org
Message-Id: <20140417230301.5294A18000C@rfc-editor.org>
Date: Thu, 17 Apr 2014 16:03:01 -0700 (PDT)
Archived-At: http://mailarchive.ietf.org/arch/msg/clue/v3voyBo7USNJS0i8Fdv1giAtirI
Cc: drafts-update-ref@iana.org, clue@ietf.org, rfc-editor@rfc-editor.org
Subject: [clue] RFC 7205 on Use Cases for Telepresence Multistreams
X-BeenThere: clue@ietf.org
X-Mailman-Version: 2.1.15
Precedence: list
List-Id: CLUE - ControLling mUltiple streams for TElepresence <clue.ietf.org>
List-Unsubscribe: <https://www.ietf.org/mailman/options/clue>, <mailto:clue-request@ietf.org?subject=unsubscribe>
List-Archive: <http://www.ietf.org/mail-archive/web/clue/>
List-Post: <mailto:clue@ietf.org>
List-Help: <mailto:clue-request@ietf.org?subject=help>
List-Subscribe: <https://www.ietf.org/mailman/listinfo/clue>, <mailto:clue-request@ietf.org?subject=subscribe>
X-List-Received-Date: Thu, 17 Apr 2014 23:03:45 -0000

A new Request for Comments is now available in online RFC libraries.

        
        RFC 7205

        Title:      Use Cases for Telepresence Multistreams 
        Author:     A. Romanow, S. Botzko,
                    M. Duckworth, R. Even, Ed.
        Status:     Informational
        Stream:     IETF
        Date:       April 2014
        Mailbox:    allyn@cisco.com, 
                    stephen.botzko@polycom.com, 
                    mark.duckworth@polycom.com, 
                    roni.even@mail01.huawei.com
        Pages:      17
        Characters: 42087
        Updates/Obsoletes/SeeAlso:   None

        I-D Tag:    draft-ietf-clue-telepresence-use-cases-09.txt

        URL:        http://www.rfc-editor.org/rfc/rfc7205.txt

Telepresence conferencing systems seek to create an environment that
gives users (or user groups) that are not co-located a feeling of
co-located presence through multimedia communication that includes at
least audio and video signals of high fidelity.  A number of
techniques for handling audio and video streams are used to create
this experience.  When these techniques are not similar,
interoperability between different systems is difficult at best, and
often not possible.  Conveying information about the relationships
between multiple streams of media would enable senders and receivers
to make choices to allow telepresence systems to interwork.  This
memo describes the most typical and important use cases for sending
multiple streams in a telepresence conference.

This document is a product of the ControLling mUltiple streams for tElepresence Working Group of the IETF.


INFORMATIONAL: This memo provides information for the Internet community.
It does not specify an Internet standard of any kind. Distribution of
this memo is unlimited.

This announcement is sent to the IETF-Announce and rfc-dist lists.
To subscribe or unsubscribe, see
  http://www.ietf.org/mailman/listinfo/ietf-announce
  http://mailman.rfc-editor.org/mailman/listinfo/rfc-dist

For searching the RFC series, see http://www.rfc-editor.org/search
For downloading RFCs, see http://www.rfc-editor.org/rfc.html

Requests for special distribution should be addressed to either the
author of the RFC in question, or to rfc-editor@rfc-editor.org.  Unless
specifically noted otherwise on the RFC itself, all RFCs are for
unlimited distribution.


The RFC Editor Team
Association Management Solutions, LLC



From nobody Thu Apr 17 19:56:09 2014
Return-Path: <pkyzivat@alum.mit.edu>
X-Original-To: clue@ietfa.amsl.com
Delivered-To: clue@ietfa.amsl.com
Received: from localhost (ietfa.amsl.com [127.0.0.1]) by ietfa.amsl.com (Postfix) with ESMTP id D92C31A01E8 for <clue@ietfa.amsl.com>; Thu, 17 Apr 2014 19:56:05 -0700 (PDT)
X-Virus-Scanned: amavisd-new at amsl.com
X-Spam-Flag: NO
X-Spam-Score: -1.235
X-Spam-Level: 
X-Spam-Status: No, score=-1.235 tagged_above=-999 required=5 tests=[BAYES_00=-1.9, DKIM_SIGNED=0.1, DKIM_VALID=-0.1, SPF_SOFTFAIL=0.665] autolearn=no
Received: from mail.ietf.org ([4.31.198.44]) by localhost (ietfa.amsl.com [127.0.0.1]) (amavisd-new, port 10024) with ESMTP id AMOI0sPBX13t for <clue@ietfa.amsl.com>; Thu, 17 Apr 2014 19:56:02 -0700 (PDT)
Received: from QMTA11.westchester.pa.mail.comcast.net (qmta11.westchester.pa.mail.comcast.net [IPv6:2001:558:fe14:44:76:96:59:211]) by ietfa.amsl.com (Postfix) with ESMTP id C81331A001E for <clue@ietf.org>; Thu, 17 Apr 2014 19:56:01 -0700 (PDT)
Received: from omta05.westchester.pa.mail.comcast.net ([76.96.62.43]) by QMTA11.westchester.pa.mail.comcast.net with comcast id rEey1n0060vyq2s5BEvxPk; Fri, 18 Apr 2014 02:55:57 +0000
Received: from Paul-Kyzivats-MacBook-Pro.local ([50.138.229.164]) by omta05.westchester.pa.mail.comcast.net with comcast id rEvx1n00R3ZTu2S3REvxkE; Fri, 18 Apr 2014 02:55:57 +0000
Message-ID: <5350943D.6080600@alum.mit.edu>
Date: Thu, 17 Apr 2014 22:55:57 -0400
From: Paul Kyzivat <pkyzivat@alum.mit.edu>
User-Agent: Mozilla/5.0 (Macintosh; Intel Mac OS X 10.7; rv:24.0) Gecko/20100101 Thunderbird/24.4.0
MIME-Version: 1.0
To: clue@ietf.org
References: <20140417230301.5294A18000C@rfc-editor.org>
In-Reply-To: <20140417230301.5294A18000C@rfc-editor.org>
Content-Type: text/plain; charset=ISO-8859-1; format=flowed
Content-Transfer-Encoding: 7bit
DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=comcast.net; s=q20140121; t=1397789757; bh=EvU+K/oL9TPAC4qhDnWHOuwPOwgJQG5ixNdgCr1O7Pw=; h=Received:Received:Message-ID:Date:From:MIME-Version:To:Subject: Content-Type; b=egeUY3e2C6Grf/HiYvDsLvpzWE9bpRV1hp87/TN6ymd8zXELynX2Z5BLVXd2WaVfc H0IVtimHSjeGLMpUnGqcfgK2QpoG71MRzjG+iNOPWgzqrs3ApPV97V7DfNxKJh4P+x ZB0YrC88hmPxNhTbNGVzMFL0+ITWc0bGZGwHjTevzrz6jxL1koj/8iL5uS2iysIZWq Iuqaj4WePLfFGI2VZv8eGSzZp5drAm9smYpIW/aKw8EPRoe8WmxIj79RMyUXQMDLmf 3pNiWY7uo5dVzzImT3HD4TNvh9JE3KYSrdx8Pqf/d/Dlb6Z+iH5K7y5wRY6ivxASCL IG/UWNG1xu+aA==
Archived-At: http://mailarchive.ietf.org/arch/msg/clue/F7yTh0hD0i3mCFbtSFxqRE-qBkY
Subject: [clue] Requirements are now an RFC!!
X-BeenThere: clue@ietf.org
X-Mailman-Version: 2.1.15
Precedence: list
List-Id: CLUE - ControLling mUltiple streams for TElepresence <clue.ietf.org>
List-Unsubscribe: <https://www.ietf.org/mailman/options/clue>, <mailto:clue-request@ietf.org?subject=unsubscribe>
List-Archive: <http://www.ietf.org/mail-archive/web/clue/>
List-Post: <mailto:clue@ietf.org>
List-Help: <mailto:clue-request@ietf.org?subject=help>
List-Subscribe: <https://www.ietf.org/mailman/listinfo/clue>, <mailto:clue-request@ietf.org?subject=subscribe>
X-List-Received-Date: Fri, 18 Apr 2014 02:56:06 -0000

On 4/17/14 7:03 PM, rfc-editor@rfc-editor.org wrote:
> A new Request for Comments is now available in online RFC libraries.
>
>
>          RFC 7205
>
>          Title:      Use Cases for Telepresence Multistreams
>          Author:     A. Romanow, S. Botzko,
>                      M. Duckworth, R. Even, Ed.
>          Status:     Informational
>          Stream:     IETF
>          Date:       April 2014
>          Mailbox:    allyn@cisco.com,
>                      stephen.botzko@polycom.com,
>                      mark.duckworth@polycom.com,
>                      roni.even@mail01.huawei.com
>          Pages:      17
>          Characters: 42087
>          Updates/Obsoletes/SeeAlso:   None
>
>          I-D Tag:    draft-ietf-clue-telepresence-use-cases-09.txt
>
>          URL:        http://www.rfc-editor.org/rfc/rfc7205.txt
>
> Telepresence conferencing systems seek to create an environment that
> gives users (or user groups) that are not co-located a feeling of
> co-located presence through multimedia communication that includes at
> least audio and video signals of high fidelity.  A number of
> techniques for handling audio and video streams are used to create
> this experience.  When these techniques are not similar,
> interoperability between different systems is difficult at best, and
> often not possible.  Conveying information about the relationships
> between multiple streams of media would enable senders and receivers
> to make choices to allow telepresence systems to interwork.  This
> memo describes the most typical and important use cases for sending
> multiple streams in a telepresence conference.
>
> This document is a product of the ControLling mUltiple streams for tElepresence Working Group of the IETF.
>
>
> INFORMATIONAL: This memo provides information for the Internet community.
> It does not specify an Internet standard of any kind. Distribution of
> this memo is unlimited.
>
> This announcement is sent to the IETF-Announce and rfc-dist lists.
> To subscribe or unsubscribe, see
>    http://www.ietf.org/mailman/listinfo/ietf-announce
>    http://mailman.rfc-editor.org/mailman/listinfo/rfc-dist
>
> For searching the RFC series, see http://www.rfc-editor.org/search
> For downloading RFCs, see http://www.rfc-editor.org/rfc.html
>
> Requests for special distribution should be addressed to either the
> author of the RFC in question, or to rfc-editor@rfc-editor.org.  Unless
> specifically noted otherwise on the RFC itself, all RFCs are for
> unlimited distribution.
>
>
> The RFC Editor Team
> Association Management Solutions, LLC
>
>
> _______________________________________________
> clue mailing list
> clue@ietf.org
> https://www.ietf.org/mailman/listinfo/clue
>


From nobody Fri Apr 18 05:50:19 2014
Return-Path: <mary.ietf.barnes@gmail.com>
X-Original-To: clue@ietfa.amsl.com
Delivered-To: clue@ietfa.amsl.com
Received: from localhost (ietfa.amsl.com [127.0.0.1]) by ietfa.amsl.com (Postfix) with ESMTP id 1AD691A01D4 for <clue@ietfa.amsl.com>; Fri, 18 Apr 2014 05:50:18 -0700 (PDT)
X-Virus-Scanned: amavisd-new at amsl.com
X-Spam-Flag: NO
X-Spam-Score: -1.999
X-Spam-Level: 
X-Spam-Status: No, score=-1.999 tagged_above=-999 required=5 tests=[BAYES_00=-1.9, DKIM_SIGNED=0.1, DKIM_VALID=-0.1, DKIM_VALID_AU=-0.1, FREEMAIL_FROM=0.001, HTML_MESSAGE=0.001, SPF_PASS=-0.001] autolearn=ham
Received: from mail.ietf.org ([4.31.198.44]) by localhost (ietfa.amsl.com [127.0.0.1]) (amavisd-new, port 10024) with ESMTP id 7qmOyx8cR909 for <clue@ietfa.amsl.com>; Fri, 18 Apr 2014 05:50:13 -0700 (PDT)
Received: from mail-we0-x235.google.com (mail-we0-x235.google.com [IPv6:2a00:1450:400c:c03::235]) by ietfa.amsl.com (Postfix) with ESMTP id A78511A01BB for <clue@ietf.org>; Fri, 18 Apr 2014 05:50:12 -0700 (PDT)
Received: by mail-we0-f181.google.com with SMTP id q58so1515805wes.40 for <clue@ietf.org>; Fri, 18 Apr 2014 05:50:08 -0700 (PDT)
DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20120113;  h=mime-version:in-reply-to:references:date:message-id:subject:from:to :cc:content-type; bh=q+BoCpew97hMxx59/MPMxDJEGv7WTbpkfHYNjHc0gG4=; b=dryTx0A75nfbbMCSvjKdJiOAqK1Uxeu53lYaavbh96+1eTS6Uj1gcPtkwEl88R1S0W Yjr65vQRYaPQNz44Jn5d+CQhOiRDC4s/yNZQJgy6R4eovw3Z69dQGs2OdBKePNVdRzmz Pusjh/gJ2VGuV13KvjdspIYCUGj9FWmF6bKcxi1nJv62gQOVqAE72QduAGpSFe+7DYTZ y8LaRySBAAg9SmexUqUmpwRt3oYSm19YMYyGQUqJDQ5HZoICq1g0g0UZy3koSa+9tqip zrUWBZR11jRlA/g+8TY5hPXMNIbkuK2zZ1bnOxYv9lQ2CoPwDXpOjdQUrqhND+lxGmve VdBw==
MIME-Version: 1.0
X-Received: by 10.194.86.7 with SMTP id l7mr16099194wjz.37.1397825408275; Fri, 18 Apr 2014 05:50:08 -0700 (PDT)
Received: by 10.216.10.6 with HTTP; Fri, 18 Apr 2014 05:50:08 -0700 (PDT)
In-Reply-To: <5350943D.6080600@alum.mit.edu>
References: <20140417230301.5294A18000C@rfc-editor.org> <5350943D.6080600@alum.mit.edu>
Date: Fri, 18 Apr 2014 07:50:08 -0500
Message-ID: <CAHBDyN4WiUuGPTwOJaECLB-9bO27XAiBaFxr8LL=00h8zyDCTw@mail.gmail.com>
From: Mary Barnes <mary.ietf.barnes@gmail.com>
To: Paul Kyzivat <pkyzivat@alum.mit.edu>
Content-Type: multipart/alternative; boundary=089e0102e4f0eb1d0a04f7509693
Archived-At: http://mailarchive.ietf.org/arch/msg/clue/ruSmoLmUvbOcDF7Hl5uUyPjc-dI
Cc: CLUE <clue@ietf.org>
Subject: Re: [clue] Requirements are now an RFC!!
X-BeenThere: clue@ietf.org
X-Mailman-Version: 2.1.15
Precedence: list
List-Id: CLUE - ControLling mUltiple streams for TElepresence <clue.ietf.org>
List-Unsubscribe: <https://www.ietf.org/mailman/options/clue>, <mailto:clue-request@ietf.org?subject=unsubscribe>
List-Archive: <http://www.ietf.org/mail-archive/web/clue/>
List-Post: <mailto:clue@ietf.org>
List-Help: <mailto:clue-request@ietf.org?subject=help>
List-Subscribe: <https://www.ietf.org/mailman/listinfo/clue>, <mailto:clue-request@ietf.org?subject=subscribe>
X-List-Received-Date: Fri, 18 Apr 2014 12:50:18 -0000

--089e0102e4f0eb1d0a04f7509693
Content-Type: text/plain; charset=UTF-8

No. It's the use cases that have just been published. The requirements
document is still in the RFC Editor's queue.


On Thu, Apr 17, 2014 at 9:55 PM, Paul Kyzivat <pkyzivat@alum.mit.edu> wrote:

> On 4/17/14 7:03 PM, rfc-editor@rfc-editor.org wrote:
>
>> A new Request for Comments is now available in online RFC libraries.
>>
>>
>>          RFC 7205
>>
>>          Title:      Use Cases for Telepresence Multistreams
>>          Author:     A. Romanow, S. Botzko,
>>                      M. Duckworth, R. Even, Ed.
>>          Status:     Informational
>>          Stream:     IETF
>>          Date:       April 2014
>>          Mailbox:    allyn@cisco.com,
>>                      stephen.botzko@polycom.com,
>>                      mark.duckworth@polycom.com,
>>                      roni.even@mail01.huawei.com
>>          Pages:      17
>>          Characters: 42087
>>          Updates/Obsoletes/SeeAlso:   None
>>
>>          I-D Tag:    draft-ietf-clue-telepresence-use-cases-09.txt
>>
>>          URL:        http://www.rfc-editor.org/rfc/rfc7205.txt
>>
>> Telepresence conferencing systems seek to create an environment that
>> gives users (or user groups) that are not co-located a feeling of
>> co-located presence through multimedia communication that includes at
>> least audio and video signals of high fidelity.  A number of
>> techniques for handling audio and video streams are used to create
>> this experience.  When these techniques are not similar,
>> interoperability between different systems is difficult at best, and
>> often not possible.  Conveying information about the relationships
>> between multiple streams of media would enable senders and receivers
>> to make choices to allow telepresence systems to interwork.  This
>> memo describes the most typical and important use cases for sending
>> multiple streams in a telepresence conference.
>>
>> This document is a product of the ControLling mUltiple streams for
>> tElepresence Working Group of the IETF.
>>
>>
>> INFORMATIONAL: This memo provides information for the Internet community.
>> It does not specify an Internet standard of any kind. Distribution of
>> this memo is unlimited.
>>
>> This announcement is sent to the IETF-Announce and rfc-dist lists.
>> To subscribe or unsubscribe, see
>>    http://www.ietf.org/mailman/listinfo/ietf-announce
>>    http://mailman.rfc-editor.org/mailman/listinfo/rfc-dist
>>
>> For searching the RFC series, see http://www.rfc-editor.org/search
>> For downloading RFCs, see http://www.rfc-editor.org/rfc.html
>>
>> Requests for special distribution should be addressed to either the
>> author of the RFC in question, or to rfc-editor@rfc-editor.org.  Unless
>> specifically noted otherwise on the RFC itself, all RFCs are for
>> unlimited distribution.
>>
>>
>> The RFC Editor Team
>> Association Management Solutions, LLC
>>
>>
>> _______________________________________________
>> clue mailing list
>> clue@ietf.org
>> https://www.ietf.org/mailman/listinfo/clue
>>
>>
> _______________________________________________
> clue mailing list
> clue@ietf.org
> https://www.ietf.org/mailman/listinfo/clue
>

--089e0102e4f0eb1d0a04f7509693
Content-Type: text/html; charset=UTF-8
Content-Transfer-Encoding: quoted-printable

<div dir=3D"ltr">No. It&#39;s the use cases that have just been published. =
The requirements document is still in the RFC Editor&#39;s queue.</div><div=
 class=3D"gmail_extra"><br><br><div class=3D"gmail_quote">On Thu, Apr 17, 2=
014 at 9:55 PM, Paul Kyzivat <span dir=3D"ltr">&lt;<a href=3D"mailto:pkyziv=
at@alum.mit.edu" target=3D"_blank">pkyzivat@alum.mit.edu</a>&gt;</span> wro=
te:<br>
<blockquote class=3D"gmail_quote" style=3D"margin:0 0 0 .8ex;border-left:1p=
x #ccc solid;padding-left:1ex">On 4/17/14 7:03 PM, <a href=3D"mailto:rfc-ed=
itor@rfc-editor.org" target=3D"_blank">rfc-editor@rfc-editor.org</a> wrote:=
<br>

<blockquote class=3D"gmail_quote" style=3D"margin:0 0 0 .8ex;border-left:1p=
x #ccc solid;padding-left:1ex">
A new Request for Comments is now available in online RFC libraries.<br>
<br>
<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0RFC 7205<br>
<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0Title: =C2=A0 =C2=A0 =C2=A0Use Cases for =
Telepresence Multistreams<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0Author: =C2=A0 =C2=A0 A. Romanow, S. Botz=
ko,<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=
=A0M. Duckworth, R. Even, Ed.<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0Status: =C2=A0 =C2=A0 Informational<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0Stream: =C2=A0 =C2=A0 IETF<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0Date: =C2=A0 =C2=A0 =C2=A0 April 2014<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0Mailbox: =C2=A0 =C2=A0<a href=3D"mailto:a=
llyn@cisco.com" target=3D"_blank">allyn@cisco.com</a>,<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=
=A0<a href=3D"mailto:stephen.botzko@polycom.com" target=3D"_blank">stephen.=
botzko@polycom.com</a>,<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=
=A0<a href=3D"mailto:mark.duckworth@polycom.com" target=3D"_blank">mark.duc=
kworth@polycom.com</a>,<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=
=A0<a href=3D"mailto:roni.even@mail01.huawei.com" target=3D"_blank">roni.ev=
en@mail01.huawei.com</a><br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0Pages: =C2=A0 =C2=A0 =C2=A017<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0Characters: 42087<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0Updates/Obsoletes/SeeAlso: =C2=A0 None<br=
>
<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0I-D Tag: =C2=A0 =C2=A0draft-ietf-clue-tel=
epresence-<u></u>use-cases-09.txt<br>
<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0URL: =C2=A0 =C2=A0 =C2=A0 =C2=A0<a href=
=3D"http://www.rfc-editor.org/rfc/rfc7205.txt" target=3D"_blank">http://www=
.rfc-editor.org/rfc/<u></u>rfc7205.txt</a><br>
<br>
Telepresence conferencing systems seek to create an environment that<br>
gives users (or user groups) that are not co-located a feeling of<br>
co-located presence through multimedia communication that includes at<br>
least audio and video signals of high fidelity. =C2=A0A number of<br>
techniques for handling audio and video streams are used to create<br>
this experience. =C2=A0When these techniques are not similar,<br>
interoperability between different systems is difficult at best, and<br>
often not possible. =C2=A0Conveying information about the relationships<br>
between multiple streams of media would enable senders and receivers<br>
to make choices to allow telepresence systems to interwork. =C2=A0This<br>
memo describes the most typical and important use cases for sending<br>
multiple streams in a telepresence conference.<br>
<br>
This document is a product of the ControLling mUltiple streams for tElepres=
ence Working Group of the IETF.<br>
<br>
<br>
INFORMATIONAL: This memo provides information for the Internet community.<b=
r>
It does not specify an Internet standard of any kind. Distribution of<br>
this memo is unlimited.<br>
<br>
This announcement is sent to the IETF-Announce and rfc-dist lists.<br>
To subscribe or unsubscribe, see<br>
=C2=A0 =C2=A0<a href=3D"http://www.ietf.org/mailman/listinfo/ietf-announce"=
 target=3D"_blank">http://www.ietf.org/mailman/<u></u>listinfo/ietf-announc=
e</a><br>
=C2=A0 =C2=A0<a href=3D"http://mailman.rfc-editor.org/mailman/listinfo/rfc-=
dist" target=3D"_blank">http://mailman.rfc-editor.org/<u></u>mailman/listin=
fo/rfc-dist</a><br>
<br>
For searching the RFC series, see <a href=3D"http://www.rfc-editor.org/sear=
ch" target=3D"_blank">http://www.rfc-editor.org/<u></u>search</a><br>
For downloading RFCs, see <a href=3D"http://www.rfc-editor.org/rfc.html" ta=
rget=3D"_blank">http://www.rfc-editor.org/rfc.<u></u>html</a><br>
<br>
Requests for special distribution should be addressed to either the<br>
author of the RFC in question, or to <a href=3D"mailto:rfc-editor@rfc-edito=
r.org" target=3D"_blank">rfc-editor@rfc-editor.org</a>. =C2=A0Unless<br>
specifically noted otherwise on the RFC itself, all RFCs are for<br>
unlimited distribution.<br>
<br>
<br>
The RFC Editor Team<br>
Association Management Solutions, LLC<br>
<br>
<br>
______________________________<u></u>_________________<br>
clue mailing list<br>
<a href=3D"mailto:clue@ietf.org" target=3D"_blank">clue@ietf.org</a><br>
<a href=3D"https://www.ietf.org/mailman/listinfo/clue" target=3D"_blank">ht=
tps://www.ietf.org/mailman/<u></u>listinfo/clue</a><br>
<br>
</blockquote>
<br>
______________________________<u></u>_________________<br>
clue mailing list<br>
<a href=3D"mailto:clue@ietf.org" target=3D"_blank">clue@ietf.org</a><br>
<a href=3D"https://www.ietf.org/mailman/listinfo/clue" target=3D"_blank">ht=
tps://www.ietf.org/mailman/<u></u>listinfo/clue</a><br>
</blockquote></div><br></div>

--089e0102e4f0eb1d0a04f7509693--


From nobody Fri Apr 18 07:58:14 2014
Return-Path: <pkyzivat@alum.mit.edu>
X-Original-To: clue@ietfa.amsl.com
Delivered-To: clue@ietfa.amsl.com
Received: from localhost (ietfa.amsl.com [127.0.0.1]) by ietfa.amsl.com (Postfix) with ESMTP id CA9D91A02FB for <clue@ietfa.amsl.com>; Fri, 18 Apr 2014 07:58:11 -0700 (PDT)
X-Virus-Scanned: amavisd-new at amsl.com
X-Spam-Flag: NO
X-Spam-Score: -1.235
X-Spam-Level: 
X-Spam-Status: No, score=-1.235 tagged_above=-999 required=5 tests=[BAYES_00=-1.9, DKIM_SIGNED=0.1, DKIM_VALID=-0.1, SPF_SOFTFAIL=0.665] autolearn=no
Received: from mail.ietf.org ([4.31.198.44]) by localhost (ietfa.amsl.com [127.0.0.1]) (amavisd-new, port 10024) with ESMTP id islDjR2kKaeI for <clue@ietfa.amsl.com>; Fri, 18 Apr 2014 07:58:07 -0700 (PDT)
Received: from qmta15.westchester.pa.mail.comcast.net (qmta15.westchester.pa.mail.comcast.net [IPv6:2001:558:fe14:44:76:96:59:228]) by ietfa.amsl.com (Postfix) with ESMTP id 36ED71A01AF for <clue@ietf.org>; Fri, 18 Apr 2014 07:58:07 -0700 (PDT)
Received: from omta20.westchester.pa.mail.comcast.net ([76.96.62.71]) by qmta15.westchester.pa.mail.comcast.net with comcast id rPeE1n0071YDfWL5FSy32g; Fri, 18 Apr 2014 14:58:03 +0000
Received: from Paul-Kyzivats-MacBook-Pro.local ([50.138.229.164]) by omta20.westchester.pa.mail.comcast.net with comcast id rSy21n01A3ZTu2S3gSy2cK; Fri, 18 Apr 2014 14:58:03 +0000
Message-ID: <53513D7A.3080204@alum.mit.edu>
Date: Fri, 18 Apr 2014 10:58:02 -0400
From: Paul Kyzivat <pkyzivat@alum.mit.edu>
User-Agent: Mozilla/5.0 (Macintosh; Intel Mac OS X 10.7; rv:24.0) Gecko/20100101 Thunderbird/24.4.0
MIME-Version: 1.0
To: Mary Barnes <mary.ietf.barnes@gmail.com>
References: <20140417230301.5294A18000C@rfc-editor.org>	<5350943D.6080600@alum.mit.edu> <CAHBDyN4WiUuGPTwOJaECLB-9bO27XAiBaFxr8LL=00h8zyDCTw@mail.gmail.com>
In-Reply-To: <CAHBDyN4WiUuGPTwOJaECLB-9bO27XAiBaFxr8LL=00h8zyDCTw@mail.gmail.com>
Content-Type: text/plain; charset=UTF-8; format=flowed
Content-Transfer-Encoding: 7bit
DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=comcast.net; s=q20140121; t=1397833083; bh=viLiyiWvxZUGofDgY0+sdPsXQwIbs3Uql8VRXehKD4o=; h=Received:Received:Message-ID:Date:From:MIME-Version:To:Subject: Content-Type; b=FM5NNmrV3k+6E6rwGi9X58jBkH59cyyK5QLROlyjDu9hsN00B++ers7sYI9Ee2Xut +diWEJGUKTgdLJHa6HiXb4dWjExyvplWc9B//SQ3gKIsb6McJ5lqj9WbhAB6Ucz8UM AV8Uibzyt+JC0N0VoeQGxqvGRXCGXTftUkJ+weOd9p211E7MGjknCeRmmdhKxl8Jz1 Gq1oNOak7iXE0Lx+uEO6WWVcEQAkFVRNKIJZWVAnfE7uZNrROWS2CKC3t4AEskO9o7 t/nlrZ2DpOC5J5n7YD8c51Ji+kWo+lhhCPDD0KwK/r6TbRPiT9Sl5ddzaGvdFjeIat +nAK+VpVgcheQ==
Archived-At: http://mailarchive.ietf.org/arch/msg/clue/3uKdCYpye3pzmO5XuhV4RV-9WIw
Cc: CLUE <clue@ietf.org>
Subject: Re: [clue] Requirements are now an RFC!!
X-BeenThere: clue@ietf.org
X-Mailman-Version: 2.1.15
Precedence: list
List-Id: CLUE - ControLling mUltiple streams for TElepresence <clue.ietf.org>
List-Unsubscribe: <https://www.ietf.org/mailman/options/clue>, <mailto:clue-request@ietf.org?subject=unsubscribe>
List-Archive: <http://www.ietf.org/mail-archive/web/clue/>
List-Post: <mailto:clue@ietf.org>
List-Help: <mailto:clue-request@ietf.org?subject=help>
List-Subscribe: <https://www.ietf.org/mailman/listinfo/clue>, <mailto:clue-request@ietf.org?subject=subscribe>
X-List-Received-Date: Fri, 18 Apr 2014 14:58:12 -0000

Oops! Sorry. :-(

On 4/18/14 8:50 AM, Mary Barnes wrote:
> No. It's the use cases that have just been published. The requirements
> document is still in the RFC Editor's queue.
>
>
> On Thu, Apr 17, 2014 at 9:55 PM, Paul Kyzivat <pkyzivat@alum.mit.edu
> <mailto:pkyzivat@alum.mit.edu>> wrote:
>
>     On 4/17/14 7:03 PM, rfc-editor@rfc-editor.org
>     <mailto:rfc-editor@rfc-editor.org> wrote:
>
>         A new Request for Comments is now available in online RFC libraries.
>
>
>                   RFC 7205
>
>                   Title:      Use Cases for Telepresence Multistreams
>                   Author:     A. Romanow, S. Botzko,
>                               M. Duckworth, R. Even, Ed.
>                   Status:     Informational
>                   Stream:     IETF
>                   Date:       April 2014
>                   Mailbox: allyn@cisco.com <mailto:allyn@cisco.com>,
>         stephen.botzko@polycom.com <mailto:stephen.botzko@polycom.com>,
>         mark.duckworth@polycom.com <mailto:mark.duckworth@polycom.com>,
>         roni.even@mail01.huawei.com <mailto:roni.even@mail01.huawei.com>
>                   Pages:      17
>                   Characters: 42087
>                   Updates/Obsoletes/SeeAlso:   None
>
>                   I-D Tag:
>           draft-ietf-clue-telepresence-__use-cases-09.txt
>
>                   URL: http://www.rfc-editor.org/rfc/__rfc7205.txt
>         <http://www.rfc-editor.org/rfc/rfc7205.txt>
>
>         Telepresence conferencing systems seek to create an environment that
>         gives users (or user groups) that are not co-located a feeling of
>         co-located presence through multimedia communication that
>         includes at
>         least audio and video signals of high fidelity.  A number of
>         techniques for handling audio and video streams are used to create
>         this experience.  When these techniques are not similar,
>         interoperability between different systems is difficult at best, and
>         often not possible.  Conveying information about the relationships
>         between multiple streams of media would enable senders and receivers
>         to make choices to allow telepresence systems to interwork.  This
>         memo describes the most typical and important use cases for sending
>         multiple streams in a telepresence conference.
>
>         This document is a product of the ControLling mUltiple streams
>         for tElepresence Working Group of the IETF.
>
>
>         INFORMATIONAL: This memo provides information for the Internet
>         community.
>         It does not specify an Internet standard of any kind.
>         Distribution of
>         this memo is unlimited.
>
>         This announcement is sent to the IETF-Announce and rfc-dist lists.
>         To subscribe or unsubscribe, see
>         http://www.ietf.org/mailman/__listinfo/ietf-announce
>         <http://www.ietf.org/mailman/listinfo/ietf-announce>
>         http://mailman.rfc-editor.org/__mailman/listinfo/rfc-dist
>         <http://mailman.rfc-editor.org/mailman/listinfo/rfc-dist>
>
>         For searching the RFC series, see
>         http://www.rfc-editor.org/__search
>         <http://www.rfc-editor.org/search>
>         For downloading RFCs, see http://www.rfc-editor.org/rfc.__html
>         <http://www.rfc-editor.org/rfc.html>
>
>         Requests for special distribution should be addressed to either the
>         author of the RFC in question, or to rfc-editor@rfc-editor.org
>         <mailto:rfc-editor@rfc-editor.org>.  Unless
>         specifically noted otherwise on the RFC itself, all RFCs are for
>         unlimited distribution.
>
>
>         The RFC Editor Team
>         Association Management Solutions, LLC
>
>
>         _________________________________________________
>         clue mailing list
>         clue@ietf.org <mailto:clue@ietf.org>
>         https://www.ietf.org/mailman/__listinfo/clue
>         <https://www.ietf.org/mailman/listinfo/clue>
>
>
>     _________________________________________________
>     clue mailing list
>     clue@ietf.org <mailto:clue@ietf.org>
>     https://www.ietf.org/mailman/__listinfo/clue
>     <https://www.ietf.org/mailman/listinfo/clue>
>
>


From nobody Mon Apr 21 17:25:40 2014
Return-Path: <pkyzivat@alum.mit.edu>
X-Original-To: clue@ietfa.amsl.com
Delivered-To: clue@ietfa.amsl.com
Received: from localhost (ietfa.amsl.com [127.0.0.1]) by ietfa.amsl.com (Postfix) with ESMTP id 3BA371A0321 for <clue@ietfa.amsl.com>; Mon, 21 Apr 2014 17:25:38 -0700 (PDT)
X-Virus-Scanned: amavisd-new at amsl.com
X-Spam-Flag: NO
X-Spam-Score: -1.235
X-Spam-Level: 
X-Spam-Status: No, score=-1.235 tagged_above=-999 required=5 tests=[BAYES_00=-1.9, DKIM_SIGNED=0.1, DKIM_VALID=-0.1, SPF_SOFTFAIL=0.665] autolearn=no
Received: from mail.ietf.org ([4.31.198.44]) by localhost (ietfa.amsl.com [127.0.0.1]) (amavisd-new, port 10024) with ESMTP id l0O4gCCyFaGX for <clue@ietfa.amsl.com>; Mon, 21 Apr 2014 17:25:37 -0700 (PDT)
Received: from qmta10.westchester.pa.mail.comcast.net (qmta10.westchester.pa.mail.comcast.net [IPv6:2001:558:fe14:43:76:96:62:17]) by ietfa.amsl.com (Postfix) with ESMTP id 135781A00C0 for <clue@ietf.org>; Mon, 21 Apr 2014 17:25:36 -0700 (PDT)
Received: from omta02.westchester.pa.mail.comcast.net ([76.96.62.19]) by qmta10.westchester.pa.mail.comcast.net with comcast id snps1n0020QuhwU5AoRXuB; Tue, 22 Apr 2014 00:25:31 +0000
Received: from Paul-Kyzivats-MacBook-Pro.local ([50.138.229.164]) by omta02.westchester.pa.mail.comcast.net with comcast id soRX1n00G3ZTu2S3NoRXoU; Tue, 22 Apr 2014 00:25:31 +0000
Message-ID: <5355B6FB.6050501@alum.mit.edu>
Date: Mon, 21 Apr 2014 20:25:31 -0400
From: Paul Kyzivat <pkyzivat@alum.mit.edu>
User-Agent: Mozilla/5.0 (Macintosh; Intel Mac OS X 10.7; rv:24.0) Gecko/20100101 Thunderbird/24.4.0
MIME-Version: 1.0
To: CLUE <clue@ietf.org>
Content-Type: text/plain; charset=ISO-8859-1; format=flowed
Content-Transfer-Encoding: 7bit
DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=comcast.net; s=q20140121; t=1398126331; bh=SXVjzblyz9aIwnRFZgTEqcMDkCHK141FAuFv1LHIg+o=; h=Received:Received:Message-ID:Date:From:MIME-Version:To:Subject: Content-Type; b=N7Z4OEdNmppa/7oZZOp8BkCPSYJQ2JGDcenH0deQ9XBZjvted9jHNfXW9ntQEahrp 20291G91JNF1Mt5Sx/dWH5rad3CDaaFSC/XoFzhs/MgrLqitU6eKq3jjPNR+zc+RsE SfOS9gND0rAW3agWzFAKwcem8lsyJnygMX459UCfZMSYYILmC9poYWgm7OqZJhD8Az v2qYcQcAAqi1d7bpMXpHs+frhagdymYMR6zHv0+ismp6G5nJ/l+z6ZaSnnPBNPVVPO Mfg0ZFOEfMjV+KxOLic+pvdNIuLELqTpJzzM9OC+AXndlaDVYz6QdcOkhAynjwckzn v7Q46NYOZ50rQ==
Archived-At: http://mailarchive.ietf.org/arch/msg/clue/-yIuUhqEhZi_zBEPgXQ3FlhOnCg
Subject: [clue] Reminder: design team meeting tomorrow Tuesday, Apr 22
X-BeenThere: clue@ietf.org
X-Mailman-Version: 2.1.15
Precedence: list
List-Id: CLUE - ControLling mUltiple streams for TElepresence <clue.ietf.org>
List-Unsubscribe: <https://www.ietf.org/mailman/options/clue>, <mailto:clue-request@ietf.org?subject=unsubscribe>
List-Archive: <http://www.ietf.org/mail-archive/web/clue/>
List-Post: <mailto:clue@ietf.org>
List-Help: <mailto:clue-request@ietf.org?subject=help>
List-Subscribe: <https://www.ietf.org/mailman/listinfo/clue>, <mailto:clue-request@ietf.org?subject=subscribe>
X-List-Received-Date: Tue, 22 Apr 2014 00:25:38 -0000

This is just a reminder. We will be having the design team meeting 
tomorrow. The primary subject will be the recent revision of the 
protocol document.

	Thanks,
	Paul


From nobody Mon Apr 21 17:45:05 2014
Return-Path: <Christian.Groves@nteczone.com>
X-Original-To: clue@ietfa.amsl.com
Delivered-To: clue@ietfa.amsl.com
Received: from localhost (ietfa.amsl.com [127.0.0.1]) by ietfa.amsl.com (Postfix) with ESMTP id 4BE221A0325 for <clue@ietfa.amsl.com>; Mon, 21 Apr 2014 17:45:04 -0700 (PDT)
X-Virus-Scanned: amavisd-new at amsl.com
X-Spam-Flag: NO
X-Spam-Score: 0.8
X-Spam-Level: 
X-Spam-Status: No, score=0.8 tagged_above=-999 required=5 tests=[BAYES_50=0.8] autolearn=ham
Received: from mail.ietf.org ([4.31.198.44]) by localhost (ietfa.amsl.com [127.0.0.1]) (amavisd-new, port 10024) with ESMTP id ZopA41rPoT_R for <clue@ietfa.amsl.com>; Mon, 21 Apr 2014 17:45:03 -0700 (PDT)
Received: from cserver5.myshophosting.com (cserver5.myshophosting.com [175.107.161.1]) by ietfa.amsl.com (Postfix) with ESMTP id 1FC111A032E for <clue@ietf.org>; Mon, 21 Apr 2014 17:44:57 -0700 (PDT)
Received: from ppp118-209-25-196.lns20.mel4.internode.on.net ([118.209.25.196]:51840 helo=[127.0.0.1]) by cserver5.myshophosting.com with esmtpsa (TLSv1:DHE-RSA-AES128-SHA:128) (Exim 4.82) (envelope-from <Christian.Groves@nteczone.com>) id 1WcOpL-0001Mh-A2 for clue@ietf.org; Tue, 22 Apr 2014 10:44:47 +1000
Message-ID: <5355BB80.7010704@nteczone.com>
Date: Tue, 22 Apr 2014 10:44:48 +1000
From: Christian Groves <Christian.Groves@nteczone.com>
User-Agent: Mozilla/5.0 (Windows NT 6.1; WOW64; rv:24.0) Gecko/20100101 Thunderbird/24.4.0
MIME-Version: 1.0
To: clue@ietf.org
References: <5355B6FB.6050501@alum.mit.edu>
In-Reply-To: <5355B6FB.6050501@alum.mit.edu>
Content-Type: text/plain; charset=ISO-8859-1; format=flowed
Content-Transfer-Encoding: 7bit
X-AntiAbuse: This header was added to track abuse, please include it with any abuse report
X-AntiAbuse: Primary Hostname - cserver5.myshophosting.com
X-AntiAbuse: Original Domain - ietf.org
X-AntiAbuse: Originator/Caller UID/GID - [47 12] / [47 12]
X-AntiAbuse: Sender Address Domain - nteczone.com
X-Get-Message-Sender-Via: cserver5.myshophosting.com: authenticated_id: christian.groves@nteczone.com
X-Source: 
X-Source-Args: 
X-Source-Dir: 
Archived-At: http://mailarchive.ietf.org/arch/msg/clue/lD5DFyKlVjPn_8IBC55W7Yt_pQc
Subject: Re: [clue] Reminder: design team meeting tomorrow Tuesday, Apr 22
X-BeenThere: clue@ietf.org
X-Mailman-Version: 2.1.15
Precedence: list
List-Id: CLUE - ControLling mUltiple streams for TElepresence <clue.ietf.org>
List-Unsubscribe: <https://www.ietf.org/mailman/options/clue>, <mailto:clue-request@ietf.org?subject=unsubscribe>
List-Archive: <http://www.ietf.org/mail-archive/web/clue/>
List-Post: <mailto:clue@ietf.org>
List-Help: <mailto:clue-request@ietf.org?subject=help>
List-Subscribe: <https://www.ietf.org/mailman/listinfo/clue>, <mailto:clue-request@ietf.org?subject=subscribe>
X-List-Received-Date: Tue, 22 Apr 2014 00:45:04 -0000

Hello Paul,

What time is it? Is there any change as a result of the doodle poll?

Christian

On 22/04/2014 10:25 AM, Paul Kyzivat wrote:
> This is just a reminder. We will be having the design team meeting 
> tomorrow. The primary subject will be the recent revision of the 
> protocol document.
>
>     Thanks,
>     Paul
>
> _______________________________________________
> clue mailing list
> clue@ietf.org
> https://www.ietf.org/mailman/listinfo/clue
>


From nobody Mon Apr 21 19:40:37 2014
Return-Path: <pkyzivat@alum.mit.edu>
X-Original-To: clue@ietfa.amsl.com
Delivered-To: clue@ietfa.amsl.com
Received: from localhost (ietfa.amsl.com [127.0.0.1]) by ietfa.amsl.com (Postfix) with ESMTP id 542871A000B for <clue@ietfa.amsl.com>; Mon, 21 Apr 2014 19:40:36 -0700 (PDT)
X-Virus-Scanned: amavisd-new at amsl.com
X-Spam-Flag: NO
X-Spam-Score: -1.235
X-Spam-Level: 
X-Spam-Status: No, score=-1.235 tagged_above=-999 required=5 tests=[BAYES_00=-1.9, DKIM_SIGNED=0.1, DKIM_VALID=-0.1, SPF_SOFTFAIL=0.665] autolearn=no
Received: from mail.ietf.org ([4.31.198.44]) by localhost (ietfa.amsl.com [127.0.0.1]) (amavisd-new, port 10024) with ESMTP id 15JAxFRniTYD for <clue@ietfa.amsl.com>; Mon, 21 Apr 2014 19:40:35 -0700 (PDT)
Received: from qmta14.westchester.pa.mail.comcast.net (qmta14.westchester.pa.mail.comcast.net [IPv6:2001:558:fe14:44:76:96:59:212]) by ietfa.amsl.com (Postfix) with ESMTP id 29AE01A0008 for <clue@ietf.org>; Mon, 21 Apr 2014 19:40:35 -0700 (PDT)
Received: from omta22.westchester.pa.mail.comcast.net ([76.96.62.73]) by qmta14.westchester.pa.mail.comcast.net with comcast id sqXH1n0021ap0As5EqgVP7; Tue, 22 Apr 2014 02:40:29 +0000
Received: from Paul-Kyzivats-MacBook-Pro.local ([50.138.229.164]) by omta22.westchester.pa.mail.comcast.net with comcast id sqgV1n00W3ZTu2S3iqgVqd; Tue, 22 Apr 2014 02:40:29 +0000
Message-ID: <5355D69D.1040603@alum.mit.edu>
Date: Mon, 21 Apr 2014 22:40:29 -0400
From: Paul Kyzivat <pkyzivat@alum.mit.edu>
User-Agent: Mozilla/5.0 (Macintosh; Intel Mac OS X 10.7; rv:24.0) Gecko/20100101 Thunderbird/24.4.0
MIME-Version: 1.0
To: clue@ietf.org
References: <5355B6FB.6050501@alum.mit.edu> <5355BB80.7010704@nteczone.com>
In-Reply-To: <5355BB80.7010704@nteczone.com>
Content-Type: text/plain; charset=ISO-8859-1; format=flowed
Content-Transfer-Encoding: 7bit
DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=comcast.net; s=q20140121; t=1398134429; bh=DAVNC/ChgAFzEVEYphFaxBVCsr1Kpy9TCypm5jKqlRA=; h=Received:Received:Message-ID:Date:From:MIME-Version:To:Subject: Content-Type; b=naTXWPCETDPRYfF51WxqgmppKkgBN3kc6LMU4GLTvXOH+9itDjKlzkm6AenbW1W8i a8VUFjDyHpPBbr5JIlKCeSGfjlrnV3ztANH9x0j+qSXgPzVnezFOBazsBMmqJ9QW5A TRgTN1iJXsbvy8aaL1dCLFxH+t+qsnwn9E4MG+dwu4H4yg0dOWzJnLgdHeOf68mkrz XAdkJlKx0VievZ4qU9wrrizH1rYEmbkmkWM4giiliioJx8LvGSl/Vnh/d3Z3JPfpUw 5oUs+aORUrqonJbML1ImeZYJ7iX7eLN+0P53hDhC2Qy7FkPdx72NlaoauKeBorNGj7 7/fSl9Fm0ZuRQ==
Archived-At: http://mailarchive.ietf.org/arch/msg/clue/0NWD5PG0_rwDd2TwOtNjeGdhPJo
Subject: Re: [clue] Reminder: design team meeting tomorrow Tuesday, Apr 22
X-BeenThere: clue@ietf.org
X-Mailman-Version: 2.1.15
Precedence: list
List-Id: CLUE - ControLling mUltiple streams for TElepresence <clue.ietf.org>
List-Unsubscribe: <https://www.ietf.org/mailman/options/clue>, <mailto:clue-request@ietf.org?subject=unsubscribe>
List-Archive: <http://www.ietf.org/mail-archive/web/clue/>
List-Post: <mailto:clue@ietf.org>
List-Help: <mailto:clue-request@ietf.org?subject=help>
List-Subscribe: <https://www.ietf.org/mailman/listinfo/clue>, <mailto:clue-request@ietf.org?subject=subscribe>
X-List-Received-Date: Tue, 22 Apr 2014 02:40:36 -0000

On 4/21/14 8:44 PM, Christian Groves wrote:
> Hello Paul,
>
> What time is it? Is there any change as a result of the doodle poll?

The schedule of design team meetings was updated. It is now a half hour 
earlier than it was - 8am central US time.

> Christian
>
> On 22/04/2014 10:25 AM, Paul Kyzivat wrote:
>> This is just a reminder. We will be having the design team meeting
>> tomorrow. The primary subject will be the recent revision of the
>> protocol document.
>>
>>     Thanks,
>>     Paul
>>
>> _______________________________________________
>> clue mailing list
>> clue@ietf.org
>> https://www.ietf.org/mailman/listinfo/clue
>>
>
> _______________________________________________
> clue mailing list
> clue@ietf.org
> https://www.ietf.org/mailman/listinfo/clue
>


From nobody Mon Apr 21 23:54:17 2014
Return-Path: <Christian.Groves@nteczone.com>
X-Original-To: clue@ietfa.amsl.com
Delivered-To: clue@ietfa.amsl.com
Received: from localhost (ietfa.amsl.com [127.0.0.1]) by ietfa.amsl.com (Postfix) with ESMTP id 661491A00D8 for <clue@ietfa.amsl.com>; Mon, 21 Apr 2014 23:54:15 -0700 (PDT)
X-Virus-Scanned: amavisd-new at amsl.com
X-Spam-Flag: NO
X-Spam-Score: -0.001
X-Spam-Level: 
X-Spam-Status: No, score=-0.001 tagged_above=-999 required=5 tests=[BAYES_40=-0.001] autolearn=ham
Received: from mail.ietf.org ([4.31.198.44]) by localhost (ietfa.amsl.com [127.0.0.1]) (amavisd-new, port 10024) with ESMTP id roYfbshOrHmZ for <clue@ietfa.amsl.com>; Mon, 21 Apr 2014 23:54:13 -0700 (PDT)
Received: from cserver5.myshophosting.com (cserver5.myshophosting.com [175.107.161.1]) by ietfa.amsl.com (Postfix) with ESMTP id 9EBDD1A00D5 for <clue@ietf.org>; Mon, 21 Apr 2014 23:54:13 -0700 (PDT)
Received: from ppp118-209-25-196.lns20.mel4.internode.on.net ([118.209.25.196]:58278 helo=[127.0.0.1]) by cserver5.myshophosting.com with esmtpsa (TLSv1:DHE-RSA-AES128-SHA:128) (Exim 4.82) (envelope-from <Christian.Groves@nteczone.com>) id 1WcUah-00018f-SB for clue@ietf.org; Tue, 22 Apr 2014 16:54:03 +1000
Message-ID: <5356120D.6070701@nteczone.com>
Date: Tue, 22 Apr 2014 16:54:05 +1000
From: Christian Groves <Christian.Groves@nteczone.com>
User-Agent: Mozilla/5.0 (Windows NT 6.1; WOW64; rv:24.0) Gecko/20100101 Thunderbird/24.4.0
MIME-Version: 1.0
To: clue@ietf.org
References: <5355B6FB.6050501@alum.mit.edu> <5355BB80.7010704@nteczone.com> <5355D69D.1040603@alum.mit.edu>
In-Reply-To: <5355D69D.1040603@alum.mit.edu>
Content-Type: text/plain; charset=ISO-8859-1; format=flowed
Content-Transfer-Encoding: 7bit
X-AntiAbuse: This header was added to track abuse, please include it with any abuse report
X-AntiAbuse: Primary Hostname - cserver5.myshophosting.com
X-AntiAbuse: Original Domain - ietf.org
X-AntiAbuse: Originator/Caller UID/GID - [47 12] / [47 12]
X-AntiAbuse: Sender Address Domain - nteczone.com
X-Get-Message-Sender-Via: cserver5.myshophosting.com: authenticated_id: christian.groves@nteczone.com
X-Source: 
X-Source-Args: 
X-Source-Dir: 
Archived-At: http://mailarchive.ietf.org/arch/msg/clue/hNFWsFLBE8XNATb95VFjCBZ1BcY
Subject: Re: [clue] Reminder: design team meeting tomorrow Tuesday, Apr 22
X-BeenThere: clue@ietf.org
X-Mailman-Version: 2.1.15
Precedence: list
List-Id: CLUE - ControLling mUltiple streams for TElepresence <clue.ietf.org>
List-Unsubscribe: <https://www.ietf.org/mailman/options/clue>, <mailto:clue-request@ietf.org?subject=unsubscribe>
List-Archive: <http://www.ietf.org/mail-archive/web/clue/>
List-Post: <mailto:clue@ietf.org>
List-Help: <mailto:clue-request@ietf.org?subject=help>
List-Subscribe: <https://www.ietf.org/mailman/listinfo/clue>, <mailto:clue-request@ietf.org?subject=subscribe>
X-List-Received-Date: Tue, 22 Apr 2014 06:54:15 -0000

Hello Paul,

Thanks, by the way the CLUE wiki page 
(http://trac.tools.ietf.org/wg/clue/trac/wiki) still indicates 8.30am 
Central. Also the webex link at 
http://trac.tools.ietf.org/wg/clue/trac/wiki/Design-Team indicates that 
the time is 8.30am. These probably should be updated.

Regards, Christian

On 22/04/2014 12:40 PM, Paul Kyzivat wrote:
> On 4/21/14 8:44 PM, Christian Groves wrote:
>> Hello Paul,
>>
>> What time is it? Is there any change as a result of the doodle poll?
>
> The schedule of design team meetings was updated. It is now a half 
> hour earlier than it was - 8am central US time.
>
>> Christian
>>
>> On 22/04/2014 10:25 AM, Paul Kyzivat wrote:
>>> This is just a reminder. We will be having the design team meeting
>>> tomorrow. The primary subject will be the recent revision of the
>>> protocol document.
>>>
>>>     Thanks,
>>>     Paul
>>>
>>> _______________________________________________
>>> clue mailing list
>>> clue@ietf.org
>>> https://www.ietf.org/mailman/listinfo/clue
>>>
>>
>> _______________________________________________
>> clue mailing list
>> clue@ietf.org
>> https://www.ietf.org/mailman/listinfo/clue
>>
>
> _______________________________________________
> clue mailing list
> clue@ietf.org
> https://www.ietf.org/mailman/listinfo/clue
>


From nobody Tue Apr 22 01:32:13 2014
Return-Path: <spromano@unina.it>
X-Original-To: clue@ietfa.amsl.com
Delivered-To: clue@ietfa.amsl.com
Received: from localhost (ietfa.amsl.com [127.0.0.1]) by ietfa.amsl.com (Postfix) with ESMTP id 99E1A1A0087 for <clue@ietfa.amsl.com>; Tue, 22 Apr 2014 01:32:11 -0700 (PDT)
X-Virus-Scanned: amavisd-new at amsl.com
X-Spam-Flag: NO
X-Spam-Score: 1.108
X-Spam-Level: *
X-Spam-Status: No, score=1.108 tagged_above=-999 required=5 tests=[BAYES_05=-0.5, HELO_EQ_IT=0.635, HOST_EQ_IT=1.245, HTML_MESSAGE=0.001, RP_MATCHES_RCVD=-0.272, SPF_PASS=-0.001] autolearn=no
Received: from mail.ietf.org ([4.31.198.44]) by localhost (ietfa.amsl.com [127.0.0.1]) (amavisd-new, port 10024) with ESMTP id A5XRyRmUVN7Z for <clue@ietfa.amsl.com>; Tue, 22 Apr 2014 01:32:07 -0700 (PDT)
Received: from smtp2.unina.it (smtp2.unina.it [192.132.34.62]) by ietfa.amsl.com (Postfix) with ESMTP id 824651A0161 for <clue@ietf.org>; Tue, 22 Apr 2014 01:32:07 -0700 (PDT)
Received: from [143.225.28.167] ([143.225.28.167]) (authenticated bits=0) by smtp2.unina.it (8.14.4/8.14.4) with ESMTP id s3M8Vv0V010347 (version=TLSv1/SSLv3 cipher=AES128-SHA bits=128 verify=NO); Tue, 22 Apr 2014 10:31:58 +0200
Mime-Version: 1.0 (Apple Message framework v1283)
Content-Type: multipart/alternative; boundary="Apple-Mail=_7E3AC2C6-D8C4-4C2B-A8C1-E6857C0514B5"
From: Simon Pietro Romano <spromano@unina.it>
In-Reply-To: <5355D69D.1040603@alum.mit.edu>
Date: Tue, 22 Apr 2014 10:32:12 +0200
Message-Id: <D5923425-CAC2-4C37-90D4-DC8DA7F42D6E@unina.it>
References: <5355B6FB.6050501@alum.mit.edu> <5355BB80.7010704@nteczone.com> <5355D69D.1040603@alum.mit.edu>
To: Paul Kyzivat <pkyzivat@alum.mit.edu>
X-Mailer: Apple Mail (2.1283)
Archived-At: http://mailarchive.ietf.org/arch/msg/clue/th3ckrAKxJB4H8zwphyDyi72DZk
Cc: clue@ietf.org
Subject: Re: [clue] Reminder: design team meeting tomorrow Tuesday, Apr 22
X-BeenThere: clue@ietf.org
X-Mailman-Version: 2.1.15
Precedence: list
List-Id: CLUE - ControLling mUltiple streams for TElepresence <clue.ietf.org>
List-Unsubscribe: <https://www.ietf.org/mailman/options/clue>, <mailto:clue-request@ietf.org?subject=unsubscribe>
List-Archive: <http://www.ietf.org/mail-archive/web/clue/>
List-Post: <mailto:clue@ietf.org>
List-Help: <mailto:clue-request@ietf.org?subject=help>
List-Subscribe: <https://www.ietf.org/mailman/listinfo/clue>, <mailto:clue-request@ietf.org?subject=subscribe>
X-List-Received-Date: Tue, 22 Apr 2014 08:32:11 -0000

--Apple-Mail=_7E3AC2C6-D8C4-4C2B-A8C1-E6857C0514B5
Content-Transfer-Encoding: quoted-printable
Content-Type: text/plain;
	charset=iso-8859-1

Hi Paul,

protocol updates are scheduled for April 29th, right? This is what we =
have on our schedule and also what appears on the wiki:

http://trac.tools.ietf.org/wg/clue/trac/wiki/Design-Team

Can you confirm this is the right agenda?

Thanx,

Simon

On 22/apr/2014, at 04:40, Paul Kyzivat wrote:

> On 4/21/14 8:44 PM, Christian Groves wrote:
>> Hello Paul,
>>=20
>> What time is it? Is there any change as a result of the doodle poll?
>=20
> The schedule of design team meetings was updated. It is now a half =
hour earlier than it was - 8am central US time.
>=20
>> Christian
>>=20
>> On 22/04/2014 10:25 AM, Paul Kyzivat wrote:
>>> This is just a reminder. We will be having the design team meeting
>>> tomorrow. The primary subject will be the recent revision of the
>>> protocol document.
>>>=20
>>>    Thanks,
>>>    Paul
>>>=20
>>> _______________________________________________
>>> clue mailing list
>>> clue@ietf.org
>>> https://www.ietf.org/mailman/listinfo/clue
>>>=20
>>=20
>> _______________________________________________
>> clue mailing list
>> clue@ietf.org
>> https://www.ietf.org/mailman/listinfo/clue
>>=20
>=20
> _______________________________________________
> clue mailing list
> clue@ietf.org
> https://www.ietf.org/mailman/listinfo/clue
>=20

                     					       _\\|//_
                           				      ( O-O )
   ~~~~~~~~~~~~~~~~~~~~~~o00~~(_)~~00o~~~~~~~~~~~~~~~~~~~~~~~~
                    				Simon Pietro Romano
             				 Universita' di Napoli Federico =
II
                		     Computer Engineering Department=20
	             Phone: +39 081 7683823 -- Fax: +39 081 7683816
                                           e-mail: spromano@unina.it

		    <<Molti mi dicono che lo scoraggiamento =E8 l'alibi =
degli=20
		    idioti. Ci rifletto un istante; e mi scoraggio>>. =
Magritte.
               			                     oooO
  ~~~~~~~~~~~~~~~~~~~~~~~(   )~~~ Oooo~~~~~~~~~~~~~~~~~~~~~~~~~
					                 \ (            =
(   )
			                                  \_)          ) =
/
                                                                       =
(_/







--Apple-Mail=_7E3AC2C6-D8C4-4C2B-A8C1-E6857C0514B5
Content-Transfer-Encoding: quoted-printable
Content-Type: text/html;
	charset=iso-8859-1

<html><head></head><body style=3D"word-wrap: break-word; =
-webkit-nbsp-mode: space; -webkit-line-break: after-white-space; ">Hi =
Paul,<div><br></div><div>protocol updates are scheduled for April 29th, =
right? This is what we have on our schedule and also what appears on the =
wiki:</div><div><br></div><div><a =
href=3D"http://trac.tools.ietf.org/wg/clue/trac/wiki/Design-Team">http://t=
rac.tools.ietf.org/wg/clue/trac/wiki/Design-Team</a></div><div><br></div><=
div>Can you confirm this is the right =
agenda?</div><div><br></div><div>Thanx,</div><div><br></div><div>Simon</di=
v><div><br><div><div>On 22/apr/2014, at 04:40, Paul Kyzivat =
wrote:</div><br class=3D"Apple-interchange-newline"><blockquote =
type=3D"cite"><div>On 4/21/14 8:44 PM, Christian Groves =
wrote:<br><blockquote type=3D"cite">Hello =
Paul,<br></blockquote><blockquote =
type=3D"cite"><br></blockquote><blockquote type=3D"cite">What time is =
it? Is there any change as a result of the doodle =
poll?<br></blockquote><br>The schedule of design team meetings was =
updated. It is now a half hour earlier than it was - 8am central US =
time.<br><br><blockquote =
type=3D"cite">Christian<br></blockquote><blockquote =
type=3D"cite"><br></blockquote><blockquote type=3D"cite">On 22/04/2014 =
10:25 AM, Paul Kyzivat wrote:<br></blockquote><blockquote =
type=3D"cite"><blockquote type=3D"cite">This is just a reminder. We will =
be having the design team =
meeting<br></blockquote></blockquote><blockquote type=3D"cite"><blockquote=
 type=3D"cite">tomorrow. The primary subject will be the recent revision =
of the<br></blockquote></blockquote><blockquote type=3D"cite"><blockquote =
type=3D"cite">protocol =
document.<br></blockquote></blockquote><blockquote =
type=3D"cite"><blockquote =
type=3D"cite"><br></blockquote></blockquote><blockquote =
type=3D"cite"><blockquote type=3D"cite"> =
&nbsp;&nbsp;&nbsp;Thanks,<br></blockquote></blockquote><blockquote =
type=3D"cite"><blockquote type=3D"cite"> =
&nbsp;&nbsp;&nbsp;Paul<br></blockquote></blockquote><blockquote =
type=3D"cite"><blockquote =
type=3D"cite"><br></blockquote></blockquote><blockquote =
type=3D"cite"><blockquote =
type=3D"cite">_______________________________________________<br></blockqu=
ote></blockquote><blockquote type=3D"cite"><blockquote type=3D"cite">clue =
mailing list<br></blockquote></blockquote><blockquote =
type=3D"cite"><blockquote type=3D"cite"><a =
href=3D"mailto:clue@ietf.org">clue@ietf.org</a><br></blockquote></blockquo=
te><blockquote type=3D"cite"><blockquote type=3D"cite"><a =
href=3D"https://www.ietf.org/mailman/listinfo/clue">https://www.ietf.org/m=
ailman/listinfo/clue</a><br></blockquote></blockquote><blockquote =
type=3D"cite"><blockquote =
type=3D"cite"><br></blockquote></blockquote><blockquote =
type=3D"cite"><br></blockquote><blockquote =
type=3D"cite">_______________________________________________<br></blockqu=
ote><blockquote type=3D"cite">clue mailing =
list<br></blockquote><blockquote type=3D"cite"><a =
href=3D"mailto:clue@ietf.org">clue@ietf.org</a><br></blockquote><blockquot=
e type=3D"cite"><a =
href=3D"https://www.ietf.org/mailman/listinfo/clue">https://www.ietf.org/m=
ailman/listinfo/clue</a><br></blockquote><blockquote =
type=3D"cite"><br></blockquote><br>_______________________________________=
________<br>clue mailing list<br><a =
href=3D"mailto:clue@ietf.org">clue@ietf.org</a><br>https://www.ietf.org/ma=
ilman/listinfo/clue<br><br></div></blockquote></div><br><div =
apple-content-edited=3D"true">
<div style=3D"word-wrap: break-word; -webkit-nbsp-mode: space; =
-webkit-line-break: after-white-space; "><span class=3D"Apple-style-span" =
style=3D"border-collapse: separate; color: rgb(0, 0, 0); font-family: =
Helvetica; font-style: normal; font-variant: normal; font-weight: =
normal; letter-spacing: normal; line-height: normal; orphans: 2; =
text-align: -webkit-auto; text-indent: 0px; text-transform: none; =
white-space: normal; widows: 2; word-spacing: 0px; =
-webkit-border-horizontal-spacing: 0px; -webkit-border-vertical-spacing: =
0px; -webkit-text-decorations-in-effect: none; -webkit-text-size-adjust: =
auto; -webkit-text-stroke-width: 0px; font-size: medium; "><div =
style=3D"word-wrap: break-word; -webkit-nbsp-mode: space; =
-webkit-line-break: after-white-space; "><span class=3D"Apple-style-span" =
style=3D"border-collapse: separate; color: rgb(0, 0, 0); font-family: =
Helvetica; font-style: normal; font-variant: normal; font-weight: =
normal; letter-spacing: normal; line-height: normal; orphans: 2; =
text-align: -webkit-auto; text-indent: 0px; text-transform: none; =
white-space: normal; widows: 2; word-spacing: 0px; =
-webkit-border-horizontal-spacing: 0px; -webkit-border-vertical-spacing: =
0px; -webkit-text-decorations-in-effect: none; -webkit-text-size-adjust: =
auto; -webkit-text-stroke-width: 0px; font-size: medium; "><div =
style=3D"word-wrap: break-word; -webkit-nbsp-mode: space; =
-webkit-line-break: after-white-space; "><div><div>&nbsp; &nbsp; &nbsp; =
&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;<span =
class=3D"Apple-tab-span" style=3D"white-space: pre; ">				=
	</span><span class=3D"Apple-converted-space">&nbsp;</span>&nbsp; =
&nbsp; &nbsp; _\\|//_</div><div>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; =
&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;<span =
class=3D"Apple-tab-span" style=3D"white-space: pre; ">				=
</span>&nbsp; &nbsp; &nbsp;&nbsp;( O-O )</div><div>&nbsp; =
&nbsp;~~~~~~~~~~~~~~~~~~~~~~o00~~(_)~~00o~~~~~~~~~~~~~~~~~~~~~~~~</div><di=
v>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; =
&nbsp;<span class=3D"Apple-converted-space">&nbsp;</span><span =
class=3D"Apple-tab-span" style=3D"white-space: pre; ">				=
</span>Simon Pietro Romano</div><div>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; =
&nbsp; &nbsp;<span class=3D"Apple-tab-span" style=3D"white-space: pre; =
">				</span><span =
class=3D"Apple-converted-space">&nbsp;</span>Universita' di Napoli =
Federico II</div><div>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; =
&nbsp;&nbsp;<span class=3D"Apple-tab-span" style=3D"white-space: pre; ">	=
	</span>&nbsp; &nbsp; &nbsp;Computer Engineering =
Department&nbsp;</div><div><span class=3D"Apple-tab-span" =
style=3D"white-space: pre; ">	</span>&nbsp; &nbsp; &nbsp;&nbsp; &nbsp; =
&nbsp; &nbsp; Phone: +39 081 7683823 -- Fax: +39 081 =
7683816</div><div>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; =
&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; =
&nbsp; &nbsp; &nbsp; &nbsp; &nbsp;e-mail: <a =
href=3D"mailto:spromano@unina.it">spromano@unina.it</a></div><div><br></di=
v><div><span class=3D"Apple-tab-span" style=3D"white-space: pre; ">		=
</span>&nbsp; &nbsp; &lt;&lt;Molti mi dicono che lo scoraggiamento =E8 =
l'alibi degli&nbsp;</div><div><span class=3D"Apple-tab-span" =
style=3D"white-space: pre; ">		</span>&nbsp;&nbsp; =
&nbsp;idioti. Ci rifletto un istante; e mi scoraggio&gt;&gt;. =
Magritte.</div><div>&nbsp; &nbsp; &nbsp; &nbsp;&nbsp; &nbsp; &nbsp; =
&nbsp;<span class=3D"Apple-converted-space">&nbsp;</span><span =
class=3D"Apple-tab-span" style=3D"white-space: pre; ">			=
</span>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; =
&nbsp; &nbsp;oooO</div><div>&nbsp; ~~~~~~~~~~~~~~~~~~~~~~~( &nbsp; =
)~~~&nbsp;Oooo~~~~~~~~~~~~~~~~~~~~~~~~~</div><div><span =
class=3D"Apple-tab-span" style=3D"white-space: pre; ">				=
	</span>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; =
&nbsp;\ ( &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;( &nbsp; =
)</div><div><span class=3D"Apple-tab-span" style=3D"white-space: pre; ">	=
		</span>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; =
&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; =
\_) &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;) /</div><div>&nbsp; &nbsp; &nbsp; =
&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; =
&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; =
&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; =
&nbsp; &nbsp; &nbsp;(_/</div></div><div><br></div></div></span><br =
class=3D"Apple-interchange-newline"></div></span><br =
class=3D"Apple-interchange-newline"></div><br =
class=3D"Apple-interchange-newline"><br =
class=3D"Apple-interchange-newline">
</div>
<br></div></body></html>=

--Apple-Mail=_7E3AC2C6-D8C4-4C2B-A8C1-E6857C0514B5--


From nobody Tue Apr 22 05:44:57 2014
Return-Path: <mary.ietf.barnes@gmail.com>
X-Original-To: clue@ietfa.amsl.com
Delivered-To: clue@ietfa.amsl.com
Received: from localhost (ietfa.amsl.com [127.0.0.1]) by ietfa.amsl.com (Postfix) with ESMTP id EEDA61A03F4 for <clue@ietfa.amsl.com>; Tue, 22 Apr 2014 05:44:53 -0700 (PDT)
X-Virus-Scanned: amavisd-new at amsl.com
X-Spam-Flag: NO
X-Spam-Score: -0.599
X-Spam-Level: 
X-Spam-Status: No, score=-0.599 tagged_above=-999 required=5 tests=[BAYES_05=-0.5, DKIM_SIGNED=0.1, DKIM_VALID=-0.1, DKIM_VALID_AU=-0.1, FREEMAIL_FROM=0.001, HTML_MESSAGE=0.001, SPF_PASS=-0.001] autolearn=ham
Received: from mail.ietf.org ([4.31.198.44]) by localhost (ietfa.amsl.com [127.0.0.1]) (amavisd-new, port 10024) with ESMTP id pf2Vakkiry3J for <clue@ietfa.amsl.com>; Tue, 22 Apr 2014 05:44:49 -0700 (PDT)
Received: from mail-wi0-x234.google.com (mail-wi0-x234.google.com [IPv6:2a00:1450:400c:c05::234]) by ietfa.amsl.com (Postfix) with ESMTP id 9CB7E1A03EC for <clue@ietf.org>; Tue, 22 Apr 2014 05:44:48 -0700 (PDT)
Received: by mail-wi0-f180.google.com with SMTP id q5so3176753wiv.7 for <clue@ietf.org>; Tue, 22 Apr 2014 05:44:42 -0700 (PDT)
DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20120113;  h=mime-version:in-reply-to:references:date:message-id:subject:from:to :cc:content-type; bh=163zMQ6SdtjI2z92R8mB3uuDkW2ixoEyYy+Ozg0c9+g=; b=xOTIePsvsXb1QQxnmTqIgT5Z7lc1DnTsu9wpsUgj0eSdB+rhXypNagy+7U1F8tGGQX yw/Rf07L7JqvWwEQJjNt+7UUi6oj6V4JMikDkJdrz4HhBtMX9UAQTIdYlRNF/OMY/fs5 Ymyr/YJVYb15Wql/QITp9C8eF4IYjegSBIz5d14IHW+ke/RR9Aq/siBZYWDErhRLZl7w PlKbmUkWcjEoS/CdVJZglH0n/hZ+JI59aWSt6tkJ+wadi/PBMdOoa1Pe3WeEpx6KYgp3 9rkJUHFZCbRFbzsP4niK3cGa48W6fcf/iqMGT1clo177qMaL0ofrgFifZGOLPglxgiVs aByQ==
MIME-Version: 1.0
X-Received: by 10.180.77.165 with SMTP id t5mr18915882wiw.38.1398170682734; Tue, 22 Apr 2014 05:44:42 -0700 (PDT)
Received: by 10.216.10.6 with HTTP; Tue, 22 Apr 2014 05:44:42 -0700 (PDT)
In-Reply-To: <D5923425-CAC2-4C37-90D4-DC8DA7F42D6E@unina.it>
References: <5355B6FB.6050501@alum.mit.edu> <5355BB80.7010704@nteczone.com> <5355D69D.1040603@alum.mit.edu> <D5923425-CAC2-4C37-90D4-DC8DA7F42D6E@unina.it>
Date: Tue, 22 Apr 2014 07:44:42 -0500
Message-ID: <CAHBDyN6qBBeMROkm6eAZR6UTaqiXyW90yv3-PV_2g0_xN5a=4A@mail.gmail.com>
From: Mary Barnes <mary.ietf.barnes@gmail.com>
To: Simon Pietro Romano <spromano@unina.it>
Content-Type: multipart/alternative; boundary=f46d043c085ae1753c04f7a0fa69
Archived-At: http://mailarchive.ietf.org/arch/msg/clue/OG0f7iLG-kae2povYA0McbKysPE
Cc: CLUE <clue@ietf.org>
Subject: Re: [clue] Reminder: design team meeting tomorrow Tuesday, Apr 22
X-BeenThere: clue@ietf.org
X-Mailman-Version: 2.1.15
Precedence: list
List-Id: CLUE - ControLling mUltiple streams for TElepresence <clue.ietf.org>
List-Unsubscribe: <https://www.ietf.org/mailman/options/clue>, <mailto:clue-request@ietf.org?subject=unsubscribe>
List-Archive: <http://www.ietf.org/mail-archive/web/clue/>
List-Post: <mailto:clue@ietf.org>
List-Help: <mailto:clue-request@ietf.org?subject=help>
List-Subscribe: <https://www.ietf.org/mailman/listinfo/clue>, <mailto:clue-request@ietf.org?subject=subscribe>
X-List-Received-Date: Tue, 22 Apr 2014 12:44:54 -0000

--f46d043c085ae1753c04f7a0fa69
Content-Type: text/plain; charset=UTF-8
Content-Transfer-Encoding: quoted-printable

That is correct.


On Tue, Apr 22, 2014 at 3:32 AM, Simon Pietro Romano <spromano@unina.it>wro=
te:

> Hi Paul,
>
> protocol updates are scheduled for April 29th, right? This is what we hav=
e
> on our schedule and also what appears on the wiki:
>
> http://trac.tools.ietf.org/wg/clue/trac/wiki/Design-Team
>
> Can you confirm this is the right agenda?
>
> Thanx,
>
> Simon
>
> On 22/apr/2014, at 04:40, Paul Kyzivat wrote:
>
> On 4/21/14 8:44 PM, Christian Groves wrote:
>
> Hello Paul,
>
>
> What time is it? Is there any change as a result of the doodle poll?
>
>
> The schedule of design team meetings was updated. It is now a half hour
> earlier than it was - 8am central US time.
>
> Christian
>
>
> On 22/04/2014 10:25 AM, Paul Kyzivat wrote:
>
> This is just a reminder. We will be having the design team meeting
>
> tomorrow. The primary subject will be the recent revision of the
>
> protocol document.
>
>
>    Thanks,
>
>    Paul
>
>
> _______________________________________________
>
> clue mailing list
>
> clue@ietf.org
>
> https://www.ietf.org/mailman/listinfo/clue
>
>
>
> _______________________________________________
>
> clue mailing list
>
> clue@ietf.org
>
> https://www.ietf.org/mailman/listinfo/clue
>
>
>
> _______________________________________________
> clue mailing list
> clue@ietf.org
> https://www.ietf.org/mailman/listinfo/clue
>
>
>                              _\\|//_
>                                   ( O-O )
>    ~~~~~~~~~~~~~~~~~~~~~~o00~~(_)~~00o~~~~~~~~~~~~~~~~~~~~~~~~
>                      Simon Pietro Romano
>                Universita' di Napoli Federico II
>                       Computer Engineering Department
>              Phone: +39 081 7683823 -- Fax: +39 081 7683816
>                                            e-mail: spromano@unina.it
>
>     <<Molti mi dicono che lo scoraggiamento =C3=A8 l'alibi degli
>     idioti. Ci rifletto un istante; e mi scoraggio>>. Magritte.
>                                      oooO
>   ~~~~~~~~~~~~~~~~~~~~~~~(   )~~~ Oooo~~~~~~~~~~~~~~~~~~~~~~~~~
>                  \ (            (   )
>                                   \_)          ) /
>                                                                        (_=
/
>
>
>
>
>
>
>
> _______________________________________________
> clue mailing list
> clue@ietf.org
> https://www.ietf.org/mailman/listinfo/clue
>
>

--f46d043c085ae1753c04f7a0fa69
Content-Type: text/html; charset=UTF-8
Content-Transfer-Encoding: quoted-printable

<div dir=3D"ltr">That is correct.=C2=A0</div><div class=3D"gmail_extra"><br=
><br><div class=3D"gmail_quote">On Tue, Apr 22, 2014 at 3:32 AM, Simon Piet=
ro Romano <span dir=3D"ltr">&lt;<a href=3D"mailto:spromano@unina.it" target=
=3D"_blank">spromano@unina.it</a>&gt;</span> wrote:<br>
<blockquote class=3D"gmail_quote" style=3D"margin:0 0 0 .8ex;border-left:1p=
x #ccc solid;padding-left:1ex"><div style=3D"word-wrap:break-word">Hi Paul,=
<div><br></div><div>protocol updates are scheduled for April 29th, right? T=
his is what we have on our schedule and also what appears on the wiki:</div=
>
<div><br></div><div><a href=3D"http://trac.tools.ietf.org/wg/clue/trac/wiki=
/Design-Team" target=3D"_blank">http://trac.tools.ietf.org/wg/clue/trac/wik=
i/Design-Team</a></div><div><br></div><div>Can you confirm this is the righ=
t agenda?</div>
<div><br></div><div>Thanx,</div><div><br></div><div>Simon</div><div><div><d=
iv class=3D"h5"><br><div><div>On 22/apr/2014, at 04:40, Paul Kyzivat wrote:=
</div><br><blockquote type=3D"cite"><div>On 4/21/14 8:44 PM, Christian Grov=
es wrote:<br>
<blockquote type=3D"cite">Hello Paul,<br></blockquote><blockquote type=3D"c=
ite"><br></blockquote><blockquote type=3D"cite">What time is it? Is there a=
ny change as a result of the doodle poll?<br></blockquote><br>The schedule =
of design team meetings was updated. It is now a half hour earlier than it =
was - 8am central US time.<br>
<br><blockquote type=3D"cite">Christian<br></blockquote><blockquote type=3D=
"cite"><br></blockquote><blockquote type=3D"cite">On 22/04/2014 10:25 AM, P=
aul Kyzivat wrote:<br></blockquote><blockquote type=3D"cite"><blockquote ty=
pe=3D"cite">
This is just a reminder. We will be having the design team meeting<br></blo=
ckquote></blockquote><blockquote type=3D"cite"><blockquote type=3D"cite">to=
morrow. The primary subject will be the recent revision of the<br></blockqu=
ote>
</blockquote><blockquote type=3D"cite"><blockquote type=3D"cite">protocol d=
ocument.<br></blockquote></blockquote><blockquote type=3D"cite"><blockquote=
 type=3D"cite"><br></blockquote></blockquote><blockquote type=3D"cite"><blo=
ckquote type=3D"cite">
 =C2=A0=C2=A0=C2=A0Thanks,<br></blockquote></blockquote><blockquote type=3D=
"cite"><blockquote type=3D"cite"> =C2=A0=C2=A0=C2=A0Paul<br></blockquote></=
blockquote><blockquote type=3D"cite"><blockquote type=3D"cite"><br></blockq=
uote></blockquote><blockquote type=3D"cite">
<blockquote type=3D"cite">_______________________________________________<b=
r></blockquote></blockquote><blockquote type=3D"cite"><blockquote type=3D"c=
ite">clue mailing list<br></blockquote></blockquote><blockquote type=3D"cit=
e"><blockquote type=3D"cite">
<a href=3D"mailto:clue@ietf.org" target=3D"_blank">clue@ietf.org</a><br></b=
lockquote></blockquote><blockquote type=3D"cite"><blockquote type=3D"cite">=
<a href=3D"https://www.ietf.org/mailman/listinfo/clue" target=3D"_blank">ht=
tps://www.ietf.org/mailman/listinfo/clue</a><br>
</blockquote></blockquote><blockquote type=3D"cite"><blockquote type=3D"cit=
e"><br></blockquote></blockquote><blockquote type=3D"cite"><br></blockquote=
><blockquote type=3D"cite">_______________________________________________<=
br></blockquote>
<blockquote type=3D"cite">clue mailing list<br></blockquote><blockquote typ=
e=3D"cite"><a href=3D"mailto:clue@ietf.org" target=3D"_blank">clue@ietf.org=
</a><br></blockquote><blockquote type=3D"cite"><a href=3D"https://www.ietf.=
org/mailman/listinfo/clue" target=3D"_blank">https://www.ietf.org/mailman/l=
istinfo/clue</a><br>
</blockquote><blockquote type=3D"cite"><br></blockquote><br>_______________=
________________________________<br>clue mailing list<br><a href=3D"mailto:=
clue@ietf.org" target=3D"_blank">clue@ietf.org</a><br><a href=3D"https://ww=
w.ietf.org/mailman/listinfo/clue" target=3D"_blank">https://www.ietf.org/ma=
ilman/listinfo/clue</a><br>
<br></div></blockquote></div><br></div></div><div>
<div style=3D"word-wrap:break-word"><span style=3D"text-indent:0px;letter-s=
pacing:normal;font-variant:normal;text-align:-webkit-auto;font-style:normal=
;font-weight:normal;line-height:normal;border-collapse:separate;text-transf=
orm:none;font-size:medium;white-space:normal;font-family:Helvetica;word-spa=
cing:0px"><div style=3D"word-wrap:break-word">
<span style=3D"text-indent:0px;letter-spacing:normal;font-variant:normal;te=
xt-align:-webkit-auto;font-style:normal;font-weight:normal;line-height:norm=
al;border-collapse:separate;text-transform:none;font-size:medium;white-spac=
e:normal;font-family:Helvetica;word-spacing:0px"><div style=3D"word-wrap:br=
eak-word">
<div><div>=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =
=C2=A0 =C2=A0<span style=3D"white-space:pre-wrap">					</span><span>=C2=A0<=
/span>=C2=A0 =C2=A0 =C2=A0 _\\|//_</div><div>=C2=A0 =C2=A0 =C2=A0 =C2=A0 =
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0<span =
style=3D"white-space:pre-wrap">				</span>=C2=A0 =C2=A0 =C2=A0=C2=A0( O-O )=
</div><div>=C2=A0 =C2=A0~~~~~~~~~~~~~~~~~~~~~~o00~~(_)~~00o~~~~~~~~~~~~~~~~=
~~~~~~~~</div>
<div>=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0<=
span>=C2=A0</span><span style=3D"white-space:pre-wrap">				</span>Simon Pie=
tro Romano</div><div>=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0<span =
style=3D"white-space:pre-wrap">				</span><span>=C2=A0</span>Universita&#39=
; di Napoli Federico II</div>
<div>=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0=C2=A0<span sty=
le=3D"white-space:pre-wrap">		</span>=C2=A0 =C2=A0 =C2=A0Computer Engineeri=
ng Department=C2=A0</div><div><span style=3D"white-space:pre-wrap">	</span>=
=C2=A0 =C2=A0 =C2=A0=C2=A0 =C2=A0 =C2=A0 =C2=A0 Phone: <a href=3D"tel:%2B39=
%20081%207683823" value=3D"+390817683823" target=3D"_blank">+39 081 7683823=
</a> -- Fax: <a href=3D"tel:%2B39%20081%207683816" value=3D"+390817683816" =
target=3D"_blank">+39 081 7683816</a></div>
<div>=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=
=A0 =C2=A0e-mail: <a href=3D"mailto:spromano@unina.it" target=3D"_blank">sp=
romano@unina.it</a></div><div><br></div><div><span style=3D"white-space:pre=
-wrap">		</span>=C2=A0 =C2=A0 &lt;&lt;Molti mi dicono che lo scoraggiamento=
 =C3=A8 l&#39;alibi degli=C2=A0</div>
<div><span style=3D"white-space:pre-wrap">		</span>=C2=A0=C2=A0 =C2=A0idiot=
i. Ci rifletto un istante; e mi scoraggio&gt;&gt;. Magritte.</div><div>=C2=
=A0 =C2=A0 =C2=A0 =C2=A0=C2=A0 =C2=A0 =C2=A0 =C2=A0<span>=C2=A0</span><span=
 style=3D"white-space:pre-wrap">			</span>=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=
=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0oooO</div>
<div>=C2=A0 ~~~~~~~~~~~~~~~~~~~~~~~( =C2=A0 )~~~=C2=A0Oooo~~~~~~~~~~~~~~~~~=
~~~~~~~~</div><div><span style=3D"white-space:pre-wrap">					</span>=C2=A0 =
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0\ ( =C2=A0 =C2=A0 =
=C2=A0 =C2=A0 =C2=A0 =C2=A0( =C2=A0 )</div><div><span style=3D"white-space:=
pre-wrap">			</span>=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0=
 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 \_) =C2=A0 =
=C2=A0 =C2=A0 =C2=A0 =C2=A0) /</div>
<div>=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=
=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0(_/</div></div><div><br></div></div></spa=
n><br></div></span><br></div><br><br>
</div>
<br></div></div><br>_______________________________________________<br>
clue mailing list<br>
<a href=3D"mailto:clue@ietf.org">clue@ietf.org</a><br>
<a href=3D"https://www.ietf.org/mailman/listinfo/clue" target=3D"_blank">ht=
tps://www.ietf.org/mailman/listinfo/clue</a><br>
<br></blockquote></div><br></div>

--f46d043c085ae1753c04f7a0fa69--


From nobody Tue Apr 22 05:48:00 2014
Return-Path: <pkyzivat@alum.mit.edu>
X-Original-To: clue@ietfa.amsl.com
Delivered-To: clue@ietfa.amsl.com
Received: from localhost (ietfa.amsl.com [127.0.0.1]) by ietfa.amsl.com (Postfix) with ESMTP id B03FB1A040B for <clue@ietfa.amsl.com>; Tue, 22 Apr 2014 05:47:58 -0700 (PDT)
X-Virus-Scanned: amavisd-new at amsl.com
X-Spam-Flag: NO
X-Spam-Score: -1.235
X-Spam-Level: 
X-Spam-Status: No, score=-1.235 tagged_above=-999 required=5 tests=[BAYES_00=-1.9, DKIM_SIGNED=0.1, DKIM_VALID=-0.1, SPF_SOFTFAIL=0.665] autolearn=no
Received: from mail.ietf.org ([4.31.198.44]) by localhost (ietfa.amsl.com [127.0.0.1]) (amavisd-new, port 10024) with ESMTP id 9lxTqupDQjSs for <clue@ietfa.amsl.com>; Tue, 22 Apr 2014 05:47:54 -0700 (PDT)
Received: from qmta13.westchester.pa.mail.comcast.net (qmta13.westchester.pa.mail.comcast.net [IPv6:2001:558:fe14:44:76:96:59:243]) by ietfa.amsl.com (Postfix) with ESMTP id 99BEE1A03EC for <clue@ietf.org>; Tue, 22 Apr 2014 05:47:54 -0700 (PDT)
Received: from omta05.westchester.pa.mail.comcast.net ([76.96.62.43]) by qmta13.westchester.pa.mail.comcast.net with comcast id szuz1n0040vyq2s5D0np5Z; Tue, 22 Apr 2014 12:47:49 +0000
Received: from Paul-Kyzivats-MacBook-Pro.local ([50.138.229.164]) by omta05.westchester.pa.mail.comcast.net with comcast id t0no1n00W3ZTu2S3R0npjh; Tue, 22 Apr 2014 12:47:49 +0000
Message-ID: <535664F4.5040405@alum.mit.edu>
Date: Tue, 22 Apr 2014 08:47:48 -0400
From: Paul Kyzivat <pkyzivat@alum.mit.edu>
User-Agent: Mozilla/5.0 (Macintosh; Intel Mac OS X 10.7; rv:24.0) Gecko/20100101 Thunderbird/24.4.0
MIME-Version: 1.0
To: clue@ietf.org
References: <5355B6FB.6050501@alum.mit.edu> <5355BB80.7010704@nteczone.com> <5355D69D.1040603@alum.mit.edu> <5356120D.6070701@nteczone.com>
In-Reply-To: <5356120D.6070701@nteczone.com>
Content-Type: text/plain; charset=ISO-8859-1; format=flowed
Content-Transfer-Encoding: 7bit
DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=comcast.net; s=q20140121; t=1398170869; bh=nPJj5RhvvINjZw2Xr31X0NwGPGy22iAdsyrOpoway/U=; h=Received:Received:Message-ID:Date:From:MIME-Version:To:Subject: Content-Type; b=lWygsMi3Lep9ExsZdr5wxVbRwc0C+E2zLv1lQtuLuiE1fs6aF1y6o25fj0UstU0ED H5nrlHWz8oj7PH/nHv7AfSEDwlUr2PpqywhCkCPnOCatK+dPx0JDRv+RIZE7klZUh+ tYZyWeje9a+VOvBzaB3Yk5MFyKY1g3NAPQNwWL52gHgyquUaiShrY7/673mx/R4pLD 4zpOO6IXbTQgknQzrqaPtM8+ZrW1daKLqmMmJ/lXyCHT9TKlEamluilBFjKN45Ww+C XjmaGFwDK8k47ddgj4O7qLQxl75+D3ePpFC6IU8CXUKsh22r+up8JXe6rWPB1R1lCL wwuCvjhByKBOw==
Archived-At: http://mailarchive.ietf.org/arch/msg/clue/hZ3AR6JO2NnOsPEGWhp2EjgkngM
Subject: Re: [clue] Reminder: design team meeting tomorrow Tuesday, Apr 22
X-BeenThere: clue@ietf.org
X-Mailman-Version: 2.1.15
Precedence: list
List-Id: CLUE - ControLling mUltiple streams for TElepresence <clue.ietf.org>
List-Unsubscribe: <https://www.ietf.org/mailman/options/clue>, <mailto:clue-request@ietf.org?subject=unsubscribe>
List-Archive: <http://www.ietf.org/mail-archive/web/clue/>
List-Post: <mailto:clue@ietf.org>
List-Help: <mailto:clue-request@ietf.org?subject=help>
List-Subscribe: <https://www.ietf.org/mailman/listinfo/clue>, <mailto:clue-request@ietf.org?subject=subscribe>
X-List-Received-Date: Tue, 22 Apr 2014 12:47:58 -0000

On 4/22/14 2:54 AM, Christian Groves wrote:
> Hello Paul,
>
> Thanks, by the way the CLUE wiki page
> (http://trac.tools.ietf.org/wg/clue/trac/wiki) still indicates 8.30am
> Central. Also the webex link at
> http://trac.tools.ietf.org/wg/clue/trac/wiki/Design-Team indicates that
> the time is 8.30am. These probably should be updated.

Hmm. I'm looking at the last url above, and it says 8am.

BUT, the webex is still set up for 8:30. :-(
I'm trying to figure out how to fix it. If I can't get it fixed, then we 
will have to start at 8:30.

	Thanks,
	Paul

> Regards, Christian
>
> On 22/04/2014 12:40 PM, Paul Kyzivat wrote:
>> On 4/21/14 8:44 PM, Christian Groves wrote:
>>> Hello Paul,
>>>
>>> What time is it? Is there any change as a result of the doodle poll?
>>
>> The schedule of design team meetings was updated. It is now a half
>> hour earlier than it was - 8am central US time.
>>
>>> Christian
>>>
>>> On 22/04/2014 10:25 AM, Paul Kyzivat wrote:
>>>> This is just a reminder. We will be having the design team meeting
>>>> tomorrow. The primary subject will be the recent revision of the
>>>> protocol document.
>>>>
>>>>     Thanks,
>>>>     Paul
>>>>
>>>> _______________________________________________
>>>> clue mailing list
>>>> clue@ietf.org
>>>> https://www.ietf.org/mailman/listinfo/clue
>>>>
>>>
>>> _______________________________________________
>>> clue mailing list
>>> clue@ietf.org
>>> https://www.ietf.org/mailman/listinfo/clue
>>>
>>
>> _______________________________________________
>> clue mailing list
>> clue@ietf.org
>> https://www.ietf.org/mailman/listinfo/clue
>>
>
> _______________________________________________
> clue mailing list
> clue@ietf.org
> https://www.ietf.org/mailman/listinfo/clue
>


From nobody Tue Apr 22 05:51:57 2014
Return-Path: <pkyzivat@alum.mit.edu>
X-Original-To: clue@ietfa.amsl.com
Delivered-To: clue@ietfa.amsl.com
Received: from localhost (ietfa.amsl.com [127.0.0.1]) by ietfa.amsl.com (Postfix) with ESMTP id C6FBF1A0415 for <clue@ietfa.amsl.com>; Tue, 22 Apr 2014 05:51:56 -0700 (PDT)
X-Virus-Scanned: amavisd-new at amsl.com
X-Spam-Flag: NO
X-Spam-Score: -1.235
X-Spam-Level: 
X-Spam-Status: No, score=-1.235 tagged_above=-999 required=5 tests=[BAYES_00=-1.9, DKIM_SIGNED=0.1, DKIM_VALID=-0.1, SPF_SOFTFAIL=0.665] autolearn=no
Received: from mail.ietf.org ([4.31.198.44]) by localhost (ietfa.amsl.com [127.0.0.1]) (amavisd-new, port 10024) with ESMTP id aVNWWrVjvDSM for <clue@ietfa.amsl.com>; Tue, 22 Apr 2014 05:51:52 -0700 (PDT)
Received: from qmta08.westchester.pa.mail.comcast.net (qmta08.westchester.pa.mail.comcast.net [IPv6:2001:558:fe14:43:76:96:62:80]) by ietfa.amsl.com (Postfix) with ESMTP id 698021A0410 for <clue@ietf.org>; Tue, 22 Apr 2014 05:51:52 -0700 (PDT)
Received: from omta01.westchester.pa.mail.comcast.net ([76.96.62.11]) by qmta08.westchester.pa.mail.comcast.net with comcast id t0rQ1n0020EZKEL580rnps; Tue, 22 Apr 2014 12:51:47 +0000
Received: from Paul-Kyzivats-MacBook-Pro.local ([50.138.229.164]) by omta01.westchester.pa.mail.comcast.net with comcast id t0rm1n01m3ZTu2S3M0rmG1; Tue, 22 Apr 2014 12:51:47 +0000
Message-ID: <535665E2.20204@alum.mit.edu>
Date: Tue, 22 Apr 2014 08:51:46 -0400
From: Paul Kyzivat <pkyzivat@alum.mit.edu>
User-Agent: Mozilla/5.0 (Macintosh; Intel Mac OS X 10.7; rv:24.0) Gecko/20100101 Thunderbird/24.4.0
MIME-Version: 1.0
To: clue@ietf.org
References: <5355B6FB.6050501@alum.mit.edu> <5355BB80.7010704@nteczone.com> <5355D69D.1040603@alum.mit.edu> <5356120D.6070701@nteczone.com> <535664F4.5040405@alum.mit.edu>
In-Reply-To: <535664F4.5040405@alum.mit.edu>
Content-Type: text/plain; charset=ISO-8859-1; format=flowed
Content-Transfer-Encoding: 7bit
DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=comcast.net; s=q20140121; t=1398171107; bh=RyA5DFAOqnTJVVT3KZhgNTuHxRvjE4hXJG2ClApqXwU=; h=Received:Received:Message-ID:Date:From:MIME-Version:To:Subject: Content-Type; b=Oi2YHs/G82hyKen9X/tvIrcp4V8Bajrw/r5UxFPqkmHnMJ8LDT+KqZzrcbpXY1Z8M vhpImTyhPxb/uYrvmPV1hU5M/4CK+nmR09PM//RhPPz11E791KPPizMxL9+iK6sF20 B00EkNn3qf3sU/DBNMQ8UZc4Za6TibR7y+tpTuF4v6DtQ9OfpxxhODYwT21KJHLNoA WzmXwuQ6yg8/OH3WhVAeVdPj5CJG8a5J7QyQX0ajXMzEI0wky2cw2X3cBzyAxZx7oY e8UPXbdte+Z9GqM4ScnvT98eET0ymHLB3NM5fnLANGMcNNoKstMrj3pvZju9VaiX3i xctf8hNH/Fmug==
Archived-At: http://mailarchive.ietf.org/arch/msg/clue/YbM44cijrxEGd8IUQJlKoSP-Fus
Subject: Re: [clue] Reminder: design team meeting tomorrow Tuesday, Apr 22
X-BeenThere: clue@ietf.org
X-Mailman-Version: 2.1.15
Precedence: list
List-Id: CLUE - ControLling mUltiple streams for TElepresence <clue.ietf.org>
List-Unsubscribe: <https://www.ietf.org/mailman/options/clue>, <mailto:clue-request@ietf.org?subject=unsubscribe>
List-Archive: <http://www.ietf.org/mail-archive/web/clue/>
List-Post: <mailto:clue@ietf.org>
List-Help: <mailto:clue-request@ietf.org?subject=help>
List-Subscribe: <https://www.ietf.org/mailman/listinfo/clue>, <mailto:clue-request@ietf.org?subject=subscribe>
X-List-Received-Date: Tue, 22 Apr 2014 12:51:56 -0000

I checked again, and even though the webex shows on the schedule for 
8:30, when clicking through to it (at 
https://ietf.webex.com/mw0307l/mywebex/default.do?siteurl=ietf), it is 
set to start at 8 central. So I guess it is ok.

And I've been corrected on the subject.

On 4/22/14 8:47 AM, Paul Kyzivat wrote:
>
>
> On 4/22/14 2:54 AM, Christian Groves wrote:
>> Hello Paul,
>>
>> Thanks, by the way the CLUE wiki page
>> (http://trac.tools.ietf.org/wg/clue/trac/wiki) still indicates 8.30am
>> Central. Also the webex link at
>> http://trac.tools.ietf.org/wg/clue/trac/wiki/Design-Team indicates that
>> the time is 8.30am. These probably should be updated.
>
> Hmm. I'm looking at the last url above, and it says 8am.
>
> BUT, the webex is still set up for 8:30. :-(
> I'm trying to figure out how to fix it. If I can't get it fixed, then we
> will have to start at 8:30.
>
>      Thanks,
>      Paul
>
>> Regards, Christian
>>
>> On 22/04/2014 12:40 PM, Paul Kyzivat wrote:
>>> On 4/21/14 8:44 PM, Christian Groves wrote:
>>>> Hello Paul,
>>>>
>>>> What time is it? Is there any change as a result of the doodle poll?
>>>
>>> The schedule of design team meetings was updated. It is now a half
>>> hour earlier than it was - 8am central US time.
>>>
>>>> Christian
>>>>
>>>> On 22/04/2014 10:25 AM, Paul Kyzivat wrote:
>>>>> This is just a reminder. We will be having the design team meeting
>>>>> tomorrow. The primary subject will be the recent revision of the
>>>>> protocol document.
>>>>>
>>>>>     Thanks,
>>>>>     Paul
>>>>>
>>>>> _______________________________________________
>>>>> clue mailing list
>>>>> clue@ietf.org
>>>>> https://www.ietf.org/mailman/listinfo/clue
>>>>>
>>>>
>>>> _______________________________________________
>>>> clue mailing list
>>>> clue@ietf.org
>>>> https://www.ietf.org/mailman/listinfo/clue
>>>>
>>>
>>> _______________________________________________
>>> clue mailing list
>>> clue@ietf.org
>>> https://www.ietf.org/mailman/listinfo/clue
>>>
>>
>> _______________________________________________
>> clue mailing list
>> clue@ietf.org
>> https://www.ietf.org/mailman/listinfo/clue
>>
>
> _______________________________________________
> clue mailing list
> clue@ietf.org
> https://www.ietf.org/mailman/listinfo/clue
>


From nobody Tue Apr 22 05:54:03 2014
Return-Path: <mary.ietf.barnes@gmail.com>
X-Original-To: clue@ietfa.amsl.com
Delivered-To: clue@ietfa.amsl.com
Received: from localhost (ietfa.amsl.com [127.0.0.1]) by ietfa.amsl.com (Postfix) with ESMTP id B373A1A041B for <clue@ietfa.amsl.com>; Tue, 22 Apr 2014 05:54:02 -0700 (PDT)
X-Virus-Scanned: amavisd-new at amsl.com
X-Spam-Flag: NO
X-Spam-Score: -1.999
X-Spam-Level: 
X-Spam-Status: No, score=-1.999 tagged_above=-999 required=5 tests=[BAYES_00=-1.9, DKIM_SIGNED=0.1, DKIM_VALID=-0.1, DKIM_VALID_AU=-0.1, FREEMAIL_FROM=0.001, HTML_MESSAGE=0.001, SPF_PASS=-0.001] autolearn=ham
Received: from mail.ietf.org ([4.31.198.44]) by localhost (ietfa.amsl.com [127.0.0.1]) (amavisd-new, port 10024) with ESMTP id 4jC4G_MHgUxN for <clue@ietfa.amsl.com>; Tue, 22 Apr 2014 05:53:58 -0700 (PDT)
Received: from mail-wi0-x235.google.com (mail-wi0-x235.google.com [IPv6:2a00:1450:400c:c05::235]) by ietfa.amsl.com (Postfix) with ESMTP id F24A21A0410 for <clue@ietf.org>; Tue, 22 Apr 2014 05:53:57 -0700 (PDT)
Received: by mail-wi0-f181.google.com with SMTP id hm4so3194469wib.14 for <clue@ietf.org>; Tue, 22 Apr 2014 05:53:52 -0700 (PDT)
DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20120113;  h=mime-version:in-reply-to:references:date:message-id:subject:from:to :cc:content-type; bh=iyQLNvUDQ0NzvrAryecE3QR2rbBh0LfDrjEiuPouzvU=; b=yFMc5YNLPTRcUo/AwLL51Z6NtxGcmZeBBPxm1wOhCZP4BaCGilMDS1ShNG+FNWJPpP BNKzODcPwzulZa1gy/Ce6EkVSr+0jkNWGMUzw2FWuJl7/PWjl9HbCCzNRO7vjD2aIjl2 iqNKorw/jwKn0t3XgJVEtPwYhwAHTFN5a5u26qSlFm2x9uBcauAHCpO1ZT4MJTgh56Ow dvyLXYpVtFNdSo5jEEsb/widOR3kTfrY4LD+tzkBS4UDyIStWwMMuwTazeTNazuNUDkX TMCzkcQZWGRRxTtqm0tI7DNgy2vDCzf3DhHf8+chANxvk+TaM3vDa41SFlpXfpphjFC/ FEcA==
MIME-Version: 1.0
X-Received: by 10.194.109.6 with SMTP id ho6mr34183817wjb.21.1398171232130; Tue, 22 Apr 2014 05:53:52 -0700 (PDT)
Received: by 10.216.10.6 with HTTP; Tue, 22 Apr 2014 05:53:52 -0700 (PDT)
In-Reply-To: <535665E2.20204@alum.mit.edu>
References: <5355B6FB.6050501@alum.mit.edu> <5355BB80.7010704@nteczone.com> <5355D69D.1040603@alum.mit.edu> <5356120D.6070701@nteczone.com> <535664F4.5040405@alum.mit.edu> <535665E2.20204@alum.mit.edu>
Date: Tue, 22 Apr 2014 07:53:52 -0500
Message-ID: <CAHBDyN7z9-0HNCzkEUjOZ4-1tsHCVNsdi0aT+QrA0ieOBx=UqQ@mail.gmail.com>
From: Mary Barnes <mary.ietf.barnes@gmail.com>
To: Paul Kyzivat <pkyzivat@alum.mit.edu>
Content-Type: multipart/alternative; boundary=047d7bf10b50a05c3504f7a11b7e
Archived-At: http://mailarchive.ietf.org/arch/msg/clue/2fvbZbFW6H6WXyPCBInxRIj7_E0
Cc: CLUE <clue@ietf.org>
Subject: Re: [clue] Reminder: design team meeting tomorrow Tuesday, Apr 22
X-BeenThere: clue@ietf.org
X-Mailman-Version: 2.1.15
Precedence: list
List-Id: CLUE - ControLling mUltiple streams for TElepresence <clue.ietf.org>
List-Unsubscribe: <https://www.ietf.org/mailman/options/clue>, <mailto:clue-request@ietf.org?subject=unsubscribe>
List-Archive: <http://www.ietf.org/mail-archive/web/clue/>
List-Post: <mailto:clue@ietf.org>
List-Help: <mailto:clue-request@ietf.org?subject=help>
List-Subscribe: <https://www.ietf.org/mailman/listinfo/clue>, <mailto:clue-request@ietf.org?subject=subscribe>
X-List-Received-Date: Tue, 22 Apr 2014 12:54:02 -0000

--047d7bf10b50a05c3504f7a11b7e
Content-Type: text/plain; charset=UTF-8

Yes, I changed the Webex start time.


On Tue, Apr 22, 2014 at 7:51 AM, Paul Kyzivat <pkyzivat@alum.mit.edu> wrote:

> I checked again, and even though the webex shows on the schedule for 8:30,
> when clicking through to it (at https://ietf.webex.com/
> mw0307l/mywebex/default.do?siteurl=ietf), it is set to start at 8
> central. So I guess it is ok.
>
> And I've been corrected on the subject.
>
>
> On 4/22/14 8:47 AM, Paul Kyzivat wrote:
>
>>
>>
>> On 4/22/14 2:54 AM, Christian Groves wrote:
>>
>>> Hello Paul,
>>>
>>> Thanks, by the way the CLUE wiki page
>>> (http://trac.tools.ietf.org/wg/clue/trac/wiki) still indicates 8.30am
>>> Central. Also the webex link at
>>> http://trac.tools.ietf.org/wg/clue/trac/wiki/Design-Team indicates that
>>> the time is 8.30am. These probably should be updated.
>>>
>>
>> Hmm. I'm looking at the last url above, and it says 8am.
>>
>> BUT, the webex is still set up for 8:30. :-(
>> I'm trying to figure out how to fix it. If I can't get it fixed, then we
>> will have to start at 8:30.
>>
>>      Thanks,
>>      Paul
>>
>>  Regards, Christian
>>>
>>> On 22/04/2014 12:40 PM, Paul Kyzivat wrote:
>>>
>>>> On 4/21/14 8:44 PM, Christian Groves wrote:
>>>>
>>>>> Hello Paul,
>>>>>
>>>>> What time is it? Is there any change as a result of the doodle poll?
>>>>>
>>>>
>>>> The schedule of design team meetings was updated. It is now a half
>>>> hour earlier than it was - 8am central US time.
>>>>
>>>>  Christian
>>>>>
>>>>> On 22/04/2014 10:25 AM, Paul Kyzivat wrote:
>>>>>
>>>>>> This is just a reminder. We will be having the design team meeting
>>>>>> tomorrow. The primary subject will be the recent revision of the
>>>>>> protocol document.
>>>>>>
>>>>>>     Thanks,
>>>>>>     Paul
>>>>>>
>>>>>> _______________________________________________
>>>>>> clue mailing list
>>>>>> clue@ietf.org
>>>>>> https://www.ietf.org/mailman/listinfo/clue
>>>>>>
>>>>>>
>>>>> _______________________________________________
>>>>> clue mailing list
>>>>> clue@ietf.org
>>>>> https://www.ietf.org/mailman/listinfo/clue
>>>>>
>>>>>
>>>> _______________________________________________
>>>> clue mailing list
>>>> clue@ietf.org
>>>> https://www.ietf.org/mailman/listinfo/clue
>>>>
>>>>
>>> _______________________________________________
>>> clue mailing list
>>> clue@ietf.org
>>> https://www.ietf.org/mailman/listinfo/clue
>>>
>>>
>> _______________________________________________
>> clue mailing list
>> clue@ietf.org
>> https://www.ietf.org/mailman/listinfo/clue
>>
>>
> _______________________________________________
> clue mailing list
> clue@ietf.org
> https://www.ietf.org/mailman/listinfo/clue
>

--047d7bf10b50a05c3504f7a11b7e
Content-Type: text/html; charset=UTF-8
Content-Transfer-Encoding: quoted-printable

<div dir=3D"ltr">Yes, I changed the Webex start time.</div><div class=3D"gm=
ail_extra"><br><br><div class=3D"gmail_quote">On Tue, Apr 22, 2014 at 7:51 =
AM, Paul Kyzivat <span dir=3D"ltr">&lt;<a href=3D"mailto:pkyzivat@alum.mit.=
edu" target=3D"_blank">pkyzivat@alum.mit.edu</a>&gt;</span> wrote:<br>
<blockquote class=3D"gmail_quote" style=3D"margin:0 0 0 .8ex;border-left:1p=
x #ccc solid;padding-left:1ex">I checked again, and even though the webex s=
hows on the schedule for 8:30, when clicking through to it (at <a href=3D"h=
ttps://ietf.webex.com/mw0307l/mywebex/default.do?siteurl=3Dietf" target=3D"=
_blank">https://ietf.webex.com/<u></u>mw0307l/mywebex/default.do?<u></u>sit=
eurl=3Dietf</a>), it is set to start at 8 central. So I guess it is ok.<br>

<br>
And I&#39;ve been corrected on the subject.<div class=3D"HOEnZb"><div class=
=3D"h5"><br>
<br>
On 4/22/14 8:47 AM, Paul Kyzivat wrote:<br>
<blockquote class=3D"gmail_quote" style=3D"margin:0 0 0 .8ex;border-left:1p=
x #ccc solid;padding-left:1ex">
<br>
<br>
On 4/22/14 2:54 AM, Christian Groves wrote:<br>
<blockquote class=3D"gmail_quote" style=3D"margin:0 0 0 .8ex;border-left:1p=
x #ccc solid;padding-left:1ex">
Hello Paul,<br>
<br>
Thanks, by the way the CLUE wiki page<br>
(<a href=3D"http://trac.tools.ietf.org/wg/clue/trac/wiki" target=3D"_blank"=
>http://trac.tools.ietf.org/<u></u>wg/clue/trac/wiki</a>) still indicates 8=
.30am<br>
Central. Also the webex link at<br>
<a href=3D"http://trac.tools.ietf.org/wg/clue/trac/wiki/Design-Team" target=
=3D"_blank">http://trac.tools.ietf.org/wg/<u></u>clue/trac/wiki/Design-Team=
</a> indicates that<br>
the time is 8.30am. These probably should be updated.<br>
</blockquote>
<br>
Hmm. I&#39;m looking at the last url above, and it says 8am.<br>
<br>
BUT, the webex is still set up for 8:30. :-(<br>
I&#39;m trying to figure out how to fix it. If I can&#39;t get it fixed, th=
en we<br>
will have to start at 8:30.<br>
<br>
=C2=A0 =C2=A0 =C2=A0Thanks,<br>
=C2=A0 =C2=A0 =C2=A0Paul<br>
<br>
<blockquote class=3D"gmail_quote" style=3D"margin:0 0 0 .8ex;border-left:1p=
x #ccc solid;padding-left:1ex">
Regards, Christian<br>
<br>
On 22/04/2014 12:40 PM, Paul Kyzivat wrote:<br>
<blockquote class=3D"gmail_quote" style=3D"margin:0 0 0 .8ex;border-left:1p=
x #ccc solid;padding-left:1ex">
On 4/21/14 8:44 PM, Christian Groves wrote:<br>
<blockquote class=3D"gmail_quote" style=3D"margin:0 0 0 .8ex;border-left:1p=
x #ccc solid;padding-left:1ex">
Hello Paul,<br>
<br>
What time is it? Is there any change as a result of the doodle poll?<br>
</blockquote>
<br>
The schedule of design team meetings was updated. It is now a half<br>
hour earlier than it was - 8am central US time.<br>
<br>
<blockquote class=3D"gmail_quote" style=3D"margin:0 0 0 .8ex;border-left:1p=
x #ccc solid;padding-left:1ex">
Christian<br>
<br>
On 22/04/2014 10:25 AM, Paul Kyzivat wrote:<br>
<blockquote class=3D"gmail_quote" style=3D"margin:0 0 0 .8ex;border-left:1p=
x #ccc solid;padding-left:1ex">
This is just a reminder. We will be having the design team meeting<br>
tomorrow. The primary subject will be the recent revision of the<br>
protocol document.<br>
<br>
=C2=A0 =C2=A0 Thanks,<br>
=C2=A0 =C2=A0 Paul<br>
<br>
______________________________<u></u>_________________<br>
clue mailing list<br>
<a href=3D"mailto:clue@ietf.org" target=3D"_blank">clue@ietf.org</a><br>
<a href=3D"https://www.ietf.org/mailman/listinfo/clue" target=3D"_blank">ht=
tps://www.ietf.org/mailman/<u></u>listinfo/clue</a><br>
<br>
</blockquote>
<br>
______________________________<u></u>_________________<br>
clue mailing list<br>
<a href=3D"mailto:clue@ietf.org" target=3D"_blank">clue@ietf.org</a><br>
<a href=3D"https://www.ietf.org/mailman/listinfo/clue" target=3D"_blank">ht=
tps://www.ietf.org/mailman/<u></u>listinfo/clue</a><br>
<br>
</blockquote>
<br>
______________________________<u></u>_________________<br>
clue mailing list<br>
<a href=3D"mailto:clue@ietf.org" target=3D"_blank">clue@ietf.org</a><br>
<a href=3D"https://www.ietf.org/mailman/listinfo/clue" target=3D"_blank">ht=
tps://www.ietf.org/mailman/<u></u>listinfo/clue</a><br>
<br>
</blockquote>
<br>
______________________________<u></u>_________________<br>
clue mailing list<br>
<a href=3D"mailto:clue@ietf.org" target=3D"_blank">clue@ietf.org</a><br>
<a href=3D"https://www.ietf.org/mailman/listinfo/clue" target=3D"_blank">ht=
tps://www.ietf.org/mailman/<u></u>listinfo/clue</a><br>
<br>
</blockquote>
<br>
______________________________<u></u>_________________<br>
clue mailing list<br>
<a href=3D"mailto:clue@ietf.org" target=3D"_blank">clue@ietf.org</a><br>
<a href=3D"https://www.ietf.org/mailman/listinfo/clue" target=3D"_blank">ht=
tps://www.ietf.org/mailman/<u></u>listinfo/clue</a><br>
<br>
</blockquote>
<br>
______________________________<u></u>_________________<br>
clue mailing list<br>
<a href=3D"mailto:clue@ietf.org" target=3D"_blank">clue@ietf.org</a><br>
<a href=3D"https://www.ietf.org/mailman/listinfo/clue" target=3D"_blank">ht=
tps://www.ietf.org/mailman/<u></u>listinfo/clue</a><br>
</div></div></blockquote></div><br></div>

--047d7bf10b50a05c3504f7a11b7e--


From nobody Tue Apr 22 06:03:52 2014
Return-Path: <mary.ietf.barnes@gmail.com>
X-Original-To: clue@ietfa.amsl.com
Delivered-To: clue@ietfa.amsl.com
Received: from localhost (ietfa.amsl.com [127.0.0.1]) by ietfa.amsl.com (Postfix) with ESMTP id 1965D1A0442 for <clue@ietfa.amsl.com>; Tue, 22 Apr 2014 06:03:51 -0700 (PDT)
X-Virus-Scanned: amavisd-new at amsl.com
X-Spam-Flag: NO
X-Spam-Score: -1.999
X-Spam-Level: 
X-Spam-Status: No, score=-1.999 tagged_above=-999 required=5 tests=[BAYES_00=-1.9, DKIM_SIGNED=0.1, DKIM_VALID=-0.1, DKIM_VALID_AU=-0.1, FREEMAIL_FROM=0.001, HTML_MESSAGE=0.001, SPF_PASS=-0.001] autolearn=ham
Received: from mail.ietf.org ([4.31.198.44]) by localhost (ietfa.amsl.com [127.0.0.1]) (amavisd-new, port 10024) with ESMTP id IrYNwd-uDhMk for <clue@ietfa.amsl.com>; Tue, 22 Apr 2014 06:03:46 -0700 (PDT)
Received: from mail-wi0-x22d.google.com (mail-wi0-x22d.google.com [IPv6:2a00:1450:400c:c05::22d]) by ietfa.amsl.com (Postfix) with ESMTP id 61C0E1A0440 for <clue@ietf.org>; Tue, 22 Apr 2014 06:03:46 -0700 (PDT)
Received: by mail-wi0-f173.google.com with SMTP id z2so3211112wiv.12 for <clue@ietf.org>; Tue, 22 Apr 2014 06:03:40 -0700 (PDT)
DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20120113;  h=mime-version:in-reply-to:references:date:message-id:subject:from:to :cc:content-type; bh=DoCruFPDJETGy3eq0PjH72lQaT7tsOGTOCrOLtqpZ08=; b=ONZ76Sw3SMUVGU4NjNm3yBBfkbgr+ISl5CsTm87BeFZZ0GhI3LbymAhXNoB5zxjrV2 MrXPwVCPCX+UOTMHB2yhzzqv8u/TnamUNbJtMSS2cGG/KBlXqVKuo5v+tnJd25R3PPUF fFuD39EFJtQ4rfC1yCpGq4KEwVGJpVTdcehbPpSrNJN6/HYWqQg5g2Netx/iPz7zqjNj f5utxOz8/Xg4CBPKt3CqQdh2Fh6k34mHWHyKyvaXYV7kc64tf1SEJ1RRuF6Ns9fgI7a8 iNT2ldZwDSfQA7VyFWe50A2t0JVqNLSvC98OzT99bQdP/Lue/hZMvGfSMfstlvcB/gfB wVZA==
MIME-Version: 1.0
X-Received: by 10.180.77.165 with SMTP id t5mr18994713wiw.38.1398171820582; Tue, 22 Apr 2014 06:03:40 -0700 (PDT)
Received: by 10.216.10.6 with HTTP; Tue, 22 Apr 2014 06:03:40 -0700 (PDT)
In-Reply-To: <CAHBDyN6qBBeMROkm6eAZR6UTaqiXyW90yv3-PV_2g0_xN5a=4A@mail.gmail.com>
References: <5355B6FB.6050501@alum.mit.edu> <5355BB80.7010704@nteczone.com> <5355D69D.1040603@alum.mit.edu> <D5923425-CAC2-4C37-90D4-DC8DA7F42D6E@unina.it> <CAHBDyN6qBBeMROkm6eAZR6UTaqiXyW90yv3-PV_2g0_xN5a=4A@mail.gmail.com>
Date: Tue, 22 Apr 2014 08:03:40 -0500
Message-ID: <CAHBDyN7UsLKZoimRWEWwUj1-HpJhfrZAoO99nOxtM5QtsOmguA@mail.gmail.com>
From: Mary Barnes <mary.ietf.barnes@gmail.com>
To: Simon Pietro Romano <spromano@unina.it>
Content-Type: multipart/alternative; boundary=f46d043c085ab3864f04f7a13eb0
Archived-At: http://mailarchive.ietf.org/arch/msg/clue/V5l9PzmZOqqagLD1ykJcCUau4as
Cc: CLUE <clue@ietf.org>
Subject: Re: [clue] Reminder: design team meeting tomorrow Tuesday, Apr 22
X-BeenThere: clue@ietf.org
X-Mailman-Version: 2.1.15
Precedence: list
List-Id: CLUE - ControLling mUltiple streams for TElepresence <clue.ietf.org>
List-Unsubscribe: <https://www.ietf.org/mailman/options/clue>, <mailto:clue-request@ietf.org?subject=unsubscribe>
List-Archive: <http://www.ietf.org/mail-archive/web/clue/>
List-Post: <mailto:clue@ietf.org>
List-Help: <mailto:clue-request@ietf.org?subject=help>
List-Subscribe: <https://www.ietf.org/mailman/listinfo/clue>, <mailto:clue-request@ietf.org?subject=subscribe>
X-List-Received-Date: Tue, 22 Apr 2014 13:03:51 -0000

--f46d043c085ab3864f04f7a13eb0
Content-Type: text/plain; charset=UTF-8
Content-Transfer-Encoding: quoted-printable

I just realized the confusion - the reference should have been to the
updated "signaling" document.


On Tue, Apr 22, 2014 at 7:44 AM, Mary Barnes <mary.ietf.barnes@gmail.com>wr=
ote:

> That is correct.
>
>
> On Tue, Apr 22, 2014 at 3:32 AM, Simon Pietro Romano <spromano@unina.it>w=
rote:
>
>> Hi Paul,
>>
>> protocol updates are scheduled for April 29th, right? This is what we
>> have on our schedule and also what appears on the wiki:
>>
>> http://trac.tools.ietf.org/wg/clue/trac/wiki/Design-Team
>>
>> Can you confirm this is the right agenda?
>>
>> Thanx,
>>
>> Simon
>>
>> On 22/apr/2014, at 04:40, Paul Kyzivat wrote:
>>
>> On 4/21/14 8:44 PM, Christian Groves wrote:
>>
>> Hello Paul,
>>
>>
>> What time is it? Is there any change as a result of the doodle poll?
>>
>>
>> The schedule of design team meetings was updated. It is now a half hour
>> earlier than it was - 8am central US time.
>>
>> Christian
>>
>>
>> On 22/04/2014 10:25 AM, Paul Kyzivat wrote:
>>
>> This is just a reminder. We will be having the design team meeting
>>
>> tomorrow. The primary subject will be the recent revision of the
>>
>> protocol document.
>>
>>
>>    Thanks,
>>
>>    Paul
>>
>>
>> _______________________________________________
>>
>> clue mailing list
>>
>> clue@ietf.org
>>
>> https://www.ietf.org/mailman/listinfo/clue
>>
>>
>>
>> _______________________________________________
>>
>> clue mailing list
>>
>> clue@ietf.org
>>
>> https://www.ietf.org/mailman/listinfo/clue
>>
>>
>>
>> _______________________________________________
>> clue mailing list
>> clue@ietf.org
>> https://www.ietf.org/mailman/listinfo/clue
>>
>>
>>                               _\\|//_
>>                                   ( O-O )
>>    ~~~~~~~~~~~~~~~~~~~~~~o00~~(_)~~00o~~~~~~~~~~~~~~~~~~~~~~~~
>>                      Simon Pietro Romano
>>                Universita' di Napoli Federico II
>>                       Computer Engineering Department
>>              Phone: +39 081 7683823 -- Fax: +39 081 7683816
>>                                            e-mail: spromano@unina.it
>>
>>     <<Molti mi dicono che lo scoraggiamento =C3=A8 l'alibi degli
>>      idioti. Ci rifletto un istante; e mi scoraggio>>. Magritte.
>>                                      oooO
>>   ~~~~~~~~~~~~~~~~~~~~~~~(   )~~~ Oooo~~~~~~~~~~~~~~~~~~~~~~~~~
>>                  \ (            (   )
>>                                   \_)          ) /
>>                                                                        (=
_/
>>
>>
>>
>>
>>
>>
>>
>> _______________________________________________
>> clue mailing list
>> clue@ietf.org
>> https://www.ietf.org/mailman/listinfo/clue
>>
>>
>

--f46d043c085ab3864f04f7a13eb0
Content-Type: text/html; charset=UTF-8
Content-Transfer-Encoding: quoted-printable

<div dir=3D"ltr">I just realized the confusion - the reference should have =
been to the updated &quot;signaling&quot; document.</div><div class=3D"gmai=
l_extra"><br><br><div class=3D"gmail_quote">On Tue, Apr 22, 2014 at 7:44 AM=
, Mary Barnes <span dir=3D"ltr">&lt;<a href=3D"mailto:mary.ietf.barnes@gmai=
l.com" target=3D"_blank">mary.ietf.barnes@gmail.com</a>&gt;</span> wrote:<b=
r>
<blockquote class=3D"gmail_quote" style=3D"margin:0 0 0 .8ex;border-left:1p=
x #ccc solid;padding-left:1ex"><div dir=3D"ltr">That is correct.=C2=A0</div=
><div class=3D"HOEnZb"><div class=3D"h5"><div class=3D"gmail_extra"><br><br=
><div class=3D"gmail_quote">
On Tue, Apr 22, 2014 at 3:32 AM, Simon Pietro Romano <span dir=3D"ltr">&lt;=
<a href=3D"mailto:spromano@unina.it" target=3D"_blank">spromano@unina.it</a=
>&gt;</span> wrote:<br>
<blockquote class=3D"gmail_quote" style=3D"margin:0 0 0 .8ex;border-left:1p=
x #ccc solid;padding-left:1ex"><div style=3D"word-wrap:break-word">Hi Paul,=
<div><br></div><div>protocol updates are scheduled for April 29th, right? T=
his is what we have on our schedule and also what appears on the wiki:</div=
>

<div><br></div><div><a href=3D"http://trac.tools.ietf.org/wg/clue/trac/wiki=
/Design-Team" target=3D"_blank">http://trac.tools.ietf.org/wg/clue/trac/wik=
i/Design-Team</a></div><div><br></div><div>Can you confirm this is the righ=
t agenda?</div>

<div><br></div><div>Thanx,</div><div><br></div><div>Simon</div><div><div><d=
iv><br><div><div>On 22/apr/2014, at 04:40, Paul Kyzivat wrote:</div><br><bl=
ockquote type=3D"cite"><div>On 4/21/14 8:44 PM, Christian Groves wrote:<br>

<blockquote type=3D"cite">Hello Paul,<br></blockquote><blockquote type=3D"c=
ite"><br></blockquote><blockquote type=3D"cite">What time is it? Is there a=
ny change as a result of the doodle poll?<br></blockquote><br>The schedule =
of design team meetings was updated. It is now a half hour earlier than it =
was - 8am central US time.<br>

<br><blockquote type=3D"cite">Christian<br></blockquote><blockquote type=3D=
"cite"><br></blockquote><blockquote type=3D"cite">On 22/04/2014 10:25 AM, P=
aul Kyzivat wrote:<br></blockquote><blockquote type=3D"cite"><blockquote ty=
pe=3D"cite">

This is just a reminder. We will be having the design team meeting<br></blo=
ckquote></blockquote><blockquote type=3D"cite"><blockquote type=3D"cite">to=
morrow. The primary subject will be the recent revision of the<br></blockqu=
ote>

</blockquote><blockquote type=3D"cite"><blockquote type=3D"cite">protocol d=
ocument.<br></blockquote></blockquote><blockquote type=3D"cite"><blockquote=
 type=3D"cite"><br></blockquote></blockquote><blockquote type=3D"cite"><blo=
ckquote type=3D"cite">

 =C2=A0=C2=A0=C2=A0Thanks,<br></blockquote></blockquote><blockquote type=3D=
"cite"><blockquote type=3D"cite"> =C2=A0=C2=A0=C2=A0Paul<br></blockquote></=
blockquote><blockquote type=3D"cite"><blockquote type=3D"cite"><br></blockq=
uote></blockquote><blockquote type=3D"cite">

<blockquote type=3D"cite">_______________________________________________<b=
r></blockquote></blockquote><blockquote type=3D"cite"><blockquote type=3D"c=
ite">clue mailing list<br></blockquote></blockquote><blockquote type=3D"cit=
e">
<blockquote type=3D"cite">
<a href=3D"mailto:clue@ietf.org" target=3D"_blank">clue@ietf.org</a><br></b=
lockquote></blockquote><blockquote type=3D"cite"><blockquote type=3D"cite">=
<a href=3D"https://www.ietf.org/mailman/listinfo/clue" target=3D"_blank">ht=
tps://www.ietf.org/mailman/listinfo/clue</a><br>

</blockquote></blockquote><blockquote type=3D"cite"><blockquote type=3D"cit=
e"><br></blockquote></blockquote><blockquote type=3D"cite"><br></blockquote=
><blockquote type=3D"cite">_______________________________________________<=
br>
</blockquote>
<blockquote type=3D"cite">clue mailing list<br></blockquote><blockquote typ=
e=3D"cite"><a href=3D"mailto:clue@ietf.org" target=3D"_blank">clue@ietf.org=
</a><br></blockquote><blockquote type=3D"cite"><a href=3D"https://www.ietf.=
org/mailman/listinfo/clue" target=3D"_blank">https://www.ietf.org/mailman/l=
istinfo/clue</a><br>

</blockquote><blockquote type=3D"cite"><br></blockquote><br>_______________=
________________________________<br>clue mailing list<br><a href=3D"mailto:=
clue@ietf.org" target=3D"_blank">clue@ietf.org</a><br><a href=3D"https://ww=
w.ietf.org/mailman/listinfo/clue" target=3D"_blank">https://www.ietf.org/ma=
ilman/listinfo/clue</a><br>

<br></div></blockquote></div><br></div></div><div>
<div style=3D"word-wrap:break-word"><span style=3D"text-indent:0px;letter-s=
pacing:normal;font-variant:normal;text-align:-webkit-auto;font-style:normal=
;font-weight:normal;line-height:normal;border-collapse:separate;text-transf=
orm:none;font-size:medium;white-space:normal;font-family:Helvetica;word-spa=
cing:0px"><div style=3D"word-wrap:break-word">

<span style=3D"text-indent:0px;letter-spacing:normal;font-variant:normal;te=
xt-align:-webkit-auto;font-style:normal;font-weight:normal;line-height:norm=
al;border-collapse:separate;text-transform:none;font-size:medium;white-spac=
e:normal;font-family:Helvetica;word-spacing:0px"><div style=3D"word-wrap:br=
eak-word">

<div><div>=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =
=C2=A0 =C2=A0<span style=3D"white-space:pre-wrap">					</span><span>=C2=A0<=
/span>=C2=A0 =C2=A0 =C2=A0 _\\|//_</div><div>=C2=A0 =C2=A0 =C2=A0 =C2=A0 =
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0<span =
style=3D"white-space:pre-wrap">				</span>=C2=A0 =C2=A0 =C2=A0=C2=A0( O-O )=
</div><div>=C2=A0 =C2=A0~~~~~~~~~~~~~~~~~~~~~~o00~~(_)~~00o~~~~~~~~~~~~~~~~=
~~~~~~~~</div>

<div>=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0<=
span>=C2=A0</span><span style=3D"white-space:pre-wrap">				</span>Simon Pie=
tro Romano</div><div>=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0<span =
style=3D"white-space:pre-wrap">				</span><span>=C2=A0</span>Universita&#39=
; di Napoli Federico II</div>

<div>=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0=C2=A0<span sty=
le=3D"white-space:pre-wrap">		</span>=C2=A0 =C2=A0 =C2=A0Computer Engineeri=
ng Department=C2=A0</div><div><span style=3D"white-space:pre-wrap">	</span>=
=C2=A0 =C2=A0 =C2=A0=C2=A0 =C2=A0 =C2=A0 =C2=A0 Phone: <a href=3D"tel:%2B39=
%20081%207683823" value=3D"+390817683823" target=3D"_blank">+39 081 7683823=
</a> -- Fax: <a href=3D"tel:%2B39%20081%207683816" value=3D"+390817683816" =
target=3D"_blank">+39 081 7683816</a></div>

<div>=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=
=A0 =C2=A0e-mail: <a href=3D"mailto:spromano@unina.it" target=3D"_blank">sp=
romano@unina.it</a></div><div><br></div><div><span style=3D"white-space:pre=
-wrap">		</span>=C2=A0 =C2=A0 &lt;&lt;Molti mi dicono che lo scoraggiamento=
 =C3=A8 l&#39;alibi degli=C2=A0</div>

<div><span style=3D"white-space:pre-wrap">		</span>=C2=A0=C2=A0 =C2=A0idiot=
i. Ci rifletto un istante; e mi scoraggio&gt;&gt;. Magritte.</div><div>=C2=
=A0 =C2=A0 =C2=A0 =C2=A0=C2=A0 =C2=A0 =C2=A0 =C2=A0<span>=C2=A0</span><span=
 style=3D"white-space:pre-wrap">			</span>=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=
=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0oooO</div>

<div>=C2=A0 ~~~~~~~~~~~~~~~~~~~~~~~( =C2=A0 )~~~=C2=A0Oooo~~~~~~~~~~~~~~~~~=
~~~~~~~~</div><div><span style=3D"white-space:pre-wrap">					</span>=C2=A0 =
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0\ ( =C2=A0 =C2=A0 =
=C2=A0 =C2=A0 =C2=A0 =C2=A0( =C2=A0 )</div><div><span style=3D"white-space:=
pre-wrap">			</span>=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0=
 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 \_) =C2=A0 =
=C2=A0 =C2=A0 =C2=A0 =C2=A0) /</div>

<div>=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=
=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0(_/</div></div><div><br></div></div></spa=
n><br></div></span><br></div><br><br>
</div>
<br></div></div><br>_______________________________________________<br>
clue mailing list<br>
<a href=3D"mailto:clue@ietf.org" target=3D"_blank">clue@ietf.org</a><br>
<a href=3D"https://www.ietf.org/mailman/listinfo/clue" target=3D"_blank">ht=
tps://www.ietf.org/mailman/listinfo/clue</a><br>
<br></blockquote></div><br></div>
</div></div></blockquote></div><br></div>

--f46d043c085ab3864f04f7a13eb0--


From nobody Tue Apr 22 07:04:09 2014
Return-Path: <pkyzivat@alum.mit.edu>
X-Original-To: clue@ietfa.amsl.com
Delivered-To: clue@ietfa.amsl.com
Received: from localhost (ietfa.amsl.com [127.0.0.1]) by ietfa.amsl.com (Postfix) with ESMTP id 8F1C01A03FF for <clue@ietfa.amsl.com>; Tue, 22 Apr 2014 07:04:08 -0700 (PDT)
X-Virus-Scanned: amavisd-new at amsl.com
X-Spam-Flag: NO
X-Spam-Score: -1.235
X-Spam-Level: 
X-Spam-Status: No, score=-1.235 tagged_above=-999 required=5 tests=[BAYES_00=-1.9, DKIM_SIGNED=0.1, DKIM_VALID=-0.1, SPF_SOFTFAIL=0.665] autolearn=no
Received: from mail.ietf.org ([4.31.198.44]) by localhost (ietfa.amsl.com [127.0.0.1]) (amavisd-new, port 10024) with ESMTP id kd8P5tANSbU2 for <clue@ietfa.amsl.com>; Tue, 22 Apr 2014 07:04:04 -0700 (PDT)
Received: from qmta02.westchester.pa.mail.comcast.net (qmta02.westchester.pa.mail.comcast.net [IPv6:2001:558:fe14:43:76:96:62:24]) by ietfa.amsl.com (Postfix) with ESMTP id E961C1A047A for <clue@ietf.org>; Tue, 22 Apr 2014 07:04:03 -0700 (PDT)
Received: from omta12.westchester.pa.mail.comcast.net ([76.96.62.44]) by qmta02.westchester.pa.mail.comcast.net with comcast id t0Fc1n0030xGWP85123ywV; Tue, 22 Apr 2014 14:03:58 +0000
Received: from Paul-Kyzivats-MacBook-Pro.local ([50.138.229.164]) by omta12.westchester.pa.mail.comcast.net with comcast id t23y1n00A3ZTu2S3Y23yHE; Tue, 22 Apr 2014 14:03:58 +0000
Message-ID: <535676CE.6040902@alum.mit.edu>
Date: Tue, 22 Apr 2014 10:03:58 -0400
From: Paul Kyzivat <pkyzivat@alum.mit.edu>
User-Agent: Mozilla/5.0 (Macintosh; Intel Mac OS X 10.7; rv:24.0) Gecko/20100101 Thunderbird/24.4.0
MIME-Version: 1.0
To: Simon Pietro Romano <spromano@unina.it>
References: <5355B6FB.6050501@alum.mit.edu> <5355BB80.7010704@nteczone.com> <5355D69D.1040603@alum.mit.edu> <D5923425-CAC2-4C37-90D4-DC8DA7F42D6E@unina.it>
In-Reply-To: <D5923425-CAC2-4C37-90D4-DC8DA7F42D6E@unina.it>
Content-Type: text/plain; charset=ISO-8859-1; format=flowed
Content-Transfer-Encoding: 8bit
DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=comcast.net; s=q20140121; t=1398175438; bh=yw4Qg2AySC3ni1Cphbsw99pI+wxu3v2mOwkNqkrCkQQ=; h=Received:Received:Message-ID:Date:From:MIME-Version:To:Subject: Content-Type; b=S1UOYQ+sJWDFTXfJHOJaOTNHq+dHk8mqBWpv//s16YXj1xBOTBpsVlD2srnXhyTPJ XRUKZo9vrKRlaoSXkwCqGeD7rn7KZrndTlJJP2I1wDFbgq09nH+LdNMv9gCO6NHSoA EmyZdgu00ipRq/skwzfEK+YQ3ZhO0N+5WzfBl1u8Adwy8MaoPFiMQ19hj6OvFxcqoS rpZ/M/yEMoO6v5//OMIp9vBOFiaPWXKwaz1mk7EgWLtZA1YL/qqw3mGOXzkHzM9GwX DD+qPdVgnqEle5JkvIg0RxATP44JGQgR+0YPuBHpQ77++APZDOVWhQy2Nn9siPrG3i 3EkIRyqAQO+bg==
Archived-At: http://mailarchive.ietf.org/arch/msg/clue/3ehoVMJ_uyDqOLkv4gtkL8dVAaE
Cc: clue@ietf.org
Subject: Re: [clue] Reminder: design team meeting tomorrow Tuesday, Apr 22
X-BeenThere: clue@ietf.org
X-Mailman-Version: 2.1.15
Precedence: list
List-Id: CLUE - ControLling mUltiple streams for TElepresence <clue.ietf.org>
List-Unsubscribe: <https://www.ietf.org/mailman/options/clue>, <mailto:clue-request@ietf.org?subject=unsubscribe>
List-Archive: <http://www.ietf.org/mail-archive/web/clue/>
List-Post: <mailto:clue@ietf.org>
List-Help: <mailto:clue-request@ietf.org?subject=help>
List-Subscribe: <https://www.ietf.org/mailman/listinfo/clue>, <mailto:clue-request@ietf.org?subject=subscribe>
X-List-Received-Date: Tue, 22 Apr 2014 14:04:08 -0000

Simon,

Yes, protocol next week. I misspoke in my reminder last nite - I was 
thinking of signaling but wrote protocol. Sorry about that.

	Thanks,
	Paul

On 4/22/14 4:32 AM, Simon Pietro Romano wrote:
> Hi Paul,
>
> protocol updates are scheduled for April 29th, right? This is what we
> have on our schedule and also what appears on the wiki:
>
> http://trac.tools.ietf.org/wg/clue/trac/wiki/Design-Team
>
> Can you confirm this is the right agenda?
>
> Thanx,
>
> Simon
>
> On 22/apr/2014, at 04:40, Paul Kyzivat wrote:
>
>> On 4/21/14 8:44 PM, Christian Groves wrote:
>>> Hello Paul,
>>>
>>> What time is it? Is there any change as a result of the doodle poll?
>>
>> The schedule of design team meetings was updated. It is now a half
>> hour earlier than it was - 8am central US time.
>>
>>> Christian
>>>
>>> On 22/04/2014 10:25 AM, Paul Kyzivat wrote:
>>>> This is just a reminder. We will be having the design team meeting
>>>> tomorrow. The primary subject will be the recent revision of the
>>>> protocol document.
>>>>
>>>>    Thanks,
>>>>    Paul
>>>>
>>>> _______________________________________________
>>>> clue mailing list
>>>> clue@ietf.org <mailto:clue@ietf.org>
>>>> https://www.ietf.org/mailman/listinfo/clue
>>>>
>>>
>>> _______________________________________________
>>> clue mailing list
>>> clue@ietf.org <mailto:clue@ietf.org>
>>> https://www.ietf.org/mailman/listinfo/clue
>>>
>>
>> _______________________________________________
>> clue mailing list
>> clue@ietf.org <mailto:clue@ietf.org>
>> https://www.ietf.org/mailman/listinfo/clue
>>
>
>        _\\|//_
>        ( O-O )
>     ~~~~~~~~~~~~~~~~~~~~~~o00~~(_)~~00o~~~~~~~~~~~~~~~~~~~~~~~~
> Simon Pietro Romano
> Universita' di Napoli Federico II
>       Computer Engineering Department
>               Phone: +39 081 7683823 -- Fax: +39 081 7683816
>                                             e-mail: spromano@unina.it
> <mailto:spromano@unina.it>
>
>      <<Molti mi dicono che lo scoraggiamento è l'alibi degli
>      idioti. Ci rifletto un istante; e mi scoraggio>>. Magritte.
>                       oooO
>    ~~~~~~~~~~~~~~~~~~~~~~~(   )~~~ Oooo~~~~~~~~~~~~~~~~~~~~~~~~~
>                   \ (            (   )
>                                    \_)          ) /
>                                                                         (_/
>
>
>
>
>
>


From nobody Tue Apr 22 07:51:01 2014
Return-Path: <john@jlc.net>
X-Original-To: clue@ietfa.amsl.com
Delivered-To: clue@ietfa.amsl.com
Received: from localhost (ietfa.amsl.com [127.0.0.1]) by ietfa.amsl.com (Postfix) with ESMTP id BC9C11A04E9 for <clue@ietfa.amsl.com>; Tue, 22 Apr 2014 07:50:59 -0700 (PDT)
X-Virus-Scanned: amavisd-new at amsl.com
X-Spam-Flag: NO
X-Spam-Score: -1.772
X-Spam-Level: 
X-Spam-Status: No, score=-1.772 tagged_above=-999 required=5 tests=[BAYES_50=0.8, RCVD_IN_DNSWL_MED=-2.3, RP_MATCHES_RCVD=-0.272] autolearn=ham
Received: from mail.ietf.org ([4.31.198.44]) by localhost (ietfa.amsl.com [127.0.0.1]) (amavisd-new, port 10024) with ESMTP id mhVT2_AUaDMQ for <clue@ietfa.amsl.com>; Tue, 22 Apr 2014 07:50:55 -0700 (PDT)
Received: from mailhost.jlc.net (mailhost.jlc.net [199.201.159.4]) by ietfa.amsl.com (Postfix) with ESMTP id 0F8A21A056D for <clue@ietf.org>; Tue, 22 Apr 2014 07:49:06 -0700 (PDT)
Received: by mailhost.jlc.net (Postfix, from userid 104) id CE6D7C94BF; Tue, 22 Apr 2014 10:48:58 -0400 (EDT)
Date: Tue, 22 Apr 2014 10:48:58 -0400
From: John Leslie <john@jlc.net>
To: CLUE <clue@ietf.org>
Message-ID: <20140422144858.GC86778@verdi>
References: <533AF351.9050201@alum.mit.edu> <20140410222705.GW39240@verdi>
Mime-Version: 1.0
Content-Type: text/plain; charset=us-ascii
Content-Disposition: inline
In-Reply-To: <20140410222705.GW39240@verdi>
User-Agent: Mutt/1.4.1i
Archived-At: http://mailarchive.ietf.org/arch/msg/clue/NWktuPhBPM9gW8gTr57LdYugvBI
Subject: Re: [clue] Improving treatment of audio
X-BeenThere: clue@ietf.org
X-Mailman-Version: 2.1.15
Precedence: list
List-Id: CLUE - ControLling mUltiple streams for TElepresence <clue.ietf.org>
List-Unsubscribe: <https://www.ietf.org/mailman/options/clue>, <mailto:clue-request@ietf.org?subject=unsubscribe>
List-Archive: <http://www.ietf.org/mail-archive/web/clue/>
List-Post: <mailto:clue@ietf.org>
List-Help: <mailto:clue-request@ietf.org?subject=help>
List-Subscribe: <https://www.ietf.org/mailman/listinfo/clue>, <mailto:clue-request@ietf.org?subject=subscribe>
X-List-Received-Date: Tue, 22 Apr 2014 14:50:59 -0000

John Leslie <john@jlc.net> wrote:
> 
> Audio 101

   continuing...

Audio 102

Q: What is Echo?
A: Echo is a reflection of sound arriving with sufficient delay to be perceived separately.

Q: How is that different from reverberation?
A: It's a question of perception: when the delay is much less than 100 milliseconds, the ear doesn't perceive it as separate.

Q: But isn't reverberation sometimes measured in seconds (thousands of milliseconds)?
A: Indeed it is. Reverberation often exceeds five seconds in a large cathedral.

Q: So why isn't that "echo"?
A: No individual reflection stands out enough to be perceived as separate. "Reverberation" comes from multiple reflections off different surfaces, which can stretch out for seconds.

Q: Why is echo considered a problem when reverberation is not?
A: But reverberation _is_ a problem when it exceeds a few hundred milliseconds.

Q: Huh?
A: Excessive reverberation "muddies" the speech, and is perceived by the speaker, who speaks more slowly.

Q: Why doesn't that work for echo?
A: The ear perceives separate sound "sources" and the brain is confused by them.

Q: What did people do about this before electronics were invented?
A: They met somewhere else. ;^)

Q: When did we decide to "fix" echo?
A: We felt a need to fix it when we couldn't "meet" somewhere else: specifically when we could only "meet" by telephone.

Q: How's that?
A: As telephones became practical over long distances, they became the only way two people could "meet".

Q: But don't electric signals travel at the speed of light?
A: Actually they never go that fast; but even if they did, 900 miles would produce audible echo. By 1960, satellites were used, giving delays over 100 milliseconds.

Q: What did telephone companies do about it?
A: They made devices to detect echo and "suppress" it by reducing the signal containing echo.

Q: Did that actually work?
A: Well enough for government work ;^) -- folks learned not to talk while the other person talked.

Q: But we've improved on that, haven't we?
A: Indeed: they figured how to do "echo cancellation" instead of just "echo suppression".

Q: How does that work?
A: They predict the echo signal and "cancel" it by inserting the exact opposite.

Q: Isn't that risky -- don't you end up amplifying if you mis-guess the phase?
A: Yes.

Q: So why did it work at all?
A: Telephony was limited to about three kilohertz -- you'd have to be off by 100 microseconds to get in trouble. Besides, their problem was much simpler: no room reverberation, just electrical reflection

Q: Have we improved on that since 1960?
A: We've improved the estimates, yes; but the acoustic echo is far more complex, and we're fighting to track the changes in reflection.

Q: How do we "track" these changes?
A: We run continuous algorithms to update the estimates.

Q: Isn't that expensive?
A: We dedicate on-chip Digital Signal Processors; so the cost is manageable.

Q: What are the practical limits of constantly-updated estimates?
A: Raw processing speed, deviations from the model, and out-of-model events all can lead to bad results. Modern devices generally add non-linear processing, beyond the scope of this tutorial.

Q: Isn't processing speed always increasing due to Moore's law?
A: Yes, but 1) it's not increasing as fast as memory capacity, and 2) equipment upgrades are pretty slow.

Q: What are some typical deviations from the model?
A: Many models assume a clear distinction between originated signal and echos, and guess badly when speech overlaps in direction.

Q: What are some out-of-model events?
A: Many models assume all "interesting" echo will come within N milliseconds. N tends to be much less than 500. :^(

Q: Is there anything we can do about these problems?
A: We task each sender to analyze and cancel the acoustic echo of any sound it receives and sends to loudspeakers in the room.
This keeps the distinction clear between sound originated in the room vs. echo of what it receives (and the number of milliseconds you need to process is kept reasonably small).

Q: Does it make a difference which particular algorithm is used?
A: Yes, but mostly in overall speed and how large N can be. There are also differences in how fast they react to changes.

Q: Is there any reason for us to recommend particular algorithm(s)?
A: No. They tend to be proprietary, and new ones are introduced every year.

--
John Leslie <john@jlc.net>


From nobody Tue Apr 22 08:15:40 2014
Return-Path: <pkyzivat@alum.mit.edu>
X-Original-To: clue@ietfa.amsl.com
Delivered-To: clue@ietfa.amsl.com
Received: from localhost (ietfa.amsl.com [127.0.0.1]) by ietfa.amsl.com (Postfix) with ESMTP id 3829E1A0656 for <clue@ietfa.amsl.com>; Tue, 22 Apr 2014 08:15:37 -0700 (PDT)
X-Virus-Scanned: amavisd-new at amsl.com
X-Spam-Flag: NO
X-Spam-Score: 0.664
X-Spam-Level: 
X-Spam-Status: No, score=0.664 tagged_above=-999 required=5 tests=[BAYES_20=-0.001, DKIM_SIGNED=0.1, DKIM_VALID=-0.1, SPF_SOFTFAIL=0.665] autolearn=no
Received: from mail.ietf.org ([4.31.198.44]) by localhost (ietfa.amsl.com [127.0.0.1]) (amavisd-new, port 10024) with ESMTP id BvWy-enwxFWV for <clue@ietfa.amsl.com>; Tue, 22 Apr 2014 08:15:33 -0700 (PDT)
Received: from qmta02.westchester.pa.mail.comcast.net (qmta02.westchester.pa.mail.comcast.net [IPv6:2001:558:fe14:43:76:96:62:24]) by ietfa.amsl.com (Postfix) with ESMTP id 3F1EC1A0229 for <clue@ietf.org>; Tue, 22 Apr 2014 08:15:33 -0700 (PDT)
Received: from omta16.westchester.pa.mail.comcast.net ([76.96.62.88]) by qmta02.westchester.pa.mail.comcast.net with comcast id sztm1n0051uE5Es513FTql; Tue, 22 Apr 2014 15:15:27 +0000
Received: from Paul-Kyzivats-MacBook-Pro.local ([50.138.229.164]) by omta16.westchester.pa.mail.comcast.net with comcast id t3FT1n00Q3ZTu2S3c3FTEx; Tue, 22 Apr 2014 15:15:27 +0000
Message-ID: <5356878F.4080504@alum.mit.edu>
Date: Tue, 22 Apr 2014 11:15:27 -0400
From: Paul Kyzivat <pkyzivat@alum.mit.edu>
User-Agent: Mozilla/5.0 (Macintosh; Intel Mac OS X 10.7; rv:24.0) Gecko/20100101 Thunderbird/24.4.0
MIME-Version: 1.0
To: clue@ietf.org
References: <533AF351.9050201@alum.mit.edu> <20140410222705.GW39240@verdi> <20140422144858.GC86778@verdi>
In-Reply-To: <20140422144858.GC86778@verdi>
Content-Type: text/plain; charset=ISO-8859-1; format=flowed
Content-Transfer-Encoding: 7bit
DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=comcast.net; s=q20140121; t=1398179727; bh=GM1TLWKq7teGxfxoxGFtIfeAnIplh6RhO2efFQqh3gM=; h=Received:Received:Message-ID:Date:From:MIME-Version:To:Subject: Content-Type; b=Q1wHrtU5xh6bA7UUot/i8YhJ+R0dDcfRfch2kHDdpSErGZ5L9Gi0Pp8fQmY7IJ3b/ 7V8CrF4n0lv/oyK2sR5MwRp6O3cU4ovHWnqJEkX74rfZ+OIsDvzyb2DSf+XG3VJUIH 2n+nNw5EDNv7h4/+ZK++6EguH/ip6VbX0Vq7kAt12wH6Sdb0Urmj/CSTRewMdyLbwz sgKvN0yMbdOyUARRrwDODABLW4Hgp8EXJ/tEpe03cGJRgpD0p2wKI6KIzmJWUQh1MW Q2KkVlggG5McTRCHeBg0viRp2C85v4Wq9WhXZkmGL/L86MkYK09Pvv7aLK8mS1FVuS RF2547fMpdpSw==
Archived-At: http://mailarchive.ietf.org/arch/msg/clue/BGGfw21mgn1BvS_a0cqIxFlxoRU
Subject: Re: [clue] Improving treatment of audio
X-BeenThere: clue@ietf.org
X-Mailman-Version: 2.1.15
Precedence: list
List-Id: CLUE - ControLling mUltiple streams for TElepresence <clue.ietf.org>
List-Unsubscribe: <https://www.ietf.org/mailman/options/clue>, <mailto:clue-request@ietf.org?subject=unsubscribe>
List-Archive: <http://www.ietf.org/mail-archive/web/clue/>
List-Post: <mailto:clue@ietf.org>
List-Help: <mailto:clue-request@ietf.org?subject=help>
List-Subscribe: <https://www.ietf.org/mailman/listinfo/clue>, <mailto:clue-request@ietf.org?subject=subscribe>
X-List-Received-Date: Tue, 22 Apr 2014 15:15:37 -0000

John,

IIUC, out of this there is one thing that you are suggesting go into our 
documents:

> We task each sender to analyze and cancel the acoustic echo of any
> sound it receives and sends to loudspeakers in the room.

And then we put this issue to bed. Is that right?

	Thanks,
	Paul

On 4/22/14 10:48 AM, John Leslie wrote:
> John Leslie <john@jlc.net> wrote:
>>
>> Audio 101
>
>     continuing...
>
> Audio 102
>
> Q: What is Echo?
> A: Echo is a reflection of sound arriving with sufficient delay to be perceived separately.
>
> Q: How is that different from reverberation?
> A: It's a question of perception: when the delay is much less than 100 milliseconds, the ear doesn't perceive it as separate.
>
> Q: But isn't reverberation sometimes measured in seconds (thousands of milliseconds)?
> A: Indeed it is. Reverberation often exceeds five seconds in a large cathedral.
>
> Q: So why isn't that "echo"?
> A: No individual reflection stands out enough to be perceived as separate. "Reverberation" comes from multiple reflections off different surfaces, which can stretch out for seconds.
>
> Q: Why is echo considered a problem when reverberation is not?
> A: But reverberation _is_ a problem when it exceeds a few hundred milliseconds.
>
> Q: Huh?
> A: Excessive reverberation "muddies" the speech, and is perceived by the speaker, who speaks more slowly.
>
> Q: Why doesn't that work for echo?
> A: The ear perceives separate sound "sources" and the brain is confused by them.
>
> Q: What did people do about this before electronics were invented?
> A: They met somewhere else. ;^)
>
> Q: When did we decide to "fix" echo?
> A: We felt a need to fix it when we couldn't "meet" somewhere else: specifically when we could only "meet" by telephone.
>
> Q: How's that?
> A: As telephones became practical over long distances, they became the only way two people could "meet".
>
> Q: But don't electric signals travel at the speed of light?
> A: Actually they never go that fast; but even if they did, 900 miles would produce audible echo. By 1960, satellites were used, giving delays over 100 milliseconds.
>
> Q: What did telephone companies do about it?
> A: They made devices to detect echo and "suppress" it by reducing the signal containing echo.
>
> Q: Did that actually work?
> A: Well enough for government work ;^) -- folks learned not to talk while the other person talked.
>
> Q: But we've improved on that, haven't we?
> A: Indeed: they figured how to do "echo cancellation" instead of just "echo suppression".
>
> Q: How does that work?
> A: They predict the echo signal and "cancel" it by inserting the exact opposite.
>
> Q: Isn't that risky -- don't you end up amplifying if you mis-guess the phase?
> A: Yes.
>
> Q: So why did it work at all?
> A: Telephony was limited to about three kilohertz -- you'd have to be off by 100 microseconds to get in trouble. Besides, their problem was much simpler: no room reverberation, just electrical reflection
>
> Q: Have we improved on that since 1960?
> A: We've improved the estimates, yes; but the acoustic echo is far more complex, and we're fighting to track the changes in reflection.
>
> Q: How do we "track" these changes?
> A: We run continuous algorithms to update the estimates.
>
> Q: Isn't that expensive?
> A: We dedicate on-chip Digital Signal Processors; so the cost is manageable.
>
> Q: What are the practical limits of constantly-updated estimates?
> A: Raw processing speed, deviations from the model, and out-of-model events all can lead to bad results. Modern devices generally add non-linear processing, beyond the scope of this tutorial.
>
> Q: Isn't processing speed always increasing due to Moore's law?
> A: Yes, but 1) it's not increasing as fast as memory capacity, and 2) equipment upgrades are pretty slow.
>
> Q: What are some typical deviations from the model?
> A: Many models assume a clear distinction between originated signal and echos, and guess badly when speech overlaps in direction.
>
> Q: What are some out-of-model events?
> A: Many models assume all "interesting" echo will come within N milliseconds. N tends to be much less than 500. :^(
>
> Q: Is there anything we can do about these problems?
> A: We task each sender to analyze and cancel the acoustic echo of any sound it receives and sends to loudspeakers in the room.
> This keeps the distinction clear between sound originated in the room vs. echo of what it receives (and the number of milliseconds you need to process is kept reasonably small).
>
> Q: Does it make a difference which particular algorithm is used?
> A: Yes, but mostly in overall speed and how large N can be. There are also differences in how fast they react to changes.
>
> Q: Is there any reason for us to recommend particular algorithm(s)?
> A: No. They tend to be proprietary, and new ones are introduced every year.
>
> --
> John Leslie <john@jlc.net>
>
> _______________________________________________
> clue mailing list
> clue@ietf.org
> https://www.ietf.org/mailman/listinfo/clue
>


From nobody Tue Apr 22 09:11:02 2014
Return-Path: <internet-drafts@ietf.org>
X-Original-To: clue@ietfa.amsl.com
Delivered-To: clue@ietfa.amsl.com
Received: from localhost (ietfa.amsl.com [127.0.0.1]) by ietfa.amsl.com (Postfix) with ESMTP id 77F781A0685; Tue, 22 Apr 2014 09:10:59 -0700 (PDT)
X-Virus-Scanned: amavisd-new at amsl.com
X-Spam-Flag: NO
X-Spam-Score: -1.9
X-Spam-Level: 
X-Spam-Status: No, score=-1.9 tagged_above=-999 required=5 tests=[BAYES_00=-1.9] autolearn=ham
Received: from mail.ietf.org ([4.31.198.44]) by localhost (ietfa.amsl.com [127.0.0.1]) (amavisd-new, port 10024) with ESMTP id w-wjXolAjSXv; Tue, 22 Apr 2014 09:10:58 -0700 (PDT)
Received: from ietfa.amsl.com (localhost [IPv6:::1]) by ietfa.amsl.com (Postfix) with ESMTP id 0BF811A0668; Tue, 22 Apr 2014 09:10:58 -0700 (PDT)
MIME-Version: 1.0
Content-Type: text/plain; charset="utf-8"
Content-Transfer-Encoding: 7bit
From: internet-drafts@ietf.org
To: i-d-announce@ietf.org
X-Test-IDTracker: no
X-IETF-IDTracker: 5.3.0
Auto-Submitted: auto-generated
Precedence: bulk
Message-ID: <20140422161058.27125.83464.idtracker@ietfa.amsl.com>
Date: Tue, 22 Apr 2014 09:10:58 -0700
Archived-At: http://mailarchive.ietf.org/arch/msg/clue/tNuffiHFj_eM4vj70SzyIswjpr0
Cc: clue@ietf.org
Subject: [clue] I-D Action: draft-ietf-clue-signaling-00.txt
X-BeenThere: clue@ietf.org
X-Mailman-Version: 2.1.15
List-Id: CLUE - ControLling mUltiple streams for TElepresence <clue.ietf.org>
List-Unsubscribe: <https://www.ietf.org/mailman/options/clue>, <mailto:clue-request@ietf.org?subject=unsubscribe>
List-Archive: <http://www.ietf.org/mail-archive/web/clue/>
List-Post: <mailto:clue@ietf.org>
List-Help: <mailto:clue-request@ietf.org?subject=help>
List-Subscribe: <https://www.ietf.org/mailman/listinfo/clue>, <mailto:clue-request@ietf.org?subject=subscribe>
X-List-Received-Date: Tue, 22 Apr 2014 16:10:59 -0000

A New Internet-Draft is available from the on-line Internet-Drafts directories.
 This draft is a work item of the ControLling mUltiple streams for tElepresence Working Group of the IETF.

        Title           : CLUE Signaling
        Authors         : Paul Kyzivat
                          Lennard Xiao
                          Christian Groves
                          Robert Hansen
	Filename        : draft-ietf-clue-signaling-00.txt
	Pages           : 41
	Date            : 2014-04-22

Abstract:
   This document specifies how CLUE-specific signaling such as the CLUE
   protocol [I-D.presta-clue-protocol] and the CLUE data channel
   [I-D.ietf-clue-datachannel] are used with each other and with
   existing signaling mechanisms such as SIP and SDP to produce a
   telepresence call.


The IETF datatracker status page for this draft is:
https://datatracker.ietf.org/doc/draft-ietf-clue-signaling/

There's also a htmlized version available at:
http://tools.ietf.org/html/draft-ietf-clue-signaling-00


Please note that it may take a couple of minutes from the time of submission
until the htmlized version and diff are available at tools.ietf.org.

Internet-Drafts are also available by anonymous FTP at:
ftp://ftp.ietf.org/internet-drafts/


From nobody Tue Apr 22 10:35:05 2014
Return-Path: <stephen.botzko@gmail.com>
X-Original-To: clue@ietfa.amsl.com
Delivered-To: clue@ietfa.amsl.com
Received: from localhost (ietfa.amsl.com [127.0.0.1]) by ietfa.amsl.com (Postfix) with ESMTP id 9BED01A0223 for <clue@ietfa.amsl.com>; Tue, 22 Apr 2014 10:35:03 -0700 (PDT)
X-Virus-Scanned: amavisd-new at amsl.com
X-Spam-Flag: NO
X-Spam-Score: -1.999
X-Spam-Level: 
X-Spam-Status: No, score=-1.999 tagged_above=-999 required=5 tests=[BAYES_00=-1.9, DKIM_SIGNED=0.1, DKIM_VALID=-0.1, DKIM_VALID_AU=-0.1, FREEMAIL_FROM=0.001, HTML_MESSAGE=0.001, SPF_PASS=-0.001] autolearn=ham
Received: from mail.ietf.org ([4.31.198.44]) by localhost (ietfa.amsl.com [127.0.0.1]) (amavisd-new, port 10024) with ESMTP id x-c4IwZqJEVI for <clue@ietfa.amsl.com>; Tue, 22 Apr 2014 10:34:58 -0700 (PDT)
Received: from mail-ve0-x229.google.com (mail-ve0-x229.google.com [IPv6:2607:f8b0:400c:c01::229]) by ietfa.amsl.com (Postfix) with ESMTP id 6A64E1A069F for <clue@ietf.org>; Tue, 22 Apr 2014 10:34:57 -0700 (PDT)
Received: by mail-ve0-f169.google.com with SMTP id pa12so9935842veb.14 for <clue@ietf.org>; Tue, 22 Apr 2014 10:34:51 -0700 (PDT)
DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20120113;  h=mime-version:in-reply-to:references:date:message-id:subject:from:to :cc:content-type; bh=cKXh+bQcjlijIPPhEfQH1zrOeBydN04joAPn3fQCd0A=; b=uLAoa59MDH7phQm7t9y0LkKbQRHG/6gycDO3Ql/7de6C4qkJNtJqA2uPux9hpmQaro CdgJFrCSunKTl5RiIYdXibAjiJoCMCA6G+F7YM7jNDZGF3ixzqjPA9gWkmCDrZW24jjU EKbBru421saNR2k07C86LarwZkpHz+rcZ4JUdDk3I64zS8bwOg68dxQEb0mwT5Ke++h7 Yqp5bolZOxFMIOSTwxIf13W4uxbnu9pQ8Y1lVafrMr9QVrNa/66BW+q5AI6ap3hyhedq 22TI8nGg+3ZQNo5twl9vns0XzG1WL9LHAS+If6tzrOeYN6wSu+d40NNdHIhgZqthFS88 hM2Q==
MIME-Version: 1.0
X-Received: by 10.221.74.200 with SMTP id yx8mr37387871vcb.3.1398188091648; Tue, 22 Apr 2014 10:34:51 -0700 (PDT)
Received: by 10.221.40.135 with HTTP; Tue, 22 Apr 2014 10:34:51 -0700 (PDT)
In-Reply-To: <5356878F.4080504@alum.mit.edu>
References: <533AF351.9050201@alum.mit.edu> <20140410222705.GW39240@verdi> <20140422144858.GC86778@verdi> <5356878F.4080504@alum.mit.edu>
Date: Tue, 22 Apr 2014 13:34:51 -0400
Message-ID: <CAMC7SJ7pu6ChnZuEaby496fQuiuk3jLHfTuLOZiifRpJaBoBXQ@mail.gmail.com>
From: Stephen Botzko <stephen.botzko@gmail.com>
To: Paul Kyzivat <pkyzivat@alum.mit.edu>
Content-Type: multipart/alternative; boundary=001a1134a56e8832db04f7a508b4
Archived-At: http://mailarchive.ietf.org/arch/msg/clue/ZJllywHKehvcIBO9a951mprbP5A
Cc: CLUE <clue@ietf.org>
Subject: Re: [clue] Improving treatment of audio
X-BeenThere: clue@ietf.org
X-Mailman-Version: 2.1.15
Precedence: list
List-Id: CLUE - ControLling mUltiple streams for TElepresence <clue.ietf.org>
List-Unsubscribe: <https://www.ietf.org/mailman/options/clue>, <mailto:clue-request@ietf.org?subject=unsubscribe>
List-Archive: <http://www.ietf.org/mail-archive/web/clue/>
List-Post: <mailto:clue@ietf.org>
List-Help: <mailto:clue-request@ietf.org?subject=help>
List-Subscribe: <https://www.ietf.org/mailman/listinfo/clue>, <mailto:clue-request@ietf.org?subject=subscribe>
X-List-Received-Date: Tue, 22 Apr 2014 17:35:03 -0000

--001a1134a56e8832db04f7a508b4
Content-Type: text/plain; charset=UTF-8

As I tried to say before, acoustic echo cancellation has been provided in
speakerphones and videoconferencing systems for decades. I am still not
seeing anything here that CLUE needs to deal with, and nothing unique to
telepresence.

But I absolutely agree with the key phrase "we put this issue to bed".

>>>We task each sender to analyze and cancel the acoustic echo of any sound
it receives and sends to loudspeakers in the room.

Paul, do you have a specific document in mind where you think this would
fit?  And John - would this really end it, or will there be other warning
labels that you will want to add somewhere?

BR
Stephen


On Tue, Apr 22, 2014 at 11:15 AM, Paul Kyzivat <pkyzivat@alum.mit.edu>wrote:

> John,
>
> IIUC, out of this there is one thing that you are suggesting go into our
> documents:
>
>
>  We task each sender to analyze and cancel the acoustic echo of any
>> sound it receives and sends to loudspeakers in the room.
>>
>
> And then we put this issue to bed. Is that right?
>
>         Thanks,
>         Paul
>
>
> On 4/22/14 10:48 AM, John Leslie wrote:
>
>> John Leslie <john@jlc.net> wrote:
>>
>>>
>>> Audio 101
>>>
>>
>>     continuing...
>>
>> Audio 102
>>
>> Q: What is Echo?
>> A: Echo is a reflection of sound arriving with sufficient delay to be
>> perceived separately.
>>
>> Q: How is that different from reverberation?
>> A: It's a question of perception: when the delay is much less than 100
>> milliseconds, the ear doesn't perceive it as separate.
>>
>> Q: But isn't reverberation sometimes measured in seconds (thousands of
>> milliseconds)?
>> A: Indeed it is. Reverberation often exceeds five seconds in a large
>> cathedral.
>>
>> Q: So why isn't that "echo"?
>> A: No individual reflection stands out enough to be perceived as
>> separate. "Reverberation" comes from multiple reflections off different
>> surfaces, which can stretch out for seconds.
>>
>> Q: Why is echo considered a problem when reverberation is not?
>> A: But reverberation _is_ a problem when it exceeds a few hundred
>> milliseconds.
>>
>> Q: Huh?
>> A: Excessive reverberation "muddies" the speech, and is perceived by the
>> speaker, who speaks more slowly.
>>
>> Q: Why doesn't that work for echo?
>> A: The ear perceives separate sound "sources" and the brain is confused
>> by them.
>>
>> Q: What did people do about this before electronics were invented?
>> A: They met somewhere else. ;^)
>>
>> Q: When did we decide to "fix" echo?
>> A: We felt a need to fix it when we couldn't "meet" somewhere else:
>> specifically when we could only "meet" by telephone.
>>
>> Q: How's that?
>> A: As telephones became practical over long distances, they became the
>> only way two people could "meet".
>>
>> Q: But don't electric signals travel at the speed of light?
>> A: Actually they never go that fast; but even if they did, 900 miles
>> would produce audible echo. By 1960, satellites were used, giving delays
>> over 100 milliseconds.
>>
>> Q: What did telephone companies do about it?
>> A: They made devices to detect echo and "suppress" it by reducing the
>> signal containing echo.
>>
>> Q: Did that actually work?
>> A: Well enough for government work ;^) -- folks learned not to talk while
>> the other person talked.
>>
>> Q: But we've improved on that, haven't we?
>> A: Indeed: they figured how to do "echo cancellation" instead of just
>> "echo suppression".
>>
>> Q: How does that work?
>> A: They predict the echo signal and "cancel" it by inserting the exact
>> opposite.
>>
>> Q: Isn't that risky -- don't you end up amplifying if you mis-guess the
>> phase?
>> A: Yes.
>>
>> Q: So why did it work at all?
>> A: Telephony was limited to about three kilohertz -- you'd have to be off
>> by 100 microseconds to get in trouble. Besides, their problem was much
>> simpler: no room reverberation, just electrical reflection
>>
>> Q: Have we improved on that since 1960?
>> A: We've improved the estimates, yes; but the acoustic echo is far more
>> complex, and we're fighting to track the changes in reflection.
>>
>> Q: How do we "track" these changes?
>> A: We run continuous algorithms to update the estimates.
>>
>> Q: Isn't that expensive?
>> A: We dedicate on-chip Digital Signal Processors; so the cost is
>> manageable.
>>
>> Q: What are the practical limits of constantly-updated estimates?
>> A: Raw processing speed, deviations from the model, and out-of-model
>> events all can lead to bad results. Modern devices generally add non-linear
>> processing, beyond the scope of this tutorial.
>>
>> Q: Isn't processing speed always increasing due to Moore's law?
>> A: Yes, but 1) it's not increasing as fast as memory capacity, and 2)
>> equipment upgrades are pretty slow.
>>
>> Q: What are some typical deviations from the model?
>> A: Many models assume a clear distinction between originated signal and
>> echos, and guess badly when speech overlaps in direction.
>>
>> Q: What are some out-of-model events?
>> A: Many models assume all "interesting" echo will come within N
>> milliseconds. N tends to be much less than 500. :^(
>>
>> Q: Is there anything we can do about these problems?
>> A: We task each sender to analyze and cancel the acoustic echo of any
>> sound it receives and sends to loudspeakers in the room.
>> This keeps the distinction clear between sound originated in the room vs.
>> echo of what it receives (and the number of milliseconds you need to
>> process is kept reasonably small).
>>
>> Q: Does it make a difference which particular algorithm is used?
>> A: Yes, but mostly in overall speed and how large N can be. There are
>> also differences in how fast they react to changes.
>>
>> Q: Is there any reason for us to recommend particular algorithm(s)?
>> A: No. They tend to be proprietary, and new ones are introduced every
>> year.
>>
>> --
>> John Leslie <john@jlc.net>
>>
>> _______________________________________________
>> clue mailing list
>> clue@ietf.org
>> https://www.ietf.org/mailman/listinfo/clue
>>
>>
> _______________________________________________
> clue mailing list
> clue@ietf.org
> https://www.ietf.org/mailman/listinfo/clue
>

--001a1134a56e8832db04f7a508b4
Content-Type: text/html; charset=UTF-8
Content-Transfer-Encoding: quoted-printable

<div dir=3D"ltr"><div>As I tried to say before, acoustic echo cancellation =
has been provided in speakerphones and videoconferencing systems for decade=
s. I am still not seeing anything here that CLUE needs to deal with, and no=
thing unique to telepresence.</div>
<div><br></div><div>But I absolutely agree with the key phrase &quot;we put=
 this issue to bed&quot;.<br></div><div><br></div><div>&gt;&gt;&gt;<span st=
yle=3D"color:rgb(80,0,80);font-family:arial,sans-serif;font-size:13px">We t=
ask each sender to analyze and cancel the acoustic echo of any=C2=A0</span>=
<span style=3D"color:rgb(80,0,80);font-family:arial,sans-serif;font-size:13=
px">sound it receives and sends to loudspeakers in the room.</span></div>
<div><br></div><div>Paul, do you have a specific document in mind where you=
 think this would fit? =C2=A0And John - would this really end it, or will t=
here be other warning labels that you will want to add somewhere? =C2=A0</d=
iv><div>
<br></div><div>BR</div><div>Stephen</div></div><div class=3D"gmail_extra"><=
br><br><div class=3D"gmail_quote">On Tue, Apr 22, 2014 at 11:15 AM, Paul Ky=
zivat <span dir=3D"ltr">&lt;<a href=3D"mailto:pkyzivat@alum.mit.edu" target=
=3D"_blank">pkyzivat@alum.mit.edu</a>&gt;</span> wrote:<br>
<blockquote class=3D"gmail_quote" style=3D"margin:0 0 0 .8ex;border-left:1p=
x #ccc solid;padding-left:1ex">John,<br>
<br>
IIUC, out of this there is one thing that you are suggesting go into our do=
cuments:<div class=3D""><br>
<br>
<blockquote class=3D"gmail_quote" style=3D"margin:0 0 0 .8ex;border-left:1p=
x #ccc solid;padding-left:1ex">
We task each sender to analyze and cancel the acoustic echo of any<br>
sound it receives and sends to loudspeakers in the room.<br>
</blockquote>
<br></div>
And then we put this issue to bed. Is that right?<br>
<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 Thanks,<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 Paul<div class=3D"HOEnZb"><div class=3D"h5"><br=
>
<br>
On 4/22/14 10:48 AM, John Leslie wrote:<br>
<blockquote class=3D"gmail_quote" style=3D"margin:0 0 0 .8ex;border-left:1p=
x #ccc solid;padding-left:1ex">
John Leslie &lt;<a href=3D"mailto:john@jlc.net" target=3D"_blank">john@jlc.=
net</a>&gt; wrote:<br>
<blockquote class=3D"gmail_quote" style=3D"margin:0 0 0 .8ex;border-left:1p=
x #ccc solid;padding-left:1ex">
<br>
Audio 101<br>
</blockquote>
<br>
=C2=A0 =C2=A0 continuing...<br>
<br>
Audio 102<br>
<br>
Q: What is Echo?<br>
A: Echo is a reflection of sound arriving with sufficient delay to be perce=
ived separately.<br>
<br>
Q: How is that different from reverberation?<br>
A: It&#39;s a question of perception: when the delay is much less than 100 =
milliseconds, the ear doesn&#39;t perceive it as separate.<br>
<br>
Q: But isn&#39;t reverberation sometimes measured in seconds (thousands of =
milliseconds)?<br>
A: Indeed it is. Reverberation often exceeds five seconds in a large cathed=
ral.<br>
<br>
Q: So why isn&#39;t that &quot;echo&quot;?<br>
A: No individual reflection stands out enough to be perceived as separate. =
&quot;Reverberation&quot; comes from multiple reflections off different sur=
faces, which can stretch out for seconds.<br>
<br>
Q: Why is echo considered a problem when reverberation is not?<br>
A: But reverberation _is_ a problem when it exceeds a few hundred milliseco=
nds.<br>
<br>
Q: Huh?<br>
A: Excessive reverberation &quot;muddies&quot; the speech, and is perceived=
 by the speaker, who speaks more slowly.<br>
<br>
Q: Why doesn&#39;t that work for echo?<br>
A: The ear perceives separate sound &quot;sources&quot; and the brain is co=
nfused by them.<br>
<br>
Q: What did people do about this before electronics were invented?<br>
A: They met somewhere else. ;^)<br>
<br>
Q: When did we decide to &quot;fix&quot; echo?<br>
A: We felt a need to fix it when we couldn&#39;t &quot;meet&quot; somewhere=
 else: specifically when we could only &quot;meet&quot; by telephone.<br>
<br>
Q: How&#39;s that?<br>
A: As telephones became practical over long distances, they became the only=
 way two people could &quot;meet&quot;.<br>
<br>
Q: But don&#39;t electric signals travel at the speed of light?<br>
A: Actually they never go that fast; but even if they did, 900 miles would =
produce audible echo. By 1960, satellites were used, giving delays over 100=
 milliseconds.<br>
<br>
Q: What did telephone companies do about it?<br>
A: They made devices to detect echo and &quot;suppress&quot; it by reducing=
 the signal containing echo.<br>
<br>
Q: Did that actually work?<br>
A: Well enough for government work ;^) -- folks learned not to talk while t=
he other person talked.<br>
<br>
Q: But we&#39;ve improved on that, haven&#39;t we?<br>
A: Indeed: they figured how to do &quot;echo cancellation&quot; instead of =
just &quot;echo suppression&quot;.<br>
<br>
Q: How does that work?<br>
A: They predict the echo signal and &quot;cancel&quot; it by inserting the =
exact opposite.<br>
<br>
Q: Isn&#39;t that risky -- don&#39;t you end up amplifying if you mis-guess=
 the phase?<br>
A: Yes.<br>
<br>
Q: So why did it work at all?<br>
A: Telephony was limited to about three kilohertz -- you&#39;d have to be o=
ff by 100 microseconds to get in trouble. Besides, their problem was much s=
impler: no room reverberation, just electrical reflection<br>
<br>
Q: Have we improved on that since 1960?<br>
A: We&#39;ve improved the estimates, yes; but the acoustic echo is far more=
 complex, and we&#39;re fighting to track the changes in reflection.<br>
<br>
Q: How do we &quot;track&quot; these changes?<br>
A: We run continuous algorithms to update the estimates.<br>
<br>
Q: Isn&#39;t that expensive?<br>
A: We dedicate on-chip Digital Signal Processors; so the cost is manageable=
.<br>
<br>
Q: What are the practical limits of constantly-updated estimates?<br>
A: Raw processing speed, deviations from the model, and out-of-model events=
 all can lead to bad results. Modern devices generally add non-linear proce=
ssing, beyond the scope of this tutorial.<br>
<br>
Q: Isn&#39;t processing speed always increasing due to Moore&#39;s law?<br>
A: Yes, but 1) it&#39;s not increasing as fast as memory capacity, and 2) e=
quipment upgrades are pretty slow.<br>
<br>
Q: What are some typical deviations from the model?<br>
A: Many models assume a clear distinction between originated signal and ech=
os, and guess badly when speech overlaps in direction.<br>
<br>
Q: What are some out-of-model events?<br>
A: Many models assume all &quot;interesting&quot; echo will come within N m=
illiseconds. N tends to be much less than 500. :^(<br>
<br>
Q: Is there anything we can do about these problems?<br>
A: We task each sender to analyze and cancel the acoustic echo of any sound=
 it receives and sends to loudspeakers in the room.<br>
This keeps the distinction clear between sound originated in the room vs. e=
cho of what it receives (and the number of milliseconds you need to process=
 is kept reasonably small).<br>
<br>
Q: Does it make a difference which particular algorithm is used?<br>
A: Yes, but mostly in overall speed and how large N can be. There are also =
differences in how fast they react to changes.<br>
<br>
Q: Is there any reason for us to recommend particular algorithm(s)?<br>
A: No. They tend to be proprietary, and new ones are introduced every year.=
<br>
<br>
--<br>
John Leslie &lt;<a href=3D"mailto:john@jlc.net" target=3D"_blank">john@jlc.=
net</a>&gt;<br>
<br>
______________________________<u></u>_________________<br>
clue mailing list<br>
<a href=3D"mailto:clue@ietf.org" target=3D"_blank">clue@ietf.org</a><br>
<a href=3D"https://www.ietf.org/mailman/listinfo/clue" target=3D"_blank">ht=
tps://www.ietf.org/mailman/<u></u>listinfo/clue</a><br>
<br>
</blockquote>
<br>
______________________________<u></u>_________________<br>
clue mailing list<br>
<a href=3D"mailto:clue@ietf.org" target=3D"_blank">clue@ietf.org</a><br>
<a href=3D"https://www.ietf.org/mailman/listinfo/clue" target=3D"_blank">ht=
tps://www.ietf.org/mailman/<u></u>listinfo/clue</a><br>
</div></div></blockquote></div><br></div>

--001a1134a56e8832db04f7a508b4--


From nobody Tue Apr 22 21:56:24 2014
Return-Path: <Christian.Groves@nteczone.com>
X-Original-To: clue@ietfa.amsl.com
Delivered-To: clue@ietfa.amsl.com
Received: from localhost (ietfa.amsl.com [127.0.0.1]) by ietfa.amsl.com (Postfix) with ESMTP id C99781A0062 for <clue@ietfa.amsl.com>; Tue, 22 Apr 2014 21:56:21 -0700 (PDT)
X-Virus-Scanned: amavisd-new at amsl.com
X-Spam-Flag: NO
X-Spam-Score: 1.5
X-Spam-Level: *
X-Spam-Status: No, score=1.5 tagged_above=-999 required=5 tests=[BAYES_60=1.5] autolearn=no
Received: from mail.ietf.org ([4.31.198.44]) by localhost (ietfa.amsl.com [127.0.0.1]) (amavisd-new, port 10024) with ESMTP id HG_fIKNgWhOf for <clue@ietfa.amsl.com>; Tue, 22 Apr 2014 21:56:19 -0700 (PDT)
Received: from cserver5.myshophosting.com (cserver5.myshophosting.com [175.107.161.1]) by ietfa.amsl.com (Postfix) with ESMTP id A7E961A0051 for <clue@ietf.org>; Tue, 22 Apr 2014 21:56:19 -0700 (PDT)
Received: from ppp118-209-17-246.lns20.mel4.internode.on.net ([118.209.17.246]:58290 helo=[127.0.0.1]) by cserver5.myshophosting.com with esmtpsa (TLSv1:DHE-RSA-AES128-SHA:128) (Exim 4.82) (envelope-from <Christian.Groves@nteczone.com>) id 1WcpEA-0003Pw-JT for clue@ietf.org; Wed, 23 Apr 2014 14:56:10 +1000
Message-ID: <535747E8.4020602@nteczone.com>
Date: Wed, 23 Apr 2014 14:56:08 +1000
From: Christian Groves <Christian.Groves@nteczone.com>
User-Agent: Mozilla/5.0 (Windows NT 6.1; WOW64; rv:24.0) Gecko/20100101 Thunderbird/24.4.0
MIME-Version: 1.0
To: "clue@ietf.org" <clue@ietf.org>
Content-Type: text/plain; charset=ISO-8859-1; format=flowed
Content-Transfer-Encoding: 7bit
X-AntiAbuse: This header was added to track abuse, please include it with any abuse report
X-AntiAbuse: Primary Hostname - cserver5.myshophosting.com
X-AntiAbuse: Original Domain - ietf.org
X-AntiAbuse: Originator/Caller UID/GID - [47 12] / [47 12]
X-AntiAbuse: Sender Address Domain - nteczone.com
X-Get-Message-Sender-Via: cserver5.myshophosting.com: authenticated_id: christian.groves@nteczone.com
X-Source: 
X-Source-Args: 
X-Source-Dir: 
Archived-At: http://mailarchive.ietf.org/arch/msg/clue/4s5sx0bk9oWeFkHWdMWk3GKULU4
Subject: [clue] Data Model Nit
X-BeenThere: clue@ietf.org
X-Mailman-Version: 2.1.15
Precedence: list
List-Id: CLUE - ControLling mUltiple streams for TElepresence <clue.ietf.org>
List-Unsubscribe: <https://www.ietf.org/mailman/options/clue>, <mailto:clue-request@ietf.org?subject=unsubscribe>
List-Archive: <http://www.ietf.org/mail-archive/web/clue/>
List-Post: <mailto:clue@ietf.org>
List-Help: <mailto:clue-request@ietf.org?subject=help>
List-Subscribe: <https://www.ietf.org/mailman/listinfo/clue>, <mailto:clue-request@ietf.org?subject=subscribe>
X-List-Received-Date: Wed, 23 Apr 2014 04:56:22 -0000

Hello Roberta,

In the example in section 22 of the data model there appears to be a nit:

<sceneEntry mediaType="audio" sceneEntryID="SE4">
                     <mediaCaptureIDs>
                         <captureIDREF>VC4</captureIDREF>
                     </mediaCaptureIDs>
                 </sceneEntry>

Should VC4 here be AC0?


Regards, Christian


From nobody Tue Apr 22 23:11:35 2014
Return-Path: <Christian.Groves@nteczone.com>
X-Original-To: clue@ietfa.amsl.com
Delivered-To: clue@ietfa.amsl.com
Received: from localhost (ietfa.amsl.com [127.0.0.1]) by ietfa.amsl.com (Postfix) with ESMTP id C34871A0322 for <clue@ietfa.amsl.com>; Tue, 22 Apr 2014 23:11:09 -0700 (PDT)
X-Virus-Scanned: amavisd-new at amsl.com
X-Spam-Flag: NO
X-Spam-Score: 0.8
X-Spam-Level: 
X-Spam-Status: No, score=0.8 tagged_above=-999 required=5 tests=[BAYES_50=0.8] autolearn=ham
Received: from mail.ietf.org ([4.31.198.44]) by localhost (ietfa.amsl.com [127.0.0.1]) (amavisd-new, port 10024) with ESMTP id pqM4q36xt6pB for <clue@ietfa.amsl.com>; Tue, 22 Apr 2014 23:10:53 -0700 (PDT)
Received: from cserver5.myshophosting.com (cserver5.myshophosting.com [175.107.161.1]) by ietfa.amsl.com (Postfix) with ESMTP id 63A881A008D for <clue@ietf.org>; Tue, 22 Apr 2014 23:10:53 -0700 (PDT)
Received: from ppp118-209-17-246.lns20.mel4.internode.on.net ([118.209.17.246]:59795 helo=[127.0.0.1]) by cserver5.myshophosting.com with esmtpsa (TLSv1:DHE-RSA-AES128-SHA:128) (Exim 4.82) (envelope-from <Christian.Groves@nteczone.com>) id 1WcqOM-0005Ll-9g for clue@ietf.org; Wed, 23 Apr 2014 16:10:46 +1000
Message-ID: <53575963.3010809@nteczone.com>
Date: Wed, 23 Apr 2014 16:10:43 +1000
From: Christian Groves <Christian.Groves@nteczone.com>
User-Agent: Mozilla/5.0 (Windows NT 6.1; WOW64; rv:24.0) Gecko/20100101 Thunderbird/24.4.0
MIME-Version: 1.0
To: "clue@ietf.org" <clue@ietf.org>
References: <535747E8.4020602@nteczone.com>
In-Reply-To: <535747E8.4020602@nteczone.com>
Content-Type: text/plain; charset=ISO-8859-1; format=flowed
Content-Transfer-Encoding: 7bit
X-AntiAbuse: This header was added to track abuse, please include it with any abuse report
X-AntiAbuse: Primary Hostname - cserver5.myshophosting.com
X-AntiAbuse: Original Domain - ietf.org
X-AntiAbuse: Originator/Caller UID/GID - [47 12] / [47 12]
X-AntiAbuse: Sender Address Domain - nteczone.com
X-Get-Message-Sender-Via: cserver5.myshophosting.com: authenticated_id: christian.groves@nteczone.com
X-Source: 
X-Source-Args: 
X-Source-Dir: 
Archived-At: http://mailarchive.ietf.org/arch/msg/clue/4wpaxlXOhSbHX962f9D5XtPC4bs
Subject: Re: [clue] Data Model Nit
X-BeenThere: clue@ietf.org
X-Mailman-Version: 2.1.15
Precedence: list
List-Id: CLUE - ControLling mUltiple streams for TElepresence <clue.ietf.org>
List-Unsubscribe: <https://www.ietf.org/mailman/options/clue>, <mailto:clue-request@ietf.org?subject=unsubscribe>
List-Archive: <http://www.ietf.org/mail-archive/web/clue/>
List-Post: <mailto:clue@ietf.org>
List-Help: <mailto:clue-request@ietf.org?subject=help>
List-Subscribe: <https://www.ietf.org/mailman/listinfo/clue>, <mailto:clue-request@ietf.org?subject=subscribe>
X-List-Received-Date: Wed, 23 Apr 2014 06:11:10 -0000

Hello Roberta,

A further question on the data model regarding the participantInfo 
element. In cl.7.3.1.1/[CLUE framework] it lists xcard information as an 
attribute of the CaptureScene. However I couldn't see that its covered 
by the data model?

Regards, Christian


On 23/04/2014 2:56 PM, Christian Groves wrote:
> Hello Roberta,
>
> In the example in section 22 of the data model there appears to be a nit:
>
> <sceneEntry mediaType="audio" sceneEntryID="SE4">
>                     <mediaCaptureIDs>
> <captureIDREF>VC4</captureIDREF>
>                     </mediaCaptureIDs>
>                 </sceneEntry>
>
> Should VC4 here be AC0?
>
>
> Regards, Christian
>
> _______________________________________________
> clue mailing list
> clue@ietf.org
> https://www.ietf.org/mailman/listinfo/clue
>


From nobody Wed Apr 23 00:18:24 2014
Return-Path: <Christian.Groves@nteczone.com>
X-Original-To: clue@ietfa.amsl.com
Delivered-To: clue@ietfa.amsl.com
Received: from localhost (ietfa.amsl.com [127.0.0.1]) by ietfa.amsl.com (Postfix) with ESMTP id D01FB1A00BF for <clue@ietfa.amsl.com>; Wed, 23 Apr 2014 00:18:23 -0700 (PDT)
X-Virus-Scanned: amavisd-new at amsl.com
X-Spam-Flag: NO
X-Spam-Score: -1.9
X-Spam-Level: 
X-Spam-Status: No, score=-1.9 tagged_above=-999 required=5 tests=[BAYES_00=-1.9] autolearn=ham
Received: from mail.ietf.org ([4.31.198.44]) by localhost (ietfa.amsl.com [127.0.0.1]) (amavisd-new, port 10024) with ESMTP id a70GMijjNz8E for <clue@ietfa.amsl.com>; Wed, 23 Apr 2014 00:18:19 -0700 (PDT)
Received: from cserver5.myshophosting.com (cserver5.myshophosting.com [175.107.161.1]) by ietfa.amsl.com (Postfix) with ESMTP id 509871A00BE for <clue@ietf.org>; Wed, 23 Apr 2014 00:18:19 -0700 (PDT)
Received: from ppp118-209-17-246.lns20.mel4.internode.on.net ([118.209.17.246]:61266 helo=[127.0.0.1]) by cserver5.myshophosting.com with esmtpsa (TLSv1:DHE-RSA-AES128-SHA:128) (Exim 4.82) (envelope-from <Christian.Groves@nteczone.com>) id 1WcrRc-0005vA-He for clue@ietf.org; Wed, 23 Apr 2014 17:18:12 +1000
Message-ID: <53576931.1000703@nteczone.com>
Date: Wed, 23 Apr 2014 17:18:09 +1000
From: Christian Groves <Christian.Groves@nteczone.com>
User-Agent: Mozilla/5.0 (Windows NT 6.1; WOW64; rv:24.0) Gecko/20100101 Thunderbird/24.4.0
MIME-Version: 1.0
To: clue@ietf.org
References: <535747E8.4020602@nteczone.com> <53575963.3010809@nteczone.com>
In-Reply-To: <53575963.3010809@nteczone.com>
Content-Type: text/plain; charset=ISO-8859-1; format=flowed
Content-Transfer-Encoding: 7bit
X-AntiAbuse: This header was added to track abuse, please include it with any abuse report
X-AntiAbuse: Primary Hostname - cserver5.myshophosting.com
X-AntiAbuse: Original Domain - ietf.org
X-AntiAbuse: Originator/Caller UID/GID - [47 12] / [47 12]
X-AntiAbuse: Sender Address Domain - nteczone.com
X-Get-Message-Sender-Via: cserver5.myshophosting.com: authenticated_id: christian.groves@nteczone.com
X-Source: 
X-Source-Args: 
X-Source-Dir: 
Archived-At: http://mailarchive.ietf.org/arch/msg/clue/OKv3D8lxQdUkvChkCOdJeI9hwao
Subject: Re: [clue] Data Model Nit
X-BeenThere: clue@ietf.org
X-Mailman-Version: 2.1.15
Precedence: list
List-Id: CLUE - ControLling mUltiple streams for TElepresence <clue.ietf.org>
List-Unsubscribe: <https://www.ietf.org/mailman/options/clue>, <mailto:clue-request@ietf.org?subject=unsubscribe>
List-Archive: <http://www.ietf.org/mail-archive/web/clue/>
List-Post: <mailto:clue@ietf.org>
List-Help: <mailto:clue-request@ietf.org?subject=help>
List-Subscribe: <https://www.ietf.org/mailman/listinfo/clue>, <mailto:clue-request@ietf.org?subject=subscribe>
X-List-Received-Date: Wed, 23 Apr 2014 07:18:24 -0000

And another nit:
cl.3 lists the Participant type as:
<!-- PARTICIPANT INFORMATION TYPE -->
<xs:complexType name="participantInfoType">
   <xs:sequence>
      <xs:element name="vcard" type="xcard:vcardType" maxOccurs="1"
                          minOccurs="0"/>
      <xs:element name="participantType" type="participantTypeType"
                          minOccurs="0"
                  maxOccurs="unbounded"/>
      <xs:any namespace="##other" processContents="lax" minOccurs="0"
                  maxOccurs="unbounded"/>
   </xs:sequence>
   <xs:attribute name="participantID" type="xs:ID"/>
   <xs:anyAttribute namespace="##other" processContents="lax"/>
</xs:complexType>

however the same structure in cl.19 appears to be missing the
   <xs:attribute name="participantID" type="xs:ID"/>
element.

Regards, Christian

On 23/04/2014 4:10 PM, Christian Groves wrote:
> Hello Roberta,
>
> A further question on the data model regarding the participantInfo 
> element. In cl.7.3.1.1/[CLUE framework] it lists xcard information as 
> an attribute of the CaptureScene. However I couldn't see that its 
> covered by the data model?
>
> Regards, Christian
>
>
> On 23/04/2014 2:56 PM, Christian Groves wrote:
>> Hello Roberta,
>>
>> In the example in section 22 of the data model there appears to be a 
>> nit:
>>
>> <sceneEntry mediaType="audio" sceneEntryID="SE4">
>>                     <mediaCaptureIDs>
>> <captureIDREF>VC4</captureIDREF>
>>                     </mediaCaptureIDs>
>>                 </sceneEntry>
>>
>> Should VC4 here be AC0?
>>
>>
>> Regards, Christian
>>
>> _______________________________________________
>> clue mailing list
>> clue@ietf.org
>> https://www.ietf.org/mailman/listinfo/clue
>>
>
> _______________________________________________
> clue mailing list
> clue@ietf.org
> https://www.ietf.org/mailman/listinfo/clue
>


From nobody Fri Apr 25 01:25:18 2014
Return-Path: <johaniel@cisco.com>
X-Original-To: clue@ietfa.amsl.com
Delivered-To: clue@ietfa.amsl.com
Received: from localhost (ietfa.amsl.com [127.0.0.1]) by ietfa.amsl.com (Postfix) with ESMTP id 5A9351A00DC for <clue@ietfa.amsl.com>; Fri, 25 Apr 2014 01:25:16 -0700 (PDT)
X-Virus-Scanned: amavisd-new at amsl.com
X-Spam-Flag: NO
X-Spam-Score: -12.873
X-Spam-Level: 
X-Spam-Status: No, score=-12.873 tagged_above=-999 required=5 tests=[BAYES_20=-0.001, DKIM_SIGNED=0.1, DKIM_VALID=-0.1, DKIM_VALID_AU=-0.1, HTML_MESSAGE=0.001, RCVD_IN_DNSWL_HI=-5, RP_MATCHES_RCVD=-0.272, SPF_PASS=-0.001, USER_IN_DEF_DKIM_WL=-7.5] autolearn=ham
Received: from mail.ietf.org ([4.31.198.44]) by localhost (ietfa.amsl.com [127.0.0.1]) (amavisd-new, port 10024) with ESMTP id Rv6qaMJ1Zr43 for <clue@ietfa.amsl.com>; Fri, 25 Apr 2014 01:25:13 -0700 (PDT)
Received: from rcdn-iport-9.cisco.com (rcdn-iport-9.cisco.com [173.37.86.80]) by ietfa.amsl.com (Postfix) with ESMTP id EAAF51A0478 for <clue@ietf.org>; Fri, 25 Apr 2014 01:25:12 -0700 (PDT)
DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/simple; d=cisco.com; i=@cisco.com; l=18113; q=dns/txt; s=iport; t=1398414308; x=1399623908; h=from:to:cc:subject:date:message-id:references: in-reply-to:mime-version; bh=YlXWtdszF/lzXj3C0FWyAQ3GEgeBoRy3q3yNPJgB9ec=; b=mpxgCicxbEARFaqER8I9wcjqGDqIqdVxUeSYBHACZIKlNUNIxEHfeH9J v8BtZsyvOPBLqZTkvgnGqm/nqM44ZDFd7iHpuQ0krBco0wdpr7FmwtU8K r1XQir3CFK0//CPxMtcuBbFdid5MY1EFFCt3GOlQd+QQPBJTIQF29gJ8+ s=;
X-IronPort-Anti-Spam-Filtered: true
X-IronPort-Anti-Spam-Result: AgYFAGkbWlOtJV2b/2dsb2JhbABZDoI0RE9XvGiHOIEPFnSCJQEBAQQBAQFrCxACAQgRAwECKAchBgsUCQgCBAENBYgtAxENw3INhmwTBIxBgUAWMQ0EB4Q5BJcVgXCNCIVUgUCBMUCBaQICHCI
X-IronPort-AV: E=Sophos;i="4.97,925,1389744000";  d="scan'208,217";a="317226667"
Received: from rcdn-core-4.cisco.com ([173.37.93.155]) by rcdn-iport-9.cisco.com with ESMTP; 25 Apr 2014 08:25:07 +0000
Received: from xhc-aln-x11.cisco.com (xhc-aln-x11.cisco.com [173.36.12.85]) by rcdn-core-4.cisco.com (8.14.5/8.14.5) with ESMTP id s3P8P6Wk029209 (version=TLSv1/SSLv3 cipher=AES128-SHA bits=128 verify=FAIL); Fri, 25 Apr 2014 08:25:06 GMT
Received: from xmb-rcd-x09.cisco.com ([169.254.9.100]) by xhc-aln-x11.cisco.com ([173.36.12.85]) with mapi id 14.03.0123.003; Fri, 25 Apr 2014 03:25:05 -0500
From: "Johan Ludvig Nielsen (johaniel)" <johaniel@cisco.com>
To: Stephen Botzko <stephen.botzko@gmail.com>, Paul Kyzivat <pkyzivat@alum.mit.edu>
Thread-Topic: [clue] Improving treatment of audio
Thread-Index: AQHPTc2c8ur938+YIkK8EAurImAl2psL0DeAgBJb+gCAAAdngIAAJvKAgAQ+7oA=
Date: Fri, 25 Apr 2014 08:25:05 +0000
Message-ID: <CF7FE526.138C6%johaniel@cisco.com>
References: <533AF351.9050201@alum.mit.edu> <20140410222705.GW39240@verdi> <20140422144858.GC86778@verdi> <5356878F.4080504@alum.mit.edu> <CAMC7SJ7pu6ChnZuEaby496fQuiuk3jLHfTuLOZiifRpJaBoBXQ@mail.gmail.com>
In-Reply-To: <CAMC7SJ7pu6ChnZuEaby496fQuiuk3jLHfTuLOZiifRpJaBoBXQ@mail.gmail.com>
Accept-Language: en-US
Content-Language: en-US
X-MS-Has-Attach: 
X-MS-TNEF-Correlator: 
user-agent: Microsoft-MacOutlook/14.3.9.131030
x-originating-ip: [10.147.113.125]
Content-Type: multipart/alternative; boundary="_000_CF7FE526138C6johanielciscocom_"
MIME-Version: 1.0
Archived-At: http://mailarchive.ietf.org/arch/msg/clue/hQeX31b6bkCLd-M9w_iXEKcMSac
Cc: CLUE <clue@ietf.org>
Subject: Re: [clue] Improving treatment of audio
X-BeenThere: clue@ietf.org
X-Mailman-Version: 2.1.15
Precedence: list
List-Id: CLUE - ControLling mUltiple streams for TElepresence <clue.ietf.org>
List-Unsubscribe: <https://www.ietf.org/mailman/options/clue>, <mailto:clue-request@ietf.org?subject=unsubscribe>
List-Archive: <http://www.ietf.org/mail-archive/web/clue/>
List-Post: <mailto:clue@ietf.org>
List-Help: <mailto:clue-request@ietf.org?subject=help>
List-Subscribe: <https://www.ietf.org/mailman/listinfo/clue>, <mailto:clue-request@ietf.org?subject=subscribe>
X-List-Received-Date: Fri, 25 Apr 2014 08:25:16 -0000

--_000_CF7FE526138C6johanielciscocom_
Content-Type: text/plain; charset="us-ascii"
Content-Transfer-Encoding: quoted-printable

On the echo issue:

I fully agree with those that argue that acoustic echo cancellation is out =
of scope for CLUE. It is a local responsibility and should be much higher t=
han CLUE on the feature list of any loudspeaking endpoint. While a warning =
label can be put in some CLUE document, it is really not necessary.

By echo I mean the audio signal received from the network, played out on yo=
ur local loudspeaker(s), reverberated through the room back to the micropho=
ne(s) where it is superimposed on the the near-end talker signal. This echo=
 should be cancelled or suppressed before the microphone signal is encoded =
and sent.

Regards
Johan

From: Stephen Botzko <stephen.botzko@gmail.com<mailto:stephen.botzko@gmail.=
com>>
Date: Tuesday 22 April 2014 19:34
To: Paul Kyzivat <pkyzivat@alum.mit.edu<mailto:pkyzivat@alum.mit.edu>>
Cc: "clue@ietf.org<mailto:clue@ietf.org>" <clue@ietf.org<mailto:clue@ietf.o=
rg>>
Subject: Re: [clue] Improving treatment of audio

As I tried to say before, acoustic echo cancellation has been provided in s=
peakerphones and videoconferencing systems for decades. I am still not seei=
ng anything here that CLUE needs to deal with, and nothing unique to telepr=
esence.

But I absolutely agree with the key phrase "we put this issue to bed".

>>>We task each sender to analyze and cancel the acoustic echo of any sound=
 it receives and sends to loudspeakers in the room.

Paul, do you have a specific document in mind where you think this would fi=
t?  And John - would this really end it, or will there be other warning lab=
els that you will want to add somewhere?

BR
Stephen


On Tue, Apr 22, 2014 at 11:15 AM, Paul Kyzivat <pkyzivat@alum.mit.edu<mailt=
o:pkyzivat@alum.mit.edu>> wrote:
John,

IIUC, out of this there is one thing that you are suggesting go into our do=
cuments:


We task each sender to analyze and cancel the acoustic echo of any
sound it receives and sends to loudspeakers in the room.

And then we put this issue to bed. Is that right?

        Thanks,
        Paul


On 4/22/14 10:48 AM, John Leslie wrote:
John Leslie <john@jlc.net<mailto:john@jlc.net>> wrote:

Audio 101

    continuing...

Audio 102

Q: What is Echo?
A: Echo is a reflection of sound arriving with sufficient delay to be perce=
ived separately.

Q: How is that different from reverberation?
A: It's a question of perception: when the delay is much less than 100 mill=
iseconds, the ear doesn't perceive it as separate.

Q: But isn't reverberation sometimes measured in seconds (thousands of mill=
iseconds)?
A: Indeed it is. Reverberation often exceeds five seconds in a large cathed=
ral.

Q: So why isn't that "echo"?
A: No individual reflection stands out enough to be perceived as separate. =
"Reverberation" comes from multiple reflections off different surfaces, whi=
ch can stretch out for seconds.

Q: Why is echo considered a problem when reverberation is not?
A: But reverberation _is_ a problem when it exceeds a few hundred milliseco=
nds.

Q: Huh?
A: Excessive reverberation "muddies" the speech, and is perceived by the sp=
eaker, who speaks more slowly.

Q: Why doesn't that work for echo?
A: The ear perceives separate sound "sources" and the brain is confused by =
them.

Q: What did people do about this before electronics were invented?
A: They met somewhere else. ;^)

Q: When did we decide to "fix" echo?
A: We felt a need to fix it when we couldn't "meet" somewhere else: specifi=
cally when we could only "meet" by telephone.

Q: How's that?
A: As telephones became practical over long distances, they became the only=
 way two people could "meet".

Q: But don't electric signals travel at the speed of light?
A: Actually they never go that fast; but even if they did, 900 miles would =
produce audible echo. By 1960, satellites were used, giving delays over 100=
 milliseconds.

Q: What did telephone companies do about it?
A: They made devices to detect echo and "suppress" it by reducing the signa=
l containing echo.

Q: Did that actually work?
A: Well enough for government work ;^) -- folks learned not to talk while t=
he other person talked.

Q: But we've improved on that, haven't we?
A: Indeed: they figured how to do "echo cancellation" instead of just "echo=
 suppression".

Q: How does that work?
A: They predict the echo signal and "cancel" it by inserting the exact oppo=
site.

Q: Isn't that risky -- don't you end up amplifying if you mis-guess the pha=
se?
A: Yes.

Q: So why did it work at all?
A: Telephony was limited to about three kilohertz -- you'd have to be off b=
y 100 microseconds to get in trouble. Besides, their problem was much simpl=
er: no room reverberation, just electrical reflection

Q: Have we improved on that since 1960?
A: We've improved the estimates, yes; but the acoustic echo is far more com=
plex, and we're fighting to track the changes in reflection.

Q: How do we "track" these changes?
A: We run continuous algorithms to update the estimates.

Q: Isn't that expensive?
A: We dedicate on-chip Digital Signal Processors; so the cost is manageable=
.

Q: What are the practical limits of constantly-updated estimates?
A: Raw processing speed, deviations from the model, and out-of-model events=
 all can lead to bad results. Modern devices generally add non-linear proce=
ssing, beyond the scope of this tutorial.

Q: Isn't processing speed always increasing due to Moore's law?
A: Yes, but 1) it's not increasing as fast as memory capacity, and 2) equip=
ment upgrades are pretty slow.

Q: What are some typical deviations from the model?
A: Many models assume a clear distinction between originated signal and ech=
os, and guess badly when speech overlaps in direction.

Q: What are some out-of-model events?
A: Many models assume all "interesting" echo will come within N millisecond=
s. N tends to be much less than 500. :^(

Q: Is there anything we can do about these problems?
A: We task each sender to analyze and cancel the acoustic echo of any sound=
 it receives and sends to loudspeakers in the room.
This keeps the distinction clear between sound originated in the room vs. e=
cho of what it receives (and the number of milliseconds you need to process=
 is kept reasonably small).

Q: Does it make a difference which particular algorithm is used?
A: Yes, but mostly in overall speed and how large N can be. There are also =
differences in how fast they react to changes.

Q: Is there any reason for us to recommend particular algorithm(s)?
A: No. They tend to be proprietary, and new ones are introduced every year.

--
John Leslie <john@jlc.net<mailto:john@jlc.net>>

_______________________________________________
clue mailing list
clue@ietf.org<mailto:clue@ietf.org>
https://www.ietf.org/mailman/listinfo/clue


_______________________________________________
clue mailing list
clue@ietf.org<mailto:clue@ietf.org>
https://www.ietf.org/mailman/listinfo/clue


--_000_CF7FE526138C6johanielciscocom_
Content-Type: text/html; charset="us-ascii"
Content-ID: <DD59DB1207FAF04EAB0C1951950BF744@emea.cisco.com>
Content-Transfer-Encoding: quoted-printable

<html>
<head>
<meta http-equiv=3D"Content-Type" content=3D"text/html; charset=3Dus-ascii"=
>
</head>
<body style=3D"word-wrap: break-word; -webkit-nbsp-mode: space; -webkit-lin=
e-break: after-white-space; color: rgb(0, 0, 0); font-size: 14px;">
<div>
<div style=3D"font-size: medium;">On the echo issue:</div>
<div style=3D"font-size: medium;"><br>
</div>
<div style=3D"font-size: medium;">I fully agree with those that argue that =
acoustic echo cancellation is out of scope for CLUE. It is a local responsi=
bility and should be much higher than CLUE on the feature list of any louds=
peaking endpoint. While a warning
 label can be put in some CLUE document, it is really not necessary.</div>
</div>
<div style=3D"font-size: medium;"><br>
</div>
<div style=3D"font-size: medium;">By echo I mean the audio signal received =
from the network, played out on your local loudspeaker(s), reverberated thr=
ough the room back to the microphone(s) where it is superimposed on the the=
 near-end talker signal. This echo
 should be cancelled or suppressed before the microphone signal is encoded =
and sent.</div>
<div style=3D"font-size: medium;"><br>
</div>
<div style=3D"font-size: medium;">Regards</div>
<div style=3D"font-size: medium;">Johan</div>
<div style=3D"font-family: Calibri, sans-serif;"><br>
</div>
<span id=3D"OLK_SRC_BODY_SECTION" style=3D"font-family: Calibri, sans-serif=
;">
<div style=3D"font-family:Calibri; font-size:11pt; text-align:left; color:b=
lack; BORDER-BOTTOM: medium none; BORDER-LEFT: medium none; PADDING-BOTTOM:=
 0in; PADDING-LEFT: 0in; PADDING-RIGHT: 0in; BORDER-TOP: #b5c4df 1pt solid;=
 BORDER-RIGHT: medium none; PADDING-TOP: 3pt">
<span style=3D"font-weight:bold">From: </span>Stephen Botzko &lt;<a href=3D=
"mailto:stephen.botzko@gmail.com">stephen.botzko@gmail.com</a>&gt;<br>
<span style=3D"font-weight:bold">Date: </span>Tuesday 22 April 2014 19:34<b=
r>
<span style=3D"font-weight:bold">To: </span>Paul Kyzivat &lt;<a href=3D"mai=
lto:pkyzivat@alum.mit.edu">pkyzivat@alum.mit.edu</a>&gt;<br>
<span style=3D"font-weight:bold">Cc: </span>&quot;<a href=3D"mailto:clue@ie=
tf.org">clue@ietf.org</a>&quot; &lt;<a href=3D"mailto:clue@ietf.org">clue@i=
etf.org</a>&gt;<br>
<span style=3D"font-weight:bold">Subject: </span>Re: [clue] Improving treat=
ment of audio<br>
</div>
<div><br>
</div>
<div>
<div>
<div dir=3D"ltr">
<div>As I tried to say before, acoustic echo cancellation has been provided=
 in speakerphones and videoconferencing systems for decades. I am still not=
 seeing anything here that CLUE needs to deal with, and nothing unique to t=
elepresence.</div>
<div><br>
</div>
<div>But I absolutely agree with the key phrase &quot;we put this issue to =
bed&quot;.<br>
</div>
<div><br>
</div>
<div>&gt;&gt;&gt;<span style=3D"color: rgb(80, 0, 80); font-family: arial, =
sans-serif; font-size: 13px;">We task each sender to analyze and cancel the=
 acoustic echo of any&nbsp;</span><span style=3D"color: rgb(80, 0, 80); fon=
t-family: arial, sans-serif; font-size: 13px;">sound
 it receives and sends to loudspeakers in the room.</span></div>
<div><br>
</div>
<div>Paul, do you have a specific document in mind where you think this wou=
ld fit? &nbsp;And John - would this really end it, or will there be other w=
arning labels that you will want to add somewhere? &nbsp;</div>
<div><br>
</div>
<div>BR</div>
<div>Stephen</div>
</div>
<div class=3D"gmail_extra"><br>
<br>
<div class=3D"gmail_quote">On Tue, Apr 22, 2014 at 11:15 AM, Paul Kyzivat <=
span dir=3D"ltr">
&lt;<a href=3D"mailto:pkyzivat@alum.mit.edu" target=3D"_blank">pkyzivat@alu=
m.mit.edu</a>&gt;</span> wrote:<br>
<blockquote class=3D"gmail_quote" style=3D"margin:0 0 0 .8ex;border-left:1p=
x #ccc solid;padding-left:1ex">
John,<br>
<br>
IIUC, out of this there is one thing that you are suggesting go into our do=
cuments:
<div class=3D""><br>
<br>
<blockquote class=3D"gmail_quote" style=3D"margin:0 0 0 .8ex;border-left:1p=
x #ccc solid;padding-left:1ex">
We task each sender to analyze and cancel the acoustic echo of any<br>
sound it receives and sends to loudspeakers in the room.<br>
</blockquote>
<br>
</div>
And then we put this issue to bed. Is that right?<br>
<br>
&nbsp; &nbsp; &nbsp; &nbsp; Thanks,<br>
&nbsp; &nbsp; &nbsp; &nbsp; Paul
<div class=3D"HOEnZb">
<div class=3D"h5"><br>
<br>
On 4/22/14 10:48 AM, John Leslie wrote:<br>
<blockquote class=3D"gmail_quote" style=3D"margin:0 0 0 .8ex;border-left:1p=
x #ccc solid;padding-left:1ex">
John Leslie &lt;<a href=3D"mailto:john@jlc.net" target=3D"_blank">john@jlc.=
net</a>&gt; wrote:<br>
<blockquote class=3D"gmail_quote" style=3D"margin:0 0 0 .8ex;border-left:1p=
x #ccc solid;padding-left:1ex">
<br>
Audio 101<br>
</blockquote>
<br>
&nbsp; &nbsp; continuing...<br>
<br>
Audio 102<br>
<br>
Q: What is Echo?<br>
A: Echo is a reflection of sound arriving with sufficient delay to be perce=
ived separately.<br>
<br>
Q: How is that different from reverberation?<br>
A: It's a question of perception: when the delay is much less than 100 mill=
iseconds, the ear doesn't perceive it as separate.<br>
<br>
Q: But isn't reverberation sometimes measured in seconds (thousands of mill=
iseconds)?<br>
A: Indeed it is. Reverberation often exceeds five seconds in a large cathed=
ral.<br>
<br>
Q: So why isn't that &quot;echo&quot;?<br>
A: No individual reflection stands out enough to be perceived as separate. =
&quot;Reverberation&quot; comes from multiple reflections off different sur=
faces, which can stretch out for seconds.<br>
<br>
Q: Why is echo considered a problem when reverberation is not?<br>
A: But reverberation _is_ a problem when it exceeds a few hundred milliseco=
nds.<br>
<br>
Q: Huh?<br>
A: Excessive reverberation &quot;muddies&quot; the speech, and is perceived=
 by the speaker, who speaks more slowly.<br>
<br>
Q: Why doesn't that work for echo?<br>
A: The ear perceives separate sound &quot;sources&quot; and the brain is co=
nfused by them.<br>
<br>
Q: What did people do about this before electronics were invented?<br>
A: They met somewhere else. ;^)<br>
<br>
Q: When did we decide to &quot;fix&quot; echo?<br>
A: We felt a need to fix it when we couldn't &quot;meet&quot; somewhere els=
e: specifically when we could only &quot;meet&quot; by telephone.<br>
<br>
Q: How's that?<br>
A: As telephones became practical over long distances, they became the only=
 way two people could &quot;meet&quot;.<br>
<br>
Q: But don't electric signals travel at the speed of light?<br>
A: Actually they never go that fast; but even if they did, 900 miles would =
produce audible echo. By 1960, satellites were used, giving delays over 100=
 milliseconds.<br>
<br>
Q: What did telephone companies do about it?<br>
A: They made devices to detect echo and &quot;suppress&quot; it by reducing=
 the signal containing echo.<br>
<br>
Q: Did that actually work?<br>
A: Well enough for government work ;^) -- folks learned not to talk while t=
he other person talked.<br>
<br>
Q: But we've improved on that, haven't we?<br>
A: Indeed: they figured how to do &quot;echo cancellation&quot; instead of =
just &quot;echo suppression&quot;.<br>
<br>
Q: How does that work?<br>
A: They predict the echo signal and &quot;cancel&quot; it by inserting the =
exact opposite.<br>
<br>
Q: Isn't that risky -- don't you end up amplifying if you mis-guess the pha=
se?<br>
A: Yes.<br>
<br>
Q: So why did it work at all?<br>
A: Telephony was limited to about three kilohertz -- you'd have to be off b=
y 100 microseconds to get in trouble. Besides, their problem was much simpl=
er: no room reverberation, just electrical reflection<br>
<br>
Q: Have we improved on that since 1960?<br>
A: We've improved the estimates, yes; but the acoustic echo is far more com=
plex, and we're fighting to track the changes in reflection.<br>
<br>
Q: How do we &quot;track&quot; these changes?<br>
A: We run continuous algorithms to update the estimates.<br>
<br>
Q: Isn't that expensive?<br>
A: We dedicate on-chip Digital Signal Processors; so the cost is manageable=
.<br>
<br>
Q: What are the practical limits of constantly-updated estimates?<br>
A: Raw processing speed, deviations from the model, and out-of-model events=
 all can lead to bad results. Modern devices generally add non-linear proce=
ssing, beyond the scope of this tutorial.<br>
<br>
Q: Isn't processing speed always increasing due to Moore's law?<br>
A: Yes, but 1) it's not increasing as fast as memory capacity, and 2) equip=
ment upgrades are pretty slow.<br>
<br>
Q: What are some typical deviations from the model?<br>
A: Many models assume a clear distinction between originated signal and ech=
os, and guess badly when speech overlaps in direction.<br>
<br>
Q: What are some out-of-model events?<br>
A: Many models assume all &quot;interesting&quot; echo will come within N m=
illiseconds. N tends to be much less than 500. :^(<br>
<br>
Q: Is there anything we can do about these problems?<br>
A: We task each sender to analyze and cancel the acoustic echo of any sound=
 it receives and sends to loudspeakers in the room.<br>
This keeps the distinction clear between sound originated in the room vs. e=
cho of what it receives (and the number of milliseconds you need to process=
 is kept reasonably small).<br>
<br>
Q: Does it make a difference which particular algorithm is used?<br>
A: Yes, but mostly in overall speed and how large N can be. There are also =
differences in how fast they react to changes.<br>
<br>
Q: Is there any reason for us to recommend particular algorithm(s)?<br>
A: No. They tend to be proprietary, and new ones are introduced every year.=
<br>
<br>
--<br>
John Leslie &lt;<a href=3D"mailto:john@jlc.net" target=3D"_blank">john@jlc.=
net</a>&gt;<br>
<br>
______________________________<u></u>_________________<br>
clue mailing list<br>
<a href=3D"mailto:clue@ietf.org" target=3D"_blank">clue@ietf.org</a><br>
<a href=3D"https://www.ietf.org/mailman/listinfo/clue" target=3D"_blank">ht=
tps://www.ietf.org/mailman/<u></u>listinfo/clue</a><br>
<br>
</blockquote>
<br>
______________________________<u></u>_________________<br>
clue mailing list<br>
<a href=3D"mailto:clue@ietf.org" target=3D"_blank">clue@ietf.org</a><br>
<a href=3D"https://www.ietf.org/mailman/listinfo/clue" target=3D"_blank">ht=
tps://www.ietf.org/mailman/<u></u>listinfo/clue</a><br>
</div>
</div>
</blockquote>
</div>
<br>
</div>
</div>
</div>
</span>
</body>
</html>

--_000_CF7FE526138C6johanielciscocom_--


From nobody Fri Apr 25 03:27:38 2014
Return-Path: <john@jlc.net>
X-Original-To: clue@ietfa.amsl.com
Delivered-To: clue@ietfa.amsl.com
Received: from localhost (ietfa.amsl.com [127.0.0.1]) by ietfa.amsl.com (Postfix) with ESMTP id D62EB1A016F for <clue@ietfa.amsl.com>; Fri, 25 Apr 2014 03:27:36 -0700 (PDT)
X-Virus-Scanned: amavisd-new at amsl.com
X-Spam-Flag: NO
X-Spam-Score: -4.472
X-Spam-Level: 
X-Spam-Status: No, score=-4.472 tagged_above=-999 required=5 tests=[BAYES_00=-1.9, RCVD_IN_DNSWL_MED=-2.3, RP_MATCHES_RCVD=-0.272] autolearn=ham
Received: from mail.ietf.org ([4.31.198.44]) by localhost (ietfa.amsl.com [127.0.0.1]) (amavisd-new, port 10024) with ESMTP id t8P2Le6ryS2x for <clue@ietfa.amsl.com>; Fri, 25 Apr 2014 03:27:35 -0700 (PDT)
Received: from mailhost.jlc.net (mailhost.jlc.net [199.201.159.4]) by ietfa.amsl.com (Postfix) with ESMTP id DB82E1A0353 for <clue@ietf.org>; Fri, 25 Apr 2014 03:27:34 -0700 (PDT)
Received: by mailhost.jlc.net (Postfix, from userid 104) id 42CB3C94BD; Fri, 25 Apr 2014 06:27:27 -0400 (EDT)
Date: Fri, 25 Apr 2014 06:27:27 -0400
From: John Leslie <john@jlc.net>
To: "Johan Ludvig Nielsen (johaniel)" <johaniel@cisco.com>
Message-ID: <20140425102727.GB44329@verdi>
References: <533AF351.9050201@alum.mit.edu> <20140410222705.GW39240@verdi> <20140422144858.GC86778@verdi> <5356878F.4080504@alum.mit.edu> <CAMC7SJ7pu6ChnZuEaby496fQuiuk3jLHfTuLOZiifRpJaBoBXQ@mail.gmail.com> <CF7FE526.138C6%johaniel@cisco.com>
Mime-Version: 1.0
Content-Type: text/plain; charset=us-ascii
Content-Disposition: inline
In-Reply-To: <CF7FE526.138C6%johaniel@cisco.com>
User-Agent: Mutt/1.4.1i
Archived-At: http://mailarchive.ietf.org/arch/msg/clue/xN-dBjEmYP5BwG1JpJkD2m1moPM
Cc: CLUE <clue@ietf.org>
Subject: Re: [clue] Improving treatment of audio
X-BeenThere: clue@ietf.org
X-Mailman-Version: 2.1.15
Precedence: list
List-Id: CLUE - ControLling mUltiple streams for TElepresence <clue.ietf.org>
List-Unsubscribe: <https://www.ietf.org/mailman/options/clue>, <mailto:clue-request@ietf.org?subject=unsubscribe>
List-Archive: <http://www.ietf.org/mail-archive/web/clue/>
List-Post: <mailto:clue@ietf.org>
List-Help: <mailto:clue-request@ietf.org?subject=help>
List-Subscribe: <https://www.ietf.org/mailman/listinfo/clue>, <mailto:clue-request@ietf.org?subject=subscribe>
X-List-Received-Date: Fri, 25 Apr 2014 10:27:37 -0000

Johan Ludvig Nielsen (johaniel) <johaniel@cisco.com> wrote:
> 
> I fully agree with those that argue that acoustic echo cancellation is
> out of scope for CLUE.

   I don't agree it's helpful to call it "out of scope," but...

> It is a local responsibility and should be much higher than CLUE on
> the feature list of any loudspeaking endpoint.

   Yes, it's a local responsibility.

   Note that not all endpoints of interest will have loudspeakers.

> While a warning label can be put in some CLUE document, it is really
> not necessary.

   I much prefer a parameter saying the responsibility is satisfied
to a warning label in some document nobody reads.

> By echo I mean the audio signal received from the network, played out
> on your local loudspeaker(s), reverberated through the room back to
> the microphone(s) where it is superimposed on the the near-end talker
> signal. This echo should be cancelled or suppressed before the
> microphone signal is encoded and sent.

   This is a rather good statement of the problem (better than mine,
certainly).

   How about a parameter saying "cancelled" vs. "suppressed"?

   BTW, audio is on the agenda for the May 13 design team call.

--
John Leslie <john@jlc.net>


From nobody Fri Apr 25 06:49:03 2014
Return-Path: <johaniel@cisco.com>
X-Original-To: clue@ietfa.amsl.com
Delivered-To: clue@ietfa.amsl.com
Received: from localhost (ietfa.amsl.com [127.0.0.1]) by ietfa.amsl.com (Postfix) with ESMTP id C591E1A04C8 for <clue@ietfa.amsl.com>; Fri, 25 Apr 2014 06:48:48 -0700 (PDT)
X-Virus-Scanned: amavisd-new at amsl.com
X-Spam-Flag: NO
X-Spam-Score: -13.373
X-Spam-Level: 
X-Spam-Status: No, score=-13.373 tagged_above=-999 required=5 tests=[BAYES_05=-0.5, DKIM_SIGNED=0.1, DKIM_VALID=-0.1, DKIM_VALID_AU=-0.1, RCVD_IN_DNSWL_HI=-5, RP_MATCHES_RCVD=-0.272, SPF_PASS=-0.001, USER_IN_DEF_DKIM_WL=-7.5] autolearn=ham
Received: from mail.ietf.org ([4.31.198.44]) by localhost (ietfa.amsl.com [127.0.0.1]) (amavisd-new, port 10024) with ESMTP id VmlXYDbKNTCU for <clue@ietfa.amsl.com>; Fri, 25 Apr 2014 06:48:41 -0700 (PDT)
Received: from rcdn-iport-4.cisco.com (rcdn-iport-4.cisco.com [173.37.86.75]) by ietfa.amsl.com (Postfix) with ESMTP id 528C71A04D2 for <clue@ietf.org>; Fri, 25 Apr 2014 06:48:37 -0700 (PDT)
DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/simple; d=cisco.com; i=@cisco.com; l=13577; q=dns/txt; s=iport; t=1398433711; x=1399643311; h=from:to:subject:date:message-id:references:in-reply-to: content-id:content-transfer-encoding:mime-version; bh=r0DXwhRd54F1hNNxpYEwgfn+3tFmFIswaea6PmHNW4w=; b=ME78DEljk8Eg8pQjfdEUWKrKrmaEuscU8Gqcq9+YOnx3BBqg2w4vZIBL Yt/xJs9WH71BXQLg1rL3BdsV+tqioXTU3ay6TMLqEtjDb4aq+Xmu8lelb EnZlYoWLJon8ptGdUoICle54C9sGq8p7AEL+IMDKnfviJ0DbtvxGEherM w=;
X-IronPort-Anti-Spam-Filtered: true
X-IronPort-Anti-Spam-Result: AgUFAOpmWlOtJA2G/2dsb2JhbABZgwZPV7xnhziBDxZ0giUBAQEEAQEBaxcEAgEIEQQBAQEnBycLFAkIAgQBEhYFBIgiDcpkEwSJN4RKLSYMAgSEMwSJM49SklyDMYFpQg
X-IronPort-AV: E=Sophos;i="4.97,927,1389744000"; d="scan'208";a="320403547"
Received: from alln-core-12.cisco.com ([173.36.13.134]) by rcdn-iport-4.cisco.com with ESMTP; 25 Apr 2014 13:48:30 +0000
Received: from xhc-rcd-x04.cisco.com (xhc-rcd-x04.cisco.com [173.37.183.78]) by alln-core-12.cisco.com (8.14.5/8.14.5) with ESMTP id s3PDmUU7019342 (version=TLSv1/SSLv3 cipher=AES128-SHA bits=128 verify=FAIL); Fri, 25 Apr 2014 13:48:30 GMT
Received: from xmb-rcd-x09.cisco.com ([169.254.9.100]) by xhc-rcd-x04.cisco.com ([fe80::200:5efe:173.37.183.34%12]) with mapi id 14.03.0123.003; Fri, 25 Apr 2014 08:48:29 -0500
From: "Johan Ludvig Nielsen (johaniel)" <johaniel@cisco.com>
To: Christian Groves <Christian.Groves@nteczone.com>, "Duckworth, Mark" <Mark.Duckworth@polycom.com>, "clue@ietf.org" <clue@ietf.org>
Thread-Topic: [clue] Improving treatment of audio
Thread-Index: AQHPTc2c8ur938+YIkK8EAurImAl2psL0DeAgAAOHwCAADfhAIAAz6yAgABJe4CAABefgIAADvsAgANu+oCAAAUZAIAADnaAgBIbUQA=
Date: Fri, 25 Apr 2014 13:48:28 +0000
Message-ID: <CF803317.13BC2%johaniel@cisco.com>
References: <533AF351.9050201@alum.mit.edu> <20140410222705.GW39240@verdi> <BLU0-SMTP6910DA20C9A9287137B585D0550@phx.gbl> <53475572.4040807@nteczone.com> <20140411150055.GE60844@verdi> <49E45C59CA48264997FEBFB29B6BC2D6215295A8AE@CRPMBOXPRD07.polycom.com> <BLU0-SMTP1804A70E3E0F666EAEAE84D0540@phx.gbl> <49E45C59CA48264997FEBFB29B6BC2D62409219636@CRPMBOXPRD07.polycom.com> <534B4303.7060707@nteczone.com> <5C4AC54BFF7A0842A6A11F554D6FB52F07152B@CRPMBOXPRD08.polycom.com> <534B536B.8000205@nteczone.com>
In-Reply-To: <534B536B.8000205@nteczone.com>
Accept-Language: en-US
Content-Language: en-US
X-MS-Has-Attach: 
X-MS-TNEF-Correlator: 
user-agent: Microsoft-MacOutlook/14.3.9.131030
x-originating-ip: [10.147.113.125]
Content-Type: text/plain; charset="iso-8859-1"
Content-ID: <3AF93B74B5EDEA49BFEBBDCFB57A6EEB@emea.cisco.com>
Content-Transfer-Encoding: quoted-printable
MIME-Version: 1.0
Archived-At: http://mailarchive.ietf.org/arch/msg/clue/KAd8EDA7FMu8u_H09673c89P2tk
Subject: Re: [clue] Improving treatment of audio
X-BeenThere: clue@ietf.org
X-Mailman-Version: 2.1.15
Precedence: list
List-Id: CLUE - ControLling mUltiple streams for TElepresence <clue.ietf.org>
List-Unsubscribe: <https://www.ietf.org/mailman/options/clue>, <mailto:clue-request@ietf.org?subject=unsubscribe>
List-Archive: <http://www.ietf.org/mail-archive/web/clue/>
List-Post: <mailto:clue@ietf.org>
List-Help: <mailto:clue-request@ietf.org?subject=help>
List-Subscribe: <https://www.ietf.org/mailman/listinfo/clue>, <mailto:clue-request@ietf.org?subject=subscribe>
X-List-Received-Date: Fri, 25 Apr 2014 13:48:49 -0000

I have not been following CLUE lately, please forgive wrong terminology.

We had a very similar discussion on this list more than two years ago
about spatial association of audio to video. At that time I suggested
relating ACs to VCs or a group of VCs, both to assist consumer selection
of streams and to assist mixing/routing for playback on the rendering
side. This was very similar to what Christian and Paul C. has mentioned.

My understanding of the current framework is that receivers calculate the
spatial association between an audio capture and the video scene from the
areas of capture. This will probably work. But I still think having an
explicit association of captures would be a much simpler, more robust and
more scalable solution.

(Many (even advanced multi-capture) systems won=B4t have a clue about the
capture area and axes of their cameras and microphones, but an association
of audio to video captures can still make perfect sense. For instance in
the education scenario.)

>From the point of view of the renderer, what is needed is info that helps
decide where to play out the audio streams it receives. Whatever the
shortcomings or unknowns of the microphone system in the sending room, and
the problems of linking physical audio and video capture areas, the
playout direction is linked to the video scene to be rendered. Normally
the aim would be to render audio spatially coherent with the video scene,
but it can also be (as John has commented well on) to be able to separate
sound sources for increased intelligibility and scene perception, or to
keep the audio scene static by choice.

An audio capture, be it a single mic signal or a mix, represents a sound
source or a collection of sound sources, which are more often than not
visible in a given video capture or scene and should therefore be
connected to that capture or scene of captures. Naturally, several audio
captures can be connected to the same video capture or scene.

Coordinates can be of interest to convey a finer granularity of the
position of a sound source in the video capture, if available. Information
about the sending room and its microphones (like point or area of capture,
sensitivity, directivity, etc.) are fundamental characteristics of the
capture system, but they are of no use in rendering direction. And
therefore do not need to be advertised or signalled.

More fancy spatial audio schemes exist and may see applications in
telepresence in the (somewhat distant) future, but will probably come in
the form of single stream multichannel formats where the spatial
information is embedded in the format like in surround sound, or
explicitly in the encoded stream like in spatial audio object coding.

Best regards
Johan



On 14/04/14 05:18, "Christian Groves" <Christian.Groves@nteczone.com>
wrote:

>Hello Mark,
>
>Sorry I didn't consider the entire 12.1.1. if I do that, according to
>that example:
>
>Video areas of capture:
>
>        bottom left    bottom right  top left         top right
>    VC0 (-2011,2850,0) (-673,3000,0) (-2011,2850,757) (-673,3000,757)
>    VC1 ( -673,3000,0) ( 673,3000,0) ( -673,3000,757) ( 673,3000,757)
>    VC2 (  673,3000,0) (2011,2850,0) (  673,3000,757) (2011,3000,757)
>    VC3 (-2011,2850,0) (2011,2850,0) (-2011,2850,757) (2011,3000,757)
>    VC4 (-2011,2850,0) (2011,2850,0) (-2011,2850,757) (2011,3000,757)
>    VC5 (-2011,2850,0) (2011,2850,0) (-2011,2850,757) (2011,3000,757)
>    VC6 none
>
>Areas of capture for audio (from 12.1.1):
>
>        bottom left    bottom right  top left         top right
>
>    AC0 (-2011,2850,0) (-673,3000,0) (-2011,2850,757) (-673,3000,757)
>    AC1 (  673,3000,0) (2011,2850,0) (  673,3000,757) (2011,3000,757)
>    AC2 ( -673,3000,0) ( 673,3000,0) ( -673,3000,757) ( 673,3000,757)
>    AC3 (-2011,2850,0) (2011,2850,0) (-2011,2850,757) (2011,3000,757)
>    AC4 none
>
>Using a reference rather than area of capture:
>     AC0 (VC0)
>     AC1 (VC2)
>     AC2 (VC1)
>     AC3 (VC3,VC4,VC5)
>I would think they are basically conveying the same information???
>
>Regards, Christian
>
>On 14/04/2014 12:26 PM, Duckworth, Mark wrote:
>> Hello Christian,
>> please see below.
>> Mark
>>
>>> -----Original Message-----
>>> From: clue [mailto:clue-bounces@ietf.org] On Behalf Of Christian Groves
>>> Sent: Sunday, April 13, 2014 10:08 PM
>>> To: clue@ietf.org
>>> Subject: Re: [clue] Improving treatment of audio
>>>
>>> Hello all,
>>>
>>> Perhaps the confusion is that some see Audio Capture area as a means to
>>> associate an Audio capture with a video capture as you've shown Mark's
>>> example. e.g. The audio spatial information isn't really used for any
>>>audio
>>> transformation other than associating it with a particular video
>>>stream.
>>>
>>> Whereas others were more thinking the audio spatial information as an
>>>input
>>> to mixing and more complicated audio processing.
>> [Duckworth, Mark] You could be right about this being a source of
>>confusion.
>>
>>> If this is the case perhaps rather than linking ACs and VCs through the
>>> physical or virtual co-ordinates of the area of capture information, we
>>> simplify things and in each audio capture we say what VC it relates to?
>>> e.g.using Mark's example AC0(VC0), AC3(V0,V1,V2)
>> [Duckworth, Mark] I'm not sure how this would really work.  Because
>>even in this simple example, we also have AC0 relates to VC3
>>(sometimes), VC4, and VC5 (at least part of it), and so on.  AC3 relates
>>to all of VC1 through VC5.  I think using area of capture works better
>>than trying to do something like this.
>>
>> Regards,
>> Mark
>>
>>> Regards, Christian
>>>
>>> On 12/04/2014 7:42 AM, Duckworth, Mark wrote:
>>>> Hi Paul,
>>>> Yes, I understand the area of capture for audio can't be nearly as
>>>>precise as
>>> it can be for video.  But for this usage I think it doesn't matter.
>>>> Mark
>>>>
>>>>> -----Original Message-----
>>>>> From: Paul Coverdale [mailto:coverdale@sympatico.ca]
>>>>> Sent: Friday, April 11, 2014 4:48 PM
>>>>> To: Duckworth, Mark; 'John Leslie'; 'Christian Groves'
>>>>> Cc: clue@ietf.org
>>>>> Subject: RE: [clue] Improving treatment of audio
>>>>>
>>>>> Hi Mark,
>>>>>
>>>>> I can see what you're trying to do in the example given in Framework
>>>>> section 12.1.1. The problem I have is that, because of the huge
>>>>> difference in wavelength between light waves and audio waves, you
>>>>> can't really define an area of capture for audio in the same way that
>>>>> you can for video. A camera can focus on a scene defined by 4
>>>>> co-planar X,Y,Z coordinates. It will capture the video inside that
>>>>> quadrilateral, and nothing outside. A microphone can capture audio
>>>>> coming from inside of the same quadrilateral, but it will also
>>>>> capture a lot of audio from outside. What this means is that we can
>>>>> never define a unique association of an audio capture with a video
>>>>> capture based on pure geometrical considerations, but we can still
>>>>>simply
>>> define that ACx is associated with VCx, but maybe with VCy and VCz too.
>>>>> ...Paul
>>>>>
>>>>>> -----Original Message-----
>>>>>> From: clue [mailto:clue-bounces@ietf.org] On Behalf Of Duckworth,
>>>>>> Mark
>>>>>> Sent: Friday, April 11, 2014 3:24 PM
>>>>>> To: John Leslie; Christian Groves
>>>>>> Cc: clue@ietf.org
>>>>>> Subject: Re: [clue] Improving treatment of audio
>>>>>>
>>>>>> Hi all,
>>>>>> The original intent of using "area of capture" for all media types,
>>>>>> including audio and video, was to provide a way to associate audio
>>>>>> with video.  So the consumer can choose captures that go together
>>>>>> and render them together.  An example of this is in the framework
>>>>>> document, section 12.1.1.  I still think this works fine for this
>>>>>>purpose.
>>>>>>
>>>>>> In that example, the consumer can choose VC0, VC1, VC2, AC0, AC1,
>>>>>> and AC2.  AC0 and VC0 have the same area of capture, indicating they
>>>>>> are capturing the same area of the scene.  If the consumer wants to
>>>>>> render
>>>>>> AC0 from a loudspeaker close to the display for VC0 it can do so.
>>>>>> Similarly, AC3 has an area of capture covering the whole scene (the
>>>>>> full extent of the areas of VC0, VC1, VC2) so the consumer knows AC3
>>>>>> includes audio associated with all of VC0, VC1, and VC2.
>>>>>>
>>>>>> So I'm puzzled why you are proposing we remove the ability to use
>>>>>> area of capture for audio for this purpose.
>>>>>>
>>>>>> Regards,
>>>>>> Mark
>>>>>>
>>>>>>> -----Original Message-----
>>>>>>> From: clue [mailto:clue-bounces@ietf.org] On Behalf Of John Leslie
>>>>>>> Sent: Friday, April 11, 2014 11:01 AM
>>>>>>> To: Christian Groves
>>>>>>> Cc: clue@ietf.org
>>>>>>> Subject: Re: [clue] Improving treatment of audio
>>>>>>>
>>>>>>> Christian Groves <Christian.Groves@nteczone.com> wrote:
>>>>>>>> "So it's not clear to me what extra we need to do from a CLUE
>>>>>>>> perspective. What is the problem?"
>>>>>>>> I guess this is the pertinent point.
>>>>>>>>
>>>>>>>>   From John's Audio 101 it seems that anything to do with spatial
>>>>>>>> information that would relate to reverberation is in the too hard
>>>>>>>> basket.
>>>>>>>      What precisely does "too hard basket" mean?
>>>>>>>
>>>>>>>      I haven't even reached the point of suggesting what metrics to
>>>>>>> define the _ability_ to send. To me, "too hard" merely means that
>>>>>>> an individual site could choose not to send them or to ignore them
>>>>>>> on
>>>>>> receipt.
>>>>>>>      But it sounds as if you're suggesting reverberation is "to
>>>>>>>hard
>>>>>>> to understand" and thus we should have no metrics about it.
>>>>>>>
>>>>>>>      I _hope_ that's not what you mean.
>>>>>>>
>>>>>>>> So it seems "area of capture" for audio could be marked "not
>>>>>>>> applicable" in the framework.
>>>>>>>      I hope so.
>>>>>>>
>>>>>>>> There doesn't seem to be any driver for having the "point of
>>>>>> capture"
>>>>>>>> apply to an audio capture either.
>>>>>>>      I don't understand this. Point of capture for a microphone may
>>>>>>> be "too hard" to track (today) for a microphone which moves;
>>>>>>> nonetheless it seems to me the most fundamental metric there could
>>> be.
>>>>>>>>   From the Audio101 there does seem to be a dependency on where
>>> the
>>>>>>>> microphone is located with respect to the person speaking (i.e.
>>>>>>>> lapel mic, desk mic, room mic) to how it is handled at
>>>>>> mixing/playout.
>>>>>>>      I don't think I really got that far...
>>>>>>>
>>>>>>>      IMHO, it's more flexible to have mixing be the responsibily of
>>>>>>> the receiver, but I haven't tried to specify that and I'm not at
>>>>>>> all sure
>>>>>> I want to specify that.
>>>>>>>      I expect the actual sound systems in different rooms to vary
>>>>>>> wildly, from one monaural speaker to stereo to full surround-sound.
>>>>>>> Mixing for these without knowing which is the actual target seems
>>>>>>> hard; but I'm sure there will be sites which prefer to do so. A
>>>>>>> question which will arise, IMHO, is how to specify the _intent_ of
>>>>>>> a mix generated in one room to be fed to other rooms. (I'd prefer
>>>>>>> not to go there yet.)
>>>>>>>
>>>>>>>> Perhaps this is useful to signal via CLUE? If it is possible to
>>>>>>>> signal this then perhaps tying a particular audio capture to a
>>>>>>>> video capture makes sense?
>>>>>>>      I'm not thinking along those lines. (That doesn't mean we
>>>>>>> shouldn't think along those lines...) I'm thinking in terms of
>>>>>>> providing several audio streams per room, associated with position
>>>>>>> information about the position of the source of those sounds, and
>>>>>>> allowing the receiver to choose how to mix them and how to present
>>>>>>> the
>>>>> mix in his/her room.
>>>>>>>> i.e. a talker giving a presentation using a lapel mic captured by
>>>>>>>> a particular video.
>>>>>>>      In fact, there's only limited tendency for humans to strictly
>>>>>>> attach the sound they hear to the video they see. Clearly, during a
>>>>>>> presentation we want to _hear_ the presenter talking, but having
>>>>>>> the sound move back and forth as the speaker walks can be
>>>>>>>confusing.
>>>>>>>
>>>>>>>      At the same time, we will want to hear the questions to which
>>>>>>> the presenter may respond. It will be far easier on the listener if
>>>>>>> these do _not_ seem to be coming from the same physical position,
>>>>>>> especially when the questioner _is_ in the same room as the
>>> presenter.
>>>>>>>      Hope this helps...
>>>>>>>
>>>>>>> --
>>>>>>> John Leslie <john@jlc.net>
>>>>>>>
>>>>>>> _______________________________________________
>>>>>>> clue mailing list
>>>>>>> clue@ietf.org
>>>>>>> https://www.ietf.org/mailman/listinfo/clue
>>>>>> _______________________________________________
>>>>>> clue mailing list
>>>>>> clue@ietf.org
>>>>>> https://www.ietf.org/mailman/listinfo/clue
>>>> _______________________________________________
>>>> clue mailing list
>>>> clue@ietf.org
>>>> https://www.ietf.org/mailman/listinfo/clue
>>>>
>>> _______________________________________________
>>> clue mailing list
>>> clue@ietf.org
>>> https://www.ietf.org/mailman/listinfo/clue
>
>_______________________________________________
>clue mailing list
>clue@ietf.org
>https://www.ietf.org/mailman/listinfo/clue


From nobody Fri Apr 25 08:05:32 2014
Return-Path: <john@jlc.net>
X-Original-To: clue@ietfa.amsl.com
Delivered-To: clue@ietfa.amsl.com
Received: from localhost (ietfa.amsl.com [127.0.0.1]) by ietfa.amsl.com (Postfix) with ESMTP id A528C1A04E7 for <clue@ietfa.amsl.com>; Fri, 25 Apr 2014 08:05:28 -0700 (PDT)
X-Virus-Scanned: amavisd-new at amsl.com
X-Spam-Flag: NO
X-Spam-Score: -4.472
X-Spam-Level: 
X-Spam-Status: No, score=-4.472 tagged_above=-999 required=5 tests=[BAYES_00=-1.9, RCVD_IN_DNSWL_MED=-2.3, RP_MATCHES_RCVD=-0.272] autolearn=ham
Received: from mail.ietf.org ([4.31.198.44]) by localhost (ietfa.amsl.com [127.0.0.1]) (amavisd-new, port 10024) with ESMTP id E9GkXNsKAEcO for <clue@ietfa.amsl.com>; Fri, 25 Apr 2014 08:05:26 -0700 (PDT)
Received: from mailhost.jlc.net (mailhost.jlc.net [199.201.159.4]) by ietfa.amsl.com (Postfix) with ESMTP id 6E2B41A04CD for <clue@ietf.org>; Fri, 25 Apr 2014 08:05:26 -0700 (PDT)
Received: by mailhost.jlc.net (Postfix, from userid 104) id 1F965C94C2; Fri, 25 Apr 2014 11:05:18 -0400 (EDT)
Date: Fri, 25 Apr 2014 11:05:18 -0400
From: John Leslie <john@jlc.net>
To: "Johan Ludvig Nielsen (johaniel)" <johaniel@cisco.com>
Message-ID: <20140425150518.GD44329@verdi>
References: <BLU0-SMTP6910DA20C9A9287137B585D0550@phx.gbl> <53475572.4040807@nteczone.com> <20140411150055.GE60844@verdi> <49E45C59CA48264997FEBFB29B6BC2D6215295A8AE@CRPMBOXPRD07.polycom.com> <BLU0-SMTP1804A70E3E0F666EAEAE84D0540@phx.gbl> <49E45C59CA48264997FEBFB29B6BC2D62409219636@CRPMBOXPRD07.polycom.com> <534B4303.7060707@nteczone.com> <5C4AC54BFF7A0842A6A11F554D6FB52F07152B@CRPMBOXPRD08.polycom.com> <534B536B.8000205@nteczone.com> <CF803317.13BC2%johaniel@cisco.com>
Mime-Version: 1.0
Content-Type: text/plain; charset=us-ascii
Content-Disposition: inline
In-Reply-To: <CF803317.13BC2%johaniel@cisco.com>
User-Agent: Mutt/1.4.1i
Archived-At: http://mailarchive.ietf.org/arch/msg/clue/2FqKJxDUKoolMPIUX1RCsq7YhUo
Cc: "clue@ietf.org" <clue@ietf.org>
Subject: Re: [clue] Improving treatment of audio
X-BeenThere: clue@ietf.org
X-Mailman-Version: 2.1.15
Precedence: list
List-Id: CLUE - ControLling mUltiple streams for TElepresence <clue.ietf.org>
List-Unsubscribe: <https://www.ietf.org/mailman/options/clue>, <mailto:clue-request@ietf.org?subject=unsubscribe>
List-Archive: <http://www.ietf.org/mail-archive/web/clue/>
List-Post: <mailto:clue@ietf.org>
List-Help: <mailto:clue-request@ietf.org?subject=help>
List-Subscribe: <https://www.ietf.org/mailman/listinfo/clue>, <mailto:clue-request@ietf.org?subject=subscribe>
X-List-Received-Date: Fri, 25 Apr 2014 15:05:28 -0000

Johan Ludvig Nielsen (johaniel) <johaniel@cisco.com> wrote:
> 
> My understanding of the current framework is that receivers calculate the
> spatial association between an audio capture and the video scene from the
> areas of capture. This will probably work.

   Poorly, at best, IMHO. I'd like to discourage it.

> But I still think having an explicit association of captures would be
> a much simpler, more robust and more scalable solution.

   I entirely agree. I'd like to propose an optional attribute of
video captures to link audio captures in the area it covers.

> From the point of view of the renderer, what is needed is info that
> helps decide where to play out the audio streams it receives.

   While I'm not enthusiastic about choosing that way, this seems an
entirely reasonable piece of information to supply.

> More fancy spatial audio schemes exist and may see applications in
> telepresence in the (somewhat distant) future, but will probably come in
> the form of single stream multichannel formats where the spatial
> information is embedded in the format like in surround sound, or
> explicitly in the encoded stream like in spatial audio object coding.

   Agreed -- that's not something to settle in our initial spec.

--
John Leslie <john@jlc.net>


From nobody Fri Apr 25 14:28:46 2014
Return-Path: <Mark.Duckworth@polycom.com>
X-Original-To: clue@ietfa.amsl.com
Delivered-To: clue@ietfa.amsl.com
Received: from localhost (ietfa.amsl.com [127.0.0.1]) by ietfa.amsl.com (Postfix) with ESMTP id 19BC41A0670 for <clue@ietfa.amsl.com>; Fri, 25 Apr 2014 14:28:44 -0700 (PDT)
X-Virus-Scanned: amavisd-new at amsl.com
X-Spam-Flag: NO
X-Spam-Score: -1.82
X-Spam-Level: 
X-Spam-Status: No, score=-1.82 tagged_above=-999 required=5 tests=[BAYES_00=-1.9, RCVD_IN_DNSWL_LOW=-0.7, SPF_NEUTRAL=0.779, UNPARSEABLE_RELAY=0.001] autolearn=ham
Received: from mail.ietf.org ([4.31.198.44]) by localhost (ietfa.amsl.com [127.0.0.1]) (amavisd-new, port 10024) with ESMTP id acXU3DTJGsu0 for <clue@ietfa.amsl.com>; Fri, 25 Apr 2014 14:28:38 -0700 (PDT)
Received: from mail1.bemta7.messagelabs.com (mail1.bemta7.messagelabs.com [216.82.254.108]) by ietfa.amsl.com (Postfix) with ESMTP id 925E31A0679 for <clue@ietf.org>; Fri, 25 Apr 2014 14:28:38 -0700 (PDT)
Received: from [216.82.254.20:52004] by server-12.bemta-7.messagelabs.com id A1/32-01335-F73DA535; Fri, 25 Apr 2014 21:28:31 +0000
X-Env-Sender: Mark.Duckworth@polycom.com
X-Msg-Ref: server-15.tower-47.messagelabs.com!1398461310!10047251!1
X-Originating-IP: [140.242.64.158]
X-StarScan-Received: 
X-StarScan-Version: 6.11.1; banners=-,-,-
X-VirusChecked: Checked
Received: (qmail 5874 invoked from network); 25 Apr 2014 21:28:31 -0000
Received: from crpehubprd01.polycom.com (HELO Crpehubprd01.polycom.com) (140.242.64.158) by server-15.tower-47.messagelabs.com with AES128-SHA encrypted SMTP; 25 Apr 2014 21:28:31 -0000
Received: from CRPMBOXPRD08.polycom.com ([169.254.1.94]) by Crpehubprd01.polycom.com ([fe80::5efe:10.236.0.158%14]) with mapi; Fri, 25 Apr 2014 14:28:30 -0700
From: "Duckworth, Mark" <Mark.Duckworth@polycom.com>
To: "Johan Ludvig Nielsen (johaniel)" <johaniel@cisco.com>, Christian Groves <Christian.Groves@nteczone.com>, "clue@ietf.org" <clue@ietf.org>
Date: Fri, 25 Apr 2014 14:28:28 -0700
Thread-Topic: [clue] Improving treatment of audio
Thread-Index: AQHPTc2c8ur938+YIkK8EAurImAl2psL0DeAgAAOHwCAADfhAIAAz6yAgABJe4CAABefgIAADvsAgANu+oCAAAUZAIAADnaAgBIbUQD//95VsA==
Message-ID: <5C4AC54BFF7A0842A6A11F554D6FB52F0B97BF@CRPMBOXPRD08.polycom.com>
References: <533AF351.9050201@alum.mit.edu> <20140410222705.GW39240@verdi> <BLU0-SMTP6910DA20C9A9287137B585D0550@phx.gbl> <53475572.4040807@nteczone.com> <20140411150055.GE60844@verdi> <49E45C59CA48264997FEBFB29B6BC2D6215295A8AE@CRPMBOXPRD07.polycom.com> <BLU0-SMTP1804A70E3E0F666EAEAE84D0540@phx.gbl> <49E45C59CA48264997FEBFB29B6BC2D62409219636@CRPMBOXPRD07.polycom.com> <534B4303.7060707@nteczone.com> <5C4AC54BFF7A0842A6A11F554D6FB52F07152B@CRPMBOXPRD08.polycom.com> <534B536B.8000205@nteczone.com> <CF803317.13BC2%johaniel@cisco.com>
In-Reply-To: <CF803317.13BC2%johaniel@cisco.com>
Accept-Language: en-US
Content-Language: en-US
X-MS-Has-Attach: 
X-MS-TNEF-Correlator: 
acceptlanguage: en-US
Content-Type: text/plain; charset="iso-8859-1"
Content-Transfer-Encoding: quoted-printable
MIME-Version: 1.0
Archived-At: http://mailarchive.ietf.org/arch/msg/clue/mwKULCOLBuSamBk4J4GmdLLKYDA
Subject: Re: [clue] Improving treatment of audio
X-BeenThere: clue@ietf.org
X-Mailman-Version: 2.1.15
Precedence: list
List-Id: CLUE - ControLling mUltiple streams for TElepresence <clue.ietf.org>
List-Unsubscribe: <https://www.ietf.org/mailman/options/clue>, <mailto:clue-request@ietf.org?subject=unsubscribe>
List-Archive: <http://www.ietf.org/mail-archive/web/clue/>
List-Post: <mailto:clue@ietf.org>
List-Help: <mailto:clue-request@ietf.org?subject=help>
List-Subscribe: <https://www.ietf.org/mailman/listinfo/clue>, <mailto:clue-request@ietf.org?subject=subscribe>
X-List-Received-Date: Fri, 25 Apr 2014 21:28:44 -0000

Hello Johan,
Thanks very much for the input.
Comments inline.
Mark

> -----Original Message-----
> From: Johan Ludvig Nielsen (johaniel) [mailto:johaniel@cisco.com]
> Sent: Friday, April 25, 2014 9:48 AM
> To: Christian Groves; Duckworth, Mark; clue@ietf.org
> Subject: Re: [clue] Improving treatment of audio
>
> I have not been following CLUE lately, please forgive wrong terminology.
>
> We had a very similar discussion on this list more than two years ago abo=
ut
> spatial association of audio to video. At that time I suggested relating =
ACs to
> VCs or a group of VCs, both to assist consumer selection of streams and t=
o
> assist mixing/routing for playback on the rendering side. This was very s=
imilar
> to what Christian and Paul C. has mentioned.
>
> My understanding of the current framework is that receivers calculate the
> spatial association between an audio capture and the video scene from the
> areas of capture. This will probably work.

[Duckworth, Mark] Yes, that is the idea

> But I still think having an explicit
> association of captures would be a much simpler, more robust and more
> scalable solution.

[Duckworth, Mark] You might be right.  Christian's proposal is something li=
ke this, but to me it still seems incomplete and would need to continue usi=
ng the spatial coordinates to make sense, which I think was not Christian's=
 intent.  Does anybody want to make further proposals along these lines?

> (Many (even advanced multi-capture) systems won=B4t have a clue about the
> capture area and axes of their cameras and microphones, but an associatio=
n
> of audio to video captures can still make perfect sense. For instance in =
the
> education scenario.)
>
> From the point of view of the renderer, what is needed is info that helps
> decide where to play out the audio streams it receives. Whatever the
> shortcomings or unknowns of the microphone system in the sending room,
> and the problems of linking physical audio and video capture areas, the
> playout direction is linked to the video scene to be rendered. Normally t=
he
> aim would be to render audio spatially coherent with the video scene, but=
 it
> can also be (as John has commented well on) to be able to separate sound
> sources for increased intelligibility and scene perception, or to keep th=
e
> audio scene static by choice.
[Duckworth, Mark] I agree

> An audio capture, be it a single mic signal or a mix, represents a sound =
source
> or a collection of sound sources, which are more often than not visible i=
n a
> given video capture or scene and should therefore be connected to that
> capture or scene of captures. Naturally, several audio captures can be
> connected to the same video capture or scene.

[Duckworth, Mark] I agree.  Today in CLUE we can explicitly associate a set=
 of audio captures with a set of video captures by putting them in the same=
 scene.  The difficulty seems to be associating audio with video when there=
 are more than one of each in the same scene.  The intent was to do that wi=
th the area of capture attribute, but maybe we can do better.

> Coordinates can be of interest to convey a finer granularity of the posit=
ion of
> a sound source in the video capture, if available. Information about the
> sending room and its microphones (like point or area of capture, sensitiv=
ity,
> directivity, etc.) are fundamental characteristics of the capture system,=
 but
> they are of no use in rendering direction.

[Duckworth, Mark] I don't understand, because those last two sentences seem=
 to contradict each other.  Wouldn't the rendering system be able to make u=
se of the more fine grain positioning information if it wanted to?  Can you=
 explain more?

> And therefore do not need to be
> advertised or signalled.
>
> More fancy spatial audio schemes exist and may see applications in
> telepresence in the (somewhat distant) future, but will probably come in =
the
> form of single stream multichannel formats where the spatial information =
is
> embedded in the format like in surround sound, or explicitly in the encod=
ed
> stream like in spatial audio object coding.

[Duckworth, Mark] I am focusing on the single stream (single capture) multi=
channel formats.  Secondarily, I am considering the multiple audio capture =
scenarios.  Both are mentioned in the use cases: "The audio may be transmit=
ted as multi-channel (stereo/surround sound) or as distinct and separate mo=
nophonic streams"

> Best regards
> Johan
>
>
>
> On 14/04/14 05:18, "Christian Groves" <Christian.Groves@nteczone.com>
> wrote:
>
> >Hello Mark,
> >
> >Sorry I didn't consider the entire 12.1.1. if I do that, according to
> >that example:
> >
> >Video areas of capture:
> >
> >        bottom left    bottom right  top left         top right
> >    VC0 (-2011,2850,0) (-673,3000,0) (-2011,2850,757) (-673,3000,757)
> >    VC1 ( -673,3000,0) ( 673,3000,0) ( -673,3000,757) ( 673,3000,757)
> >    VC2 (  673,3000,0) (2011,2850,0) (  673,3000,757) (2011,3000,757)
> >    VC3 (-2011,2850,0) (2011,2850,0) (-2011,2850,757) (2011,3000,757)
> >    VC4 (-2011,2850,0) (2011,2850,0) (-2011,2850,757) (2011,3000,757)
> >    VC5 (-2011,2850,0) (2011,2850,0) (-2011,2850,757) (2011,3000,757)
> >    VC6 none
> >
> >Areas of capture for audio (from 12.1.1):
> >
> >        bottom left    bottom right  top left         top right
> >
> >    AC0 (-2011,2850,0) (-673,3000,0) (-2011,2850,757) (-673,3000,757)
> >    AC1 (  673,3000,0) (2011,2850,0) (  673,3000,757) (2011,3000,757)
> >    AC2 ( -673,3000,0) ( 673,3000,0) ( -673,3000,757) ( 673,3000,757)
> >    AC3 (-2011,2850,0) (2011,2850,0) (-2011,2850,757) (2011,3000,757)
> >    AC4 none
> >
> >Using a reference rather than area of capture:
> >     AC0 (VC0)
> >     AC1 (VC2)
> >     AC2 (VC1)
> >     AC3 (VC3,VC4,VC5)
> >I would think they are basically conveying the same information???
> >
> >Regards, Christian
> >
> >On 14/04/2014 12:26 PM, Duckworth, Mark wrote:
> >> Hello Christian,
> >> please see below.
> >> Mark
> >>
> >>> -----Original Message-----
> >>> From: clue [mailto:clue-bounces@ietf.org] On Behalf Of Christian
> >>> Groves
> >>> Sent: Sunday, April 13, 2014 10:08 PM
> >>> To: clue@ietf.org
> >>> Subject: Re: [clue] Improving treatment of audio
> >>>
> >>> Hello all,
> >>>
> >>> Perhaps the confusion is that some see Audio Capture area as a means
> >>>to  associate an Audio capture with a video capture as you've shown
> >>>Mark's  example. e.g. The audio spatial information isn't really used
> >>>for any audio  transformation other than associating it with a
> >>>particular video stream.
> >>>
> >>> Whereas others were more thinking the audio spatial information as
> >>>an input  to mixing and more complicated audio processing.
> >> [Duckworth, Mark] You could be right about this being a source of
> >>confusion.
> >>
> >>> If this is the case perhaps rather than linking ACs and VCs through
> >>> the physical or virtual co-ordinates of the area of capture
> >>> information, we simplify things and in each audio capture we say what
> VC it relates to?
> >>> e.g.using Mark's example AC0(VC0), AC3(V0,V1,V2)
> >> [Duckworth, Mark] I'm not sure how this would really work.  Because
> >>even in this simple example, we also have AC0 relates to VC3
> >>(sometimes), VC4, and VC5 (at least part of it), and so on.  AC3
> >>relates to all of VC1 through VC5.  I think using area of capture
> >>works better than trying to do something like this.
> >>
> >> Regards,
> >> Mark
> >>
> >>> Regards, Christian
> >>>
> >>> On 12/04/2014 7:42 AM, Duckworth, Mark wrote:
> >>>> Hi Paul,
> >>>> Yes, I understand the area of capture for audio can't be nearly as
> >>>>precise as
> >>> it can be for video.  But for this usage I think it doesn't matter.
> >>>> Mark
> >>>>
> >>>>> -----Original Message-----
> >>>>> From: Paul Coverdale [mailto:coverdale@sympatico.ca]
> >>>>> Sent: Friday, April 11, 2014 4:48 PM
> >>>>> To: Duckworth, Mark; 'John Leslie'; 'Christian Groves'
> >>>>> Cc: clue@ietf.org
> >>>>> Subject: RE: [clue] Improving treatment of audio
> >>>>>
> >>>>> Hi Mark,
> >>>>>
> >>>>> I can see what you're trying to do in the example given in
> >>>>>Framework  section 12.1.1. The problem I have is that, because of
> >>>>>the huge  difference in wavelength between light waves and audio
> >>>>>waves, you  can't really define an area of capture for audio in the
> >>>>>same way that  you can for video. A camera can focus on a scene
> >>>>>defined by 4  co-planar X,Y,Z coordinates. It will capture the
> >>>>>video inside that  quadrilateral, and nothing outside. A microphone
> >>>>>can capture audio  coming from inside of the same quadrilateral,
> >>>>>but it will also  capture a lot of audio from outside. What this
> >>>>>means is that we can  never define a unique association of an audio
> >>>>>capture with a video  capture based on pure geometrical
> >>>>>considerations, but we can still simply
> >>> define that ACx is associated with VCx, but maybe with VCy and VCz to=
o.
> >>>>> ...Paul
> >>>>>
> >>>>>> -----Original Message-----
> >>>>>> From: clue [mailto:clue-bounces@ietf.org] On Behalf Of Duckworth,
> >>>>>> Mark
> >>>>>> Sent: Friday, April 11, 2014 3:24 PM
> >>>>>> To: John Leslie; Christian Groves
> >>>>>> Cc: clue@ietf.org
> >>>>>> Subject: Re: [clue] Improving treatment of audio
> >>>>>>
> >>>>>> Hi all,
> >>>>>> The original intent of using "area of capture" for all media
> >>>>>>types,  including audio and video, was to provide a way to
> >>>>>>associate audio  with video.  So the consumer can choose captures
> >>>>>>that go together  and render them together.  An example of this is
> >>>>>>in the framework  document, section 12.1.1.  I still think this
> >>>>>>works fine for this purpose.
> >>>>>>
> >>>>>> In that example, the consumer can choose VC0, VC1, VC2, AC0, AC1,
> >>>>>> and AC2.  AC0 and VC0 have the same area of capture, indicating
> >>>>>> they are capturing the same area of the scene.  If the consumer
> >>>>>> wants to render
> >>>>>> AC0 from a loudspeaker close to the display for VC0 it can do so.
> >>>>>> Similarly, AC3 has an area of capture covering the whole scene
> >>>>>> (the full extent of the areas of VC0, VC1, VC2) so the consumer
> >>>>>> knows AC3 includes audio associated with all of VC0, VC1, and VC2.
> >>>>>>
> >>>>>> So I'm puzzled why you are proposing we remove the ability to use
> >>>>>> area of capture for audio for this purpose.
> >>>>>>
> >>>>>> Regards,
> >>>>>> Mark
> >>>>>>
> >>>>>>> -----Original Message-----
> >>>>>>> From: clue [mailto:clue-bounces@ietf.org] On Behalf Of John
> >>>>>>> Leslie
> >>>>>>> Sent: Friday, April 11, 2014 11:01 AM
> >>>>>>> To: Christian Groves
> >>>>>>> Cc: clue@ietf.org
> >>>>>>> Subject: Re: [clue] Improving treatment of audio
> >>>>>>>
> >>>>>>> Christian Groves <Christian.Groves@nteczone.com> wrote:
> >>>>>>>> "So it's not clear to me what extra we need to do from a CLUE
> >>>>>>>> perspective. What is the problem?"
> >>>>>>>> I guess this is the pertinent point.
> >>>>>>>>
> >>>>>>>>   From John's Audio 101 it seems that anything to do with
> >>>>>>>> spatial information that would relate to reverberation is in
> >>>>>>>> the too hard basket.
> >>>>>>>      What precisely does "too hard basket" mean?
> >>>>>>>
> >>>>>>>      I haven't even reached the point of suggesting what metrics
> >>>>>>> to define the _ability_ to send. To me, "too hard" merely means
> >>>>>>> that an individual site could choose not to send them or to
> >>>>>>> ignore them on
> >>>>>> receipt.
> >>>>>>>      But it sounds as if you're suggesting reverberation is "to
> >>>>>>>hard  to understand" and thus we should have no metrics about it.
> >>>>>>>
> >>>>>>>      I _hope_ that's not what you mean.
> >>>>>>>
> >>>>>>>> So it seems "area of capture" for audio could be marked "not
> >>>>>>>> applicable" in the framework.
> >>>>>>>      I hope so.
> >>>>>>>
> >>>>>>>> There doesn't seem to be any driver for having the "point of
> >>>>>> capture"
> >>>>>>>> apply to an audio capture either.
> >>>>>>>      I don't understand this. Point of capture for a microphone
> >>>>>>> may be "too hard" to track (today) for a microphone which moves;
> >>>>>>> nonetheless it seems to me the most fundamental metric there
> >>>>>>> could
> >>> be.
> >>>>>>>>   From the Audio101 there does seem to be a dependency on
> where
> >>> the
> >>>>>>>> microphone is located with respect to the person speaking (i.e.
> >>>>>>>> lapel mic, desk mic, room mic) to how it is handled at
> >>>>>> mixing/playout.
> >>>>>>>      I don't think I really got that far...
> >>>>>>>
> >>>>>>>      IMHO, it's more flexible to have mixing be the responsibily
> >>>>>>> of the receiver, but I haven't tried to specify that and I'm not
> >>>>>>> at all sure
> >>>>>> I want to specify that.
> >>>>>>>      I expect the actual sound systems in different rooms to
> >>>>>>> vary wildly, from one monaural speaker to stereo to full surround=
-
> sound.
> >>>>>>> Mixing for these without knowing which is the actual target
> >>>>>>> seems hard; but I'm sure there will be sites which prefer to do
> >>>>>>> so. A question which will arise, IMHO, is how to specify the
> >>>>>>> _intent_ of a mix generated in one room to be fed to other
> >>>>>>> rooms. (I'd prefer not to go there yet.)
> >>>>>>>
> >>>>>>>> Perhaps this is useful to signal via CLUE? If it is possible to
> >>>>>>>> signal this then perhaps tying a particular audio capture to a
> >>>>>>>> video capture makes sense?
> >>>>>>>      I'm not thinking along those lines. (That doesn't mean we
> >>>>>>> shouldn't think along those lines...) I'm thinking in terms of
> >>>>>>> providing several audio streams per room, associated with
> >>>>>>> position information about the position of the source of those
> >>>>>>> sounds, and allowing the receiver to choose how to mix them and
> >>>>>>> how to present the
> >>>>> mix in his/her room.
> >>>>>>>> i.e. a talker giving a presentation using a lapel mic captured
> >>>>>>>> by a particular video.
> >>>>>>>      In fact, there's only limited tendency for humans to
> >>>>>>>strictly  attach the sound they hear to the video they see.
> >>>>>>>Clearly, during a  presentation we want to _hear_ the presenter
> >>>>>>>talking, but having  the sound move back and forth as the speaker
> >>>>>>>walks can be confusing.
> >>>>>>>
> >>>>>>>      At the same time, we will want to hear the questions to
> >>>>>>> which the presenter may respond. It will be far easier on the
> >>>>>>> listener if these do _not_ seem to be coming from the same
> >>>>>>> physical position, especially when the questioner _is_ in the
> >>>>>>> same room as the
> >>> presenter.
> >>>>>>>      Hope this helps...
> >>>>>>>
> >>>>>>> --
> >>>>>>> John Leslie <john@jlc.net>
> >>>>>>>
> >>>>>>> _______________________________________________
> >>>>>>> clue mailing list
> >>>>>>> clue@ietf.org
> >>>>>>> https://www.ietf.org/mailman/listinfo/clue
> >>>>>> _______________________________________________
> >>>>>> clue mailing list
> >>>>>> clue@ietf.org
> >>>>>> https://www.ietf.org/mailman/listinfo/clue
> >>>> _______________________________________________
> >>>> clue mailing list
> >>>> clue@ietf.org
> >>>> https://www.ietf.org/mailman/listinfo/clue
> >>>>
> >>> _______________________________________________
> >>> clue mailing list
> >>> clue@ietf.org
> >>> https://www.ietf.org/mailman/listinfo/clue
> >
> >_______________________________________________
> >clue mailing list
> >clue@ietf.org
> >https://www.ietf.org/mailman/listinfo/clue


From nobody Fri Apr 25 14:42:31 2014
Return-Path: <Mark.Duckworth@polycom.com>
X-Original-To: clue@ietfa.amsl.com
Delivered-To: clue@ietfa.amsl.com
Received: from localhost (ietfa.amsl.com [127.0.0.1]) by ietfa.amsl.com (Postfix) with ESMTP id 4A5C01A03D8 for <clue@ietfa.amsl.com>; Fri, 25 Apr 2014 14:42:27 -0700 (PDT)
X-Virus-Scanned: amavisd-new at amsl.com
X-Spam-Flag: NO
X-Spam-Score: -1.82
X-Spam-Level: 
X-Spam-Status: No, score=-1.82 tagged_above=-999 required=5 tests=[BAYES_00=-1.9, RCVD_IN_DNSWL_LOW=-0.7, SPF_NEUTRAL=0.779, UNPARSEABLE_RELAY=0.001] autolearn=ham
Received: from mail.ietf.org ([4.31.198.44]) by localhost (ietfa.amsl.com [127.0.0.1]) (amavisd-new, port 10024) with ESMTP id iHQZSCK9ltjV for <clue@ietfa.amsl.com>; Fri, 25 Apr 2014 14:42:25 -0700 (PDT)
Received: from mail1.bemta7.messagelabs.com (mail1.bemta7.messagelabs.com [216.82.254.111]) by ietfa.amsl.com (Postfix) with ESMTP id D1FBB1A03D0 for <clue@ietf.org>; Fri, 25 Apr 2014 14:42:25 -0700 (PDT)
Received: from [216.82.254.20:36076] by server-15.bemta-7.messagelabs.com id 04/CC-32283-BB6DA535; Fri, 25 Apr 2014 21:42:19 +0000
X-Env-Sender: Mark.Duckworth@polycom.com
X-Msg-Ref: server-11.tower-47.messagelabs.com!1398462138!9971519!1
X-Originating-IP: [140.242.64.158]
X-StarScan-Received: 
X-StarScan-Version: 6.11.1; banners=-,-,-
X-VirusChecked: Checked
Received: (qmail 27447 invoked from network); 25 Apr 2014 21:42:18 -0000
Received: from crpehubprd01.polycom.com (HELO Crpehubprd01.polycom.com) (140.242.64.158) by server-11.tower-47.messagelabs.com with AES128-SHA encrypted SMTP; 25 Apr 2014 21:42:18 -0000
Received: from CRPMBOXPRD08.polycom.com ([169.254.1.94]) by Crpehubprd01.polycom.com ([fe80::5efe:10.236.0.158%14]) with mapi; Fri, 25 Apr 2014 14:42:17 -0700
From: "Duckworth, Mark" <Mark.Duckworth@polycom.com>
To: John Leslie <john@jlc.net>, "Johan Ludvig Nielsen (johaniel)" <johaniel@cisco.com>
Date: Fri, 25 Apr 2014 14:42:16 -0700
Thread-Topic: [clue] Improving treatment of audio
Thread-Index: Ac9gl8eqvxAOrKdFSwqHz/E4D2aA7wANz/kQ
Message-ID: <5C4AC54BFF7A0842A6A11F554D6FB52F0B97C9@CRPMBOXPRD08.polycom.com>
References: <BLU0-SMTP6910DA20C9A9287137B585D0550@phx.gbl> <53475572.4040807@nteczone.com> <20140411150055.GE60844@verdi> <49E45C59CA48264997FEBFB29B6BC2D6215295A8AE@CRPMBOXPRD07.polycom.com> <BLU0-SMTP1804A70E3E0F666EAEAE84D0540@phx.gbl> <49E45C59CA48264997FEBFB29B6BC2D62409219636@CRPMBOXPRD07.polycom.com> <534B4303.7060707@nteczone.com> <5C4AC54BFF7A0842A6A11F554D6FB52F07152B@CRPMBOXPRD08.polycom.com> <534B536B.8000205@nteczone.com> <CF803317.13BC2%johaniel@cisco.com> <20140425150518.GD44329@verdi>
In-Reply-To: <20140425150518.GD44329@verdi>
Accept-Language: en-US
Content-Language: en-US
X-MS-Has-Attach: 
X-MS-TNEF-Correlator: 
acceptlanguage: en-US
Content-Type: text/plain; charset="us-ascii"
Content-Transfer-Encoding: quoted-printable
MIME-Version: 1.0
Archived-At: http://mailarchive.ietf.org/arch/msg/clue/mUk08VpZWcf1do90wM-wBLF46gc
Cc: "clue@ietf.org" <clue@ietf.org>
Subject: Re: [clue] Improving treatment of audio
X-BeenThere: clue@ietf.org
X-Mailman-Version: 2.1.15
Precedence: list
List-Id: CLUE - ControLling mUltiple streams for TElepresence <clue.ietf.org>
List-Unsubscribe: <https://www.ietf.org/mailman/options/clue>, <mailto:clue-request@ietf.org?subject=unsubscribe>
List-Archive: <http://www.ietf.org/mail-archive/web/clue/>
List-Post: <mailto:clue@ietf.org>
List-Help: <mailto:clue-request@ietf.org?subject=help>
List-Subscribe: <https://www.ietf.org/mailman/listinfo/clue>, <mailto:clue-request@ietf.org?subject=subscribe>
X-List-Received-Date: Fri, 25 Apr 2014 21:42:27 -0000

Hi John,
Can you elaborate on your proposal about linking audio and video captures? =
 I'd like to see more detail.

Mark

> -----Original Message-----
> From: John Leslie [mailto:john@jlc.net]
> Sent: Friday, April 25, 2014 11:05 AM
> To: Johan Ludvig Nielsen (johaniel)
> Cc: Christian Groves; Duckworth, Mark; clue@ietf.org
> Subject: Re: [clue] Improving treatment of audio
>=20
> Johan Ludvig Nielsen (johaniel) <johaniel@cisco.com> wrote:
> >
> > My understanding of the current framework is that receivers calculate
> > the spatial association between an audio capture and the video scene
> > from the areas of capture. This will probably work.
>=20
>    Poorly, at best, IMHO. I'd like to discourage it.
>=20
> > But I still think having an explicit association of captures would be
> > a much simpler, more robust and more scalable solution.
>=20
>    I entirely agree. I'd like to propose an optional attribute of video c=
aptures to
> link audio captures in the area it covers.
>=20
> > From the point of view of the renderer, what is needed is info that
> > helps decide where to play out the audio streams it receives.
>=20
>    While I'm not enthusiastic about choosing that way, this seems an enti=
rely
> reasonable piece of information to supply.
>=20
> > More fancy spatial audio schemes exist and may see applications in
> > telepresence in the (somewhat distant) future, but will probably come
> > in the form of single stream multichannel formats where the spatial
> > information is embedded in the format like in surround sound, or
> > explicitly in the encoded stream like in spatial audio object coding.
>=20
>    Agreed -- that's not something to settle in our initial spec.
>=20
> --
> John Leslie <john@jlc.net>


From nobody Fri Apr 25 16:53:01 2014
Return-Path: <stephen.botzko@gmail.com>
X-Original-To: clue@ietfa.amsl.com
Delivered-To: clue@ietfa.amsl.com
Received: from localhost (ietfa.amsl.com [127.0.0.1]) by ietfa.amsl.com (Postfix) with ESMTP id 876E91A06D4 for <clue@ietfa.amsl.com>; Fri, 25 Apr 2014 16:53:00 -0700 (PDT)
X-Virus-Scanned: amavisd-new at amsl.com
X-Spam-Flag: NO
X-Spam-Score: -1.999
X-Spam-Level: 
X-Spam-Status: No, score=-1.999 tagged_above=-999 required=5 tests=[BAYES_00=-1.9, DKIM_SIGNED=0.1, DKIM_VALID=-0.1, DKIM_VALID_AU=-0.1, FREEMAIL_FROM=0.001, HTML_MESSAGE=0.001, SPF_PASS=-0.001] autolearn=ham
Received: from mail.ietf.org ([4.31.198.44]) by localhost (ietfa.amsl.com [127.0.0.1]) (amavisd-new, port 10024) with ESMTP id pFMLa2Ibt3Be for <clue@ietfa.amsl.com>; Fri, 25 Apr 2014 16:52:58 -0700 (PDT)
Received: from mail-ve0-x233.google.com (mail-ve0-x233.google.com [IPv6:2607:f8b0:400c:c01::233]) by ietfa.amsl.com (Postfix) with ESMTP id 88DBF1A06B9 for <clue@ietf.org>; Fri, 25 Apr 2014 16:52:58 -0700 (PDT)
Received: by mail-ve0-f179.google.com with SMTP id db12so5514743veb.38 for <clue@ietf.org>; Fri, 25 Apr 2014 16:52:51 -0700 (PDT)
DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20120113;  h=mime-version:in-reply-to:references:date:message-id:subject:from:to :cc:content-type; bh=S5buZcSV6ArpnRX0gQH2DUNShTsxySnWtuMbEGGWLL8=; b=sHmA+8o7yBh//o3ghHm0Iw9o+LOPARRU9aMGailSHNLoqerMAsb/cqB0iCoQ9niEFt Qj65LGoqoVkBF2OUUb6LLmJhHhoJOOXWohLVKxcme93OguKivQ6Dbvj3/BeP/oHLYzWM YHT+0zX8O6V7zbE7WPMT3XR9/hgUIZB6IKSyps2MXhp8VS++9Fq0JjTOmB75sBSCahD8 N+ALamItKm9jtOkI57qUYKfFQuqmSfMXLjWIRAwfbkn/+e6jpjVLSiWvIO1WHIHYAvcK 2YxDAlDImo5ui6ksytoP9AXoi/fJMUbSH3QY4McIvJ+wWH3B806Kj1Wfdn0y/9P3iOgV I2Yw==
MIME-Version: 1.0
X-Received: by 10.220.105.130 with SMTP id t2mr9220818vco.18.1398469971872; Fri, 25 Apr 2014 16:52:51 -0700 (PDT)
Received: by 10.221.40.135 with HTTP; Fri, 25 Apr 2014 16:52:51 -0700 (PDT)
In-Reply-To: <20140425102727.GB44329@verdi>
References: <533AF351.9050201@alum.mit.edu> <20140410222705.GW39240@verdi> <20140422144858.GC86778@verdi> <5356878F.4080504@alum.mit.edu> <CAMC7SJ7pu6ChnZuEaby496fQuiuk3jLHfTuLOZiifRpJaBoBXQ@mail.gmail.com> <CF7FE526.138C6%johaniel@cisco.com> <20140425102727.GB44329@verdi>
Date: Fri, 25 Apr 2014 19:52:51 -0400
Message-ID: <CAMC7SJ5TnD5i8TcQhLjcFf0LatawYhyFE6ogTE-xoEP62MjTyw@mail.gmail.com>
From: Stephen Botzko <stephen.botzko@gmail.com>
To: John Leslie <john@jlc.net>
Content-Type: multipart/alternative; boundary=047d7b343752e710f304f7e6a905
Archived-At: http://mailarchive.ietf.org/arch/msg/clue/ETIQRVqfJv0mmKhu4HLLMVOe3p8
Cc: CLUE <clue@ietf.org>, "Johan Ludvig Nielsen \(johaniel\)" <johaniel@cisco.com>
Subject: Re: [clue] Improving treatment of audio
X-BeenThere: clue@ietf.org
X-Mailman-Version: 2.1.15
Precedence: list
List-Id: CLUE - ControLling mUltiple streams for TElepresence <clue.ietf.org>
List-Unsubscribe: <https://www.ietf.org/mailman/options/clue>, <mailto:clue-request@ietf.org?subject=unsubscribe>
List-Archive: <http://www.ietf.org/mail-archive/web/clue/>
List-Post: <mailto:clue@ietf.org>
List-Help: <mailto:clue-request@ietf.org?subject=help>
List-Subscribe: <https://www.ietf.org/mailman/listinfo/clue>, <mailto:clue-request@ietf.org?subject=subscribe>
X-List-Received-Date: Fri, 25 Apr 2014 23:53:00 -0000

--047d7b343752e710f304f7e6a905
Content-Type: text/plain; charset=UTF-8

On Fri, Apr 25, 2014 at 6:27 AM, John Leslie <john@jlc.net> wrote:

> Johan Ludvig Nielsen (johaniel) <johaniel@cisco.com> wrote:
> >
> > I fully agree with those that argue that acoustic echo cancellation is
> > out of scope for CLUE.
>
>    I don't agree it's helpful to call it "out of scope," but...
>
> > It is a local responsibility and should be much higher than CLUE on
> > the feature list of any loudspeaking endpoint.
>
>    Yes, it's a local responsibility.
>
>    Note that not all endpoints of interest will have loudspeakers.
>
> > While a warning label can be put in some CLUE document, it is really
> > not necessary.
>
>    I much prefer a parameter saying the responsibility is satisfied
> to a warning label in some document nobody reads.
>

How could such a parameter be used?   If my system isn't using loudspeakers
then there is need for my system to do echo cancellation, so it would
presumably set the parameter to "no".  If my system does have loudspeakers,
then it does need to do it, and it would presumably set the parameter to
"yes".  In either case there is no action needed anywhere else.  So I don't
see how any other system could use such a parameter.



> > By echo I mean the audio signal received from the network, played out
> > on your local loudspeaker(s), reverberated through the room back to
> > the microphone(s) where it is superimposed on the the near-end talker
> > signal. This echo should be cancelled or suppressed before the
> > microphone signal is encoded and sent.
>
>    This is a rather good statement of the problem (better than mine,
> certainly).
>
>    How about a parameter saying "cancelled" vs. "suppressed"?
>
What would the use case be? Either way my rendering doesn't change.  Also,
echo cancellers all do some amount of suppression, so one would need to be
very careful on the definition.

>
>    BTW, audio is on the agenda for the May 13 design team call.
>
> --
> John Leslie <john@jlc.net>
>

--047d7b343752e710f304f7e6a905
Content-Type: text/html; charset=UTF-8
Content-Transfer-Encoding: quoted-printable

<div dir=3D"ltr"><br><div class=3D"gmail_extra"><br><br><div class=3D"gmail=
_quote">On Fri, Apr 25, 2014 at 6:27 AM, John Leslie <span dir=3D"ltr">&lt;=
<a href=3D"mailto:john@jlc.net" target=3D"_blank">john@jlc.net</a>&gt;</spa=
n> wrote:<br>
<blockquote class=3D"gmail_quote" style=3D"margin:0 0 0 .8ex;border-left:1p=
x #ccc solid;padding-left:1ex"><div class=3D"">Johan Ludvig Nielsen (johani=
el) &lt;<a href=3D"mailto:johaniel@cisco.com">johaniel@cisco.com</a>&gt; wr=
ote:<br>

&gt;<br>
&gt; I fully agree with those that argue that acoustic echo cancellation is=
<br>
&gt; out of scope for CLUE.<br>
<br>
</div>=C2=A0 =C2=A0I don&#39;t agree it&#39;s helpful to call it &quot;out =
of scope,&quot; but...<br>
<div class=3D""><br>
&gt; It is a local responsibility and should be much higher than CLUE on<br=
>
&gt; the feature list of any loudspeaking endpoint.<br>
<br>
</div>=C2=A0 =C2=A0Yes, it&#39;s a local responsibility.<br>
<br>
=C2=A0 =C2=A0Note that not all endpoints of interest will have loudspeakers=
.<br>
<div class=3D""><br>
&gt; While a warning label can be put in some CLUE document, it is really<b=
r>
&gt; not necessary.<br>
<br>
</div>=C2=A0 =C2=A0I much prefer a parameter saying the responsibility is s=
atisfied<br>
to a warning label in some document nobody reads.<br></blockquote><div><br>=
</div><div>How could such a parameter be used?=C2=A0=C2=A0 If my system isn=
&#39;t using loudspeakers then there is need for my system to do echo cance=
llation, so it would presumably set the parameter to &quot;no&quot;.=C2=A0 =
If my system does have loudspeakers, then it does need to do it, and it wou=
ld presumably set the parameter to &quot;yes&quot;.=C2=A0 In either case th=
ere is no action needed anywhere else.=C2=A0 So I don&#39;t see how any oth=
er system could use such a parameter.<br>
<br><br></div><blockquote class=3D"gmail_quote" style=3D"margin:0 0 0 .8ex;=
border-left:1px #ccc solid;padding-left:1ex">
<div class=3D""><br>
&gt; By echo I mean the audio signal received from the network, played out<=
br>
&gt; on your local loudspeaker(s), reverberated through the room back to<br=
>
&gt; the microphone(s) where it is superimposed on the the near-end talker<=
br>
&gt; signal. This echo should be cancelled or suppressed before the<br>
&gt; microphone signal is encoded and sent.<br>
<br>
</div>=C2=A0 =C2=A0This is a rather good statement of the problem (better t=
han mine,<br>
certainly).<br>
<br>
=C2=A0 =C2=A0How about a parameter saying &quot;cancelled&quot; vs. &quot;s=
uppressed&quot;?<br></blockquote><div>What would the use case be? Either wa=
y my rendering doesn&#39;t change.=C2=A0 Also, echo cancellers all do some =
amount of suppression, so one would need to be very careful on the definiti=
on.<br>
</div><blockquote class=3D"gmail_quote" style=3D"margin:0 0 0 .8ex;border-l=
eft:1px #ccc solid;padding-left:1ex">
<br>
=C2=A0 =C2=A0BTW, audio is on the agenda for the May 13 design team call.<b=
r>
<br>
--<br>
John Leslie &lt;<a href=3D"mailto:john@jlc.net">john@jlc.net</a>&gt;<br>
</blockquote></div><br></div></div>

--047d7b343752e710f304f7e6a905--


From nobody Sun Apr 27 18:19:17 2014
Return-Path: <Christian.Groves@nteczone.com>
X-Original-To: clue@ietfa.amsl.com
Delivered-To: clue@ietfa.amsl.com
Received: from localhost (ietfa.amsl.com [127.0.0.1]) by ietfa.amsl.com (Postfix) with ESMTP id 0388F1A0825 for <clue@ietfa.amsl.com>; Sun, 27 Apr 2014 18:19:15 -0700 (PDT)
X-Virus-Scanned: amavisd-new at amsl.com
X-Spam-Flag: NO
X-Spam-Score: -0.001
X-Spam-Level: 
X-Spam-Status: No, score=-0.001 tagged_above=-999 required=5 tests=[BAYES_40=-0.001] autolearn=ham
Received: from mail.ietf.org ([4.31.198.44]) by localhost (ietfa.amsl.com [127.0.0.1]) (amavisd-new, port 10024) with ESMTP id tZJ4GbcmZurK for <clue@ietfa.amsl.com>; Sun, 27 Apr 2014 18:19:13 -0700 (PDT)
Received: from cserver5.myshophosting.com (cserver5.myshophosting.com [175.107.161.1]) by ietfa.amsl.com (Postfix) with ESMTP id 22E101A082C for <clue@ietf.org>; Sun, 27 Apr 2014 18:19:12 -0700 (PDT)
Received: from ppp118-209-175-111.lns20.mel6.internode.on.net ([118.209.175.111]:51016 helo=[127.0.0.1]) by cserver5.myshophosting.com with esmtpsa (TLSv1:DHE-RSA-AES128-SHA:128) (Exim 4.82) (envelope-from <Christian.Groves@nteczone.com>) id 1WeaDu-0003VB-1t; Mon, 28 Apr 2014 11:19:10 +1000
Message-ID: <535DAC8E.7000403@nteczone.com>
Date: Mon, 28 Apr 2014 11:19:10 +1000
From: Christian Groves <Christian.Groves@nteczone.com>
User-Agent: Mozilla/5.0 (Windows NT 6.1; WOW64; rv:24.0) Gecko/20100101 Thunderbird/24.4.0
MIME-Version: 1.0
To: "Duckworth, Mark" <Mark.Duckworth@polycom.com>,  "Johan Ludvig Nielsen (johaniel)" <johaniel@cisco.com>, "clue@ietf.org" <clue@ietf.org>
References: <533AF351.9050201@alum.mit.edu> <20140410222705.GW39240@verdi> <BLU0-SMTP6910DA20C9A9287137B585D0550@phx.gbl> <53475572.4040807@nteczone.com> <20140411150055.GE60844@verdi> <49E45C59CA48264997FEBFB29B6BC2D6215295A8AE@CRPMBOXPRD07.polycom.com> <BLU0-SMTP1804A70E3E0F666EAEAE84D0540@phx.gbl> <49E45C59CA48264997FEBFB29B6BC2D62409219636@CRPMBOXPRD07.polycom.com> <534B4303.7060707@nteczone.com> <5C4AC54BFF7A0842A6A11F554D6FB52F07152B@CRPMBOXPRD08.polycom.com> <534B536B.8000205@nteczone.com> <CF803317.13BC2%johaniel@cisco.com> <5C4AC54BFF7A0842A6A11F554D6FB52F0B97BF@CRPMBOXPRD08.polycom.com>
In-Reply-To: <5C4AC54BFF7A0842A6A11F554D6FB52F0B97BF@CRPMBOXPRD08.polycom.com>
Content-Type: text/plain; charset=ISO-8859-1; format=flowed
Content-Transfer-Encoding: 7bit
X-AntiAbuse: This header was added to track abuse, please include it with any abuse report
X-AntiAbuse: Primary Hostname - cserver5.myshophosting.com
X-AntiAbuse: Original Domain - ietf.org
X-AntiAbuse: Originator/Caller UID/GID - [47 12] / [47 12]
X-AntiAbuse: Sender Address Domain - nteczone.com
X-Get-Message-Sender-Via: cserver5.myshophosting.com: authenticated_id: christian.groves@nteczone.com
X-Source: 
X-Source-Args: 
X-Source-Dir: 
Archived-At: http://mailarchive.ietf.org/arch/msg/clue/_RYBfrVzOn7lTwYcBiiQaKYaupM
Subject: Re: [clue] Improving treatment of audio
X-BeenThere: clue@ietf.org
X-Mailman-Version: 2.1.15
Precedence: list
List-Id: CLUE - ControLling mUltiple streams for TElepresence <clue.ietf.org>
List-Unsubscribe: <https://www.ietf.org/mailman/options/clue>, <mailto:clue-request@ietf.org?subject=unsubscribe>
List-Archive: <http://www.ietf.org/mail-archive/web/clue/>
List-Post: <mailto:clue@ietf.org>
List-Help: <mailto:clue-request@ietf.org?subject=help>
List-Subscribe: <https://www.ietf.org/mailman/listinfo/clue>, <mailto:clue-request@ietf.org?subject=subscribe>
X-List-Received-Date: Mon, 28 Apr 2014 01:19:15 -0000

Hello Mark,

Please see my comments below.

Regards, Christian
On 26/04/2014 7:28 AM, Duckworth, Mark wrote:
..snip..
>> But I still think having an explicit
>> association of captures would be a much simpler, more robust and more
>> scalable solution.
> [Duckworth, Mark] You might be right.  Christian's proposal is something like this, but to me it still seems incomplete and would need to continue using the spatial coordinates to make sense, which I think was not Christian's intent.  Does anybody want to make further proposals along these lines?
[CNG] Yes my intent was only to use one set of spatial co-ordinates i.e. 
the ones on the video captures. The reference in the audio captures 
indicates which video captures it relates to. Could you elaborate what 
seems to be incomplete?


From nobody Mon Apr 28 13:24:13 2014
Return-Path: <pkyzivat@alum.mit.edu>
X-Original-To: clue@ietfa.amsl.com
Delivered-To: clue@ietfa.amsl.com
Received: from localhost (ietfa.amsl.com [127.0.0.1]) by ietfa.amsl.com (Postfix) with ESMTP id 3862E1A702F for <clue@ietfa.amsl.com>; Mon, 28 Apr 2014 13:24:07 -0700 (PDT)
X-Virus-Scanned: amavisd-new at amsl.com
X-Spam-Flag: NO
X-Spam-Score: -1.235
X-Spam-Level: 
X-Spam-Status: No, score=-1.235 tagged_above=-999 required=5 tests=[BAYES_00=-1.9, DKIM_SIGNED=0.1, DKIM_VALID=-0.1, SPF_SOFTFAIL=0.665] autolearn=no
Received: from mail.ietf.org ([4.31.198.44]) by localhost (ietfa.amsl.com [127.0.0.1]) (amavisd-new, port 10024) with ESMTP id ZSx4Fzu8i_U7 for <clue@ietfa.amsl.com>; Mon, 28 Apr 2014 13:24:06 -0700 (PDT)
Received: from qmta10.westchester.pa.mail.comcast.net (qmta10.westchester.pa.mail.comcast.net [IPv6:2001:558:fe14:43:76:96:62:17]) by ietfa.amsl.com (Postfix) with ESMTP id C6C251A6FBA for <clue@ietf.org>; Mon, 28 Apr 2014 13:24:01 -0700 (PDT)
Received: from omta23.westchester.pa.mail.comcast.net ([76.96.62.74]) by qmta10.westchester.pa.mail.comcast.net with comcast id vY4e1n0031c6gX85AYQ02c; Mon, 28 Apr 2014 20:24:00 +0000
Received: from Paul-Kyzivats-MacBook-Pro.local ([50.138.229.164]) by omta23.westchester.pa.mail.comcast.net with comcast id vYQ01n00n3ZTu2S3jYQ0U6; Mon, 28 Apr 2014 20:24:00 +0000
Message-ID: <535EB8E0.3080605@alum.mit.edu>
Date: Mon, 28 Apr 2014 16:24:00 -0400
From: Paul Kyzivat <pkyzivat@alum.mit.edu>
User-Agent: Mozilla/5.0 (Macintosh; Intel Mac OS X 10.7; rv:24.0) Gecko/20100101 Thunderbird/24.4.0
MIME-Version: 1.0
To: clue@ietf.org
References: <BLU0-SMTP6910DA20C9A9287137B585D0550@phx.gbl> <53475572.4040807@nteczone.com> <20140411150055.GE60844@verdi> <49E45C59CA48264997FEBFB29B6BC2D6215295A8AE@CRPMBOXPRD07.polycom.com> <BLU0-SMTP1804A70E3E0F666EAEAE84D0540@phx.gbl> <49E45C59CA48264997FEBFB29B6BC2D62409219636@CRPMBOXPRD07.polycom.com> <534B4303.7060707@nteczone.com> <5C4AC54BFF7A0842A6A11F554D6FB52F07152B@CRPMBOXPRD08.polycom.com> <534B536B.8000205@nteczone.com> <CF803317.13BC2%johaniel@cisco.com> <20140425150518.GD44329@verdi>
In-Reply-To: <20140425150518.GD44329@verdi>
Content-Type: text/plain; charset=ISO-8859-1; format=flowed
Content-Transfer-Encoding: 7bit
DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=comcast.net; s=q20140121; t=1398716640; bh=oeHQNzD3VZYBxazUDSmIl7UvxwDsOJHxx91nlsjQgRE=; h=Received:Received:Message-ID:Date:From:MIME-Version:To:Subject: Content-Type; b=W5cGOtDE+/vJhjiodGsUfZnzEpMJ0Q4EmjE6e7/2RAw4gAqKkLt8ezAAe8T4xNRD7 UT8eh3uNzJR9b1YtAo5D7AM5qzztnT2oV2ko5wV4Iog1+KUCYK4tJs5OycPjctAUpM cFhpSt0OWlk+ZhtvPSb1UjLOSiAxWaxSk03AiiicLmi79Wzx4aarM7peqsx69bscEc KlgGXalNTdYsEk8yV9mpydlGm5BJ30w/HXIghbpEpoj0GbeEtO5ZUYgMwP5Iwg0rqs VZgwW1cLFeEQgqx0lxwrJd1AFVFbW+SLsNzCbjeL1Efd1xp1llONuZYjLJOhvlqnsq tJr1RNCT03NFQ==
Archived-At: http://mailarchive.ietf.org/arch/msg/clue/EN35PIpyc_TtxBp0Y6XfGxnMR1o
Subject: Re: [clue] Improving treatment of audio
X-BeenThere: clue@ietf.org
X-Mailman-Version: 2.1.15
Precedence: list
List-Id: CLUE - ControLling mUltiple streams for TElepresence <clue.ietf.org>
List-Unsubscribe: <https://www.ietf.org/mailman/options/clue>, <mailto:clue-request@ietf.org?subject=unsubscribe>
List-Archive: <http://www.ietf.org/mail-archive/web/clue/>
List-Post: <mailto:clue@ietf.org>
List-Help: <mailto:clue-request@ietf.org?subject=help>
List-Subscribe: <https://www.ietf.org/mailman/listinfo/clue>, <mailto:clue-request@ietf.org?subject=subscribe>
X-List-Received-Date: Mon, 28 Apr 2014 20:24:07 -0000

On 4/25/14 11:05 AM, John Leslie wrote:
> Johan Ludvig Nielsen (johaniel) <johaniel@cisco.com> wrote:
>>
>> My understanding of the current framework is that receivers calculate the
>> spatial association between an audio capture and the video scene from the
>> areas of capture. This will probably work.
>
>     Poorly, at best, IMHO. I'd like to discourage it.

I still don't understand how it would work.

The area of capture is a planar quadrilateral. But of course neither a 
camera nor a microphone capture their media solely from a planar 
quadrilateral in three-space. (Though this might be a reasonable 
representation of a powerpoint or screen capture.)

The camera captures roughly a pyramid shaped region of the scene, with 
the point of capture being the peak of the pyramid, and the area of 
capture bounding the intersection of the pyramid and a plane. Not all of 
that will be in focus, and some of it may be occluded. I guess typically 
the content of interest will fall in the portion of the pyramid behind 
the area of capture.

How to relate the area of capture of an audio capture to what is going 
on in the scene seems much less obvious. Does it mean it only captures 
audio emitted within the pyramid? Or that it captures audio that may be 
heard from within the pyramid?

*How* do I relate the area of capture of an audio and video capture to 
decide what audio captures I need to make sense of the video capture? Do 
I intersect the planar areas of capture? (Which yields nothing if they 
aren't coplanar.) Or do I intersect the pyramids?

Or, as John has suggested, will I always need all the audio from a scene 
to understand any video from the scene?

If this isn't defined well we may get very poor interoperation.

	Thanks,
	Paul


>> But I still think having an explicit association of captures would be
>> a much simpler, more robust and more scalable solution.
>
>     I entirely agree. I'd like to propose an optional attribute of
> video captures to link audio captures in the area it covers.
>
>>  From the point of view of the renderer, what is needed is info that
>> helps decide where to play out the audio streams it receives.
>
>     While I'm not enthusiastic about choosing that way, this seems an
> entirely reasonable piece of information to supply.
>
>> More fancy spatial audio schemes exist and may see applications in
>> telepresence in the (somewhat distant) future, but will probably come in
>> the form of single stream multichannel formats where the spatial
>> information is embedded in the format like in surround sound, or
>> explicitly in the encoded stream like in spatial audio object coding.
>
>     Agreed -- that's not something to settle in our initial spec.
>
> --
> John Leslie <john@jlc.net>
>
> _______________________________________________
> clue mailing list
> clue@ietf.org
> https://www.ietf.org/mailman/listinfo/clue
>


From nobody Mon Apr 28 16:09:42 2014
Return-Path: <coverdale@sympatico.ca>
X-Original-To: clue@ietfa.amsl.com
Delivered-To: clue@ietfa.amsl.com
Received: from localhost (ietfa.amsl.com [127.0.0.1]) by ietfa.amsl.com (Postfix) with ESMTP id 03E9F1A6FB3 for <clue@ietfa.amsl.com>; Mon, 28 Apr 2014 16:09:41 -0700 (PDT)
X-Virus-Scanned: amavisd-new at amsl.com
X-Spam-Flag: NO
X-Spam-Score: -0.5
X-Spam-Level: 
X-Spam-Status: No, score=-0.5 tagged_above=-999 required=5 tests=[BAYES_05=-0.5, MSGID_FROM_MTA_HEADER=0.001, RCVD_IN_DNSWL_NONE=-0.0001, SPF_PASS=-0.001] autolearn=ham
Received: from mail.ietf.org ([4.31.198.44]) by localhost (ietfa.amsl.com [127.0.0.1]) (amavisd-new, port 10024) with ESMTP id rQN79dYZnTUG for <clue@ietfa.amsl.com>; Mon, 28 Apr 2014 16:09:38 -0700 (PDT)
Received: from blu0-omc1-s36.blu0.hotmail.com (blu0-omc1-s36.blu0.hotmail.com [65.55.116.47]) by ietfa.amsl.com (Postfix) with ESMTP id 711021A6F17 for <clue@ietf.org>; Mon, 28 Apr 2014 16:09:38 -0700 (PDT)
Received: from BLU0-SMTP62 ([65.55.116.7]) by blu0-omc1-s36.blu0.hotmail.com with Microsoft SMTPSVC(6.0.3790.4675);  Mon, 28 Apr 2014 16:09:37 -0700
X-TMN: [XbZ3upG83h2wbyx9eArSpQBTCHggdJVP]
X-Originating-Email: [coverdale@sympatico.ca]
Message-ID: <BLU0-SMTP6222323588C548A1847EF2D0470@phx.gbl>
Received: from PaulNewPC ([184.147.38.66]) by BLU0-SMTP62.phx.gbl over TLS secured channel with Microsoft SMTPSVC(6.0.3790.4675);  Mon, 28 Apr 2014 16:09:36 -0700
From: Paul Coverdale <coverdale@sympatico.ca>
To: "'Paul Kyzivat'" <pkyzivat@alum.mit.edu>, <clue@ietf.org>
References: <BLU0-SMTP6910DA20C9A9287137B585D0550@phx.gbl> <53475572.4040807@nteczone.com> <20140411150055.GE60844@verdi> <49E45C59CA48264997FEBFB29B6BC2D6215295A8AE@CRPMBOXPRD07.polycom.com> <BLU0-SMTP1804A70E3E0F666EAEAE84D0540@phx.gbl> <49E45C59CA48264997FEBFB29B6BC2D62409219636@CRPMBOXPRD07.polycom.com> <534B4303.7060707@nteczone.com> <5C4AC54BFF7A0842A6A11F554D6FB52F07152B@CRPMBOXPRD08.polycom.com> <534B536B.8000205@nteczone.com> <CF803317.13BC2%johaniel@cisco.com> <20140425150518.GD44329@verdi> <535EB8E0.3080605@alum.mit.edu>
In-Reply-To: <535EB8E0.3080605@alum.mit.edu>
Date: Mon, 28 Apr 2014 19:09:33 -0400
MIME-Version: 1.0
Content-Type: text/plain; charset="us-ascii"
Content-Transfer-Encoding: 7bit
X-Mailer: Microsoft Office Outlook 12.0
Thread-Index: Ac9jH9B8kF79dGMoTuuSCsEU2x9S+gAFmsMg
Content-Language: en-us
X-OriginalArrivalTime: 28 Apr 2014 23:09:36.0964 (UTC) FILETIME=[ECD29040:01CF6336]
Archived-At: http://mailarchive.ietf.org/arch/msg/clue/zP6VyVCfj9ZhoTDlLRINZ8QWNZ8
Subject: Re: [clue] Improving treatment of audio
X-BeenThere: clue@ietf.org
X-Mailman-Version: 2.1.15
Precedence: list
List-Id: CLUE - ControLling mUltiple streams for TElepresence <clue.ietf.org>
List-Unsubscribe: <https://www.ietf.org/mailman/options/clue>, <mailto:clue-request@ietf.org?subject=unsubscribe>
List-Archive: <http://www.ietf.org/mail-archive/web/clue/>
List-Post: <mailto:clue@ietf.org>
List-Help: <mailto:clue-request@ietf.org?subject=help>
List-Subscribe: <https://www.ietf.org/mailman/listinfo/clue>, <mailto:clue-request@ietf.org?subject=subscribe>
X-List-Received-Date: Mon, 28 Apr 2014 23:09:41 -0000

Comments in-line...

>-----Original Message-----
>From: clue [mailto:clue-bounces@ietf.org] On Behalf Of Paul Kyzivat
>Sent: Monday, April 28, 2014 4:24 PM
>To: clue@ietf.org
>Subject: Re: [clue] Improving treatment of audio
>
>On 4/25/14 11:05 AM, John Leslie wrote:
>> Johan Ludvig Nielsen (johaniel) <johaniel@cisco.com> wrote:
>>>
>>> My understanding of the current framework is that receivers calculate
>>> the spatial association between an audio capture and the video scene
>>> from the areas of capture. This will probably work.
>>
>>     Poorly, at best, IMHO. I'd like to discourage it.
>
>I still don't understand how it would work.

[PVC] Me neither.

>
>The area of capture is a planar quadrilateral. But of course neither a
>camera nor a microphone capture their media solely from a planar
>quadrilateral in three-space. (Though this might be a reasonable
>representation of a powerpoint or screen capture.)
>
>The camera captures roughly a pyramid shaped region of the scene, with
>the point of capture being the peak of the pyramid, and the area of
>capture bounding the intersection of the pyramid and a plane. Not all of
>that will be in focus, and some of it may be occluded. I guess typically
>the content of interest will fall in the portion of the pyramid behind
>the area of capture.
>
>How to relate the area of capture of an audio capture to what is going
>on in the scene seems much less obvious. Does it mean it only captures
>audio emitted within the pyramid? Or that it captures audio that may be
>heard from within the pyramid?

[PVC] I think the sooner we give up trying to define an audio area of
capture in the same manner as a video area of capture, the better. It just
doesn't work.

>
>*How* do I relate the area of capture of an audio and video capture to
>decide what audio captures I need to make sense of the video capture? Do
>I intersect the planar areas of capture? (Which yields nothing if they
>aren't coplanar.) Or do I intersect the pyramids?
>
>Or, as John has suggested, will I always need all the audio from a scene
>to understand any video from the scene?
>
>If this isn't defined well we may get very poor interoperation.
>
>	Thanks,
>	Paul
>
>



From nobody Mon Apr 28 18:37:03 2014
Return-Path: <Mark.Duckworth@polycom.com>
X-Original-To: clue@ietfa.amsl.com
Delivered-To: clue@ietfa.amsl.com
Received: from localhost (ietfa.amsl.com [127.0.0.1]) by ietfa.amsl.com (Postfix) with ESMTP id A18951A6FA9 for <clue@ietfa.amsl.com>; Mon, 28 Apr 2014 18:37:01 -0700 (PDT)
X-Virus-Scanned: amavisd-new at amsl.com
X-Spam-Flag: NO
X-Spam-Score: -1.82
X-Spam-Level: 
X-Spam-Status: No, score=-1.82 tagged_above=-999 required=5 tests=[BAYES_00=-1.9, RCVD_IN_DNSWL_LOW=-0.7, SPF_NEUTRAL=0.779, UNPARSEABLE_RELAY=0.001] autolearn=ham
Received: from mail.ietf.org ([4.31.198.44]) by localhost (ietfa.amsl.com [127.0.0.1]) (amavisd-new, port 10024) with ESMTP id cOlCzgiDMAQh for <clue@ietfa.amsl.com>; Mon, 28 Apr 2014 18:37:00 -0700 (PDT)
Received: from mail1.bemta7.messagelabs.com (mail1.bemta7.messagelabs.com [216.82.254.112]) by ietfa.amsl.com (Postfix) with ESMTP id 06BED1A886B for <clue@ietf.org>; Mon, 28 Apr 2014 18:36:59 -0700 (PDT)
Received: from [216.82.254.19:12964] by server-16.bemta-7.messagelabs.com id F2/A6-25346-B320F535; Tue, 29 Apr 2014 01:36:59 +0000
X-Env-Sender: Mark.Duckworth@polycom.com
X-Msg-Ref: server-7.tower-96.messagelabs.com!1398735417!1505434!1
X-Originating-IP: [140.242.64.158]
X-StarScan-Received: 
X-StarScan-Version: 6.11.3; banners=-,-,-
X-VirusChecked: Checked
Received: (qmail 10377 invoked from network); 29 Apr 2014 01:36:58 -0000
Received: from crpehubprd01.polycom.com (HELO crpehubprd02.polycom.com) (140.242.64.158) by server-7.tower-96.messagelabs.com with AES128-SHA encrypted SMTP; 29 Apr 2014 01:36:58 -0000
Received: from CRPMBOXPRD08.polycom.com ([169.254.1.94]) by crpehubprd02.polycom.com ([::1]) with mapi; Mon, 28 Apr 2014 18:36:10 -0700
From: "Duckworth, Mark" <Mark.Duckworth@polycom.com>
To: Christian Groves <Christian.Groves@nteczone.com>, "clue@ietf.org" <clue@ietf.org>
Date: Mon, 28 Apr 2014 18:36:06 -0700
Thread-Topic: [clue] Improving treatment of audio
Thread-Index: Ac9XkCn3b7gv5XkmR4+PhgMa7p16/wLuJidw
Message-ID: <5C4AC54BFF7A0842A6A11F554D6FB52F13D9B2@CRPMBOXPRD08.polycom.com>
References: <533AF351.9050201@alum.mit.edu> <20140410222705.GW39240@verdi> <BLU0-SMTP6910DA20C9A9287137B585D0550@phx.gbl> <53475572.4040807@nteczone.com> <20140411150055.GE60844@verdi> <49E45C59CA48264997FEBFB29B6BC2D6215295A8AE@CRPMBOXPRD07.polycom.com> <BLU0-SMTP1804A70E3E0F666EAEAE84D0540@phx.gbl> <49E45C59CA48264997FEBFB29B6BC2D62409219636@CRPMBOXPRD07.polycom.com> <534B4303.7060707@nteczone.com> <5C4AC54BFF7A0842A6A11F554D6FB52F07152B@CRPMBOXPRD08.polycom.com> <534B536B.8000205@nteczone.com>
In-Reply-To: <534B536B.8000205@nteczone.com>
Accept-Language: en-US
Content-Language: en-US
X-MS-Has-Attach: 
X-MS-TNEF-Correlator: 
acceptlanguage: en-US
Content-Type: text/plain; charset="us-ascii"
Content-Transfer-Encoding: quoted-printable
MIME-Version: 1.0
Archived-At: http://mailarchive.ietf.org/arch/msg/clue/IsXndevkIflO-OrcCSKzxpNxT48
Subject: Re: [clue] Improving treatment of audio
X-BeenThere: clue@ietf.org
X-Mailman-Version: 2.1.15
Precedence: list
List-Id: CLUE - ControLling mUltiple streams for TElepresence <clue.ietf.org>
List-Unsubscribe: <https://www.ietf.org/mailman/options/clue>, <mailto:clue-request@ietf.org?subject=unsubscribe>
List-Archive: <http://www.ietf.org/mail-archive/web/clue/>
List-Post: <mailto:clue@ietf.org>
List-Help: <mailto:clue-request@ietf.org?subject=help>
List-Subscribe: <https://www.ietf.org/mailman/listinfo/clue>, <mailto:clue-request@ietf.org?subject=subscribe>
X-List-Received-Date: Tue, 29 Apr 2014 01:37:01 -0000

Hi Christian,
I'm coming back to this message, because I'm trying to understand your prop=
osal.  You proposed "in each audio capture we say what VC it relates to."  =
What does this mean?  What does it mean for an audio capture to relate to a=
 video capture?  What exactly is the producer telling the consumer?  Can an=
 audio capture relate to more than one video capture?

More comments inline below.

Regards,
Mark

> -----Original Message-----
> From: Christian Groves [mailto:Christian.Groves@nteczone.com]
> Sent: Sunday, April 13, 2014 11:18 PM
> To: Duckworth, Mark; clue@ietf.org
> Subject: Re: [clue] Improving treatment of audio
>=20
> Hello Mark,
>=20
> Sorry I didn't consider the entire 12.1.1. if I do that, according to tha=
t
> example:
>=20
> Video areas of capture:
>=20
>         bottom left    bottom right  top left         top right
>     VC0 (-2011,2850,0) (-673,3000,0) (-2011,2850,757) (-673,3000,757)
>     VC1 ( -673,3000,0) ( 673,3000,0) ( -673,3000,757) ( 673,3000,757)
>     VC2 (  673,3000,0) (2011,2850,0) (  673,3000,757) (2011,3000,757)
>     VC3 (-2011,2850,0) (2011,2850,0) (-2011,2850,757) (2011,3000,757)
>     VC4 (-2011,2850,0) (2011,2850,0) (-2011,2850,757) (2011,3000,757)
>     VC5 (-2011,2850,0) (2011,2850,0) (-2011,2850,757) (2011,3000,757)
>     VC6 none
>=20
> Areas of capture for audio (from 12.1.1):
>=20
>         bottom left    bottom right  top left         top right
>=20
>     AC0 (-2011,2850,0) (-673,3000,0) (-2011,2850,757) (-673,3000,757)
>     AC1 (  673,3000,0) (2011,2850,0) (  673,3000,757) (2011,3000,757)
>     AC2 ( -673,3000,0) ( 673,3000,0) ( -673,3000,757) ( 673,3000,757)
>     AC3 (-2011,2850,0) (2011,2850,0) (-2011,2850,757) (2011,3000,757)
>     AC4 none
>=20
> Using a reference rather than area of capture:
>      AC0 (VC0)
>      AC1 (VC2)
>      AC2 (VC1)
>      AC3 (VC3,VC4,VC5)
> I would think they are basically conveying the same information???

[Duckworth, Mark] Looking from the consumer point of view, in trying to und=
erstand the advertisement, what does the consumer know from the advertiseme=
nt?  Suppose the consumer wants to get all three of VC0, VC1, VC2, but it o=
nly wants a single audio capture.  It should probably ask for AC3, but acco=
rding to the advertisement AC3 doesn't "relate to" VC0, VC1, or VC2.

[Duckworth, Mark] For a case with a different consumer, suppose the consume=
r wants the single video capture VC5, but it would like to receive and rend=
er spatial audio.  There is no audio capture with multiple channels, so it =
could ask for AC0, AC1, and AC2 if it knew how they were spatially "related=
 to" VC5.  But according to the advertisement, AC0, AC1 and AC2 aren't "rel=
ated to" VC5 at all.

>=20
> Regards, Christian
>=20
> On 14/04/2014 12:26 PM, Duckworth, Mark wrote:
> > Hello Christian,
> > please see below.
> > Mark
> >
> >> -----Original Message-----
> >> From: clue [mailto:clue-bounces@ietf.org] On Behalf Of Christian
> >> Groves
> >> Sent: Sunday, April 13, 2014 10:08 PM
> >> To: clue@ietf.org
> >> Subject: Re: [clue] Improving treatment of audio
> >>
> >> Hello all,
> >>
> >> Perhaps the confusion is that some see Audio Capture area as a means
> >> to associate an Audio capture with a video capture as you've shown
> >> Mark's example. e.g. The audio spatial information isn't really used
> >> for any audio transformation other than associating it with a particul=
ar
> video stream.
> >>
> >> Whereas others were more thinking the audio spatial information as an
> >> input to mixing and more complicated audio processing.
> > [Duckworth, Mark] You could be right about this being a source of
> confusion.
> >
> >> If this is the case perhaps rather than linking ACs and VCs through
> >> the physical or virtual co-ordinates of the area of capture
> >> information, we simplify things and in each audio capture we say what =
VC
> it relates to?
> >> e.g.using Mark's example AC0(VC0), AC3(V0,V1,V2)
> > [Duckworth, Mark] I'm not sure how this would really work.  Because eve=
n
> in this simple example, we also have AC0 relates to VC3 (sometimes), VC4,
> and VC5 (at least part of it), and so on.  AC3 relates to all of VC1 thro=
ugh VC5.
> I think using area of capture works better than trying to do something li=
ke
> this.
...snip...=20


From nobody Mon Apr 28 19:15:32 2014
Return-Path: <Christian.Groves@nteczone.com>
X-Original-To: clue@ietfa.amsl.com
Delivered-To: clue@ietfa.amsl.com
Received: from localhost (ietfa.amsl.com [127.0.0.1]) by ietfa.amsl.com (Postfix) with ESMTP id 407621A8870 for <clue@ietfa.amsl.com>; Mon, 28 Apr 2014 19:15:30 -0700 (PDT)
X-Virus-Scanned: amavisd-new at amsl.com
X-Spam-Flag: NO
X-Spam-Score: -1.9
X-Spam-Level: 
X-Spam-Status: No, score=-1.9 tagged_above=-999 required=5 tests=[BAYES_00=-1.9] autolearn=ham
Received: from mail.ietf.org ([4.31.198.44]) by localhost (ietfa.amsl.com [127.0.0.1]) (amavisd-new, port 10024) with ESMTP id LTDgLqlYBY-u for <clue@ietfa.amsl.com>; Mon, 28 Apr 2014 19:15:28 -0700 (PDT)
Received: from cserver5.myshophosting.com (cserver5.myshophosting.com [175.107.161.1]) by ietfa.amsl.com (Postfix) with ESMTP id 592AD1A885E for <clue@ietf.org>; Mon, 28 Apr 2014 19:15:28 -0700 (PDT)
Received: from ppp118-209-168-86.lns20.mel6.internode.on.net ([118.209.168.86]:52492 helo=[127.0.0.1]) by cserver5.myshophosting.com with esmtpsa (TLSv1:DHE-RSA-AES128-SHA:128) (Exim 4.82) (envelope-from <Christian.Groves@nteczone.com>) id 1WexZs-0003hY-AD; Tue, 29 Apr 2014 12:15:24 +1000
Message-ID: <535F0B3A.70808@nteczone.com>
Date: Tue, 29 Apr 2014 12:15:22 +1000
From: Christian Groves <Christian.Groves@nteczone.com>
User-Agent: Mozilla/5.0 (Windows NT 6.1; WOW64; rv:24.0) Gecko/20100101 Thunderbird/24.4.0
MIME-Version: 1.0
To: "Duckworth, Mark" <Mark.Duckworth@polycom.com>,  "clue@ietf.org" <clue@ietf.org>
References: <533AF351.9050201@alum.mit.edu> <20140410222705.GW39240@verdi> <BLU0-SMTP6910DA20C9A9287137B585D0550@phx.gbl> <53475572.4040807@nteczone.com> <20140411150055.GE60844@verdi> <49E45C59CA48264997FEBFB29B6BC2D6215295A8AE@CRPMBOXPRD07.polycom.com> <BLU0-SMTP1804A70E3E0F666EAEAE84D0540@phx.gbl> <49E45C59CA48264997FEBFB29B6BC2D62409219636@CRPMBOXPRD07.polycom.com> <534B4303.7060707@nteczone.com> <5C4AC54BFF7A0842A6A11F554D6FB52F07152B@CRPMBOXPRD08.polycom.com> <534B536B.8000205@nteczone.com> <5C4AC54BFF7A0842A6A11F554D6FB52F13D9B2@CRPMBOXPRD08.polycom.com>
In-Reply-To: <5C4AC54BFF7A0842A6A11F554D6FB52F13D9B2@CRPMBOXPRD08.polycom.com>
Content-Type: text/plain; charset=ISO-8859-1; format=flowed
Content-Transfer-Encoding: 7bit
X-AntiAbuse: This header was added to track abuse, please include it with any abuse report
X-AntiAbuse: Primary Hostname - cserver5.myshophosting.com
X-AntiAbuse: Original Domain - ietf.org
X-AntiAbuse: Originator/Caller UID/GID - [47 12] / [47 12]
X-AntiAbuse: Sender Address Domain - nteczone.com
X-Get-Message-Sender-Via: cserver5.myshophosting.com: authenticated_id: christian.groves@nteczone.com
X-Source: 
X-Source-Args: 
X-Source-Dir: 
Archived-At: http://mailarchive.ietf.org/arch/msg/clue/bz9I-n6tW5zy8yATaHnvUhlt4Wk
Subject: Re: [clue] Improving treatment of audio
X-BeenThere: clue@ietf.org
X-Mailman-Version: 2.1.15
Precedence: list
List-Id: CLUE - ControLling mUltiple streams for TElepresence <clue.ietf.org>
List-Unsubscribe: <https://www.ietf.org/mailman/options/clue>, <mailto:clue-request@ietf.org?subject=unsubscribe>
List-Archive: <http://www.ietf.org/mail-archive/web/clue/>
List-Post: <mailto:clue@ietf.org>
List-Help: <mailto:clue-request@ietf.org?subject=help>
List-Subscribe: <https://www.ietf.org/mailman/listinfo/clue>, <mailto:clue-request@ietf.org?subject=subscribe>
X-List-Received-Date: Tue, 29 Apr 2014 02:15:30 -0000

Hello Mark,

Please see below.

Regards, Christian

On 29/04/2014 11:36 AM, Duckworth, Mark wrote:
> Hi Christian,
> I'm coming back to this message, because I'm trying to understand your proposal.  You proposed "in each audio capture we say what VC it relates to."  What does this mean?  What does it mean for an audio capture to relate to a video capture?  What exactly is the producer telling the consumer?  Can an audio capture relate to more than one video capture?
[CNG] The meaning has largely the same semantic as if the ADV used the 
same capture area on a VC and AC. Basically its saying that the audio 
comes from the same "region" as what the video capture does. I think the 
difference is that its not trying to put a mathematical certainty about 
what that region is. Its a way for the producer to say to the consumer 
if you choose video capture A you probably want to choose audio capture 
B for a good experience. In terms of whether an audio capture can relate 
to more than one video capture yes I think so. You could have one audio 
capture for a room and three video captures.
>
> More comments inline below.
>
> Regards,
> Mark
>
>> -----Original Message-----
>> From: Christian Groves [mailto:Christian.Groves@nteczone.com]
>> Sent: Sunday, April 13, 2014 11:18 PM
>> To: Duckworth, Mark; clue@ietf.org
>> Subject: Re: [clue] Improving treatment of audio
>>
>> Hello Mark,
>>
>> Sorry I didn't consider the entire 12.1.1. if I do that, according to that
>> example:
>>
>> Video areas of capture:
>>
>>          bottom left    bottom right  top left         top right
>>      VC0 (-2011,2850,0) (-673,3000,0) (-2011,2850,757) (-673,3000,757)
>>      VC1 ( -673,3000,0) ( 673,3000,0) ( -673,3000,757) ( 673,3000,757)
>>      VC2 (  673,3000,0) (2011,2850,0) (  673,3000,757) (2011,3000,757)
>>      VC3 (-2011,2850,0) (2011,2850,0) (-2011,2850,757) (2011,3000,757)
>>      VC4 (-2011,2850,0) (2011,2850,0) (-2011,2850,757) (2011,3000,757)
>>      VC5 (-2011,2850,0) (2011,2850,0) (-2011,2850,757) (2011,3000,757)
>>      VC6 none
>>
>> Areas of capture for audio (from 12.1.1):
>>
>>          bottom left    bottom right  top left         top right
>>
>>      AC0 (-2011,2850,0) (-673,3000,0) (-2011,2850,757) (-673,3000,757)
>>      AC1 (  673,3000,0) (2011,2850,0) (  673,3000,757) (2011,3000,757)
>>      AC2 ( -673,3000,0) ( 673,3000,0) ( -673,3000,757) ( 673,3000,757)
>>      AC3 (-2011,2850,0) (2011,2850,0) (-2011,2850,757) (2011,3000,757)
>>      AC4 none
>>
>> Using a reference rather than area of capture:
>>       AC0 (VC0)
>>       AC1 (VC2)
>>       AC2 (VC1)
>>       AC3 (VC3,VC4,VC5)
>> I would think they are basically conveying the same information???
> [Duckworth, Mark] Looking from the consumer point of view, in trying to understand the advertisement, what does the consumer know from the advertisement?  Suppose the consumer wants to get all three of VC0, VC1, VC2, but it only wants a single audio capture.  It should probably ask for AC3, but according to the advertisement AC3 doesn't "relate to" VC0, VC1, or VC2.
[CNG] I agree. If the provider wanted to indicate that AC3 could also be 
used for VC0,VC1,VC2 it could advertise:
      AC0 (VC0)
      AC1 (VC2)
      AC2 (VC1)
      AC3 (VC0,VC1,VC2,VC3,VC4,VC5)
>
> [Duckworth, Mark] For a case with a different consumer, suppose the consumer wants the single video capture VC5, but it would like to receive and render spatial audio.  There is no audio capture with multiple channels, so it could ask for AC0, AC1, and AC2 if it knew how they were spatially "related to" VC5.  But according to the advertisement, AC0, AC1 and AC2 aren't "related to" VC5 at all.
[CNG] So in this case there is no AC3? In that case the advertisement 
could look like:
      AC0 (VC0,VC5)
      AC1 (VC2,VC5)
      AC2 (VC1,VC5)

Each audio capture could relate to a particular video or be part of the 
zoomed out video capture if there was no room audio.
>
>> Regards, Christian
>>
>> On 14/04/2014 12:26 PM, Duckworth, Mark wrote:
>>> Hello Christian,
>>> please see below.
>>> Mark
>>>
>>>> -----Original Message-----
>>>> From: clue [mailto:clue-bounces@ietf.org] On Behalf Of Christian
>>>> Groves
>>>> Sent: Sunday, April 13, 2014 10:08 PM
>>>> To: clue@ietf.org
>>>> Subject: Re: [clue] Improving treatment of audio
>>>>
>>>> Hello all,
>>>>
>>>> Perhaps the confusion is that some see Audio Capture area as a means
>>>> to associate an Audio capture with a video capture as you've shown
>>>> Mark's example. e.g. The audio spatial information isn't really used
>>>> for any audio transformation other than associating it with a particular
>> video stream.
>>>> Whereas others were more thinking the audio spatial information as an
>>>> input to mixing and more complicated audio processing.
>>> [Duckworth, Mark] You could be right about this being a source of
>> confusion.
>>>> If this is the case perhaps rather than linking ACs and VCs through
>>>> the physical or virtual co-ordinates of the area of capture
>>>> information, we simplify things and in each audio capture we say what VC
>> it relates to?
>>>> e.g.using Mark's example AC0(VC0), AC3(V0,V1,V2)
>>> [Duckworth, Mark] I'm not sure how this would really work.  Because even
>> in this simple example, we also have AC0 relates to VC3 (sometimes), VC4,
>> and VC5 (at least part of it), and so on.  AC3 relates to all of VC1 through VC5.
>> I think using area of capture works better than trying to do something like
>> this.
> ...snip...
>


From nobody Mon Apr 28 21:10:28 2014
Return-Path: <pkyzivat@alum.mit.edu>
X-Original-To: clue@ietfa.amsl.com
Delivered-To: clue@ietfa.amsl.com
Received: from localhost (ietfa.amsl.com [127.0.0.1]) by ietfa.amsl.com (Postfix) with ESMTP id B0FC41A085A for <clue@ietfa.amsl.com>; Mon, 28 Apr 2014 21:10:23 -0700 (PDT)
X-Virus-Scanned: amavisd-new at amsl.com
X-Spam-Flag: NO
X-Spam-Score: -1.235
X-Spam-Level: 
X-Spam-Status: No, score=-1.235 tagged_above=-999 required=5 tests=[BAYES_00=-1.9, DKIM_SIGNED=0.1, DKIM_VALID=-0.1, SPF_SOFTFAIL=0.665] autolearn=no
Received: from mail.ietf.org ([4.31.198.44]) by localhost (ietfa.amsl.com [127.0.0.1]) (amavisd-new, port 10024) with ESMTP id yo3y1b2Ipi4S for <clue@ietfa.amsl.com>; Mon, 28 Apr 2014 21:10:22 -0700 (PDT)
Received: from qmta02.westchester.pa.mail.comcast.net (qmta02.westchester.pa.mail.comcast.net [IPv6:2001:558:fe14:43:76:96:62:24]) by ietfa.amsl.com (Postfix) with ESMTP id 2478D1A084F for <clue@ietf.org>; Mon, 28 Apr 2014 21:10:22 -0700 (PDT)
Received: from omta21.westchester.pa.mail.comcast.net ([76.96.62.72]) by qmta02.westchester.pa.mail.comcast.net with comcast id vg8K1n0031ZXKqc51gAMg0; Tue, 29 Apr 2014 04:10:21 +0000
Received: from Paul-Kyzivats-MacBook-Pro.local ([50.138.229.164]) by omta21.westchester.pa.mail.comcast.net with comcast id vgAL1n00e3ZTu2S3hgALNN; Tue, 29 Apr 2014 04:10:21 +0000
Message-ID: <535F262C.6090009@alum.mit.edu>
Date: Tue, 29 Apr 2014 00:10:20 -0400
From: Paul Kyzivat <pkyzivat@alum.mit.edu>
User-Agent: Mozilla/5.0 (Macintosh; Intel Mac OS X 10.7; rv:24.0) Gecko/20100101 Thunderbird/24.4.0
MIME-Version: 1.0
To: clue@ietf.org
References: <533AF351.9050201@alum.mit.edu> <20140410222705.GW39240@verdi> <BLU0-SMTP6910DA20C9A9287137B585D0550@phx.gbl> <53475572.4040807@nteczone.com> <20140411150055.GE60844@verdi> <49E45C59CA48264997FEBFB29B6BC2D6215295A8AE@CRPMBOXPRD07.polycom.com> <BLU0-SMTP1804A70E3E0F666EAEAE84D0540@phx.gbl> <49E45C59CA48264997FEBFB29B6BC2D62409219636@CRPMBOXPRD07.polycom.com> <534B4303.7060707@nteczone.com> <5C4AC54BFF7A0842A6A11F554D6FB52F07152B@CRPMBOXPRD08.polycom.com> <534B536B.8000205@nteczone.com> <5C4AC54BFF7A0842A6A11F554D6FB52F13D9B2@CRPMBOXPRD08.polycom.com> <535F0B3A.70808@nteczone.com>
In-Reply-To: <535F0B3A.70808@nteczone.com>
Content-Type: text/plain; charset=ISO-8859-1; format=flowed
Content-Transfer-Encoding: 7bit
DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=comcast.net; s=q20140121; t=1398744621; bh=TkKZgNA5l8Wx461W1G4BhPxtmvPhDT5PEj72yBPoDio=; h=Received:Received:Message-ID:Date:From:MIME-Version:To:Subject: Content-Type; b=VkJ3TJdByQBVFE6jMw8nVCxp9J0vo+oYl4aMfFJ1/dZkEUC+U0sbjo14w/1laXmVa RxPuPIK4tIMxhsTCfQUaGAnO/mxdPnPgKPFo0yWGhtHmOlYbUYGawFhTBL8A1fV+AE LdojY4smQ1gKfsjy8ZhfKf87xIopP6nJ5uH8U+aYvYAHt7/tnMrnusia6MKaX4HaiZ /4zC3p88+JsAU0in2fHiL2rsbBFvST48OTYvVPhuVXAlzFMD1XkYTGdQR5FO5L0G2a VB81JD+x2tdbcjJqCtUS3KXAHPHTulhETOHgGHVGevxc/DUffsP3VDIPQAMqF4J9bV HwfZHUMRB25Jg==
Archived-At: http://mailarchive.ietf.org/arch/msg/clue/JQt07OuHq3DuXUwZ_C2zonuWCEg
Subject: Re: [clue] Improving treatment of audio
X-BeenThere: clue@ietf.org
X-Mailman-Version: 2.1.15
Precedence: list
List-Id: CLUE - ControLling mUltiple streams for TElepresence <clue.ietf.org>
List-Unsubscribe: <https://www.ietf.org/mailman/options/clue>, <mailto:clue-request@ietf.org?subject=unsubscribe>
List-Archive: <http://www.ietf.org/mail-archive/web/clue/>
List-Post: <mailto:clue@ietf.org>
List-Help: <mailto:clue-request@ietf.org?subject=help>
List-Subscribe: <https://www.ietf.org/mailman/listinfo/clue>, <mailto:clue-request@ietf.org?subject=subscribe>
X-List-Received-Date: Tue, 29 Apr 2014 04:10:23 -0000

On 4/28/14 10:15 PM, Christian Groves wrote:
> Hello Mark,
>
> Please see below.
>
> Regards, Christian
>
> On 29/04/2014 11:36 AM, Duckworth, Mark wrote:
>> Hi Christian,
>> I'm coming back to this message, because I'm trying to understand your
>> proposal.  You proposed "in each audio capture we say what VC it
>> relates to."  What does this mean?  What does it mean for an audio
>> capture to relate to a video capture?  What exactly is the producer
>> telling the consumer?  Can an audio capture relate to more than one
>> video capture?
> [CNG] The meaning has largely the same semantic as if the ADV used the
> same capture area on a VC and AC. Basically its saying that the audio
> comes from the same "region" as what the video capture does. I think the
> difference is that its not trying to put a mathematical certainty about
> what that region is. Its a way for the producer to say to the consumer
> if you choose video capture A you probably want to choose audio capture
> B for a good experience. In terms of whether an audio capture can relate
> to more than one video capture yes I think so. You could have one audio
> capture for a room and three video captures.

Do you think this can be defined well enough that two people, given the 
same configuration of equipment, would reach the same conclusion about 
which audios to associate with which videos?

	Thanks,
	Paul

>> More comments inline below.
>>
>> Regards,
>> Mark
>>
>>> -----Original Message-----
>>> From: Christian Groves [mailto:Christian.Groves@nteczone.com]
>>> Sent: Sunday, April 13, 2014 11:18 PM
>>> To: Duckworth, Mark; clue@ietf.org
>>> Subject: Re: [clue] Improving treatment of audio
>>>
>>> Hello Mark,
>>>
>>> Sorry I didn't consider the entire 12.1.1. if I do that, according to
>>> that
>>> example:
>>>
>>> Video areas of capture:
>>>
>>>          bottom left    bottom right  top left         top right
>>>      VC0 (-2011,2850,0) (-673,3000,0) (-2011,2850,757) (-673,3000,757)
>>>      VC1 ( -673,3000,0) ( 673,3000,0) ( -673,3000,757) ( 673,3000,757)
>>>      VC2 (  673,3000,0) (2011,2850,0) (  673,3000,757) (2011,3000,757)
>>>      VC3 (-2011,2850,0) (2011,2850,0) (-2011,2850,757) (2011,3000,757)
>>>      VC4 (-2011,2850,0) (2011,2850,0) (-2011,2850,757) (2011,3000,757)
>>>      VC5 (-2011,2850,0) (2011,2850,0) (-2011,2850,757) (2011,3000,757)
>>>      VC6 none
>>>
>>> Areas of capture for audio (from 12.1.1):
>>>
>>>          bottom left    bottom right  top left         top right
>>>
>>>      AC0 (-2011,2850,0) (-673,3000,0) (-2011,2850,757) (-673,3000,757)
>>>      AC1 (  673,3000,0) (2011,2850,0) (  673,3000,757) (2011,3000,757)
>>>      AC2 ( -673,3000,0) ( 673,3000,0) ( -673,3000,757) ( 673,3000,757)
>>>      AC3 (-2011,2850,0) (2011,2850,0) (-2011,2850,757) (2011,3000,757)
>>>      AC4 none
>>>
>>> Using a reference rather than area of capture:
>>>       AC0 (VC0)
>>>       AC1 (VC2)
>>>       AC2 (VC1)
>>>       AC3 (VC3,VC4,VC5)
>>> I would think they are basically conveying the same information???
>> [Duckworth, Mark] Looking from the consumer point of view, in trying
>> to understand the advertisement, what does the consumer know from the
>> advertisement?  Suppose the consumer wants to get all three of VC0,
>> VC1, VC2, but it only wants a single audio capture.  It should
>> probably ask for AC3, but according to the advertisement AC3 doesn't
>> "relate to" VC0, VC1, or VC2.
> [CNG] I agree. If the provider wanted to indicate that AC3 could also be
> used for VC0,VC1,VC2 it could advertise:
>       AC0 (VC0)
>       AC1 (VC2)
>       AC2 (VC1)
>       AC3 (VC0,VC1,VC2,VC3,VC4,VC5)
>>
>> [Duckworth, Mark] For a case with a different consumer, suppose the
>> consumer wants the single video capture VC5, but it would like to
>> receive and render spatial audio.  There is no audio capture with
>> multiple channels, so it could ask for AC0, AC1, and AC2 if it knew
>> how they were spatially "related to" VC5.  But according to the
>> advertisement, AC0, AC1 and AC2 aren't "related to" VC5 at all.
> [CNG] So in this case there is no AC3? In that case the advertisement
> could look like:
>       AC0 (VC0,VC5)
>       AC1 (VC2,VC5)
>       AC2 (VC1,VC5)
>
> Each audio capture could relate to a particular video or be part of the
> zoomed out video capture if there was no room audio.
>>
>>> Regards, Christian
>>>
>>> On 14/04/2014 12:26 PM, Duckworth, Mark wrote:
>>>> Hello Christian,
>>>> please see below.
>>>> Mark
>>>>
>>>>> -----Original Message-----
>>>>> From: clue [mailto:clue-bounces@ietf.org] On Behalf Of Christian
>>>>> Groves
>>>>> Sent: Sunday, April 13, 2014 10:08 PM
>>>>> To: clue@ietf.org
>>>>> Subject: Re: [clue] Improving treatment of audio
>>>>>
>>>>> Hello all,
>>>>>
>>>>> Perhaps the confusion is that some see Audio Capture area as a means
>>>>> to associate an Audio capture with a video capture as you've shown
>>>>> Mark's example. e.g. The audio spatial information isn't really used
>>>>> for any audio transformation other than associating it with a
>>>>> particular
>>> video stream.
>>>>> Whereas others were more thinking the audio spatial information as an
>>>>> input to mixing and more complicated audio processing.
>>>> [Duckworth, Mark] You could be right about this being a source of
>>> confusion.
>>>>> If this is the case perhaps rather than linking ACs and VCs through
>>>>> the physical or virtual co-ordinates of the area of capture
>>>>> information, we simplify things and in each audio capture we say
>>>>> what VC
>>> it relates to?
>>>>> e.g.using Mark's example AC0(VC0), AC3(V0,V1,V2)
>>>> [Duckworth, Mark] I'm not sure how this would really work.  Because
>>>> even
>>> in this simple example, we also have AC0 relates to VC3 (sometimes),
>>> VC4,
>>> and VC5 (at least part of it), and so on.  AC3 relates to all of VC1
>>> through VC5.
>>> I think using area of capture works better than trying to do
>>> something like
>>> this.
>> ...snip...
>>
>
> _______________________________________________
> clue mailing list
> clue@ietf.org
> https://www.ietf.org/mailman/listinfo/clue
>


From nobody Tue Apr 29 02:41:20 2014
Return-Path: <Christian.Groves@nteczone.com>
X-Original-To: clue@ietfa.amsl.com
Delivered-To: clue@ietfa.amsl.com
Received: from localhost (ietfa.amsl.com [127.0.0.1]) by ietfa.amsl.com (Postfix) with ESMTP id 5DE2E1A0707 for <clue@ietfa.amsl.com>; Tue, 29 Apr 2014 02:41:18 -0700 (PDT)
X-Virus-Scanned: amavisd-new at amsl.com
X-Spam-Flag: NO
X-Spam-Score: -1.9
X-Spam-Level: 
X-Spam-Status: No, score=-1.9 tagged_above=-999 required=5 tests=[BAYES_00=-1.9] autolearn=ham
Received: from mail.ietf.org ([4.31.198.44]) by localhost (ietfa.amsl.com [127.0.0.1]) (amavisd-new, port 10024) with ESMTP id KmshPTNJVCDg for <clue@ietfa.amsl.com>; Tue, 29 Apr 2014 02:41:16 -0700 (PDT)
Received: from cserver5.myshophosting.com (cserver5.myshophosting.com [175.107.161.1]) by ietfa.amsl.com (Postfix) with ESMTP id 92FB71A0118 for <clue@ietf.org>; Tue, 29 Apr 2014 02:41:15 -0700 (PDT)
Received: from ppp118-209-127-77.lns20.mel4.internode.on.net ([118.209.127.77]:49431 helo=[127.0.0.1]) by cserver5.myshophosting.com with esmtpsa (TLSv1:DHE-RSA-AES128-SHA:128) (Exim 4.82) (envelope-from <Christian.Groves@nteczone.com>) id 1Wf4XH-0001F7-Li for clue@ietf.org; Tue, 29 Apr 2014 19:41:11 +1000
Message-ID: <535F73B6.9030106@nteczone.com>
Date: Tue, 29 Apr 2014 19:41:10 +1000
From: Christian Groves <Christian.Groves@nteczone.com>
User-Agent: Mozilla/5.0 (Windows NT 6.1; WOW64; rv:24.0) Gecko/20100101 Thunderbird/24.4.0
MIME-Version: 1.0
To: clue@ietf.org
References: <533AF351.9050201@alum.mit.edu> <20140410222705.GW39240@verdi> <BLU0-SMTP6910DA20C9A9287137B585D0550@phx.gbl> <53475572.4040807@nteczone.com> <20140411150055.GE60844@verdi> <49E45C59CA48264997FEBFB29B6BC2D6215295A8AE@CRPMBOXPRD07.polycom.com> <BLU0-SMTP1804A70E3E0F666EAEAE84D0540@phx.gbl> <49E45C59CA48264997FEBFB29B6BC2D62409219636@CRPMBOXPRD07.polycom.com> <534B4303.7060707@nteczone.com> <5C4AC54BFF7A0842A6A11F554D6FB52F07152B@CRPMBOXPRD08.polycom.com> <534B536B.8000205@nteczone.com> <5C4AC54BFF7A0842A6A11F554D6FB52F13D9B2@CRPMBOXPRD08.polycom.com> <535F0B3A.70808@nteczone.com> <535F262C.6090009@alum.mit.edu>
In-Reply-To: <535F262C.6090009@alum.mit.edu>
Content-Type: text/plain; charset=ISO-8859-1; format=flowed
Content-Transfer-Encoding: 7bit
X-AntiAbuse: This header was added to track abuse, please include it with any abuse report
X-AntiAbuse: Primary Hostname - cserver5.myshophosting.com
X-AntiAbuse: Original Domain - ietf.org
X-AntiAbuse: Originator/Caller UID/GID - [47 12] / [47 12]
X-AntiAbuse: Sender Address Domain - nteczone.com
X-Get-Message-Sender-Via: cserver5.myshophosting.com: authenticated_id: christian.groves@nteczone.com
X-Source: 
X-Source-Args: 
X-Source-Dir: 
Archived-At: http://mailarchive.ietf.org/arch/msg/clue/uWeJrfaYdAhaPoXX_CDA6xzK7PE
Subject: Re: [clue] Improving treatment of audio
X-BeenThere: clue@ietf.org
X-Mailman-Version: 2.1.15
Precedence: list
List-Id: CLUE - ControLling mUltiple streams for TElepresence <clue.ietf.org>
List-Unsubscribe: <https://www.ietf.org/mailman/options/clue>, <mailto:clue-request@ietf.org?subject=unsubscribe>
List-Archive: <http://www.ietf.org/mail-archive/web/clue/>
List-Post: <mailto:clue@ietf.org>
List-Help: <mailto:clue-request@ietf.org?subject=help>
List-Subscribe: <https://www.ietf.org/mailman/listinfo/clue>, <mailto:clue-request@ietf.org?subject=subscribe>
X-List-Received-Date: Tue, 29 Apr 2014 09:41:18 -0000

On 29/04/2014 2:10 PM, Paul Kyzivat wrote:
> On 4/28/14 10:15 PM, Christian Groves wrote:
>> Hello Mark,
>>
>> Please see below.
>>
>> Regards, Christian
>>
>> On 29/04/2014 11:36 AM, Duckworth, Mark wrote:
>>> Hi Christian,
>>> I'm coming back to this message, because I'm trying to understand your
>>> proposal.  You proposed "in each audio capture we say what VC it
>>> relates to."  What does this mean?  What does it mean for an audio
>>> capture to relate to a video capture?  What exactly is the producer
>>> telling the consumer?  Can an audio capture relate to more than one
>>> video capture?
>> [CNG] The meaning has largely the same semantic as if the ADV used the
>> same capture area on a VC and AC. Basically its saying that the audio
>> comes from the same "region" as what the video capture does. I think the
>> difference is that its not trying to put a mathematical certainty about
>> what that region is. Its a way for the producer to say to the consumer
>> if you choose video capture A you probably want to choose audio capture
>> B for a good experience. In terms of whether an audio capture can relate
>> to more than one video capture yes I think so. You could have one audio
>> capture for a room and three video captures.
>
> Do you think this can be defined well enough that two people, given 
> the same configuration of equipment, would reach the same conclusion 
> about which audios to associate with which videos?

I can't see why not? CLUE already suggests relationships between 
captures i.e.through CSE and GCSEs. You can't assume that the spatial 
co-ordinates will give the relationships in these cases as the spatial 
attributes aren't mandatory.

Christian

>
>     Thanks,
>     Paul
>
>>> More comments inline below.
>>>
>>> Regards,
>>> Mark
>>>
>>>> -----Original Message-----
>>>> From: Christian Groves [mailto:Christian.Groves@nteczone.com]
>>>> Sent: Sunday, April 13, 2014 11:18 PM
>>>> To: Duckworth, Mark; clue@ietf.org
>>>> Subject: Re: [clue] Improving treatment of audio
>>>>
>>>> Hello Mark,
>>>>
>>>> Sorry I didn't consider the entire 12.1.1. if I do that, according to
>>>> that
>>>> example:
>>>>
>>>> Video areas of capture:
>>>>
>>>>          bottom left    bottom right  top left         top right
>>>>      VC0 (-2011,2850,0) (-673,3000,0) (-2011,2850,757) (-673,3000,757)
>>>>      VC1 ( -673,3000,0) ( 673,3000,0) ( -673,3000,757) ( 673,3000,757)
>>>>      VC2 (  673,3000,0) (2011,2850,0) (  673,3000,757) (2011,3000,757)
>>>>      VC3 (-2011,2850,0) (2011,2850,0) (-2011,2850,757) (2011,3000,757)
>>>>      VC4 (-2011,2850,0) (2011,2850,0) (-2011,2850,757) (2011,3000,757)
>>>>      VC5 (-2011,2850,0) (2011,2850,0) (-2011,2850,757) (2011,3000,757)
>>>>      VC6 none
>>>>
>>>> Areas of capture for audio (from 12.1.1):
>>>>
>>>>          bottom left    bottom right  top left         top right
>>>>
>>>>      AC0 (-2011,2850,0) (-673,3000,0) (-2011,2850,757) (-673,3000,757)
>>>>      AC1 (  673,3000,0) (2011,2850,0) (  673,3000,757) (2011,3000,757)
>>>>      AC2 ( -673,3000,0) ( 673,3000,0) ( -673,3000,757) ( 673,3000,757)
>>>>      AC3 (-2011,2850,0) (2011,2850,0) (-2011,2850,757) (2011,3000,757)
>>>>      AC4 none
>>>>
>>>> Using a reference rather than area of capture:
>>>>       AC0 (VC0)
>>>>       AC1 (VC2)
>>>>       AC2 (VC1)
>>>>       AC3 (VC3,VC4,VC5)
>>>> I would think they are basically conveying the same information???
>>> [Duckworth, Mark] Looking from the consumer point of view, in trying
>>> to understand the advertisement, what does the consumer know from the
>>> advertisement?  Suppose the consumer wants to get all three of VC0,
>>> VC1, VC2, but it only wants a single audio capture.  It should
>>> probably ask for AC3, but according to the advertisement AC3 doesn't
>>> "relate to" VC0, VC1, or VC2.
>> [CNG] I agree. If the provider wanted to indicate that AC3 could also be
>> used for VC0,VC1,VC2 it could advertise:
>>       AC0 (VC0)
>>       AC1 (VC2)
>>       AC2 (VC1)
>>       AC3 (VC0,VC1,VC2,VC3,VC4,VC5)
>>>
>>> [Duckworth, Mark] For a case with a different consumer, suppose the
>>> consumer wants the single video capture VC5, but it would like to
>>> receive and render spatial audio.  There is no audio capture with
>>> multiple channels, so it could ask for AC0, AC1, and AC2 if it knew
>>> how they were spatially "related to" VC5.  But according to the
>>> advertisement, AC0, AC1 and AC2 aren't "related to" VC5 at all.
>> [CNG] So in this case there is no AC3? In that case the advertisement
>> could look like:
>>       AC0 (VC0,VC5)
>>       AC1 (VC2,VC5)
>>       AC2 (VC1,VC5)
>>
>> Each audio capture could relate to a particular video or be part of the
>> zoomed out video capture if there was no room audio.
>>>
>>>> Regards, Christian
>>>>
>>>> On 14/04/2014 12:26 PM, Duckworth, Mark wrote:
>>>>> Hello Christian,
>>>>> please see below.
>>>>> Mark
>>>>>
>>>>>> -----Original Message-----
>>>>>> From: clue [mailto:clue-bounces@ietf.org] On Behalf Of Christian
>>>>>> Groves
>>>>>> Sent: Sunday, April 13, 2014 10:08 PM
>>>>>> To: clue@ietf.org
>>>>>> Subject: Re: [clue] Improving treatment of audio
>>>>>>
>>>>>> Hello all,
>>>>>>
>>>>>> Perhaps the confusion is that some see Audio Capture area as a means
>>>>>> to associate an Audio capture with a video capture as you've shown
>>>>>> Mark's example. e.g. The audio spatial information isn't really used
>>>>>> for any audio transformation other than associating it with a
>>>>>> particular
>>>> video stream.
>>>>>> Whereas others were more thinking the audio spatial information 
>>>>>> as an
>>>>>> input to mixing and more complicated audio processing.
>>>>> [Duckworth, Mark] You could be right about this being a source of
>>>> confusion.
>>>>>> If this is the case perhaps rather than linking ACs and VCs through
>>>>>> the physical or virtual co-ordinates of the area of capture
>>>>>> information, we simplify things and in each audio capture we say
>>>>>> what VC
>>>> it relates to?
>>>>>> e.g.using Mark's example AC0(VC0), AC3(V0,V1,V2)
>>>>> [Duckworth, Mark] I'm not sure how this would really work.  Because
>>>>> even
>>>> in this simple example, we also have AC0 relates to VC3 (sometimes),
>>>> VC4,
>>>> and VC5 (at least part of it), and so on.  AC3 relates to all of VC1
>>>> through VC5.
>>>> I think using area of capture works better than trying to do
>>>> something like
>>>> this.
>>> ...snip...
>>>
>>
>> _______________________________________________
>> clue mailing list
>> clue@ietf.org
>> https://www.ietf.org/mailman/listinfo/clue
>>
>
> _______________________________________________
> clue mailing list
> clue@ietf.org
> https://www.ietf.org/mailman/listinfo/clue
>


From nobody Tue Apr 29 07:35:28 2014
Return-Path: <stephen.botzko@gmail.com>
X-Original-To: clue@ietfa.amsl.com
Delivered-To: clue@ietfa.amsl.com
Received: from localhost (ietfa.amsl.com [127.0.0.1]) by ietfa.amsl.com (Postfix) with ESMTP id 631211A0905 for <clue@ietfa.amsl.com>; Tue, 29 Apr 2014 07:35:27 -0700 (PDT)
X-Virus-Scanned: amavisd-new at amsl.com
X-Spam-Flag: NO
X-Spam-Score: -1.999
X-Spam-Level: 
X-Spam-Status: No, score=-1.999 tagged_above=-999 required=5 tests=[BAYES_00=-1.9, DKIM_SIGNED=0.1, DKIM_VALID=-0.1, DKIM_VALID_AU=-0.1, FREEMAIL_FROM=0.001, HTML_MESSAGE=0.001, SPF_PASS=-0.001] autolearn=ham
Received: from mail.ietf.org ([4.31.198.44]) by localhost (ietfa.amsl.com [127.0.0.1]) (amavisd-new, port 10024) with ESMTP id sDoGzQeABLE8 for <clue@ietfa.amsl.com>; Tue, 29 Apr 2014 07:35:20 -0700 (PDT)
Received: from mail-ve0-x233.google.com (mail-ve0-x233.google.com [IPv6:2607:f8b0:400c:c01::233]) by ietfa.amsl.com (Postfix) with ESMTP id CDAA01A04AC for <clue@ietf.org>; Tue, 29 Apr 2014 07:35:19 -0700 (PDT)
Received: by mail-ve0-f179.google.com with SMTP id db12so364919veb.38 for <clue@ietf.org>; Tue, 29 Apr 2014 07:35:18 -0700 (PDT)
DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20120113;  h=mime-version:in-reply-to:references:date:message-id:subject:from:to :cc:content-type; bh=IVn/zYPqOK8uc8ZrjLs/Y8hVFjmqeSrkG/QQGGyGzaE=; b=Hcotcn61QqLSnavKb6L6xi1UAqxovMjlggmF8Q4zBdHQkFRPp/NvOOYXIbGyva6hy2 I9LTUpLx+i8U3BuXLu+Iw/hP7DV4Xidwr+Crc+JRFMNf1X6zkG6lbnL6knwbuCxTvyqA oNb5OReasKOUQAUaQTbvlNWG94GPzTAwMuSnp20ODSzgGA+E2I0PASJ34CucWBlThdcS GefkF/ICVpAf0tSHdc7TgUhVU2wlmZHvA7rkXX7euaB4iz8gTpFmK1jPo9lSMo6BRCPj NAVbLpzi/RrYg7b31wwh0xuhu7emZShA83k2QwWXa691QVrfUi9NMGzUi8A55xTrIeOQ fEGg==
MIME-Version: 1.0
X-Received: by 10.52.238.161 with SMTP id vl1mr99071vdc.88.1398782118473; Tue, 29 Apr 2014 07:35:18 -0700 (PDT)
Received: by 10.221.40.135 with HTTP; Tue, 29 Apr 2014 07:35:18 -0700 (PDT)
In-Reply-To: <535F73B6.9030106@nteczone.com>
References: <533AF351.9050201@alum.mit.edu> <20140410222705.GW39240@verdi> <BLU0-SMTP6910DA20C9A9287137B585D0550@phx.gbl> <53475572.4040807@nteczone.com> <20140411150055.GE60844@verdi> <49E45C59CA48264997FEBFB29B6BC2D6215295A8AE@CRPMBOXPRD07.polycom.com> <BLU0-SMTP1804A70E3E0F666EAEAE84D0540@phx.gbl> <49E45C59CA48264997FEBFB29B6BC2D62409219636@CRPMBOXPRD07.polycom.com> <534B4303.7060707@nteczone.com> <5C4AC54BFF7A0842A6A11F554D6FB52F07152B@CRPMBOXPRD08.polycom.com> <534B536B.8000205@nteczone.com> <5C4AC54BFF7A0842A6A11F554D6FB52F13D9B2@CRPMBOXPRD08.polycom.com> <535F0B3A.70808@nteczone.com> <535F262C.6090009@alum.mit.edu> <535F73B6.9030106@nteczone.com>
Date: Tue, 29 Apr 2014 10:35:18 -0400
Message-ID: <CAMC7SJ7ZNW+=PDB-39gQ3MkQpTh5W1LZLPz7Y8VaFc7OofRviw@mail.gmail.com>
From: Stephen Botzko <stephen.botzko@gmail.com>
To: Christian Groves <Christian.Groves@nteczone.com>
Content-Type: multipart/alternative; boundary=001a1134c7844a366b04f82f57de
Archived-At: http://mailarchive.ietf.org/arch/msg/clue/OeQBbk6v3b-s5Mb9EgsZ5lGMK5g
Cc: CLUE <clue@ietf.org>
Subject: Re: [clue] Improving treatment of audio
X-BeenThere: clue@ietf.org
X-Mailman-Version: 2.1.15
Precedence: list
List-Id: CLUE - ControLling mUltiple streams for TElepresence <clue.ietf.org>
List-Unsubscribe: <https://www.ietf.org/mailman/options/clue>, <mailto:clue-request@ietf.org?subject=unsubscribe>
List-Archive: <http://www.ietf.org/mail-archive/web/clue/>
List-Post: <mailto:clue@ietf.org>
List-Help: <mailto:clue-request@ietf.org?subject=help>
List-Subscribe: <https://www.ietf.org/mailman/listinfo/clue>, <mailto:clue-request@ietf.org?subject=subscribe>
X-List-Received-Date: Tue, 29 Apr 2014 14:35:27 -0000

--001a1134c7844a366b04f82f57de
Content-Type: text/plain; charset=UTF-8

This may be a digression, but I think it would be useful to identify
specific scenarios that work (and don't work) before we go in and try and
improve things.  I'm providing my assessment below for comments/feedback.

BR
Stephen


-Room to Room rendering works.

In the normal telepresence point-to-point case, there is a large video wall
in the front, with some number of speakers also in the front.  The size of
the wall, the number of cameras/screens, and the number of loudspeakers is
mismatched.

On the video side, there is enough spatial information to allow the
receivers to render the far end room "properly", and several strategies
that receivers can use.  Once the receiver knows the mapping between its
local coordinates and the far end, it can then make corresponding
adjustments to the received audio.  For instance, for each received audio
channel, it can identify the nearest two loudspeakers, and can use normal
panning to create a sound source at a point in between.  The spatial audio
is "soft" - similar to the results you get when you derive a center channel
from two stereo speakers.  A center speaker in the right place certainly
sounds more realistic, but using two stereo speakers to simulate it works
acceptably.

For users, even the "soft" spatial audio improves the user experience.  For
instance, when someone starts talking, even the "soft" spatial audio is
enough to allow the participants to automatically look in the right
direction when someone new starts talking.

If you wanted to use rear speakers to place some sounds behind the local
audience, that would not work.  But I don't think that's a very important
case.  It is probably more common to use any rear speakers for normal sound
reinforcement (exploiting the Haas effect).

So I believe this scenario is already enabled.  Industry experience with
interoperability today (using TIP) pretty much supports that view.

Could we do better?  In theory some approaches (MPEG SAOC) might be able to
create a better spatial sound. Although SAOC proponents sometimes refer to
interactive conferencing as an application, I don't think it is
well-enabled in practice.  Locating and isolating each sound source in the
room in real time is not an easy task, and SAOC gives no hints on how to do
it.  SAOC's real application is gaming, where pre-recorded and synthetic
sounds are being blended into a virtual reality.  Also, SAOC uses in-band
transmission of audio object information, so it could be used with CLUE in
a future system.  There would likely need to be some form of tagging the
objects, so that the video rendering could also take advantage of them. But
that could be added in the future.  SAOC is likely encumbered (just a guess
I haven't checked).

So at some point we could perhaps do better, but it is still a bit of a
research project.  I think the current functionality is enough.



-Whole Room Composition works

Similarly, in a multipoint case you frequently end up giving up on full
size rendering and eye-contact because you simply don't have the screen
real estate.  In those cases, you are scaling the room video down (usually
to something a lot smaller) and placing it somewhere in the composition.

On the audio side, the panning techniques in the full room-to-room case
still apply, and give a satisfactory (though soft") spatial audio
experience.  Sounds that are off-camera in the far-end room might be placed
in an adjacent room in the composition.  However the existing spatial
system allows receivers to detect that the far-end sound stage is wider
than the video scene.  And there is a remedy - receivers can also choose to
down-mix the far-end room audio to a single mono capture, and align that
with the video.  In many cases, that works just as well (since the room
view on the screen is relatively small), and it ensures that all the audio
from the room is aligned with the video composition.

So I also think the current functionality is enough.

-Rendering in a deliberately different spatial alignment does not work.

One could chose to totally ignore the spatial information for video.  For
instance, you might wish to present independent video tiles for the various
talkers, putting the current talker in the center, and other recent talkers
around the side - totally ignoring the spatial information and what room
the talker is in.

You can do this with video.  However, re-processing the sound from the far
end rooms in this case doesn't work well.  The spatial sound would not be
coherent or natural, because each audio capture includes some sound from
the other parts of the room, and in general there is no way to remove that
sound.  So in this scenario, receivers probably need to fall back to
monophonic audio (rendering all the audio from all senders monophonically).

It would be nice to do better here.  But I think the problem is really in
the audio capture itself.  Changing the signaling won't fix it.  So the
current functionality is not enough for this case, but I think for now
commercial systems can't do better than monophonic rendering.

-Talker identification does not work.

It is common in videoconferencing to use "voice-activated switching".
That works by detecting speech in the far-end audio channel, and then
automatically switching to the corresponding video channel.

This is a problem area for CLUE (and it is also a problem area I've seen
with systems using TIP).

Part of the problem grounded in the audio capture itself.  A talker in the
room is to some extent picked up by all the microphones.  You might think
it is easy to identify the right capture from the volume level, but in many
microphone arrangements that doesn't work reliably.  Participants sitting
between microphones are hard to locate.  Participants also turn towards
other people in the local room when they talk, which often changes the
microphone that picks them up best.  And often there are multiple
simultaneous talkers in the room - we are not always polite.  So even
detecting the approximate talker location(s) from the far end is
problematic.  My company's products can locate the talker(s) within the
room accurately, but they are using information that is not available to
far-end systems, and which would be hard to send in a product-independent
way.

Personally I favor explicit signaling from the sender on which capture(s)
carry the current talker(s).  Middle boxes and far end systems can
determine which rooms have active talkers by down-mixing to mono.  To get
finer granularity within the room, you could then use the explicit
signaling.

The second aspect of the problem is associating the talker with a video
capture.  Again, this seems somewhat difficult, especially if the audio
captures are not precisely aligned to the video captures. We can probably
improve this - but if we need explicit talker signaling anyway, it could
potentially identify the best video capture. In fact, since I personally
think that whole-room audio rendering should always be done, I am much more
interested in the video capture than the audio capture.

Other approaches might work, the main point I am making here is that remote
identification of talker locations in the room is probably broken.

-Partial Room Rendering will not always work well.

If you are selecting some audio captures from a room, and discarding
others, you might not get natural sounding audio.  How good it sounds will
depend on the details of the microphone pickup in the far end room, how
controlled participant seating is, and which captures you happen to discard.

Since the goal for CLUE is interoperability across disparate room designs,
we can't make a lot of assumptions about microphone pickup.  I think we can
assume that the full room audio will sound appropriate - if it doesn't,
then that means that the sender doesn't have a well-designed system.  But
once you start filtering out captures in the middle, the results will vary.
 I'm not seeing much CLUE can do about that, other than point it out.  As
noted above, I think that rendering the entire room's transmitted sound
field is going to sound the best.  If there is a need for more conditioning
(say noise reduction) it is probably better to do it locally at the sender.

--001a1134c7844a366b04f82f57de
Content-Type: text/html; charset=UTF-8
Content-Transfer-Encoding: quoted-printable

<div dir=3D"ltr"><div class=3D"gmail_extra">This may be a digression, but I=
 think it would be useful to identify specific scenarios that work (and don=
&#39;t work) before we go in and try and improve things. =C2=A0I&#39;m prov=
iding my assessment below for comments/feedback.</div>
<div class=3D"gmail_extra"><br></div><div class=3D"gmail_extra">BR</div><di=
v class=3D"gmail_extra">Stephen</div>
<div class=3D"gmail_extra"><br></div><div class=3D"gmail_extra"><br></div><=
div class=3D"gmail_extra">-Room to Room rendering works.</div><blockquote s=
tyle=3D"margin:0px 0px 0px 40px;border:none;padding:0px"><div class=3D"gmai=
l_extra">

In the normal telepresence point-to-point case, there is a large video wall=
 in the front, with some number of speakers also in the front. =C2=A0The si=
ze of the wall, the number of cameras/screens, and the number of loudspeake=
rs is mismatched.</div>

<div class=3D"gmail_extra"><br></div><div class=3D"gmail_extra">On the vide=
o side, there is enough spatial information to allow the receivers to rende=
r the far end room &quot;properly&quot;, and several strategies that receiv=
ers can use. =C2=A0Once the receiver knows the mapping between its local co=
ordinates and the far end, it can then make corresponding adjustments to th=
e received audio. =C2=A0For instance, for each received audio channel, it c=
an identify the nearest two loudspeakers, and can use normal panning to cre=
ate a sound source at a point in between. =C2=A0The spatial audio is &quot;=
soft&quot; - similar to the results you get when you derive a center channe=
l from two stereo speakers. =C2=A0A center speaker in the right place certa=
inly sounds more realistic, but using two stereo speakers to simulate it wo=
rks acceptably.</div>

<div class=3D"gmail_extra"><br></div><div class=3D"gmail_extra">For users, =
even the &quot;soft&quot; spatial audio improves the user experience. =C2=
=A0For instance, when someone starts talking, even the &quot;soft&quot; spa=
tial audio is enough to allow the participants to automatically look in the=
 right direction when someone new starts talking.</div>

<div class=3D"gmail_extra"><br></div><div class=3D"gmail_extra">If you want=
ed to use rear speakers to place some sounds behind the local audience, tha=
t would not work. =C2=A0But I don&#39;t think that&#39;s a very important c=
ase. =C2=A0It is probably more common to use any rear speakers for normal s=
ound reinforcement (exploiting the Haas effect).</div>

<div class=3D"gmail_extra"><br></div><div class=3D"gmail_extra">So I believ=
e this scenario is already enabled. =C2=A0Industry experience with interope=
rability today (using TIP) pretty much supports that view.</div><div class=
=3D"gmail_extra">

<br></div></blockquote><blockquote style=3D"margin:0px 0px 0px 40px;border:=
none;padding:0px"><div class=3D"gmail_extra">Could we do better? =C2=A0In t=
heory some approaches (MPEG SAOC) might be able to create a better spatial =
sound. Although SAOC proponents sometimes refer to interactive conferencing=
 as an application, I don&#39;t think it is well-enabled in practice. =C2=
=A0Locating and isolating each sound source in the room in real time is not=
 an easy task, and SAOC gives no hints on how to do it. =C2=A0SAOC&#39;s re=
al application is gaming, where pre-recorded and synthetic sounds are being=
 blended into a virtual reality. =C2=A0Also, SAOC uses in-band transmission=
 of audio object information, so it could be used with CLUE in a future sys=
tem. =C2=A0There would likely need to be some form of tagging the objects, =
so that the video rendering could also take advantage of them. But that cou=
ld be added in the future. =C2=A0SAOC is likely encumbered (just a guess I =
haven&#39;t checked).</div>

<div class=3D"gmail_extra"><br></div></blockquote><blockquote style=3D"marg=
in:0px 0px 0px 40px;border:none;padding:0px"><div class=3D"gmail_extra">So =
at some point we could perhaps do better, but it is still a bit of a resear=
ch project. =C2=A0I think the current functionality is enough.</div>

</blockquote><blockquote style=3D"margin:0px 0px 0px 40px;border:none;paddi=
ng:0px"><div class=3D"gmail_extra"><br></div><div class=3D"gmail_extra"><br=
></div></blockquote><div class=3D"gmail_extra">-Whole Room Composition work=
s</div>

<blockquote style=3D"margin:0px 0px 0px 40px;border:none;padding:0px"><div =
class=3D"gmail_extra">Similarly, in a multipoint case you frequently end up=
 giving up on full size rendering and eye-contact because you simply don&#3=
9;t have the screen real estate. =C2=A0In those cases, you are scaling the =
room video down (usually to something a lot smaller) and placing it somewhe=
re in the composition.</div>

<div class=3D"gmail_extra"><br></div><div class=3D"gmail_extra">On the audi=
o side, the panning techniques in the full room-to-room case still apply, a=
nd give a satisfactory (though soft&quot;) spatial audio experience. =C2=A0=
Sounds that are off-camera in the far-end room might be placed in an adjace=
nt room in the composition. =C2=A0However the existing spatial system allow=
s receivers to detect that the far-end sound stage is wider than the video =
scene. =C2=A0And there is a remedy - receivers can also choose to down-mix =
the far-end room audio to a single mono capture, and align that with the vi=
deo. =C2=A0In many cases, that works just as well (since the room view on t=
he screen is relatively small), and it ensures that all the audio from the =
room is aligned with the video composition.</div>

<div class=3D"gmail_extra"><br></div><div class=3D"gmail_extra">So I also t=
hink the current functionality is enough. =C2=A0</div><div class=3D"gmail_e=
xtra"><br></div></blockquote><div class=3D"gmail_extra">-Rendering in a del=
iberately different spatial alignment does not work.</div>

<blockquote style=3D"margin:0px 0px 0px 40px;border:none;padding:0px"><div =
class=3D"gmail_extra">One could chose to totally ignore the spatial informa=
tion for video. =C2=A0For instance, you might wish to present independent v=
ideo tiles for the various talkers, putting the current talker in the cente=
r, and other recent talkers around the side - totally ignoring the spatial =
information and what room the talker is in.</div>

<div class=3D"gmail_extra"><br></div><div class=3D"gmail_extra">You can do =
this with video. =C2=A0However, re-processing the sound from the far end ro=
oms in this case doesn&#39;t work well. =C2=A0The spatial sound would not b=
e coherent or natural, because each audio capture includes some sound from =
the other parts of the room, and in general there is no way to remove that =
sound. =C2=A0So in this scenario, receivers probably need to fall back to m=
onophonic audio (rendering all the audio from all senders monophonically).<=
/div>

<div class=3D"gmail_extra"><br></div><div class=3D"gmail_extra">It would be=
 nice to do better here. =C2=A0But I think the problem is really in the aud=
io capture itself. =C2=A0Changing the signaling won&#39;t fix it. =C2=A0So =
the current functionality is not enough for this case, but I think for now =
commercial systems can&#39;t do better than monophonic rendering.</div>

<div class=3D"gmail_extra"><br></div></blockquote>-Talker identification do=
es not work.<blockquote style=3D"margin:0px 0px 0px 40px;border:none;paddin=
g:0px"><div>It is common in videoconferencing to use &quot;voice-activated =
switching&quot;. =C2=A0 That works by detecting speech in the far-end audio=
 channel, and then automatically switching to the corresponding video chann=
el.</div>
<div><br></div><div>This is a problem area for CLUE (and it is also a probl=
em area I&#39;ve seen with systems using TIP).</div><div><br></div><div>Par=
t of the problem grounded in the audio capture itself. =C2=A0A talker in th=
e room is to some extent picked up by all the microphones. =C2=A0You might =
think it is easy to identify the right capture from the volume level, but i=
n many microphone arrangements that doesn&#39;t work reliably. =C2=A0Partic=
ipants sitting between microphones are hard to locate. =C2=A0Participants a=
lso turn towards other people in the local room when they talk, which often=
 changes the microphone that picks them up best. =C2=A0And often there are =
multiple simultaneous talkers in the room - we are not always polite. =C2=
=A0So even detecting the approximate talker location(s) from the far end is=
 problematic. =C2=A0My company&#39;s products can locate the talker(s) with=
in the room accurately, but they are using information that is not availabl=
e to far-end systems, and which would be hard to send in a product-independ=
ent way.</div>
<div><br></div><div>Personally I favor explicit signaling from the sender o=
n which capture(s) carry the current talker(s). =C2=A0Middle boxes and far =
end systems can determine which rooms have active talkers by down-mixing to=
 mono. =C2=A0To get finer granularity within the room, you could then use t=
he explicit signaling. =C2=A0</div>
<div><br></div><div>The second aspect of the problem is associating the tal=
ker with a video capture. =C2=A0Again, this seems somewhat difficult, espec=
ially if the audio captures are not precisely aligned to the video captures=
. We can probably improve this - but if we need explicit talker signaling a=
nyway, it could potentially identify the best video capture. In fact, since=
 I personally think that whole-room audio rendering should always be done, =
I am much more interested in the video capture than the audio capture.</div=
>
<div><br></div><div>Other approaches might work, the main point I am making=
 here is that remote identification of talker locations in the room is prob=
ably broken.</div><div><br></div></blockquote>-Partial Room Rendering will =
not always work well.<blockquote style=3D"margin:0 0 0 40px;border:none;pad=
ding:0px">
<div>If you are selecting some audio captures from a room, and discarding o=
thers, you might not get natural sounding audio. =C2=A0How good it sounds w=
ill depend on the details of the microphone pickup in the far end room, how=
 controlled participant seating is, and which captures you happen to discar=
d.</div>
<div><br></div><div>Since the goal for CLUE is interoperability across disp=
arate room designs, we can&#39;t make a lot of assumptions about microphone=
 pickup. =C2=A0I think we can assume that the full room audio will sound ap=
propriate - if it doesn&#39;t, then that means that the sender doesn&#39;t =
have a well-designed system. =C2=A0But once you start filtering out capture=
s in the middle, the results will vary. =C2=A0I&#39;m not seeing much CLUE =
can do about that, other than point it out. =C2=A0As noted above, I think t=
hat rendering the entire room&#39;s transmitted sound field is going to sou=
nd the best. =C2=A0If there is a need for more conditioning (say noise redu=
ction) it is probably better to do it locally at the sender.</div>
</blockquote><div><div><br></div><div><br>

<div><br></div><div><br></div><div><div class=3D"gmail_extra"><br></div></d=
iv></div></div></div>

--001a1134c7844a366b04f82f57de--


From nobody Tue Apr 29 09:49:50 2014
Return-Path: <pkyzivat@alum.mit.edu>
X-Original-To: clue@ietfa.amsl.com
Delivered-To: clue@ietfa.amsl.com
Received: from localhost (ietfa.amsl.com [127.0.0.1]) by ietfa.amsl.com (Postfix) with ESMTP id 7B8B31A08CA for <clue@ietfa.amsl.com>; Tue, 29 Apr 2014 09:49:49 -0700 (PDT)
X-Virus-Scanned: amavisd-new at amsl.com
X-Spam-Flag: NO
X-Spam-Score: -1.235
X-Spam-Level: 
X-Spam-Status: No, score=-1.235 tagged_above=-999 required=5 tests=[BAYES_00=-1.9, DKIM_SIGNED=0.1, DKIM_VALID=-0.1, SPF_SOFTFAIL=0.665] autolearn=no
Received: from mail.ietf.org ([4.31.198.44]) by localhost (ietfa.amsl.com [127.0.0.1]) (amavisd-new, port 10024) with ESMTP id RtTwKcLsWgoN for <clue@ietfa.amsl.com>; Tue, 29 Apr 2014 09:49:48 -0700 (PDT)
Received: from qmta14.westchester.pa.mail.comcast.net (qmta14.westchester.pa.mail.comcast.net [IPv6:2001:558:fe14:44:76:96:59:212]) by ietfa.amsl.com (Postfix) with ESMTP id C598D1A0939 for <clue@ietf.org>; Tue, 29 Apr 2014 09:49:47 -0700 (PDT)
Received: from omta02.westchester.pa.mail.comcast.net ([76.96.62.19]) by qmta14.westchester.pa.mail.comcast.net with comcast id vsN81n0070QuhwU5EspmUL; Tue, 29 Apr 2014 16:49:46 +0000
Received: from Paul-Kyzivats-MacBook-Pro.local ([50.138.229.164]) by omta02.westchester.pa.mail.comcast.net with comcast id vspm1n0093ZTu2S3NspmXP; Tue, 29 Apr 2014 16:49:46 +0000
Message-ID: <535FD82A.80708@alum.mit.edu>
Date: Tue, 29 Apr 2014 12:49:46 -0400
From: Paul Kyzivat <pkyzivat@alum.mit.edu>
User-Agent: Mozilla/5.0 (Macintosh; Intel Mac OS X 10.7; rv:24.0) Gecko/20100101 Thunderbird/24.4.0
MIME-Version: 1.0
To: clue@ietf.org
References: <533AF351.9050201@alum.mit.edu> <20140410222705.GW39240@verdi> <BLU0-SMTP6910DA20C9A9287137B585D0550@phx.gbl> <53475572.4040807@nteczone.com> <20140411150055.GE60844@verdi> <49E45C59CA48264997FEBFB29B6BC2D6215295A8AE@CRPMBOXPRD07.polycom.com> <BLU0-SMTP1804A70E3E0F666EAEAE84D0540@phx.gbl> <49E45C59CA48264997FEBFB29B6BC2D62409219636@CRPMBOXPRD07.polycom.com> <534B4303.7060707@nteczone.com> <5C4AC54BFF7A0842A6A11F554D6FB52F07152B@CRPMBOXPRD08.polycom.com> <534B536B.8000205@nteczone.com> <5C4AC54BFF7A0842A6A11F554D6FB52F13D9B2@CRPMBOXPRD08.polycom.com> <535F0B3A.70808@nteczone.com> <535F262C.6090009@alum.mit.edu> <535F73B6.9030106@nteczone.com> <CAMC7SJ7ZNW+=PDB-39gQ3MkQpTh5W1LZLPz7Y8VaFc7OofRviw@mail.gmail.com>
In-Reply-To: <CAMC7SJ7ZNW+=PDB-39gQ3MkQpTh5W1LZLPz7Y8VaFc7OofRviw@mail.gmail.com>
Content-Type: text/plain; charset=ISO-8859-1; format=flowed
Content-Transfer-Encoding: 7bit
DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=comcast.net; s=q20140121; t=1398790186; bh=AXuLKPovDVSr2yNeKWfuSp/UmWfFF6JTuEa8UBiAwXA=; h=Received:Received:Message-ID:Date:From:MIME-Version:To:Subject: Content-Type; b=iFnc8epAl89tXN7YhQI59HDibmUNZuoMqwm8LEi0lGkFkF2zsaU3AVILME53oI4zv TvAinw1EPN0hos98zIsVuz/0k3UcFTmpmqtfL6c7XHJe2dfDw01ZVC3C3gU8iONMyU GMsr8v8FPjb1UbbzaKOT5PBQk068+VsMzRTnZPTyRVBQ34tQYO951NW/IifHUF1kuk J1Wlf+d9Cqn604x9L0EG743V7TCQKY9iv5jnIYPjc0Hsw/xo6Dq/eWDJhPVupMsAVu zGhI1zvXHCV4LSBUnPxEBiFzgAPoQAVdhSkSNuX+R1qy9W8R6Rhr9zGDIl9/fwXx2L J3kdXTjctdDYQ==
Archived-At: http://mailarchive.ietf.org/arch/msg/clue/yUwqgkuF_bhkDQyHYZR5wivfz7k
Subject: Re: [clue] Improving treatment of audio
X-BeenThere: clue@ietf.org
X-Mailman-Version: 2.1.15
Precedence: list
List-Id: CLUE - ControLling mUltiple streams for TElepresence <clue.ietf.org>
List-Unsubscribe: <https://www.ietf.org/mailman/options/clue>, <mailto:clue-request@ietf.org?subject=unsubscribe>
List-Archive: <http://www.ietf.org/mail-archive/web/clue/>
List-Post: <mailto:clue@ietf.org>
List-Help: <mailto:clue-request@ietf.org?subject=help>
List-Subscribe: <https://www.ietf.org/mailman/listinfo/clue>, <mailto:clue-request@ietf.org?subject=subscribe>
X-List-Received-Date: Tue, 29 Apr 2014 16:49:49 -0000

Stephen,

(First my usual disclaimer that I know nothing about audio, so these may 
be naive comments.)

I more-or-less follow what you are saying below. My main concern is the 
same as I mentioned yesterday. You aren't giving me enough detail about 
how you are assuming certain things are determined from the advertisement.

E.g., you say "for each received audio channel, it can identify the 
nearest two loudspeakers". But nearest to *what*? Nearest to the center 
of the Area of Capture of the audio after coordinate mapping? Or nearest 
to the Point of Capture of the audio?

	Thanks,
	Paul

On 4/29/14 10:35 AM, Stephen Botzko wrote:
> This may be a digression, but I think it would be useful to identify
> specific scenarios that work (and don't work) before we go in and try
> and improve things.  I'm providing my assessment below for
> comments/feedback.
>
> BR
> Stephen
>
>
> -Room to Room rendering works.
>
>     In the normal telepresence point-to-point case, there is a large
>     video wall in the front, with some number of speakers also in the
>     front.  The size of the wall, the number of cameras/screens, and the
>     number of loudspeakers is mismatched.
>
>     On the video side, there is enough spatial information to allow the
>     receivers to render the far end room "properly", and several
>     strategies that receivers can use.  Once the receiver knows the
>     mapping between its local coordinates and the far end, it can then
>     make corresponding adjustments to the received audio.  For instance,
>     for each received audio channel, it can identify the nearest two
>     loudspeakers, and can use normal panning to create a sound source at
>     a point in between.  The spatial audio is "soft" - similar to the
>     results you get when you derive a center channel from two stereo
>     speakers.  A center speaker in the right place certainly sounds more
>     realistic, but using two stereo speakers to simulate it works
>     acceptably.
>
>     For users, even the "soft" spatial audio improves the user
>     experience.  For instance, when someone starts talking, even the
>     "soft" spatial audio is enough to allow the participants to
>     automatically look in the right direction when someone new starts
>     talking.
>
>     If you wanted to use rear speakers to place some sounds behind the
>     local audience, that would not work.  But I don't think that's a
>     very important case.  It is probably more common to use any rear
>     speakers for normal sound reinforcement (exploiting the Haas effect).
>
>     So I believe this scenario is already enabled.  Industry experience
>     with interoperability today (using TIP) pretty much supports that view.
>
>     Could we do better?  In theory some approaches (MPEG SAOC) might be
>     able to create a better spatial sound. Although SAOC proponents
>     sometimes refer to interactive conferencing as an application, I
>     don't think it is well-enabled in practice.  Locating and isolating
>     each sound source in the room in real time is not an easy task, and
>     SAOC gives no hints on how to do it.  SAOC's real application is
>     gaming, where pre-recorded and synthetic sounds are being blended
>     into a virtual reality.  Also, SAOC uses in-band transmission of
>     audio object information, so it could be used with CLUE in a future
>     system.  There would likely need to be some form of tagging the
>     objects, so that the video rendering could also take advantage of
>     them. But that could be added in the future.  SAOC is likely
>     encumbered (just a guess I haven't checked).
>
>     So at some point we could perhaps do better, but it is still a bit
>     of a research project.  I think the current functionality is enough.
>
>
>
> -Whole Room Composition works
>
>     Similarly, in a multipoint case you frequently end up giving up on
>     full size rendering and eye-contact because you simply don't have
>     the screen real estate.  In those cases, you are scaling the room
>     video down (usually to something a lot smaller) and placing it
>     somewhere in the composition.
>
>     On the audio side, the panning techniques in the full room-to-room
>     case still apply, and give a satisfactory (though soft") spatial
>     audio experience.  Sounds that are off-camera in the far-end room
>     might be placed in an adjacent room in the composition.  However the
>     existing spatial system allows receivers to detect that the far-end
>     sound stage is wider than the video scene.  And there is a remedy -
>     receivers can also choose to down-mix the far-end room audio to a
>     single mono capture, and align that with the video.  In many cases,
>     that works just as well (since the room view on the screen is
>     relatively small), and it ensures that all the audio from the room
>     is aligned with the video composition.
>
>     So I also think the current functionality is enough.
>
> -Rendering in a deliberately different spatial alignment does not work.
>
>     One could chose to totally ignore the spatial information for video.
>       For instance, you might wish to present independent video tiles
>     for the various talkers, putting the current talker in the center,
>     and other recent talkers around the side - totally ignoring the
>     spatial information and what room the talker is in.
>
>     You can do this with video.  However, re-processing the sound from
>     the far end rooms in this case doesn't work well.  The spatial sound
>     would not be coherent or natural, because each audio capture
>     includes some sound from the other parts of the room, and in general
>     there is no way to remove that sound.  So in this scenario,
>     receivers probably need to fall back to monophonic audio (rendering
>     all the audio from all senders monophonically).
>
>     It would be nice to do better here.  But I think the problem is
>     really in the audio capture itself.  Changing the signaling won't
>     fix it.  So the current functionality is not enough for this case,
>     but I think for now commercial systems can't do better than
>     monophonic rendering.
>
> -Talker identification does not work.
>
>     It is common in videoconferencing to use "voice-activated
>     switching".   That works by detecting speech in the far-end audio
>     channel, and then automatically switching to the corresponding video
>     channel.
>
>     This is a problem area for CLUE (and it is also a problem area I've
>     seen with systems using TIP).
>
>     Part of the problem grounded in the audio capture itself.  A talker
>     in the room is to some extent picked up by all the microphones.  You
>     might think it is easy to identify the right capture from the volume
>     level, but in many microphone arrangements that doesn't work
>     reliably.  Participants sitting between microphones are hard to
>     locate.  Participants also turn towards other people in the local
>     room when they talk, which often changes the microphone that picks
>     them up best.  And often there are multiple simultaneous talkers in
>     the room - we are not always polite.  So even detecting the
>     approximate talker location(s) from the far end is problematic.  My
>     company's products can locate the talker(s) within the room
>     accurately, but they are using information that is not available to
>     far-end systems, and which would be hard to send in a
>     product-independent way.
>
>     Personally I favor explicit signaling from the sender on which
>     capture(s) carry the current talker(s).  Middle boxes and far end
>     systems can determine which rooms have active talkers by down-mixing
>     to mono.  To get finer granularity within the room, you could then
>     use the explicit signaling.
>
>     The second aspect of the problem is associating the talker with a
>     video capture.  Again, this seems somewhat difficult, especially if
>     the audio captures are not precisely aligned to the video captures.
>     We can probably improve this - but if we need explicit talker
>     signaling anyway, it could potentially identify the best video
>     capture. In fact, since I personally think that whole-room audio
>     rendering should always be done, I am much more interested in the
>     video capture than the audio capture.
>
>     Other approaches might work, the main point I am making here is that
>     remote identification of talker locations in the room is probably
>     broken.
>
> -Partial Room Rendering will not always work well.
>
>     If you are selecting some audio captures from a room, and discarding
>     others, you might not get natural sounding audio.  How good it
>     sounds will depend on the details of the microphone pickup in the
>     far end room, how controlled participant seating is, and which
>     captures you happen to discard.
>
>     Since the goal for CLUE is interoperability across disparate room
>     designs, we can't make a lot of assumptions about microphone pickup.
>       I think we can assume that the full room audio will sound
>     appropriate - if it doesn't, then that means that the sender doesn't
>     have a well-designed system.  But once you start filtering out
>     captures in the middle, the results will vary.  I'm not seeing much
>     CLUE can do about that, other than point it out.  As noted above, I
>     think that rendering the entire room's transmitted sound field is
>     going to sound the best.  If there is a need for more conditioning
>     (say noise reduction) it is probably better to do it locally at the
>     sender.
>
>
>
>
>
>
>
>
> _______________________________________________
> clue mailing list
> clue@ietf.org
> https://www.ietf.org/mailman/listinfo/clue
>


From nobody Tue Apr 29 11:48:57 2014
Return-Path: <stephen.botzko@gmail.com>
X-Original-To: clue@ietfa.amsl.com
Delivered-To: clue@ietfa.amsl.com
Received: from localhost (ietfa.amsl.com [127.0.0.1]) by ietfa.amsl.com (Postfix) with ESMTP id 44BB61A07D9 for <clue@ietfa.amsl.com>; Tue, 29 Apr 2014 11:48:56 -0700 (PDT)
X-Virus-Scanned: amavisd-new at amsl.com
X-Spam-Flag: NO
X-Spam-Score: -1.999
X-Spam-Level: 
X-Spam-Status: No, score=-1.999 tagged_above=-999 required=5 tests=[BAYES_00=-1.9, DKIM_SIGNED=0.1, DKIM_VALID=-0.1, DKIM_VALID_AU=-0.1, FREEMAIL_FROM=0.001, HTML_MESSAGE=0.001, SPF_PASS=-0.001] autolearn=ham
Received: from mail.ietf.org ([4.31.198.44]) by localhost (ietfa.amsl.com [127.0.0.1]) (amavisd-new, port 10024) with ESMTP id nTmZWr8RNR19 for <clue@ietfa.amsl.com>; Tue, 29 Apr 2014 11:48:52 -0700 (PDT)
Received: from mail-ve0-x22a.google.com (mail-ve0-x22a.google.com [IPv6:2607:f8b0:400c:c01::22a]) by ietfa.amsl.com (Postfix) with ESMTP id 21A0E1A0515 for <clue@ietf.org>; Tue, 29 Apr 2014 11:48:52 -0700 (PDT)
Received: by mail-ve0-f170.google.com with SMTP id sa20so815342veb.1 for <clue@ietf.org>; Tue, 29 Apr 2014 11:48:50 -0700 (PDT)
DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20120113;  h=mime-version:in-reply-to:references:date:message-id:subject:from:to :cc:content-type; bh=l5BS28PB9MH6p7CnFNP933B5WjKtDI8dOzOUPhSJVOM=; b=tlwY4oOcVj/XWX0WpioVXheI2Vhfs9OPNjMmEfELdkGPDN9ycObNBRtUgXe3JzzYkX GamP1yOWR+Q37jCAnbKeVHxQgsuwKpYscW4+F7uGFox4XWf0YpafRVtyXQ3+SD8DShNp x/Uy7d2h2FjigRi2Hy04BkB1931ld3Kcm7PcD59DqAB24puTC4OjxE3kx7ueBJGxyfML 8JAybhMUdg866l4vaA/0s93URcuVmIEAp4qg/9FImtSQUE8GbTpk1p9013B8ji/nu4GG PRymyGdmjsRMrhJXBMWhhhiHAmcLwAwJVv1TpXQhfLB75HXpT3KfllrA+macZJApqoQl C+ug==
MIME-Version: 1.0
X-Received: by 10.58.23.6 with SMTP id i6mr620495vef.12.1398797330712; Tue, 29 Apr 2014 11:48:50 -0700 (PDT)
Received: by 10.221.40.135 with HTTP; Tue, 29 Apr 2014 11:48:50 -0700 (PDT)
In-Reply-To: <535FD82A.80708@alum.mit.edu>
References: <533AF351.9050201@alum.mit.edu> <20140410222705.GW39240@verdi> <BLU0-SMTP6910DA20C9A9287137B585D0550@phx.gbl> <53475572.4040807@nteczone.com> <20140411150055.GE60844@verdi> <49E45C59CA48264997FEBFB29B6BC2D6215295A8AE@CRPMBOXPRD07.polycom.com> <BLU0-SMTP1804A70E3E0F666EAEAE84D0540@phx.gbl> <49E45C59CA48264997FEBFB29B6BC2D62409219636@CRPMBOXPRD07.polycom.com> <534B4303.7060707@nteczone.com> <5C4AC54BFF7A0842A6A11F554D6FB52F07152B@CRPMBOXPRD08.polycom.com> <534B536B.8000205@nteczone.com> <5C4AC54BFF7A0842A6A11F554D6FB52F13D9B2@CRPMBOXPRD08.polycom.com> <535F0B3A.70808@nteczone.com> <535F262C.6090009@alum.mit.edu> <535F73B6.9030106@nteczone.com> <CAMC7SJ7ZNW+=PDB-39gQ3MkQpTh5W1LZLPz7Y8VaFc7OofRviw@mail.gmail.com> <535FD82A.80708@alum.mit.edu>
Date: Tue, 29 Apr 2014 14:48:50 -0400
Message-ID: <CAMC7SJ6TGUYQVNcf__At7rkwgBQkoYBXAOfmnnLPW56=-=96wA@mail.gmail.com>
From: Stephen Botzko <stephen.botzko@gmail.com>
To: Paul Kyzivat <pkyzivat@alum.mit.edu>
Content-Type: multipart/alternative; boundary=047d7b339db1028e4104f832e24b
Archived-At: http://mailarchive.ietf.org/arch/msg/clue/XGksE5XI_vL2exyu8i0bv84wzc4
Cc: CLUE <clue@ietf.org>
Subject: Re: [clue] Improving treatment of audio
X-BeenThere: clue@ietf.org
X-Mailman-Version: 2.1.15
Precedence: list
List-Id: CLUE - ControLling mUltiple streams for TElepresence <clue.ietf.org>
List-Unsubscribe: <https://www.ietf.org/mailman/options/clue>, <mailto:clue-request@ietf.org?subject=unsubscribe>
List-Archive: <http://www.ietf.org/mail-archive/web/clue/>
List-Post: <mailto:clue@ietf.org>
List-Help: <mailto:clue-request@ietf.org?subject=help>
List-Subscribe: <https://www.ietf.org/mailman/listinfo/clue>, <mailto:clue-request@ietf.org?subject=subscribe>
X-List-Received-Date: Tue, 29 Apr 2014 18:48:56 -0000

--047d7b339db1028e4104f832e24b
Content-Type: text/plain; charset=UTF-8

In-line


On Tue, Apr 29, 2014 at 12:49 PM, Paul Kyzivat <pkyzivat@alum.mit.edu>wrote:

> Stephen,
>
> (First my usual disclaimer that I know nothing about audio, so these may
> be naive comments.)
>
> I more-or-less follow what you are saying below. My main concern is the
> same as I mentioned yesterday. You aren't giving me enough detail about how
> you are assuming certain things are determined from the advertisement.
>
> E.g., you say "for each received audio channel, it can identify the
> nearest two loudspeakers". But nearest to *what*? Nearest to the center of
> the Area of Capture of the audio after coordinate mapping? Or nearest to
> the Point of Capture of the audio?
>
[sb]Center of the area was what I was thinking, but Point of Capture could
probably also be used in practice.. Either way, that point is mapped onto a
corresponding pixel on the video display (which is showing that same point
in the far end room's video).  If the point isn't displayed because it is
off camera, then you can still compute where it would have been displayed
if you had a big enough screen.

The rendered audio capture is panned so it sounds like it is coming from
that pixel.  To do that you can use the two nearest speakers to that pixel,
and you send the right proportion of sound to each (proportionally more
sound to the closer speaker - just linear interpolation).  In practice
people won't be able to localize the sound very precisely with this
approach, which is one of things I was trying to communicate with the
"soft" spatial audio description.  Most multi-channel loudspeaker systems
don't allow humans to localize the sound all that accurately - which I
think was one of the points John Leslie also made.  But it is good enough
to be useful.

What I mean by localize: imagine blindfolding people and ask them to point
to the apparent sound source.  They won't be spot-on with the CLUE
approach, and they might feel tentative on the direction they choose (this
is often the case with spatial audio rendered through multiple
loudspeakers).  But they should be pointing in approximately the right
direction most of the time.

I can provide more details if you ask more questions.  Though I am also
wanting folks who are conversant with audio to provide their own thoughts
on what works and what doesn't.[/sb]



>
>         Thanks,
>         Paul
>
>
> On 4/29/14 10:35 AM, Stephen Botzko wrote:
>
>> This may be a digression, but I think it would be useful to identify
>> specific scenarios that work (and don't work) before we go in and try
>> and improve things.  I'm providing my assessment below for
>> comments/feedback.
>>
>> BR
>> Stephen
>>
>>
>> -Room to Room rendering works.
>>
>>     In the normal telepresence point-to-point case, there is a large
>>     video wall in the front, with some number of speakers also in the
>>     front.  The size of the wall, the number of cameras/screens, and the
>>     number of loudspeakers is mismatched.
>>
>>     On the video side, there is enough spatial information to allow the
>>     receivers to render the far end room "properly", and several
>>     strategies that receivers can use.  Once the receiver knows the
>>     mapping between its local coordinates and the far end, it can then
>>     make corresponding adjustments to the received audio.  For instance,
>>     for each received audio channel, it can identify the nearest two
>>     loudspeakers, and can use normal panning to create a sound source at
>>     a point in between.  The spatial audio is "soft" - similar to the
>>     results you get when you derive a center channel from two stereo
>>     speakers.  A center speaker in the right place certainly sounds more
>>     realistic, but using two stereo speakers to simulate it works
>>     acceptably.
>>
>>     For users, even the "soft" spatial audio improves the user
>>     experience.  For instance, when someone starts talking, even the
>>     "soft" spatial audio is enough to allow the participants to
>>     automatically look in the right direction when someone new starts
>>     talking.
>>
>>     If you wanted to use rear speakers to place some sounds behind the
>>     local audience, that would not work.  But I don't think that's a
>>     very important case.  It is probably more common to use any rear
>>     speakers for normal sound reinforcement (exploiting the Haas effect).
>>
>>     So I believe this scenario is already enabled.  Industry experience
>>     with interoperability today (using TIP) pretty much supports that
>> view.
>>
>>     Could we do better?  In theory some approaches (MPEG SAOC) might be
>>     able to create a better spatial sound. Although SAOC proponents
>>     sometimes refer to interactive conferencing as an application, I
>>     don't think it is well-enabled in practice.  Locating and isolating
>>     each sound source in the room in real time is not an easy task, and
>>     SAOC gives no hints on how to do it.  SAOC's real application is
>>     gaming, where pre-recorded and synthetic sounds are being blended
>>     into a virtual reality.  Also, SAOC uses in-band transmission of
>>     audio object information, so it could be used with CLUE in a future
>>     system.  There would likely need to be some form of tagging the
>>     objects, so that the video rendering could also take advantage of
>>     them. But that could be added in the future.  SAOC is likely
>>     encumbered (just a guess I haven't checked).
>>
>>     So at some point we could perhaps do better, but it is still a bit
>>     of a research project.  I think the current functionality is enough.
>>
>>
>>
>> -Whole Room Composition works
>>
>>     Similarly, in a multipoint case you frequently end up giving up on
>>     full size rendering and eye-contact because you simply don't have
>>     the screen real estate.  In those cases, you are scaling the room
>>     video down (usually to something a lot smaller) and placing it
>>     somewhere in the composition.
>>
>>     On the audio side, the panning techniques in the full room-to-room
>>     case still apply, and give a satisfactory (though soft") spatial
>>     audio experience.  Sounds that are off-camera in the far-end room
>>     might be placed in an adjacent room in the composition.  However the
>>     existing spatial system allows receivers to detect that the far-end
>>     sound stage is wider than the video scene.  And there is a remedy -
>>     receivers can also choose to down-mix the far-end room audio to a
>>     single mono capture, and align that with the video.  In many cases,
>>     that works just as well (since the room view on the screen is
>>     relatively small), and it ensures that all the audio from the room
>>     is aligned with the video composition.
>>
>>     So I also think the current functionality is enough.
>>
>> -Rendering in a deliberately different spatial alignment does not work.
>>
>>     One could chose to totally ignore the spatial information for video.
>>       For instance, you might wish to present independent video tiles
>>     for the various talkers, putting the current talker in the center,
>>     and other recent talkers around the side - totally ignoring the
>>     spatial information and what room the talker is in.
>>
>>     You can do this with video.  However, re-processing the sound from
>>     the far end rooms in this case doesn't work well.  The spatial sound
>>     would not be coherent or natural, because each audio capture
>>     includes some sound from the other parts of the room, and in general
>>     there is no way to remove that sound.  So in this scenario,
>>     receivers probably need to fall back to monophonic audio (rendering
>>     all the audio from all senders monophonically).
>>
>>     It would be nice to do better here.  But I think the problem is
>>     really in the audio capture itself.  Changing the signaling won't
>>     fix it.  So the current functionality is not enough for this case,
>>     but I think for now commercial systems can't do better than
>>     monophonic rendering.
>>
>> -Talker identification does not work.
>>
>>     It is common in videoconferencing to use "voice-activated
>>     switching".   That works by detecting speech in the far-end audio
>>     channel, and then automatically switching to the corresponding video
>>     channel.
>>
>>     This is a problem area for CLUE (and it is also a problem area I've
>>     seen with systems using TIP).
>>
>>     Part of the problem grounded in the audio capture itself.  A talker
>>     in the room is to some extent picked up by all the microphones.  You
>>     might think it is easy to identify the right capture from the volume
>>     level, but in many microphone arrangements that doesn't work
>>     reliably.  Participants sitting between microphones are hard to
>>     locate.  Participants also turn towards other people in the local
>>     room when they talk, which often changes the microphone that picks
>>     them up best.  And often there are multiple simultaneous talkers in
>>     the room - we are not always polite.  So even detecting the
>>     approximate talker location(s) from the far end is problematic.  My
>>     company's products can locate the talker(s) within the room
>>     accurately, but they are using information that is not available to
>>     far-end systems, and which would be hard to send in a
>>     product-independent way.
>>
>>     Personally I favor explicit signaling from the sender on which
>>     capture(s) carry the current talker(s).  Middle boxes and far end
>>     systems can determine which rooms have active talkers by down-mixing
>>     to mono.  To get finer granularity within the room, you could then
>>     use the explicit signaling.
>>
>>     The second aspect of the problem is associating the talker with a
>>     video capture.  Again, this seems somewhat difficult, especially if
>>     the audio captures are not precisely aligned to the video captures.
>>     We can probably improve this - but if we need explicit talker
>>     signaling anyway, it could potentially identify the best video
>>     capture. In fact, since I personally think that whole-room audio
>>     rendering should always be done, I am much more interested in the
>>     video capture than the audio capture.
>>
>>     Other approaches might work, the main point I am making here is that
>>     remote identification of talker locations in the room is probably
>>     broken.
>>
>> -Partial Room Rendering will not always work well.
>>
>>     If you are selecting some audio captures from a room, and discarding
>>     others, you might not get natural sounding audio.  How good it
>>     sounds will depend on the details of the microphone pickup in the
>>     far end room, how controlled participant seating is, and which
>>     captures you happen to discard.
>>
>>     Since the goal for CLUE is interoperability across disparate room
>>     designs, we can't make a lot of assumptions about microphone pickup.
>>       I think we can assume that the full room audio will sound
>>     appropriate - if it doesn't, then that means that the sender doesn't
>>     have a well-designed system.  But once you start filtering out
>>     captures in the middle, the results will vary.  I'm not seeing much
>>     CLUE can do about that, other than point it out.  As noted above, I
>>     think that rendering the entire room's transmitted sound field is
>>     going to sound the best.  If there is a need for more conditioning
>>     (say noise reduction) it is probably better to do it locally at the
>>     sender.
>>
>>
>>
>>
>>
>>
>>
>>
>> _______________________________________________
>> clue mailing list
>> clue@ietf.org
>> https://www.ietf.org/mailman/listinfo/clue
>>
>>
> _______________________________________________
> clue mailing list
> clue@ietf.org
> https://www.ietf.org/mailman/listinfo/clue
>

--047d7b339db1028e4104f832e24b
Content-Type: text/html; charset=UTF-8
Content-Transfer-Encoding: quoted-printable

<div dir=3D"ltr">In-line<br><div class=3D"gmail_extra"><br><br><div class=
=3D"gmail_quote">On Tue, Apr 29, 2014 at 12:49 PM, Paul Kyzivat <span dir=
=3D"ltr">&lt;<a href=3D"mailto:pkyzivat@alum.mit.edu" target=3D"_blank">pky=
zivat@alum.mit.edu</a>&gt;</span> wrote:<br>
<blockquote class=3D"gmail_quote" style=3D"margin:0 0 0 .8ex;border-left:1p=
x #ccc solid;padding-left:1ex">Stephen,<br>
<br>
(First my usual disclaimer that I know nothing about audio, so these may be=
 naive comments.)<br>
<br>
I more-or-less follow what you are saying below. My main concern is the sam=
e as I mentioned yesterday. You aren&#39;t giving me enough detail about ho=
w you are assuming certain things are determined from the advertisement.<br=
>

<br>
E.g., you say &quot;for each received audio channel, it can identify the ne=
arest two loudspeakers&quot;. But nearest to *what*? Nearest to the center =
of the Area of Capture of the audio after coordinate mapping? Or nearest to=
 the Point of Capture of the audio?<br>
</blockquote><div>[sb]Center of the area was what I was thinking, but Point=
 of Capture could probably also be used in practice.. Either way, that poin=
t is mapped onto a corresponding pixel on the video display (which is showi=
ng that same point in the far end room&#39;s video). =C2=A0If the point isn=
&#39;t displayed because it is off camera, then you can still compute where=
 it would have been displayed if you had a big enough screen.</div>
<div><br></div><div>The rendered audio capture is panned so it sounds like =
it is coming from that pixel. =C2=A0To do that you can use the two nearest =
speakers to that pixel, and you send the right proportion of sound to each =
(proportionally more sound to the closer speaker - just linear interpolatio=
n). =C2=A0In practice people won&#39;t be able to localize the sound very p=
recisely with this approach, which is one of things I was trying to communi=
cate with the &quot;soft&quot; spatial audio description. =C2=A0Most multi-=
channel loudspeaker systems don&#39;t allow humans to localize the sound al=
l that accurately - which I think was one of the points John Leslie also ma=
de. =C2=A0But it is good enough to be useful.</div>
<div><br></div><div>What I mean by localize: imagine blindfolding people an=
d ask them to point to the apparent sound source. =C2=A0They won&#39;t be s=
pot-on with the CLUE approach, and they might feel tentative on the directi=
on they choose (this is often the case with spatial audio rendered through =
multiple loudspeakers). =C2=A0But they should be pointing in approximately =
the right direction most of the time. =C2=A0</div>
<div><br></div><div>I can provide more details if you ask more questions. =
=C2=A0Though I am also wanting folks who are conversant with audio to provi=
de their own thoughts on what works and what doesn&#39;t.[/sb]</div><div><b=
r>
</div><div>=C2=A0</div><blockquote class=3D"gmail_quote" style=3D"margin:0 =
0 0 .8ex;border-left:1px #ccc solid;padding-left:1ex">
<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 Thanks,<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 Paul<div><div class=3D"h5"><br>
<br>
On 4/29/14 10:35 AM, Stephen Botzko wrote:<br>
</div></div><blockquote class=3D"gmail_quote" style=3D"margin:0 0 0 .8ex;bo=
rder-left:1px #ccc solid;padding-left:1ex"><div><div class=3D"h5">
This may be a digression, but I think it would be useful to identify<br>
specific scenarios that work (and don&#39;t work) before we go in and try<b=
r>
and improve things. =C2=A0I&#39;m providing my assessment below for<br>
comments/feedback.<br>
<br>
BR<br>
Stephen<br>
<br>
<br>
-Room to Room rendering works.<br>
<br>
=C2=A0 =C2=A0 In the normal telepresence point-to-point case, there is a la=
rge<br>
=C2=A0 =C2=A0 video wall in the front, with some number of speakers also in=
 the<br>
=C2=A0 =C2=A0 front. =C2=A0The size of the wall, the number of cameras/scre=
ens, and the<br>
=C2=A0 =C2=A0 number of loudspeakers is mismatched.<br>
<br>
=C2=A0 =C2=A0 On the video side, there is enough spatial information to all=
ow the<br>
=C2=A0 =C2=A0 receivers to render the far end room &quot;properly&quot;, an=
d several<br>
=C2=A0 =C2=A0 strategies that receivers can use. =C2=A0Once the receiver kn=
ows the<br>
=C2=A0 =C2=A0 mapping between its local coordinates and the far end, it can=
 then<br>
=C2=A0 =C2=A0 make corresponding adjustments to the received audio. =C2=A0F=
or instance,<br>
=C2=A0 =C2=A0 for each received audio channel, it can identify the nearest =
two<br>
=C2=A0 =C2=A0 loudspeakers, and can use normal panning to create a sound so=
urce at<br>
=C2=A0 =C2=A0 a point in between. =C2=A0The spatial audio is &quot;soft&quo=
t; - similar to the<br>
=C2=A0 =C2=A0 results you get when you derive a center channel from two ste=
reo<br>
=C2=A0 =C2=A0 speakers. =C2=A0A center speaker in the right place certainly=
 sounds more<br>
=C2=A0 =C2=A0 realistic, but using two stereo speakers to simulate it works=
<br>
=C2=A0 =C2=A0 acceptably.<br>
<br>
=C2=A0 =C2=A0 For users, even the &quot;soft&quot; spatial audio improves t=
he user<br>
=C2=A0 =C2=A0 experience. =C2=A0For instance, when someone starts talking, =
even the<br>
=C2=A0 =C2=A0 &quot;soft&quot; spatial audio is enough to allow the partici=
pants to<br>
=C2=A0 =C2=A0 automatically look in the right direction when someone new st=
arts<br>
=C2=A0 =C2=A0 talking.<br>
<br>
=C2=A0 =C2=A0 If you wanted to use rear speakers to place some sounds behin=
d the<br>
=C2=A0 =C2=A0 local audience, that would not work. =C2=A0But I don&#39;t th=
ink that&#39;s a<br>
=C2=A0 =C2=A0 very important case. =C2=A0It is probably more common to use =
any rear<br>
=C2=A0 =C2=A0 speakers for normal sound reinforcement (exploiting the Haas =
effect).<br>
<br>
=C2=A0 =C2=A0 So I believe this scenario is already enabled. =C2=A0Industry=
 experience<br>
=C2=A0 =C2=A0 with interoperability today (using TIP) pretty much supports =
that view.<br>
<br>
=C2=A0 =C2=A0 Could we do better? =C2=A0In theory some approaches (MPEG SAO=
C) might be<br>
=C2=A0 =C2=A0 able to create a better spatial sound. Although SAOC proponen=
ts<br>
=C2=A0 =C2=A0 sometimes refer to interactive conferencing as an application=
, I<br>
=C2=A0 =C2=A0 don&#39;t think it is well-enabled in practice. =C2=A0Locatin=
g and isolating<br>
=C2=A0 =C2=A0 each sound source in the room in real time is not an easy tas=
k, and<br>
=C2=A0 =C2=A0 SAOC gives no hints on how to do it. =C2=A0SAOC&#39;s real ap=
plication is<br>
=C2=A0 =C2=A0 gaming, where pre-recorded and synthetic sounds are being ble=
nded<br>
=C2=A0 =C2=A0 into a virtual reality. =C2=A0Also, SAOC uses in-band transmi=
ssion of<br>
=C2=A0 =C2=A0 audio object information, so it could be used with CLUE in a =
future<br>
=C2=A0 =C2=A0 system. =C2=A0There would likely need to be some form of tagg=
ing the<br>
=C2=A0 =C2=A0 objects, so that the video rendering could also take advantag=
e of<br>
=C2=A0 =C2=A0 them. But that could be added in the future. =C2=A0SAOC is li=
kely<br>
=C2=A0 =C2=A0 encumbered (just a guess I haven&#39;t checked).<br>
<br>
=C2=A0 =C2=A0 So at some point we could perhaps do better, but it is still =
a bit<br>
=C2=A0 =C2=A0 of a research project. =C2=A0I think the current functionalit=
y is enough.<br>
<br>
<br>
<br>
-Whole Room Composition works<br>
<br>
=C2=A0 =C2=A0 Similarly, in a multipoint case you frequently end up giving =
up on<br>
=C2=A0 =C2=A0 full size rendering and eye-contact because you simply don&#3=
9;t have<br>
=C2=A0 =C2=A0 the screen real estate. =C2=A0In those cases, you are scaling=
 the room<br>
=C2=A0 =C2=A0 video down (usually to something a lot smaller) and placing i=
t<br>
=C2=A0 =C2=A0 somewhere in the composition.<br>
<br>
=C2=A0 =C2=A0 On the audio side, the panning techniques in the full room-to=
-room<br>
=C2=A0 =C2=A0 case still apply, and give a satisfactory (though soft&quot;)=
 spatial<br>
=C2=A0 =C2=A0 audio experience. =C2=A0Sounds that are off-camera in the far=
-end room<br>
=C2=A0 =C2=A0 might be placed in an adjacent room in the composition. =C2=
=A0However the<br>
=C2=A0 =C2=A0 existing spatial system allows receivers to detect that the f=
ar-end<br>
=C2=A0 =C2=A0 sound stage is wider than the video scene. =C2=A0And there is=
 a remedy -<br>
=C2=A0 =C2=A0 receivers can also choose to down-mix the far-end room audio =
to a<br>
=C2=A0 =C2=A0 single mono capture, and align that with the video. =C2=A0In =
many cases,<br>
=C2=A0 =C2=A0 that works just as well (since the room view on the screen is=
<br>
=C2=A0 =C2=A0 relatively small), and it ensures that all the audio from the=
 room<br>
=C2=A0 =C2=A0 is aligned with the video composition.<br>
<br>
=C2=A0 =C2=A0 So I also think the current functionality is enough.<br>
<br>
-Rendering in a deliberately different spatial alignment does not work.<br>
<br>
=C2=A0 =C2=A0 One could chose to totally ignore the spatial information for=
 video.<br>
=C2=A0 =C2=A0 =C2=A0 For instance, you might wish to present independent vi=
deo tiles<br>
=C2=A0 =C2=A0 for the various talkers, putting the current talker in the ce=
nter,<br>
=C2=A0 =C2=A0 and other recent talkers around the side - totally ignoring t=
he<br>
=C2=A0 =C2=A0 spatial information and what room the talker is in.<br>
<br>
=C2=A0 =C2=A0 You can do this with video. =C2=A0However, re-processing the =
sound from<br>
=C2=A0 =C2=A0 the far end rooms in this case doesn&#39;t work well. =C2=A0T=
he spatial sound<br>
=C2=A0 =C2=A0 would not be coherent or natural, because each audio capture<=
br>
=C2=A0 =C2=A0 includes some sound from the other parts of the room, and in =
general<br>
=C2=A0 =C2=A0 there is no way to remove that sound. =C2=A0So in this scenar=
io,<br>
=C2=A0 =C2=A0 receivers probably need to fall back to monophonic audio (ren=
dering<br>
=C2=A0 =C2=A0 all the audio from all senders monophonically).<br>
<br>
=C2=A0 =C2=A0 It would be nice to do better here. =C2=A0But I think the pro=
blem is<br>
=C2=A0 =C2=A0 really in the audio capture itself. =C2=A0Changing the signal=
ing won&#39;t<br>
=C2=A0 =C2=A0 fix it. =C2=A0So the current functionality is not enough for =
this case,<br>
=C2=A0 =C2=A0 but I think for now commercial systems can&#39;t do better th=
an<br>
=C2=A0 =C2=A0 monophonic rendering.<br>
<br>
-Talker identification does not work.<br>
<br>
=C2=A0 =C2=A0 It is common in videoconferencing to use &quot;voice-activate=
d<br>
=C2=A0 =C2=A0 switching&quot;. =C2=A0 That works by detecting speech in the=
 far-end audio<br>
=C2=A0 =C2=A0 channel, and then automatically switching to the correspondin=
g video<br>
=C2=A0 =C2=A0 channel.<br>
<br>
=C2=A0 =C2=A0 This is a problem area for CLUE (and it is also a problem are=
a I&#39;ve<br>
=C2=A0 =C2=A0 seen with systems using TIP).<br>
<br>
=C2=A0 =C2=A0 Part of the problem grounded in the audio capture itself. =C2=
=A0A talker<br>
=C2=A0 =C2=A0 in the room is to some extent picked up by all the microphone=
s. =C2=A0You<br>
=C2=A0 =C2=A0 might think it is easy to identify the right capture from the=
 volume<br>
=C2=A0 =C2=A0 level, but in many microphone arrangements that doesn&#39;t w=
ork<br>
=C2=A0 =C2=A0 reliably. =C2=A0Participants sitting between microphones are =
hard to<br>
=C2=A0 =C2=A0 locate. =C2=A0Participants also turn towards other people in =
the local<br>
=C2=A0 =C2=A0 room when they talk, which often changes the microphone that =
picks<br>
=C2=A0 =C2=A0 them up best. =C2=A0And often there are multiple simultaneous=
 talkers in<br>
=C2=A0 =C2=A0 the room - we are not always polite. =C2=A0So even detecting =
the<br>
=C2=A0 =C2=A0 approximate talker location(s) from the far end is problemati=
c. =C2=A0My<br>
=C2=A0 =C2=A0 company&#39;s products can locate the talker(s) within the ro=
om<br>
=C2=A0 =C2=A0 accurately, but they are using information that is not availa=
ble to<br>
=C2=A0 =C2=A0 far-end systems, and which would be hard to send in a<br>
=C2=A0 =C2=A0 product-independent way.<br>
<br>
=C2=A0 =C2=A0 Personally I favor explicit signaling from the sender on whic=
h<br>
=C2=A0 =C2=A0 capture(s) carry the current talker(s). =C2=A0Middle boxes an=
d far end<br>
=C2=A0 =C2=A0 systems can determine which rooms have active talkers by down=
-mixing<br>
=C2=A0 =C2=A0 to mono. =C2=A0To get finer granularity within the room, you =
could then<br>
=C2=A0 =C2=A0 use the explicit signaling.<br>
<br>
=C2=A0 =C2=A0 The second aspect of the problem is associating the talker wi=
th a<br>
=C2=A0 =C2=A0 video capture. =C2=A0Again, this seems somewhat difficult, es=
pecially if<br>
=C2=A0 =C2=A0 the audio captures are not precisely aligned to the video cap=
tures.<br>
=C2=A0 =C2=A0 We can probably improve this - but if we need explicit talker=
<br>
=C2=A0 =C2=A0 signaling anyway, it could potentially identify the best vide=
o<br>
=C2=A0 =C2=A0 capture. In fact, since I personally think that whole-room au=
dio<br>
=C2=A0 =C2=A0 rendering should always be done, I am much more interested in=
 the<br>
=C2=A0 =C2=A0 video capture than the audio capture.<br>
<br>
=C2=A0 =C2=A0 Other approaches might work, the main point I am making here =
is that<br>
=C2=A0 =C2=A0 remote identification of talker locations in the room is prob=
ably<br>
=C2=A0 =C2=A0 broken.<br>
<br>
-Partial Room Rendering will not always work well.<br>
<br>
=C2=A0 =C2=A0 If you are selecting some audio captures from a room, and dis=
carding<br>
=C2=A0 =C2=A0 others, you might not get natural sounding audio. =C2=A0How g=
ood it<br>
=C2=A0 =C2=A0 sounds will depend on the details of the microphone pickup in=
 the<br>
=C2=A0 =C2=A0 far end room, how controlled participant seating is, and whic=
h<br>
=C2=A0 =C2=A0 captures you happen to discard.<br>
<br>
=C2=A0 =C2=A0 Since the goal for CLUE is interoperability across disparate =
room<br>
=C2=A0 =C2=A0 designs, we can&#39;t make a lot of assumptions about microph=
one pickup.<br>
=C2=A0 =C2=A0 =C2=A0 I think we can assume that the full room audio will so=
und<br>
=C2=A0 =C2=A0 appropriate - if it doesn&#39;t, then that means that the sen=
der doesn&#39;t<br>
=C2=A0 =C2=A0 have a well-designed system. =C2=A0But once you start filteri=
ng out<br>
=C2=A0 =C2=A0 captures in the middle, the results will vary. =C2=A0I&#39;m =
not seeing much<br>
=C2=A0 =C2=A0 CLUE can do about that, other than point it out. =C2=A0As not=
ed above, I<br>
=C2=A0 =C2=A0 think that rendering the entire room&#39;s transmitted sound =
field is<br>
=C2=A0 =C2=A0 going to sound the best. =C2=A0If there is a need for more co=
nditioning<br>
=C2=A0 =C2=A0 (say noise reduction) it is probably better to do it locally =
at the<br>
=C2=A0 =C2=A0 sender.<br>
<br>
<br>
<br>
<br>
<br>
<br>
<br>
<br></div></div><div class=3D"">
______________________________<u></u>_________________<br>
clue mailing list<br>
<a href=3D"mailto:clue@ietf.org" target=3D"_blank">clue@ietf.org</a><br>
<a href=3D"https://www.ietf.org/mailman/listinfo/clue" target=3D"_blank">ht=
tps://www.ietf.org/mailman/<u></u>listinfo/clue</a><br>
<br>
</div></blockquote><div class=3D"HOEnZb"><div class=3D"h5">
<br>
______________________________<u></u>_________________<br>
clue mailing list<br>
<a href=3D"mailto:clue@ietf.org" target=3D"_blank">clue@ietf.org</a><br>
<a href=3D"https://www.ietf.org/mailman/listinfo/clue" target=3D"_blank">ht=
tps://www.ietf.org/mailman/<u></u>listinfo/clue</a><br>
</div></div></blockquote></div><br></div></div>

--047d7b339db1028e4104f832e24b--


From nobody Tue Apr 29 12:44:16 2014
Return-Path: <pkyzivat@alum.mit.edu>
X-Original-To: clue@ietfa.amsl.com
Delivered-To: clue@ietfa.amsl.com
Received: from localhost (ietfa.amsl.com [127.0.0.1]) by ietfa.amsl.com (Postfix) with ESMTP id 3094D1A09C4 for <clue@ietfa.amsl.com>; Tue, 29 Apr 2014 12:44:15 -0700 (PDT)
X-Virus-Scanned: amavisd-new at amsl.com
X-Spam-Flag: NO
X-Spam-Score: -1.235
X-Spam-Level: 
X-Spam-Status: No, score=-1.235 tagged_above=-999 required=5 tests=[BAYES_00=-1.9, DKIM_SIGNED=0.1, DKIM_VALID=-0.1, SPF_SOFTFAIL=0.665] autolearn=no
Received: from mail.ietf.org ([4.31.198.44]) by localhost (ietfa.amsl.com [127.0.0.1]) (amavisd-new, port 10024) with ESMTP id 1tbumvzsvsZT for <clue@ietfa.amsl.com>; Tue, 29 Apr 2014 12:44:13 -0700 (PDT)
Received: from qmta13.westchester.pa.mail.comcast.net (qmta13.westchester.pa.mail.comcast.net [IPv6:2001:558:fe14:44:76:96:59:243]) by ietfa.amsl.com (Postfix) with ESMTP id 2E0B81A09B6 for <clue@ietf.org>; Tue, 29 Apr 2014 12:44:12 -0700 (PDT)
Received: from omta10.westchester.pa.mail.comcast.net ([76.96.62.28]) by qmta13.westchester.pa.mail.comcast.net with comcast id vtAM1n0090cZkys5DvkBkV; Tue, 29 Apr 2014 19:44:11 +0000
Received: from Paul-Kyzivats-MacBook-Pro.local ([50.138.229.164]) by omta10.westchester.pa.mail.comcast.net with comcast id vvkA1n00z3ZTu2S3WvkAWg; Tue, 29 Apr 2014 19:44:11 +0000
Message-ID: <5360010A.1000904@alum.mit.edu>
Date: Tue, 29 Apr 2014 15:44:10 -0400
From: Paul Kyzivat <pkyzivat@alum.mit.edu>
User-Agent: Mozilla/5.0 (Macintosh; Intel Mac OS X 10.7; rv:24.0) Gecko/20100101 Thunderbird/24.4.0
MIME-Version: 1.0
To: Stephen Botzko <stephen.botzko@gmail.com>
References: <533AF351.9050201@alum.mit.edu>	<20140410222705.GW39240@verdi>	<BLU0-SMTP6910DA20C9A9287137B585D0550@phx.gbl>	<53475572.4040807@nteczone.com>	<20140411150055.GE60844@verdi>	<49E45C59CA48264997FEBFB29B6BC2D6215295A8AE@CRPMBOXPRD07.polycom.com>	<BLU0-SMTP1804A70E3E0F666EAEAE84D0540@phx.gbl>	<49E45C59CA48264997FEBFB29B6BC2D62409219636@CRPMBOXPRD07.polycom.com>	<534B4303.7060707@nteczone.com>	<5C4AC54BFF7A0842A6A11F554D6FB52F07152B@CRPMBOXPRD08.polycom.com>	<534B536B.8000205@nteczone.com>	<5C4AC54BFF7A0842A6A11F554D6FB52F13D9B2@CRPMBOXPRD08.polycom.com>	<535F0B3A.70808@nteczone.com>	<535F262C.6090009@alum.mit.edu>	<535F73B6.9030106@nteczone.com>	<CAMC7SJ7ZNW+=PDB-39gQ3MkQpTh5W1LZLPz7Y8VaFc7OofRviw@mail.gmail.com>	<535FD82A.80708@alum.mit.edu> <CAMC7SJ6TGUYQVNcf__At7rkwgBQkoYBXAOfmnnLPW56=-=96wA@mail.gmail.com>
In-Reply-To: <CAMC7SJ6TGUYQVNcf__At7rkwgBQkoYBXAOfmnnLPW56=-=96wA@mail.gmail.com>
Content-Type: text/plain; charset=UTF-8; format=flowed
Content-Transfer-Encoding: 7bit
DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=comcast.net; s=q20140121; t=1398800651; bh=Tk/uYXlDvSywSdvyuhz4SBzwPJrO89aXN4lU5QnNy5c=; h=Received:Received:Message-ID:Date:From:MIME-Version:To:Subject: Content-Type; b=ifzJYvekh1+AG8oDSBahBOXvcvkSSvcvLifnlVdEzDVfQUNNUbudU4N7MvYiXMoMQ IkqNBdMQlzS0D8g9de0uuSts2TJ4sW0XcMZDF5wJzrQRaO4qcE0DONQhxIVcBvrAjz O4xbgOmjiYMzAJJ1y/iyp7BxxY0ggn40nuEVtLYUBFh3GCPr+m8hwYV5fY1Y2c2DE6 cjXGxMbX0g4uHGlSQZvGwGGJgvDPX5wIT43lvZ56W+Gl5C7z0kPY4x0+JLFVR7SMCD wJaHvn2V03YiRu7YOK5p1+wVkVl/6IXrt0F6E6kAkbefgp514U/bzc40iQrzeBIFA2 M0gxkH44W2wrA==
Archived-At: http://mailarchive.ietf.org/arch/msg/clue/act2-XEC21t_DJ7Pp6-UI100Ai0
Cc: CLUE <clue@ietf.org>
Subject: Re: [clue] Improving treatment of audio
X-BeenThere: clue@ietf.org
X-Mailman-Version: 2.1.15
Precedence: list
List-Id: CLUE - ControLling mUltiple streams for TElepresence <clue.ietf.org>
List-Unsubscribe: <https://www.ietf.org/mailman/options/clue>, <mailto:clue-request@ietf.org?subject=unsubscribe>
List-Archive: <http://www.ietf.org/mail-archive/web/clue/>
List-Post: <mailto:clue@ietf.org>
List-Help: <mailto:clue-request@ietf.org?subject=help>
List-Subscribe: <https://www.ietf.org/mailman/listinfo/clue>, <mailto:clue-request@ietf.org?subject=subscribe>
X-List-Received-Date: Tue, 29 Apr 2014 19:44:15 -0000

I realize this is probably not a topic that ought to be normative. But 
ISTM that *something* should be said about it. I think the advertiser 
and the consumer need to be in agreement on what this information means 
or how it is intended to be used in order to get interoperability. There 
seems to be no one obvious way to pick an Area of Capture for a 
microphone, but different decisions about what to choose may well effect 
the way the consumer's algorithm works.

 From what you describe, perhaps there is no value in area of capture 
for audio, and axis of capture would be more useful. And I assume that 
in the absence of axis of capture, one may approximate it with a line 
through the center of the area of capture and perpendicular to it.

	Thanks,
	Paul

On 4/29/14 2:48 PM, Stephen Botzko wrote:
> In-line
>
>
> On Tue, Apr 29, 2014 at 12:49 PM, Paul Kyzivat <pkyzivat@alum.mit.edu
> <mailto:pkyzivat@alum.mit.edu>> wrote:
>
>     Stephen,
>
>     (First my usual disclaimer that I know nothing about audio, so these
>     may be naive comments.)
>
>     I more-or-less follow what you are saying below. My main concern is
>     the same as I mentioned yesterday. You aren't giving me enough
>     detail about how you are assuming certain things are determined from
>     the advertisement.
>
>     E.g., you say "for each received audio channel, it can identify the
>     nearest two loudspeakers". But nearest to *what*? Nearest to the
>     center of the Area of Capture of the audio after coordinate mapping?
>     Or nearest to the Point of Capture of the audio?
>
> [sb]Center of the area was what I was thinking, but Point of Capture
> could probably also be used in practice.. Either way, that point is
> mapped onto a corresponding pixel on the video display (which is showing
> that same point in the far end room's video).  If the point isn't
> displayed because it is off camera, then you can still compute where it
> would have been displayed if you had a big enough screen.
>
> The rendered audio capture is panned so it sounds like it is coming from
> that pixel.  To do that you can use the two nearest speakers to that
> pixel, and you send the right proportion of sound to each
> (proportionally more sound to the closer speaker - just linear
> interpolation).  In practice people won't be able to localize the sound
> very precisely with this approach, which is one of things I was trying
> to communicate with the "soft" spatial audio description.  Most
> multi-channel loudspeaker systems don't allow humans to localize the
> sound all that accurately - which I think was one of the points John
> Leslie also made.  But it is good enough to be useful.
>
> What I mean by localize: imagine blindfolding people and ask them to
> point to the apparent sound source.  They won't be spot-on with the CLUE
> approach, and they might feel tentative on the direction they choose
> (this is often the case with spatial audio rendered through multiple
> loudspeakers).  But they should be pointing in approximately the right
> direction most of the time.
>
> I can provide more details if you ask more questions.  Though I am also
> wanting folks who are conversant with audio to provide their own
> thoughts on what works and what doesn't.[/sb]
>
>
>              Thanks,
>              Paul
>
>
>     On 4/29/14 10:35 AM, Stephen Botzko wrote:
>
>         This may be a digression, but I think it would be useful to identify
>         specific scenarios that work (and don't work) before we go in
>         and try
>         and improve things.  I'm providing my assessment below for
>         comments/feedback.
>
>         BR
>         Stephen
>
>
>         -Room to Room rendering works.
>
>              In the normal telepresence point-to-point case, there is a
>         large
>              video wall in the front, with some number of speakers also
>         in the
>              front.  The size of the wall, the number of
>         cameras/screens, and the
>              number of loudspeakers is mismatched.
>
>              On the video side, there is enough spatial information to
>         allow the
>              receivers to render the far end room "properly", and several
>              strategies that receivers can use.  Once the receiver knows the
>              mapping between its local coordinates and the far end, it
>         can then
>              make corresponding adjustments to the received audio.  For
>         instance,
>              for each received audio channel, it can identify the
>         nearest two
>              loudspeakers, and can use normal panning to create a sound
>         source at
>              a point in between.  The spatial audio is "soft" - similar
>         to the
>              results you get when you derive a center channel from two
>         stereo
>              speakers.  A center speaker in the right place certainly
>         sounds more
>              realistic, but using two stereo speakers to simulate it works
>              acceptably.
>
>              For users, even the "soft" spatial audio improves the user
>              experience.  For instance, when someone starts talking,
>         even the
>              "soft" spatial audio is enough to allow the participants to
>              automatically look in the right direction when someone new
>         starts
>              talking.
>
>              If you wanted to use rear speakers to place some sounds
>         behind the
>              local audience, that would not work.  But I don't think
>         that's a
>              very important case.  It is probably more common to use any
>         rear
>              speakers for normal sound reinforcement (exploiting the
>         Haas effect).
>
>              So I believe this scenario is already enabled.  Industry
>         experience
>              with interoperability today (using TIP) pretty much
>         supports that view.
>
>              Could we do better?  In theory some approaches (MPEG SAOC)
>         might be
>              able to create a better spatial sound. Although SAOC proponents
>              sometimes refer to interactive conferencing as an
>         application, I
>              don't think it is well-enabled in practice.  Locating and
>         isolating
>              each sound source in the room in real time is not an easy
>         task, and
>              SAOC gives no hints on how to do it.  SAOC's real
>         application is
>              gaming, where pre-recorded and synthetic sounds are being
>         blended
>              into a virtual reality.  Also, SAOC uses in-band
>         transmission of
>              audio object information, so it could be used with CLUE in
>         a future
>              system.  There would likely need to be some form of tagging the
>              objects, so that the video rendering could also take
>         advantage of
>              them. But that could be added in the future.  SAOC is likely
>              encumbered (just a guess I haven't checked).
>
>              So at some point we could perhaps do better, but it is
>         still a bit
>              of a research project.  I think the current functionality
>         is enough.
>
>
>
>         -Whole Room Composition works
>
>              Similarly, in a multipoint case you frequently end up
>         giving up on
>              full size rendering and eye-contact because you simply
>         don't have
>              the screen real estate.  In those cases, you are scaling
>         the room
>              video down (usually to something a lot smaller) and placing it
>              somewhere in the composition.
>
>              On the audio side, the panning techniques in the full
>         room-to-room
>              case still apply, and give a satisfactory (though soft")
>         spatial
>              audio experience.  Sounds that are off-camera in the
>         far-end room
>              might be placed in an adjacent room in the composition.
>           However the
>              existing spatial system allows receivers to detect that the
>         far-end
>              sound stage is wider than the video scene.  And there is a
>         remedy -
>              receivers can also choose to down-mix the far-end room
>         audio to a
>              single mono capture, and align that with the video.  In
>         many cases,
>              that works just as well (since the room view on the screen is
>              relatively small), and it ensures that all the audio from
>         the room
>              is aligned with the video composition.
>
>              So I also think the current functionality is enough.
>
>         -Rendering in a deliberately different spatial alignment does
>         not work.
>
>              One could chose to totally ignore the spatial information
>         for video.
>                For instance, you might wish to present independent video
>         tiles
>              for the various talkers, putting the current talker in the
>         center,
>              and other recent talkers around the side - totally ignoring the
>              spatial information and what room the talker is in.
>
>              You can do this with video.  However, re-processing the
>         sound from
>              the far end rooms in this case doesn't work well.  The
>         spatial sound
>              would not be coherent or natural, because each audio capture
>              includes some sound from the other parts of the room, and
>         in general
>              there is no way to remove that sound.  So in this scenario,
>              receivers probably need to fall back to monophonic audio
>         (rendering
>              all the audio from all senders monophonically).
>
>              It would be nice to do better here.  But I think the problem is
>              really in the audio capture itself.  Changing the signaling
>         won't
>              fix it.  So the current functionality is not enough for
>         this case,
>              but I think for now commercial systems can't do better than
>              monophonic rendering.
>
>         -Talker identification does not work.
>
>              It is common in videoconferencing to use "voice-activated
>              switching".   That works by detecting speech in the far-end
>         audio
>              channel, and then automatically switching to the
>         corresponding video
>              channel.
>
>              This is a problem area for CLUE (and it is also a problem
>         area I've
>              seen with systems using TIP).
>
>              Part of the problem grounded in the audio capture itself.
>           A talker
>              in the room is to some extent picked up by all the
>         microphones.  You
>              might think it is easy to identify the right capture from
>         the volume
>              level, but in many microphone arrangements that doesn't work
>              reliably.  Participants sitting between microphones are hard to
>              locate.  Participants also turn towards other people in the
>         local
>              room when they talk, which often changes the microphone
>         that picks
>              them up best.  And often there are multiple simultaneous
>         talkers in
>              the room - we are not always polite.  So even detecting the
>              approximate talker location(s) from the far end is
>         problematic.  My
>              company's products can locate the talker(s) within the room
>              accurately, but they are using information that is not
>         available to
>              far-end systems, and which would be hard to send in a
>              product-independent way.
>
>              Personally I favor explicit signaling from the sender on which
>              capture(s) carry the current talker(s).  Middle boxes and
>         far end
>              systems can determine which rooms have active talkers by
>         down-mixing
>              to mono.  To get finer granularity within the room, you
>         could then
>              use the explicit signaling.
>
>              The second aspect of the problem is associating the talker
>         with a
>              video capture.  Again, this seems somewhat difficult,
>         especially if
>              the audio captures are not precisely aligned to the video
>         captures.
>              We can probably improve this - but if we need explicit talker
>              signaling anyway, it could potentially identify the best video
>              capture. In fact, since I personally think that whole-room
>         audio
>              rendering should always be done, I am much more interested
>         in the
>              video capture than the audio capture.
>
>              Other approaches might work, the main point I am making
>         here is that
>              remote identification of talker locations in the room is
>         probably
>              broken.
>
>         -Partial Room Rendering will not always work well.
>
>              If you are selecting some audio captures from a room, and
>         discarding
>              others, you might not get natural sounding audio.  How good it
>              sounds will depend on the details of the microphone pickup
>         in the
>              far end room, how controlled participant seating is, and which
>              captures you happen to discard.
>
>              Since the goal for CLUE is interoperability across
>         disparate room
>              designs, we can't make a lot of assumptions about
>         microphone pickup.
>                I think we can assume that the full room audio will sound
>              appropriate - if it doesn't, then that means that the
>         sender doesn't
>              have a well-designed system.  But once you start filtering out
>              captures in the middle, the results will vary.  I'm not
>         seeing much
>              CLUE can do about that, other than point it out.  As noted
>         above, I
>              think that rendering the entire room's transmitted sound
>         field is
>              going to sound the best.  If there is a need for more
>         conditioning
>              (say noise reduction) it is probably better to do it
>         locally at the
>              sender.
>
>
>
>
>
>
>
>
>         _________________________________________________
>         clue mailing list
>         clue@ietf.org <mailto:clue@ietf.org>
>         https://www.ietf.org/mailman/__listinfo/clue
>         <https://www.ietf.org/mailman/listinfo/clue>
>
>
>     _________________________________________________
>     clue mailing list
>     clue@ietf.org <mailto:clue@ietf.org>
>     https://www.ietf.org/mailman/__listinfo/clue
>     <https://www.ietf.org/mailman/listinfo/clue>
>
>


From nobody Tue Apr 29 13:07:06 2014
Return-Path: <stephen.botzko@gmail.com>
X-Original-To: clue@ietfa.amsl.com
Delivered-To: clue@ietfa.amsl.com
Received: from localhost (ietfa.amsl.com [127.0.0.1]) by ietfa.amsl.com (Postfix) with ESMTP id EA6C91A0979 for <clue@ietfa.amsl.com>; Tue, 29 Apr 2014 13:07:02 -0700 (PDT)
X-Virus-Scanned: amavisd-new at amsl.com
X-Spam-Flag: NO
X-Spam-Score: -1.999
X-Spam-Level: 
X-Spam-Status: No, score=-1.999 tagged_above=-999 required=5 tests=[BAYES_00=-1.9, DKIM_SIGNED=0.1, DKIM_VALID=-0.1, DKIM_VALID_AU=-0.1, FREEMAIL_FROM=0.001, HTML_MESSAGE=0.001, SPF_PASS=-0.001] autolearn=ham
Received: from mail.ietf.org ([4.31.198.44]) by localhost (ietfa.amsl.com [127.0.0.1]) (amavisd-new, port 10024) with ESMTP id 9jGvKIkAMCKg for <clue@ietfa.amsl.com>; Tue, 29 Apr 2014 13:06:58 -0700 (PDT)
Received: from mail-ve0-x22d.google.com (mail-ve0-x22d.google.com [IPv6:2607:f8b0:400c:c01::22d]) by ietfa.amsl.com (Postfix) with ESMTP id 47E131A0927 for <clue@ietf.org>; Tue, 29 Apr 2014 13:06:58 -0700 (PDT)
Received: by mail-ve0-f173.google.com with SMTP id oy12so935581veb.18 for <clue@ietf.org>; Tue, 29 Apr 2014 13:06:56 -0700 (PDT)
DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20120113;  h=mime-version:in-reply-to:references:date:message-id:subject:from:to :cc:content-type; bh=t63p9Iuhofo2uGwMattrG3TaTKg35WDPOOE16zNLXvM=; b=BTX7CuqrqHlzQJ77wOLyMXTe8GvW/DAOIhC66/v4P4aTm8WeYwM506L0WVPXR88xKX qbGGy1S77MiB5Slw3ZRa9hBfZ7ybOElN4aRx4vZsHE+a8HzQxkxcvxThzu3HtM26cyP3 xtYlrcsTbBqi79Wl7svfOCU9YFofWyG7KX0tUQqdQTzb2oXezzhVBJv9AUQK83ZvXZZY MRmTMIT8Mxl+ydMab9LPBuqRkKN11+T7pfD/PXaqsY91N77+dSyxur4jKHtXqD1zg3vP FUsotnpKsaYewV76KzPlkcKipzO388P3zHvqkGIa8LAnxgLYvjzBHcBAlIAI8LLzMJp9 Smew==
MIME-Version: 1.0
X-Received: by 10.58.171.229 with SMTP id ax5mr75830vec.24.1398802016866; Tue, 29 Apr 2014 13:06:56 -0700 (PDT)
Received: by 10.221.40.135 with HTTP; Tue, 29 Apr 2014 13:06:56 -0700 (PDT)
In-Reply-To: <5360010A.1000904@alum.mit.edu>
References: <533AF351.9050201@alum.mit.edu> <20140410222705.GW39240@verdi> <BLU0-SMTP6910DA20C9A9287137B585D0550@phx.gbl> <53475572.4040807@nteczone.com> <20140411150055.GE60844@verdi> <49E45C59CA48264997FEBFB29B6BC2D6215295A8AE@CRPMBOXPRD07.polycom.com> <BLU0-SMTP1804A70E3E0F666EAEAE84D0540@phx.gbl> <49E45C59CA48264997FEBFB29B6BC2D62409219636@CRPMBOXPRD07.polycom.com> <534B4303.7060707@nteczone.com> <5C4AC54BFF7A0842A6A11F554D6FB52F07152B@CRPMBOXPRD08.polycom.com> <534B536B.8000205@nteczone.com> <5C4AC54BFF7A0842A6A11F554D6FB52F13D9B2@CRPMBOXPRD08.polycom.com> <535F0B3A.70808@nteczone.com> <535F262C.6090009@alum.mit.edu> <535F73B6.9030106@nteczone.com> <CAMC7SJ7ZNW+=PDB-39gQ3MkQpTh5W1LZLPz7Y8VaFc7OofRviw@mail.gmail.com> <535FD82A.80708@alum.mit.edu> <CAMC7SJ6TGUYQVNcf__At7rkwgBQkoYBXAOfmnnLPW56=-=96wA@mail.gmail.com> <5360010A.1000904@alum.mit.edu>
Date: Tue, 29 Apr 2014 16:06:56 -0400
Message-ID: <CAMC7SJ6ARW5Z2AX526AiabO1AMi5qd0C8uMEc63PRdHWv_b8Kw@mail.gmail.com>
From: Stephen Botzko <stephen.botzko@gmail.com>
To: Paul Kyzivat <pkyzivat@alum.mit.edu>
Content-Type: multipart/alternative; boundary=047d7b675e6653986b04f833f988
Archived-At: http://mailarchive.ietf.org/arch/msg/clue/qkXO1bWw0Aysbk3MQBEdeu9JPwM
Cc: CLUE <clue@ietf.org>
Subject: Re: [clue] Improving treatment of audio
X-BeenThere: clue@ietf.org
X-Mailman-Version: 2.1.15
Precedence: list
List-Id: CLUE - ControLling mUltiple streams for TElepresence <clue.ietf.org>
List-Unsubscribe: <https://www.ietf.org/mailman/options/clue>, <mailto:clue-request@ietf.org?subject=unsubscribe>
List-Archive: <http://www.ietf.org/mail-archive/web/clue/>
List-Post: <mailto:clue@ietf.org>
List-Help: <mailto:clue-request@ietf.org?subject=help>
List-Subscribe: <https://www.ietf.org/mailman/listinfo/clue>, <mailto:clue-request@ietf.org?subject=subscribe>
X-List-Received-Date: Tue, 29 Apr 2014 20:07:03 -0000

--047d7b675e6653986b04f833f988
Content-Type: text/plain; charset=UTF-8

On Tue, Apr 29, 2014 at 3:44 PM, Paul Kyzivat <pkyzivat@alum.mit.edu> wrote:

> I realize this is probably not a topic that ought to be normative. But
> ISTM that *something* should be said about it. I think the advertiser and
> the consumer need to be in agreement on what this information means or how
> it is intended to be used in order to get interoperability. There seems to
> be no one obvious way to pick an Area of Capture for a microphone, but
> different decisions about what to choose may well effect the way the
> consumer's algorithm works.

I agree that some information is needed. Perhaps continuing to drill down
on some of these cases will help clarify what needs to be said.

>
>
>From what you describe, perhaps there is no value in area of capture for
> audio, and axis of capture would be more useful. And I assume that in the
> absence of axis of capture, one may approximate it with a line through the
> center of the area of capture and perpendicular to it.
>
Axis of capture could be more useful, though I am not aware of any
techniques that would let you align your output audio along a different
axis from the loudspeaker's natural orientation.


>         Thanks,
>         Paul
>
>
> On 4/29/14 2:48 PM, Stephen Botzko wrote:
>
>> In-line
>>
>>
>> On Tue, Apr 29, 2014 at 12:49 PM, Paul Kyzivat <pkyzivat@alum.mit.edu
>> <mailto:pkyzivat@alum.mit.edu>> wrote:
>>
>>     Stephen,
>>
>>     (First my usual disclaimer that I know nothing about audio, so these
>>     may be naive comments.)
>>
>>     I more-or-less follow what you are saying below. My main concern is
>>     the same as I mentioned yesterday. You aren't giving me enough
>>     detail about how you are assuming certain things are determined from
>>     the advertisement.
>>
>>     E.g., you say "for each received audio channel, it can identify the
>>     nearest two loudspeakers". But nearest to *what*? Nearest to the
>>     center of the Area of Capture of the audio after coordinate mapping?
>>     Or nearest to the Point of Capture of the audio?
>>
>> [sb]Center of the area was what I was thinking, but Point of Capture
>> could probably also be used in practice.. Either way, that point is
>> mapped onto a corresponding pixel on the video display (which is showing
>> that same point in the far end room's video).  If the point isn't
>> displayed because it is off camera, then you can still compute where it
>> would have been displayed if you had a big enough screen.
>>
>> The rendered audio capture is panned so it sounds like it is coming from
>> that pixel.  To do that you can use the two nearest speakers to that
>> pixel, and you send the right proportion of sound to each
>> (proportionally more sound to the closer speaker - just linear
>> interpolation).  In practice people won't be able to localize the sound
>> very precisely with this approach, which is one of things I was trying
>> to communicate with the "soft" spatial audio description.  Most
>> multi-channel loudspeaker systems don't allow humans to localize the
>> sound all that accurately - which I think was one of the points John
>> Leslie also made.  But it is good enough to be useful.
>>
>> What I mean by localize: imagine blindfolding people and ask them to
>> point to the apparent sound source.  They won't be spot-on with the CLUE
>> approach, and they might feel tentative on the direction they choose
>> (this is often the case with spatial audio rendered through multiple
>> loudspeakers).  But they should be pointing in approximately the right
>> direction most of the time.
>>
>> I can provide more details if you ask more questions.  Though I am also
>> wanting folks who are conversant with audio to provide their own
>> thoughts on what works and what doesn't.[/sb]
>>
>>
>>              Thanks,
>>              Paul
>>
>>
>>     On 4/29/14 10:35 AM, Stephen Botzko wrote:
>>
>>         This may be a digression, but I think it would be useful to
>> identify
>>         specific scenarios that work (and don't work) before we go in
>>         and try
>>         and improve things.  I'm providing my assessment below for
>>         comments/feedback.
>>
>>         BR
>>         Stephen
>>
>>
>>         -Room to Room rendering works.
>>
>>              In the normal telepresence point-to-point case, there is a
>>         large
>>              video wall in the front, with some number of speakers also
>>         in the
>>              front.  The size of the wall, the number of
>>         cameras/screens, and the
>>              number of loudspeakers is mismatched.
>>
>>              On the video side, there is enough spatial information to
>>         allow the
>>              receivers to render the far end room "properly", and several
>>              strategies that receivers can use.  Once the receiver knows
>> the
>>              mapping between its local coordinates and the far end, it
>>         can then
>>              make corresponding adjustments to the received audio.  For
>>         instance,
>>              for each received audio channel, it can identify the
>>         nearest two
>>              loudspeakers, and can use normal panning to create a sound
>>         source at
>>              a point in between.  The spatial audio is "soft" - similar
>>         to the
>>              results you get when you derive a center channel from two
>>         stereo
>>              speakers.  A center speaker in the right place certainly
>>         sounds more
>>              realistic, but using two stereo speakers to simulate it works
>>              acceptably.
>>
>>              For users, even the "soft" spatial audio improves the user
>>              experience.  For instance, when someone starts talking,
>>         even the
>>              "soft" spatial audio is enough to allow the participants to
>>              automatically look in the right direction when someone new
>>         starts
>>              talking.
>>
>>              If you wanted to use rear speakers to place some sounds
>>         behind the
>>              local audience, that would not work.  But I don't think
>>         that's a
>>              very important case.  It is probably more common to use any
>>         rear
>>              speakers for normal sound reinforcement (exploiting the
>>         Haas effect).
>>
>>              So I believe this scenario is already enabled.  Industry
>>         experience
>>              with interoperability today (using TIP) pretty much
>>         supports that view.
>>
>>              Could we do better?  In theory some approaches (MPEG SAOC)
>>         might be
>>              able to create a better spatial sound. Although SAOC
>> proponents
>>              sometimes refer to interactive conferencing as an
>>         application, I
>>              don't think it is well-enabled in practice.  Locating and
>>         isolating
>>              each sound source in the room in real time is not an easy
>>         task, and
>>              SAOC gives no hints on how to do it.  SAOC's real
>>         application is
>>              gaming, where pre-recorded and synthetic sounds are being
>>         blended
>>              into a virtual reality.  Also, SAOC uses in-band
>>         transmission of
>>              audio object information, so it could be used with CLUE in
>>         a future
>>              system.  There would likely need to be some form of tagging
>> the
>>              objects, so that the video rendering could also take
>>         advantage of
>>              them. But that could be added in the future.  SAOC is likely
>>              encumbered (just a guess I haven't checked).
>>
>>              So at some point we could perhaps do better, but it is
>>         still a bit
>>              of a research project.  I think the current functionality
>>         is enough.
>>
>>
>>
>>         -Whole Room Composition works
>>
>>              Similarly, in a multipoint case you frequently end up
>>         giving up on
>>              full size rendering and eye-contact because you simply
>>         don't have
>>              the screen real estate.  In those cases, you are scaling
>>         the room
>>              video down (usually to something a lot smaller) and placing
>> it
>>              somewhere in the composition.
>>
>>              On the audio side, the panning techniques in the full
>>         room-to-room
>>              case still apply, and give a satisfactory (though soft")
>>         spatial
>>              audio experience.  Sounds that are off-camera in the
>>         far-end room
>>              might be placed in an adjacent room in the composition.
>>           However the
>>              existing spatial system allows receivers to detect that the
>>         far-end
>>              sound stage is wider than the video scene.  And there is a
>>         remedy -
>>              receivers can also choose to down-mix the far-end room
>>         audio to a
>>              single mono capture, and align that with the video.  In
>>         many cases,
>>              that works just as well (since the room view on the screen is
>>              relatively small), and it ensures that all the audio from
>>         the room
>>              is aligned with the video composition.
>>
>>              So I also think the current functionality is enough.
>>
>>         -Rendering in a deliberately different spatial alignment does
>>         not work.
>>
>>              One could chose to totally ignore the spatial information
>>         for video.
>>                For instance, you might wish to present independent video
>>         tiles
>>              for the various talkers, putting the current talker in the
>>         center,
>>              and other recent talkers around the side - totally ignoring
>> the
>>              spatial information and what room the talker is in.
>>
>>              You can do this with video.  However, re-processing the
>>         sound from
>>              the far end rooms in this case doesn't work well.  The
>>         spatial sound
>>              would not be coherent or natural, because each audio capture
>>              includes some sound from the other parts of the room, and
>>         in general
>>              there is no way to remove that sound.  So in this scenario,
>>              receivers probably need to fall back to monophonic audio
>>         (rendering
>>              all the audio from all senders monophonically).
>>
>>              It would be nice to do better here.  But I think the problem
>> is
>>              really in the audio capture itself.  Changing the signaling
>>         won't
>>              fix it.  So the current functionality is not enough for
>>         this case,
>>              but I think for now commercial systems can't do better than
>>              monophonic rendering.
>>
>>         -Talker identification does not work.
>>
>>              It is common in videoconferencing to use "voice-activated
>>              switching".   That works by detecting speech in the far-end
>>         audio
>>              channel, and then automatically switching to the
>>         corresponding video
>>              channel.
>>
>>              This is a problem area for CLUE (and it is also a problem
>>         area I've
>>              seen with systems using TIP).
>>
>>              Part of the problem grounded in the audio capture itself.
>>           A talker
>>              in the room is to some extent picked up by all the
>>         microphones.  You
>>              might think it is easy to identify the right capture from
>>         the volume
>>              level, but in many microphone arrangements that doesn't work
>>              reliably.  Participants sitting between microphones are hard
>> to
>>              locate.  Participants also turn towards other people in the
>>         local
>>              room when they talk, which often changes the microphone
>>         that picks
>>              them up best.  And often there are multiple simultaneous
>>         talkers in
>>              the room - we are not always polite.  So even detecting the
>>              approximate talker location(s) from the far end is
>>         problematic.  My
>>              company's products can locate the talker(s) within the room
>>              accurately, but they are using information that is not
>>         available to
>>              far-end systems, and which would be hard to send in a
>>              product-independent way.
>>
>>              Personally I favor explicit signaling from the sender on
>> which
>>              capture(s) carry the current talker(s).  Middle boxes and
>>         far end
>>              systems can determine which rooms have active talkers by
>>         down-mixing
>>              to mono.  To get finer granularity within the room, you
>>         could then
>>              use the explicit signaling.
>>
>>              The second aspect of the problem is associating the talker
>>         with a
>>              video capture.  Again, this seems somewhat difficult,
>>         especially if
>>              the audio captures are not precisely aligned to the video
>>         captures.
>>              We can probably improve this - but if we need explicit talker
>>              signaling anyway, it could potentially identify the best
>> video
>>              capture. In fact, since I personally think that whole-room
>>         audio
>>              rendering should always be done, I am much more interested
>>         in the
>>              video capture than the audio capture.
>>
>>              Other approaches might work, the main point I am making
>>         here is that
>>              remote identification of talker locations in the room is
>>         probably
>>              broken.
>>
>>         -Partial Room Rendering will not always work well.
>>
>>              If you are selecting some audio captures from a room, and
>>         discarding
>>              others, you might not get natural sounding audio.  How good
>> it
>>              sounds will depend on the details of the microphone pickup
>>         in the
>>              far end room, how controlled participant seating is, and
>> which
>>              captures you happen to discard.
>>
>>              Since the goal for CLUE is interoperability across
>>         disparate room
>>              designs, we can't make a lot of assumptions about
>>         microphone pickup.
>>                I think we can assume that the full room audio will sound
>>              appropriate - if it doesn't, then that means that the
>>         sender doesn't
>>              have a well-designed system.  But once you start filtering
>> out
>>              captures in the middle, the results will vary.  I'm not
>>         seeing much
>>              CLUE can do about that, other than point it out.  As noted
>>         above, I
>>              think that rendering the entire room's transmitted sound
>>         field is
>>              going to sound the best.  If there is a need for more
>>         conditioning
>>              (say noise reduction) it is probably better to do it
>>         locally at the
>>              sender.
>>
>>
>>
>>
>>
>>
>>
>>
>>         _________________________________________________
>>         clue mailing list
>>         clue@ietf.org <mailto:clue@ietf.org>
>>         https://www.ietf.org/mailman/__listinfo/clue
>>         <https://www.ietf.org/mailman/listinfo/clue>
>>
>>
>>     _________________________________________________
>>     clue mailing list
>>     clue@ietf.org <mailto:clue@ietf.org>
>>     https://www.ietf.org/mailman/__listinfo/clue
>>     <https://www.ietf.org/mailman/listinfo/clue>
>>
>>
>>
>

--047d7b675e6653986b04f833f988
Content-Type: text/html; charset=UTF-8
Content-Transfer-Encoding: quoted-printable

<div dir=3D"ltr"><br><div class=3D"gmail_extra"><br><br><div class=3D"gmail=
_quote">On Tue, Apr 29, 2014 at 3:44 PM, Paul Kyzivat <span dir=3D"ltr">&lt=
;<a href=3D"mailto:pkyzivat@alum.mit.edu" target=3D"_blank">pkyzivat@alum.m=
it.edu</a>&gt;</span> wrote:<br>
<blockquote class=3D"gmail_quote" style=3D"margin:0 0 0 .8ex;border-left:1p=
x #ccc solid;padding-left:1ex">I realize this is probably not a topic that =
ought to be normative. But ISTM that *something* should be said about it. I=
 think the advertiser and the consumer need to be in agreement on what this=
 information means or how it is intended to be used in order to get interop=
erability. There seems to be no one obvious way to pick an Area of Capture =
for a microphone, but different decisions about what to choose may well eff=
ect the way the consumer&#39;s algorithm works.</blockquote>
<div>I agree that some information is needed. Perhaps continuing to drill d=
own on some of these cases will help clarify what needs to be said.</div><b=
lockquote class=3D"gmail_quote" style=3D"margin:0 0 0 .8ex;border-left:1px =
#ccc solid;padding-left:1ex">
=C2=A0<br></blockquote><blockquote class=3D"gmail_quote" style=3D"margin:0 =
0 0 .8ex;border-left:1px #ccc solid;padding-left:1ex">
>From what you describe, perhaps there is no value in area of capture for au=
dio, and axis of capture would be more useful. And I assume that in the abs=
ence of axis of capture, one may approximate it with a line through the cen=
ter of the area of capture and perpendicular to it.<br>
</blockquote><div>Axis of capture could be more useful, though I am not awa=
re of any techniques that would let you align your output audio along a dif=
ferent axis from the loudspeaker&#39;s natural orientation.=C2=A0</div><div=
>
<br></div><blockquote class=3D"gmail_quote" style=3D"margin:0 0 0 .8ex;bord=
er-left:1px #ccc solid;padding-left:1ex">
<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 Thanks,<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 Paul<div class=3D""><br>
<br>
On 4/29/14 2:48 PM, Stephen Botzko wrote:<br>
</div><blockquote class=3D"gmail_quote" style=3D"margin:0 0 0 .8ex;border-l=
eft:1px #ccc solid;padding-left:1ex"><div class=3D"">
In-line<br>
<br>
<br>
On Tue, Apr 29, 2014 at 12:49 PM, Paul Kyzivat &lt;<a href=3D"mailto:pkyziv=
at@alum.mit.edu" target=3D"_blank">pkyzivat@alum.mit.edu</a><br></div><div>=
<div class=3D"h5">
&lt;mailto:<a href=3D"mailto:pkyzivat@alum.mit.edu" target=3D"_blank">pkyzi=
vat@alum.mit.edu</a>&gt;<u></u>&gt; wrote:<br>
<br>
=C2=A0 =C2=A0 Stephen,<br>
<br>
=C2=A0 =C2=A0 (First my usual disclaimer that I know nothing about audio, s=
o these<br>
=C2=A0 =C2=A0 may be naive comments.)<br>
<br>
=C2=A0 =C2=A0 I more-or-less follow what you are saying below. My main conc=
ern is<br>
=C2=A0 =C2=A0 the same as I mentioned yesterday. You aren&#39;t giving me e=
nough<br>
=C2=A0 =C2=A0 detail about how you are assuming certain things are determin=
ed from<br>
=C2=A0 =C2=A0 the advertisement.<br>
<br>
=C2=A0 =C2=A0 E.g., you say &quot;for each received audio channel, it can i=
dentify the<br>
=C2=A0 =C2=A0 nearest two loudspeakers&quot;. But nearest to *what*? Neares=
t to the<br>
=C2=A0 =C2=A0 center of the Area of Capture of the audio after coordinate m=
apping?<br>
=C2=A0 =C2=A0 Or nearest to the Point of Capture of the audio?<br>
<br>
[sb]Center of the area was what I was thinking, but Point of Capture<br>
could probably also be used in practice.. Either way, that point is<br>
mapped onto a corresponding pixel on the video display (which is showing<br=
>
that same point in the far end room&#39;s video). =C2=A0If the point isn&#3=
9;t<br>
displayed because it is off camera, then you can still compute where it<br>
would have been displayed if you had a big enough screen.<br>
<br>
The rendered audio capture is panned so it sounds like it is coming from<br=
>
that pixel. =C2=A0To do that you can use the two nearest speakers to that<b=
r>
pixel, and you send the right proportion of sound to each<br>
(proportionally more sound to the closer speaker - just linear<br>
interpolation). =C2=A0In practice people won&#39;t be able to localize the =
sound<br>
very precisely with this approach, which is one of things I was trying<br>
to communicate with the &quot;soft&quot; spatial audio description. =C2=A0M=
ost<br>
multi-channel loudspeaker systems don&#39;t allow humans to localize the<br=
>
sound all that accurately - which I think was one of the points John<br>
Leslie also made. =C2=A0But it is good enough to be useful.<br>
<br>
What I mean by localize: imagine blindfolding people and ask them to<br>
point to the apparent sound source. =C2=A0They won&#39;t be spot-on with th=
e CLUE<br>
approach, and they might feel tentative on the direction they choose<br>
(this is often the case with spatial audio rendered through multiple<br>
loudspeakers). =C2=A0But they should be pointing in approximately the right=
<br>
direction most of the time.<br>
<br>
I can provide more details if you ask more questions. =C2=A0Though I am als=
o<br>
wanting folks who are conversant with audio to provide their own<br>
thoughts on what works and what doesn&#39;t.[/sb]<br>
<br>
<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0Thanks,<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0Paul<br>
<br>
<br>
=C2=A0 =C2=A0 On 4/29/14 10:35 AM, Stephen Botzko wrote:<br>
<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 This may be a digression, but I think it would =
be useful to identify<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 specific scenarios that work (and don&#39;t wor=
k) before we go in<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 and try<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 and improve things. =C2=A0I&#39;m providing my =
assessment below for<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 comments/feedback.<br>
<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 BR<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 Stephen<br>
<br>
<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 -Room to Room rendering works.<br>
<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0In the normal telepresence =
point-to-point case, there is a<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 large<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0video wall in the front, wi=
th some number of speakers also<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 in the<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0front. =C2=A0The size of th=
e wall, the number of<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 cameras/screens, and the<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0number of loudspeakers is m=
ismatched.<br>
<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0On the video side, there is=
 enough spatial information to<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 allow the<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0receivers to render the far=
 end room &quot;properly&quot;, and several<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0strategies that receivers c=
an use. =C2=A0Once the receiver knows the<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0mapping between its local c=
oordinates and the far end, it<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 can then<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0make corresponding adjustme=
nts to the received audio. =C2=A0For<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 instance,<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0for each received audio cha=
nnel, it can identify the<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 nearest two<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0loudspeakers, and can use n=
ormal panning to create a sound<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 source at<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0a point in between. =C2=A0T=
he spatial audio is &quot;soft&quot; - similar<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 to the<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0results you get when you de=
rive a center channel from two<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 stereo<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0speakers. =C2=A0A center sp=
eaker in the right place certainly<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 sounds more<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0realistic, but using two st=
ereo speakers to simulate it works<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0acceptably.<br>
<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0For users, even the &quot;s=
oft&quot; spatial audio improves the user<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0experience. =C2=A0For insta=
nce, when someone starts talking,<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 even the<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0&quot;soft&quot; spatial au=
dio is enough to allow the participants to<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0automatically look in the r=
ight direction when someone new<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 starts<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0talking.<br>
<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0If you wanted to use rear s=
peakers to place some sounds<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 behind the<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0local audience, that would =
not work. =C2=A0But I don&#39;t think<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 that&#39;s a<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0very important case. =C2=A0=
It is probably more common to use any<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 rear<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0speakers for normal sound r=
einforcement (exploiting the<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 Haas effect).<br>
<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0So I believe this scenario =
is already enabled. =C2=A0Industry<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 experience<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0with interoperability today=
 (using TIP) pretty much<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 supports that view.<br>
<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0Could we do better? =C2=A0I=
n theory some approaches (MPEG SAOC)<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 might be<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0able to create a better spa=
tial sound. Although SAOC proponents<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0sometimes refer to interact=
ive conferencing as an<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 application, I<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0don&#39;t think it is well-=
enabled in practice. =C2=A0Locating and<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 isolating<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0each sound source in the ro=
om in real time is not an easy<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 task, and<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0SAOC gives no hints on how =
to do it. =C2=A0SAOC&#39;s real<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 application is<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0gaming, where pre-recorded =
and synthetic sounds are being<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 blended<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0into a virtual reality. =C2=
=A0Also, SAOC uses in-band<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 transmission of<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0audio object information, s=
o it could be used with CLUE in<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 a future<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0system. =C2=A0There would l=
ikely need to be some form of tagging the<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0objects, so that the video =
rendering could also take<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 advantage of<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0them. But that could be add=
ed in the future. =C2=A0SAOC is likely<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0encumbered (just a guess I =
haven&#39;t checked).<br>
<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0So at some point we could p=
erhaps do better, but it is<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 still a bit<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0of a research project. =C2=
=A0I think the current functionality<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 is enough.<br>
<br>
<br>
<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 -Whole Room Composition works<br>
<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0Similarly, in a multipoint =
case you frequently end up<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 giving up on<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0full size rendering and eye=
-contact because you simply<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 don&#39;t have<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0the screen real estate. =C2=
=A0In those cases, you are scaling<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 the room<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0video down (usually to some=
thing a lot smaller) and placing it<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0somewhere in the compositio=
n.<br>
<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0On the audio side, the pann=
ing techniques in the full<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 room-to-room<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0case still apply, and give =
a satisfactory (though soft&quot;)<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 spatial<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0audio experience. =C2=A0Sou=
nds that are off-camera in the<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 far-end room<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0might be placed in an adjac=
ent room in the composition.<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 However the<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0existing spatial system all=
ows receivers to detect that the<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 far-end<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0sound stage is wider than t=
he video scene. =C2=A0And there is a<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 remedy -<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0receivers can also choose t=
o down-mix the far-end room<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 audio to a<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0single mono capture, and al=
ign that with the video. =C2=A0In<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 many cases,<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0that works just as well (si=
nce the room view on the screen is<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0relatively small), and it e=
nsures that all the audio from<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 the room<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0is aligned with the video c=
omposition.<br>
<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0So I also think the current=
 functionality is enough.<br>
<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 -Rendering in a deliberately different spatial =
alignment does<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 not work.<br>
<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0One could chose to totally =
ignore the spatial information<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 for video.<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0For instance, you mi=
ght wish to present independent video<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 tiles<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0for the various talkers, pu=
tting the current talker in the<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 center,<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0and other recent talkers ar=
ound the side - totally ignoring the<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0spatial information and wha=
t room the talker is in.<br>
<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0You can do this with video.=
 =C2=A0However, re-processing the<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 sound from<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0the far end rooms in this c=
ase doesn&#39;t work well. =C2=A0The<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 spatial sound<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0would not be coherent or na=
tural, because each audio capture<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0includes some sound from th=
e other parts of the room, and<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 in general<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0there is no way to remove t=
hat sound. =C2=A0So in this scenario,<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0receivers probably need to =
fall back to monophonic audio<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 (rendering<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0all the audio from all send=
ers monophonically).<br>
<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0It would be nice to do bett=
er here. =C2=A0But I think the problem is<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0really in the audio capture=
 itself. =C2=A0Changing the signaling<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 won&#39;t<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0fix it. =C2=A0So the curren=
t functionality is not enough for<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 this case,<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0but I think for now commerc=
ial systems can&#39;t do better than<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0monophonic rendering.<br>
<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 -Talker identification does not work.<br>
<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0It is common in videoconfer=
encing to use &quot;voice-activated<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0switching&quot;. =C2=A0 Tha=
t works by detecting speech in the far-end<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 audio<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0channel, and then automatic=
ally switching to the<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 corresponding video<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0channel.<br>
<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0This is a problem area for =
CLUE (and it is also a problem<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 area I&#39;ve<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0seen with systems using TIP=
).<br>
<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0Part of the problem grounde=
d in the audio capture itself.<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 A talker<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0in the room is to some exte=
nt picked up by all the<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 microphones. =C2=A0You<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0might think it is easy to i=
dentify the right capture from<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 the volume<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0level, but in many micropho=
ne arrangements that doesn&#39;t work<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0reliably. =C2=A0Participant=
s sitting between microphones are hard to<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0locate. =C2=A0Participants =
also turn towards other people in the<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 local<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0room when they talk, which =
often changes the microphone<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 that picks<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0them up best. =C2=A0And oft=
en there are multiple simultaneous<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 talkers in<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0the room - we are not alway=
s polite. =C2=A0So even detecting the<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0approximate talker location=
(s) from the far end is<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 problematic. =C2=A0My<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0company&#39;s products can =
locate the talker(s) within the room<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0accurately, but they are us=
ing information that is not<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 available to<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0far-end systems, and which =
would be hard to send in a<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0product-independent way.<br=
>
<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0Personally I favor explicit=
 signaling from the sender on which<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0capture(s) carry the curren=
t talker(s). =C2=A0Middle boxes and<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 far end<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0systems can determine which=
 rooms have active talkers by<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 down-mixing<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0to mono. =C2=A0To get finer=
 granularity within the room, you<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 could then<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0use the explicit signaling.=
<br>
<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0The second aspect of the pr=
oblem is associating the talker<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 with a<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0video capture. =C2=A0Again,=
 this seems somewhat difficult,<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 especially if<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0the audio captures are not =
precisely aligned to the video<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 captures.<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0We can probably improve thi=
s - but if we need explicit talker<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0signaling anyway, it could =
potentially identify the best video<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0capture. In fact, since I p=
ersonally think that whole-room<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 audio<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0rendering should always be =
done, I am much more interested<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 in the<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0video capture than the audi=
o capture.<br>
<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0Other approaches might work=
, the main point I am making<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 here is that<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0remote identification of ta=
lker locations in the room is<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 probably<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0broken.<br>
<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 -Partial Room Rendering will not always work we=
ll.<br>
<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0If you are selecting some a=
udio captures from a room, and<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 discarding<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0others, you might not get n=
atural sounding audio. =C2=A0How good it<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0sounds will depend on the d=
etails of the microphone pickup<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 in the<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0far end room, how controlle=
d participant seating is, and which<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0captures you happen to disc=
ard.<br>
<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0Since the goal for CLUE is =
interoperability across<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 disparate room<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0designs, we can&#39;t make =
a lot of assumptions about<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 microphone pickup.<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0I think we can assum=
e that the full room audio will sound<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0appropriate - if it doesn&#=
39;t, then that means that the<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 sender doesn&#39;t<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0have a well-designed system=
. =C2=A0But once you start filtering out<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0captures in the middle, the=
 results will vary. =C2=A0I&#39;m not<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 seeing much<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0CLUE can do about that, oth=
er than point it out. =C2=A0As noted<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 above, I<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0think that rendering the en=
tire room&#39;s transmitted sound<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 field is<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0going to sound the best. =
=C2=A0If there is a need for more<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 conditioning<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0(say noise reduction) it is=
 probably better to do it<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 locally at the<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0sender.<br>
<br>
<br>
<br>
<br>
<br>
<br>
<br>
<br></div></div>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 ______________________________<u></u>__________=
_________<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 clue mailing list<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 <a href=3D"mailto:clue@ietf.org" target=3D"_bla=
nk">clue@ietf.org</a> &lt;mailto:<a href=3D"mailto:clue@ietf.org" target=3D=
"_blank">clue@ietf.org</a>&gt;<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 <a href=3D"https://www.ietf.org/mailman/__listi=
nfo/clue" target=3D"_blank">https://www.ietf.org/mailman/_<u></u>_listinfo/=
clue</a><br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 &lt;<a href=3D"https://www.ietf.org/mailman/lis=
tinfo/clue" target=3D"_blank">https://www.ietf.org/mailman/<u></u>listinfo/=
clue</a>&gt;<br>
<br>
<br>
=C2=A0 =C2=A0 ______________________________<u></u>___________________<br>
=C2=A0 =C2=A0 clue mailing list<br>
=C2=A0 =C2=A0 <a href=3D"mailto:clue@ietf.org" target=3D"_blank">clue@ietf.=
org</a> &lt;mailto:<a href=3D"mailto:clue@ietf.org" target=3D"_blank">clue@=
ietf.org</a>&gt;<br>
=C2=A0 =C2=A0 <a href=3D"https://www.ietf.org/mailman/__listinfo/clue" targ=
et=3D"_blank">https://www.ietf.org/mailman/_<u></u>_listinfo/clue</a><br>
=C2=A0 =C2=A0 &lt;<a href=3D"https://www.ietf.org/mailman/listinfo/clue" ta=
rget=3D"_blank">https://www.ietf.org/mailman/<u></u>listinfo/clue</a>&gt;<b=
r>
<br>
<br>
</blockquote>
<br>
</blockquote></div><br></div></div>

--047d7b675e6653986b04f833f988--


From nobody Tue Apr 29 16:50:38 2014
Return-Path: <coverdale@sympatico.ca>
X-Original-To: clue@ietfa.amsl.com
Delivered-To: clue@ietfa.amsl.com
Received: from localhost (ietfa.amsl.com [127.0.0.1]) by ietfa.amsl.com (Postfix) with ESMTP id 7ED7E1A0989 for <clue@ietfa.amsl.com>; Tue, 29 Apr 2014 16:50:36 -0700 (PDT)
X-Virus-Scanned: amavisd-new at amsl.com
X-Spam-Flag: NO
X-Spam-Score: -1.899
X-Spam-Level: 
X-Spam-Status: No, score=-1.899 tagged_above=-999 required=5 tests=[BAYES_00=-1.9, HTML_MESSAGE=0.001, MSGID_FROM_MTA_HEADER=0.001, RCVD_IN_DNSWL_NONE=-0.0001, SPF_PASS=-0.001] autolearn=ham
Received: from mail.ietf.org ([4.31.198.44]) by localhost (ietfa.amsl.com [127.0.0.1]) (amavisd-new, port 10024) with ESMTP id bQgsN4t1okCW for <clue@ietfa.amsl.com>; Tue, 29 Apr 2014 16:50:34 -0700 (PDT)
Received: from blu0-omc1-s29.blu0.hotmail.com (blu0-omc1-s29.blu0.hotmail.com [65.55.116.40]) by ietfa.amsl.com (Postfix) with ESMTP id 5A4B61A096B for <clue@ietf.org>; Tue, 29 Apr 2014 16:50:34 -0700 (PDT)
Received: from BLU0-SMTP21 ([65.55.116.7]) by blu0-omc1-s29.blu0.hotmail.com with Microsoft SMTPSVC(6.0.3790.4675);  Tue, 29 Apr 2014 16:50:32 -0700
X-TMN: [V6bnAekJ91ZL8qMkTxSw9X05ULYl/ied]
X-Originating-Email: [coverdale@sympatico.ca]
Message-ID: <BLU0-SMTP215C180D3356A66A02AD04D0460@phx.gbl>
Received: from PaulNewPC ([184.147.38.66]) by BLU0-SMTP21.phx.gbl over TLS secured channel with Microsoft SMTPSVC(6.0.3790.4675);  Tue, 29 Apr 2014 16:50:32 -0700
From: Paul Coverdale <coverdale@sympatico.ca>
To: "'Stephen Botzko'" <stephen.botzko@gmail.com>, "'Paul Kyzivat'" <pkyzivat@alum.mit.edu>
References: <533AF351.9050201@alum.mit.edu> <20140410222705.GW39240@verdi> <BLU0-SMTP6910DA20C9A9287137B585D0550@phx.gbl> <53475572.4040807@nteczone.com> <20140411150055.GE60844@verdi> <49E45C59CA48264997FEBFB29B6BC2D6215295A8AE@CRPMBOXPRD07.polycom.com> <BLU0-SMTP1804A70E3E0F666EAEAE84D0540@phx.gbl> <49E45C59CA48264997FEBFB29B6BC2D62409219636@CRPMBOXPRD07.polycom.com> <534B4303.7060707@nteczone.com> <5C4AC54BFF7A0842A6A11F554D6FB52F07152B@CRPMBOXPRD08.polycom.com> <534B536B.8000205@nteczone.com> <5C4AC54BFF7A0842A6A11F554D6FB52F13D9B2@CRPMBOXPRD08.polycom.com> <535F0B3A.70808@nteczone.com> <535F262C.6090009@alum.mit.edu> <535F73B6.9030106@nteczone.com> <CAMC7SJ7ZNW+=PDB-39gQ3MkQpTh5W1LZLPz7Y8VaFc7OofRviw@mail.gmail.com> <535FD82A.80708@alum.mit.edu> <CAMC7SJ6TGUYQVNcf__At7rkwgBQkoYBXAOfmnnLPW56=-=96wA@mail.gmail.com> <5360010A.1000904@alum.mit.edu> <CAMC7SJ6ARW5Z2AX526AiabO1AMi5qd0C8uMEc63PRdHWv_b8Kw@mail.gmail.com>
In-Reply-To: <CAMC7SJ6ARW5Z2AX526AiabO1AMi5qd0C8uMEc63PRdHWv_b8Kw@mail.gmail.com>
Date: Tue, 29 Apr 2014 19:50:27 -0400
MIME-Version: 1.0
Content-Type: multipart/alternative; boundary="----=_NextPart_000_0103_01CF63E4.4584AD10"
X-Mailer: Microsoft Office Outlook 12.0
Thread-Index: Ac9j5pdbyL55yfVmTsqiez//ziMhigAG3WgQ
Content-Language: en-us
X-OriginalArrivalTime: 29 Apr 2014 23:50:32.0241 (UTC) FILETIME=[CEB1DA10:01CF6405]
Archived-At: http://mailarchive.ietf.org/arch/msg/clue/6_I1HuPLcAJS1HGJDmixY8Ks6uY
Cc: 'CLUE' <clue@ietf.org>
Subject: Re: [clue] Improving treatment of audio
X-BeenThere: clue@ietf.org
X-Mailman-Version: 2.1.15
Precedence: list
List-Id: CLUE - ControLling mUltiple streams for TElepresence <clue.ietf.org>
List-Unsubscribe: <https://www.ietf.org/mailman/options/clue>, <mailto:clue-request@ietf.org?subject=unsubscribe>
List-Archive: <http://www.ietf.org/mail-archive/web/clue/>
List-Post: <mailto:clue@ietf.org>
List-Help: <mailto:clue-request@ietf.org?subject=help>
List-Subscribe: <https://www.ietf.org/mailman/listinfo/clue>, <mailto:clue-request@ietf.org?subject=subscribe>
X-List-Received-Date: Tue, 29 Apr 2014 23:50:36 -0000

------=_NextPart_000_0103_01CF63E4.4584AD10
Content-Type: text/plain; charset="utf-8"
Content-Transfer-Encoding: quoted-printable

=20

On Tue, Apr 29, 2014 at 3:44 PM, Paul Kyzivat <pkyzivat@alum.mit.edu> =
wrote:

I realize this is probably not a topic that ought to be normative. But =
ISTM that *something* should be said about it. I think the advertiser =
and the consumer need to be in agreement on what this information means =
or how it is intended to be used in order to get interoperability. There =
seems to be no one obvious way to pick an Area of Capture for a =
microphone, but different decisions about what to choose may well effect =
the way the consumer's algorithm works.

=20

=20

I agree that some information is needed. Perhaps continuing to drill =
down on some of these cases will help clarify what needs to be said.

=20

>From what you describe, perhaps there is no value in area of capture =
for audio, and axis of capture would be more useful. And I assume that =
in the absence of axis of capture, one may approximate it with a line =
through the center of the area of capture and perpendicular to it.

=20

=20

Axis of capture could be more useful, though I am not aware of any =
techniques that would let you align your output audio along a different =
axis from the loudspeaker's natural orientation.=20

=20

[Paul Coverdale]  I=E2=80=99m not aware of such a technique either, but =
are you saying that the idea is that the axis of capture might be used =
to align the output audio at the receiving side to where the microphone =
at the sending side is located? Is this useful? What happens if this =
microphone is not in a very good location anyway? It seems to me that =
the =E2=80=9Cbest=E2=80=9D audio rendering direction at the receiving =
side would be perpendicular to the display, for a given video capture. =
In other words, microphone types and their physical location should be a =
matter for the sending side only, the only criterion being that they =
should be able to be associated with any possible video capture in that =
room, and provide a high signal to noise ratio for sending to the =
receiving side. This would cover the use of a single omni, multiple =
directional, headsets, beam-steering arrays etc.

=20

=20

=20

=20


------=_NextPart_000_0103_01CF63E4.4584AD10
Content-Type: text/html; charset="utf-8"
Content-Transfer-Encoding: quoted-printable

<html xmlns:v=3D"urn:schemas-microsoft-com:vml" =
xmlns:o=3D"urn:schemas-microsoft-com:office:office" =
xmlns:w=3D"urn:schemas-microsoft-com:office:word" =
xmlns:m=3D"http://schemas.microsoft.com/office/2004/12/omml" =
xmlns=3D"http://www.w3.org/TR/REC-html40"><head><meta =
http-equiv=3DContent-Type content=3D"text/html; charset=3Dutf-8"><meta =
name=3DGenerator content=3D"Microsoft Word 12 (filtered =
medium)"><style><!--
/* Font Definitions */
@font-face
	{font-family:"Cambria Math";
	panose-1:2 4 5 3 5 4 6 3 2 4;}
@font-face
	{font-family:Calibri;
	panose-1:2 15 5 2 2 2 4 3 2 4;}
/* Style Definitions */
p.MsoNormal, li.MsoNormal, div.MsoNormal
	{margin:0in;
	margin-bottom:.0001pt;
	font-size:12.0pt;
	font-family:"Times New Roman","serif";}
a:link, span.MsoHyperlink
	{mso-style-priority:99;
	color:blue;
	text-decoration:underline;}
a:visited, span.MsoHyperlinkFollowed
	{mso-style-priority:99;
	color:purple;
	text-decoration:underline;}
span.EmailStyle17
	{mso-style-type:personal-reply;
	font-family:"Calibri","sans-serif";
	color:#1F497D;}
.MsoChpDefault
	{mso-style-type:export-only;}
@page WordSection1
	{size:8.5in 11.0in;
	margin:1.0in 1.0in 1.0in 1.0in;}
div.WordSection1
	{page:WordSection1;}
--></style><!--[if gte mso 9]><xml>
<o:shapedefaults v:ext=3D"edit" spidmax=3D"1026" />
</xml><![endif]--><!--[if gte mso 9]><xml>
<o:shapelayout v:ext=3D"edit">
<o:idmap v:ext=3D"edit" data=3D"1" />
</o:shapelayout></xml><![endif]--></head><body lang=3DEN-US link=3Dblue =
vlink=3Dpurple><div class=3DWordSection1><div =
style=3D'border:none;border-left:solid blue 1.5pt;padding:0in 0in 0in =
4.0pt'><div><div><p class=3DMsoNormal =
style=3D'margin-bottom:12.0pt'><o:p>&nbsp;</o:p></p><div><p =
class=3DMsoNormal>On Tue, Apr 29, 2014 at 3:44 PM, Paul Kyzivat &lt;<a =
href=3D"mailto:pkyzivat@alum.mit.edu" =
target=3D"_blank">pkyzivat@alum.mit.edu</a>&gt; wrote:<o:p></o:p></p><p =
class=3DMsoNormal>I realize this is probably not a topic that ought to =
be normative. But ISTM that *something* should be said about it. I think =
the advertiser and the consumer need to be in agreement on what this =
information means or how it is intended to be used in order to get =
interoperability. There seems to be no one obvious way to pick an Area =
of Capture for a microphone, but different decisions about what to =
choose may well effect the way the consumer's algorithm =
works.<o:p></o:p></p><div><p class=3DMsoNormal><span =
style=3D'color:#1F497D'><o:p>&nbsp;</o:p></span></p><p =
class=3DMsoNormal><b><i><span =
style=3D'font-size:11.0pt;font-family:"Calibri","sans-serif";color:#1F497=
D'><o:p>&nbsp;</o:p></span></i></b></p><p class=3DMsoNormal>I agree that =
some information is needed. Perhaps continuing to drill down on some of =
these cases will help clarify what needs to be =
said.<o:p></o:p></p></div><blockquote =
style=3D'border:none;border-left:solid #CCCCCC 1.0pt;padding:0in 0in 0in =
6.0pt;margin-left:4.8pt;margin-right:0in'><p =
class=3DMsoNormal>&nbsp;<o:p></o:p></p></blockquote><blockquote =
style=3D'border:none;border-left:solid #CCCCCC 1.0pt;padding:0in 0in 0in =
6.0pt;margin-left:4.8pt;margin-right:0in'><p class=3DMsoNormal>&gt;From =
what you describe, perhaps there is no value in area of capture for =
audio, and axis of capture would be more useful. And I assume that in =
the absence of axis of capture, one may approximate it with a line =
through the center of the area of capture and perpendicular to =
it.<o:p></o:p></p></blockquote><div><p class=3DMsoNormal><span =
style=3D'color:#1F497D'><o:p>&nbsp;</o:p></span></p><p =
class=3DMsoNormal><b><i><span =
style=3D'font-size:11.0pt;font-family:"Calibri","sans-serif";color:#1F497=
D'><o:p>&nbsp;</o:p></span></i></b></p><p class=3DMsoNormal>Axis of =
capture could be more useful, though I am not aware of any techniques =
that would let you align your output audio along a different axis from =
the loudspeaker's natural orientation.&nbsp;<o:p></o:p></p></div><div><p =
class=3DMsoNormal><o:p>&nbsp;</o:p></p></div><blockquote =
style=3D'border:none;border-left:solid #CCCCCC 1.0pt;padding:0in 0in 0in =
6.0pt;margin-left:4.8pt;margin-right:0in'><p =
class=3DMsoNormal><b><i><span style=3D'color:#1F497D'>[Paul Coverdale] =
</span></i></b><span style=3D'color:#1F497D'>=C2=A0I=E2=80=99m not aware =
of such a technique either, but are you saying that the idea is that the =
axis of capture might be used to align the output audio at the receiving =
side to where the microphone at the sending side is located? Is this =
useful? What happens if this microphone is not in a very good location =
anyway? It seems to me that the =E2=80=9Cbest=E2=80=9D audio rendering =
direction at the receiving side would be perpendicular to the display, =
for a given video capture. In other words, microphone types and their =
physical location should be a matter for the sending side only, the only =
criterion being that they should be able to be associated with any =
possible video capture in that room, and provide a high signal to noise =
ratio for sending to the receiving side. This would cover the use of a =
single omni, multiple directional, headsets, beam-steering arrays =
etc.<o:p></o:p></span></p><p class=3DMsoNormal><span =
style=3D'color:#1F497D'><o:p>&nbsp;</o:p></span></p><p class=3DMsoNormal =
style=3D'margin-left:16.8pt'><o:p>&nbsp;</o:p></p><p =
class=3DMsoNormal><o:p>&nbsp;</o:p></p></blockquote></div><p =
class=3DMsoNormal><o:p>&nbsp;</o:p></p></div></div></div></div></body></h=
tml>
------=_NextPart_000_0103_01CF63E4.4584AD10--


From nobody Tue Apr 29 18:51:55 2014
Return-Path: <stephen.botzko@gmail.com>
X-Original-To: clue@ietfa.amsl.com
Delivered-To: clue@ietfa.amsl.com
Received: from localhost (ietfa.amsl.com [127.0.0.1]) by ietfa.amsl.com (Postfix) with ESMTP id 6760D1A09D1 for <clue@ietfa.amsl.com>; Tue, 29 Apr 2014 18:51:47 -0700 (PDT)
X-Virus-Scanned: amavisd-new at amsl.com
X-Spam-Flag: NO
X-Spam-Score: -1.999
X-Spam-Level: 
X-Spam-Status: No, score=-1.999 tagged_above=-999 required=5 tests=[BAYES_00=-1.9, DKIM_SIGNED=0.1, DKIM_VALID=-0.1, DKIM_VALID_AU=-0.1, FREEMAIL_FROM=0.001, HTML_MESSAGE=0.001, SPF_PASS=-0.001] autolearn=ham
Received: from mail.ietf.org ([4.31.198.44]) by localhost (ietfa.amsl.com [127.0.0.1]) (amavisd-new, port 10024) with ESMTP id m_3M9kIKPNUO for <clue@ietfa.amsl.com>; Tue, 29 Apr 2014 18:51:46 -0700 (PDT)
Received: from mail-ve0-x234.google.com (mail-ve0-x234.google.com [IPv6:2607:f8b0:400c:c01::234]) by ietfa.amsl.com (Postfix) with ESMTP id C09BF1A092F for <clue@ietf.org>; Tue, 29 Apr 2014 18:51:45 -0700 (PDT)
Received: by mail-ve0-f180.google.com with SMTP id jz11so1292511veb.25 for <clue@ietf.org>; Tue, 29 Apr 2014 18:51:44 -0700 (PDT)
DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20120113;  h=mime-version:in-reply-to:references:date:message-id:subject:from:to :cc:content-type; bh=ueteuojXtb0CLj+jdgf3Rpp7PEk5oVCS0+YtLaKFS2E=; b=FezTZQECkWXX4DvQ+qII4UVZf6JqPl3t8E8yYZGahNjNKTKxW80yD2udx+o5bUttj6 d6Qr/tXytWTtoidOWX7dxGHeJi6f+8ye8AIyPabHho7H5KK+OwVUQ/kackPASwuuZPaF D1RIN3mh+OCUBN/cOZbaQAV+WWtqYbho7pbO5f/6y0rQG0/q4nHtedqwBwsGJBzwi2f/ X1waPiS0+AmkQNTIXyT32kzTjZr6BLrdzBJSTq9wDWQg7H1mRbjKOCQY2hke3hKNPaN6 8kt+zMQR4C+lwdO6RmOV3Nerm1stSLKD7b0WK4HvQZgkoraw+FQzwwisewx6dbKfwAOY rs+Q==
MIME-Version: 1.0
X-Received: by 10.221.26.10 with SMTP id rk10mr1195711vcb.0.1398822704246; Tue, 29 Apr 2014 18:51:44 -0700 (PDT)
Received: by 10.221.40.135 with HTTP; Tue, 29 Apr 2014 18:51:44 -0700 (PDT)
In-Reply-To: <BLU0-SMTP215C180D3356A66A02AD04D0460@phx.gbl>
References: <533AF351.9050201@alum.mit.edu> <20140410222705.GW39240@verdi> <BLU0-SMTP6910DA20C9A9287137B585D0550@phx.gbl> <53475572.4040807@nteczone.com> <20140411150055.GE60844@verdi> <49E45C59CA48264997FEBFB29B6BC2D6215295A8AE@CRPMBOXPRD07.polycom.com> <BLU0-SMTP1804A70E3E0F666EAEAE84D0540@phx.gbl> <49E45C59CA48264997FEBFB29B6BC2D62409219636@CRPMBOXPRD07.polycom.com> <534B4303.7060707@nteczone.com> <5C4AC54BFF7A0842A6A11F554D6FB52F07152B@CRPMBOXPRD08.polycom.com> <534B536B.8000205@nteczone.com> <5C4AC54BFF7A0842A6A11F554D6FB52F13D9B2@CRPMBOXPRD08.polycom.com> <535F0B3A.70808@nteczone.com> <535F262C.6090009@alum.mit.edu> <535F73B6.9030106@nteczone.com> <CAMC7SJ7ZNW+=PDB-39gQ3MkQpTh5W1LZLPz7Y8VaFc7OofRviw@mail.gmail.com> <535FD82A.80708@alum.mit.edu> <CAMC7SJ6TGUYQVNcf__At7rkwgBQkoYBXAOfmnnLPW56=-=96wA@mail.gmail.com> <5360010A.1000904@alum.mit.edu> <CAMC7SJ6ARW5Z2AX526AiabO1AMi5qd0C8uMEc63PRdHWv_b8Kw@mail.gmail.com> <BLU0-SMTP215C180D3356A66A02AD04D0460@phx.gbl>
Date: Tue, 29 Apr 2014 21:51:44 -0400
Message-ID: <CAMC7SJ6VDbOegVayPVa7zoOZ6Q7t1hUg=dCtwSNkyaUzxdjKoA@mail.gmail.com>
From: Stephen Botzko <stephen.botzko@gmail.com>
To: Paul Coverdale <coverdale@sympatico.ca>
Content-Type: multipart/alternative; boundary=001a11339ae463f3ca04f838ca1b
Archived-At: http://mailarchive.ietf.org/arch/msg/clue/Ut3GldW2p74patULcSPjUHHbTvA
Cc: CLUE <clue@ietf.org>
Subject: Re: [clue] Improving treatment of audio
X-BeenThere: clue@ietf.org
X-Mailman-Version: 2.1.15
Precedence: list
List-Id: CLUE - ControLling mUltiple streams for TElepresence <clue.ietf.org>
List-Unsubscribe: <https://www.ietf.org/mailman/options/clue>, <mailto:clue-request@ietf.org?subject=unsubscribe>
List-Archive: <http://www.ietf.org/mail-archive/web/clue/>
List-Post: <mailto:clue@ietf.org>
List-Help: <mailto:clue-request@ietf.org?subject=help>
List-Subscribe: <https://www.ietf.org/mailman/listinfo/clue>, <mailto:clue-request@ietf.org?subject=subscribe>
X-List-Received-Date: Wed, 30 Apr 2014 01:51:47 -0000

--001a11339ae463f3ca04f838ca1b
Content-Type: text/plain; charset=UTF-8
Content-Transfer-Encoding: quoted-printable

On Tue, Apr 29, 2014 at 7:50 PM, Paul Coverdale <coverdale@sympatico.ca>wro=
te:

>
>
> On Tue, Apr 29, 2014 at 3:44 PM, Paul Kyzivat <pkyzivat@alum.mit.edu>
> wrote:
>
> I realize this is probably not a topic that ought to be normative. But
> ISTM that *something* should be said about it. I think the advertiser and
> the consumer need to be in agreement on what this information means or ho=
w
> it is intended to be used in order to get interoperability. There seems t=
o
> be no one obvious way to pick an Area of Capture for a microphone, but
> different decisions about what to choose may well effect the way the
> consumer's algorithm works.
>
>
>
>
>
> I agree that some information is needed. Perhaps continuing to drill down
> on some of these cases will help clarify what needs to be said.
>
>
>
> >From what you describe, perhaps there is no value in area of capture for
> audio, and axis of capture would be more useful. And I assume that in the
> absence of axis of capture, one may approximate it with a line through th=
e
> center of the area of capture and perpendicular to it.
>
>
>
>
>
> Axis of capture could be more useful, though I am not aware of any
> techniques that would let you align your output audio along a different
> axis from the loudspeaker's natural orientation.
>
>
>
> *[Paul Coverdale] * I=E2=80=99m not aware of such a technique either, but=
 are you
> saying that the idea is that the axis of capture might be used to align t=
he
> output audio at the receiving side to where the microphone at the sending
> side is located? Is this useful? What happens if this microphone is not i=
n
> a very good location anyway? It seems to me that the =E2=80=9Cbest=E2=80=
=9D audio rendering
> direction at the receiving side would be perpendicular to the display, fo=
r
> a given video capture. In other words, microphone types and their physica=
l
> location should be a matter for the sending side only, the only criterion
> being that they should be able to be associated with any possible video
> capture in that room, and provide a high signal to noise ratio for sendin=
g
> to the receiving side. This would cover the use of a single omni, multipl=
e
> directional, headsets, beam-steering arrays etc.
>
>
>
>
>
>   I was thinking that if you had a way to align the audio capture axis in
sender and receiver, that you would have the audio analog of a
line-of-sight.  I don't see any way to use the audio capture axis
information without such a method.

I agree that for most systems the rendering axis of the speakers will be
perpendicular to their displays, and that it would be sensible for senders
to take that into account when they design their audio captures.

I also agree that microphone types and physical location need be a matter
for the sending side only (and that loudspeaker types and loudspeaker
locations also need to be a matter for the receiving side only).  There is
a price for that design constraint - if you have full control over the both
sender and receiver equipment you can design a more convincing sense of
presence (both for audio and for video).  But that's a price worth paying
to get broad interoperability.

>
>
>
>

--001a11339ae463f3ca04f838ca1b
Content-Type: text/html; charset=UTF-8
Content-Transfer-Encoding: quoted-printable

<div dir=3D"ltr"><br><div class=3D"gmail_extra"><br><br><div class=3D"gmail=
_quote">On Tue, Apr 29, 2014 at 7:50 PM, Paul Coverdale <span dir=3D"ltr">&=
lt;<a href=3D"mailto:coverdale@sympatico.ca" target=3D"_blank">coverdale@sy=
mpatico.ca</a>&gt;</span> wrote:<br>
<blockquote class=3D"gmail_quote" style=3D"margin:0px 0px 0px 0.8ex;border-=
left-width:1px;border-left-color:rgb(204,204,204);border-left-style:solid;p=
adding-left:1ex"><div lang=3D"EN-US" link=3D"blue" vlink=3D"purple"><div><d=
iv style=3D"border-style:none none none solid;border-left-color:blue;border=
-left-width:1.5pt;padding:0in 0in 0in 4pt">
<div><div><p class=3D"MsoNormal" style=3D"margin-bottom:12pt"><u></u>=C2=A0=
<u></u></p><div><div class=3D""><p class=3D"MsoNormal">On Tue, Apr 29, 2014=
 at 3:44 PM, Paul Kyzivat &lt;<a href=3D"mailto:pkyzivat@alum.mit.edu" targ=
et=3D"_blank">pkyzivat@alum.mit.edu</a>&gt; wrote:<u></u><u></u></p>
<p class=3D"MsoNormal">I realize this is probably not a topic that ought to=
 be normative. But ISTM that *something* should be said about it. I think t=
he advertiser and the consumer need to be in agreement on what this informa=
tion means or how it is intended to be used in order to get interoperabilit=
y. There seems to be no one obvious way to pick an Area of Capture for a mi=
crophone, but different decisions about what to choose may well effect the =
way the consumer&#39;s algorithm works.<u></u><u></u></p>
<div><p class=3D"MsoNormal"><span style=3D"color:rgb(31,73,125)"><u></u>=C2=
=A0<u></u></span></p><p class=3D"MsoNormal"><b><i><span style=3D"font-size:=
11pt;font-family:Calibri,sans-serif;color:rgb(31,73,125)"><u></u>=C2=A0<u><=
/u></span></i></b></p>
<p class=3D"MsoNormal">I agree that some information is needed. Perhaps con=
tinuing to drill down on some of these cases will help clarify what needs t=
o be said.<u></u><u></u></p></div><blockquote style=3D"border-style:none no=
ne none solid;border-left-color:rgb(204,204,204);border-left-width:1pt;padd=
ing:0in 0in 0in 6pt;margin-left:4.8pt;margin-right:0in">
<p class=3D"MsoNormal">=C2=A0<u></u><u></u></p></blockquote><blockquote sty=
le=3D"border-style:none none none solid;border-left-color:rgb(204,204,204);=
border-left-width:1pt;padding:0in 0in 0in 6pt;margin-left:4.8pt;margin-righ=
t:0in">
<p class=3D"MsoNormal">&gt;From what you describe, perhaps there is no valu=
e in area of capture for audio, and axis of capture would be more useful. A=
nd I assume that in the absence of axis of capture, one may approximate it =
with a line through the center of the area of capture and perpendicular to =
it.<u></u><u></u></p>
</blockquote><div><p class=3D"MsoNormal"><span style=3D"color:rgb(31,73,125=
)"><u></u>=C2=A0<u></u></span></p><p class=3D"MsoNormal"><b><i><span style=
=3D"font-size:11pt;font-family:Calibri,sans-serif;color:rgb(31,73,125)"><u>=
</u>=C2=A0<u></u></span></i></b></p>
<p class=3D"MsoNormal">Axis of capture could be more useful, though I am no=
t aware of any techniques that would let you align your output audio along =
a different axis from the loudspeaker&#39;s natural orientation.=C2=A0<u></=
u><u></u></p>
</div><div><p class=3D"MsoNormal"><u></u>=C2=A0<u></u></p></div></div><bloc=
kquote style=3D"border-style:none none none solid;border-left-color:rgb(204=
,204,204);border-left-width:1pt;padding:0in 0in 0in 6pt;margin-left:4.8pt;m=
argin-right:0in">
<p class=3D"MsoNormal"><b><i><span style=3D"color:rgb(31,73,125)">[Paul Cov=
erdale] </span></i></b><span style=3D"color:rgb(31,73,125)">=C2=A0I=E2=80=
=99m not aware of such a technique either, but are you saying that the idea=
 is that the axis of capture might be used to align the output audio at the=
 receiving side to where the microphone at the sending side is located? Is =
this useful? What happens if this microphone is not in a very good location=
 anyway? It seems to me that the =E2=80=9Cbest=E2=80=9D audio rendering dir=
ection at the receiving side would be perpendicular to the display, for a g=
iven video capture. In other words, microphone types and their physical loc=
ation should be a matter for the sending side only, the only criterion bein=
g that they should be able to be associated with any possible video capture=
 in that room, and provide a high signal to noise ratio for sending to the =
receiving side. This would cover the use of a single omni, multiple directi=
onal, headsets, beam-steering arrays etc.<u></u><u></u></span></p>
<p class=3D"MsoNormal"><span style=3D"color:rgb(31,73,125)"><u></u>=C2=A0<u=
></u></span></p><p class=3D"MsoNormal" style=3D"margin-left:16.8pt"><u></u>=
=C2=A0<u></u></p></blockquote><blockquote style=3D"border-style:none none n=
one solid;border-left-color:rgb(204,204,204);border-left-width:1pt;padding:=
0in 0in 0in 6pt;margin-left:4.8pt;margin-right:0in">
<p class=3D"MsoNormal"><u></u></p></blockquote></div></div></div></div></di=
v></div></blockquote><div>=C2=A0 I was thinking that if you had a way to al=
ign the audio capture axis in sender and receiver, that you would have the =
audio analog of a line-of-sight. =C2=A0I don&#39;t see any way to use the a=
udio capture axis information without such a method.</div>
<div><br></div><div>I agree that for most systems the rendering axis of the=
 speakers will be perpendicular to their displays, and that it would be sen=
sible for senders to take that into account when they design their audio ca=
ptures. =C2=A0</div>
<div><br></div><div>I also agree that microphone types and physical locatio=
n need be a matter for the sending side only (and that loudspeaker types an=
d loudspeaker locations also need to be a matter for the receiving side onl=
y). =C2=A0There is a price for that design constraint - if you have full co=
ntrol over the both sender and receiver equipment you can design a more con=
vincing sense of presence (both for audio and for video). =C2=A0But that&#3=
9;s a price worth paying to get broad interoperability.</div>
<blockquote class=3D"gmail_quote" style=3D"margin:0px 0px 0px 0.8ex;border-=
left-width:1px;border-left-color:rgb(204,204,204);border-left-style:solid;p=
adding-left:1ex"><div lang=3D"EN-US" link=3D"blue" vlink=3D"purple"><div><d=
iv style=3D"border-style:none none none solid;border-left-color:blue;border=
-left-width:1.5pt;padding:0in 0in 0in 4pt">
<div><div><div><blockquote style=3D"border-style:none none none solid;borde=
r-left-color:rgb(204,204,204);border-left-width:1pt;padding:0in 0in 0in 6pt=
;margin-left:4.8pt;margin-right:0in"><p class=3D"MsoNormal">=C2=A0<u></u></=
p></blockquote>
</div><p class=3D"MsoNormal"><u></u>=C2=A0<u></u></p></div></div></div></di=
v></div></blockquote></div><br></div></div>

--001a11339ae463f3ca04f838ca1b--


From nobody Wed Apr 30 06:28:43 2014
Return-Path: <Mark.Duckworth@polycom.com>
X-Original-To: clue@ietfa.amsl.com
Delivered-To: clue@ietfa.amsl.com
Received: from localhost (ietfa.amsl.com [127.0.0.1]) by ietfa.amsl.com (Postfix) with ESMTP id 6F4041A0684 for <clue@ietfa.amsl.com>; Wed, 30 Apr 2014 06:28:41 -0700 (PDT)
X-Virus-Scanned: amavisd-new at amsl.com
X-Spam-Flag: NO
X-Spam-Score: -1.82
X-Spam-Level: 
X-Spam-Status: No, score=-1.82 tagged_above=-999 required=5 tests=[BAYES_00=-1.9, RCVD_IN_DNSWL_LOW=-0.7, SPF_NEUTRAL=0.779, UNPARSEABLE_RELAY=0.001] autolearn=ham
Received: from mail.ietf.org ([4.31.198.44]) by localhost (ietfa.amsl.com [127.0.0.1]) (amavisd-new, port 10024) with ESMTP id 7JTO6KRseSd6 for <clue@ietfa.amsl.com>; Wed, 30 Apr 2014 06:28:39 -0700 (PDT)
Received: from mail1.bemta7.messagelabs.com (mail1.bemta7.messagelabs.com [216.82.254.105]) by ietfa.amsl.com (Postfix) with ESMTP id 6576D1A08DF for <clue@ietf.org>; Wed, 30 Apr 2014 06:28:39 -0700 (PDT)
Received: from [216.82.254.20:27397] by server-9.bemta-7.messagelabs.com id A5/60-28934-58AF0635; Wed, 30 Apr 2014 13:28:37 +0000
X-Env-Sender: Mark.Duckworth@polycom.com
X-Msg-Ref: server-10.tower-47.messagelabs.com!1398864509!974447!2
X-Originating-IP: [140.242.64.158]
X-StarScan-Received: 
X-StarScan-Version: 6.11.3; banners=-,-,-
X-VirusChecked: Checked
Received: (qmail 5851 invoked from network); 30 Apr 2014 13:28:37 -0000
Received: from crpehubprd01.polycom.com (HELO Crpehubprd01.polycom.com) (140.242.64.158) by server-10.tower-47.messagelabs.com with AES128-SHA encrypted SMTP; 30 Apr 2014 13:28:37 -0000
Received: from PWEHUB01.polycom.com (10.236.2.221) by Crpehubprd01.polycom.com (10.236.0.158) with Microsoft SMTP Server (TLS) id 8.3.192.1; Wed, 30 Apr 2014 06:28:36 -0700
Received: from CRPMBOXPRD08.polycom.com ([169.254.1.94]) by PWEHUB01.polycom.com ([fe80::99a8:f785:3f0c:2bb6%17]) with mapi; Wed, 30 Apr 2014 06:28:36 -0700
From: "Duckworth, Mark" <Mark.Duckworth@polycom.com>
To: Christian Groves <Christian.Groves@nteczone.com>, "clue@ietf.org" <clue@ietf.org>
Date: Wed, 30 Apr 2014 06:28:33 -0700
Thread-Topic: [clue] Improving treatment of audio
Thread-Index: Ac9jUOfTrGNIJ7ehSEm2XCjQ6gPaMgBIatYg
Message-ID: <5C4AC54BFF7A0842A6A11F554D6FB52F13DAE6@CRPMBOXPRD08.polycom.com>
References: <533AF351.9050201@alum.mit.edu> <20140410222705.GW39240@verdi> <BLU0-SMTP6910DA20C9A9287137B585D0550@phx.gbl> <53475572.4040807@nteczone.com> <20140411150055.GE60844@verdi> <49E45C59CA48264997FEBFB29B6BC2D6215295A8AE@CRPMBOXPRD07.polycom.com> <BLU0-SMTP1804A70E3E0F666EAEAE84D0540@phx.gbl> <49E45C59CA48264997FEBFB29B6BC2D62409219636@CRPMBOXPRD07.polycom.com> <534B4303.7060707@nteczone.com> <5C4AC54BFF7A0842A6A11F554D6FB52F07152B@CRPMBOXPRD08.polycom.com> <534B536B.8000205@nteczone.com> <5C4AC54BFF7A0842A6A11F554D6FB52F13D9B2@CRPMBOXPRD08.polycom.com> <535F0B3A.70808@nteczone.com>
In-Reply-To: <535F0B3A.70808@nteczone.com>
Accept-Language: en-US
Content-Language: en-US
X-MS-Has-Attach: 
X-MS-TNEF-Correlator: 
acceptlanguage: en-US
Content-Type: text/plain; charset="us-ascii"
Content-Transfer-Encoding: quoted-printable
MIME-Version: 1.0
Archived-At: http://mailarchive.ietf.org/arch/msg/clue/xZrZ68bQ1dwhV82q64bweNHi8sc
Subject: Re: [clue] Improving treatment of audio
X-BeenThere: clue@ietf.org
X-Mailman-Version: 2.1.15
Precedence: list
List-Id: CLUE - ControLling mUltiple streams for TElepresence <clue.ietf.org>
List-Unsubscribe: <https://www.ietf.org/mailman/options/clue>, <mailto:clue-request@ietf.org?subject=unsubscribe>
List-Archive: <http://www.ietf.org/mail-archive/web/clue/>
List-Post: <mailto:clue@ietf.org>
List-Help: <mailto:clue-request@ietf.org?subject=help>
List-Subscribe: <https://www.ietf.org/mailman/listinfo/clue>, <mailto:clue-request@ietf.org?subject=subscribe>
X-List-Received-Date: Wed, 30 Apr 2014 13:28:41 -0000

Hi Christian,
Thanks for elaborating on the idea.  I have some more questions inline.
I see basically what you are trying to do, but I still think it doesn't wor=
k in practice any better than what we already have, as I explain below.
Mark

> -----Original Message-----
> From: Christian Groves [mailto:Christian.Groves@nteczone.com]
> Sent: Monday, April 28, 2014 10:15 PM
> To: Duckworth, Mark; clue@ietf.org
> Subject: Re: [clue] Improving treatment of audio
>=20
> Hello Mark,
>=20
> Please see below.
>=20
> Regards, Christian
>=20
> On 29/04/2014 11:36 AM, Duckworth, Mark wrote:
> > Hi Christian,
> > I'm coming back to this message, because I'm trying to understand your
> proposal.  You proposed "in each audio capture we say what VC it relates =
to."
> What does this mean?  What does it mean for an audio capture to relate to=
 a
> video capture?  What exactly is the producer telling the consumer?  Can a=
n
> audio capture relate to more than one video capture?
> [CNG] The meaning has largely the same semantic as if the ADV used the
> same capture area on a VC and AC. Basically its saying that the audio com=
es
> from the same "region" as what the video capture does. I think the
> difference is that its not trying to put a mathematical certainty about w=
hat
> that region is. Its a way for the producer to say to the consumer if you =
choose
> video capture A you probably want to choose audio capture B for a good
> experience.

[Duckworth, Mark] If that is the meaning, then would it be better for the p=
rovider to advertise a list of audio captures that are "related to" each vi=
deo capture?  Rather than the other way around?  Maybe I'm not understandin=
g the problem you are trying to solve.

> In terms of whether an audio capture can relate to more than
> one video capture yes I think so. You could have one audio capture for a
> room and three video captures.
> >
> > More comments inline below.
> >
> > Regards,
> > Mark
> >
> >> -----Original Message-----
> >> From: Christian Groves [mailto:Christian.Groves@nteczone.com]
> >> Sent: Sunday, April 13, 2014 11:18 PM
> >> To: Duckworth, Mark; clue@ietf.org
> >> Subject: Re: [clue] Improving treatment of audio
> >>
> >> Hello Mark,
> >>
> >> Sorry I didn't consider the entire 12.1.1. if I do that, according to
> >> that
> >> example:
> >>
> >> Video areas of capture:
> >>
> >>          bottom left    bottom right  top left         top right
> >>      VC0 (-2011,2850,0) (-673,3000,0) (-2011,2850,757) (-673,3000,757)
> >>      VC1 ( -673,3000,0) ( 673,3000,0) ( -673,3000,757) ( 673,3000,757)
> >>      VC2 (  673,3000,0) (2011,2850,0) (  673,3000,757) (2011,3000,757)
> >>      VC3 (-2011,2850,0) (2011,2850,0) (-2011,2850,757) (2011,3000,757)
> >>      VC4 (-2011,2850,0) (2011,2850,0) (-2011,2850,757) (2011,3000,757)
> >>      VC5 (-2011,2850,0) (2011,2850,0) (-2011,2850,757) (2011,3000,757)
> >>      VC6 none
> >>
> >> Areas of capture for audio (from 12.1.1):
> >>
> >>          bottom left    bottom right  top left         top right
> >>
> >>      AC0 (-2011,2850,0) (-673,3000,0) (-2011,2850,757) (-673,3000,757)
> >>      AC1 (  673,3000,0) (2011,2850,0) (  673,3000,757) (2011,3000,757)
> >>      AC2 ( -673,3000,0) ( 673,3000,0) ( -673,3000,757) ( 673,3000,757)
> >>      AC3 (-2011,2850,0) (2011,2850,0) (-2011,2850,757) (2011,3000,757)
> >>      AC4 none
> >>
> >> Using a reference rather than area of capture:
> >>       AC0 (VC0)
> >>       AC1 (VC2)
> >>       AC2 (VC1)
> >>       AC3 (VC3,VC4,VC5)
> >> I would think they are basically conveying the same information???
> > [Duckworth, Mark] Looking from the consumer point of view, in trying to
> understand the advertisement, what does the consumer know from the
> advertisement?  Suppose the consumer wants to get all three of VC0, VC1,
> VC2, but it only wants a single audio capture.  It should probably ask fo=
r AC3,
> but according to the advertisement AC3 doesn't "relate to" VC0, VC1, or V=
C2.
> [CNG] I agree. If the provider wanted to indicate that AC3 could also be =
used
> for VC0,VC1,VC2 it could advertise:
>       AC0 (VC0)
>       AC1 (VC2)
>       AC2 (VC1)
>       AC3 (VC0,VC1,VC2,VC3,VC4,VC5)

[Duckworth, Mark] I agree that  AC3 (the whole room audio) "relates to" all=
 those video captures.  But earlier you said "the audio comes from the same=
 'region' as what the video capture does."  So now do we have to modify tha=
t to say something like "the audio comes from the same region as the video,=
 or a subset of that region, or a superset, or an overlapping region"?  In =
this example, AC3 covers a region that is a superset of the VC0 region, a s=
uperset of the VC1 region, and a superset of the VC2 region.

> > [Duckworth, Mark] For a case with a different consumer, suppose the
> consumer wants the single video capture VC5, but it would like to receive
> and render spatial audio.  There is no audio capture with multiple channe=
ls,
> so it could ask for AC0, AC1, and AC2 if it knew how they were spatially
> "related to" VC5.  But according to the advertisement, AC0, AC1 and AC2
> aren't "related to" VC5 at all.
> [CNG] So in this case there is no AC3? In that case the advertisement cou=
ld
> look like:
>       AC0 (VC0,VC5)
>       AC1 (VC2,VC5)
>       AC2 (VC1,VC5)

[Duckworth, Mark] I didn't mean to remove AC3.  That would still be there i=
n this example.  So now since the provider forms it's advertisement without=
 knowing what the consumer wants to do, the provider has to give all this i=
nformation:
  AC0 relates to (VC0, VC3, VC4, VC5)
  AC1 relates to (VC2, VC3, VC4, VC5)
  AC2 relates to (VC1, VC3, VC4, VC5)
  AC3 relates to (VC0, VC1, VC2, VC3, VC4, VC5)
[Duckworth, Mark] Does this information really help the consumer?  I don't =
think it is any better than what we already have with area of capture.  Wit=
h my two different consumer examples here, the first consumer would simply =
pick AC3 because it appears by itself in a CSE, and the consumer would rend=
er it as "whole scene" audio.  It doesn't need any more information than th=
at.  The second consumer has a tougher job.  It would choose AC0, AC1, AC2 =
because they appear in a CSE.  The tricky part is figuring out how to rende=
r them spatially.  With your proposal, I think the consumer has to look at =
the list of "relates to" video, make assumptions about how each AC relates =
to those VCs in the list, assuming for example that AC0 is more focused on =
the region of VC0 rather than the regions of the others, and then associate=
 the VC0 area of capture as a subset of the VC5 area (because it is really =
receiving VC5) and then associate that with the AC0 rendering area.  It see=
ms to me it is more straightforward for the provider to associate AC0 direc=
tly with the same area of capture as VC0 in the first place, as we already =
have today.

[Duckworth, Mark] For another provider example, suppose the provider also o=
ffered mono audio captures AC5 and AC6, capturing the left and right halves=
 of the room.  How would your proposal handle that?  Would it be like this?
  AC5 relates to (VC0, VC1, VC3, VC4, VC5)
  AC6 relates to (VC1, VC2, VC3, VC4, VC5)

> Each audio capture could relate to a particular video or be part of the z=
oomed
> out video capture if there was no room audio.
> >
> >> Regards, Christian
> >>
> >> On 14/04/2014 12:26 PM, Duckworth, Mark wrote:
> >>> Hello Christian,
> >>> please see below.
> >>> Mark
> >>>
> >>>> -----Original Message-----
> >>>> From: clue [mailto:clue-bounces@ietf.org] On Behalf Of Christian
> >>>> Groves
> >>>> Sent: Sunday, April 13, 2014 10:08 PM
> >>>> To: clue@ietf.org
> >>>> Subject: Re: [clue] Improving treatment of audio
> >>>>
> >>>> Hello all,
> >>>>
> >>>> Perhaps the confusion is that some see Audio Capture area as a
> >>>> means to associate an Audio capture with a video capture as you've
> >>>> shown Mark's example. e.g. The audio spatial information isn't
> >>>> really used for any audio transformation other than associating it
> >>>> with a particular
> >> video stream.
> >>>> Whereas others were more thinking the audio spatial information as
> >>>> an input to mixing and more complicated audio processing.
> >>> [Duckworth, Mark] You could be right about this being a source of
> >> confusion.
> >>>> If this is the case perhaps rather than linking ACs and VCs through
> >>>> the physical or virtual co-ordinates of the area of capture
> >>>> information, we simplify things and in each audio capture we say
> >>>> what VC
> >> it relates to?
> >>>> e.g.using Mark's example AC0(VC0), AC3(V0,V1,V2)
> >>> [Duckworth, Mark] I'm not sure how this would really work.  Because
> >>> even
> >> in this simple example, we also have AC0 relates to VC3 (sometimes),
> >> VC4, and VC5 (at least part of it), and so on.  AC3 relates to all of =
VC1
> through VC5.
> >> I think using area of capture works better than trying to do
> >> something like this.
> > ...snip...
> >


From nobody Wed Apr 30 08:40:23 2014
Return-Path: <john@jlc.net>
X-Original-To: clue@ietfa.amsl.com
Delivered-To: clue@ietfa.amsl.com
Received: from localhost (ietfa.amsl.com [127.0.0.1]) by ietfa.amsl.com (Postfix) with ESMTP id D0B791A08E8 for <clue@ietfa.amsl.com>; Wed, 30 Apr 2014 08:40:21 -0700 (PDT)
X-Virus-Scanned: amavisd-new at amsl.com
X-Spam-Flag: NO
X-Spam-Score: -4.851
X-Spam-Level: 
X-Spam-Status: No, score=-4.851 tagged_above=-999 required=5 tests=[BAYES_00=-1.9, RCVD_IN_DNSWL_MED=-2.3, RP_MATCHES_RCVD=-0.651] autolearn=ham
Received: from mail.ietf.org ([4.31.198.44]) by localhost (ietfa.amsl.com [127.0.0.1]) (amavisd-new, port 10024) with ESMTP id fsVemFBUsts7 for <clue@ietfa.amsl.com>; Wed, 30 Apr 2014 08:40:19 -0700 (PDT)
Received: from mailhost.jlc.net (mailhost.jlc.net [199.201.159.4]) by ietfa.amsl.com (Postfix) with ESMTP id 609E21A08E6 for <clue@ietf.org>; Wed, 30 Apr 2014 08:40:19 -0700 (PDT)
Received: by mailhost.jlc.net (Postfix, from userid 104) id 88FFFC94BD; Wed, 30 Apr 2014 11:40:15 -0400 (EDT)
Date: Wed, 30 Apr 2014 11:40:15 -0400
From: John Leslie <john@jlc.net>
To: Stephen Botzko <stephen.botzko@gmail.com>
Message-ID: <20140430154015.GE44329@verdi>
References: <535F0B3A.70808@nteczone.com> <535F262C.6090009@alum.mit.edu> <535F73B6.9030106@nteczone.com> <CAMC7SJ7ZNW+=PDB-39gQ3MkQpTh5W1LZLPz7Y8VaFc7OofRviw@mail.gmail.com> <535FD82A.80708@alum.mit.edu> <CAMC7SJ6TGUYQVNcf__At7rkwgBQkoYBXAOfmnnLPW56=-=96wA@mail.gmail.com> <5360010A.1000904@alum.mit.edu> <CAMC7SJ6ARW5Z2AX526AiabO1AMi5qd0C8uMEc63PRdHWv_b8Kw@mail.gmail.com> <BLU0-SMTP215C180D3356A66A02AD04D0460@phx.gbl> <CAMC7SJ6VDbOegVayPVa7zoOZ6Q7t1hUg=dCtwSNkyaUzxdjKoA@mail.gmail.com>
Mime-Version: 1.0
Content-Type: text/plain; charset=us-ascii
Content-Disposition: inline
In-Reply-To: <CAMC7SJ6VDbOegVayPVa7zoOZ6Q7t1hUg=dCtwSNkyaUzxdjKoA@mail.gmail.com>
User-Agent: Mutt/1.4.1i
Archived-At: http://mailarchive.ietf.org/arch/msg/clue/GybHkujOHLr_8KYX36nWa6BI6d0
Cc: CLUE <clue@ietf.org>
Subject: Re: [clue] Improving treatment of audio
X-BeenThere: clue@ietf.org
X-Mailman-Version: 2.1.15
Precedence: list
List-Id: CLUE - ControLling mUltiple streams for TElepresence <clue.ietf.org>
List-Unsubscribe: <https://www.ietf.org/mailman/options/clue>, <mailto:clue-request@ietf.org?subject=unsubscribe>
List-Archive: <http://www.ietf.org/mail-archive/web/clue/>
List-Post: <mailto:clue@ietf.org>
List-Help: <mailto:clue-request@ietf.org?subject=help>
List-Subscribe: <https://www.ietf.org/mailman/listinfo/clue>, <mailto:clue-request@ietf.org?subject=subscribe>
X-List-Received-Date: Wed, 30 Apr 2014 15:40:22 -0000

   (I realize I'm falling behind in replies -- I'm afraid that will
continue for the rest of this week. I do mean to get things out well
before May 13th.)

Stephen Botzko <stephen.botzko@gmail.com> wrote:
> On Tue, Apr 29, 2014 at 7:50 PM, Paul Coverdale <coverdale@sympatico.ca>
> wrote:
>>
>>... In other words, microphone types and their physical location should
>> be a matter for the sending side only, the only criterion being that
>> they should be able to be associated with any possible video capture
>> in that room, and provide a high signal to noise ratio for sending to
>> the receiving side. This would cover the use of a single omni, multiple
>> directional, headsets, beam-steering arrays etc.

   This sounds like hand-waving to me.

> I was thinking that if you had a way to align the audio capture axis
> in sender and receiver, that you would have the audio analog of a
> line-of-sight.

   I can't imagine what such an "analog" would be.

> I don't see any way to use the audio capture axis information without
> such a method.

   Audio capture axis isn't intended to solve the problem I think you're
trying to solve. It's intended to convey the _meaning_ of the microphone
capture.

   If I understand correctly, you're trying to find a way to direct the
listener's attention to the screen of the person talking. This is trivial
if every person speaking has an individual microphone, and nearly
impossible otherwise. The best you can do is compare the relative audio
levels of various microphones to place the person on a line perpendicular
to the line between two microphones; guess from the video where on that
line the person lies; and mix the sound between speakers to match the
position on the screen.

   This may be worth doing, for all I know; but myself, I wouldn't bother.

> I agree that for most systems the rendering axis of the speakers will be
> perpendicular to their displays, and that it would be sensible for senders
> to take that into account when they design their audio captures.

   Telepresence system designers can't reasonably "take that into account"
for the wide spectrum of future telepresence systems. Nor do I believe
there's anything we could write in our spec to enable that.

> I also agree that microphone types and physical location need be a matter
> for the sending side only

   In practice, we will often see a single audio stream from a room, which
is the output of a mixer. To a first approximation, what inputs go into
that mixer are indeed a matter for the sending side only.

   But the actual location, line-of-capture, and directional-pattern of
a microphone whose audio is sent as a separate stream is _critical_ to
intelligent use of that stream. (IMHO, there's plenty of bandwidth
available to typical users of telepresence equipment to send one audio
stream per person expected to talk in addition to the mixer output.)

   Of course, when there is only a single mixer output sent, the tricks
being discussed here won't help anyway...

> (and that loudspeaker types and loudspeaker locations also need to be
> a matter for the receiving side only).

   I don't believe it could ever be helpful to the sender to know the
placement of loudspeakers. The closest we might come is for the sender
to announce to the receiver what placement it _intends_ when it sends
stereo or multi-channel streams.

> There is a price for that design constraint - if you have full control
> over both sender and receiver equipment you can design a more convincing
> sense of presence (both for audio and for video).

   That, IMHO, can only happen if both ends use the same vendor. We're
writing standards for interoperation of different vendors' equipment.

> But that's a price worth paying to get broad interoperability.

   Selecting the same vendor will indeed make economic sense for many
customers -- quite regardless of anything we specify here.

--
John Leslie <john@jlc.net>


From nobody Wed Apr 30 09:35:00 2014
Return-Path: <pkyzivat@alum.mit.edu>
X-Original-To: clue@ietfa.amsl.com
Delivered-To: clue@ietfa.amsl.com
Received: from localhost (ietfa.amsl.com [127.0.0.1]) by ietfa.amsl.com (Postfix) with ESMTP id 48F361A8824 for <clue@ietfa.amsl.com>; Wed, 30 Apr 2014 09:34:59 -0700 (PDT)
X-Virus-Scanned: amavisd-new at amsl.com
X-Spam-Flag: NO
X-Spam-Score: -1.235
X-Spam-Level: 
X-Spam-Status: No, score=-1.235 tagged_above=-999 required=5 tests=[BAYES_00=-1.9, DKIM_SIGNED=0.1, DKIM_VALID=-0.1, SPF_SOFTFAIL=0.665] autolearn=no
Received: from mail.ietf.org ([4.31.198.44]) by localhost (ietfa.amsl.com [127.0.0.1]) (amavisd-new, port 10024) with ESMTP id KDJsD16EF5s0 for <clue@ietfa.amsl.com>; Wed, 30 Apr 2014 09:34:58 -0700 (PDT)
Received: from qmta03.westchester.pa.mail.comcast.net (qmta03.westchester.pa.mail.comcast.net [IPv6:2001:558:fe14:43:76:96:62:32]) by ietfa.amsl.com (Postfix) with ESMTP id 44EC51A802D for <clue@ietf.org>; Wed, 30 Apr 2014 09:34:58 -0700 (PDT)
Received: from omta10.westchester.pa.mail.comcast.net ([76.96.62.28]) by qmta03.westchester.pa.mail.comcast.net with comcast id wGXl1n0030cZkys53Gawqc; Wed, 30 Apr 2014 16:34:56 +0000
Received: from Paul-Kyzivats-MacBook-Pro.local ([50.138.229.164]) by omta10.westchester.pa.mail.comcast.net with comcast id wGaw1n00t3ZTu2S3WGawD5; Wed, 30 Apr 2014 16:34:56 +0000
Message-ID: <53612630.7050802@alum.mit.edu>
Date: Wed, 30 Apr 2014 12:34:56 -0400
From: Paul Kyzivat <pkyzivat@alum.mit.edu>
User-Agent: Mozilla/5.0 (Macintosh; Intel Mac OS X 10.7; rv:24.0) Gecko/20100101 Thunderbird/24.5.0
MIME-Version: 1.0
To: clue@ietf.org
References: <535F0B3A.70808@nteczone.com> <535F262C.6090009@alum.mit.edu> <535F73B6.9030106@nteczone.com> <CAMC7SJ7ZNW+=PDB-39gQ3MkQpTh5W1LZLPz7Y8VaFc7OofRviw@mail.gmail.com> <535FD82A.80708@alum.mit.edu> <CAMC7SJ6TGUYQVNcf__At7rkwgBQkoYBXAOfmnnLPW56=-=96wA@mail.gmail.com> <5360010A.1000904@alum.mit.edu> <CAMC7SJ6ARW5Z2AX526AiabO1AMi5qd0C8uMEc63PRdHWv_b8Kw@mail.gmail.com> <BLU0-SMTP215C180D3356A66A02AD04D0460@phx.gbl> <CAMC7SJ6VDbOegVayPVa7zoOZ6Q7t1hUg=dCtwSNkyaUzxdjKoA@mail.gmail.com> <20140430154015.GE44329@verdi>
In-Reply-To: <20140430154015.GE44329@verdi>
Content-Type: text/plain; charset=ISO-8859-1; format=flowed
Content-Transfer-Encoding: 7bit
DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=comcast.net; s=q20140121; t=1398875696; bh=+XddD74/Sj4b43IiwyabUF+DPX9hDAvbTBVDWYC5ChM=; h=Received:Received:Message-ID:Date:From:MIME-Version:To:Subject: Content-Type; b=sWvSWqhbYJ5Inn1pYi43OQafW/w0ahR0UGVHvav6huB0bmeGgMbMejhlZisqqY3Zs fZDFIMH7PxykYPUdLnsyVsn5fdH9rfiFdWGDx6hZLlgLs6uVtUzVYz7JQItuLmT04g qBhnK66gGOhByc+9u4A4wFrdBvWCVBixs8+2ge4+aX8aEOhjv0BzW33ktD/+6RoJbY TPWNp+wFb6I5yWxAI7/9BThL0AcDrkvi9QQAHeRkgyAhyFP+HghTe0fJ6nJhgVLhof vXj2vhqCQB2+q6Ktfviwyd7L8VVkkSJenZnuwWe/PXcY7TpJpNoNS0ECSeFWz5soy1 1vldFSb53SWbw==
Archived-At: http://mailarchive.ietf.org/arch/msg/clue/p_MeTwexZUhZA-iNYyIcd8_0gTY
Subject: Re: [clue] Improving treatment of audio
X-BeenThere: clue@ietf.org
X-Mailman-Version: 2.1.15
Precedence: list
List-Id: CLUE - ControLling mUltiple streams for TElepresence <clue.ietf.org>
List-Unsubscribe: <https://www.ietf.org/mailman/options/clue>, <mailto:clue-request@ietf.org?subject=unsubscribe>
List-Archive: <http://www.ietf.org/mail-archive/web/clue/>
List-Post: <mailto:clue@ietf.org>
List-Help: <mailto:clue-request@ietf.org?subject=help>
List-Subscribe: <https://www.ietf.org/mailman/listinfo/clue>, <mailto:clue-request@ietf.org?subject=subscribe>
X-List-Received-Date: Wed, 30 Apr 2014 16:34:59 -0000

Replying generally to this thread:

When I asked whether axis of capture might be more useful than area of 
capture I was responding to Steven's comment about computing which video 
capture was "closest" to the audio capture. Steven seemed to be 
suggesting that that distance would be computed between the centers of 
the audio and video areas of capture. I was thinking that computing 
which video area of capture an audio axis of capture came closest to. 
Also, ISTM that point of capture and axis of capture are more natural 
attributes to provide for audio than is an area of capture.

This is moot if we instead explicitly indicate associations between 
audio and video captures. (If we do that, I don't think it makes a lot 
of difference whether audio points to video or visa versa.)

More below.

On 4/30/14 11:40 AM, John Leslie wrote:
[snip]
> Stephen Botzko <stephen.botzko@gmail.com> wrote:
[snip]
>> I also agree that microphone types and physical location need be a matter
>> for the sending side only
>
>     In practice, we will often see a single audio stream from a room, which
> is the output of a mixer. To a first approximation, what inputs go into
> that mixer are indeed a matter for the sending side only.
>
>     But the actual location, line-of-capture, and directional-pattern of
> a microphone whose audio is sent as a separate stream is _critical_ to
> intelligent use of that stream. (IMHO, there's plenty of bandwidth
> available to typical users of telepresence equipment to send one audio
> stream per person expected to talk in addition to the mixer output.)
>
>     Of course, when there is only a single mixer output sent, the tricks
> being discussed here won't help anyway...
>
>> (and that loudspeaker types and loudspeaker locations also need to be
>> a matter for the receiving side only).
>
>     I don't believe it could ever be helpful to the sender to know the
> placement of loudspeakers. The closest we might come is for the sender
> to announce to the receiver what placement it _intends_ when it sends
> stereo or multi-channel streams.
>
>> There is a price for that design constraint - if you have full control
>> over both sender and receiver equipment you can design a more convincing
>> sense of presence (both for audio and for video).
>
>     That, IMHO, can only happen if both ends use the same vendor. We're
> writing standards for interoperation of different vendors' equipment.
>
>> But that's a price worth paying to get broad interoperability.
>
>     Selecting the same vendor will indeed make economic sense for many
> customers -- quite regardless of anything we specify here.

ISTM that there are some truths here that may be more apparent to some 
than others - that a certain division of responsibility is being made 
between the advertiser and the consumer, and a limited ability to share 
information via the advertisement. This limits how faithful the 
rendition can be.

I think we all recognize that if the equipment in two rooms is very 
inconsistent then it may be impossible to get a really good rendering.

But it is also the case that we could have two rooms that are quite 
different, but both very capable. If each had total knowledge of the 
other, and they could do arbitrary negotiation, then the advertiser 
might be able to construct (using the equipment it has) a set of 
captures that the consumer could process (using equipment it has) those 
captures to provide an excellent rendering.

However, the advertise/configure mechanism won't be sufficient to allow 
that to happen in an optimal way in all cases. We are making tradeoffs 
to simplify the process and make it practical to implement and deploy. 
It is unclear *to me* how are we are ending up from the ideal.

I *hope* that what we provide is sufficient so that implementations 
don't need to employ different algorithms when interoperating with known 
and unknown equipment.

	Thanks,
	Paul


From nobody Wed Apr 30 12:38:24 2014
Return-Path: <ietf-ipr@ietf.org>
X-Original-To: clue@ietfa.amsl.com
Delivered-To: clue@ietfa.amsl.com
Received: from localhost (ietfa.amsl.com [127.0.0.1]) by ietfa.amsl.com (Postfix) with ESMTP id 69FF81A88AA; Wed, 30 Apr 2014 12:38:15 -0700 (PDT)
X-Virus-Scanned: amavisd-new at amsl.com
X-Spam-Flag: NO
X-Spam-Score: -1.9
X-Spam-Level: 
X-Spam-Status: No, score=-1.9 tagged_above=-999 required=5 tests=[BAYES_00=-1.9] autolearn=ham
Received: from mail.ietf.org ([4.31.198.44]) by localhost (ietfa.amsl.com [127.0.0.1]) (amavisd-new, port 10024) with ESMTP id 2vp5YuhpfO8T; Wed, 30 Apr 2014 12:38:11 -0700 (PDT)
Received: from ietfa.amsl.com (localhost [IPv6:::1]) by ietfa.amsl.com (Postfix) with ESMTP id DE19B1A889E; Wed, 30 Apr 2014 12:38:10 -0700 (PDT)
MIME-Version: 1.0
Content-Type: text/plain; charset="utf-8"
Content-Transfer-Encoding: 7bit
From: IETF Secretariat <ietf-ipr@ietf.org>
To: apeppere@gmail.com,mark.duckworth@polycom.com,stewe@stewe.org
X-Test-IDTracker: no
X-IETF-IDTracker: 5.4.1
Auto-Submitted: auto-generated
Precedence: bulk
Message-ID: <20140430193810.6483.37869.idtracker@ietfa.amsl.com>
Date: Wed, 30 Apr 2014 12:38:10 -0700
Archived-At: http://mailarchive.ietf.org/arch/msg/clue/S6DpXpsmM3sl60yNg8IZthOWJlY
Cc: clue@ietf.org, ipr-announce@ietf.org
Subject: [clue] IPR Disclosure: Huawei Technologies Co., Ltd's Statement about IPR related to draft-ietf-clue-framework-14
X-BeenThere: clue@ietf.org
X-Mailman-Version: 2.1.15
List-Id: CLUE - ControLling mUltiple streams for TElepresence <clue.ietf.org>
List-Unsubscribe: <https://www.ietf.org/mailman/options/clue>, <mailto:clue-request@ietf.org?subject=unsubscribe>
List-Archive: <http://www.ietf.org/mail-archive/web/clue/>
List-Post: <mailto:clue@ietf.org>
List-Help: <mailto:clue-request@ietf.org?subject=help>
List-Subscribe: <https://www.ietf.org/mailman/listinfo/clue>, <mailto:clue-request@ietf.org?subject=subscribe>
X-List-Received-Date: Wed, 30 Apr 2014 19:38:15 -0000

Dear Andrew Pepperell, Mark Duckworth, Stephan Wenger:

 An IPR disclosure that pertains to your Internet-Draft entitled "Framework for
Telepresence Multi-Streams" (draft-ietf-clue-framework) was submitted to the
IETF Secretariat on 2014-04-28 and has been posted on the "IETF Page of
Intellectual Property Rights Disclosures"
(https://datatracker.ietf.org/ipr/2347/). The title of the IPR disclosure is
"Huawei Technologies Co.,Ltd's Statement about IPR related to draft-ietf-clue-
framework-14."");

The IETF Secretariat


From nobody Wed Apr 30 19:41:05 2014
Return-Path: <stephen.botzko@gmail.com>
X-Original-To: clue@ietfa.amsl.com
Delivered-To: clue@ietfa.amsl.com
Received: from localhost (ietfa.amsl.com [127.0.0.1]) by ietfa.amsl.com (Postfix) with ESMTP id 035331A09DD for <clue@ietfa.amsl.com>; Wed, 30 Apr 2014 19:41:04 -0700 (PDT)
X-Virus-Scanned: amavisd-new at amsl.com
X-Spam-Flag: NO
X-Spam-Score: -1.999
X-Spam-Level: 
X-Spam-Status: No, score=-1.999 tagged_above=-999 required=5 tests=[BAYES_00=-1.9, DKIM_SIGNED=0.1, DKIM_VALID=-0.1, DKIM_VALID_AU=-0.1, FREEMAIL_FROM=0.001, HTML_MESSAGE=0.001, SPF_PASS=-0.001] autolearn=ham
Received: from mail.ietf.org ([4.31.198.44]) by localhost (ietfa.amsl.com [127.0.0.1]) (amavisd-new, port 10024) with ESMTP id OLA7mxaBOv48 for <clue@ietfa.amsl.com>; Wed, 30 Apr 2014 19:41:01 -0700 (PDT)
Received: from mail-ve0-x229.google.com (mail-ve0-x229.google.com [IPv6:2607:f8b0:400c:c01::229]) by ietfa.amsl.com (Postfix) with ESMTP id 7A7FD1A092F for <clue@ietf.org>; Wed, 30 Apr 2014 19:41:01 -0700 (PDT)
Received: by mail-ve0-f169.google.com with SMTP id jx11so3229222veb.14 for <clue@ietf.org>; Wed, 30 Apr 2014 19:40:59 -0700 (PDT)
DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20120113;  h=mime-version:in-reply-to:references:date:message-id:subject:from:to :cc:content-type; bh=ubP4W9qjDWSybvuRRGFWYc5H68/FkLirC+kZfa7mnyE=; b=meFNMqyGBllbNhPJ1PJndDrBpCK0MDzqJQ+5CdnHTKDG8B6eRmC8hbANr5B/nz41nQ 2VntUM2OYyEieN5VZHdu6FNY+OQ9VNYf9rEPJRULrmQIhFZFFNQOLJ0LqJWqVEIDoqBi EMeu+mTlJ7EYLI0DaMxiqm4urCL4eP38chg0Pohju288HXOOGsNNdNZyIeR9lGqnJH0T rCLN8C9FnWXseqUeTf+tdARacXZm+raq5+3sJZaJ+AvgK2L9bVtCZg4LymiaAX6J6l6q M3oYp7OF0+Y7Kcxbg9HDc/g2OriArwa8e8dHSWXEP+uW8urmgydygNKHdrOwpqr5Nr9I +6vg==
MIME-Version: 1.0
X-Received: by 10.58.123.71 with SMTP id ly7mr6449903veb.11.1398912059519; Wed, 30 Apr 2014 19:40:59 -0700 (PDT)
Received: by 10.221.40.135 with HTTP; Wed, 30 Apr 2014 19:40:59 -0700 (PDT)
In-Reply-To: <53612630.7050802@alum.mit.edu>
References: <535F0B3A.70808@nteczone.com> <535F262C.6090009@alum.mit.edu> <535F73B6.9030106@nteczone.com> <CAMC7SJ7ZNW+=PDB-39gQ3MkQpTh5W1LZLPz7Y8VaFc7OofRviw@mail.gmail.com> <535FD82A.80708@alum.mit.edu> <CAMC7SJ6TGUYQVNcf__At7rkwgBQkoYBXAOfmnnLPW56=-=96wA@mail.gmail.com> <5360010A.1000904@alum.mit.edu> <CAMC7SJ6ARW5Z2AX526AiabO1AMi5qd0C8uMEc63PRdHWv_b8Kw@mail.gmail.com> <BLU0-SMTP215C180D3356A66A02AD04D0460@phx.gbl> <CAMC7SJ6VDbOegVayPVa7zoOZ6Q7t1hUg=dCtwSNkyaUzxdjKoA@mail.gmail.com> <20140430154015.GE44329@verdi> <53612630.7050802@alum.mit.edu>
Date: Wed, 30 Apr 2014 22:40:59 -0400
Message-ID: <CAMC7SJ6bCAZ0sw7yuf3Y2jGmWE13o6gKsCh=H=Zhw36G=3+YnQ@mail.gmail.com>
From: Stephen Botzko <stephen.botzko@gmail.com>
To: Paul Kyzivat <pkyzivat@alum.mit.edu>
Content-Type: multipart/alternative; boundary=089e0115f0a261365b04f84d98ff
Archived-At: http://mailarchive.ietf.org/arch/msg/clue/eO00WK0PKR45vciWoRJTIpDnZfA
Cc: CLUE <clue@ietf.org>
Subject: Re: [clue] Improving treatment of audio
X-BeenThere: clue@ietf.org
X-Mailman-Version: 2.1.15
Precedence: list
List-Id: CLUE - ControLling mUltiple streams for TElepresence <clue.ietf.org>
List-Unsubscribe: <https://www.ietf.org/mailman/options/clue>, <mailto:clue-request@ietf.org?subject=unsubscribe>
List-Archive: <http://www.ietf.org/mail-archive/web/clue/>
List-Post: <mailto:clue@ietf.org>
List-Help: <mailto:clue-request@ietf.org?subject=help>
List-Subscribe: <https://www.ietf.org/mailman/listinfo/clue>, <mailto:clue-request@ietf.org?subject=subscribe>
X-List-Received-Date: Thu, 01 May 2014 02:41:04 -0000

--089e0115f0a261365b04f84d98ff
Content-Type: text/plain; charset=UTF-8

On Wed, Apr 30, 2014 at 12:34 PM, Paul Kyzivat <pkyzivat@alum.mit.edu>wrote:

> Replying generally to this thread:
>
> When I asked whether axis of capture might be more useful than area of
> capture I was responding to Steven's comment about computing which video
> capture was "closest" to the audio capture. Steven seemed to be suggesting
> that that distance would be computed between the centers of the audio and
> video areas of capture. I was thinking that computing which video area of
> capture an audio axis of capture came closest to. Also, ISTM that point of
> capture and axis of capture are more natural attributes to provide for
> audio than is an area of capture.
>
I believe I was actually describing how to pan an audio capture so that
when rendering it is placed at the correct point in the rendered video
scene.  I wasn't attempting to describe how to identify which video capture
is closest to which audio capture.

>
> This is moot if we instead explicitly indicate associations between audio
> and video captures. (If we do that, I don't think it makes a lot of
> difference whether audio points to video or visa versa.)
>
> More below.
>
> On 4/30/14 11:40 AM, John Leslie wrote:
> [snip]
>
>> Stephen Botzko <stephen.botzko@gmail.com> wrote:
>>
> [snip]
>
>  I also agree that microphone types and physical location need be a matter
>>> for the sending side only
>>>
>>
>>     In practice, we will often see a single audio stream from a room,
>> which
>> is the output of a mixer. To a first approximation, what inputs go into
>> that mixer are indeed a matter for the sending side only.
>>
>>     But the actual location, line-of-capture, and directional-pattern of
>> a microphone whose audio is sent as a separate stream is _critical_ to
>> intelligent use of that stream. (IMHO, there's plenty of bandwidth
>> available to typical users of telepresence equipment to send one audio
>> stream per person expected to talk in addition to the mixer output.)
>>
>>     Of course, when there is only a single mixer output sent, the tricks
>> being discussed here won't help anyway...
>>
>>  (and that loudspeaker types and loudspeaker locations also need to be
>>> a matter for the receiving side only).
>>>
>>
>>     I don't believe it could ever be helpful to the sender to know the
>> placement of loudspeakers. The closest we might come is for the sender
>> to announce to the receiver what placement it _intends_ when it sends
>> stereo or multi-channel streams.
>>
>>  There is a price for that design constraint - if you have full control
>>> over both sender and receiver equipment you can design a more convincing
>>> sense of presence (both for audio and for video).
>>>
>>
>>     That, IMHO, can only happen if both ends use the same vendor. We're
>> writing standards for interoperation of different vendors' equipment.
>>
>>  But that's a price worth paying to get broad interoperability.
>>>
>>
>>     Selecting the same vendor will indeed make economic sense for many
>> customers -- quite regardless of anything we specify here.
>>
>
> ISTM that there are some truths here that may be more apparent to some
> than others - that a certain division of responsibility is being made
> between the advertiser and the consumer, and a limited ability to share
> information via the advertisement. This limits how faithful the rendition
> can be.
>
> I think we all recognize that if the equipment in two rooms is very
> inconsistent then it may be impossible to get a really good rendering.
>
> But it is also the case that we could have two rooms that are quite
> different, but both very capable. If each had total knowledge of the other,
> and they could do arbitrary negotiation, then the advertiser might be able
> to construct (using the equipment it has) a set of captures that the
> consumer could process (using equipment it has) those captures to provide
> an excellent rendering.
>
> However, the advertise/configure mechanism won't be sufficient to allow
> that to happen in an optimal way in all cases. We are making tradeoffs to
> simplify the process and make it practical to implement and deploy. It is
> unclear *to me* how are we are ending up from the ideal.
>
> I believe your overall assessment is correct - there are trade-offs here.
 In my view, a simple but adequate framework is more practical than one
that turns rendering into a research problem in acoustics.  Also, we need
to ensure that the information that is sent is something that can be easily
understood and verified by the folks that are doing the signaling
implementations - otherwise the information will almost certainly be wrong.


I do understand that the impact of the compromises might be unclear.
 Hopefully the design team meetings will help clarify.  Though we should
also keep in mind that some amount of telepresence interoperability between
disparate systems already happens.  Advertisement/configures are not being
used, but the basic audio techniques are,


> I *hope* that what we provide is sufficient so that implementations don't
> need to employ different algorithms when interoperating with known and
> unknown equipment.
>
>         Thanks,
>         Paul
>
>
> _______________________________________________
> clue mailing list
> clue@ietf.org
> https://www.ietf.org/mailman/listinfo/clue
>

--089e0115f0a261365b04f84d98ff
Content-Type: text/html; charset=UTF-8
Content-Transfer-Encoding: quoted-printable

<div dir=3D"ltr"><br><div class=3D"gmail_extra"><br><br><div class=3D"gmail=
_quote">On Wed, Apr 30, 2014 at 12:34 PM, Paul Kyzivat <span dir=3D"ltr">&l=
t;<a href=3D"mailto:pkyzivat@alum.mit.edu" target=3D"_blank">pkyzivat@alum.=
mit.edu</a>&gt;</span> wrote:<br>
<blockquote class=3D"gmail_quote" style=3D"margin:0 0 0 .8ex;border-left:1p=
x #ccc solid;padding-left:1ex">Replying generally to this thread:<br>
<br>
When I asked whether axis of capture might be more useful than area of capt=
ure I was responding to Steven&#39;s comment about computing which video ca=
pture was &quot;closest&quot; to the audio capture. Steven seemed to be sug=
gesting that that distance would be computed between the centers of the aud=
io and video areas of capture. I was thinking that computing which video ar=
ea of capture an audio axis of capture came closest to. Also, ISTM that poi=
nt of capture and axis of capture are more natural attributes to provide fo=
r audio than is an area of capture.<br>
</blockquote><div>I believe I was actually describing how to pan an audio c=
apture so that when rendering it is placed at the correct point in the rend=
ered video scene. =C2=A0I wasn&#39;t attempting to describe how to identify=
 which video capture is closest to which audio capture.=C2=A0</div>
<blockquote class=3D"gmail_quote" style=3D"margin:0 0 0 .8ex;border-left:1p=
x #ccc solid;padding-left:1ex">
<br>
This is moot if we instead explicitly indicate associations between audio a=
nd video captures. (If we do that, I don&#39;t think it makes a lot of diff=
erence whether audio points to video or visa versa.)<br>
<br>
More below.<br>
<br>
On 4/30/14 11:40 AM, John Leslie wrote:<br>
[snip]<br>
<blockquote class=3D"gmail_quote" style=3D"margin:0 0 0 .8ex;border-left:1p=
x #ccc solid;padding-left:1ex">
Stephen Botzko &lt;<a href=3D"mailto:stephen.botzko@gmail.com" target=3D"_b=
lank">stephen.botzko@gmail.com</a>&gt; wrote:<br>
</blockquote>
[snip]<div><div class=3D"h5"><br>
<blockquote class=3D"gmail_quote" style=3D"margin:0 0 0 .8ex;border-left:1p=
x #ccc solid;padding-left:1ex"><blockquote class=3D"gmail_quote" style=3D"m=
argin:0 0 0 .8ex;border-left:1px #ccc solid;padding-left:1ex">
I also agree that microphone types and physical location need be a matter<b=
r>
for the sending side only<br>
</blockquote>
<br>
=C2=A0 =C2=A0 In practice, we will often see a single audio stream from a r=
oom, which<br>
is the output of a mixer. To a first approximation, what inputs go into<br>
that mixer are indeed a matter for the sending side only.<br>
<br>
=C2=A0 =C2=A0 But the actual location, line-of-capture, and directional-pat=
tern of<br>
a microphone whose audio is sent as a separate stream is _critical_ to<br>
intelligent use of that stream. (IMHO, there&#39;s plenty of bandwidth<br>
available to typical users of telepresence equipment to send one audio<br>
stream per person expected to talk in addition to the mixer output.)<br>
<br>
=C2=A0 =C2=A0 Of course, when there is only a single mixer output sent, the=
 tricks<br>
being discussed here won&#39;t help anyway...<br>
<br>
<blockquote class=3D"gmail_quote" style=3D"margin:0 0 0 .8ex;border-left:1p=
x #ccc solid;padding-left:1ex">
(and that loudspeaker types and loudspeaker locations also need to be<br>
a matter for the receiving side only).<br>
</blockquote>
<br>
=C2=A0 =C2=A0 I don&#39;t believe it could ever be helpful to the sender to=
 know the<br>
placement of loudspeakers. The closest we might come is for the sender<br>
to announce to the receiver what placement it _intends_ when it sends<br>
stereo or multi-channel streams.<br>
<br>
<blockquote class=3D"gmail_quote" style=3D"margin:0 0 0 .8ex;border-left:1p=
x #ccc solid;padding-left:1ex">
There is a price for that design constraint - if you have full control<br>
over both sender and receiver equipment you can design a more convincing<br=
>
sense of presence (both for audio and for video).<br>
</blockquote>
<br>
=C2=A0 =C2=A0 That, IMHO, can only happen if both ends use the same vendor.=
 We&#39;re<br>
writing standards for interoperation of different vendors&#39; equipment.<b=
r>
<br>
<blockquote class=3D"gmail_quote" style=3D"margin:0 0 0 .8ex;border-left:1p=
x #ccc solid;padding-left:1ex">
But that&#39;s a price worth paying to get broad interoperability.<br>
</blockquote>
<br>
=C2=A0 =C2=A0 Selecting the same vendor will indeed make economic sense for=
 many<br>
customers -- quite regardless of anything we specify here.<br>
</blockquote>
<br></div></div>
ISTM that there are some truths here that may be more apparent to some than=
 others - that a certain division of responsibility is being made between t=
he advertiser and the consumer, and a limited ability to share information =
via the advertisement. This limits how faithful the rendition can be.<br>

<br>
I think we all recognize that if the equipment in two rooms is very inconsi=
stent then it may be impossible to get a really good rendering.<br>
<br>
But it is also the case that we could have two rooms that are quite differe=
nt, but both very capable. If each had total knowledge of the other, and th=
ey could do arbitrary negotiation, then the advertiser might be able to con=
struct (using the equipment it has) a set of captures that the consumer cou=
ld process (using equipment it has) those captures to provide an excellent =
rendering.<br>

<br>
However, the advertise/configure mechanism won&#39;t be sufficient to allow=
 that to happen in an optimal way in all cases. We are making tradeoffs to =
simplify the process and make it practical to implement and deploy. It is u=
nclear *to me* how are we are ending up from the ideal.<br>

<br></blockquote><div>I believe your overall assessment is correct - there =
are trade-offs here. =C2=A0In my view, a simple but adequate framework is m=
ore practical than one that turns rendering into a research problem in acou=
stics. =C2=A0Also, we need to ensure that the information that is sent is s=
omething that can be easily understood and verified by the folks that are d=
oing the signaling implementations - otherwise the information will almost =
certainly be wrong. =C2=A0=C2=A0</div>
<div><br></div><div>I do understand that the impact of the compromises migh=
t be unclear. =C2=A0Hopefully the design team meetings will help clarify. =
=C2=A0Though we should also keep in mind that some amount of telepresence i=
nteroperability between disparate systems already happens. =C2=A0Advertisem=
ent/configures are not being used, but the basic audio techniques are, =C2=
=A0</div>
<div>=C2=A0</div><blockquote class=3D"gmail_quote" style=3D"margin:0 0 0 .8=
ex;border-left:1px #ccc solid;padding-left:1ex">
I *hope* that what we provide is sufficient so that implementations don&#39=
;t need to employ different algorithms when interoperating with known and u=
nknown equipment.<br>
<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 Thanks,<br>
=C2=A0 =C2=A0 =C2=A0 =C2=A0 Paul<div class=3D"HOEnZb"><div class=3D"h5"><br=
>
<br>
______________________________<u></u>_________________<br>
clue mailing list<br>
<a href=3D"mailto:clue@ietf.org" target=3D"_blank">clue@ietf.org</a><br>
<a href=3D"https://www.ietf.org/mailman/listinfo/clue" target=3D"_blank">ht=
tps://www.ietf.org/mailman/<u></u>listinfo/clue</a><br>
</div></div></blockquote></div><br></div></div>

--089e0115f0a261365b04f84d98ff--


From nobody Wed Apr 30 20:06:16 2014
Return-Path: <Christian.Groves@nteczone.com>
X-Original-To: clue@ietfa.amsl.com
Delivered-To: clue@ietfa.amsl.com
Received: from localhost (ietfa.amsl.com [127.0.0.1]) by ietfa.amsl.com (Postfix) with ESMTP id 9D1D31A88A6 for <clue@ietfa.amsl.com>; Wed, 30 Apr 2014 20:06:13 -0700 (PDT)
X-Virus-Scanned: amavisd-new at amsl.com
X-Spam-Flag: NO
X-Spam-Score: -1.9
X-Spam-Level: 
X-Spam-Status: No, score=-1.9 tagged_above=-999 required=5 tests=[BAYES_00=-1.9] autolearn=ham
Received: from mail.ietf.org ([4.31.198.44]) by localhost (ietfa.amsl.com [127.0.0.1]) (amavisd-new, port 10024) with ESMTP id aKOIrNO7BBZv for <clue@ietfa.amsl.com>; Wed, 30 Apr 2014 20:06:12 -0700 (PDT)
Received: from cserver5.myshophosting.com (cserver5.myshophosting.com [175.107.161.1]) by ietfa.amsl.com (Postfix) with ESMTP id 482A81A88A2 for <clue@ietf.org>; Wed, 30 Apr 2014 20:06:12 -0700 (PDT)
Received: from ppp118-209-119-137.lns20.mel4.internode.on.net ([118.209.119.137]:53865 helo=[127.0.0.1]) by cserver5.myshophosting.com with esmtpsa (TLSv1:DHE-RSA-AES128-SHA:128) (Exim 4.82) (envelope-from <Christian.Groves@nteczone.com>) id 1WfhK4-0003WY-5r for clue@ietf.org; Thu, 01 May 2014 13:06:08 +1000
Message-ID: <5361BA1C.7030509@nteczone.com>
Date: Thu, 01 May 2014 13:06:04 +1000
From: Christian Groves <Christian.Groves@nteczone.com>
User-Agent: Mozilla/5.0 (Windows NT 6.1; WOW64; rv:24.0) Gecko/20100101 Thunderbird/24.5.0
MIME-Version: 1.0
To: "clue@ietf.org" <clue@ietf.org>
Content-Type: text/plain; charset=ISO-8859-1; format=flowed
Content-Transfer-Encoding: 7bit
X-AntiAbuse: This header was added to track abuse, please include it with any abuse report
X-AntiAbuse: Primary Hostname - cserver5.myshophosting.com
X-AntiAbuse: Original Domain - ietf.org
X-AntiAbuse: Originator/Caller UID/GID - [47 12] / [47 12]
X-AntiAbuse: Sender Address Domain - nteczone.com
X-Get-Message-Sender-Via: cserver5.myshophosting.com: authenticated_id: christian.groves@nteczone.com
X-Source: 
X-Source-Args: 
X-Source-Dir: 
Archived-At: http://mailarchive.ietf.org/arch/msg/clue/z7GLHXh_BLxWLX5dAOJg_31w9YI
Subject: [clue] Taxonomy
X-BeenThere: clue@ietf.org
X-Mailman-Version: 2.1.15
Precedence: list
List-Id: CLUE - ControLling mUltiple streams for TElepresence <clue.ietf.org>
List-Unsubscribe: <https://www.ietf.org/mailman/options/clue>, <mailto:clue-request@ietf.org?subject=unsubscribe>
List-Archive: <http://www.ietf.org/mail-archive/web/clue/>
List-Post: <mailto:clue@ietf.org>
List-Help: <mailto:clue-request@ietf.org?subject=help>
List-Subscribe: <https://www.ietf.org/mailman/listinfo/clue>, <mailto:clue-request@ietf.org?subject=subscribe>
X-List-Received-Date: Thu, 01 May 2014 03:06:13 -0000

Hello all,

At the London meeting it was mentioned there is a potential conflict in 
the way that CLUE uses some of the terminology from the RTP taxonomy 
draft (draft-ietf-avtext-rtp-grouping-taxonomy).

I think this mainly relates to the use of the term "Participant".

In cl.2.2.3 of the taxonomy draft there is a definition for the term 
"participant":
/"A participant is an entity reachable by a single signaling address,//
//   and is thus related more to the signaling context than to the media//
//   context."

/In the XCON data model cl.4.6.5/[RFC6501] it uses the term "user" to 
describe the concept outlined in the draft definition. RFC4575 also uses 
the term "user" for this.

In CLUE we use participant as per the definition above (i.e. in the 
definition of Endpoint and MCU) but we also use "participant" to mean 
someone captured by the endpoint, i.e. the people in the room.

In order to minimise any confusion between the two uses I propose that 
we change the later use to "person" in our documents.

For example: in the framework:
- "Participant Information" becomes "Person Information",
- "participant type" becomes "person type".
- Where "participant" is used in the text it becomes "person" or "people".
   i.e.   Table - Captures the conference table with seated people.
     Individual - Captures an individual person.


Christian


From nobody Wed Apr 30 20:50:30 2014
Return-Path: <stephen.botzko@gmail.com>
X-Original-To: clue@ietfa.amsl.com
Delivered-To: clue@ietfa.amsl.com
Received: from localhost (ietfa.amsl.com [127.0.0.1]) by ietfa.amsl.com (Postfix) with ESMTP id 1B0771A090D for <clue@ietfa.amsl.com>; Wed, 30 Apr 2014 20:50:20 -0700 (PDT)
X-Virus-Scanned: amavisd-new at amsl.com
X-Spam-Flag: NO
X-Spam-Score: -1.999
X-Spam-Level: 
X-Spam-Status: No, score=-1.999 tagged_above=-999 required=5 tests=[BAYES_00=-1.9, DKIM_SIGNED=0.1, DKIM_VALID=-0.1, DKIM_VALID_AU=-0.1, FREEMAIL_FROM=0.001, HTML_MESSAGE=0.001, SPF_PASS=-0.001] autolearn=ham
Received: from mail.ietf.org ([4.31.198.44]) by localhost (ietfa.amsl.com [127.0.0.1]) (amavisd-new, port 10024) with ESMTP id PF1EwRdh1Z5F for <clue@ietfa.amsl.com>; Wed, 30 Apr 2014 20:50:15 -0700 (PDT)
Received: from mail-ve0-x231.google.com (mail-ve0-x231.google.com [IPv6:2607:f8b0:400c:c01::231]) by ietfa.amsl.com (Postfix) with ESMTP id B48871A09A0 for <clue@ietf.org>; Wed, 30 Apr 2014 20:50:14 -0700 (PDT)
Received: by mail-ve0-f177.google.com with SMTP id sa20so3286062veb.8 for <clue@ietf.org>; Wed, 30 Apr 2014 20:50:12 -0700 (PDT)
DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20120113;  h=mime-version:in-reply-to:references:date:message-id:subject:from:to :cc:content-type; bh=6kHsmHrYx3r9bF67qzOfT2+jRuQ4NsbTz6raX/UGvzQ=; b=rAokqLqJzjRtunRXoXYxlCfkWnu/DkY8Nbg4Y+cvclUE//f1/5VGndhOnb4CTN94RC xD50QEk2+U8zhkjp8ka4gdu1N4hwKe1NKx4Pa4K8q0hSK/1XOgF1ckhzrhLWQJwtfnEH QJwB1myc4zMa5BFDyCYZ5xwo6XOMJVl/f9PkRZBLzoT16OFmvLQcuew7oZnG/MWKOuRa UP9G/kDjabZVV4F0wTODMzdnZXAEMu1oJ42A5ckMa2jK5JqU+gJqf20GJKI1MUWDFz93 6U6Tk02aMXf+IsuVF2ke8meEZV0WyOvqbab/rCLdAagY3xlvAjTLaO6AqFEeYDFfECmA auRA==
MIME-Version: 1.0
X-Received: by 10.52.90.37 with SMTP id bt5mr5625476vdb.7.1398916212750; Wed, 30 Apr 2014 20:50:12 -0700 (PDT)
Received: by 10.221.40.135 with HTTP; Wed, 30 Apr 2014 20:50:12 -0700 (PDT)
In-Reply-To: <20140430154015.GE44329@verdi>
References: <535F0B3A.70808@nteczone.com> <535F262C.6090009@alum.mit.edu> <535F73B6.9030106@nteczone.com> <CAMC7SJ7ZNW+=PDB-39gQ3MkQpTh5W1LZLPz7Y8VaFc7OofRviw@mail.gmail.com> <535FD82A.80708@alum.mit.edu> <CAMC7SJ6TGUYQVNcf__At7rkwgBQkoYBXAOfmnnLPW56=-=96wA@mail.gmail.com> <5360010A.1000904@alum.mit.edu> <CAMC7SJ6ARW5Z2AX526AiabO1AMi5qd0C8uMEc63PRdHWv_b8Kw@mail.gmail.com> <BLU0-SMTP215C180D3356A66A02AD04D0460@phx.gbl> <CAMC7SJ6VDbOegVayPVa7zoOZ6Q7t1hUg=dCtwSNkyaUzxdjKoA@mail.gmail.com> <20140430154015.GE44329@verdi>
Date: Wed, 30 Apr 2014 23:50:12 -0400
Message-ID: <CAMC7SJ7aDCooaRLL_3DDaQdkUEXwZNz-Am1Xx5vq2zfOiO5qHg@mail.gmail.com>
From: Stephen Botzko <stephen.botzko@gmail.com>
To: John Leslie <john@jlc.net>
Content-Type: multipart/alternative; boundary=001a11369b06ee81eb04f84e8f2f
Archived-At: http://mailarchive.ietf.org/arch/msg/clue/zw-mxAE0yF_EzIOVP2HF5_Roj2k
Cc: CLUE <clue@ietf.org>
Subject: Re: [clue] Improving treatment of audio
X-BeenThere: clue@ietf.org
X-Mailman-Version: 2.1.15
Precedence: list
List-Id: CLUE - ControLling mUltiple streams for TElepresence <clue.ietf.org>
List-Unsubscribe: <https://www.ietf.org/mailman/options/clue>, <mailto:clue-request@ietf.org?subject=unsubscribe>
List-Archive: <http://www.ietf.org/mail-archive/web/clue/>
List-Post: <mailto:clue@ietf.org>
List-Help: <mailto:clue-request@ietf.org?subject=help>
List-Subscribe: <https://www.ietf.org/mailman/listinfo/clue>, <mailto:clue-request@ietf.org?subject=subscribe>
X-List-Received-Date: Thu, 01 May 2014 03:50:20 -0000

--001a11369b06ee81eb04f84e8f2f
Content-Type: text/plain; charset=UTF-8

On Wed, Apr 30, 2014 at 11:40 AM, John Leslie <john@jlc.net> wrote:

>    (I realize I'm falling behind in replies -- I'm afraid that will
> continue for the rest of this week. I do mean to get things out well
> before May 13th.)
>
> Stephen Botzko <stephen.botzko@gmail.com> wrote:
> > On Tue, Apr 29, 2014 at 7:50 PM, Paul Coverdale <coverdale@sympatico.ca>
> > wrote:
> >>
> >>... In other words, microphone types and their physical location should
> >> be a matter for the sending side only, the only criterion being that
> >> they should be able to be associated with any possible video capture
> >> in that room, and provide a high signal to noise ratio for sending to
> >> the receiving side. This would cover the use of a single omni, multiple
> >> directional, headsets, beam-steering arrays etc.
>
>    This sounds like hand-waving to me.
>
> > I was thinking that if you had a way to align the audio capture axis
> > in sender and receiver, that you would have the audio analog of a
> > line-of-sight.
>
>    I can't imagine what such an "analog" would be.

It's simple enough.  A camera/display pair is aligned to provide a line of
sight between a pair of participants, maintaining eye-contact.  That's
described in the CLUE tutorial we did some time ago.  A microphone and
speaker pair can potentially be configured to do something similar, if you
align the microphone capture axis with the loudspeaker axis in the remote
room.

>

> I don't see any way to use the audio capture axis information without
> > such a method.
>
>    Audio capture axis isn't intended to solve the problem I think you're
> trying to solve. It's intended to convey the _meaning_ of the microphone
> capture.
>
Why does the rendering system need to know that meaning?  How might it use
that information?  If it doesn't need to know it, why should I take the
trouble to measure it and signal it?

>
>    If I understand correctly, you're trying to find a way to direct the
> listener's attention to the screen of the person talking. This is trivial
> if every person speaking has an individual microphone, and nearly
> impossible otherwise. The best you can do is compare the relative audio
> levels of various microphones to place the person on a line perpendicular
> to the line between two microphones; guess from the video where on that
> line the person lies; and mix the sound between speakers to match the
> position on the screen.


>    This may be worth doing, for all I know; but myself, I wouldn't bother.
>
(a) it can be done without individual miking
(b) Panning audio captures to align them with the video is worth doing.

>
> > I agree that for most systems the rendering axis of the speakers will be
> > perpendicular to their displays, and that it would be sensible for
> senders
> > to take that into account when they design their audio captures.
>
>    Telepresence system designers can't reasonably "take that into account"
> for the wide spectrum of future telepresence systems. Nor do I believe
> there's anything we could write in our spec to enable that.



>
> Personally I believe it is reasonable for senders to make that particular
assumption about rendering.  If you want to make different assumptions when
you set up your microphone arrangements and capture advertisements, you are
free to do so.  I wasn't proposing that my assumptions be written into a
spec or that they needed to be.


> > I also agree that microphone types and physical location need be a matter
> > for the sending side only
>
>    In practice, we will often see a single audio stream from a room, which
> is the output of a mixer. To a first approximation, what inputs go into
> that mixer are indeed a matter for the sending side only.
>
>    But the actual location, line-of-capture, and directional-pattern of
> a microphone whose audio is sent as a separate stream is _critical_ to
> intelligent use of that stream.


Here we disagree.  I am not understanding why you think any of those things
are critical, or how you are proposing that a renderer or middle box would
use them.
Why, for instance, does a receiver need to know that my system is using
ceiling microphones instead of boundary microphones on the table?


> (IMHO, there's plenty of bandwidth
> available to typical users of telepresence equipment to send one audio
> stream per person expected to talk in addition to the mixer output.)
>
>    Of course, when there is only a single mixer output sent, the tricks
> being discussed here won't help anyway...
>
> > (and that loudspeaker types and loudspeaker locations also need to be
> > a matter for the receiving side only).
>
>    I don't believe it could ever be helpful to the sender to know the
> placement of loudspeakers. The closest we might come is for the sender
> to announce to the receiver what placement it _intends_ when it sends
> stereo or multi-channel streams

I think we are agreeing that loudspeaker placement shouldn't be signaled.

>
> > There is a price for that design constraint - if you have full control
> > over both sender and receiver equipment you can design a more convincing
> > sense of presence (both for audio and for video).
>
>    That, IMHO, can only happen if both ends use the same vendor. We're
> writing standards for interoperation of different vendors' equipment.
>
Which was my point. And that means there will be inevitably be some
compromises in the user experience

>
> > But that's a price worth paying to get broad interoperability.
>
>    Selecting the same vendor will indeed make economic sense for many
> customers -- quite regardless of anything we specify here.
>
> Certainly.  But that doesn't mean that those customers don't value or need
interoperability.
In some cases they will be talking to people in a different organization
which chose a different vendor. Also, even a single vendor can market more
than one type of system.  We do.

> --
> John Leslie <john@jlc.net>
>

--001a11369b06ee81eb04f84e8f2f
Content-Type: text/html; charset=UTF-8
Content-Transfer-Encoding: quoted-printable

<div dir=3D"ltr"><br><div class=3D"gmail_extra"><br><br><div class=3D"gmail=
_quote">On Wed, Apr 30, 2014 at 11:40 AM, John Leslie <span dir=3D"ltr">&lt=
;<a href=3D"mailto:john@jlc.net" target=3D"_blank">john@jlc.net</a>&gt;</sp=
an> wrote:<br>
<blockquote class=3D"gmail_quote" style=3D"margin:0px 0px 0px 0.8ex;border-=
left-width:1px;border-left-color:rgb(204,204,204);border-left-style:solid;p=
adding-left:1ex">=C2=A0 =C2=A0(I realize I&#39;m falling behind in replies =
-- I&#39;m afraid that will<br>

continue for the rest of this week. I do mean to get things out well<br>
before May 13th.)<br>
<div class=3D""><br>
Stephen Botzko &lt;<a href=3D"mailto:stephen.botzko@gmail.com">stephen.botz=
ko@gmail.com</a>&gt; wrote:<br>
&gt; On Tue, Apr 29, 2014 at 7:50 PM, Paul Coverdale &lt;<a href=3D"mailto:=
coverdale@sympatico.ca">coverdale@sympatico.ca</a>&gt;<br>
&gt; wrote:<br>
&gt;&gt;<br>
</div>&gt;&gt;... In other words, microphone types and their physical locat=
ion should<br>
<div class=3D"">&gt;&gt; be a matter for the sending side only, the only cr=
iterion being that<br>
&gt;&gt; they should be able to be associated with any possible video captu=
re<br>
&gt;&gt; in that room, and provide a high signal to noise ratio for sending=
 to<br>
&gt;&gt; the receiving side. This would cover the use of a single omni, mul=
tiple<br>
&gt;&gt; directional, headsets, beam-steering arrays etc.<br>
<br>
</div>=C2=A0 =C2=A0This sounds like hand-waving to me.<br>
<div class=3D""><br>
&gt; I was thinking that if you had a way to align the audio capture axis<b=
r>
&gt; in sender and receiver, that you would have the audio analog of a<br>
&gt; line-of-sight.<br>
<br>
</div>=C2=A0 =C2=A0I can&#39;t imagine what such an &quot;analog&quot; woul=
d be.</blockquote><div>It&#39;s simple enough. =C2=A0A camera/display pair =
is aligned to provide a line of sight between a pair of participants, maint=
aining eye-contact. =C2=A0That&#39;s described in the CLUE tutorial we did =
some time ago. =C2=A0A microphone and speaker pair can potentially be confi=
gured to do something similar, if you align the microphone capture axis wit=
h the loudspeaker axis in the remote room. =C2=A0 =C2=A0</div>
<blockquote class=3D"gmail_quote" style=3D"margin:0px 0px 0px 0.8ex;border-=
left-width:1px;border-left-color:rgb(204,204,204);border-left-style:solid;p=
adding-left:1ex">=C2=A0</blockquote><blockquote class=3D"gmail_quote" style=
=3D"margin:0px 0px 0px 0.8ex;border-left-width:1px;border-left-color:rgb(20=
4,204,204);border-left-style:solid;padding-left:1ex">
<div class=3D"">
&gt; I don&#39;t see any way to use the audio capture axis information with=
out<br>
&gt; such a method.<br>
<br>
</div>=C2=A0 =C2=A0Audio capture axis isn&#39;t intended to solve the probl=
em I think you&#39;re<br>
trying to solve. It&#39;s intended to convey the _meaning_ of the microphon=
e<br>
capture.<br></blockquote><div>Why does the rendering system need to know th=
at meaning? =C2=A0How might it use that information? =C2=A0If it doesn&#39;=
t need to know it, why should I take the trouble to measure it and signal i=
t?</div>
<blockquote class=3D"gmail_quote" style=3D"margin:0px 0px 0px 0.8ex;border-=
left-width:1px;border-left-color:rgb(204,204,204);border-left-style:solid;p=
adding-left:1ex">
<br>
=C2=A0 =C2=A0If I understand correctly, you&#39;re trying to find a way to =
direct the<br>
listener&#39;s attention to the screen of the person talking. This is trivi=
al<br>
if every person speaking has an individual microphone, and nearly<br>
impossible otherwise. The best you can do is compare the relative audio<br>
levels of various microphones to place the person on a line perpendicular<b=
r>
to the line between two microphones; guess from the video where on that<br>
line the person lies; and mix the sound between speakers to match the<br>
position on the screen.=C2=A0</blockquote><blockquote class=3D"gmail_quote"=
 style=3D"margin:0px 0px 0px 0.8ex;border-left-width:1px;border-left-color:=
rgb(204,204,204);border-left-style:solid;padding-left:1ex">
<br>
=C2=A0 =C2=A0This may be worth doing, for all I know; but myself, I wouldn&=
#39;t bother.<br></blockquote><div>(a) it can be done without individual mi=
king</div><div>(b) Panning audio captures to align them with the video is w=
orth doing.=C2=A0</div>
<blockquote class=3D"gmail_quote" style=3D"margin:0px 0px 0px 0.8ex;border-=
left-width:1px;border-left-color:rgb(204,204,204);border-left-style:solid;p=
adding-left:1ex">
<div class=3D""><br>
&gt; I agree that for most systems the rendering axis of the speakers will =
be<br>
&gt; perpendicular to their displays, and that it would be sensible for sen=
ders<br>
&gt; to take that into account when they design their audio captures.<br>
<br>
</div>=C2=A0 =C2=A0Telepresence system designers can&#39;t reasonably &quot=
;take that into account&quot;<br>
for the wide spectrum of future telepresence systems. Nor do I believe<br>
there&#39;s anything we could write in our spec to enable that.</blockquote=
><div>=C2=A0</div><blockquote class=3D"gmail_quote" style=3D"margin:0px 0px=
 0px 0.8ex;border-left-width:1px;border-left-color:rgb(204,204,204);border-=
left-style:solid;padding-left:1ex">
<br></blockquote><div>Personally I believe it is reasonable for senders to =
make that particular assumption about rendering. =C2=A0If you want to make =
different assumptions when you set up your microphone arrangements and capt=
ure advertisements, you are free to do so. =C2=A0I wasn&#39;t proposing tha=
t my assumptions be written into a spec or that they needed to be.=C2=A0</d=
iv>
<div><br></div><blockquote class=3D"gmail_quote" style=3D"margin:0px 0px 0p=
x 0.8ex;border-left-width:1px;border-left-color:rgb(204,204,204);border-lef=
t-style:solid;padding-left:1ex">
<div class=3D""><br>
&gt; I also agree that microphone types and physical location need be a mat=
ter<br>
&gt; for the sending side only<br>
<br>
</div>=C2=A0 =C2=A0In practice, we will often see a single audio stream fro=
m a room, which<br>
is the output of a mixer. To a first approximation, what inputs go into<br>
that mixer are indeed a matter for the sending side only.<br>
<br>
=C2=A0 =C2=A0But the actual location, line-of-capture, and directional-patt=
ern of<br>
a microphone whose audio is sent as a separate stream is _critical_ to<br>
intelligent use of that stream.</blockquote><div>=C2=A0</div><div>Here we d=
isagree. =C2=A0I am not understanding why you think any of those things are=
 critical, or how you are proposing that a renderer or middle box would use=
 them.=C2=A0</div>
<div>Why, for instance, does a receiver need to know that my system is usin=
g ceiling microphones instead of boundary microphones on the table?=C2=A0</=
div><div>=C2=A0</div><blockquote class=3D"gmail_quote" style=3D"margin:0px =
0px 0px 0.8ex;border-left-width:1px;border-left-color:rgb(204,204,204);bord=
er-left-style:solid;padding-left:1ex">
(IMHO, there&#39;s plenty of bandwidth<br>
available to typical users of telepresence equipment to send one audio<br>
stream per person expected to talk in addition to the mixer output.)<br>
<br>
=C2=A0 =C2=A0Of course, when there is only a single mixer output sent, the =
tricks<br>
being discussed here won&#39;t help anyway...<br>
<div class=3D""><br>
&gt; (and that loudspeaker types and loudspeaker locations also need to be<=
br>
&gt; a matter for the receiving side only).<br>
<br>
</div>=C2=A0 =C2=A0I don&#39;t believe it could ever be helpful to the send=
er to know the<br>
placement of loudspeakers. The closest we might come is for the sender<br>
to announce to the receiver what placement it _intends_ when it sends<br>
stereo or multi-channel streams</blockquote><div>I think we are agreeing th=
at loudspeaker placement shouldn&#39;t be signaled.=C2=A0</div><blockquote =
class=3D"gmail_quote" style=3D"margin:0px 0px 0px 0.8ex;border-left-width:1=
px;border-left-color:rgb(204,204,204);border-left-style:solid;padding-left:=
1ex">

<div class=3D""><br>
&gt; There is a price for that design constraint - if you have full control=
<br>
</div>&gt; over both sender and receiver equipment you can design a more co=
nvincing<br>
<div class=3D"">&gt; sense of presence (both for audio and for video).<br>
<br>
</div>=C2=A0 =C2=A0That, IMHO, can only happen if both ends use the same ve=
ndor. We&#39;re<br>
writing standards for interoperation of different vendors&#39; equipment.<b=
r></blockquote><div>Which was my point. And that means there will be inevit=
ably be some compromises in the user experience</div><blockquote class=3D"g=
mail_quote" style=3D"margin:0px 0px 0px 0.8ex;border-left-width:1px;border-=
left-color:rgb(204,204,204);border-left-style:solid;padding-left:1ex">

<div class=3D""><br>
&gt; But that&#39;s a price worth paying to get broad interoperability.<br>
<br>
</div>=C2=A0 =C2=A0Selecting the same vendor will indeed make economic sens=
e for many<br>
customers -- quite regardless of anything we specify here.<br>
<br></blockquote><div>Certainly. =C2=A0But that doesn&#39;t mean that those=
 customers don&#39;t value or need interoperability.=C2=A0</div><div>In som=
e cases they will be talking to people in a different organization which ch=
ose a different vendor. Also, even a single vendor can market more than one=
 type of system. =C2=A0We do.</div>
<blockquote class=3D"gmail_quote" style=3D"margin:0px 0px 0px 0.8ex;border-=
left-width:1px;border-left-color:rgb(204,204,204);border-left-style:solid;p=
adding-left:1ex">
--<br>
John Leslie &lt;<a href=3D"mailto:john@jlc.net">john@jlc.net</a>&gt;<br>
</blockquote></div><br></div></div>

--001a11369b06ee81eb04f84e8f2f--


From nobody Wed Apr 30 21:09:26 2014
Return-Path: <Christian.Groves@nteczone.com>
X-Original-To: clue@ietfa.amsl.com
Delivered-To: clue@ietfa.amsl.com
Received: from localhost (ietfa.amsl.com [127.0.0.1]) by ietfa.amsl.com (Postfix) with ESMTP id D98211A09B9 for <clue@ietfa.amsl.com>; Wed, 30 Apr 2014 21:09:24 -0700 (PDT)
X-Virus-Scanned: amavisd-new at amsl.com
X-Spam-Flag: NO
X-Spam-Score: -1.9
X-Spam-Level: 
X-Spam-Status: No, score=-1.9 tagged_above=-999 required=5 tests=[BAYES_00=-1.9] autolearn=ham
Received: from mail.ietf.org ([4.31.198.44]) by localhost (ietfa.amsl.com [127.0.0.1]) (amavisd-new, port 10024) with ESMTP id z6_m2wwZo6l5 for <clue@ietfa.amsl.com>; Wed, 30 Apr 2014 21:09:23 -0700 (PDT)
Received: from cserver5.myshophosting.com (cserver5.myshophosting.com [175.107.161.1]) by ietfa.amsl.com (Postfix) with ESMTP id 974841A09AD for <clue@ietf.org>; Wed, 30 Apr 2014 21:09:22 -0700 (PDT)
Received: from ppp118-209-119-137.lns20.mel4.internode.on.net ([118.209.119.137]:54139 helo=[127.0.0.1]) by cserver5.myshophosting.com with esmtpsa (TLSv1:DHE-RSA-AES128-SHA:128) (Exim 4.82) (envelope-from <Christian.Groves@nteczone.com>) id 1WfiJC-0006DC-1o; Thu, 01 May 2014 14:09:18 +1000
Message-ID: <5361C8EA.2060206@nteczone.com>
Date: Thu, 01 May 2014 14:09:14 +1000
From: Christian Groves <Christian.Groves@nteczone.com>
User-Agent: Mozilla/5.0 (Windows NT 6.1; WOW64; rv:24.0) Gecko/20100101 Thunderbird/24.5.0
MIME-Version: 1.0
To: "Duckworth, Mark" <Mark.Duckworth@polycom.com>,  "clue@ietf.org" <clue@ietf.org>
References: <533AF351.9050201@alum.mit.edu> <20140410222705.GW39240@verdi> <BLU0-SMTP6910DA20C9A9287137B585D0550@phx.gbl> <53475572.4040807@nteczone.com> <20140411150055.GE60844@verdi> <49E45C59CA48264997FEBFB29B6BC2D6215295A8AE@CRPMBOXPRD07.polycom.com> <BLU0-SMTP1804A70E3E0F666EAEAE84D0540@phx.gbl> <49E45C59CA48264997FEBFB29B6BC2D62409219636@CRPMBOXPRD07.polycom.com> <534B4303.7060707@nteczone.com> <5C4AC54BFF7A0842A6A11F554D6FB52F07152B@CRPMBOXPRD08.polycom.com> <534B536B.8000205@nteczone.com> <5C4AC54BFF7A0842A6A11F554D6FB52F13D9B2@CRPMBOXPRD08.polycom.com> <535F0B3A.70808@nteczone.com> <5C4AC54BFF7A0842A6A11F554D6FB52F13DAE6@CRPMBOXPRD08.polycom.com>
In-Reply-To: <5C4AC54BFF7A0842A6A11F554D6FB52F13DAE6@CRPMBOXPRD08.polycom.com>
Content-Type: text/plain; charset=ISO-8859-1; format=flowed
Content-Transfer-Encoding: 7bit
X-AntiAbuse: This header was added to track abuse, please include it with any abuse report
X-AntiAbuse: Primary Hostname - cserver5.myshophosting.com
X-AntiAbuse: Original Domain - ietf.org
X-AntiAbuse: Originator/Caller UID/GID - [47 12] / [47 12]
X-AntiAbuse: Sender Address Domain - nteczone.com
X-Get-Message-Sender-Via: cserver5.myshophosting.com: authenticated_id: christian.groves@nteczone.com
X-Source: 
X-Source-Args: 
X-Source-Dir: 
Archived-At: http://mailarchive.ietf.org/arch/msg/clue/FZZWEJU-EXmifdOUEqfnkHcbrHA
Subject: Re: [clue] Improving treatment of audio
X-BeenThere: clue@ietf.org
X-Mailman-Version: 2.1.15
Precedence: list
List-Id: CLUE - ControLling mUltiple streams for TElepresence <clue.ietf.org>
List-Unsubscribe: <https://www.ietf.org/mailman/options/clue>, <mailto:clue-request@ietf.org?subject=unsubscribe>
List-Archive: <http://www.ietf.org/mail-archive/web/clue/>
List-Post: <mailto:clue@ietf.org>
List-Help: <mailto:clue-request@ietf.org?subject=help>
List-Subscribe: <https://www.ietf.org/mailman/listinfo/clue>, <mailto:clue-request@ietf.org?subject=subscribe>
X-List-Received-Date: Thu, 01 May 2014 04:09:25 -0000

Hello Mark,

Please see my replies below.

Regards, Christian

On 30/04/2014 11:28 PM, Duckworth, Mark wrote:
> Hi Christian,
> Thanks for elaborating on the idea.  I have some more questions inline.
> I see basically what you are trying to do, but I still think it doesn't work in practice any better than what we already have, as I explain below.
> Mark
>
>> -----Original Message-----
>> From: Christian Groves [mailto:Christian.Groves@nteczone.com]
>> Sent: Monday, April 28, 2014 10:15 PM
>> To: Duckworth, Mark; clue@ietf.org
>> Subject: Re: [clue] Improving treatment of audio
>>
>> Hello Mark,
>>
>> Please see below.
>>
>> Regards, Christian
>>
>> On 29/04/2014 11:36 AM, Duckworth, Mark wrote:
>>> Hi Christian,
>>> I'm coming back to this message, because I'm trying to understand your
>> proposal.  You proposed "in each audio capture we say what VC it relates to."
>> What does this mean?  What does it mean for an audio capture to relate to a
>> video capture?  What exactly is the producer telling the consumer?  Can an
>> audio capture relate to more than one video capture?
>> [CNG] The meaning has largely the same semantic as if the ADV used the
>> same capture area on a VC and AC. Basically its saying that the audio comes
>> from the same "region" as what the video capture does. I think the
>> difference is that its not trying to put a mathematical certainty about what
>> that region is. Its a way for the producer to say to the consumer if you choose
>> video capture A you probably want to choose audio capture B for a good
>> experience.
> [Duckworth, Mark] If that is the meaning, then would it be better for the provider to advertise a list of audio captures that are "related to" each video capture?  Rather than the other way around?  Maybe I'm not understanding the problem you are trying to solve.
[CNG] I don't think it matters the way it is defined. video related to 
audio or audio related to video is essentially the same thing. The 
problem I'm trying to solve is how does the media consumer know which 
audio captures to choose when presented with a list of captures.

As I mentioned in an earlier email I think people where seeing the "area 
of capture" in two separate ways. The first way was using the area of 
capture to determine which video captures the audio related to (i.e. 
capture selection). The second was concentrating on how to use the area 
of capture information for audio rendering (i.e. capture rendering). I'm 
focusing on the first, capture selection.



>
>> In terms of whether an audio capture can relate to more than
>> one video capture yes I think so. You could have one audio capture for a
>> room and three video captures.
>>> More comments inline below.
>>>
>>> Regards,
>>> Mark
>>>
>>>> -----Original Message-----
>>>> From: Christian Groves [mailto:Christian.Groves@nteczone.com]
>>>> Sent: Sunday, April 13, 2014 11:18 PM
>>>> To: Duckworth, Mark; clue@ietf.org
>>>> Subject: Re: [clue] Improving treatment of audio
>>>>
>>>> Hello Mark,
>>>>
>>>> Sorry I didn't consider the entire 12.1.1. if I do that, according to
>>>> that
>>>> example:
>>>>
>>>> Video areas of capture:
>>>>
>>>>           bottom left    bottom right  top left         top right
>>>>       VC0 (-2011,2850,0) (-673,3000,0) (-2011,2850,757) (-673,3000,757)
>>>>       VC1 ( -673,3000,0) ( 673,3000,0) ( -673,3000,757) ( 673,3000,757)
>>>>       VC2 (  673,3000,0) (2011,2850,0) (  673,3000,757) (2011,3000,757)
>>>>       VC3 (-2011,2850,0) (2011,2850,0) (-2011,2850,757) (2011,3000,757)
>>>>       VC4 (-2011,2850,0) (2011,2850,0) (-2011,2850,757) (2011,3000,757)
>>>>       VC5 (-2011,2850,0) (2011,2850,0) (-2011,2850,757) (2011,3000,757)
>>>>       VC6 none
>>>>
>>>> Areas of capture for audio (from 12.1.1):
>>>>
>>>>           bottom left    bottom right  top left         top right
>>>>
>>>>       AC0 (-2011,2850,0) (-673,3000,0) (-2011,2850,757) (-673,3000,757)
>>>>       AC1 (  673,3000,0) (2011,2850,0) (  673,3000,757) (2011,3000,757)
>>>>       AC2 ( -673,3000,0) ( 673,3000,0) ( -673,3000,757) ( 673,3000,757)
>>>>       AC3 (-2011,2850,0) (2011,2850,0) (-2011,2850,757) (2011,3000,757)
>>>>       AC4 none
>>>>
>>>> Using a reference rather than area of capture:
>>>>        AC0 (VC0)
>>>>        AC1 (VC2)
>>>>        AC2 (VC1)
>>>>        AC3 (VC3,VC4,VC5)
>>>> I would think they are basically conveying the same information???
>>> [Duckworth, Mark] Looking from the consumer point of view, in trying to
>> understand the advertisement, what does the consumer know from the
>> advertisement?  Suppose the consumer wants to get all three of VC0, VC1,
>> VC2, but it only wants a single audio capture.  It should probably ask for AC3,
>> but according to the advertisement AC3 doesn't "relate to" VC0, VC1, or VC2.
>> [CNG] I agree. If the provider wanted to indicate that AC3 could also be used
>> for VC0,VC1,VC2 it could advertise:
>>        AC0 (VC0)
>>        AC1 (VC2)
>>        AC2 (VC1)
>>        AC3 (VC0,VC1,VC2,VC3,VC4,VC5)
> [Duckworth, Mark] I agree that  AC3 (the whole room audio) "relates to" all those video captures.  But earlier you said "the audio comes from the same 'region' as what the video capture does."  So now do we have to modify that to say something like "the audio comes from the same region as the video, or a subset of that region, or a superset, or an overlapping region"?  In this example, AC3 covers a region that is a superset of the VC0 region, a superset of the VC1 region, and a superset of the VC2 region.
[CNG] I was trying to illustrate the concept when I used the term 
"region". I didn't mean for that to relate to a specific area. I prefer 
"relates to" because that is essentially I'm trying specify. For the 
purposes of capture selection you could specify the relation in terms of 
a mathematical area or you can simply say the captures are related.

>
>>> [Duckworth, Mark] For a case with a different consumer, suppose the
>> consumer wants the single video capture VC5, but it would like to receive
>> and render spatial audio.  There is no audio capture with multiple channels,
>> so it could ask for AC0, AC1, and AC2 if it knew how they were spatially
>> "related to" VC5.  But according to the advertisement, AC0, AC1 and AC2
>> aren't "related to" VC5 at all.
>> [CNG] So in this case there is no AC3? In that case the advertisement could
>> look like:
>>        AC0 (VC0,VC5)
>>        AC1 (VC2,VC5)
>>        AC2 (VC1,VC5)
> [Duckworth, Mark] I didn't mean to remove AC3.  That would still be there in this example.  So now since the provider forms it's advertisement without knowing what the consumer wants to do, the provider has to give all this information:
>    AC0 relates to (VC0, VC3, VC4, VC5)
>    AC1 relates to (VC2, VC3, VC4, VC5)
>    AC2 relates to (VC1, VC3, VC4, VC5)
>    AC3 relates to (VC0, VC1, VC2, VC3, VC4, VC5)
> [Duckworth, Mark] Does this information really help the consumer?  I don't think it is any better than what we already have with area of capture.
[CNG] Well it does say to the MC a number of things, for exmaple:
- if you're going to choose VC0 then either AC0 and AC3 would be the 
best choice.
- AC3 is the best choice if you want one audio capture as it relates to 
all video captures.

My proposal was only duplicate what is supported by area of capture in 
terms of relating audio and video captures for capture selection. The 
way the discussion is going it seems that people are dubious about 
specifying capture area for audio. The proposal is trying to get around 
that.
Another option would be to indicate that the capture area when used for 
audio actually relates to a video capture region. However that seems to 
be confusing to people.

