From speechsc-bounces@ietf.org Tue Jul 05 03:26:17 2005
Received: from localhost.localdomain ([127.0.0.1] helo=megatron.ietf.org)
	by megatron.ietf.org with esmtp (Exim 4.32)
	id 1Dphov-0005zs-LJ; Tue, 05 Jul 2005 03:26:17 -0400
Received: from odin.ietf.org ([132.151.1.176] helo=ietf.org)
	by megatron.ietf.org with esmtp (Exim 4.32) id 1Dphol-0005y6-3E
	for speechsc@megatron.ietf.org; Tue, 05 Jul 2005 03:26:07 -0400
Received: from ietf-mx.ietf.org (ietf-mx [132.151.6.1])
	by ietf.org (8.9.1a/8.9.1a) with ESMTP id DAA07353
	for <speechsc@ietf.org>; Tue, 5 Jul 2005 03:26:05 -0400 (EDT)
Received: from pb-exchcon2.scansoft.com ([199.4.160.64])
	by ietf-mx.ietf.org with esmtp (Exim 4.33) id 1DpiFT-0001p7-LC
	for speechsc@ietf.org; Tue, 05 Jul 2005 03:53:44 -0400
Received: by pb-exchcon2.pb.scansoft.com with Internet Mail Service
	(5.5.2658.27) id <N5A2A3JZ>; Tue, 5 Jul 2005 03:25:47 -0400
Message-ID: <BBF29C9B95E52E4DB5C29A0ACC94E83B016AA095@ac-exch1.eu.scansoft.com>
From: "Reifenrath, Klaus" <Klaus.Reifenrath@Scansoft.com>
To: "'sarvi@cisco.com'" <sarvi@cisco.com>
Date: Tue, 5 Jul 2005 03:25:25 -0400 
MIME-Version: 1.0
X-Mailer: Internet Mail Service (5.5.2658.27)
X-Spam-Score: 0.7 (/)
X-Scan-Signature: 39bd8f8cbb76cae18b7e23f7cf6b2b9f
Cc: "'speechsc@ietf.org'" <speechsc@ietf.org>
Subject: [Speechsc] Obsolete reference to START-TIMERS
X-BeenThere: speechsc@ietf.org
X-Mailman-Version: 2.1.5
Precedence: list
List-Id: Speech Services Control Working Group <speechsc.ietf.org>
List-Unsubscribe: <https://www1.ietf.org/mailman/listinfo/speechsc>,
	<mailto:speechsc-request@ietf.org?subject=unsubscribe>
List-Post: <mailto:speechsc@ietf.org>
List-Help: <mailto:speechsc-request@ietf.org?subject=help>
List-Subscribe: <https://www1.ietf.org/mailman/listinfo/speechsc>,
	<mailto:speechsc-request@ietf.org?subject=subscribe>
Content-Type: multipart/mixed; boundary="===============0386563461=="
Sender: speechsc-bounces@ietf.org
Errors-To: speechsc-bounces@ietf.org

This message is in MIME format. Since your mail reader does not understand
this format, some or all of this message may not be legible.

--===============0386563461==
Content-Type: multipart/alternative;
	boundary="----_=_NextPart_001_01C58132.B6307D8A"

This message is in MIME format. Since your mail reader does not understand
this format, some or all of this message may not be legible.

------_=_NextPart_001_01C58132.B6307D8A
Content-Type: text/plain

In section 10.4 "(C) - START-TIMERS" need to be replaced by "(C) -
START-INPUT-TIMERS".  

Klaus

------_=_NextPart_001_01C58132.B6307D8A
Content-Type: text/html
Content-Transfer-Encoding: quoted-printable

<!DOCTYPE HTML PUBLIC "-//W3C//DTD HTML 3.2//EN">
<HTML>
<HEAD>
<META HTTP-EQUIV=3D"Content-Type" CONTENT=3D"text/html; =
charset=3Dus-ascii">
<META NAME=3D"Generator" CONTENT=3D"MS Exchange Server version =
5.5.2655.35">
<TITLE>Obsolete reference to START-TIMERS</TITLE>
</HEAD>
<BODY>

<P><FONT SIZE=3D2 FACE=3D"Arial">In section 10.4 &quot;(C) - =
START-TIMERS&quot; need to be replaced by &quot;(C) - =
START-INPUT-TIMERS&quot;.&nbsp; </FONT>
</P>

<P><FONT SIZE=3D2 FACE=3D"Arial">Klaus</FONT>
</P>

</BODY>
</HTML>
------_=_NextPart_001_01C58132.B6307D8A--


--===============0386563461==
Content-Type: text/plain; charset="us-ascii"
MIME-Version: 1.0
Content-Transfer-Encoding: 7bit
Content-Disposition: inline

_______________________________________________
Speechsc mailing list
Speechsc@ietf.org
https://www1.ietf.org/mailman/listinfo/speechsc

--===============0386563461==--




From speechsc-bounces@ietf.org Tue Jul 05 03:46:59 2005
Received: from localhost.localdomain ([127.0.0.1] helo=megatron.ietf.org)
	by megatron.ietf.org with esmtp (Exim 4.32)
	id 1Dpi8w-0002dX-Tw; Tue, 05 Jul 2005 03:46:58 -0400
Received: from odin.ietf.org ([132.151.1.176] helo=ietf.org)
	by megatron.ietf.org with esmtp (Exim 4.32) id 1Dpi8u-0002dQ-7B
	for speechsc@megatron.ietf.org; Tue, 05 Jul 2005 03:46:57 -0400
Received: from ietf-mx.ietf.org (ietf-mx [132.151.6.1])
	by ietf.org (8.9.1a/8.9.1a) with ESMTP id DAA11859
	for <speechsc@ietf.org>; Tue, 5 Jul 2005 03:46:54 -0400 (EDT)
Received: from pb-exchcon2.scansoft.com ([199.4.160.64])
	by ietf-mx.ietf.org with esmtp (Exim 4.33) id 1DpiZe-0003dD-68
	for speechsc@ietf.org; Tue, 05 Jul 2005 04:14:34 -0400
Received: by pb-exchcon2.pb.scansoft.com with Internet Mail Service
	(5.5.2658.27) id <N5A2A3MY>; Tue, 5 Jul 2005 03:46:47 -0400
Message-ID: <BBF29C9B95E52E4DB5C29A0ACC94E83B016AA098@ac-exch1.eu.scansoft.com>
From: "Reifenrath, Klaus" <Klaus.Reifenrath@Scansoft.com>
To: "'speechsc@ietf.org'" <speechsc@ietf.org>
Date: Tue, 5 Jul 2005 03:46:31 -0400 
MIME-Version: 1.0
X-Mailer: Internet Mail Service (5.5.2658.27)
X-Spam-Score: 0.9 (/)
X-Scan-Signature: f607d15ccc2bc4eaf3ade8ffa8af02a0
Subject: [Speechsc] START-OF-SPEECH in DTMF-only mode
X-BeenThere: speechsc@ietf.org
X-Mailman-Version: 2.1.5
Precedence: list
List-Id: Speech Services Control Working Group <speechsc.ietf.org>
List-Unsubscribe: <https://www1.ietf.org/mailman/listinfo/speechsc>,
	<mailto:speechsc-request@ietf.org?subject=unsubscribe>
List-Post: <mailto:speechsc@ietf.org>
List-Help: <mailto:speechsc-request@ietf.org?subject=help>
List-Subscribe: <https://www1.ietf.org/mailman/listinfo/speechsc>,
	<mailto:speechsc-request@ietf.org?subject=subscribe>
Content-Type: multipart/mixed; boundary="===============1328829717=="
Sender: speechsc-bounces@ietf.org
Errors-To: speechsc-bounces@ietf.org

This message is in MIME format. Since your mail reader does not understand
this format, some or all of this message may not be legible.

--===============1328829717==
Content-Type: multipart/alternative;
	boundary="----_=_NextPart_001_01C58135.79661952"

This message is in MIME format. Since your mail reader does not understand
this format, some or all of this message may not be legible.

------_=_NextPart_001_01C58135.79661952
Content-Type: text/plain

The current spec is not clear when START-OF-SPEECH need to be send in the
following scenarios:
A) The client requested a DTMF Recognizer. Is the START-OF-SPEECH event send
to the client also if speech was detected? 
B) The client requested a Speech Recognizer, but only activated DTMF
grammars. Is the START-OF-SPEECH event send to the client also if speech was
detected?
I think in both cases START-OF-SPEECH should only be send after detecting a
DTMF digit (see Figure 12 of VoiceXML 2.0:
http://www.w3.org/TR/voicexml20/#dmlATiming
<http://www.w3.org/TR/voicexml20/> ).

Klaus


------_=_NextPart_001_01C58135.79661952
Content-Type: text/html
Content-Transfer-Encoding: quoted-printable

<!DOCTYPE HTML PUBLIC "-//W3C//DTD HTML 3.2//EN">
<HTML>
<HEAD>
<META HTTP-EQUIV=3D"Content-Type" CONTENT=3D"text/html; =
charset=3Dus-ascii">
<META NAME=3D"Generator" CONTENT=3D"MS Exchange Server version =
5.5.2655.35">
<TITLE>START-OF-SPEECH in DTMF-only mode</TITLE>
</HEAD>
<BODY>

<P><FONT SIZE=3D2 FACE=3D"Arial">The current spec is not clear when =
START-OF-SPEECH need to be send in the following scenarios:</FONT>
<BR><FONT SIZE=3D2 FACE=3D"Arial">A) The client requested a DTMF =
Recognizer. Is the START-OF-SPEECH event send to the client also if =
speech was detected? </FONT></P>

<P><FONT SIZE=3D2 FACE=3D"Arial">B) The client requested a Speech =
Recognizer, but only activated DTMF grammars. Is the START-OF-SPEECH =
event send to the client also if speech was detected?</FONT></P>

<P><FONT SIZE=3D2 FACE=3D"Arial">I think in both cases START-OF-SPEECH =
should only be send after detecting a DTMF digit (see </FONT><FONT =
COLOR=3D"#0000FF" SIZE=3D2 FACE=3D"Arial">F</FONT><A =
HREF=3D"http://www.w3.org/TR/voicexml20/"><U></U><U></U><U><FONT =
COLOR=3D"#0000FF" SIZE=3D2 FACE=3D"Arial">igure 12 of VoiceXML 2.0: =
http://www.w3.org/TR/voicexml20/#dmlATiming</FONT></U></A><FONT =
FACE=3D"Times New Roman">).</FONT></P>

<P><FONT SIZE=3D2 FACE=3D"Arial">Klaus</FONT>
</P>

</BODY>
</HTML>
------_=_NextPart_001_01C58135.79661952--


--===============1328829717==
Content-Type: text/plain; charset="us-ascii"
MIME-Version: 1.0
Content-Transfer-Encoding: 7bit
Content-Disposition: inline

_______________________________________________
Speechsc mailing list
Speechsc@ietf.org
https://www1.ietf.org/mailman/listinfo/speechsc

--===============1328829717==--




From speechsc-bounces@ietf.org Tue Jul 05 08:49:07 2005
Received: from localhost.localdomain ([127.0.0.1] helo=megatron.ietf.org)
	by megatron.ietf.org with esmtp (Exim 4.32)
	id 1DpmrK-0000mw-NG; Tue, 05 Jul 2005 08:49:06 -0400
Received: from odin.ietf.org ([132.151.1.176] helo=ietf.org)
	by megatron.ietf.org with esmtp (Exim 4.32) id 1DpmrI-0000mG-LG
	for speechsc@megatron.ietf.org; Tue, 05 Jul 2005 08:49:05 -0400
Received: from ietf-mx.ietf.org (ietf-mx [132.151.6.1])
	by ietf.org (8.9.1a/8.9.1a) with ESMTP id IAA24848
	for <speechsc@ietf.org>; Tue, 5 Jul 2005 08:49:00 -0400 (EDT)
Received: from sj-iport-2-in.cisco.com ([171.71.176.71]
	helo=sj-iport-2.cisco.com) by ietf-mx.ietf.org with esmtp (Exim 4.33)
	id 1DpmuO-0004aj-0j
	for speechsc@ietf.org; Tue, 05 Jul 2005 08:52:17 -0400
Received: from sj-core-2.cisco.com (171.71.177.254)
	by sj-iport-2.cisco.com with ESMTP; 05 Jul 2005 05:24:26 -0700
Received: from imail.cisco.com (imail.cisco.com [128.107.200.91])
	by sj-core-2.cisco.com (8.12.10/8.12.6) with ESMTP id j65COLod014829;
	Tue, 5 Jul 2005 05:24:21 -0700 (PDT)
Received: from [10.32.245.152] (stealth-10-32-245-152.cisco.com
	[10.32.245.152])
	by imail.cisco.com (8.12.11/8.12.10) with SMTP id j65CNYV1032067;
	Tue, 5 Jul 2005 05:23:35 -0700
In-Reply-To: <BBF29C9B95E52E4DB5C29A0ACC94E83B016AA098@ac-exch1.eu.scansoft.com>
References: <BBF29C9B95E52E4DB5C29A0ACC94E83B016AA098@ac-exch1.eu.scansoft.com>
Mime-Version: 1.0 (Apple Message framework v730)
Content-Type: text/plain; charset=US-ASCII; delsp=yes; format=flowed
Message-Id: <D5BDBA1C-FA06-4FC5-A471-5FCCC786E9C8@cisco.com>
Content-Transfer-Encoding: 7bit
From: David R Oran <oran@cisco.com>
Subject: Re: [Speechsc] START-OF-SPEECH in DTMF-only mode
Date: Tue, 5 Jul 2005 08:24:23 -0400
To: Klaus Reifenrath <Klaus.Reifenrath@Scansoft.com>
X-Mailer: Apple Mail (2.730)
IIM-SIG: v:"1.1"; h:"imail.cisco.com"; d:"cisco.com"; z:"home"; m:"krs";
	t:"1120566215.561866"; x:"432200"; a:"rsa-sha1"; b:"nofws:1020";
	e:"Iw=="; n:"sQYarK2E51MdcTiUqeif3F7cWdxIfoCiXhdfb9vD5ee/j0jXL15gbFxF2p"
	"XIweAblu0N6XAgK7k+wrbr7bQDJaCDqOmzqpRUBjIRQAXQ7NzadpmR3pUL6wxaRUtW+c43sl9jC"
	"50Qg1sXHpPjt8Y+Y16ioyQAQAdSunM4YhevURc=";
	s:"gHG7lLJoD+JURa73wdD9Kptv7NH6lVON2bXrFnYVMGMRCh92frsfWoqB1pWWzhQB0kzmrGmx"
	"NJMJQ1yNXBrcJv1kZk56DxdWvBbclPrRRqIDkC8Ou6yI1c/gNd3LBs8C4WtaAJh0pODskqWXqP1"
	"eLIdLUb6toFKxxivW9Nal1Uk="; c:"From: David R Oran <oran@cisco.com>";
	c:"Subject: Re: [Speechsc] START-OF-SPEECH in DTMF-only mode";
	c:"Date: Tue, 5 Jul 2005 08:24:23 -0400"
IIM-VERIFY: s:"y"; v:"y"; r:"60"; h:"imail.cisco.com";
	c:"message from imail.cisco.com verified; "
X-Spam-Score: 0.0 (/)
X-Scan-Signature: e5ba305d0e64821bf3d8bc5d3bb07228
Content-Transfer-Encoding: 7bit
Cc: "'speechsc@ietf.org'" <speechsc@ietf.org>
X-BeenThere: speechsc@ietf.org
X-Mailman-Version: 2.1.5
Precedence: list
List-Id: Speech Services Control Working Group <speechsc.ietf.org>
List-Unsubscribe: <https://www1.ietf.org/mailman/listinfo/speechsc>,
	<mailto:speechsc-request@ietf.org?subject=unsubscribe>
List-Post: <mailto:speechsc@ietf.org>
List-Help: <mailto:speechsc-request@ietf.org?subject=help>
List-Subscribe: <https://www1.ietf.org/mailman/listinfo/speechsc>,
	<mailto:speechsc-request@ietf.org?subject=subscribe>
Sender: speechsc-bounces@ietf.org
Errors-To: speechsc-bounces@ietf.org


On Jul 5, 2005, at 3:46 AM, Reifenrath, Klaus wrote:

> The current spec is not clear when START-OF-SPEECH need to be send  
> in the following scenarios:
> A) The client requested a DTMF Recognizer. Is the START-OF-SPEECH  
> event send to the client also if speech was detected?
I suspect so, since one of the prime purposes is to enable client- 
mediated barge-in handling. However, if the recognizer is in fact  
only capable of recognizing DTMF then it may in fact not report  
anythin unless it's using some primitive thresholding machinery, like  
a SN threshold.
> B) The client requested a Speech Recognizer, but only activated  
> DTMF grammars. Is the START-OF-SPEECH event send to the client also  
> if speech was detected?
Again, I'd say yes, for the same reason as above.
> I think in both cases START-OF-SPEECH should only be send after  
> detecting a DTMF digit (see Figure 12 of VoiceXML 2.0: http:// 
> www.w3.org/TR/voicexml20/#dmlATiming).
We seem to have reached different conclusions. I'd be interested in  
why you think my analysis above is wrong.

Dave.

> Klaus
>
> _______________________________________________
> Speechsc mailing list
> Speechsc@ietf.org
> https://www1.ietf.org/mailman/listinfo/speechsc
>

_______________________________________________
Speechsc mailing list
Speechsc@ietf.org
https://www1.ietf.org/mailman/listinfo/speechsc



From speechsc-bounces@ietf.org Tue Jul 05 09:19:12 2005
Received: from localhost.localdomain ([127.0.0.1] helo=megatron.ietf.org)
	by megatron.ietf.org with esmtp (Exim 4.32)
	id 1DpnKR-0006hF-PD; Tue, 05 Jul 2005 09:19:11 -0400
Received: from odin.ietf.org ([132.151.1.176] helo=ietf.org)
	by megatron.ietf.org with esmtp (Exim 4.32) id 1DpnK8-0006Z0-PW
	for speechsc@megatron.ietf.org; Tue, 05 Jul 2005 09:18:53 -0400
Received: from ietf-mx.ietf.org (ietf-mx [132.151.6.1])
	by ietf.org (8.9.1a/8.9.1a) with ESMTP id JAA14334
	for <speechsc@ietf.org>; Tue, 5 Jul 2005 09:18:48 -0400 (EDT)
Received: from fw01.db01.voxpilot.com ([212.17.54.82] helo=mail.voxpilot.com)
	by ietf-mx.ietf.org with esmtp (Exim 4.33) id 1DpjRo-00052n-GZ
	for speechsc@ietf.org; Tue, 05 Jul 2005 05:10:33 -0400
Received: from daburkewxp (unknown [10.0.0.203])
	by mail.voxpilot.com (Postfix) with ESMTP
	id 1F58D214042; Tue,  5 Jul 2005 08:42:38 +0000 (GMT)
Message-ID: <034b01c5813d$7efbb360$cb00000a@db01.voxpilot.com>
From: "Dave Burke" <david.burke@voxpilot.com>
To: "Reifenrath, Klaus" <Klaus.Reifenrath@Scansoft.com>, <speechsc@ietf.org>
References: <BBF29C9B95E52E4DB5C29A0ACC94E83B016AA098@ac-exch1.eu.scansoft.com>
Subject: Re: [Speechsc] START-OF-SPEECH in DTMF-only mode
Date: Tue, 5 Jul 2005 09:42:37 +0100
MIME-Version: 1.0
X-Priority: 3
X-MSMail-Priority: Normal
X-Mailer: Microsoft Outlook Express 6.00.2900.2180
X-MimeOLE: Produced By Microsoft MimeOLE V6.00.2900.2180
X-Spam-Score: 0.9 (/)
X-Scan-Signature: 3fbd9b434023f8abfcb1532abaec7a21
Cc: 
X-BeenThere: speechsc@ietf.org
X-Mailman-Version: 2.1.5
Precedence: list
List-Id: Speech Services Control Working Group <speechsc.ietf.org>
List-Unsubscribe: <https://www1.ietf.org/mailman/listinfo/speechsc>,
	<mailto:speechsc-request@ietf.org?subject=unsubscribe>
List-Post: <mailto:speechsc@ietf.org>
List-Help: <mailto:speechsc-request@ietf.org?subject=help>
List-Subscribe: <https://www1.ietf.org/mailman/listinfo/speechsc>,
	<mailto:speechsc-request@ietf.org?subject=subscribe>
Content-Type: multipart/mixed; boundary="===============1544096530=="
Sender: speechsc-bounces@ietf.org
Errors-To: speechsc-bounces@ietf.org

This is a multi-part message in MIME format.

--===============1544096530==
Content-Type: multipart/alternative;
	boundary="----=_NextPart_000_0348_01C58145.E09B7C60"

This is a multi-part message in MIME format.

------=_NextPart_000_0348_01C58145.E09B7C60
Content-Type: text/plain;
	charset="iso-8859-1"
Content-Transfer-Encoding: quoted-printable

START-OF-SPEECH in DTMF-only modeGood catch.=20

There has to be a way to select the input mode to dtmf, speech, or both. =
For example, an application in a noisy environment might fall back to =
DTMF (but still use its speechrecog resource). If (B) were true, then =
the fall back would not work due to bargin from noise!

For me, choosing a dtmfrecog implies you only want a DTMF input mode. =
Choosing a speechrecog is ambigious - there are now 3 choices of input =
mode.=20

Using the mode attribute of the activated grammars to decide the input =
mode will work but this is NOT what VoiceXML does. For example, I am not =
sure how MRCP would then let me implement "Grammar activation is not =
affected by the inputmodes property. For instance, if the inputmodes =
property restricts input to just voice, DTMF grammars will still be =
activated, but cannot be matched."

Sounds like a case of YAMHR (yet another MRCP header required).

Dave
  ----- Original Message -----=20
  From: Reifenrath, Klaus=20
  To: 'speechsc@ietf.org'=20
  Sent: Tuesday, July 05, 2005 8:46 AM
  Subject: [Speechsc] START-OF-SPEECH in DTMF-only mode


  The current spec is not clear when START-OF-SPEECH need to be send in =
the following scenarios:=20
  A) The client requested a DTMF Recognizer. Is the START-OF-SPEECH =
event send to the client also if speech was detected?=20

  B) The client requested a Speech Recognizer, but only activated DTMF =
grammars. Is the START-OF-SPEECH event send to the client also if speech =
was detected?

  I think in both cases START-OF-SPEECH should only be send after =
detecting a DTMF digit (see Figure 12 of VoiceXML 2.0: =
http://www.w3.org/TR/voicexml20/#dmlATiming).

  Klaus=20



-------------------------------------------------------------------------=
-----


  _______________________________________________
  Speechsc mailing list
  Speechsc@ietf.org
  https://www1.ietf.org/mailman/listinfo/speechsc

------=_NextPart_000_0348_01C58145.E09B7C60
Content-Type: text/html;
	charset="iso-8859-1"
Content-Transfer-Encoding: quoted-printable

<!DOCTYPE HTML PUBLIC "-//W3C//DTD HTML 4.0 Transitional//EN">
<HTML><HEAD><TITLE>START-OF-SPEECH in DTMF-only mode</TITLE>
<META http-equiv=3DContent-Type content=3D"text/html; =
charset=3Diso-8859-1">
<META content=3D"MSHTML 6.00.2900.2668" name=3DGENERATOR>
<STYLE></STYLE>
</HEAD>
<BODY bgColor=3D#ffffff>
<DIV><FONT face=3DArial size=3D2>Good catch. </FONT></DIV>
<DIV><FONT face=3DArial size=3D2></FONT>&nbsp;</DIV>
<DIV><FONT face=3DArial size=3D2>There has to be a way to select the =
input mode to=20
dtmf, speech, or both. For example, an application in a noisy =
environment might=20
fall back to DTMF (but still use its speechrecog resource). If (B) were =
true,=20
then the fall back would not work due to bargin from noise!</FONT></DIV>
<DIV><FONT face=3DArial size=3D2></FONT>&nbsp;</DIV>
<DIV><FONT face=3DArial size=3D2>For me,&nbsp;choosing a dtmfrecog =
implies you only=20
want a DTMF input mode. Choosing a speechrecog is ambigious - there are =
now 3=20
choices of input mode. </FONT></DIV>
<DIV><FONT face=3DArial size=3D2></FONT>&nbsp;</DIV>
<DIV><FONT face=3DArial size=3D2>Using the mode attribute&nbsp;of the =
activated=20
grammars to decide the input mode will work but this is NOT what =
VoiceXML does.=20
For example, I am not sure how MRCP would then let me implement "Grammar =

activation is not affected by the inputmodes property. For instance, if =
the=20
inputmodes property restricts input to just voice, DTMF grammars will =
still be=20
activated, but cannot be matched."</FONT></DIV>
<DIV><FONT face=3DArial size=3D2></FONT>&nbsp;</DIV>
<DIV><FONT face=3DArial size=3D2>Sounds like a case of YAMHR (yet =
another MRCP=20
header required).</FONT></DIV>
<DIV>&nbsp;</DIV>
<DIV><FONT face=3DArial size=3D2>Dave</FONT></DIV>
<BLOCKQUOTE=20
style=3D"PADDING-RIGHT: 0px; PADDING-LEFT: 5px; MARGIN-LEFT: 5px; =
BORDER-LEFT: #000000 2px solid; MARGIN-RIGHT: 0px">
  <DIV style=3D"FONT: 10pt arial">----- Original Message ----- </DIV>
  <DIV=20
  style=3D"BACKGROUND: #e4e4e4; FONT: 10pt arial; font-color: =
black"><B>From:</B>=20
  <A title=3DKlaus.Reifenrath@Scansoft.com=20
  href=3D"mailto:Klaus.Reifenrath@Scansoft.com">Reifenrath, Klaus</A> =
</DIV>
  <DIV style=3D"FONT: 10pt arial"><B>To:</B> <A =
title=3Dspeechsc@ietf.org=20
  href=3D"mailto:'speechsc@ietf.org'">'speechsc@ietf.org'</A> </DIV>
  <DIV style=3D"FONT: 10pt arial"><B>Sent:</B> Tuesday, July 05, 2005 =
8:46=20
AM</DIV>
  <DIV style=3D"FONT: 10pt arial"><B>Subject:</B> [Speechsc] =
START-OF-SPEECH in=20
  DTMF-only mode</DIV>
  <DIV><FONT face=3DArial size=3D2></FONT><BR></DIV>
  <P><FONT face=3DArial size=3D2>The current spec is not clear when =
START-OF-SPEECH=20
  need to be send in the following scenarios:</FONT> <BR><FONT =
face=3DArial=20
  size=3D2>A) The client requested a DTMF Recognizer. Is the =
START-OF-SPEECH event=20
  send to the client also if speech was detected? </FONT></P>
  <P><FONT face=3DArial size=3D2>B) The client requested a Speech =
Recognizer, but=20
  only activated DTMF grammars. Is the START-OF-SPEECH event send to the =
client=20
  also if speech was detected?</FONT></P>
  <P><FONT face=3DArial size=3D2>I think in both cases START-OF-SPEECH =
should only=20
  be send after detecting a DTMF digit (see </FONT><FONT face=3DArial=20
  color=3D#0000ff size=3D2>F</FONT><A=20
  href=3D"http://www.w3.org/TR/voicexml20/"><U></U><U></U><U><FONT =
face=3DArial=20
  color=3D#0000ff size=3D2>igure 12 of VoiceXML 2.0:=20
  http://www.w3.org/TR/voicexml20/#dmlATiming</FONT></U></A><FONT=20
  face=3D"Times New Roman">).</FONT></P>
  <P><FONT face=3DArial size=3D2>Klaus</FONT> </P>
  <P>
  <HR>

  <P></P>_______________________________________________<BR>Speechsc =
mailing=20
  =
list<BR>Speechsc@ietf.org<BR>https://www1.ietf.org/mailman/listinfo/speec=
hsc<BR></BLOCKQUOTE></BODY></HTML>

------=_NextPart_000_0348_01C58145.E09B7C60--



--===============1544096530==
Content-Type: text/plain; charset="us-ascii"
MIME-Version: 1.0
Content-Transfer-Encoding: 7bit
Content-Disposition: inline

_______________________________________________
Speechsc mailing list
Speechsc@ietf.org
https://www1.ietf.org/mailman/listinfo/speechsc

--===============1544096530==--





From speechsc-bounces@ietf.org Tue Jul 05 11:40:43 2005
Received: from localhost.localdomain ([127.0.0.1] helo=megatron.ietf.org)
	by megatron.ietf.org with esmtp (Exim 4.32)
	id 1DppXM-0000VH-CL; Tue, 05 Jul 2005 11:40:40 -0400
Received: from odin.ietf.org ([132.151.1.176] helo=ietf.org)
	by megatron.ietf.org with esmtp (Exim 4.32) id 1DppX9-0000OH-V0
	for speechsc@megatron.ietf.org; Tue, 05 Jul 2005 11:40:30 -0400
Received: from ietf-mx.ietf.org (ietf-mx [132.151.6.1])
	by ietf.org (8.9.1a/8.9.1a) with ESMTP id LAA11146
	for <speechsc@ietf.org>; Tue, 5 Jul 2005 11:40:25 -0400 (EDT)
Received: from letter.nuance.com ([207.107.210.132])
	by ietf-mx.ietf.org with esmtp (Exim 4.33) id 1Dppw2-0001eP-BM
	for speechsc@ietf.org; Tue, 05 Jul 2005 12:06:11 -0400
Received: from postcard.nuance.com ([10.3.6.20]:16689)
	by letter.nuance.com with esmtp id 1DppV0-0007sj-Jx;
	Tue, 05 Jul 2005 08:38:14 -0700
Received: from mtb1exch01.nuance.com ([10.3.2.6]) by postcard.nuance.com with
	Microsoft SMTPSVC(6.0.3790.0); Tue, 5 Jul 2005 11:38:09 -0400
X-MimeOLE: Produced By Microsoft Exchange V6.5.7226.0
Content-class: urn:content-classes:message
MIME-Version: 1.0
Content-Type: text/plain;
	charset="us-ascii"
Content-Transfer-Encoding: quoted-printable
Subject: RE: [Speechsc] START-OF-SPEECH in DTMF-only mode
Date: Tue, 5 Jul 2005 11:38:07 -0400
Message-ID: <7DE7C4EF3B7C8B4B82955191378290D802ED3CE2@mtb1exch01.nuance.com>
Thread-Topic: [Speechsc] START-OF-SPEECH in DTMF-only mode
Thread-Index: AcWBYODQKOULysVUQeeQX5iGXSpUGQAFlHkA
From: "Pierre Forgues" <forgues@nuance.com>
To: "David R Oran" <oran@cisco.com>,
	"Klaus Reifenrath" <Klaus.Reifenrath@Scansoft.com>
X-OriginalArrivalTime: 05 Jul 2005 15:38:09.0412 (UTC)
	FILETIME=[8BA8E440:01C58177]
X-FromHost: postcard.nuance.com [10.3.6.20]:16689
Lines: 57
X-Spam-Score: 0.0 (/)
X-Scan-Signature: 5a9a1bd6c2d06a21d748b7d0070ddcb8
Content-Transfer-Encoding: quoted-printable
Cc: speechsc@ietf.org
X-BeenThere: speechsc@ietf.org
X-Mailman-Version: 2.1.5
Precedence: list
List-Id: Speech Services Control Working Group <speechsc.ietf.org>
List-Unsubscribe: <https://www1.ietf.org/mailman/listinfo/speechsc>,
	<mailto:speechsc-request@ietf.org?subject=unsubscribe>
List-Post: <mailto:speechsc@ietf.org>
List-Help: <mailto:speechsc-request@ietf.org?subject=help>
List-Subscribe: <https://www1.ietf.org/mailman/listinfo/speechsc>,
	<mailto:speechsc-request@ietf.org?subject=subscribe>
Sender: speechsc-bounces@ietf.org
Errors-To: speechsc-bounces@ietf.org

The START-OF-SPEECH event should be sent to the client whether the user
spoke or pressed DTMF.  I agree with Dave Oran's comment that it applies
to both scenarios mentioned below.

The action to take on the client side is likely going to be the same in
both cases, i.e. to stop the prompt.

Pierre

-----Original Message-----
From: speechsc-bounces@ietf.org [mailto:speechsc-bounces@ietf.org] On
Behalf Of David R Oran
Sent: Tuesday, July 05, 2005 8:24 AM
To: Klaus Reifenrath
Cc: 'speechsc@ietf.org'
Subject: Re: [Speechsc] START-OF-SPEECH in DTMF-only mode


On Jul 5, 2005, at 3:46 AM, Reifenrath, Klaus wrote:

> The current spec is not clear when START-OF-SPEECH need to be send =20
> in the following scenarios:
> A) The client requested a DTMF Recognizer. Is the START-OF-SPEECH =20
> event send to the client also if speech was detected?
I suspect so, since one of the prime purposes is to enable client-=20
mediated barge-in handling. However, if the recognizer is in fact =20
only capable of recognizing DTMF then it may in fact not report =20
anythin unless it's using some primitive thresholding machinery, like =20
a SN threshold.
> B) The client requested a Speech Recognizer, but only activated =20
> DTMF grammars. Is the START-OF-SPEECH event send to the client also =20
> if speech was detected?
Again, I'd say yes, for the same reason as above.
> I think in both cases START-OF-SPEECH should only be send after =20
> detecting a DTMF digit (see Figure 12 of VoiceXML 2.0: http://=20
> www.w3.org/TR/voicexml20/#dmlATiming).
We seem to have reached different conclusions. I'd be interested in =20
why you think my analysis above is wrong.

Dave.

> Klaus
>
> _______________________________________________
> Speechsc mailing list
> Speechsc@ietf.org
> https://www1.ietf.org/mailman/listinfo/speechsc
>

_______________________________________________
Speechsc mailing list
Speechsc@ietf.org
https://www1.ietf.org/mailman/listinfo/speechsc

=20
 =20


_______________________________________________
Speechsc mailing list
Speechsc@ietf.org
https://www1.ietf.org/mailman/listinfo/speechsc



From speechsc-bounces@ietf.org Tue Jul 05 12:36:59 2005
Received: from localhost.localdomain ([127.0.0.1] helo=megatron.ietf.org)
	by megatron.ietf.org with esmtp (Exim 4.32)
	id 1DpqPr-0002dz-2k; Tue, 05 Jul 2005 12:36:59 -0400
Received: from odin.ietf.org ([132.151.1.176] helo=ietf.org)
	by megatron.ietf.org with esmtp (Exim 4.32) id 1DpqJx-0008Sc-SQ
	for speechsc@megatron.ietf.org; Tue, 05 Jul 2005 12:30:54 -0400
Received: from ietf-mx.ietf.org (ietf-mx [132.151.6.1])
	by ietf.org (8.9.1a/8.9.1a) with ESMTP id MAA18251
	for <speechsc@ietf.org>; Tue, 5 Jul 2005 12:30:46 -0400 (EDT)
Received: from sj-iport-4.cisco.com ([171.68.10.86])
	by ietf-mx.ietf.org with esmtp (Exim 4.33) id 1DpqcZ-0005Tk-MZ
	for speechsc@ietf.org; Tue, 05 Jul 2005 12:50:10 -0400
Received: from sj-core-3.cisco.com (171.68.223.137)
	by sj-iport-4.cisco.com with ESMTP; 05 Jul 2005 09:22:15 -0700
Received: from vtg-um-e2k6.sj21ad.cisco.com (vtg-um-e2k6.cisco.com
	[171.70.93.77])
	by sj-core-3.cisco.com (8.12.10/8.12.6) with ESMTP id j65GMC6p019439;
	Tue, 5 Jul 2005 09:22:13 -0700 (PDT)
Content-class: urn:content-classes:message
MIME-Version: 1.0
Content-Type: text/plain;
	charset="us-ascii"
Content-Transfer-Encoding: quoted-printable
X-MimeOLE: Produced By Microsoft Exchange V6.5.7226.0
Subject: RE: [Speechsc] START-OF-SPEECH in DTMF-only mode
Date: Tue, 5 Jul 2005 09:22:12 -0700
Message-ID: <03772D1EC8DE624A863058C75874A75C112692@vtg-um-e2k6.sj21ad.cisco.com>
Thread-Topic: [Speechsc] START-OF-SPEECH in DTMF-only mode
Thread-Index: AcWBYQywOyk/8ZziSciIVsVduk7X1AAHB1WQ
From: "Shanmugham, Saravanan" <sarvi@cisco.com>
To: "David R Oran" <oran@cisco.com>,
	"Klaus Reifenrath" <Klaus.Reifenrath@Scansoft.com>
X-Spam-Score: 0.0 (/)
X-Scan-Signature: f607d15ccc2bc4eaf3ade8ffa8af02a0
Content-Transfer-Encoding: quoted-printable
Cc: speechsc@ietf.org
X-BeenThere: speechsc@ietf.org
X-Mailman-Version: 2.1.5
Precedence: list
List-Id: Speech Services Control Working Group <speechsc.ietf.org>
List-Unsubscribe: <https://www1.ietf.org/mailman/listinfo/speechsc>,
	<mailto:speechsc-request@ietf.org?subject=unsubscribe>
List-Post: <mailto:speechsc@ietf.org>
List-Help: <mailto:speechsc-request@ietf.org?subject=help>
List-Subscribe: <https://www1.ietf.org/mailman/listinfo/speechsc>,
	<mailto:speechsc-request@ietf.org?subject=subscribe>
Sender: speechsc-bounces@ietf.org
Errors-To: speechsc-bounces@ietf.org

I agree with Dave's analysis. The purpose of this event was barge-in.=20
And barge-in should happen for both DTMF and speech.

Is there a case where you think it should not behave this way. If soe,
please provide a scenario where you think
    1. Barge-in should happen for DTMF and not voice or vice-versa.
    2. You would benefit from the client knowing what caused the
barge-in, DTMF Vs speech.

Sarvi

     -----Original Message-----
     From: speechsc-bounces@ietf.org=20
     [mailto:speechsc-bounces@ietf.org] On Behalf Of David R Oran
     Sent: Tuesday, July 05, 2005 5:24 AM
     To: Klaus Reifenrath
     Cc: 'speechsc@ietf.org'
     Subject: Re: [Speechsc] START-OF-SPEECH in DTMF-only mode
    =20
    =20
     On Jul 5, 2005, at 3:46 AM, Reifenrath, Klaus wrote:
    =20
     > The current spec is not clear when START-OF-SPEECH need=20
     to be send in=20
     > the following scenarios:
     > A) The client requested a DTMF Recognizer. Is the=20
     START-OF-SPEECH=20
     > event send to the client also if speech was detected?
     I suspect so, since one of the prime purposes is to enable=20
     client- mediated barge-in handling. However, if the=20
     recognizer is in fact only capable of recognizing DTMF=20
     then it may in fact not report anythin unless it's using=20
     some primitive thresholding machinery, like a SN threshold.
     > B) The client requested a Speech Recognizer, but only=20
     activated DTMF=20
     > grammars. Is the START-OF-SPEECH event send to the=20
     client also if=20
     > speech was detected?
     Again, I'd say yes, for the same reason as above.
     > I think in both cases START-OF-SPEECH should only be send after=20
     > detecting a DTMF digit (see Figure 12 of VoiceXML 2.0: http://=20
     > www.w3.org/TR/voicexml20/#dmlATiming).
     We seem to have reached different conclusions. I'd be=20
     interested in why you think my analysis above is wrong.
    =20
     Dave.
    =20
     > Klaus
     >
     > _______________________________________________
     > Speechsc mailing list
     > Speechsc@ietf.org
     > https://www1.ietf.org/mailman/listinfo/speechsc
     >
    =20
     _______________________________________________
     Speechsc mailing list
     Speechsc@ietf.org
     https://www1.ietf.org/mailman/listinfo/speechsc
    =20

_______________________________________________
Speechsc mailing list
Speechsc@ietf.org
https://www1.ietf.org/mailman/listinfo/speechsc



From speechsc-bounces@ietf.org Tue Jul 05 13:04:20 2005
Received: from localhost.localdomain ([127.0.0.1] helo=megatron.ietf.org)
	by megatron.ietf.org with esmtp (Exim 4.32)
	id 1DpqqJ-0003Tu-DK; Tue, 05 Jul 2005 13:04:19 -0400
Received: from odin.ietf.org ([132.151.1.176] helo=ietf.org)
	by megatron.ietf.org with esmtp (Exim 4.32) id 1Dpqq8-0003Qm-A2
	for speechsc@megatron.ietf.org; Tue, 05 Jul 2005 13:04:08 -0400
Received: from ietf-mx.ietf.org (ietf-mx [132.151.6.1])
	by ietf.org (8.9.1a/8.9.1a) with ESMTP id NAA23855
	for <speechsc@ietf.org>; Tue, 5 Jul 2005 13:04:03 -0400 (EDT)
Received: from fw01.db01.voxpilot.com ([212.17.54.82] helo=mail.voxpilot.com)
	by ietf-mx.ietf.org with esmtp (Exim 4.33) id 1Dpr7T-0008J5-1J
	for speechsc@ietf.org; Tue, 05 Jul 2005 13:22:04 -0400
Received: from daburkewxp (unknown [10.0.0.203])
	by mail.voxpilot.com (Postfix) with ESMTP
	id D4944214042; Tue,  5 Jul 2005 16:54:15 +0000 (GMT)
Message-ID: <052f01c58182$2d78f300$cb00000a@db01.voxpilot.com>
From: "Dave Burke" <david.burke@voxpilot.com>
To: "Shanmugham, Saravanan" <sarvi@cisco.com>, "David R Oran" <oran@cisco.com>,
	"Klaus Reifenrath" <Klaus.Reifenrath@Scansoft.com>
References: <03772D1EC8DE624A863058C75874A75C112692@vtg-um-e2k6.sj21ad.cisco.com>
Subject: Re: [Speechsc] START-OF-SPEECH in DTMF-only mode
Date: Tue, 5 Jul 2005 17:54:15 +0100
MIME-Version: 1.0
Content-Type: text/plain; format=flowed; charset="iso-8859-1";
	reply-type=original
Content-Transfer-Encoding: 7bit
X-Priority: 3
X-MSMail-Priority: Normal
X-Mailer: Microsoft Outlook Express 6.00.2900.2180
X-MimeOLE: Produced By Microsoft MimeOLE V6.00.2900.2180
X-Spam-Score: 0.0 (/)
X-Scan-Signature: f66b12316365a3fe519e75911daf28a8
Content-Transfer-Encoding: 7bit
Cc: speechsc@ietf.org
X-BeenThere: speechsc@ietf.org
X-Mailman-Version: 2.1.5
Precedence: list
List-Id: Speech Services Control Working Group <speechsc.ietf.org>
List-Unsubscribe: <https://www1.ietf.org/mailman/listinfo/speechsc>,
	<mailto:speechsc-request@ietf.org?subject=unsubscribe>
List-Post: <mailto:speechsc@ietf.org>
List-Help: <mailto:speechsc-request@ietf.org?subject=help>
List-Subscribe: <https://www1.ietf.org/mailman/listinfo/speechsc>,
	<mailto:speechsc-request@ietf.org?subject=subscribe>
Sender: speechsc-bounces@ietf.org
Errors-To: speechsc-bounces@ietf.org

Inline.

Dave

----- Original Message ----- 
From: "Shanmugham, Saravanan" <sarvi@cisco.com>
To: "David R Oran" <oran@cisco.com>; "Klaus Reifenrath" 
<Klaus.Reifenrath@Scansoft.com>
Cc: <speechsc@ietf.org>
Sent: Tuesday, July 05, 2005 5:22 PM
Subject: RE: [Speechsc] START-OF-SPEECH in DTMF-only mode


I agree with Dave's analysis. The purpose of this event was barge-in.
And barge-in should happen for both DTMF and speech.

Is there a case where you think it should not behave this way. If soe,
please provide a scenario where you think
    1. Barge-in should happen for DTMF and not voice or vice-versa.

DB> You want to do a DTMF recognition only because it is noisy. While 
waiting for DTMF input, the speechrecog resource (or advanced dtmfrecog) 
generates a START-OF-SPEECH because it heard some speech. The client does 
not want to stop prompt playing unless DTMF was heard but it can't tell by 
the START-OF-SPEECH whether speech or DTMF was heard. Similarly vice versa.

    2. You would benefit from the client knowing what caused the
barge-in, DTMF Vs speech.

DB> See previous comment. And previous e-mail: either add an inputmodes 
header (taking value speech, dtmf, both) to the RECOGNIZE request or add a 
header to the START-OF-SPEECH event indicating DTMF or speech.

Sarvi

     -----Original Message-----
     From: speechsc-bounces@ietf.org
     [mailto:speechsc-bounces@ietf.org] On Behalf Of David R Oran
     Sent: Tuesday, July 05, 2005 5:24 AM
     To: Klaus Reifenrath
     Cc: 'speechsc@ietf.org'
     Subject: Re: [Speechsc] START-OF-SPEECH in DTMF-only mode


     On Jul 5, 2005, at 3:46 AM, Reifenrath, Klaus wrote:

     > The current spec is not clear when START-OF-SPEECH need
     to be send in
     > the following scenarios:
     > A) The client requested a DTMF Recognizer. Is the
     START-OF-SPEECH
     > event send to the client also if speech was detected?
     I suspect so, since one of the prime purposes is to enable
     client- mediated barge-in handling. However, if the
     recognizer is in fact only capable of recognizing DTMF
     then it may in fact not report anythin unless it's using
     some primitive thresholding machinery, like a SN threshold.
     > B) The client requested a Speech Recognizer, but only
     activated DTMF
     > grammars. Is the START-OF-SPEECH event send to the
     client also if
     > speech was detected?
     Again, I'd say yes, for the same reason as above.
     > I think in both cases START-OF-SPEECH should only be send after
     > detecting a DTMF digit (see Figure 12 of VoiceXML 2.0: http://
     > www.w3.org/TR/voicexml20/#dmlATiming).
     We seem to have reached different conclusions. I'd be
     interested in why you think my analysis above is wrong.

     Dave.

     > Klaus
     >
     > _______________________________________________
     > Speechsc mailing list
     > Speechsc@ietf.org
     > https://www1.ietf.org/mailman/listinfo/speechsc
     >

     _______________________________________________
     Speechsc mailing list
     Speechsc@ietf.org
     https://www1.ietf.org/mailman/listinfo/speechsc


_______________________________________________
Speechsc mailing list
Speechsc@ietf.org
https://www1.ietf.org/mailman/listinfo/speechsc


_______________________________________________
Speechsc mailing list
Speechsc@ietf.org
https://www1.ietf.org/mailman/listinfo/speechsc



From speechsc-bounces@ietf.org Tue Jul 05 13:20:05 2005
Received: from localhost.localdomain ([127.0.0.1] helo=megatron.ietf.org)
	by megatron.ietf.org with esmtp (Exim 4.32)
	id 1Dpr5Y-0006nH-QM; Tue, 05 Jul 2005 13:20:04 -0400
Received: from odin.ietf.org ([132.151.1.176] helo=ietf.org)
	by megatron.ietf.org with esmtp (Exim 4.32) id 1Dpr5U-0006hO-5z
	for speechsc@megatron.ietf.org; Tue, 05 Jul 2005 13:20:01 -0400
Received: from ietf-mx.ietf.org (ietf-mx [132.151.6.1])
	by ietf.org (8.9.1a/8.9.1a) with ESMTP id NAA25491
	for <speechsc@ietf.org>; Tue, 5 Jul 2005 13:19:56 -0400 (EDT)
Received: from sj-iport-5.cisco.com ([171.68.10.87])
	by ietf-mx.ietf.org with esmtp (Exim 4.33) id 1DprUe-0001i8-4S
	for speechsc@ietf.org; Tue, 05 Jul 2005 13:46:00 -0400
Received: from sj-core-3.cisco.com (171.68.223.137)
	by sj-iport-5.cisco.com with ESMTP; 05 Jul 2005 10:18:07 -0700
X-IronPort-AV: i="3.93,261,1115017200"; 
	d="scan'208"; a="196414407:sNHT53070004"
Received: from vtg-um-e2k6.sj21ad.cisco.com (vtg-um-e2k6.cisco.com
	[171.70.93.77])
	by sj-core-3.cisco.com (8.12.10/8.12.6) with ESMTP id j65HI46p013563;
	Tue, 5 Jul 2005 10:18:04 -0700 (PDT)
Content-class: urn:content-classes:message
MIME-Version: 1.0
Content-Type: text/plain;
	charset="us-ascii"
Content-Transfer-Encoding: quoted-printable
X-MimeOLE: Produced By Microsoft Exchange V6.5.7226.0
Subject: RE: [Speechsc] START-OF-SPEECH in DTMF-only mode
Date: Tue, 5 Jul 2005 10:18:03 -0700
Message-ID: <03772D1EC8DE624A863058C75874A75C1126AD@vtg-um-e2k6.sj21ad.cisco.com>
Thread-Topic: [Speechsc] START-OF-SPEECH in DTMF-only mode
Thread-Index: AcWBgjwBvP5/iTu5R72fvf7+5Qgx4wAAH1JA
From: "Shanmugham, Saravanan" <sarvi@cisco.com>
To: "Dave Burke" <david.burke@voxpilot.com>, "David R Oran" <oran@cisco.com>, 
	"Klaus Reifenrath" <Klaus.Reifenrath@Scansoft.com>
X-Spam-Score: 0.0 (/)
X-Scan-Signature: 36b1f8810cb91289d885dc8ab4fc8172
Content-Transfer-Encoding: quoted-printable
Cc: speechsc@ietf.org
X-BeenThere: speechsc@ietf.org
X-Mailman-Version: 2.1.5
Precedence: list
List-Id: Speech Services Control Working Group <speechsc.ietf.org>
List-Unsubscribe: <https://www1.ietf.org/mailman/listinfo/speechsc>,
	<mailto:speechsc-request@ietf.org?subject=unsubscribe>
List-Post: <mailto:speechsc@ietf.org>
List-Help: <mailto:speechsc-request@ietf.org?subject=help>
List-Subscribe: <https://www1.ietf.org/mailman/listinfo/speechsc>,
	<mailto:speechsc-request@ietf.org?subject=subscribe>
Sender: speechsc-bounces@ietf.org
Errors-To: speechsc-bounces@ietf.org

Barge-in does not necessarily happen only through the START-OF-SPEECH
event. Note that if both the ASR and TTS resources are on the same box
they could optimize internally for better performance.

So I think what you are suggesting that is that the recognizer should be
sensitive to what input is being chosen, when deciding when to generate
the START-OF-SPEECH event.

I was going to say, this could be based on the grammar and whether it
contains tokens other than digits or not. But I guess even if they
contained only digits we could still either speak it or press it and the
recognizer should be able to deal with both.

In which case the problem you are raising becomes less of a barge-in
issue. It more about restricting the recognizer to do recognition in a
speech-only or dtmf-only or both modes and whether we want to allow it.=20

If we want to DTMF-only, we could always allocate a DTMF-only recognizer
and thus not waste precious speech resources for that channel. If we
want to do both a regular recognizer resource can do both. The question
then becomes, is there a case where you may want to do speech only
recognition or not. Is there a use case for this restriction. The noisy
channel argument does not apply here I am guessing.=20

To address the barge-in problem we could just add some text to clarify
that a dtmf-recog resource generates the barge-in event only for DTMF
digits while the speech-recog resource generates the barge-in event for
both speech or DTMF.

Sarvi

     -----Original Message-----
     From: Dave Burke [mailto:david.burke@voxpilot.com]=20
     Sent: Tuesday, July 05, 2005 9:54 AM
     To: Shanmugham, Saravanan; David R Oran; Klaus Reifenrath
     Cc: speechsc@ietf.org
     Subject: Re: [Speechsc] START-OF-SPEECH in DTMF-only mode
    =20
     Inline.
    =20
     Dave
    =20
     ----- Original Message -----
     From: "Shanmugham, Saravanan" <sarvi@cisco.com>
     To: "David R Oran" <oran@cisco.com>; "Klaus Reifenrath"=20
     <Klaus.Reifenrath@Scansoft.com>
     Cc: <speechsc@ietf.org>
     Sent: Tuesday, July 05, 2005 5:22 PM
     Subject: RE: [Speechsc] START-OF-SPEECH in DTMF-only mode
    =20
    =20
     I agree with Dave's analysis. The purpose of this event=20
     was barge-in.
     And barge-in should happen for both DTMF and speech.
    =20
     Is there a case where you think it should not behave this=20
     way. If soe,
     please provide a scenario where you think
         1. Barge-in should happen for DTMF and not voice or vice-versa.
    =20
     DB> You want to do a DTMF recognition only because it is=20
     noisy. While=20
     waiting for DTMF input, the speechrecog resource (or=20
     advanced dtmfrecog)=20
     generates a START-OF-SPEECH because it heard some speech.=20
     The client does=20
     not want to stop prompt playing unless DTMF was heard but=20
     it can't tell by=20
     the START-OF-SPEECH whether speech or DTMF was heard.=20
     Similarly vice versa.
    =20
         2. You would benefit from the client knowing what caused the
     barge-in, DTMF Vs speech.
    =20
     DB> See previous comment. And previous e-mail: either add=20
     an inputmodes=20
     header (taking value speech, dtmf, both) to the RECOGNIZE=20
     request or add a=20
     header to the START-OF-SPEECH event indicating DTMF or speech.
    =20
     Sarvi
    =20
          -----Original Message-----
          From: speechsc-bounces@ietf.org
          [mailto:speechsc-bounces@ietf.org] On Behalf Of David R Oran
          Sent: Tuesday, July 05, 2005 5:24 AM
          To: Klaus Reifenrath
          Cc: 'speechsc@ietf.org'
          Subject: Re: [Speechsc] START-OF-SPEECH in DTMF-only mode
    =20
    =20
          On Jul 5, 2005, at 3:46 AM, Reifenrath, Klaus wrote:
    =20
          > The current spec is not clear when START-OF-SPEECH need
          to be send in
          > the following scenarios:
          > A) The client requested a DTMF Recognizer. Is the
          START-OF-SPEECH
          > event send to the client also if speech was detected?
          I suspect so, since one of the prime purposes is to enable
          client- mediated barge-in handling. However, if the
          recognizer is in fact only capable of recognizing DTMF
          then it may in fact not report anythin unless it's using
          some primitive thresholding machinery, like a SN threshold.
          > B) The client requested a Speech Recognizer, but only
          activated DTMF
          > grammars. Is the START-OF-SPEECH event send to the
          client also if
          > speech was detected?
          Again, I'd say yes, for the same reason as above.
          > I think in both cases START-OF-SPEECH should only=20
     be send after
          > detecting a DTMF digit (see Figure 12 of VoiceXML=20
     2.0: http://
          > www.w3.org/TR/voicexml20/#dmlATiming).
          We seem to have reached different conclusions. I'd be
          interested in why you think my analysis above is wrong.
    =20
          Dave.
    =20
          > Klaus
          >
          > _______________________________________________
          > Speechsc mailing list
          > Speechsc@ietf.org
          > https://www1.ietf.org/mailman/listinfo/speechsc
          >
    =20
          _______________________________________________
          Speechsc mailing list
          Speechsc@ietf.org
          https://www1.ietf.org/mailman/listinfo/speechsc
    =20
    =20
     _______________________________________________
     Speechsc mailing list
     Speechsc@ietf.org
     https://www1.ietf.org/mailman/listinfo/speechsc
    =20

_______________________________________________
Speechsc mailing list
Speechsc@ietf.org
https://www1.ietf.org/mailman/listinfo/speechsc



From speechsc-bounces@ietf.org Tue Jul 05 14:08:52 2005
Received: from localhost.localdomain ([127.0.0.1] helo=megatron.ietf.org)
	by megatron.ietf.org with esmtp (Exim 4.32)
	id 1Dprqm-0005q8-Pc; Tue, 05 Jul 2005 14:08:52 -0400
Received: from odin.ietf.org ([132.151.1.176] helo=ietf.org)
	by megatron.ietf.org with esmtp (Exim 4.32) id 1Dprqi-0005ow-KE
	for speechsc@megatron.ietf.org; Tue, 05 Jul 2005 14:08:51 -0400
Received: from ietf-mx.ietf.org (ietf-mx [132.151.6.1])
	by ietf.org (8.9.1a/8.9.1a) with ESMTP id OAA01499
	for <speechsc@ietf.org>; Tue, 5 Jul 2005 14:08:47 -0400 (EDT)
Received: from sj-iport-2-in.cisco.com ([171.71.176.71]
	helo=sj-iport-2.cisco.com) by ietf-mx.ietf.org with esmtp (Exim 4.33)
	id 1DpsHL-0008Mo-8u
	for speechsc@ietf.org; Tue, 05 Jul 2005 14:36:21 -0400
Received: from sj-core-2.cisco.com (171.71.177.254)
	by sj-iport-2.cisco.com with ESMTP; 05 Jul 2005 11:08:26 -0700
Received: from imail.cisco.com (imail.cisco.com [128.107.200.91])
	by sj-core-2.cisco.com (8.12.10/8.12.6) with ESMTP id j65I8Lod010674;
	Tue, 5 Jul 2005 11:08:21 -0700 (PDT)
Received: from [10.32.245.152] (stealth-10-32-245-152.cisco.com
	[10.32.245.152])
	by imail.cisco.com (8.12.11/8.12.10) with SMTP id j65I7XvD002261;
	Tue, 5 Jul 2005 11:07:33 -0700
In-Reply-To: <052f01c58182$2d78f300$cb00000a@db01.voxpilot.com>
References: <03772D1EC8DE624A863058C75874A75C112692@vtg-um-e2k6.sj21ad.cisco.com>
	<052f01c58182$2d78f300$cb00000a@db01.voxpilot.com>
Mime-Version: 1.0 (Apple Message framework v730)
X-Priority: 3
Content-Type: text/plain; charset=US-ASCII; delsp=yes; format=flowed
Message-Id: <DC5207D9-F31A-4D70-BE24-F694AE8B83A2@cisco.com>
Content-Transfer-Encoding: 7bit
From: David R Oran <oran@cisco.com>
Subject: Re: [Speechsc] START-OF-SPEECH in DTMF-only mode
Date: Tue, 5 Jul 2005 14:08:22 -0400
To: Dave Burke <david.burke@voxpilot.com>
X-Mailer: Apple Mail (2.730)
IIM-SIG: v:"1.1"; h:"imail.cisco.com"; d:"cisco.com"; z:"home"; m:"krs";
	t:"1120586854.625253"; x:"432200"; a:"rsa-sha1"; b:"nofws:4040";
	e:"Iw=="; n:"sQYarK2E51MdcTiUqeif3F7cWdxIfoCiXhdfb9vD5ee/j0jXL15gbFxF2p"
	"XIweAblu0N6XAgK7k+wrbr7bQDJaCDqOmzqpRUBjIRQAXQ7NzadpmR3pUL6wxaRUtW+c43sl9jC"
	"50Qg1sXHpPjt8Y+Y16ioyQAQAdSunM4YhevURc=";
	s:"AbFyUerOK3WlIVi9M59G81DqcQAzyS2gdPqEYeWd+Y5y3uUpIXbwrCAwClmHZ6pNnDY8pN7O"
	"Zi++oNDwMwcsEQ6rX/EJiSSJjCLjlRmJgI+eJblXiM8ZDXMPv+33pI3Wfvbl5f0nsahfIM2+Tq+"
	"OTAaQFYNGHz95lLfOANrVJnE="; c:"From: David R Oran <oran@cisco.com>";
	c:"Subject: Re: [Speechsc] START-OF-SPEECH in DTMF-only mode";
	c:"Date: Tue, 5 Jul 2005 14:08:22 -0400"
IIM-VERIFY: s:"y"; v:"y"; r:"60"; h:"imail.cisco.com";
	c:"message from imail.cisco.com verified; "
X-Spam-Score: 0.0 (/)
X-Scan-Signature: b2809b6f39decc6de467dcf252f42af1
Content-Transfer-Encoding: 7bit
Cc: speechsc@ietf.org, "Shanmugham, Saravanan" <sarvi@cisco.com>,
	Klaus Reifenrath <Klaus.Reifenrath@Scansoft.com>
X-BeenThere: speechsc@ietf.org
X-Mailman-Version: 2.1.5
Precedence: list
List-Id: Speech Services Control Working Group <speechsc.ietf.org>
List-Unsubscribe: <https://www1.ietf.org/mailman/listinfo/speechsc>,
	<mailto:speechsc-request@ietf.org?subject=unsubscribe>
List-Post: <mailto:speechsc@ietf.org>
List-Help: <mailto:speechsc-request@ietf.org?subject=help>
List-Subscribe: <https://www1.ietf.org/mailman/listinfo/speechsc>,
	<mailto:speechsc-request@ietf.org?subject=subscribe>
Sender: speechsc-bounces@ietf.org
Errors-To: speechsc-bounces@ietf.org


On Jul 5, 2005, at 12:54 PM, Dave Burke wrote:

> Inline.
>
> Dave
>
> ----- Original Message ----- From: "Shanmugham, Saravanan"  
> <sarvi@cisco.com>
> To: "David R Oran" <oran@cisco.com>; "Klaus Reifenrath"  
> <Klaus.Reifenrath@Scansoft.com>
> Cc: <speechsc@ietf.org>
> Sent: Tuesday, July 05, 2005 5:22 PM
> Subject: RE: [Speechsc] START-OF-SPEECH in DTMF-only mode
>
>
> I agree with Dave's analysis. The purpose of this event was barge-in.
> And barge-in should happen for both DTMF and speech.
>
> Is there a case where you think it should not behave this way. If soe,
> please provide a scenario where you think
>    1. Barge-in should happen for DTMF and not voice or vice-versa.
>
> DB> You want to do a DTMF recognition only because it is noisy.  
> While waiting for DTMF input, the speechrecog resource (or advanced  
> dtmfrecog) generates a START-OF-SPEECH because it heard some  
> speech. The client does not want to stop prompt playing unless DTMF  
> was heard but it can't tell by the START-OF-SPEECH whether speech  
> or DTMF was heard. Similarly vice versa.
>
It's an interesting design question what part of the policy resides  
at the client and what at the server, and who makes the "final  
decision" about whether what was heard was relevant to the control  
channel. Right now we (IMO) have a weird partitioning in many cases  
where the client basically says "do what I mean", but there are no  
constraints of what the server actually does, and no normalized basis  
for the client to figure out what to set various magic numbers to  
(e.g. sensitivity).

In this case the only thing the client needs to decide is whether to  
kill the prompt because the server thinks something that would  
interfere with the feedback ear/mouth/finger control happened. What  
this says to me is that it isn't necessarily a good idea for the  
client to have more knobs to control the server (especially if those  
knows are just more value/policy input ungrounded in any physics/ 
acoustics). On the other hand, having the server tell the client more  
about what it thinks is going on is probably valuable.

So, Coming to the point after this long rambling introduction, I  
think it would in fact be useful for the START-Of-SPEECH event to  
indicate some extra information, for example:
a) I got something enough above the noise floor to qualify for  
exceeding the "Sensisitvity" parameter you sent in on the request but  
I really can't tell what it is (could be a hippopatmus fart, or a  
siren in the background, or captain crunch trying to whistle DTMF).
b) I think I'm hearing speech
c) I think I'm hearing DTMF



>    2. You would benefit from the client knowing what caused the
> barge-in, DTMF Vs speech.
>
> DB> See previous comment. And previous e-mail: either add an  
> inputmodes header (taking value speech, dtmf, both) to the  
> RECOGNIZE request or add a header to the START-OF-SPEECH event  
> indicating DTMF or speech.
>
I'm leaning in your direction on this latter point - as should be  
evident from what I wrote above.

> Sarvi
>
>     -----Original Message-----
>     From: speechsc-bounces@ietf.org
>     [mailto:speechsc-bounces@ietf.org] On Behalf Of David R Oran
>     Sent: Tuesday, July 05, 2005 5:24 AM
>     To: Klaus Reifenrath
>     Cc: 'speechsc@ietf.org'
>     Subject: Re: [Speechsc] START-OF-SPEECH in DTMF-only mode
>
>
>     On Jul 5, 2005, at 3:46 AM, Reifenrath, Klaus wrote:
>
>     > The current spec is not clear when START-OF-SPEECH need
>     to be send in
>     > the following scenarios:
>     > A) The client requested a DTMF Recognizer. Is the
>     START-OF-SPEECH
>     > event send to the client also if speech was detected?
>     I suspect so, since one of the prime purposes is to enable
>     client- mediated barge-in handling. However, if the
>     recognizer is in fact only capable of recognizing DTMF
>     then it may in fact not report anythin unless it's using
>     some primitive thresholding machinery, like a SN threshold.
>     > B) The client requested a Speech Recognizer, but only
>     activated DTMF
>     > grammars. Is the START-OF-SPEECH event send to the
>     client also if
>     > speech was detected?
>     Again, I'd say yes, for the same reason as above.
>     > I think in both cases START-OF-SPEECH should only be send after
>     > detecting a DTMF digit (see Figure 12 of VoiceXML 2.0: http://
>     > www.w3.org/TR/voicexml20/#dmlATiming).
>     We seem to have reached different conclusions. I'd be
>     interested in why you think my analysis above is wrong.
>
>     Dave.
>
>     > Klaus
>     >
>     > _______________________________________________
>     > Speechsc mailing list
>     > Speechsc@ietf.org
>     > https://www1.ietf.org/mailman/listinfo/speechsc
>     >
>
>     _______________________________________________
>     Speechsc mailing list
>     Speechsc@ietf.org
>     https://www1.ietf.org/mailman/listinfo/speechsc
>
>
> _______________________________________________
> Speechsc mailing list
> Speechsc@ietf.org
> https://www1.ietf.org/mailman/listinfo/speechsc
>

_______________________________________________
Speechsc mailing list
Speechsc@ietf.org
https://www1.ietf.org/mailman/listinfo/speechsc



From speechsc-bounces@ietf.org Tue Jul 05 16:04:27 2005
Received: from localhost.localdomain ([127.0.0.1] helo=megatron.ietf.org)
	by megatron.ietf.org with esmtp (Exim 4.32)
	id 1Dptec-0004Mz-Se; Tue, 05 Jul 2005 16:04:26 -0400
Received: from odin.ietf.org ([132.151.1.176] helo=ietf.org)
	by megatron.ietf.org with esmtp (Exim 4.32) id 1Dptea-0004Ky-Cd
	for speechsc@megatron.ietf.org; Tue, 05 Jul 2005 16:04:25 -0400
Received: from ietf-mx.ietf.org (ietf-mx [132.151.6.1])
	by ietf.org (8.9.1a/8.9.1a) with ESMTP id QAA16366
	for <speechsc@ietf.org>; Tue, 5 Jul 2005 16:04:18 -0400 (EDT)
Received: from sj-iport-5.cisco.com ([171.68.10.87])
	by ietf-mx.ietf.org with esmtp (Exim 4.33) id 1Dptta-0006ky-VS
	for speechsc@ietf.org; Tue, 05 Jul 2005 16:19:55 -0400
Received: from sj-core-3.cisco.com (171.68.223.137)
	by sj-iport-5.cisco.com with ESMTP; 05 Jul 2005 12:52:01 -0700
X-IronPort-AV: i="3.93,262,1115017200"; 
	d="scan'208"; a="196455833:sNHT31745464"
Received: from vtg-um-e2k6.sj21ad.cisco.com (vtg-um-e2k6.cisco.com
	[171.70.93.77])
	by sj-core-3.cisco.com (8.12.10/8.12.6) with ESMTP id j65Jpw6p023568;
	Tue, 5 Jul 2005 12:51:58 -0700 (PDT)
Content-class: urn:content-classes:message
MIME-Version: 1.0
Content-Type: text/plain;
	charset="us-ascii"
Content-Transfer-Encoding: quoted-printable
X-MimeOLE: Produced By Microsoft Exchange V6.5.7226.0
Subject: RE: [Speechsc] START-OF-SPEECH in DTMF-only mode
Date: Tue, 5 Jul 2005 12:51:57 -0700
Message-ID: <03772D1EC8DE624A863058C75874A75C1126EE@vtg-um-e2k6.sj21ad.cisco.com>
Thread-Topic: [Speechsc] START-OF-SPEECH in DTMF-only mode
Thread-Index: AcWBjIp7vAuEXNbLT1iALkA9BtfHuwADPUKQ
From: "Shanmugham, Saravanan" <sarvi@cisco.com>
To: "David R Oran" <oran@cisco.com>, "Dave Burke" <david.burke@voxpilot.com>
X-Spam-Score: 0.0 (/)
X-Scan-Signature: 0e9ebc0cbd700a87c0637ad0e2c91610
Content-Transfer-Encoding: quoted-printable
Cc: speechsc@ietf.org, Klaus Reifenrath <Klaus.Reifenrath@Scansoft.com>
X-BeenThere: speechsc@ietf.org
X-Mailman-Version: 2.1.5
Precedence: list
List-Id: Speech Services Control Working Group <speechsc.ietf.org>
List-Unsubscribe: <https://www1.ietf.org/mailman/listinfo/speechsc>,
	<mailto:speechsc-request@ietf.org?subject=unsubscribe>
List-Post: <mailto:speechsc@ietf.org>
List-Help: <mailto:speechsc-request@ietf.org?subject=help>
List-Subscribe: <https://www1.ietf.org/mailman/listinfo/speechsc>,
	<mailto:speechsc-request@ietf.org?subject=subscribe>
Sender: speechsc-bounces@ietf.org
Errors-To: speechsc-bounces@ietf.org

 inline.

     -----Original Message-----
     From: David R Oran [mailto:oran@cisco.com]=20
     Sent: Tuesday, July 05, 2005 11:08 AM
     To: Dave Burke
     Cc: Shanmugham, Saravanan; Klaus Reifenrath; speechsc@ietf.org
     Subject: Re: [Speechsc] START-OF-SPEECH in DTMF-only mode
    =20
    =20
     On Jul 5, 2005, at 12:54 PM, Dave Burke wrote:
    =20
     > Inline.
     >
     > Dave
     >
     > ----- Original Message ----- From: "Shanmugham, Saravanan" =20
     > <sarvi@cisco.com>
     > To: "David R Oran" <oran@cisco.com>; "Klaus Reifenrath" =20
     > <Klaus.Reifenrath@Scansoft.com>
     > Cc: <speechsc@ietf.org>
     > Sent: Tuesday, July 05, 2005 5:22 PM
     > Subject: RE: [Speechsc] START-OF-SPEECH in DTMF-only mode
     >
     >
     > I agree with Dave's analysis. The purpose of this event=20
     was barge-in.
     > And barge-in should happen for both DTMF and speech.
     >
     > Is there a case where you think it should not behave=20
     this way. If soe,=20
     > please provide a scenario where you think
     >    1. Barge-in should happen for DTMF and not voice or=20
     vice-versa.
     >
     > DB> You want to do a DTMF recognition only because it is noisy. =20
     > While waiting for DTMF input, the speechrecog resource=20
     (or advanced
     > dtmfrecog) generates a START-OF-SPEECH because it heard=20
     some speech.=20
     > The client does not want to stop prompt playing unless=20
     DTMF was heard=20
     > but it can't tell by the START-OF-SPEECH whether speech=20
     or DTMF was=20
     > heard. Similarly vice versa.
     >
     It's an interesting design question what part of the=20
     policy resides at the client and what at the server, and=20
     who makes the "final decision" about whether what was=20
     heard was relevant to the control channel. Right now we=20
     (IMO) have a weird partitioning in many cases where the=20
     client basically says "do what I mean", but there are no=20
     constraints of what the server actually does, and no=20
     normalized basis for the client to figure out what to set=20
     various magic numbers to (e.g. sensitivity).
    =20
     In this case the only thing the client needs to decide is=20
     whether to kill the prompt because the server thinks=20
     something that would interfere with the feedback=20
     ear/mouth/finger control happened. What this says to me is=20
     that it isn't necessarily a good idea for the client to=20
     have more knobs to control the server (especially if those=20
     knows are just more value/policy input ungrounded in any=20
     physics/ acoustics). On the other hand, having the server=20
     tell the client more about what it thinks is going on is=20
     probably valuable.
    =20
     So, Coming to the point after this long rambling=20
     introduction, I think it would in fact be useful for the=20
     START-Of-SPEECH event to indicate some extra information,=20
     for example:
     a) I got something enough above the noise floor to qualify=20
     for exceeding the "Sensisitvity" parameter you sent in on=20
     the request but I really can't tell what it is (could be a=20
     hippopatmus fart, or a siren in the background, or captain=20
     crunch trying to whistle DTMF).
     b) I think I'm hearing speech
     c) I think I'm hearing DTMF

Though I agree with your former part of your response. I am not sure I
agree with your proposed solution.=20
The way I see this problem is that, it is more of what constitues a
barge-in event. This boils down to whether it is speech, DTMF or both.
This is inturn boils down to what type of recognizer resource we are
using, dtmf-recog, speech-recog, and speech-only-recog(we don't have
this and I don't think we should add it, but think of this as a place
holder that explains the concept).

A client knowing what type of barge-in happenned, does not impact the
barge-in operation itself as it may be too late(for the optimized
barge-in case). It may have other use cases, and if we can identify
them, I don't mind adding support for the START-OF-SPEECH event to say
what type of barge-in happenned. But that itself does not solve the
original problem raised. Refer to my previous response.

The solution lies in defining what what is a barge-in event. That boils
down to what type of recognition is happenning, dtmf-only, speech-dtmf
or speech-only. We do not support speech-only as a resource today, the
question is do we need a header to force it.=20

Sarvi    =20
    =20
    =20
     >    2. You would benefit from the client knowing what caused the=20
     > barge-in, DTMF Vs speech.
     >
     > DB> See previous comment. And previous e-mail: either add an
     > inputmodes header (taking value speech, dtmf, both) to=20
     the RECOGNIZE=20
     > request or add a header to the START-OF-SPEECH event=20
     indicating DTMF=20
     > or speech.
     >
     I'm leaning in your direction on this latter point - as=20
     should be evident from what I wrote above.
    =20
     > Sarvi
     >
     >     -----Original Message-----
     >     From: speechsc-bounces@ietf.org
     >     [mailto:speechsc-bounces@ietf.org] On Behalf Of David R Oran
     >     Sent: Tuesday, July 05, 2005 5:24 AM
     >     To: Klaus Reifenrath
     >     Cc: 'speechsc@ietf.org'
     >     Subject: Re: [Speechsc] START-OF-SPEECH in DTMF-only mode
     >
     >
     >     On Jul 5, 2005, at 3:46 AM, Reifenrath, Klaus wrote:
     >
     >     > The current spec is not clear when START-OF-SPEECH need
     >     to be send in
     >     > the following scenarios:
     >     > A) The client requested a DTMF Recognizer. Is the
     >     START-OF-SPEECH
     >     > event send to the client also if speech was detected?
     >     I suspect so, since one of the prime purposes is to enable
     >     client- mediated barge-in handling. However, if the
     >     recognizer is in fact only capable of recognizing DTMF
     >     then it may in fact not report anythin unless it's using
     >     some primitive thresholding machinery, like a SN threshold.
     >     > B) The client requested a Speech Recognizer, but only
     >     activated DTMF
     >     > grammars. Is the START-OF-SPEECH event send to the
     >     client also if
     >     > speech was detected?
     >     Again, I'd say yes, for the same reason as above.
     >     > I think in both cases START-OF-SPEECH should only=20
     be send after
     >     > detecting a DTMF digit (see Figure 12 of VoiceXML=20
     2.0: http://
     >     > www.w3.org/TR/voicexml20/#dmlATiming).
     >     We seem to have reached different conclusions. I'd be
     >     interested in why you think my analysis above is wrong.
     >
     >     Dave.
     >
     >     > Klaus
     >     >
     >     > _______________________________________________
     >     > Speechsc mailing list
     >     > Speechsc@ietf.org
     >     > https://www1.ietf.org/mailman/listinfo/speechsc
     >     >
     >
     >     _______________________________________________
     >     Speechsc mailing list
     >     Speechsc@ietf.org
     >     https://www1.ietf.org/mailman/listinfo/speechsc
     >
     >
     > _______________________________________________
     > Speechsc mailing list
     > Speechsc@ietf.org
     > https://www1.ietf.org/mailman/listinfo/speechsc
     >
    =20

_______________________________________________
Speechsc mailing list
Speechsc@ietf.org
https://www1.ietf.org/mailman/listinfo/speechsc



From speechsc-bounces@ietf.org Tue Jul 05 16:27:50 2005
Received: from localhost.localdomain ([127.0.0.1] helo=megatron.ietf.org)
	by megatron.ietf.org with esmtp (Exim 4.32)
	id 1Dpu1G-0005h8-2B; Tue, 05 Jul 2005 16:27:50 -0400
Received: from odin.ietf.org ([132.151.1.176] helo=ietf.org)
	by megatron.ietf.org with esmtp (Exim 4.32) id 1Dptzr-0003fI-TP
	for speechsc@megatron.ietf.org; Tue, 05 Jul 2005 16:26:24 -0400
Received: from ietf-mx.ietf.org (ietf-mx [132.151.6.1])
	by ietf.org (8.9.1a/8.9.1a) with ESMTP id QAA19100
	for <speechsc@ietf.org>; Tue, 5 Jul 2005 16:21:39 -0400 (EDT)
Received: from letter.nuance.com ([207.107.210.132])
	by ietf-mx.ietf.org with esmtp (Exim 4.33) id 1DpuDU-0002zx-O4
	for speechsc@ietf.org; Tue, 05 Jul 2005 16:40:30 -0400
Received: from postcard.nuance.com ([10.3.6.20]:20768)
	by letter.nuance.com with esmtp id 1DptmT-0003Xm-PO;
	Tue, 05 Jul 2005 13:12:33 -0700
Received: from mtb1exch01.nuance.com ([10.3.2.6]) by postcard.nuance.com with
	Microsoft SMTPSVC(6.0.3790.0); Tue, 5 Jul 2005 16:12:27 -0400
X-MimeOLE: Produced By Microsoft Exchange V6.5.7226.0
Content-class: urn:content-classes:message
MIME-Version: 1.0
Content-Type: text/plain;
	charset="us-ascii"
Content-Transfer-Encoding: quoted-printable
Subject: RE: [Speechsc] START-OF-SPEECH in DTMF-only mode
Date: Tue, 5 Jul 2005 16:12:27 -0400
Message-ID: <7DE7C4EF3B7C8B4B82955191378290D802ED3D53@mtb1exch01.nuance.com>
Thread-Topic: [Speechsc] START-OF-SPEECH in DTMF-only mode
Thread-Index: AcWBjIp7vAuEXNbLT1iALkA9BtfHuwADPUKQAAEATEA=
From: "Pierre Forgues" <forgues@nuance.com>
To: "Shanmugham, Saravanan" <sarvi@cisco.com>, "David R Oran" <oran@cisco.com>,
	"Dave Burke" <david.burke@voxpilot.com>
X-OriginalArrivalTime: 05 Jul 2005 20:12:27.0852 (UTC)
	FILETIME=[DDA788C0:01C5819D]
X-FromHost: postcard.nuance.com [10.3.6.20]:20768
Lines: 201
X-Spam-Score: 0.0 (/)
X-Scan-Signature: 311e798ce51dbeacf5cdfcc8e9fda21b
Content-Transfer-Encoding: quoted-printable
Cc: speechsc@ietf.org, Klaus Reifenrath <Klaus.Reifenrath@Scansoft.com>
X-BeenThere: speechsc@ietf.org
X-Mailman-Version: 2.1.5
Precedence: list
List-Id: Speech Services Control Working Group <speechsc.ietf.org>
List-Unsubscribe: <https://www1.ietf.org/mailman/listinfo/speechsc>,
	<mailto:speechsc-request@ietf.org?subject=unsubscribe>
List-Post: <mailto:speechsc@ietf.org>
List-Help: <mailto:speechsc-request@ietf.org?subject=help>
List-Subscribe: <https://www1.ietf.org/mailman/listinfo/speechsc>,
	<mailto:speechsc-request@ietf.org?subject=subscribe>
Sender: speechsc-bounces@ietf.org
Errors-To: speechsc-bounces@ietf.org

The client will be told in the recognition results exactly what
triggered the START-OF-SPEECH event (dtmf or speech).  At the time the
SOS event is reported, the knowledge of what triggered the event is not
needed to make a decision to stop a prompt (nor does the server even
know at that time what triggered it).

Pierre

-----Original Message-----
From: speechsc-bounces@ietf.org [mailto:speechsc-bounces@ietf.org] On
Behalf Of Shanmugham, Saravanan
Sent: Tuesday, July 05, 2005 3:52 PM
To: David R Oran; Dave Burke
Cc: speechsc@ietf.org; Klaus Reifenrath
Subject: RE: [Speechsc] START-OF-SPEECH in DTMF-only mode

 inline.

     -----Original Message-----
     From: David R Oran [mailto:oran@cisco.com]=20
     Sent: Tuesday, July 05, 2005 11:08 AM
     To: Dave Burke
     Cc: Shanmugham, Saravanan; Klaus Reifenrath; speechsc@ietf.org
     Subject: Re: [Speechsc] START-OF-SPEECH in DTMF-only mode
    =20
    =20
     On Jul 5, 2005, at 12:54 PM, Dave Burke wrote:
    =20
     > Inline.
     >
     > Dave
     >
     > ----- Original Message ----- From: "Shanmugham, Saravanan" =20
     > <sarvi@cisco.com>
     > To: "David R Oran" <oran@cisco.com>; "Klaus Reifenrath" =20
     > <Klaus.Reifenrath@Scansoft.com>
     > Cc: <speechsc@ietf.org>
     > Sent: Tuesday, July 05, 2005 5:22 PM
     > Subject: RE: [Speechsc] START-OF-SPEECH in DTMF-only mode
     >
     >
     > I agree with Dave's analysis. The purpose of this event=20
     was barge-in.
     > And barge-in should happen for both DTMF and speech.
     >
     > Is there a case where you think it should not behave=20
     this way. If soe,=20
     > please provide a scenario where you think
     >    1. Barge-in should happen for DTMF and not voice or=20
     vice-versa.
     >
     > DB> You want to do a DTMF recognition only because it is noisy. =20
     > While waiting for DTMF input, the speechrecog resource=20
     (or advanced
     > dtmfrecog) generates a START-OF-SPEECH because it heard=20
     some speech.=20
     > The client does not want to stop prompt playing unless=20
     DTMF was heard=20
     > but it can't tell by the START-OF-SPEECH whether speech=20
     or DTMF was=20
     > heard. Similarly vice versa.
     >
     It's an interesting design question what part of the=20
     policy resides at the client and what at the server, and=20
     who makes the "final decision" about whether what was=20
     heard was relevant to the control channel. Right now we=20
     (IMO) have a weird partitioning in many cases where the=20
     client basically says "do what I mean", but there are no=20
     constraints of what the server actually does, and no=20
     normalized basis for the client to figure out what to set=20
     various magic numbers to (e.g. sensitivity).
    =20
     In this case the only thing the client needs to decide is=20
     whether to kill the prompt because the server thinks=20
     something that would interfere with the feedback=20
     ear/mouth/finger control happened. What this says to me is=20
     that it isn't necessarily a good idea for the client to=20
     have more knobs to control the server (especially if those=20
     knows are just more value/policy input ungrounded in any=20
     physics/ acoustics). On the other hand, having the server=20
     tell the client more about what it thinks is going on is=20
     probably valuable.
    =20
     So, Coming to the point after this long rambling=20
     introduction, I think it would in fact be useful for the=20
     START-Of-SPEECH event to indicate some extra information,=20
     for example:
     a) I got something enough above the noise floor to qualify=20
     for exceeding the "Sensisitvity" parameter you sent in on=20
     the request but I really can't tell what it is (could be a=20
     hippopatmus fart, or a siren in the background, or captain=20
     crunch trying to whistle DTMF).
     b) I think I'm hearing speech
     c) I think I'm hearing DTMF

Though I agree with your former part of your response. I am not sure I
agree with your proposed solution.=20
The way I see this problem is that, it is more of what constitues a
barge-in event. This boils down to whether it is speech, DTMF or both.
This is inturn boils down to what type of recognizer resource we are
using, dtmf-recog, speech-recog, and speech-only-recog(we don't have
this and I don't think we should add it, but think of this as a place
holder that explains the concept).

A client knowing what type of barge-in happenned, does not impact the
barge-in operation itself as it may be too late(for the optimized
barge-in case). It may have other use cases, and if we can identify
them, I don't mind adding support for the START-OF-SPEECH event to say
what type of barge-in happenned. But that itself does not solve the
original problem raised. Refer to my previous response.

The solution lies in defining what what is a barge-in event. That boils
down to what type of recognition is happenning, dtmf-only, speech-dtmf
or speech-only. We do not support speech-only as a resource today, the
question is do we need a header to force it.=20

Sarvi    =20
    =20
    =20
     >    2. You would benefit from the client knowing what caused the=20
     > barge-in, DTMF Vs speech.
     >
     > DB> See previous comment. And previous e-mail: either add an
     > inputmodes header (taking value speech, dtmf, both) to=20
     the RECOGNIZE=20
     > request or add a header to the START-OF-SPEECH event=20
     indicating DTMF=20
     > or speech.
     >
     I'm leaning in your direction on this latter point - as=20
     should be evident from what I wrote above.
    =20
     > Sarvi
     >
     >     -----Original Message-----
     >     From: speechsc-bounces@ietf.org
     >     [mailto:speechsc-bounces@ietf.org] On Behalf Of David R Oran
     >     Sent: Tuesday, July 05, 2005 5:24 AM
     >     To: Klaus Reifenrath
     >     Cc: 'speechsc@ietf.org'
     >     Subject: Re: [Speechsc] START-OF-SPEECH in DTMF-only mode
     >
     >
     >     On Jul 5, 2005, at 3:46 AM, Reifenrath, Klaus wrote:
     >
     >     > The current spec is not clear when START-OF-SPEECH need
     >     to be send in
     >     > the following scenarios:
     >     > A) The client requested a DTMF Recognizer. Is the
     >     START-OF-SPEECH
     >     > event send to the client also if speech was detected?
     >     I suspect so, since one of the prime purposes is to enable
     >     client- mediated barge-in handling. However, if the
     >     recognizer is in fact only capable of recognizing DTMF
     >     then it may in fact not report anythin unless it's using
     >     some primitive thresholding machinery, like a SN threshold.
     >     > B) The client requested a Speech Recognizer, but only
     >     activated DTMF
     >     > grammars. Is the START-OF-SPEECH event send to the
     >     client also if
     >     > speech was detected?
     >     Again, I'd say yes, for the same reason as above.
     >     > I think in both cases START-OF-SPEECH should only=20
     be send after
     >     > detecting a DTMF digit (see Figure 12 of VoiceXML=20
     2.0: http://
     >     > www.w3.org/TR/voicexml20/#dmlATiming).
     >     We seem to have reached different conclusions. I'd be
     >     interested in why you think my analysis above is wrong.
     >
     >     Dave.
     >
     >     > Klaus
     >     >
     >     > _______________________________________________
     >     > Speechsc mailing list
     >     > Speechsc@ietf.org
     >     > https://www1.ietf.org/mailman/listinfo/speechsc
     >     >
     >
     >     _______________________________________________
     >     Speechsc mailing list
     >     Speechsc@ietf.org
     >     https://www1.ietf.org/mailman/listinfo/speechsc
     >
     >
     > _______________________________________________
     > Speechsc mailing list
     > Speechsc@ietf.org
     > https://www1.ietf.org/mailman/listinfo/speechsc
     >
    =20

_______________________________________________
Speechsc mailing list
Speechsc@ietf.org
https://www1.ietf.org/mailman/listinfo/speechsc

=20
 =20


_______________________________________________
Speechsc mailing list
Speechsc@ietf.org
https://www1.ietf.org/mailman/listinfo/speechsc



From speechsc-bounces@ietf.org Wed Jul 06 05:32:17 2005
Received: from localhost.localdomain ([127.0.0.1] helo=megatron.ietf.org)
	by megatron.ietf.org with esmtp (Exim 4.32)
	id 1Dq6GP-0005OW-EI; Wed, 06 Jul 2005 05:32:17 -0400
Received: from odin.ietf.org ([132.151.1.176] helo=ietf.org)
	by megatron.ietf.org with esmtp (Exim 4.32) id 1Dq6GN-0005Lg-2q
	for speechsc@megatron.ietf.org; Wed, 06 Jul 2005 05:32:15 -0400
Received: from ietf-mx.ietf.org (ietf-mx [132.151.6.1])
	by ietf.org (8.9.1a/8.9.1a) with ESMTP id FAA04570
	for <speechsc@ietf.org>; Wed, 6 Jul 2005 05:32:12 -0400 (EDT)
Received: from fw01.db01.voxpilot.com ([212.17.54.82] helo=mail.voxpilot.com)
	by ietf-mx.ietf.org with esmtp (Exim 4.33) id 1Dq6hJ-00054t-CF
	for speechsc@ietf.org; Wed, 06 Jul 2005 06:00:06 -0400
Received: from daburkewxp (unknown [10.0.0.203])
	by mail.voxpilot.com (Postfix) with ESMTP
	id 6E784214041; Wed,  6 Jul 2005 09:31:51 +0000 (GMT)
Message-ID: <073a01c5820d$8a2c4530$cb00000a@db01.voxpilot.com>
From: "Dave Burke" <david.burke@voxpilot.com>
To: "Shanmugham, Saravanan" <sarvi@cisco.com>, "David R Oran" <oran@cisco.com>
References: <03772D1EC8DE624A863058C75874A75C1126EE@vtg-um-e2k6.sj21ad.cisco.com>
Subject: Re: [Speechsc] START-OF-SPEECH in DTMF-only mode -> summary(?)
Date: Wed, 6 Jul 2005 10:31:51 +0100
MIME-Version: 1.0
Content-Type: text/plain; format=flowed; charset="iso-8859-1";
	reply-type=original
Content-Transfer-Encoding: 7bit
X-Priority: 3
X-MSMail-Priority: Normal
X-Mailer: Microsoft Outlook Express 6.00.2900.2180
X-MimeOLE: Produced By Microsoft MimeOLE V6.00.2900.2180
X-Spam-Score: 0.0 (/)
X-Scan-Signature: 223e3c753032a50d5dc4443c921c3fcd
Content-Transfer-Encoding: 7bit
Cc: speechsc@ietf.org, Klaus Reifenrath <Klaus.Reifenrath@Scansoft.com>
X-BeenThere: speechsc@ietf.org
X-Mailman-Version: 2.1.5
Precedence: list
List-Id: Speech Services Control Working Group <speechsc.ietf.org>
List-Unsubscribe: <https://www1.ietf.org/mailman/listinfo/speechsc>,
	<mailto:speechsc-request@ietf.org?subject=unsubscribe>
List-Post: <mailto:speechsc@ietf.org>
List-Help: <mailto:speechsc-request@ietf.org?subject=help>
List-Subscribe: <https://www1.ietf.org/mailman/listinfo/speechsc>,
	<mailto:speechsc-request@ietf.org?subject=subscribe>
Sender: speechsc-bounces@ietf.org
Errors-To: speechsc-bounces@ietf.org

+ Attempting to summarise:

1. START-OF-SPEECH is useful for the client to know when to stop playing 
prompts the in non-optimised case
2. START-OF-SPEECH is useful for the client to calculate the bargin time 
(e.g. VoiceXML 2.1 <mark>)
3. In the optimised case, a bargin automatically stops prompt playing 
(assuming prompts barginable)
4. Because of the previous point, the question of what input type caused 
bargin is different and less important to what input type(s) the recogniser 
is listening for
5. Currently, MRCPv2 has no way of indicating what input type(s) a 
recogniser is listening for

+ Why implement 5?

i. Noisy case: Need DTMF-only recognition (and may only have a speechrecog)
ii. Flexibility: Want speech-only recognition (because a second recogniser 
is doing hotword on DTMF)

+ How to implement 5?

a. Implicitly:
    - dtmfrecog: always DTMF-only recognition
    - speechrecog: depends on active grammar type
        > if a dtmf grammar is active then DTMF input is "on"
        > if a speech grammar is active then speech input is "on"

b. Explicitly:
    - Add inputmodes header to RECOGNIZE

Option a is David's "do what I mean case"; option b is the extra dial for 
the client.

It is worth noting that VoiceXML uses option b. This allows one to activate 
both speech grammars and DTMF grammars (and therefore be informed of any 
errors in the grammars at activation time) but independently turn on 
whichever input mode you like e.g. perhaps start with "both" then change to 
"dtmf".

Dave












Sarvi makes a good point that adding the reason why the START-OF-SPEECH 
occurred does not fix the optimised bargin case.

dtmfrecog - listens for DTMF only
speechrecog - listens for DTMF only, or speech only, or speech & DTMF








----- Original Message ----- 
From: "Shanmugham, Saravanan" <sarvi@cisco.com>
To: "David R Oran" <oran@cisco.com>; "Dave Burke" <david.burke@voxpilot.com>
Cc: <speechsc@ietf.org>; "Klaus Reifenrath" <Klaus.Reifenrath@Scansoft.com>
Sent: Tuesday, July 05, 2005 8:51 PM
Subject: RE: [Speechsc] START-OF-SPEECH in DTMF-only mode


inline.

     -----Original Message-----
     From: David R Oran [mailto:oran@cisco.com]
     Sent: Tuesday, July 05, 2005 11:08 AM
     To: Dave Burke
     Cc: Shanmugham, Saravanan; Klaus Reifenrath; speechsc@ietf.org
     Subject: Re: [Speechsc] START-OF-SPEECH in DTMF-only mode


     On Jul 5, 2005, at 12:54 PM, Dave Burke wrote:

     > Inline.
     >
     > Dave
     >
     > ----- Original Message ----- From: "Shanmugham, Saravanan"
     > <sarvi@cisco.com>
     > To: "David R Oran" <oran@cisco.com>; "Klaus Reifenrath"
     > <Klaus.Reifenrath@Scansoft.com>
     > Cc: <speechsc@ietf.org>
     > Sent: Tuesday, July 05, 2005 5:22 PM
     > Subject: RE: [Speechsc] START-OF-SPEECH in DTMF-only mode
     >
     >
     > I agree with Dave's analysis. The purpose of this event
     was barge-in.
     > And barge-in should happen for both DTMF and speech.
     >
     > Is there a case where you think it should not behave
     this way. If soe,
     > please provide a scenario where you think
     >    1. Barge-in should happen for DTMF and not voice or
     vice-versa.
     >
     > DB> You want to do a DTMF recognition only because it is noisy.
     > While waiting for DTMF input, the speechrecog resource
     (or advanced
     > dtmfrecog) generates a START-OF-SPEECH because it heard
     some speech.
     > The client does not want to stop prompt playing unless
     DTMF was heard
     > but it can't tell by the START-OF-SPEECH whether speech
     or DTMF was
     > heard. Similarly vice versa.
     >
     It's an interesting design question what part of the
     policy resides at the client and what at the server, and
     who makes the "final decision" about whether what was
     heard was relevant to the control channel. Right now we
     (IMO) have a weird partitioning in many cases where the
     client basically says "do what I mean", but there are no
     constraints of what the server actually does, and no
     normalized basis for the client to figure out what to set
     various magic numbers to (e.g. sensitivity).

     In this case the only thing the client needs to decide is
     whether to kill the prompt because the server thinks
     something that would interfere with the feedback
     ear/mouth/finger control happened. What this says to me is
     that it isn't necessarily a good idea for the client to
     have more knobs to control the server (especially if those
     knows are just more value/policy input ungrounded in any
     physics/ acoustics). On the other hand, having the server
     tell the client more about what it thinks is going on is
     probably valuable.

     So, Coming to the point after this long rambling
     introduction, I think it would in fact be useful for the
     START-Of-SPEECH event to indicate some extra information,
     for example:
     a) I got something enough above the noise floor to qualify
     for exceeding the "Sensisitvity" parameter you sent in on
     the request but I really can't tell what it is (could be a
     hippopatmus fart, or a siren in the background, or captain
     crunch trying to whistle DTMF).
     b) I think I'm hearing speech
     c) I think I'm hearing DTMF

Though I agree with your former part of your response. I am not sure I
agree with your proposed solution.
The way I see this problem is that, it is more of what constitues a
barge-in event. This boils down to whether it is speech, DTMF or both.
This is inturn boils down to what type of recognizer resource we are
using, dtmf-recog, speech-recog, and speech-only-recog(we don't have
this and I don't think we should add it, but think of this as a place
holder that explains the concept).

A client knowing what type of barge-in happenned, does not impact the
barge-in operation itself as it may be too late(for the optimized
barge-in case). It may have other use cases, and if we can identify
them, I don't mind adding support for the START-OF-SPEECH event to say
what type of barge-in happenned. But that itself does not solve the
original problem raised. Refer to my previous response.

The solution lies in defining what what is a barge-in event. That boils
down to what type of recognition is happenning, dtmf-only, speech-dtmf
or speech-only. We do not support speech-only as a resource today, the
question is do we need a header to force it.

Sarvi


     >    2. You would benefit from the client knowing what caused the
     > barge-in, DTMF Vs speech.
     >
     > DB> See previous comment. And previous e-mail: either add an
     > inputmodes header (taking value speech, dtmf, both) to
     the RECOGNIZE
     > request or add a header to the START-OF-SPEECH event
     indicating DTMF
     > or speech.
     >
     I'm leaning in your direction on this latter point - as
     should be evident from what I wrote above.

     > Sarvi
     >
     >     -----Original Message-----
     >     From: speechsc-bounces@ietf.org
     >     [mailto:speechsc-bounces@ietf.org] On Behalf Of David R Oran
     >     Sent: Tuesday, July 05, 2005 5:24 AM
     >     To: Klaus Reifenrath
     >     Cc: 'speechsc@ietf.org'
     >     Subject: Re: [Speechsc] START-OF-SPEECH in DTMF-only mode
     >
     >
     >     On Jul 5, 2005, at 3:46 AM, Reifenrath, Klaus wrote:
     >
     >     > The current spec is not clear when START-OF-SPEECH need
     >     to be send in
     >     > the following scenarios:
     >     > A) The client requested a DTMF Recognizer. Is the
     >     START-OF-SPEECH
     >     > event send to the client also if speech was detected?
     >     I suspect so, since one of the prime purposes is to enable
     >     client- mediated barge-in handling. However, if the
     >     recognizer is in fact only capable of recognizing DTMF
     >     then it may in fact not report anythin unless it's using
     >     some primitive thresholding machinery, like a SN threshold.
     >     > B) The client requested a Speech Recognizer, but only
     >     activated DTMF
     >     > grammars. Is the START-OF-SPEECH event send to the
     >     client also if
     >     > speech was detected?
     >     Again, I'd say yes, for the same reason as above.
     >     > I think in both cases START-OF-SPEECH should only
     be send after
     >     > detecting a DTMF digit (see Figure 12 of VoiceXML
     2.0: http://
     >     > www.w3.org/TR/voicexml20/#dmlATiming).
     >     We seem to have reached different conclusions. I'd be
     >     interested in why you think my analysis above is wrong.
     >
     >     Dave.
     >
     >     > Klaus
     >     >
     >     > _______________________________________________
     >     > Speechsc mailing list
     >     > Speechsc@ietf.org
     >     > https://www1.ietf.org/mailman/listinfo/speechsc
     >     >
     >
     >     _______________________________________________
     >     Speechsc mailing list
     >     Speechsc@ietf.org
     >     https://www1.ietf.org/mailman/listinfo/speechsc
     >
     >
     > _______________________________________________
     > Speechsc mailing list
     > Speechsc@ietf.org
     > https://www1.ietf.org/mailman/listinfo/speechsc
     >


_______________________________________________
Speechsc mailing list
Speechsc@ietf.org
https://www1.ietf.org/mailman/listinfo/speechsc


_______________________________________________
Speechsc mailing list
Speechsc@ietf.org
https://www1.ietf.org/mailman/listinfo/speechsc



From speechsc-bounces@ietf.org Wed Jul 06 08:27:07 2005
Received: from localhost.localdomain ([127.0.0.1] helo=megatron.ietf.org)
	by megatron.ietf.org with esmtp (Exim 4.32)
	id 1Dq8zb-0004W4-PZ; Wed, 06 Jul 2005 08:27:07 -0400
Received: from odin.ietf.org ([132.151.1.176] helo=ietf.org)
	by megatron.ietf.org with esmtp (Exim 4.32) id 1Dq8xa-0002Y2-8L
	for speechsc@megatron.ietf.org; Wed, 06 Jul 2005 08:25:02 -0400
Received: from ietf-mx.ietf.org (ietf-mx [132.151.6.1])
	by ietf.org (8.9.1a/8.9.1a) with ESMTP id IAA11206
	for <speechsc@ietf.org>; Wed, 6 Jul 2005 08:23:35 -0400 (EDT)
Received: from sj-iport-3-in.cisco.com ([171.71.176.72]
	helo=sj-iport-3.cisco.com) by ietf-mx.ietf.org with esmtp (Exim 4.33)
	id 1Dq9A3-0001QT-H2
	for speechsc@ietf.org; Wed, 06 Jul 2005 08:37:57 -0400
Received: from sj-core-2.cisco.com (171.71.177.254)
	by sj-iport-3.cisco.com with ESMTP; 06 Jul 2005 05:09:53 -0700
X-IronPort-AV: i="3.93,264,1115017200"; 
	d="scan'208"; a="291190975:sNHT107149376"
Received: from imail.cisco.com (imail.cisco.com [128.107.200.91])
	by sj-core-2.cisco.com (8.12.10/8.12.6) with ESMTP id j66C9lod018559;
	Wed, 6 Jul 2005 05:09:47 -0700 (PDT)
Received: from [10.32.245.152] (stealth-10-32-245-152.cisco.com
	[10.32.245.152])
	by imail.cisco.com (8.12.11/8.12.10) with SMTP id j66C78hi007631;
	Wed, 6 Jul 2005 05:07:45 -0700
In-Reply-To: <073a01c5820d$8a2c4530$cb00000a@db01.voxpilot.com>
References: <03772D1EC8DE624A863058C75874A75C1126EE@vtg-um-e2k6.sj21ad.cisco.com>
	<073a01c5820d$8a2c4530$cb00000a@db01.voxpilot.com>
Mime-Version: 1.0 (Apple Message framework v730)
X-Priority: 3
Content-Type: text/plain; charset=US-ASCII; delsp=yes; format=flowed
Message-Id: <1CC816CD-E142-421F-BB6C-BBE8572A41FD@cisco.com>
Content-Transfer-Encoding: 7bit
From: David R Oran <oran@cisco.com>
Subject: Re: [Speechsc] START-OF-SPEECH in DTMF-only mode -> summary(?)
Date: Wed, 6 Jul 2005 08:07:37 -0400
To: "Dave Burke" <david.burke@voxpilot.com>
X-Mailer: Apple Mail (2.730)
IIM-SIG: v:"1.1"; h:"imail.cisco.com"; d:"cisco.com"; z:"home"; m:"krs";
	t:"1120651738.615169"; x:"432200"; a:"rsa-sha1"; b:"nofws:11557";
	e:"Iw=="; n:"sQYarK2E51MdcTiUqeif3F7cWdxIfoCiXhdfb9vD5ee/j0jXL15gbFxF2p"
	"XIweAblu0N6XAgK7k+wrbr7bQDJaCDqOmzqpRUBjIRQAXQ7NzadpmR3pUL6wxaRUtW+c43sl9jC"
	"50Qg1sXHpPjt8Y+Y16ioyQAQAdSunM4YhevURc=";
	s:"BIygREnmXUZTlAsdioB4L61G7FyL2Q+F2FS74SxRWGqKo9e0KFUQE6HH224m7I5crsVyL7JR"
	"ZldzF3PxpQKPCMimv6bDlrt66tb6/AiNbTn7BC9jyP9O0Z6uzzcWus9xvzzbecBCnBGIx/uFac6"
	"xhseQjTOmj6yVDPCdkiULrk0="; c:"From: David R Oran <oran@cisco.com>";
	c:"Subject: Re: [Speechsc] START-OF-SPEECH in DTMF-only mode ->
	summary" "(?)"; c:"Date: Wed, 6 Jul 2005 08:07:37 -0400"
IIM-VERIFY: s:"y"; v:"y"; r:"60"; h:"imail.cisco.com";
	c:"message from imail.cisco.com verified; "
X-Spam-Score: 0.0 (/)
X-Scan-Signature: e95407604bef3289cd27cb4f3b3a35b4
Content-Transfer-Encoding: 7bit
Cc: speechsc@ietf.org, "Shanmugham, Saravanan" <sarvi@cisco.com>,
	Klaus Reifenrath <Klaus.Reifenrath@Scansoft.com>
X-BeenThere: speechsc@ietf.org
X-Mailman-Version: 2.1.5
Precedence: list
List-Id: Speech Services Control Working Group <speechsc.ietf.org>
List-Unsubscribe: <https://www1.ietf.org/mailman/listinfo/speechsc>,
	<mailto:speechsc-request@ietf.org?subject=unsubscribe>
List-Post: <mailto:speechsc@ietf.org>
List-Help: <mailto:speechsc-request@ietf.org?subject=help>
List-Subscribe: <https://www1.ietf.org/mailman/listinfo/speechsc>,
	<mailto:speechsc-request@ietf.org?subject=subscribe>
Sender: speechsc-bounces@ietf.org
Errors-To: speechsc-bounces@ietf.org


On Jul 6, 2005, at 5:31 AM, Dave Burke wrote:

> + Attempting to summarise:
>
> 1. START-OF-SPEECH is useful for the client to know when to stop  
> playing prompts the in non-optimised case
> 2. START-OF-SPEECH is useful for the client to calculate the bargin  
> time (e.g. VoiceXML 2.1 <mark>)
> 3. In the optimised case, a bargin automatically stops prompt  
> playing (assuming prompts barginable)
> 4. Because of the previous point, the question of what input type  
> caused bargin is different and less important to what input type(s)  
> the recogniser is listening for
I'm not sure I follow point 4. Could you elaborate?

> 5. Currently, MRCPv2 has no way of indicating what input type(s) a  
> recogniser is listening for
>
Do you mean exactly this, or do you mean "for the client to indicate  
to the resource what input types it should look for"?

> + Why implement 5?
>
> i. Noisy case: Need DTMF-only recognition (and may only have a  
> speechrecog)
I'm having difficulty following the logic of why noise would  
necessarily trigger START-OF-SPEECH if you were listening for speech  
but not DTMF. I suppose you can use a more forgiving discriminator if  
all you need to tell is if you're getting DTMF, but I've had a number  
of real-world cases where wind noise was detected as DTMF, and  
there's always the ambiguity when you have Captain Crunch on the  
phone. In either case in the non-optimized case it's the client who  
gets to decide whether an event should be interpreted as barge-in or  
not, so it seems an aesthetic protocol design decision whether the  
client tells the server ahead of time what circumstances to generate  
the START-OF-SPEECH event for, or whether the event gets generated  
and the client decides based on what's in the event whether it should  
be treated as barge- in.

I suppose one could make the argument that because the spec implies  
that the event can only be generated once per request that if a DTMF/ 
speech capable recognizer first hears enough noise to think it's  
hearing speech and later hears DTMF, the client will declare barge-in  
when the event comes and not when he DTMF actually gets heard.

If that's deemed a problem, we can still handle that in the design  
where the server just reports what it's hearing by allowing multiple  
events to be generated during a single request.

Between the approach just outlined above, and an approach where the  
client provides a filter for whether to generate the event or not, I  
have a mild preference (based on aesthetics rather than some hard  
engineering tradeoff) for the approach where the server just reports  
what it's hearing.


Having had some useful exchanges on this topic, it also is becoming  
apparent to me that this event is poorly named, and we should  
consider renaming it to "INTERESTING-INPUT-HEARD" or something akin  
to that, because as others have pointed out, a DTMF-only recognizer  
will never detect "start of speech".

Another consideration to fold into the design choice is  
extensibility. Bear with me through a little gedankenexperiment.

Suppose we want to define a new recognizer type, which I'll call the  
"name that tune" recognizer. The client plays music to the server and  
the server recognizes musical notes. The grammar is a standard  
musical notation, augmented with a semantic interpretation that  
transforms the notes into the title of the tune and provides that as  
an answer.

First, there's no speech involved (or is there...hang on a minute).  
Second, in order to accommodate the "name that tune" recognizer, we'd  
have to extend both the client and the server to undetstand a  
directive as to whether to recognize music or now, inaddition to what  
the server already knows what to do based on the grammar. If you  
follow my logic above, whether or not we do that, we have to extend  
"start-of-speech" to say "I'm hearing music". So far fairly  
straightforward, but let me now throw in the pathological twist.

Suppose what I feed to a  combined music/speech recognizer is a work  
in sprechstimme (spoken music), like the "Geographical Fugue" (aside:  
this is a wonderful piece of music I highly recommend to anyone  
interested in small ensemble singing). In this case, the tune could  
be named by either doing speech or music recognition. Why is there  
any need for the client to constrain the server as to which it tries  
to do when it's already told the server what it wants through the  
grammar?

A few other comments below

> ii. Flexibility: Want speech-only recognition (because a second  
> recogniser is doing hotword on DTMF)
>
I don't see how flexibility is affected by this deisgn choice. If  
that's what you want, feed the speech-only recognizer a grammar  
without any DTMF rules.

> + How to implement 5?
>
> a. Implicitly:
>    - dtmfrecog: always DTMF-only recognition
>    - speechrecog: depends on active grammar type
>        > if a dtmf grammar is active then DTMF input is "on"
>        > if a speech grammar is active then speech input is "on"
>
> b. Explicitly:
>    - Add inputmodes header to RECOGNIZE
>
> Option a is David's "do what I mean case"; option b is the extra  
> dial for the client.
>
Actually, that's not the point I was making with "do what I mean",  
but it's not essential to the discussion so let's move on.

> It is worth noting that VoiceXML uses option b. This allows one to  
> activate both speech grammars and DTMF grammars (and therefore be  
> informed of any errors in the grammars at activation time) but  
> independently turn on whichever input mode you like e.g. perhaps  
> start with "both" then change to "dtmf".
>
I'm not sure the VXML precedent is relevant here, because the  
application behind VXML is working a different part of the problem -  
how to traverse a TUI dialog based on different parts of the input  
space. In fact, I suspect that the VXML: choice was conditioned more  
by limitations at the time it was specified than an underlying good  
design choice. Clearly having to specify this in VXML make the job of  
handling a TUI with nodes like "Say or press 5" harder rather than  
easier.

Summing up, while I don't feel strongly one way or the other, I have  
a preference for handling this as follows:

a) Rename "START-OF-SPEECH" to "INTERESTING-INPUT-RECEIVED" or  
something equivalent.
b) Include a parameter in the event saying what was interesting about  
the input you received, with a registry of values which includes:
     - signal above noise floor
     - speech
     - dtmf
     - (possibly) music
c) allow the event to be generated multiple times during a request

Note that all of the above I'm saying with my technical hat on and my  
chair hat off.
Putting my chair hat on for a moment, we really need to get this spec  
to last call, so at some point Eric or I is going to declare rough  
consensus so we can move on.

Dave Oran.
> Dave
>
>
>
>
>
>
>
>
>
>
>
>
> Sarvi makes a good point that adding the reason why the START-OF- 
> SPEECH occurred does not fix the optimised bargin case.
>
> dtmfrecog - listens for DTMF only
> speechrecog - listens for DTMF only, or speech only, or speech & DTMF
>
>
>
>
>
>
>
>
> ----- Original Message ----- From: "Shanmugham, Saravanan"  
> <sarvi@cisco.com>
> To: "David R Oran" <oran@cisco.com>; "Dave Burke"  
> <david.burke@voxpilot.com>
> Cc: <speechsc@ietf.org>; "Klaus Reifenrath"  
> <Klaus.Reifenrath@Scansoft.com>
> Sent: Tuesday, July 05, 2005 8:51 PM
> Subject: RE: [Speechsc] START-OF-SPEECH in DTMF-only mode
>
>
> inline.
>
>     -----Original Message-----
>     From: David R Oran [mailto:oran@cisco.com]
>     Sent: Tuesday, July 05, 2005 11:08 AM
>     To: Dave Burke
>     Cc: Shanmugham, Saravanan; Klaus Reifenrath; speechsc@ietf.org
>     Subject: Re: [Speechsc] START-OF-SPEECH in DTMF-only mode
>
>
>     On Jul 5, 2005, at 12:54 PM, Dave Burke wrote:
>
>     > Inline.
>     >
>     > Dave
>     >
>     > ----- Original Message ----- From: "Shanmugham, Saravanan"
>     > <sarvi@cisco.com>
>     > To: "David R Oran" <oran@cisco.com>; "Klaus Reifenrath"
>     > <Klaus.Reifenrath@Scansoft.com>
>     > Cc: <speechsc@ietf.org>
>     > Sent: Tuesday, July 05, 2005 5:22 PM
>     > Subject: RE: [Speechsc] START-OF-SPEECH in DTMF-only mode
>     >
>     >
>     > I agree with Dave's analysis. The purpose of this event
>     was barge-in.
>     > And barge-in should happen for both DTMF and speech.
>     >
>     > Is there a case where you think it should not behave
>     this way. If soe,
>     > please provide a scenario where you think
>     >    1. Barge-in should happen for DTMF and not voice or
>     vice-versa.
>     >
>     > DB> You want to do a DTMF recognition only because it is noisy.
>     > While waiting for DTMF input, the speechrecog resource
>     (or advanced
>     > dtmfrecog) generates a START-OF-SPEECH because it heard
>     some speech.
>     > The client does not want to stop prompt playing unless
>     DTMF was heard
>     > but it can't tell by the START-OF-SPEECH whether speech
>     or DTMF was
>     > heard. Similarly vice versa.
>     >
>     It's an interesting design question what part of the
>     policy resides at the client and what at the server, and
>     who makes the "final decision" about whether what was
>     heard was relevant to the control channel. Right now we
>     (IMO) have a weird partitioning in many cases where the
>     client basically says "do what I mean", but there are no
>     constraints of what the server actually does, and no
>     normalized basis for the client to figure out what to set
>     various magic numbers to (e.g. sensitivity).
>
>     In this case the only thing the client needs to decide is
>     whether to kill the prompt because the server thinks
>     something that would interfere with the feedback
>     ear/mouth/finger control happened. What this says to me is
>     that it isn't necessarily a good idea for the client to
>     have more knobs to control the server (especially if those
>     knows are just more value/policy input ungrounded in any
>     physics/ acoustics). On the other hand, having the server
>     tell the client more about what it thinks is going on is
>     probably valuable.
>
>     So, Coming to the point after this long rambling
>     introduction, I think it would in fact be useful for the
>     START-Of-SPEECH event to indicate some extra information,
>     for example:
>     a) I got something enough above the noise floor to qualify
>     for exceeding the "Sensisitvity" parameter you sent in on
>     the request but I really can't tell what it is (could be a
>     hippopatmus fart, or a siren in the background, or captain
>     crunch trying to whistle DTMF).
>     b) I think I'm hearing speech
>     c) I think I'm hearing DTMF
>
> Though I agree with your former part of your response. I am not sure I
> agree with your proposed solution.
> The way I see this problem is that, it is more of what constitues a
> barge-in event. This boils down to whether it is speech, DTMF or both.
> This is inturn boils down to what type of recognizer resource we are
> using, dtmf-recog, speech-recog, and speech-only-recog(we don't have
> this and I don't think we should add it, but think of this as a place
> holder that explains the concept).
>
> A client knowing what type of barge-in happenned, does not impact the
> barge-in operation itself as it may be too late(for the optimized
> barge-in case). It may have other use cases, and if we can identify
> them, I don't mind adding support for the START-OF-SPEECH event to say
> what type of barge-in happenned. But that itself does not solve the
> original problem raised. Refer to my previous response.
>
> The solution lies in defining what what is a barge-in event. That  
> boils
> down to what type of recognition is happenning, dtmf-only, speech-dtmf
> or speech-only. We do not support speech-only as a resource today, the
> question is do we need a header to force it.
>
> Sarvi
>
>
>     >    2. You would benefit from the client knowing what caused the
>     > barge-in, DTMF Vs speech.
>     >
>     > DB> See previous comment. And previous e-mail: either add an
>     > inputmodes header (taking value speech, dtmf, both) to
>     the RECOGNIZE
>     > request or add a header to the START-OF-SPEECH event
>     indicating DTMF
>     > or speech.
>     >
>     I'm leaning in your direction on this latter point - as
>     should be evident from what I wrote above.
>
>     > Sarvi
>     >
>     >     -----Original Message-----
>     >     From: speechsc-bounces@ietf.org
>     >     [mailto:speechsc-bounces@ietf.org] On Behalf Of David R Oran
>     >     Sent: Tuesday, July 05, 2005 5:24 AM
>     >     To: Klaus Reifenrath
>     >     Cc: 'speechsc@ietf.org'
>     >     Subject: Re: [Speechsc] START-OF-SPEECH in DTMF-only mode
>     >
>     >
>     >     On Jul 5, 2005, at 3:46 AM, Reifenrath, Klaus wrote:
>     >
>     >     > The current spec is not clear when START-OF-SPEECH need
>     >     to be send in
>     >     > the following scenarios:
>     >     > A) The client requested a DTMF Recognizer. Is the
>     >     START-OF-SPEECH
>     >     > event send to the client also if speech was detected?
>     >     I suspect so, since one of the prime purposes is to enable
>     >     client- mediated barge-in handling. However, if the
>     >     recognizer is in fact only capable of recognizing DTMF
>     >     then it may in fact not report anythin unless it's using
>     >     some primitive thresholding machinery, like a SN threshold.
>     >     > B) The client requested a Speech Recognizer, but only
>     >     activated DTMF
>     >     > grammars. Is the START-OF-SPEECH event send to the
>     >     client also if
>     >     > speech was detected?
>     >     Again, I'd say yes, for the same reason as above.
>     >     > I think in both cases START-OF-SPEECH should only
>     be send after
>     >     > detecting a DTMF digit (see Figure 12 of VoiceXML
>     2.0: http://
>     >     > www.w3.org/TR/voicexml20/#dmlATiming).
>     >     We seem to have reached different conclusions. I'd be
>     >     interested in why you think my analysis above is wrong.
>     >
>     >     Dave.
>     >
>     >     > Klaus
>     >     >
>     >     > _______________________________________________
>     >     > Speechsc mailing list
>     >     > Speechsc@ietf.org
>     >     > https://www1.ietf.org/mailman/listinfo/speechsc
>     >     >
>     >
>     >     _______________________________________________
>     >     Speechsc mailing list
>     >     Speechsc@ietf.org
>     >     https://www1.ietf.org/mailman/listinfo/speechsc
>     >
>     >
>     > _______________________________________________
>     > Speechsc mailing list
>     > Speechsc@ietf.org
>     > https://www1.ietf.org/mailman/listinfo/speechsc
>     >
>
>
> _______________________________________________
> Speechsc mailing list
> Speechsc@ietf.org
> https://www1.ietf.org/mailman/listinfo/speechsc
>

_______________________________________________
Speechsc mailing list
Speechsc@ietf.org
https://www1.ietf.org/mailman/listinfo/speechsc



From speechsc-bounces@ietf.org Wed Jul 06 10:39:38 2005
Received: from localhost.localdomain ([127.0.0.1] helo=megatron.ietf.org)
	by megatron.ietf.org with esmtp (Exim 4.32)
	id 1DqB3q-0000WR-9J; Wed, 06 Jul 2005 10:39:38 -0400
Received: from odin.ietf.org ([132.151.1.176] helo=ietf.org)
	by megatron.ietf.org with esmtp (Exim 4.32) id 1DqB3n-0000Rf-6N
	for speechsc@megatron.ietf.org; Wed, 06 Jul 2005 10:39:37 -0400
Received: from ietf-mx.ietf.org (ietf-mx [132.151.6.1])
	by ietf.org (8.9.1a/8.9.1a) with ESMTP id KAA04579
	for <speechsc@ietf.org>; Wed, 6 Jul 2005 10:39:28 -0400 (EDT)
Received: from fw01.db01.voxpilot.com ([212.17.54.82] helo=mail.voxpilot.com)
	by ietf-mx.ietf.org with esmtp (Exim 4.33) id 1DqBKr-0001UD-Qb
	for speechsc@ietf.org; Wed, 06 Jul 2005 10:57:17 -0400
Received: from daburkewxp (unknown [10.0.0.203])
	by mail.voxpilot.com (Postfix) with ESMTP
	id B408E2140F0; Wed,  6 Jul 2005 14:29:12 +0000 (GMT)
Message-ID: <087701c58237$147262a0$cb00000a@db01.voxpilot.com>
From: "Dave Burke" <david.burke@voxpilot.com>
To: "David R Oran" <oran@cisco.com>
References: <03772D1EC8DE624A863058C75874A75C1126EE@vtg-um-e2k6.sj21ad.cisco.com><073a01c5820d$8a2c4530$cb00000a@db01.voxpilot.com>
	<1CC816CD-E142-421F-BB6C-BBE8572A41FD@cisco.com>
Subject: Re: [Speechsc] START-OF-SPEECH in DTMF-only mode -> summary(?)
Date: Wed, 6 Jul 2005 15:29:12 +0100
MIME-Version: 1.0
Content-Type: text/plain; format=flowed; charset="iso-8859-1";
	reply-type=response
Content-Transfer-Encoding: 7bit
X-Priority: 3
X-MSMail-Priority: Normal
X-Mailer: Microsoft Outlook Express 6.00.2900.2180
X-MimeOLE: Produced By Microsoft MimeOLE V6.00.2900.2180
X-Spam-Score: 0.0 (/)
X-Scan-Signature: 14278aea5bdd1edf35ec09ffb7b61f9d
Content-Transfer-Encoding: 7bit
Cc: speechsc@ietf.org, "Shanmugham, Saravanan" <sarvi@cisco.com>,
	Klaus Reifenrath <Klaus.Reifenrath@Scansoft.com>
X-BeenThere: speechsc@ietf.org
X-Mailman-Version: 2.1.5
Precedence: list
List-Id: Speech Services Control Working Group <speechsc.ietf.org>
List-Unsubscribe: <https://www1.ietf.org/mailman/listinfo/speechsc>,
	<mailto:speechsc-request@ietf.org?subject=unsubscribe>
List-Post: <mailto:speechsc@ietf.org>
List-Help: <mailto:speechsc-request@ietf.org?subject=help>
List-Subscribe: <https://www1.ietf.org/mailman/listinfo/speechsc>,
	<mailto:speechsc-request@ietf.org?subject=subscribe>
Sender: speechsc-bounces@ietf.org
Errors-To: speechsc-bounces@ietf.org

Inline.

Dave

----- Original Message ----- 
From: "David R Oran" <oran@cisco.com>
To: "Dave Burke" <david.burke@voxpilot.com>
Cc: <speechsc@ietf.org>; "Shanmugham, Saravanan" <sarvi@cisco.com>; "Klaus 
Reifenrath" <Klaus.Reifenrath@Scansoft.com>
Sent: Wednesday, July 06, 2005 1:07 PM
Subject: Re: [Speechsc] START-OF-SPEECH in DTMF-only mode -> summary(?)


>
> On Jul 6, 2005, at 5:31 AM, Dave Burke wrote:
>
>> + Attempting to summarise:
>>
>> 1. START-OF-SPEECH is useful for the client to know when to stop  playing 
>> prompts the in non-optimised case
>> 2. START-OF-SPEECH is useful for the client to calculate the bargin  time 
>> (e.g. VoiceXML 2.1 <mark>)
>> 3. In the optimised case, a bargin automatically stops prompt  playing 
>> (assuming prompts barginable)
>> 4. Because of the previous point, the question of what input type  caused 
>> bargin is different and less important to what input type(s)  the 
>> recogniser is listening for
> I'm not sure I follow point 4. Could you elaborate?

DB> Adding a parameter to the START-OF-SPEECH event would certainly allow 
the client to ignore (i.e. let prompts continue playing) the event if the 
event type is not of interest (e.g. the client would ignore speech start 
events when it is interested only in  DTMF start events). This _only_ works 
for the non-optimised case, however. For the optimised case, assuming 
START-OF-SPEECH coincides with the bargin signal to the speechsynth, prompts 
will stop playing for inputs that the client might not be interested in 
(e.g. a speech input will stop prompts playing even if the client is only 
interested in DTMF).

>
>> 5. Currently, MRCPv2 has no way of indicating what input type(s) a 
>> recogniser is listening for
>>
> Do you mean exactly this, or do you mean "for the client to indicate  to 
> the resource what input types it should look for"?

DB> Yes exactly - apologies for not being clear.

>
>> + Why implement 5?
>>
>> i. Noisy case: Need DTMF-only recognition (and may only have a 
>> speechrecog)
> I'm having difficulty following the logic of why noise would  necessarily 
> trigger START-OF-SPEECH if you were listening for speech  but not DTMF. I 
> suppose you can use a more forgiving discriminator if  all you need to 
> tell is if you're getting DTMF, but I've had a number  of real-world cases 
> where wind noise was detected as DTMF, and  there's always the ambiguity 
> when you have Captain Crunch on the  phone. In either case in the 
> non-optimized case it's the client who  gets to decide whether an event 
> should be interpreted as barge-in or  not, so it seems an aesthetic 
> protocol design decision whether the  client tells the server ahead of 
> time what circumstances to generate  the START-OF-SPEECH event for, or 
> whether the event gets generated  and the client decides based on what's 
> in the event whether it should  be treated as barge- in.
>
> I suppose one could make the argument that because the spec implies  that 
> the event can only be generated once per request that if a DTMF/ speech 
> capable recognizer first hears enough noise to think it's  hearing speech 
> and later hears DTMF, the client will declare barge-in  when the event 
> comes and not when he DTMF actually gets heard.
>
> If that's deemed a problem, we can still handle that in the design  where 
> the server just reports what it's hearing by allowing multiple  events to 
> be generated during a single request.
>
> Between the approach just outlined above, and an approach where the 
> client provides a filter for whether to generate the event or not, I  have 
> a mild preference (based on aesthetics rather than some hard  engineering 
> tradeoff) for the approach where the server just reports  what it's 
> hearing.
>
>
> Having had some useful exchanges on this topic, it also is becoming 
> apparent to me that this event is poorly named, and we should  consider 
> renaming it to "INTERESTING-INPUT-HEARD" or something akin  to that, 
> because as others have pointed out, a DTMF-only recognizer  will never 
> detect "start of speech".
>
> Another consideration to fold into the design choice is  extensibility. 
> Bear with me through a little gedankenexperiment.
>
> Suppose we want to define a new recognizer type, which I'll call the 
> "name that tune" recognizer. The client plays music to the server and  the 
> server recognizes musical notes. The grammar is a standard  musical 
> notation, augmented with a semantic interpretation that  transforms the 
> notes into the title of the tune and provides that as  an answer.
>
> First, there's no speech involved (or is there...hang on a minute). 
> Second, in order to accommodate the "name that tune" recognizer, we'd 
> have to extend both the client and the server to undetstand a  directive 
> as to whether to recognize music or now, inaddition to what  the server 
> already knows what to do based on the grammar. If you  follow my logic 
> above, whether or not we do that, we have to extend  "start-of-speech" to 
> say "I'm hearing music". So far fairly  straightforward, but let me now 
> throw in the pathological twist.
>
> Suppose what I feed to a  combined music/speech recognizer is a work  in 
> sprechstimme (spoken music), like the "Geographical Fugue" (aside:  this 
> is a wonderful piece of music I highly recommend to anyone  interested in 
> small ensemble singing). In this case, the tune could  be named by either 
> doing speech or music recognition. Why is there  any need for the client 
> to constrain the server as to which it tries  to do when it's already told 
> the server what it wants through the  grammar?
>
> A few other comments below
>
>> ii. Flexibility: Want speech-only recognition (because a second 
>> recogniser is doing hotword on DTMF)
>>
> I don't see how flexibility is affected by this deisgn choice. If  that's 
> what you want, feed the speech-only recognizer a grammar  without any DTMF 
> rules.
>
>> + How to implement 5?
>>
>> a. Implicitly:
>>    - dtmfrecog: always DTMF-only recognition
>>    - speechrecog: depends on active grammar type
>>        > if a dtmf grammar is active then DTMF input is "on"
>>        > if a speech grammar is active then speech input is "on"
>>
>> b. Explicitly:
>>    - Add inputmodes header to RECOGNIZE
>>
>> Option a is David's "do what I mean case"; option b is the extra  dial 
>> for the client.
>>
> Actually, that's not the point I was making with "do what I mean",  but 
> it's not essential to the discussion so let's move on.
>
>> It is worth noting that VoiceXML uses option b. This allows one to 
>> activate both speech grammars and DTMF grammars (and therefore be 
>> informed of any errors in the grammars at activation time) but 
>> independently turn on whichever input mode you like e.g. perhaps  start 
>> with "both" then change to "dtmf".
>>
> I'm not sure the VXML precedent is relevant here, because the  application 
> behind VXML is working a different part of the problem -  how to traverse 
> a TUI dialog based on different parts of the input  space. In fact, I 
> suspect that the VXML: choice was conditioned more  by limitations at the 
> time it was specified than an underlying good  design choice. Clearly 
> having to specify this in VXML make the job of  handling a TUI with nodes 
> like "Say or press 5" harder rather than  easier.

DB> The VoiceXML edge-case is pretty weird so it's not a major concern. My 
main concern is that the client can indicate, somehow, what the input modes 
are.

>
> Summing up, while I don't feel strongly one way or the other, I have  a 
> preference for handling this as follows:
>
> a) Rename "START-OF-SPEECH" to "INTERESTING-INPUT-RECEIVED" or  something 
> equivalent.
> b) Include a parameter in the event saying what was interesting about  the 
> input you received, with a registry of values which includes:
>     - signal above noise floor
>     - speech
>     - dtmf
>     - (possibly) music
> c) allow the event to be generated multiple times during a request
>

DB> I like these suggestions (START-OF-INPUT?). However, I don't see how the 
problem of the optimised case is not solved by them. I think the optimised 
case is fine if we have the following rules:

1. START-OF-SPEECH (and optimised bargin) is only generated for the input 
type that is being listened for
2. A speechrecog listens for DTMF if DTMF grammars are active, speech if 
speech grammars are active, or speech and DTMF if both grammar types are 
active.

> Note that all of the above I'm saying with my technical hat on and my 
> chair hat off.
> Putting my chair hat on for a moment, we really need to get this spec  to 
> last call, so at some point Eric or I is going to declare rough  consensus 
> so we can move on.

DB> Agreed!

>
> Dave Oran.
>> Dave
>>
>>
>>
>>
>>
>>
>>
>>
>>
>>
>>
>>
>> Sarvi makes a good point that adding the reason why the START-OF- SPEECH 
>> occurred does not fix the optimised bargin case.
>>
>> dtmfrecog - listens for DTMF only
>> speechrecog - listens for DTMF only, or speech only, or speech & DTMF
>>
>>
>>
>>
>>
>>
>>
>>
>> ----- Original Message ----- From: "Shanmugham, Saravanan" 
>> <sarvi@cisco.com>
>> To: "David R Oran" <oran@cisco.com>; "Dave Burke" 
>> <david.burke@voxpilot.com>
>> Cc: <speechsc@ietf.org>; "Klaus Reifenrath" 
>> <Klaus.Reifenrath@Scansoft.com>
>> Sent: Tuesday, July 05, 2005 8:51 PM
>> Subject: RE: [Speechsc] START-OF-SPEECH in DTMF-only mode
>>
>>
>> inline.
>>
>>     -----Original Message-----
>>     From: David R Oran [mailto:oran@cisco.com]
>>     Sent: Tuesday, July 05, 2005 11:08 AM
>>     To: Dave Burke
>>     Cc: Shanmugham, Saravanan; Klaus Reifenrath; speechsc@ietf.org
>>     Subject: Re: [Speechsc] START-OF-SPEECH in DTMF-only mode
>>
>>
>>     On Jul 5, 2005, at 12:54 PM, Dave Burke wrote:
>>
>>     > Inline.
>>     >
>>     > Dave
>>     >
>>     > ----- Original Message ----- From: "Shanmugham, Saravanan"
>>     > <sarvi@cisco.com>
>>     > To: "David R Oran" <oran@cisco.com>; "Klaus Reifenrath"
>>     > <Klaus.Reifenrath@Scansoft.com>
>>     > Cc: <speechsc@ietf.org>
>>     > Sent: Tuesday, July 05, 2005 5:22 PM
>>     > Subject: RE: [Speechsc] START-OF-SPEECH in DTMF-only mode
>>     >
>>     >
>>     > I agree with Dave's analysis. The purpose of this event
>>     was barge-in.
>>     > And barge-in should happen for both DTMF and speech.
>>     >
>>     > Is there a case where you think it should not behave
>>     this way. If soe,
>>     > please provide a scenario where you think
>>     >    1. Barge-in should happen for DTMF and not voice or
>>     vice-versa.
>>     >
>>     > DB> You want to do a DTMF recognition only because it is noisy.
>>     > While waiting for DTMF input, the speechrecog resource
>>     (or advanced
>>     > dtmfrecog) generates a START-OF-SPEECH because it heard
>>     some speech.
>>     > The client does not want to stop prompt playing unless
>>     DTMF was heard
>>     > but it can't tell by the START-OF-SPEECH whether speech
>>     or DTMF was
>>     > heard. Similarly vice versa.
>>     >
>>     It's an interesting design question what part of the
>>     policy resides at the client and what at the server, and
>>     who makes the "final decision" about whether what was
>>     heard was relevant to the control channel. Right now we
>>     (IMO) have a weird partitioning in many cases where the
>>     client basically says "do what I mean", but there are no
>>     constraints of what the server actually does, and no
>>     normalized basis for the client to figure out what to set
>>     various magic numbers to (e.g. sensitivity).
>>
>>     In this case the only thing the client needs to decide is
>>     whether to kill the prompt because the server thinks
>>     something that would interfere with the feedback
>>     ear/mouth/finger control happened. What this says to me is
>>     that it isn't necessarily a good idea for the client to
>>     have more knobs to control the server (especially if those
>>     knows are just more value/policy input ungrounded in any
>>     physics/ acoustics). On the other hand, having the server
>>     tell the client more about what it thinks is going on is
>>     probably valuable.
>>
>>     So, Coming to the point after this long rambling
>>     introduction, I think it would in fact be useful for the
>>     START-Of-SPEECH event to indicate some extra information,
>>     for example:
>>     a) I got something enough above the noise floor to qualify
>>     for exceeding the "Sensisitvity" parameter you sent in on
>>     the request but I really can't tell what it is (could be a
>>     hippopatmus fart, or a siren in the background, or captain
>>     crunch trying to whistle DTMF).
>>     b) I think I'm hearing speech
>>     c) I think I'm hearing DTMF
>>
>> Though I agree with your former part of your response. I am not sure I
>> agree with your proposed solution.
>> The way I see this problem is that, it is more of what constitues a
>> barge-in event. This boils down to whether it is speech, DTMF or both.
>> This is inturn boils down to what type of recognizer resource we are
>> using, dtmf-recog, speech-recog, and speech-only-recog(we don't have
>> this and I don't think we should add it, but think of this as a place
>> holder that explains the concept).
>>
>> A client knowing what type of barge-in happenned, does not impact the
>> barge-in operation itself as it may be too late(for the optimized
>> barge-in case). It may have other use cases, and if we can identify
>> them, I don't mind adding support for the START-OF-SPEECH event to say
>> what type of barge-in happenned. But that itself does not solve the
>> original problem raised. Refer to my previous response.
>>
>> The solution lies in defining what what is a barge-in event. That  boils
>> down to what type of recognition is happenning, dtmf-only, speech-dtmf
>> or speech-only. We do not support speech-only as a resource today, the
>> question is do we need a header to force it.
>>
>> Sarvi
>>
>>
>>     >    2. You would benefit from the client knowing what caused the
>>     > barge-in, DTMF Vs speech.
>>     >
>>     > DB> See previous comment. And previous e-mail: either add an
>>     > inputmodes header (taking value speech, dtmf, both) to
>>     the RECOGNIZE
>>     > request or add a header to the START-OF-SPEECH event
>>     indicating DTMF
>>     > or speech.
>>     >
>>     I'm leaning in your direction on this latter point - as
>>     should be evident from what I wrote above.
>>
>>     > Sarvi
>>     >
>>     >     -----Original Message-----
>>     >     From: speechsc-bounces@ietf.org
>>     >     [mailto:speechsc-bounces@ietf.org] On Behalf Of David R Oran
>>     >     Sent: Tuesday, July 05, 2005 5:24 AM
>>     >     To: Klaus Reifenrath
>>     >     Cc: 'speechsc@ietf.org'
>>     >     Subject: Re: [Speechsc] START-OF-SPEECH in DTMF-only mode
>>     >
>>     >
>>     >     On Jul 5, 2005, at 3:46 AM, Reifenrath, Klaus wrote:
>>     >
>>     >     > The current spec is not clear when START-OF-SPEECH need
>>     >     to be send in
>>     >     > the following scenarios:
>>     >     > A) The client requested a DTMF Recognizer. Is the
>>     >     START-OF-SPEECH
>>     >     > event send to the client also if speech was detected?
>>     >     I suspect so, since one of the prime purposes is to enable
>>     >     client- mediated barge-in handling. However, if the
>>     >     recognizer is in fact only capable of recognizing DTMF
>>     >     then it may in fact not report anythin unless it's using
>>     >     some primitive thresholding machinery, like a SN threshold.
>>     >     > B) The client requested a Speech Recognizer, but only
>>     >     activated DTMF
>>     >     > grammars. Is the START-OF-SPEECH event send to the
>>     >     client also if
>>     >     > speech was detected?
>>     >     Again, I'd say yes, for the same reason as above.
>>     >     > I think in both cases START-OF-SPEECH should only
>>     be send after
>>     >     > detecting a DTMF digit (see Figure 12 of VoiceXML
>>     2.0: http://
>>     >     > www.w3.org/TR/voicexml20/#dmlATiming).
>>     >     We seem to have reached different conclusions. I'd be
>>     >     interested in why you think my analysis above is wrong.
>>     >
>>     >     Dave.
>>     >
>>     >     > Klaus
>>     >     >
>>     >     > _______________________________________________
>>     >     > Speechsc mailing list
>>     >     > Speechsc@ietf.org
>>     >     > https://www1.ietf.org/mailman/listinfo/speechsc
>>     >     >
>>     >
>>     >     _______________________________________________
>>     >     Speechsc mailing list
>>     >     Speechsc@ietf.org
>>     >     https://www1.ietf.org/mailman/listinfo/speechsc
>>     >
>>     >
>>     > _______________________________________________
>>     > Speechsc mailing list
>>     > Speechsc@ietf.org
>>     > https://www1.ietf.org/mailman/listinfo/speechsc
>>     >
>>
>>
>> _______________________________________________
>> Speechsc mailing list
>> Speechsc@ietf.org
>> https://www1.ietf.org/mailman/listinfo/speechsc
>>
>
> _______________________________________________
> Speechsc mailing list
> Speechsc@ietf.org
> https://www1.ietf.org/mailman/listinfo/speechsc
> 


_______________________________________________
Speechsc mailing list
Speechsc@ietf.org
https://www1.ietf.org/mailman/listinfo/speechsc



From speechsc-bounces@ietf.org Wed Jul 06 10:55:32 2005
Received: from localhost.localdomain ([127.0.0.1] helo=megatron.ietf.org)
	by megatron.ietf.org with esmtp (Exim 4.32)
	id 1DqBJE-0004RF-4y; Wed, 06 Jul 2005 10:55:32 -0400
Received: from odin.ietf.org ([132.151.1.176] helo=ietf.org)
	by megatron.ietf.org with esmtp (Exim 4.32) id 1DqBJB-0004Nt-EX
	for speechsc@megatron.ietf.org; Wed, 06 Jul 2005 10:55:30 -0400
Received: from ietf-mx.ietf.org (ietf-mx [132.151.6.1])
	by ietf.org (8.9.1a/8.9.1a) with ESMTP id KAA06077
	for <speechsc@ietf.org>; Wed, 6 Jul 2005 10:55:22 -0400 (EDT)
Received: from sj-iport-1-in.cisco.com ([171.71.176.70]
	helo=sj-iport-1.cisco.com) by ietf-mx.ietf.org with esmtp (Exim 4.33)
	id 1DqBaX-0005Ao-TR
	for speechsc@ietf.org; Wed, 06 Jul 2005 11:13:26 -0400
Received: from sj-core-5.cisco.com (171.71.177.238)
	by sj-iport-1.cisco.com with ESMTP; 06 Jul 2005 07:45:23 -0700
X-IronPort-AV: i="3.93,265,1115017200"; 
	d="scan'208"; a="647275400:sNHT35807128"
Received: from imail.cisco.com (imail.cisco.com [128.107.200.91])
	by sj-core-5.cisco.com (8.12.10/8.12.6) with ESMTP id j66EjHVX008081;
	Wed, 6 Jul 2005 07:45:17 -0700 (PDT)
Received: from [10.32.245.152] (stealth-10-32-245-152.cisco.com
	[10.32.245.152])
	by imail.cisco.com (8.12.11/8.12.10) with SMTP id j66EiFFR008438;
	Wed, 6 Jul 2005 07:44:18 -0700
In-Reply-To: <087701c58237$147262a0$cb00000a@db01.voxpilot.com>
References: <03772D1EC8DE624A863058C75874A75C1126EE@vtg-um-e2k6.sj21ad.cisco.com><073a01c5820d$8a2c4530$cb00000a@db01.voxpilot.com>
	<1CC816CD-E142-421F-BB6C-BBE8572A41FD@cisco.com>
	<087701c58237$147262a0$cb00000a@db01.voxpilot.com>
Mime-Version: 1.0 (Apple Message framework v730)
X-Priority: 3
Content-Type: text/plain; charset=US-ASCII; delsp=yes; format=flowed
Message-Id: <11CA998B-CA62-494E-9868-9DC2EEF2DB04@cisco.com>
Content-Transfer-Encoding: 7bit
From: David R Oran <oran@cisco.com>
Subject: Re: [Speechsc] START-OF-SPEECH in DTMF-only mode -> summary(?)
Date: Wed, 6 Jul 2005 10:45:06 -0400
To: "Dave Burke" <david.burke@voxpilot.com>
X-Mailer: Apple Mail (2.730)
IIM-SIG: v:"1.1"; h:"imail.cisco.com"; d:"cisco.com"; z:"home"; m:"krs";
	t:"1120661063.176080"; x:"432200"; a:"rsa-sha1"; b:"nofws:14378";
	e:"Iw=="; n:"sQYarK2E51MdcTiUqeif3F7cWdxIfoCiXhdfb9vD5ee/j0jXL15gbFxF2p"
	"XIweAblu0N6XAgK7k+wrbr7bQDJaCDqOmzqpRUBjIRQAXQ7NzadpmR3pUL6wxaRUtW+c43sl9jC"
	"50Qg1sXHpPjt8Y+Y16ioyQAQAdSunM4YhevURc=";
	s:"PGO8a8D3UPb5IC/w63SFaZ1tM/Y8O1fPN40XgZkniOMDr/qQtFYmfj5bGJbXYkMuyUYQmztp"
	"v9mB4I/rX7mZozwQY6wukWFkdevEEJdZHN2oZtSCZ9A4k33RbkPhtjCqwy8LodyhNBqonK3/7qP"
	"RdUi/12RZPGrfNHNpZ702b34="; c:"From: David R Oran <oran@cisco.com>";
	c:"Subject: Re: [Speechsc] START-OF-SPEECH in DTMF-only mode ->
	summary" "(?)"; c:"Date: Wed, 6 Jul 2005 10:45:06 -0400"
IIM-VERIFY: s:"y"; v:"y"; r:"60"; h:"imail.cisco.com";
	c:"message from imail.cisco.com verified; "
X-Spam-Score: 0.0 (/)
X-Scan-Signature: d2bb8c7db26f2d12109e0cd8e454db52
Content-Transfer-Encoding: 7bit
Cc: speechsc@ietf.org, "Shanmugham, Saravanan" <sarvi@cisco.com>,
	Klaus Reifenrath <Klaus.Reifenrath@Scansoft.com>
X-BeenThere: speechsc@ietf.org
X-Mailman-Version: 2.1.5
Precedence: list
List-Id: Speech Services Control Working Group <speechsc.ietf.org>
List-Unsubscribe: <https://www1.ietf.org/mailman/listinfo/speechsc>,
	<mailto:speechsc-request@ietf.org?subject=unsubscribe>
List-Post: <mailto:speechsc@ietf.org>
List-Help: <mailto:speechsc-request@ietf.org?subject=help>
List-Subscribe: <https://www1.ietf.org/mailman/listinfo/speechsc>,
	<mailto:speechsc-request@ietf.org?subject=subscribe>
Sender: speechsc-bounces@ietf.org
Errors-To: speechsc-bounces@ietf.org

I think we're getting close. I though about snipping out some pieces  
to cut down the text, but I realized the context is still needed. See  
inline.

On Jul 6, 2005, at 10:29 AM, Dave Burke wrote:

> Inline.
>
> Dave
>
> ----- Original Message ----- From: "David R Oran" <oran@cisco.com>
> To: "Dave Burke" <david.burke@voxpilot.com>
> Cc: <speechsc@ietf.org>; "Shanmugham, Saravanan" <sarvi@cisco.com>;  
> "Klaus Reifenrath" <Klaus.Reifenrath@Scansoft.com>
> Sent: Wednesday, July 06, 2005 1:07 PM
> Subject: Re: [Speechsc] START-OF-SPEECH in DTMF-only mode -> summary 
> (?)
>
>
>
>>
>> On Jul 6, 2005, at 5:31 AM, Dave Burke wrote:
>>
>>
>>> + Attempting to summarise:
>>>
>>> 1. START-OF-SPEECH is useful for the client to know when to stop   
>>> playing prompts the in non-optimised case
>>> 2. START-OF-SPEECH is useful for the client to calculate the  
>>> bargin  time (e.g. VoiceXML 2.1 <mark>)
>>> 3. In the optimised case, a bargin automatically stops prompt   
>>> playing (assuming prompts barginable)
>>> 4. Because of the previous point, the question of what input  
>>> type  caused bargin is different and less important to what input  
>>> type(s)  the recogniser is listening for
>>>
>> I'm not sure I follow point 4. Could you elaborate?
>>
>
> DB> Adding a parameter to the START-OF-SPEECH event would certainly  
> allow the client to ignore (i.e. let prompts continue playing) the  
> event if the event type is not of interest (e.g. the client would  
> ignore speech start events when it is interested only in  DTMF  
> start events). This _only_ works for the non-optimised case,  
> however. For the optimised case, assuming START-OF-SPEECH coincides  
> with the bargin signal to the speechsynth, prompts will stop  
> playing for inputs that the client might not be interested in (e.g.  
> a speech input will stop prompts playing even if the client is only  
> interested in DTMF).
>
OK, now I get it. There's a need for the client to both handle the  
non-optimized case itself, and influence or at least have a clue what  
the server is going to do in the optimized case.

>
>>
>>
>>> 5. Currently, MRCPv2 has no way of indicating what input type(s)  
>>> a recogniser is listening for
>>>
>>>
>> Do you mean exactly this, or do you mean "for the client to  
>> indicate  to the resource what input types it should look for"?
>>
>
> DB> Yes exactly - apologies for not being clear.
>
>
>>
>>
>>> + Why implement 5?
>>>
>>> i. Noisy case: Need DTMF-only recognition (and may only have a  
>>> speechrecog)
>>>
>> I'm having difficulty following the logic of why noise would   
>> necessarily trigger START-OF-SPEECH if you were listening for  
>> speech  but not DTMF. I suppose you can use a more forgiving  
>> discriminator if  all you need to tell is if you're getting DTMF,  
>> but I've had a number  of real-world cases where wind noise was  
>> detected as DTMF, and  there's always the ambiguity when you have  
>> Captain Crunch on the  phone. In either case in the non-optimized  
>> case it's the client who  gets to decide whether an event should  
>> be interpreted as barge-in or  not, so it seems an aesthetic  
>> protocol design decision whether the  client tells the server  
>> ahead of time what circumstances to generate  the START-OF-SPEECH  
>> event for, or whether the event gets generated  and the client  
>> decides based on what's in the event whether it should  be treated  
>> as barge- in.
>>
>> I suppose one could make the argument that because the spec  
>> implies  that the event can only be generated once per request  
>> that if a DTMF/ speech capable recognizer first hears enough noise  
>> to think it's  hearing speech and later hears DTMF, the client  
>> will declare barge-in  when the event comes and not when he DTMF  
>> actually gets heard.
>>
>> If that's deemed a problem, we can still handle that in the  
>> design  where the server just reports what it's hearing by  
>> allowing multiple  events to be generated during a single request.
>>
>> Between the approach just outlined above, and an approach where  
>> the client provides a filter for whether to generate the event or  
>> not, I  have a mild preference (based on aesthetics rather than  
>> some hard  engineering tradeoff) for the approach where the server  
>> just reports  what it's hearing.
>>
>>
>> Having had some useful exchanges on this topic, it also is  
>> becoming apparent to me that this event is poorly named, and we  
>> should  consider renaming it to "INTERESTING-INPUT-HEARD" or  
>> something akin  to that, because as others have pointed out, a  
>> DTMF-only recognizer  will never detect "start of speech".
>>
>> Another consideration to fold into the design choice is   
>> extensibility. Bear with me through a little gedankenexperiment.
>>
>> Suppose we want to define a new recognizer type, which I'll call  
>> the "name that tune" recognizer. The client plays music to the  
>> server and  the server recognizes musical notes. The grammar is a  
>> standard  musical notation, augmented with a semantic  
>> interpretation that  transforms the notes into the title of the  
>> tune and provides that as  an answer.
>>
>> First, there's no speech involved (or is there...hang on a  
>> minute). Second, in order to accommodate the "name that tune"  
>> recognizer, we'd have to extend both the client and the server to  
>> undetstand a  directive as to whether to recognize music or now,  
>> inaddition to what  the server already knows what to do based on  
>> the grammar. If you  follow my logic above, whether or not we do  
>> that, we have to extend  "start-of-speech" to say "I'm hearing  
>> music". So far fairly  straightforward, but let me now throw in  
>> the pathological twist.
>>
>> Suppose what I feed to a  combined music/speech recognizer is a  
>> work  in sprechstimme (spoken music), like the "Geographical  
>> Fugue" (aside:  this is a wonderful piece of music I highly  
>> recommend to anyone  interested in small ensemble singing). In  
>> this case, the tune could  be named by either doing speech or  
>> music recognition. Why is there  any need for the client to  
>> constrain the server as to which it tries  to do when it's already  
>> told the server what it wants through the  grammar?
>>
>> A few other comments below
>>
>>
>>> ii. Flexibility: Want speech-only recognition (because a second  
>>> recogniser is doing hotword on DTMF)
>>>
>>>
>> I don't see how flexibility is affected by this deisgn choice. If   
>> that's what you want, feed the speech-only recognizer a grammar   
>> without any DTMF rules.
>>
>>
>>> + How to implement 5?
>>>
>>> a. Implicitly:
>>>    - dtmfrecog: always DTMF-only recognition
>>>    - speechrecog: depends on active grammar type
>>>        > if a dtmf grammar is active then DTMF input is "on"
>>>        > if a speech grammar is active then speech input is "on"
>>>
>>> b. Explicitly:
>>>    - Add inputmodes header to RECOGNIZE
>>>
>>> Option a is David's "do what I mean case"; option b is the extra   
>>> dial for the client.
>>>
>>>
>> Actually, that's not the point I was making with "do what I  
>> mean",  but it's not essential to the discussion so let's move on.
>>
>>
>>> It is worth noting that VoiceXML uses option b. This allows one  
>>> to activate both speech grammars and DTMF grammars (and therefore  
>>> be informed of any errors in the grammars at activation time) but  
>>> independently turn on whichever input mode you like e.g. perhaps   
>>> start with "both" then change to "dtmf".
>>>
>>>
>> I'm not sure the VXML precedent is relevant here, because the   
>> application behind VXML is working a different part of the problem  
>> -  how to traverse a TUI dialog based on different parts of the  
>> input  space. In fact, I suspect that the VXML: choice was  
>> conditioned more  by limitations at the time it was specified than  
>> an underlying good  design choice. Clearly having to specify this  
>> in VXML make the job of  handling a TUI with nodes like "Say or  
>> press 5" harder rather than  easier.
>>
>
> DB> The VoiceXML edge-case is pretty weird so it's not a major  
> concern. My main concern is that the client can indicate, somehow,  
> what the input modes are.
>
>
>>
>> Summing up, while I don't feel strongly one way or the other, I  
>> have  a preference for handling this as follows:
>>
>> a) Rename "START-OF-SPEECH" to "INTERESTING-INPUT-RECEIVED" or   
>> something equivalent.
>> b) Include a parameter in the event saying what was interesting  
>> about  the input you received, with a registry of values which  
>> includes:
>>     - signal above noise floor
>>     - speech
>>     - dtmf
>>     - (possibly) music
>> c) allow the event to be generated multiple times during a request
>>
>>
>
> DB> I like these suggestions (START-OF-INPUT?). However, I don't  
> see how the problem of the optimised case is not solved by them. I  
> think the optimised case is fine if we have the following rules:
>
Yes, I hadn't thought through the optimized case as thoroughly as  
you. Your suggested method name is fine by me as well

> 1. START-OF-SPEECH (and optimised bargin) is only generated for the  
> input type that is being listened for
> 2. A speechrecog listens for DTMF if DTMF grammars are active,  
> speech if speech grammars are active, or speech and DTMF if both  
> grammar types are active.
>
Works for me.

>
>> Note that all of the above I'm saying with my technical hat on and  
>> my chair hat off.
>> Putting my chair hat on for a moment, we really need to get this  
>> spec  to last call, so at some point Eric or I is going to declare  
>> rough  consensus so we can move on.
>>
>
> DB> Agreed!
>
>
>>
>> Dave Oran.
>>
>>> Dave
>>>
>>>
>>>
>>>
>>>
>>>
>>>
>>>
>>>
>>>
>>>
>>>
>>> Sarvi makes a good point that adding the reason why the START-OF-  
>>> SPEECH occurred does not fix the optimised bargin case.
>>>
>>> dtmfrecog - listens for DTMF only
>>> speechrecog - listens for DTMF only, or speech only, or speech &  
>>> DTMF
>>>
>>>
>>>
>>>
>>>
>>>
>>>
>>>
>>> ----- Original Message ----- From: "Shanmugham, Saravanan"  
>>> <sarvi@cisco.com>
>>> To: "David R Oran" <oran@cisco.com>; "Dave Burke"  
>>> <david.burke@voxpilot.com>
>>> Cc: <speechsc@ietf.org>; "Klaus Reifenrath"  
>>> <Klaus.Reifenrath@Scansoft.com>
>>> Sent: Tuesday, July 05, 2005 8:51 PM
>>> Subject: RE: [Speechsc] START-OF-SPEECH in DTMF-only mode
>>>
>>>
>>> inline.
>>>
>>>     -----Original Message-----
>>>     From: David R Oran [mailto:oran@cisco.com]
>>>     Sent: Tuesday, July 05, 2005 11:08 AM
>>>     To: Dave Burke
>>>     Cc: Shanmugham, Saravanan; Klaus Reifenrath; speechsc@ietf.org
>>>     Subject: Re: [Speechsc] START-OF-SPEECH in DTMF-only mode
>>>
>>>
>>>     On Jul 5, 2005, at 12:54 PM, Dave Burke wrote:
>>>
>>>     > Inline.
>>>     >
>>>     > Dave
>>>     >
>>>     > ----- Original Message ----- From: "Shanmugham, Saravanan"
>>>     > <sarvi@cisco.com>
>>>     > To: "David R Oran" <oran@cisco.com>; "Klaus Reifenrath"
>>>     > <Klaus.Reifenrath@Scansoft.com>
>>>     > Cc: <speechsc@ietf.org>
>>>     > Sent: Tuesday, July 05, 2005 5:22 PM
>>>     > Subject: RE: [Speechsc] START-OF-SPEECH in DTMF-only mode
>>>     >
>>>     >
>>>     > I agree with Dave's analysis. The purpose of this event
>>>     was barge-in.
>>>     > And barge-in should happen for both DTMF and speech.
>>>     >
>>>     > Is there a case where you think it should not behave
>>>     this way. If soe,
>>>     > please provide a scenario where you think
>>>     >    1. Barge-in should happen for DTMF and not voice or
>>>     vice-versa.
>>>     >
>>>     > DB> You want to do a DTMF recognition only because it is  
>>> noisy.
>>>     > While waiting for DTMF input, the speechrecog resource
>>>     (or advanced
>>>     > dtmfrecog) generates a START-OF-SPEECH because it heard
>>>     some speech.
>>>     > The client does not want to stop prompt playing unless
>>>     DTMF was heard
>>>     > but it can't tell by the START-OF-SPEECH whether speech
>>>     or DTMF was
>>>     > heard. Similarly vice versa.
>>>     >
>>>     It's an interesting design question what part of the
>>>     policy resides at the client and what at the server, and
>>>     who makes the "final decision" about whether what was
>>>     heard was relevant to the control channel. Right now we
>>>     (IMO) have a weird partitioning in many cases where the
>>>     client basically says "do what I mean", but there are no
>>>     constraints of what the server actually does, and no
>>>     normalized basis for the client to figure out what to set
>>>     various magic numbers to (e.g. sensitivity).
>>>
>>>     In this case the only thing the client needs to decide is
>>>     whether to kill the prompt because the server thinks
>>>     something that would interfere with the feedback
>>>     ear/mouth/finger control happened. What this says to me is
>>>     that it isn't necessarily a good idea for the client to
>>>     have more knobs to control the server (especially if those
>>>     knows are just more value/policy input ungrounded in any
>>>     physics/ acoustics). On the other hand, having the server
>>>     tell the client more about what it thinks is going on is
>>>     probably valuable.
>>>
>>>     So, Coming to the point after this long rambling
>>>     introduction, I think it would in fact be useful for the
>>>     START-Of-SPEECH event to indicate some extra information,
>>>     for example:
>>>     a) I got something enough above the noise floor to qualify
>>>     for exceeding the "Sensisitvity" parameter you sent in on
>>>     the request but I really can't tell what it is (could be a
>>>     hippopatmus fart, or a siren in the background, or captain
>>>     crunch trying to whistle DTMF).
>>>     b) I think I'm hearing speech
>>>     c) I think I'm hearing DTMF
>>>
>>> Though I agree with your former part of your response. I am not  
>>> sure I
>>> agree with your proposed solution.
>>> The way I see this problem is that, it is more of what constitues a
>>> barge-in event. This boils down to whether it is speech, DTMF or  
>>> both.
>>> This is inturn boils down to what type of recognizer resource we are
>>> using, dtmf-recog, speech-recog, and speech-only-recog(we don't have
>>> this and I don't think we should add it, but think of this as a  
>>> place
>>> holder that explains the concept).
>>>
>>> A client knowing what type of barge-in happenned, does not impact  
>>> the
>>> barge-in operation itself as it may be too late(for the optimized
>>> barge-in case). It may have other use cases, and if we can identify
>>> them, I don't mind adding support for the START-OF-SPEECH event  
>>> to say
>>> what type of barge-in happenned. But that itself does not solve the
>>> original problem raised. Refer to my previous response.
>>>
>>> The solution lies in defining what what is a barge-in event.  
>>> That  boils
>>> down to what type of recognition is happenning, dtmf-only, speech- 
>>> dtmf
>>> or speech-only. We do not support speech-only as a resource  
>>> today, the
>>> question is do we need a header to force it.
>>>
>>> Sarvi
>>>
>>>
>>>     >    2. You would benefit from the client knowing what caused  
>>> the
>>>     > barge-in, DTMF Vs speech.
>>>     >
>>>     > DB> See previous comment. And previous e-mail: either add an
>>>     > inputmodes header (taking value speech, dtmf, both) to
>>>     the RECOGNIZE
>>>     > request or add a header to the START-OF-SPEECH event
>>>     indicating DTMF
>>>     > or speech.
>>>     >
>>>     I'm leaning in your direction on this latter point - as
>>>     should be evident from what I wrote above.
>>>
>>>     > Sarvi
>>>     >
>>>     >     -----Original Message-----
>>>     >     From: speechsc-bounces@ietf.org
>>>     >     [mailto:speechsc-bounces@ietf.org] On Behalf Of David R  
>>> Oran
>>>     >     Sent: Tuesday, July 05, 2005 5:24 AM
>>>     >     To: Klaus Reifenrath
>>>     >     Cc: 'speechsc@ietf.org'
>>>     >     Subject: Re: [Speechsc] START-OF-SPEECH in DTMF-only mode
>>>     >
>>>     >
>>>     >     On Jul 5, 2005, at 3:46 AM, Reifenrath, Klaus wrote:
>>>     >
>>>     >     > The current spec is not clear when START-OF-SPEECH need
>>>     >     to be send in
>>>     >     > the following scenarios:
>>>     >     > A) The client requested a DTMF Recognizer. Is the
>>>     >     START-OF-SPEECH
>>>     >     > event send to the client also if speech was detected?
>>>     >     I suspect so, since one of the prime purposes is to enable
>>>     >     client- mediated barge-in handling. However, if the
>>>     >     recognizer is in fact only capable of recognizing DTMF
>>>     >     then it may in fact not report anythin unless it's using
>>>     >     some primitive thresholding machinery, like a SN  
>>> threshold.
>>>     >     > B) The client requested a Speech Recognizer, but only
>>>     >     activated DTMF
>>>     >     > grammars. Is the START-OF-SPEECH event send to the
>>>     >     client also if
>>>     >     > speech was detected?
>>>     >     Again, I'd say yes, for the same reason as above.
>>>     >     > I think in both cases START-OF-SPEECH should only
>>>     be send after
>>>     >     > detecting a DTMF digit (see Figure 12 of VoiceXML
>>>     2.0: http://
>>>     >     > www.w3.org/TR/voicexml20/#dmlATiming).
>>>     >     We seem to have reached different conclusions. I'd be
>>>     >     interested in why you think my analysis above is wrong.
>>>     >
>>>     >     Dave.
>>>     >
>>>     >     > Klaus
>>>     >     >
>>>     >     > _______________________________________________
>>>     >     > Speechsc mailing list
>>>     >     > Speechsc@ietf.org
>>>     >     > https://www1.ietf.org/mailman/listinfo/speechsc
>>>     >     >
>>>     >
>>>     >     _______________________________________________
>>>     >     Speechsc mailing list
>>>     >     Speechsc@ietf.org
>>>     >     https://www1.ietf.org/mailman/listinfo/speechsc
>>>     >
>>>     >
>>>     > _______________________________________________
>>>     > Speechsc mailing list
>>>     > Speechsc@ietf.org
>>>     > https://www1.ietf.org/mailman/listinfo/speechsc
>>>     >
>>>
>>>
>>> _______________________________________________
>>> Speechsc mailing list
>>> Speechsc@ietf.org
>>> https://www1.ietf.org/mailman/listinfo/speechsc
>>>
>>>
>>
>> _______________________________________________
>> Speechsc mailing list
>> Speechsc@ietf.org
>> https://www1.ietf.org/mailman/listinfo/speechsc
>

_______________________________________________
Speechsc mailing list
Speechsc@ietf.org
https://www1.ietf.org/mailman/listinfo/speechsc



From speechsc-bounces@ietf.org Wed Jul 06 11:27:31 2005
Received: from localhost.localdomain ([127.0.0.1] helo=megatron.ietf.org)
	by megatron.ietf.org with esmtp (Exim 4.32)
	id 1DqBoB-00042u-Bg; Wed, 06 Jul 2005 11:27:31 -0400
Received: from odin.ietf.org ([132.151.1.176] helo=ietf.org)
	by megatron.ietf.org with esmtp (Exim 4.32) id 1DqBny-0003vk-EG
	for speechsc@megatron.ietf.org; Wed, 06 Jul 2005 11:27:18 -0400
Received: from ietf-mx.ietf.org (ietf-mx [132.151.6.1])
	by ietf.org (8.9.1a/8.9.1a) with ESMTP id LAA09596
	for <speechsc@ietf.org>; Wed, 6 Jul 2005 11:27:13 -0400 (EDT)
Received: from e31.co.us.ibm.com ([32.97.110.129])
	by ietf-mx.ietf.org with esmtp (Exim 4.33) id 1DqCEU-0006y6-AP
	for speechsc@ietf.org; Wed, 06 Jul 2005 11:54:44 -0400
Received: from westrelay02.boulder.ibm.com (westrelay02.boulder.ibm.com
	[9.17.195.11])
	by e31.co.us.ibm.com (8.12.10/8.12.9) with ESMTP id j66FQb8D014394
	for <speechsc@ietf.org>; Wed, 6 Jul 2005 11:26:38 -0400
Received: from d03av04.boulder.ibm.com (d03av04.boulder.ibm.com [9.17.195.170])
	by westrelay02.boulder.ibm.com (8.12.10/NCO/VER6.6) with ESMTP id
	j66FQboO233824
	for <speechsc@ietf.org>; Wed, 6 Jul 2005 09:26:37 -0600
Received: from d03av04.boulder.ibm.com (loopback [127.0.0.1])
	by d03av04.boulder.ibm.com (8.12.11/8.13.3) with ESMTP id
	j66FQb1V014191
	for <speechsc@ietf.org>; Wed, 6 Jul 2005 09:26:37 -0600
Received: from d03nm119.boulder.ibm.com (d03nm119.boulder.ibm.com
	[9.17.195.145])
	by d03av04.boulder.ibm.com (8.12.11/8.12.11) with ESMTP id
	j66FQb40014182
	for <speechsc@ietf.org>; Wed, 6 Jul 2005 09:26:37 -0600
To: speechsc@ietf.org
MIME-Version: 1.0
X-Mailer: Lotus Notes Release 6.0.2CF1 June 9, 2003
Message-ID: <OF6F597EBB.0D67D8C6-ON87257036.0053BB4A-85257036.0054D5B9@us.ibm.com>
From: Brett Gavagni <gavagni@us.ibm.com>
Date: Wed, 6 Jul 2005 11:26:35 -0400
X-MIMETrack: Serialize by Router on D03NM119/03/M/IBM(Release 6.5.4|March 27,
	2005) at 07/06/2005 09:26:37,
	Serialize complete at 07/06/2005 09:26:37
Content-Type: text/plain; charset="US-ASCII"
X-Spam-Score: 0.0 (/)
X-Scan-Signature: 21c69d3cfc2dd19218717dbe1d974352
Subject: [Speechsc] Speed-Vs-Accuracy inconsistency with MRCPv2 and VoiceXML
	2.0
X-BeenThere: speechsc@ietf.org
X-Mailman-Version: 2.1.5
Precedence: list
List-Id: Speech Services Control Working Group <speechsc.ietf.org>
List-Unsubscribe: <https://www1.ietf.org/mailman/listinfo/speechsc>,
	<mailto:speechsc-request@ietf.org?subject=unsubscribe>
List-Post: <mailto:speechsc@ietf.org>
List-Help: <mailto:speechsc-request@ietf.org?subject=help>
List-Subscribe: <https://www1.ietf.org/mailman/listinfo/speechsc>,
	<mailto:speechsc-request@ietf.org?subject=subscribe>
Sender: speechsc-bounces@ietf.org
Errors-To: speechsc-bounces@ietf.org

Hi,

Currently there exists some inconsistent interpretation of the values for 
Speed-Vs-Accuracy in the MRCPv2 specification.

The current MRCPv2 draft states the following:
Speed Vs Accuracy 
 
   Depending on the implementation and capability of the recognizer 
   resource it may be tunable towards Performance or Accuracy. Higher 
   accuracy may mean more processing and higher CPU utilization, 
   meaning less calls per server and vice versa. This header is a float 
   value between 0.0 and 1.0 and allows this field to be tuned by the 
   speed-vs-accuracy header. This header field MAY occur in RECOGNIZE, 
   SET-PARAMS or GET-PARAMS. A higher value for this field means higher 
   speed. The default value for this field is platform specific. 
 
     speed-vs-accuracy   =     "Speed-Vs-Accuracy" ":" FLOAT CRLF

The VoiceXML 2.0 specification states the following:
http://www.w3.org/TR/voicexml20/#dml6.3.2
6.3.2 Generic Speech Recognizer Properties 

speedvsaccuracy 
A hint specifying the desired balance between speed vs. accuracy. 
A value of 0.0 means fastest recognition. A value of 1.0 means best 
accuracy. 
The value is a Real Number Designation (see Section 6.5). The default is 
value 0.5.

The VoiceXML spec, states that a 1.0 value means best accuracy, and the 
MRCPv2 draft states that the higher the value means higher speed.

Was it intentional to have the MRCPv2 draft convey an inconsistent 
interpretation of the values in comparison to VoiceXML 2.0?

Thanks,

Brett Gavagni 
WebSphere Voice Server Development 
http://www-306.ibm.com/software/pervasive/voice_server/
(561) 862-2097 T/L (975) 
gavagni@us.ibm.com


_______________________________________________
Speechsc mailing list
Speechsc@ietf.org
https://www1.ietf.org/mailman/listinfo/speechsc



From speechsc-bounces@ietf.org Wed Jul 06 11:46:01 2005
Received: from localhost.localdomain ([127.0.0.1] helo=megatron.ietf.org)
	by megatron.ietf.org with esmtp (Exim 4.32)
	id 1DqC4M-0007YJ-9e; Wed, 06 Jul 2005 11:44:14 -0400
Received: from odin.ietf.org ([132.151.1.176] helo=ietf.org)
	by megatron.ietf.org with esmtp (Exim 4.32) id 1DqC4I-0007Rw-Rz
	for speechsc@megatron.ietf.org; Wed, 06 Jul 2005 11:44:13 -0400
Received: from ietf-mx.ietf.org (ietf-mx [132.151.6.1])
	by ietf.org (8.9.1a/8.9.1a) with ESMTP id LAA12356
	for <speechsc@ietf.org>; Wed, 6 Jul 2005 11:44:07 -0400 (EDT)
Received: from sj-iport-1-in.cisco.com ([171.71.176.70]
	helo=sj-iport-1.cisco.com) by ietf-mx.ietf.org with esmtp (Exim 4.33)
	id 1DqCOM-0000kN-DY
	for speechsc@ietf.org; Wed, 06 Jul 2005 12:04:54 -0400
Received: from sj-core-1.cisco.com (171.71.177.237)
	by sj-iport-1.cisco.com with ESMTP; 06 Jul 2005 08:36:51 -0700
X-IronPort-AV: i="3.93,265,1115017200"; 
	d="scan'208"; a="647288291:sNHT29836508"
Received: from imail.cisco.com (imail.cisco.com [128.107.200.91])
	by sj-core-1.cisco.com (8.12.10/8.12.6) with ESMTP id j66FaovM004025;
	Wed, 6 Jul 2005 08:36:50 -0700 (PDT)
Received: from [10.32.245.152] (stealth-10-32-245-152.cisco.com
	[10.32.245.152])
	by imail.cisco.com (8.12.11/8.12.10) with SMTP id j66FZsrE008859;
	Wed, 6 Jul 2005 08:35:55 -0700
In-Reply-To: <OF6F597EBB.0D67D8C6-ON87257036.0053BB4A-85257036.0054D5B9@us.ibm.com>
References: <OF6F597EBB.0D67D8C6-ON87257036.0053BB4A-85257036.0054D5B9@us.ibm.com>
Mime-Version: 1.0 (Apple Message framework v730)
Content-Type: text/plain; charset=US-ASCII; delsp=yes; format=flowed
Message-Id: <EB43B7BF-E42F-472D-B867-DC3BEFD0A4FB@cisco.com>
Content-Transfer-Encoding: 7bit
From: David R Oran <oran@cisco.com>
Subject: Re: [Speechsc] Speed-Vs-Accuracy inconsistency with MRCPv2 and
	VoiceXML 2.0
Date: Wed, 6 Jul 2005 11:36:47 -0400
To: Brett Gavagni <gavagni@us.ibm.com>
X-Mailer: Apple Mail (2.730)
IIM-SIG: v:"1.1"; h:"imail.cisco.com"; d:"cisco.com"; z:"home"; m:"krs";
	t:"1120664155.684038"; x:"432200"; a:"rsa-sha1"; b:"nofws:1644";
	e:"Iw=="; n:"sQYarK2E51MdcTiUqeif3F7cWdxIfoCiXhdfb9vD5ee/j0jXL15gbFxF2p"
	"XIweAblu0N6XAgK7k+wrbr7bQDJaCDqOmzqpRUBjIRQAXQ7NzadpmR3pUL6wxaRUtW+c43sl9jC"
	"50Qg1sXHpPjt8Y+Y16ioyQAQAdSunM4YhevURc=";
	s:"n9LsV68C6cw9bPgat9f8O1GRoQA9g6bcCKUsHLmHVFEtMsREwwvT23oHqR2bNSLtkrvMz+WY"
	"4/evPdzL2+BuXCXAwwQKGQ7Gwrure5SYfaMOY8FnBUDZT8PujmC347HEzzXA2CMrPRy35pwVGyk"
	"+fuFux96SDNmJUoE+UR25aDs="; c:"From: David R Oran <oran@cisco.com>";
	c:"Subject: Re: [Speechsc] Speed-Vs-Accuracy inconsistency with MRCPv2
	" "and VoiceXML 2.0"; c:"Date: Wed, 6 Jul 2005 11:36:47 -0400"
IIM-VERIFY: s:"y"; v:"y"; r:"60"; h:"imail.cisco.com";
	c:"message from imail.cisco.com verified; "
X-Spam-Score: 0.0 (/)
X-Scan-Signature: 02ec665d00de228c50c93ed6b5e4fc1a
Content-Transfer-Encoding: 7bit
Cc: speechsc@ietf.org
X-BeenThere: speechsc@ietf.org
X-Mailman-Version: 2.1.5
Precedence: list
List-Id: Speech Services Control Working Group <speechsc.ietf.org>
List-Unsubscribe: <https://www1.ietf.org/mailman/listinfo/speechsc>,
	<mailto:speechsc-request@ietf.org?subject=unsubscribe>
List-Post: <mailto:speechsc@ietf.org>
List-Help: <mailto:speechsc-request@ietf.org?subject=help>
List-Subscribe: <https://www1.ietf.org/mailman/listinfo/speechsc>,
	<mailto:speechsc-request@ietf.org?subject=subscribe>
Sender: speechsc-bounces@ietf.org
Errors-To: speechsc-bounces@ietf.org


On Jul 6, 2005, at 11:26 AM, Brett Gavagni wrote:

> Hi,
>
> Currently there exists some inconsistent interpretation of the  
> values for
> Speed-Vs-Accuracy in the MRCPv2 specification.
>
> The current MRCPv2 draft states the following:
> Speed Vs Accuracy
>
>    Depending on the implementation and capability of the recognizer
>    resource it may be tunable towards Performance or Accuracy. Higher
>    accuracy may mean more processing and higher CPU utilization,
>    meaning less calls per server and vice versa. This header is a  
> float
>    value between 0.0 and 1.0 and allows this field to be tuned by the
>    speed-vs-accuracy header. This header field MAY occur in RECOGNIZE,
>    SET-PARAMS or GET-PARAMS. A higher value for this field means  
> higher
>    speed. The default value for this field is platform specific.
>
>      speed-vs-accuracy   =     "Speed-Vs-Accuracy" ":" FLOAT CRLF
>
> The VoiceXML 2.0 specification states the following:
> http://www.w3.org/TR/voicexml20/#dml6.3.2
> 6.3.2 Generic Speech Recognizer Properties
>
> speedvsaccuracy
> A hint specifying the desired balance between speed vs. accuracy.
> A value of 0.0 means fastest recognition. A value of 1.0 means best
> accuracy.
> The value is a Real Number Designation (see Section 6.5). The  
> default is
> value 0.5.
>
> The VoiceXML spec, states that a 1.0 value means best accuracy, and  
> the
> MRCPv2 draft states that the higher the value means higher speed.
>
> Was it intentional to have the MRCPv2 draft convey an inconsistent
> interpretation of the values in comparison to VoiceXML 2.0?
>
I don't think so. I have no objection to reversing it. Anyone else  
want to chime in?

> Thanks,
>
> Brett Gavagni
> WebSphere Voice Server Development
> http://www-306.ibm.com/software/pervasive/voice_server/
> (561) 862-2097 T/L (975)
> gavagni@us.ibm.com
>
>
> _______________________________________________
> Speechsc mailing list
> Speechsc@ietf.org
> https://www1.ietf.org/mailman/listinfo/speechsc
>

_______________________________________________
Speechsc mailing list
Speechsc@ietf.org
https://www1.ietf.org/mailman/listinfo/speechsc



From speechsc-bounces@ietf.org Thu Jul 07 08:40:26 2005
Received: from localhost.localdomain ([127.0.0.1] helo=megatron.ietf.org)
	by megatron.ietf.org with esmtp (Exim 4.32)
	id 1DqVg2-0000R0-Fp; Thu, 07 Jul 2005 08:40:26 -0400
Received: from odin.ietf.org ([132.151.1.176] helo=ietf.org)
	by megatron.ietf.org with esmtp (Exim 4.32) id 1DqVg0-0000NT-Gr
	for speechsc@megatron.ietf.org; Thu, 07 Jul 2005 08:40:24 -0400
Received: from ietf-mx.ietf.org (ietf-mx [132.151.6.1])
	by ietf.org (8.9.1a/8.9.1a) with ESMTP id IAA01572
	for <speechsc@ietf.org>; Thu, 7 Jul 2005 08:40:22 -0400 (EDT)
Received: from salvelinus.brooktrout.com ([204.176.205.6])
	by ietf-mx.ietf.org with esmtp (Exim 4.43) id 1DqW7B-00040z-DD
	for speechsc@ietf.org; Thu, 07 Jul 2005 09:08:30 -0400
Received: from nhmail2.needham.brooktrout.com (nhmail2.brooktrout.com
	[204.176.205.242])
	by salvelinus.brooktrout.com (8.12.5/8.12.5) with ESMTP id
	j67CbAhM016425
	for <speechsc@ietf.org>; Thu, 7 Jul 2005 08:37:10 -0400 (EDT)
Received: by nhmail2.brooktrout.com with Internet Mail Service (5.5.2653.19)
	id <NF1KLN1D>; Thu, 7 Jul 2005 08:32:09 -0400
Message-ID: <EDD694D47377D7119C8400D0B77FD331016AFA75@nhmail2.brooktrout.com>
From: Eric Burger <eburger@brooktrout.com>
To: speechsc@ietf.org
Subject: RE: [Speechsc] START-OF-SPEECH in DTMF-only mode -> summary(?)
Date: Thu, 7 Jul 2005 08:32:06 -0400 
MIME-Version: 1.0
X-Mailer: Internet Mail Service (5.5.2653.19)
Content-Type: text/plain
X-Spam-Score: 0.0 (/)
X-Scan-Signature: 9b157e6e8a3799aef911c0bc37fc93a6
X-BeenThere: speechsc@ietf.org
X-Mailman-Version: 2.1.5
Precedence: list
List-Id: Speech Services Control Working Group <speechsc.ietf.org>
List-Unsubscribe: <https://www1.ietf.org/mailman/listinfo/speechsc>,
	<mailto:speechsc-request@ietf.org?subject=unsubscribe>
List-Post: <mailto:speechsc@ietf.org>
List-Help: <mailto:speechsc-request@ietf.org?subject=help>
List-Subscribe: <https://www1.ietf.org/mailman/listinfo/speechsc>,
	<mailto:speechsc-request@ietf.org?subject=subscribe>
Sender: speechsc-bounces@ietf.org
Errors-To: speechsc-bounces@ietf.org

Can we declare consensus? 

> -----Original Message-----
> From: speechsc-bounces@ietf.org [mailto:speechsc-bounces@ietf.org] 
> Sent: Wednesday, July 06, 2005 10:45 AM
> To: Dave Burke
> Cc: speechsc@ietf.org; Shanmugham, Saravanan; Klaus Reifenrath
> Subject: Re: [Speechsc] START-OF-SPEECH in DTMF-only mode -> 
> summary(?)
> 
> I think we're getting close. I though about snipping out some 
> pieces to cut down the text, but I realized the context is 
> still needed. See inline.
> 
> On Jul 6, 2005, at 10:29 AM, Dave Burke wrote:
> 
> > Inline.
> >
> > Dave
> >
> > ----- Original Message ----- From: "David R Oran" <oran@cisco.com>
> > To: "Dave Burke" <david.burke@voxpilot.com>
> > Cc: <speechsc@ietf.org>; "Shanmugham, Saravanan" <sarvi@cisco.com>; 
> > "Klaus Reifenrath" <Klaus.Reifenrath@Scansoft.com>
> > Sent: Wednesday, July 06, 2005 1:07 PM
> > Subject: Re: [Speechsc] START-OF-SPEECH in DTMF-only mode -> summary
> > (?)
> >
> >
> >
> >>
> >> On Jul 6, 2005, at 5:31 AM, Dave Burke wrote:
> >>
> >>
> >>> + Attempting to summarise:
> >>>
> >>> 1. START-OF-SPEECH is useful for the client to know when 
> to stop   
> >>> playing prompts the in non-optimised case
> >>> 2. START-OF-SPEECH is useful for the client to calculate the  
> >>> bargin  time (e.g. VoiceXML 2.1 <mark>)
> >>> 3. In the optimised case, a bargin automatically stops prompt   
> >>> playing (assuming prompts barginable)
> >>> 4. Because of the previous point, the question of what input  
> >>> type  caused bargin is different and less important to 
> what input  
> >>> type(s)  the recogniser is listening for
> >>>
> >> I'm not sure I follow point 4. Could you elaborate?
> >>
> >
> > DB> Adding a parameter to the START-OF-SPEECH event would 
> certainly  
> > allow the client to ignore (i.e. let prompts continue playing) the  
> > event if the event type is not of interest (e.g. the client would  
> > ignore speech start events when it is interested only in  DTMF  
> > start events). This _only_ works for the non-optimised case,  
> > however. For the optimised case, assuming START-OF-SPEECH 
> coincides  
> > with the bargin signal to the speechsynth, prompts will stop  
> > playing for inputs that the client might not be interested 
> in (e.g.  
> > a speech input will stop prompts playing even if the client 
> is only  
> > interested in DTMF).
> >
> OK, now I get it. There's a need for the client to both handle the  
> non-optimized case itself, and influence or at least have a 
> clue what  
> the server is going to do in the optimized case.
> 
> >
> >>
> >>
> >>> 5. Currently, MRCPv2 has no way of indicating what input type(s)  
> >>> a recogniser is listening for
> >>>
> >>>
> >> Do you mean exactly this, or do you mean "for the client to  
> >> indicate  to the resource what input types it should look for"?
> >>
> >
> > DB> Yes exactly - apologies for not being clear.
> >
> >
> >>
> >>
> >>> + Why implement 5?
> >>>
> >>> i. Noisy case: Need DTMF-only recognition (and may only have a  
> >>> speechrecog)
> >>>
> >> I'm having difficulty following the logic of why noise would   
> >> necessarily trigger START-OF-SPEECH if you were listening for  
> >> speech  but not DTMF. I suppose you can use a more forgiving  
> >> discriminator if  all you need to tell is if you're getting DTMF,  
> >> but I've had a number  of real-world cases where wind noise was  
> >> detected as DTMF, and  there's always the ambiguity when you have  
> >> Captain Crunch on the  phone. In either case in the non-optimized  
> >> case it's the client who  gets to decide whether an event should  
> >> be interpreted as barge-in or  not, so it seems an aesthetic  
> >> protocol design decision whether the  client tells the server  
> >> ahead of time what circumstances to generate  the START-OF-SPEECH  
> >> event for, or whether the event gets generated  and the client  
> >> decides based on what's in the event whether it should  be 
> treated  
> >> as barge- in.
> >>
> >> I suppose one could make the argument that because the spec  
> >> implies  that the event can only be generated once per request  
> >> that if a DTMF/ speech capable recognizer first hears 
> enough noise  
> >> to think it's  hearing speech and later hears DTMF, the client  
> >> will declare barge-in  when the event comes and not when he DTMF  
> >> actually gets heard.
> >>
> >> If that's deemed a problem, we can still handle that in the  
> >> design  where the server just reports what it's hearing by  
> >> allowing multiple  events to be generated during a single request.
> >>
> >> Between the approach just outlined above, and an approach where  
> >> the client provides a filter for whether to generate the event or  
> >> not, I  have a mild preference (based on aesthetics rather than  
> >> some hard  engineering tradeoff) for the approach where 
> the server  
> >> just reports  what it's hearing.
> >>
> >>
> >> Having had some useful exchanges on this topic, it also is  
> >> becoming apparent to me that this event is poorly named, and we  
> >> should  consider renaming it to "INTERESTING-INPUT-HEARD" or  
> >> something akin  to that, because as others have pointed out, a  
> >> DTMF-only recognizer  will never detect "start of speech".
> >>
> >> Another consideration to fold into the design choice is   
> >> extensibility. Bear with me through a little gedankenexperiment.
> >>
> >> Suppose we want to define a new recognizer type, which I'll call  
> >> the "name that tune" recognizer. The client plays music to the  
> >> server and  the server recognizes musical notes. The grammar is a  
> >> standard  musical notation, augmented with a semantic  
> >> interpretation that  transforms the notes into the title of the  
> >> tune and provides that as  an answer.
> >>
> >> First, there's no speech involved (or is there...hang on a  
> >> minute). Second, in order to accommodate the "name that tune"  
> >> recognizer, we'd have to extend both the client and the server to  
> >> undetstand a  directive as to whether to recognize music or now,  
> >> inaddition to what  the server already knows what to do based on  
> >> the grammar. If you  follow my logic above, whether or not we do  
> >> that, we have to extend  "start-of-speech" to say "I'm hearing  
> >> music". So far fairly  straightforward, but let me now throw in  
> >> the pathological twist.
> >>
> >> Suppose what I feed to a  combined music/speech recognizer is a  
> >> work  in sprechstimme (spoken music), like the "Geographical  
> >> Fugue" (aside:  this is a wonderful piece of music I highly  
> >> recommend to anyone  interested in small ensemble singing). In  
> >> this case, the tune could  be named by either doing speech or  
> >> music recognition. Why is there  any need for the client to  
> >> constrain the server as to which it tries  to do when it's 
> already  
> >> told the server what it wants through the  grammar?
> >>
> >> A few other comments below
> >>
> >>
> >>> ii. Flexibility: Want speech-only recognition (because a second  
> >>> recogniser is doing hotword on DTMF)
> >>>
> >>>
> >> I don't see how flexibility is affected by this deisgn 
> choice. If   
> >> that's what you want, feed the speech-only recognizer a grammar   
> >> without any DTMF rules.
> >>
> >>
> >>> + How to implement 5?
> >>>
> >>> a. Implicitly:
> >>>    - dtmfrecog: always DTMF-only recognition
> >>>    - speechrecog: depends on active grammar type
> >>>        > if a dtmf grammar is active then DTMF input is "on"
> >>>        > if a speech grammar is active then speech input is "on"
> >>>
> >>> b. Explicitly:
> >>>    - Add inputmodes header to RECOGNIZE
> >>>
> >>> Option a is David's "do what I mean case"; option b is 
> the extra   
> >>> dial for the client.
> >>>
> >>>
> >> Actually, that's not the point I was making with "do what I  
> >> mean",  but it's not essential to the discussion so let's move on.
> >>
> >>
> >>> It is worth noting that VoiceXML uses option b. This allows one  
> >>> to activate both speech grammars and DTMF grammars (and 
> therefore  
> >>> be informed of any errors in the grammars at activation 
> time) but  
> >>> independently turn on whichever input mode you like e.g. 
> perhaps   
> >>> start with "both" then change to "dtmf".
> >>>
> >>>
> >> I'm not sure the VXML precedent is relevant here, because the   
> >> application behind VXML is working a different part of the 
> problem  
> >> -  how to traverse a TUI dialog based on different parts of the  
> >> input  space. In fact, I suspect that the VXML: choice was  
> >> conditioned more  by limitations at the time it was 
> specified than  
> >> an underlying good  design choice. Clearly having to specify this  
> >> in VXML make the job of  handling a TUI with nodes like "Say or  
> >> press 5" harder rather than  easier.
> >>
> >
> > DB> The VoiceXML edge-case is pretty weird so it's not a major  
> > concern. My main concern is that the client can indicate, somehow,  
> > what the input modes are.
> >
> >
> >>
> >> Summing up, while I don't feel strongly one way or the other, I  
> >> have  a preference for handling this as follows:
> >>
> >> a) Rename "START-OF-SPEECH" to "INTERESTING-INPUT-RECEIVED" or   
> >> something equivalent.
> >> b) Include a parameter in the event saying what was interesting  
> >> about  the input you received, with a registry of values which  
> >> includes:
> >>     - signal above noise floor
> >>     - speech
> >>     - dtmf
> >>     - (possibly) music
> >> c) allow the event to be generated multiple times during a request
> >>
> >>
> >
> > DB> I like these suggestions (START-OF-INPUT?). However, I don't  
> > see how the problem of the optimised case is not solved by them. I  
> > think the optimised case is fine if we have the following rules:
> >
> Yes, I hadn't thought through the optimized case as thoroughly as  
> you. Your suggested method name is fine by me as well
> 
> > 1. START-OF-SPEECH (and optimised bargin) is only generated 
> for the  
> > input type that is being listened for
> > 2. A speechrecog listens for DTMF if DTMF grammars are active,  
> > speech if speech grammars are active, or speech and DTMF if both  
> > grammar types are active.
> >
> Works for me.
> 
> >
> >> Note that all of the above I'm saying with my technical 
> hat on and  
> >> my chair hat off.
> >> Putting my chair hat on for a moment, we really need to get this  
> >> spec  to last call, so at some point Eric or I is going to 
> declare  
> >> rough  consensus so we can move on.
> >>
> >
> > DB> Agreed!
> >
> >
> >>
> >> Dave Oran.
> >>
> >>> Dave
> >>>
> >>>
> >>>
> >>>
> >>>
> >>>
> >>>
> >>>
> >>>
> >>>
> >>>
> >>>
> >>> Sarvi makes a good point that adding the reason why the 
> START-OF-  
> >>> SPEECH occurred does not fix the optimised bargin case.
> >>>
> >>> dtmfrecog - listens for DTMF only
> >>> speechrecog - listens for DTMF only, or speech only, or speech &  
> >>> DTMF
> >>>
> >>>
> >>>
> >>>
> >>>
> >>>
> >>>
> >>>
> >>> ----- Original Message ----- From: "Shanmugham, Saravanan"  
> >>> <sarvi@cisco.com>
> >>> To: "David R Oran" <oran@cisco.com>; "Dave Burke"  
> >>> <david.burke@voxpilot.com>
> >>> Cc: <speechsc@ietf.org>; "Klaus Reifenrath"  
> >>> <Klaus.Reifenrath@Scansoft.com>
> >>> Sent: Tuesday, July 05, 2005 8:51 PM
> >>> Subject: RE: [Speechsc] START-OF-SPEECH in DTMF-only mode
> >>>
> >>>
> >>> inline.
> >>>
> >>>     -----Original Message-----
> >>>     From: David R Oran [mailto:oran@cisco.com]
> >>>     Sent: Tuesday, July 05, 2005 11:08 AM
> >>>     To: Dave Burke
> >>>     Cc: Shanmugham, Saravanan; Klaus Reifenrath; speechsc@ietf.org
> >>>     Subject: Re: [Speechsc] START-OF-SPEECH in DTMF-only mode
> >>>
> >>>
> >>>     On Jul 5, 2005, at 12:54 PM, Dave Burke wrote:
> >>>
> >>>     > Inline.
> >>>     >
> >>>     > Dave
> >>>     >
> >>>     > ----- Original Message ----- From: "Shanmugham, Saravanan"
> >>>     > <sarvi@cisco.com>
> >>>     > To: "David R Oran" <oran@cisco.com>; "Klaus Reifenrath"
> >>>     > <Klaus.Reifenrath@Scansoft.com>
> >>>     > Cc: <speechsc@ietf.org>
> >>>     > Sent: Tuesday, July 05, 2005 5:22 PM
> >>>     > Subject: RE: [Speechsc] START-OF-SPEECH in DTMF-only mode
> >>>     >
> >>>     >
> >>>     > I agree with Dave's analysis. The purpose of this event
> >>>     was barge-in.
> >>>     > And barge-in should happen for both DTMF and speech.
> >>>     >
> >>>     > Is there a case where you think it should not behave
> >>>     this way. If soe,
> >>>     > please provide a scenario where you think
> >>>     >    1. Barge-in should happen for DTMF and not voice or
> >>>     vice-versa.
> >>>     >
> >>>     > DB> You want to do a DTMF recognition only because it is  
> >>> noisy.
> >>>     > While waiting for DTMF input, the speechrecog resource
> >>>     (or advanced
> >>>     > dtmfrecog) generates a START-OF-SPEECH because it heard
> >>>     some speech.
> >>>     > The client does not want to stop prompt playing unless
> >>>     DTMF was heard
> >>>     > but it can't tell by the START-OF-SPEECH whether speech
> >>>     or DTMF was
> >>>     > heard. Similarly vice versa.
> >>>     >
> >>>     It's an interesting design question what part of the
> >>>     policy resides at the client and what at the server, and
> >>>     who makes the "final decision" about whether what was
> >>>     heard was relevant to the control channel. Right now we
> >>>     (IMO) have a weird partitioning in many cases where the
> >>>     client basically says "do what I mean", but there are no
> >>>     constraints of what the server actually does, and no
> >>>     normalized basis for the client to figure out what to set
> >>>     various magic numbers to (e.g. sensitivity).
> >>>
> >>>     In this case the only thing the client needs to decide is
> >>>     whether to kill the prompt because the server thinks
> >>>     something that would interfere with the feedback
> >>>     ear/mouth/finger control happened. What this says to me is
> >>>     that it isn't necessarily a good idea for the client to
> >>>     have more knobs to control the server (especially if those
> >>>     knows are just more value/policy input ungrounded in any
> >>>     physics/ acoustics). On the other hand, having the server
> >>>     tell the client more about what it thinks is going on is
> >>>     probably valuable.
> >>>
> >>>     So, Coming to the point after this long rambling
> >>>     introduction, I think it would in fact be useful for the
> >>>     START-Of-SPEECH event to indicate some extra information,
> >>>     for example:
> >>>     a) I got something enough above the noise floor to qualify
> >>>     for exceeding the "Sensisitvity" parameter you sent in on
> >>>     the request but I really can't tell what it is (could be a
> >>>     hippopatmus fart, or a siren in the background, or captain
> >>>     crunch trying to whistle DTMF).
> >>>     b) I think I'm hearing speech
> >>>     c) I think I'm hearing DTMF
> >>>
> >>> Though I agree with your former part of your response. I am not  
> >>> sure I
> >>> agree with your proposed solution.
> >>> The way I see this problem is that, it is more of what 
> constitues a
> >>> barge-in event. This boils down to whether it is speech, DTMF or  
> >>> both.
> >>> This is inturn boils down to what type of recognizer 
> resource we are
> >>> using, dtmf-recog, speech-recog, and speech-only-recog(we 
> don't have
> >>> this and I don't think we should add it, but think of this as a  
> >>> place
> >>> holder that explains the concept).
> >>>
> >>> A client knowing what type of barge-in happenned, does 
> not impact  
> >>> the
> >>> barge-in operation itself as it may be too late(for the optimized
> >>> barge-in case). It may have other use cases, and if we 
> can identify
> >>> them, I don't mind adding support for the START-OF-SPEECH event  
> >>> to say
> >>> what type of barge-in happenned. But that itself does not 
> solve the
> >>> original problem raised. Refer to my previous response.
> >>>
> >>> The solution lies in defining what what is a barge-in event.  
> >>> That  boils
> >>> down to what type of recognition is happenning, 
> dtmf-only, speech- 
> >>> dtmf
> >>> or speech-only. We do not support speech-only as a resource  
> >>> today, the
> >>> question is do we need a header to force it.
> >>>
> >>> Sarvi
> >>>
> >>>
> >>>     >    2. You would benefit from the client knowing 
> what caused  
> >>> the
> >>>     > barge-in, DTMF Vs speech.
> >>>     >
> >>>     > DB> See previous comment. And previous e-mail: either add an
> >>>     > inputmodes header (taking value speech, dtmf, both) to
> >>>     the RECOGNIZE
> >>>     > request or add a header to the START-OF-SPEECH event
> >>>     indicating DTMF
> >>>     > or speech.
> >>>     >
> >>>     I'm leaning in your direction on this latter point - as
> >>>     should be evident from what I wrote above.
> >>>
> >>>     > Sarvi
> >>>     >
> >>>     >     -----Original Message-----
> >>>     >     From: speechsc-bounces@ietf.org
> >>>     >     [mailto:speechsc-bounces@ietf.org] On Behalf Of 
> David R  
> >>> Oran
> >>>     >     Sent: Tuesday, July 05, 2005 5:24 AM
> >>>     >     To: Klaus Reifenrath
> >>>     >     Cc: 'speechsc@ietf.org'
> >>>     >     Subject: Re: [Speechsc] START-OF-SPEECH in 
> DTMF-only mode
> >>>     >
> >>>     >
> >>>     >     On Jul 5, 2005, at 3:46 AM, Reifenrath, Klaus wrote:
> >>>     >
> >>>     >     > The current spec is not clear when 
> START-OF-SPEECH need
> >>>     >     to be send in
> >>>     >     > the following scenarios:
> >>>     >     > A) The client requested a DTMF Recognizer. Is the
> >>>     >     START-OF-SPEECH
> >>>     >     > event send to the client also if speech was detected?
> >>>     >     I suspect so, since one of the prime purposes 
> is to enable
> >>>     >     client- mediated barge-in handling. However, if the
> >>>     >     recognizer is in fact only capable of recognizing DTMF
> >>>     >     then it may in fact not report anythin unless it's using
> >>>     >     some primitive thresholding machinery, like a SN  
> >>> threshold.
> >>>     >     > B) The client requested a Speech Recognizer, but only
> >>>     >     activated DTMF
> >>>     >     > grammars. Is the START-OF-SPEECH event send to the
> >>>     >     client also if
> >>>     >     > speech was detected?
> >>>     >     Again, I'd say yes, for the same reason as above.
> >>>     >     > I think in both cases START-OF-SPEECH should only
> >>>     be send after
> >>>     >     > detecting a DTMF digit (see Figure 12 of VoiceXML
> >>>     2.0: http://
> >>>     >     > www.w3.org/TR/voicexml20/#dmlATiming).
> >>>     >     We seem to have reached different conclusions. I'd be
> >>>     >     interested in why you think my analysis above is wrong.
> >>>     >
> >>>     >     Dave.
> >>>     >
> >>>     >     > Klaus
> >>>     >     >
> >>>     >     > _______________________________________________
> >>>     >     > Speechsc mailing list
> >>>     >     > Speechsc@ietf.org
> >>>     >     > https://www1.ietf.org/mailman/listinfo/speechsc
> >>>     >     >
> >>>     >
> >>>     >     _______________________________________________
> >>>     >     Speechsc mailing list
> >>>     >     Speechsc@ietf.org
> >>>     >     https://www1.ietf.org/mailman/listinfo/speechsc
> >>>     >
> >>>     >
> >>>     > _______________________________________________
> >>>     > Speechsc mailing list
> >>>     > Speechsc@ietf.org
> >>>     > https://www1.ietf.org/mailman/listinfo/speechsc
> >>>     >
> >>>
> >>>
> >>> _______________________________________________
> >>> Speechsc mailing list
> >>> Speechsc@ietf.org
> >>> https://www1.ietf.org/mailman/listinfo/speechsc
> >>>
> >>>
> >>
> >> _______________________________________________
> >> Speechsc mailing list
> >> Speechsc@ietf.org
> >> https://www1.ietf.org/mailman/listinfo/speechsc
> >
> 
> _______________________________________________
> Speechsc mailing list
> Speechsc@ietf.org
> https://www1.ietf.org/mailman/listinfo/speechsc
> 


_______________________________________________
Speechsc mailing list
Speechsc@ietf.org
https://www1.ietf.org/mailman/listinfo/speechsc



From speechsc-bounces@ietf.org Thu Jul 07 08:40:28 2005
Received: from localhost.localdomain ([127.0.0.1] helo=megatron.ietf.org)
	by megatron.ietf.org with esmtp (Exim 4.32)
	id 1DqVg4-0000Ss-Q2; Thu, 07 Jul 2005 08:40:28 -0400
Received: from odin.ietf.org ([132.151.1.176] helo=ietf.org)
	by megatron.ietf.org with esmtp (Exim 4.32) id 1DqVg4-0000Sn-1p
	for speechsc@megatron.ietf.org; Thu, 07 Jul 2005 08:40:28 -0400
Received: from ietf-mx.ietf.org (ietf-mx [132.151.6.1])
	by ietf.org (8.9.1a/8.9.1a) with ESMTP id IAA01575
	for <speechsc@ietf.org>; Thu, 7 Jul 2005 08:40:26 -0400 (EDT)
Received: from salvelinus.brooktrout.com ([204.176.205.6])
	by ietf-mx.ietf.org with esmtp (Exim 4.43) id 1DqW7F-00040z-NQ
	for speechsc@ietf.org; Thu, 07 Jul 2005 09:08:34 -0400
Received: from nhmail2.needham.brooktrout.com (nhmail2.brooktrout.com
	[204.176.205.242])
	by salvelinus.brooktrout.com (8.12.5/8.12.5) with ESMTP id
	j67CcKhM016459
	for <speechsc@ietf.org>; Thu, 7 Jul 2005 08:38:20 -0400 (EDT)
Received: by nhmail2.brooktrout.com with Internet Mail Service (5.5.2653.19)
	id <NF1KLN11>; Thu, 7 Jul 2005 08:33:19 -0400
Message-ID: <EDD694D47377D7119C8400D0B77FD331016AFA76@nhmail2.brooktrout.com>
From: Eric Burger <eburger@brooktrout.com>
To: speechsc@ietf.org
Subject: RE: [Speechsc] Speed-Vs-Accuracy inconsistency with MRCPv2 andVoi
	ceXML 2.0
Date: Thu, 7 Jul 2005 08:33:12 -0400 
MIME-Version: 1.0
X-Mailer: Internet Mail Service (5.5.2653.19)
Content-Type: text/plain
X-Spam-Score: 0.0 (/)
X-Scan-Signature: b280b4db656c3ca28dd62e5e0b03daa8
X-BeenThere: speechsc@ietf.org
X-Mailman-Version: 2.1.5
Precedence: list
List-Id: Speech Services Control Working Group <speechsc.ietf.org>
List-Unsubscribe: <https://www1.ietf.org/mailman/listinfo/speechsc>,
	<mailto:speechsc-request@ietf.org?subject=unsubscribe>
List-Post: <mailto:speechsc@ietf.org>
List-Help: <mailto:speechsc-request@ietf.org?subject=help>
List-Subscribe: <https://www1.ietf.org/mailman/listinfo/speechsc>,
	<mailto:speechsc-request@ietf.org?subject=subscribe>
Sender: speechsc-bounces@ietf.org
Errors-To: speechsc-bounces@ietf.org

Does silence here mean, "OK with me" or "I haven't bothered to read the
mail"? 

> -----Original Message-----
> From: speechsc-bounces@ietf.org [mailto:speechsc-bounces@ietf.org] 
> Sent: Wednesday, July 06, 2005 11:37 AM
> To: Brett Gavagni
> Cc: speechsc@ietf.org
> Subject: Re: [Speechsc] Speed-Vs-Accuracy inconsistency with 
> MRCPv2 andVoiceXML 2.0
> 
> 
> On Jul 6, 2005, at 11:26 AM, Brett Gavagni wrote:
> 
> > Hi,
> >
> > Currently there exists some inconsistent interpretation of 
> the values 
> > for Speed-Vs-Accuracy in the MRCPv2 specification.
> >
> > The current MRCPv2 draft states the following:
> > Speed Vs Accuracy
> >
> >    Depending on the implementation and capability of the recognizer
> >    resource it may be tunable towards Performance or 
> Accuracy. Higher
> >    accuracy may mean more processing and higher CPU utilization,
> >    meaning less calls per server and vice versa. This header is a 
> > float
> >    value between 0.0 and 1.0 and allows this field to be 
> tuned by the
> >    speed-vs-accuracy header. This header field MAY occur in 
> RECOGNIZE,
> >    SET-PARAMS or GET-PARAMS. A higher value for this field means 
> > higher
> >    speed. The default value for this field is platform specific.
> >
> >      speed-vs-accuracy   =     "Speed-Vs-Accuracy" ":" FLOAT CRLF
> >
> > The VoiceXML 2.0 specification states the following:
> > http://www.w3.org/TR/voicexml20/#dml6.3.2
> > 6.3.2 Generic Speech Recognizer Properties
> >
> > speedvsaccuracy
> > A hint specifying the desired balance between speed vs. accuracy.
> > A value of 0.0 means fastest recognition. A value of 1.0 means best 
> > accuracy.
> > The value is a Real Number Designation (see Section 6.5). 
> The default 
> > is value 0.5.
> >
> > The VoiceXML spec, states that a 1.0 value means best accuracy, and 
> > the
> > MRCPv2 draft states that the higher the value means higher speed.
> >
> > Was it intentional to have the MRCPv2 draft convey an inconsistent 
> > interpretation of the values in comparison to VoiceXML 2.0?
> >
> I don't think so. I have no objection to reversing it. Anyone 
> else want to chime in?
> 
> > Thanks,
> >
> > Brett Gavagni
> > WebSphere Voice Server Development
> > http://www-306.ibm.com/software/pervasive/voice_server/
> > (561) 862-2097 T/L (975)
> > gavagni@us.ibm.com
> >
> >
> > _______________________________________________
> > Speechsc mailing list
> > Speechsc@ietf.org
> > https://www1.ietf.org/mailman/listinfo/speechsc
> >
> 
> _______________________________________________
> Speechsc mailing list
> Speechsc@ietf.org
> https://www1.ietf.org/mailman/listinfo/speechsc
> 


_______________________________________________
Speechsc mailing list
Speechsc@ietf.org
https://www1.ietf.org/mailman/listinfo/speechsc



From speechsc-bounces@ietf.org Thu Jul 07 09:04:39 2005
Received: from localhost.localdomain ([127.0.0.1] helo=megatron.ietf.org)
	by megatron.ietf.org with esmtp (Exim 4.32)
	id 1DqW3T-0001CL-Gw; Thu, 07 Jul 2005 09:04:39 -0400
Received: from odin.ietf.org ([132.151.1.176] helo=ietf.org)
	by megatron.ietf.org with esmtp (Exim 4.32) id 1DqW3N-00017H-Li
	for speechsc@megatron.ietf.org; Thu, 07 Jul 2005 09:04:35 -0400
Received: from ietf-mx.ietf.org (ietf-mx [132.151.6.1])
	by ietf.org (8.9.1a/8.9.1a) with ESMTP id JAA03567
	for <speechsc@ietf.org>; Thu, 7 Jul 2005 09:04:31 -0400 (EDT)
Received: from pb-exchcon2.scansoft.com ([199.4.160.64])
	by ietf-mx.ietf.org with esmtp (Exim 4.43) id 1DqWUU-0006FV-DZ
	for speechsc@ietf.org; Thu, 07 Jul 2005 09:32:39 -0400
Received: by pb-exchcon2.pb.scansoft.com with Internet Mail Service
	(5.5.2658.27) id <N5A2A96Z>; Thu, 7 Jul 2005 09:04:08 -0400
Message-ID: <BBF29C9B95E52E4DB5C29A0ACC94E83B016AA0AE@ac-exch1.eu.scansoft.com>
From: "Reifenrath, Klaus" <Klaus.Reifenrath@Scansoft.com>
To: "'Eric Burger'" <eburger@brooktrout.com>, speechsc@ietf.org
Subject: RE: [Speechsc] START-OF-SPEECH in DTMF-only mode -> summary(?)
Date: Thu, 7 Jul 2005 09:03:56 -0400 
MIME-Version: 1.0
X-Mailer: Internet Mail Service (5.5.2658.27)
Content-Type: text/plain
X-Spam-Score: 0.0 (/)
X-Scan-Signature: d67762704726a1bed57e7f4595960d34
Cc: 
X-BeenThere: speechsc@ietf.org
X-Mailman-Version: 2.1.5
Precedence: list
List-Id: Speech Services Control Working Group <speechsc.ietf.org>
List-Unsubscribe: <https://www1.ietf.org/mailman/listinfo/speechsc>,
	<mailto:speechsc-request@ietf.org?subject=unsubscribe>
List-Post: <mailto:speechsc@ietf.org>
List-Help: <mailto:speechsc-request@ietf.org?subject=help>
List-Subscribe: <https://www1.ietf.org/mailman/listinfo/speechsc>,
	<mailto:speechsc-request@ietf.org?subject=subscribe>
Sender: speechsc-bounces@ietf.org
Errors-To: speechsc-bounces@ietf.org

I) START-XYZ for the new event is misleading if we allow the event to be
generated multiple times.
Do we really need to introduce a new (optional) event? Why should a DTMF
only resource do speech detection? 

II)
  1. START-OF-SPEECH (and optimised bargin) is only generated for the  
  input type that is being listened for
  2. A speechrecog listens for DTMF if DTMF grammars are active,  
  speech if speech grammars are active, or speech and DTMF if both  
  grammar types are active.
I personally like this, BUT doesn't it conflict with VoiceXML 2.0 section
4.1.5.1 (bargein type speech):
"The prompt will be stopped as soon as speech or DTMF input is detected. The
prompt is stopped irrespective of whether or not the input matches a grammar
and irrespective of which grammars are active."  

Klaus

-----Original Message-----
From: Eric Burger [mailto:eburger@brooktrout.com] 
Sent: Donnerstag, 7. Juli 2005 14:32
To: speechsc@ietf.org
Subject: RE: [Speechsc] START-OF-SPEECH in DTMF-only mode -> summary(?)

Can we declare consensus? 

> -----Original Message-----
> From: speechsc-bounces@ietf.org [mailto:speechsc-bounces@ietf.org]
> Sent: Wednesday, July 06, 2005 10:45 AM
> To: Dave Burke
> Cc: speechsc@ietf.org; Shanmugham, Saravanan; Klaus Reifenrath
> Subject: Re: [Speechsc] START-OF-SPEECH in DTMF-only mode ->
> summary(?)
> 
> I think we're getting close. I though about snipping out some pieces 
> to cut down the text, but I realized the context is still needed. See 
> inline.
> 
> On Jul 6, 2005, at 10:29 AM, Dave Burke wrote:
> 
> > Inline.
> >
> > Dave
> >
> > ----- Original Message ----- From: "David R Oran" <oran@cisco.com>
> > To: "Dave Burke" <david.burke@voxpilot.com>
> > Cc: <speechsc@ietf.org>; "Shanmugham, Saravanan" <sarvi@cisco.com>; 
> > "Klaus Reifenrath" <Klaus.Reifenrath@Scansoft.com>
> > Sent: Wednesday, July 06, 2005 1:07 PM
> > Subject: Re: [Speechsc] START-OF-SPEECH in DTMF-only mode -> summary
> > (?)
> >
> >
> >
> >>
> >> On Jul 6, 2005, at 5:31 AM, Dave Burke wrote:
> >>
> >>
> >>> + Attempting to summarise:
> >>>
> >>> 1. START-OF-SPEECH is useful for the client to know when
> to stop   
> >>> playing prompts the in non-optimised case 2. START-OF-SPEECH is 
> >>> useful for the client to calculate the bargin  time (e.g. VoiceXML 
> >>> 2.1 <mark>)
> >>> 3. In the optimised case, a bargin automatically stops prompt   
> >>> playing (assuming prompts barginable) 4. Because of the previous 
> >>> point, the question of what input type  caused bargin is different 
> >>> and less important to
> what input
> >>> type(s)  the recogniser is listening for
> >>>
> >> I'm not sure I follow point 4. Could you elaborate?
> >>
> >
> > DB> Adding a parameter to the START-OF-SPEECH event would
> certainly
> > allow the client to ignore (i.e. let prompts continue playing) the 
> > event if the event type is not of interest (e.g. the client would 
> > ignore speech start events when it is interested only in  DTMF start 
> > events). This _only_ works for the non-optimised case, however. For 
> > the optimised case, assuming START-OF-SPEECH
> coincides
> > with the bargin signal to the speechsynth, prompts will stop playing 
> > for inputs that the client might not be interested
> in (e.g.  
> > a speech input will stop prompts playing even if the client
> is only
> > interested in DTMF).
> >
> OK, now I get it. There's a need for the client to both handle the 
> non-optimized case itself, and influence or at least have a clue what 
> the server is going to do in the optimized case.
> 
> >
> >>
> >>
> >>> 5. Currently, MRCPv2 has no way of indicating what input type(s) a 
> >>> recogniser is listening for
> >>>
> >>>
> >> Do you mean exactly this, or do you mean "for the client to 
> >> indicate  to the resource what input types it should look for"?
> >>
> >
> > DB> Yes exactly - apologies for not being clear.
> >
> >
> >>
> >>
> >>> + Why implement 5?
> >>>
> >>> i. Noisy case: Need DTMF-only recognition (and may only have a
> >>> speechrecog)
> >>>
> >> I'm having difficulty following the logic of why noise would   
> >> necessarily trigger START-OF-SPEECH if you were listening for 
> >> speech  but not DTMF. I suppose you can use a more forgiving 
> >> discriminator if  all you need to tell is if you're getting DTMF, 
> >> but I've had a number  of real-world cases where wind noise was 
> >> detected as DTMF, and  there's always the ambiguity when you have 
> >> Captain Crunch on the  phone. In either case in the non-optimized 
> >> case it's the client who  gets to decide whether an event should be 
> >> interpreted as barge-in or  not, so it seems an aesthetic protocol 
> >> design decision whether the  client tells the server ahead of time 
> >> what circumstances to generate  the START-OF-SPEECH event for, or 
> >> whether the event gets generated  and the client decides based on 
> >> what's in the event whether it should  be
> treated
> >> as barge- in.
> >>
> >> I suppose one could make the argument that because the spec implies  
> >> that the event can only be generated once per request that if a 
> >> DTMF/ speech capable recognizer first hears
> enough noise
> >> to think it's  hearing speech and later hears DTMF, the client will 
> >> declare barge-in  when the event comes and not when he DTMF 
> >> actually gets heard.
> >>
> >> If that's deemed a problem, we can still handle that in the design  
> >> where the server just reports what it's hearing by allowing 
> >> multiple  events to be generated during a single request.
> >>
> >> Between the approach just outlined above, and an approach where the 
> >> client provides a filter for whether to generate the event or not, 
> >> I  have a mild preference (based on aesthetics rather than some 
> >> hard  engineering tradeoff) for the approach where
> the server
> >> just reports  what it's hearing.
> >>
> >>
> >> Having had some useful exchanges on this topic, it also is becoming 
> >> apparent to me that this event is poorly named, and we should  
> >> consider renaming it to "INTERESTING-INPUT-HEARD" or something akin  
> >> to that, because as others have pointed out, a DTMF-only recognizer  
> >> will never detect "start of speech".
> >>
> >> Another consideration to fold into the design choice is   
> >> extensibility. Bear with me through a little gedankenexperiment.
> >>
> >> Suppose we want to define a new recognizer type, which I'll call 
> >> the "name that tune" recognizer. The client plays music to the 
> >> server and  the server recognizes musical notes. The grammar is a 
> >> standard  musical notation, augmented with a semantic 
> >> interpretation that  transforms the notes into the title of the 
> >> tune and provides that as  an answer.
> >>
> >> First, there's no speech involved (or is there...hang on a minute). 
> >> Second, in order to accommodate the "name that tune"
> >> recognizer, we'd have to extend both the client and the server to 
> >> undetstand a  directive as to whether to recognize music or now, 
> >> inaddition to what  the server already knows what to do based on 
> >> the grammar. If you  follow my logic above, whether or not we do 
> >> that, we have to extend  "start-of-speech" to say "I'm hearing 
> >> music". So far fairly  straightforward, but let me now throw in the 
> >> pathological twist.
> >>
> >> Suppose what I feed to a  combined music/speech recognizer is a 
> >> work  in sprechstimme (spoken music), like the "Geographical Fugue" 
> >> (aside:  this is a wonderful piece of music I highly recommend to 
> >> anyone  interested in small ensemble singing). In this case, the 
> >> tune could  be named by either doing speech or music recognition. 
> >> Why is there  any need for the client to constrain the server as to 
> >> which it tries  to do when it's
> already
> >> told the server what it wants through the  grammar?
> >>
> >> A few other comments below
> >>
> >>
> >>> ii. Flexibility: Want speech-only recognition (because a second 
> >>> recogniser is doing hotword on DTMF)
> >>>
> >>>
> >> I don't see how flexibility is affected by this deisgn
> choice. If   
> >> that's what you want, feed the speech-only recognizer a grammar   
> >> without any DTMF rules.
> >>
> >>
> >>> + How to implement 5?
> >>>
> >>> a. Implicitly:
> >>>    - dtmfrecog: always DTMF-only recognition
> >>>    - speechrecog: depends on active grammar type
> >>>        > if a dtmf grammar is active then DTMF input is "on"
> >>>        > if a speech grammar is active then speech input is "on"
> >>>
> >>> b. Explicitly:
> >>>    - Add inputmodes header to RECOGNIZE
> >>>
> >>> Option a is David's "do what I mean case"; option b is
> the extra   
> >>> dial for the client.
> >>>
> >>>
> >> Actually, that's not the point I was making with "do what I mean",  
> >> but it's not essential to the discussion so let's move on.
> >>
> >>
> >>> It is worth noting that VoiceXML uses option b. This allows one to 
> >>> activate both speech grammars and DTMF grammars (and
> therefore
> >>> be informed of any errors in the grammars at activation
> time) but
> >>> independently turn on whichever input mode you like e.g. 
> perhaps   
> >>> start with "both" then change to "dtmf".
> >>>
> >>>
> >> I'm not sure the VXML precedent is relevant here, because the   
> >> application behind VXML is working a different part of the
> problem
> >> -  how to traverse a TUI dialog based on different parts of the 
> >> input  space. In fact, I suspect that the VXML: choice was 
> >> conditioned more  by limitations at the time it was
> specified than
> >> an underlying good  design choice. Clearly having to specify this 
> >> in VXML make the job of  handling a TUI with nodes like "Say or 
> >> press 5" harder rather than  easier.
> >>
> >
> > DB> The VoiceXML edge-case is pretty weird so it's not a major
> > concern. My main concern is that the client can indicate, somehow, 
> > what the input modes are.
> >
> >
> >>
> >> Summing up, while I don't feel strongly one way or the other, I 
> >> have  a preference for handling this as follows:
> >>
> >> a) Rename "START-OF-SPEECH" to "INTERESTING-INPUT-RECEIVED" or   
> >> something equivalent.
> >> b) Include a parameter in the event saying what was interesting 
> >> about  the input you received, with a registry of values which
> >> includes:
> >>     - signal above noise floor
> >>     - speech
> >>     - dtmf
> >>     - (possibly) music
> >> c) allow the event to be generated multiple times during a request
> >>
> >>
> >
> > DB> I like these suggestions (START-OF-INPUT?). However, I don't
> > see how the problem of the optimised case is not solved by them. I 
> > think the optimised case is fine if we have the following rules:
> >
> Yes, I hadn't thought through the optimized case as thoroughly as you. 
> Your suggested method name is fine by me as well
> 
> > 1. START-OF-SPEECH (and optimised bargin) is only generated
> for the
> > input type that is being listened for 2. A speechrecog listens for 
> > DTMF if DTMF grammars are active, speech if speech grammars are 
> > active, or speech and DTMF if both grammar types are active.
> >
> Works for me.
> 
> >
> >> Note that all of the above I'm saying with my technical
> hat on and
> >> my chair hat off.
> >> Putting my chair hat on for a moment, we really need to get this 
> >> spec  to last call, so at some point Eric or I is going to
> declare
> >> rough  consensus so we can move on.
> >>
> >
> > DB> Agreed!
> >
> >
> >>
> >> Dave Oran.
> >>
> >>> Dave
> >>>
> >>>
> >>>
> >>>
> >>>
> >>>
> >>>
> >>>
> >>>
> >>>
> >>>
> >>>
> >>> Sarvi makes a good point that adding the reason why the
> START-OF-
> >>> SPEECH occurred does not fix the optimised bargin case.
> >>>
> >>> dtmfrecog - listens for DTMF only
> >>> speechrecog - listens for DTMF only, or speech only, or speech & 
> >>> DTMF
> >>>
> >>>
> >>>
> >>>
> >>>
> >>>
> >>>
> >>>
> >>> ----- Original Message ----- From: "Shanmugham, Saravanan"  
> >>> <sarvi@cisco.com>
> >>> To: "David R Oran" <oran@cisco.com>; "Dave Burke"  
> >>> <david.burke@voxpilot.com>
> >>> Cc: <speechsc@ietf.org>; "Klaus Reifenrath"  
> >>> <Klaus.Reifenrath@Scansoft.com>
> >>> Sent: Tuesday, July 05, 2005 8:51 PM
> >>> Subject: RE: [Speechsc] START-OF-SPEECH in DTMF-only mode
> >>>
> >>>
> >>> inline.
> >>>
> >>>     -----Original Message-----
> >>>     From: David R Oran [mailto:oran@cisco.com]
> >>>     Sent: Tuesday, July 05, 2005 11:08 AM
> >>>     To: Dave Burke
> >>>     Cc: Shanmugham, Saravanan; Klaus Reifenrath; speechsc@ietf.org
> >>>     Subject: Re: [Speechsc] START-OF-SPEECH in DTMF-only mode
> >>>
> >>>
> >>>     On Jul 5, 2005, at 12:54 PM, Dave Burke wrote:
> >>>
> >>>     > Inline.
> >>>     >
> >>>     > Dave
> >>>     >
> >>>     > ----- Original Message ----- From: "Shanmugham, Saravanan"
> >>>     > <sarvi@cisco.com>
> >>>     > To: "David R Oran" <oran@cisco.com>; "Klaus Reifenrath"
> >>>     > <Klaus.Reifenrath@Scansoft.com>
> >>>     > Cc: <speechsc@ietf.org>
> >>>     > Sent: Tuesday, July 05, 2005 5:22 PM
> >>>     > Subject: RE: [Speechsc] START-OF-SPEECH in DTMF-only mode
> >>>     >
> >>>     >
> >>>     > I agree with Dave's analysis. The purpose of this event
> >>>     was barge-in.
> >>>     > And barge-in should happen for both DTMF and speech.
> >>>     >
> >>>     > Is there a case where you think it should not behave
> >>>     this way. If soe,
> >>>     > please provide a scenario where you think
> >>>     >    1. Barge-in should happen for DTMF and not voice or
> >>>     vice-versa.
> >>>     >
> >>>     > DB> You want to do a DTMF recognition only because it is 
> >>> noisy.
> >>>     > While waiting for DTMF input, the speechrecog resource
> >>>     (or advanced
> >>>     > dtmfrecog) generates a START-OF-SPEECH because it heard
> >>>     some speech.
> >>>     > The client does not want to stop prompt playing unless
> >>>     DTMF was heard
> >>>     > but it can't tell by the START-OF-SPEECH whether speech
> >>>     or DTMF was
> >>>     > heard. Similarly vice versa.
> >>>     >
> >>>     It's an interesting design question what part of the
> >>>     policy resides at the client and what at the server, and
> >>>     who makes the "final decision" about whether what was
> >>>     heard was relevant to the control channel. Right now we
> >>>     (IMO) have a weird partitioning in many cases where the
> >>>     client basically says "do what I mean", but there are no
> >>>     constraints of what the server actually does, and no
> >>>     normalized basis for the client to figure out what to set
> >>>     various magic numbers to (e.g. sensitivity).
> >>>
> >>>     In this case the only thing the client needs to decide is
> >>>     whether to kill the prompt because the server thinks
> >>>     something that would interfere with the feedback
> >>>     ear/mouth/finger control happened. What this says to me is
> >>>     that it isn't necessarily a good idea for the client to
> >>>     have more knobs to control the server (especially if those
> >>>     knows are just more value/policy input ungrounded in any
> >>>     physics/ acoustics). On the other hand, having the server
> >>>     tell the client more about what it thinks is going on is
> >>>     probably valuable.
> >>>
> >>>     So, Coming to the point after this long rambling
> >>>     introduction, I think it would in fact be useful for the
> >>>     START-Of-SPEECH event to indicate some extra information,
> >>>     for example:
> >>>     a) I got something enough above the noise floor to qualify
> >>>     for exceeding the "Sensisitvity" parameter you sent in on
> >>>     the request but I really can't tell what it is (could be a
> >>>     hippopatmus fart, or a siren in the background, or captain
> >>>     crunch trying to whistle DTMF).
> >>>     b) I think I'm hearing speech
> >>>     c) I think I'm hearing DTMF
> >>>
> >>> Though I agree with your former part of your response. I am not 
> >>> sure I agree with your proposed solution.
> >>> The way I see this problem is that, it is more of what
> constitues a
> >>> barge-in event. This boils down to whether it is speech, DTMF or 
> >>> both.
> >>> This is inturn boils down to what type of recognizer
> resource we are
> >>> using, dtmf-recog, speech-recog, and speech-only-recog(we
> don't have
> >>> this and I don't think we should add it, but think of this as a 
> >>> place holder that explains the concept).
> >>>
> >>> A client knowing what type of barge-in happenned, does
> not impact
> >>> the
> >>> barge-in operation itself as it may be too late(for the optimized 
> >>> barge-in case). It may have other use cases, and if we
> can identify
> >>> them, I don't mind adding support for the START-OF-SPEECH event to 
> >>> say what type of barge-in happenned. But that itself does not
> solve the
> >>> original problem raised. Refer to my previous response.
> >>>
> >>> The solution lies in defining what what is a barge-in event.  
> >>> That  boils
> >>> down to what type of recognition is happenning,
> dtmf-only, speech-
> >>> dtmf
> >>> or speech-only. We do not support speech-only as a resource today, 
> >>> the question is do we need a header to force it.
> >>>
> >>> Sarvi
> >>>
> >>>
> >>>     >    2. You would benefit from the client knowing 
> what caused
> >>> the
> >>>     > barge-in, DTMF Vs speech.
> >>>     >
> >>>     > DB> See previous comment. And previous e-mail: either add an
> >>>     > inputmodes header (taking value speech, dtmf, both) to
> >>>     the RECOGNIZE
> >>>     > request or add a header to the START-OF-SPEECH event
> >>>     indicating DTMF
> >>>     > or speech.
> >>>     >
> >>>     I'm leaning in your direction on this latter point - as
> >>>     should be evident from what I wrote above.
> >>>
> >>>     > Sarvi
> >>>     >
> >>>     >     -----Original Message-----
> >>>     >     From: speechsc-bounces@ietf.org
> >>>     >     [mailto:speechsc-bounces@ietf.org] On Behalf Of 
> David R
> >>> Oran
> >>>     >     Sent: Tuesday, July 05, 2005 5:24 AM
> >>>     >     To: Klaus Reifenrath
> >>>     >     Cc: 'speechsc@ietf.org'
> >>>     >     Subject: Re: [Speechsc] START-OF-SPEECH in 
> DTMF-only mode
> >>>     >
> >>>     >
> >>>     >     On Jul 5, 2005, at 3:46 AM, Reifenrath, Klaus wrote:
> >>>     >
> >>>     >     > The current spec is not clear when 
> START-OF-SPEECH need
> >>>     >     to be send in
> >>>     >     > the following scenarios:
> >>>     >     > A) The client requested a DTMF Recognizer. Is the
> >>>     >     START-OF-SPEECH
> >>>     >     > event send to the client also if speech was detected?
> >>>     >     I suspect so, since one of the prime purposes 
> is to enable
> >>>     >     client- mediated barge-in handling. However, if the
> >>>     >     recognizer is in fact only capable of recognizing DTMF
> >>>     >     then it may in fact not report anythin unless it's using
> >>>     >     some primitive thresholding machinery, like a SN  
> >>> threshold.
> >>>     >     > B) The client requested a Speech Recognizer, but only
> >>>     >     activated DTMF
> >>>     >     > grammars. Is the START-OF-SPEECH event send to the
> >>>     >     client also if
> >>>     >     > speech was detected?
> >>>     >     Again, I'd say yes, for the same reason as above.
> >>>     >     > I think in both cases START-OF-SPEECH should only
> >>>     be send after
> >>>     >     > detecting a DTMF digit (see Figure 12 of VoiceXML
> >>>     2.0: http://
> >>>     >     > www.w3.org/TR/voicexml20/#dmlATiming).
> >>>     >     We seem to have reached different conclusions. I'd be
> >>>     >     interested in why you think my analysis above is wrong.
> >>>     >
> >>>     >     Dave.
> >>>     >
> >>>     >     > Klaus
> >>>     >     >
> >>>     >     > _______________________________________________
> >>>     >     > Speechsc mailing list
> >>>     >     > Speechsc@ietf.org
> >>>     >     > https://www1.ietf.org/mailman/listinfo/speechsc
> >>>     >     >
> >>>     >
> >>>     >     _______________________________________________
> >>>     >     Speechsc mailing list
> >>>     >     Speechsc@ietf.org
> >>>     >     https://www1.ietf.org/mailman/listinfo/speechsc
> >>>     >
> >>>     >
> >>>     > _______________________________________________
> >>>     > Speechsc mailing list
> >>>     > Speechsc@ietf.org
> >>>     > https://www1.ietf.org/mailman/listinfo/speechsc
> >>>     >
> >>>
> >>>
> >>> _______________________________________________
> >>> Speechsc mailing list
> >>> Speechsc@ietf.org
> >>> https://www1.ietf.org/mailman/listinfo/speechsc
> >>>
> >>>
> >>
> >> _______________________________________________
> >> Speechsc mailing list
> >> Speechsc@ietf.org
> >> https://www1.ietf.org/mailman/listinfo/speechsc
> >
> 
> _______________________________________________
> Speechsc mailing list
> Speechsc@ietf.org
> https://www1.ietf.org/mailman/listinfo/speechsc
> 


_______________________________________________
Speechsc mailing list
Speechsc@ietf.org
https://www1.ietf.org/mailman/listinfo/speechsc

_______________________________________________
Speechsc mailing list
Speechsc@ietf.org
https://www1.ietf.org/mailman/listinfo/speechsc



From speechsc-bounces@ietf.org Thu Jul 07 09:15:15 2005
Received: from localhost.localdomain ([127.0.0.1] helo=megatron.ietf.org)
	by megatron.ietf.org with esmtp (Exim 4.32)
	id 1DqWDj-0004M1-7J; Thu, 07 Jul 2005 09:15:15 -0400
Received: from odin.ietf.org ([132.151.1.176] helo=ietf.org)
	by megatron.ietf.org with esmtp (Exim 4.32) id 1DqWDh-0004IV-GY
	for speechsc@megatron.ietf.org; Thu, 07 Jul 2005 09:15:13 -0400
Received: from ietf-mx.ietf.org (ietf-mx [132.151.6.1])
	by ietf.org (8.9.1a/8.9.1a) with ESMTP id JAA04839
	for <speechsc@ietf.org>; Thu, 7 Jul 2005 09:15:11 -0400 (EDT)
Received: from fw01.db01.voxpilot.com ([212.17.54.82] helo=mail.voxpilot.com)
	by ietf-mx.ietf.org with esmtp (Exim 4.43) id 1DqWer-0006wJ-K1
	for speechsc@ietf.org; Thu, 07 Jul 2005 09:43:19 -0400
Received: from daburkewxp (unknown [10.0.0.203])
	by mail.voxpilot.com (Postfix) with ESMTP
	id 09924214041; Thu,  7 Jul 2005 13:14:49 +0000 (GMT)
Message-ID: <00c901c582f5$da41f040$cb00000a@db01.voxpilot.com>
From: "Dave Burke" <david.burke@voxpilot.com>
To: "Eric Burger" <eburger@brooktrout.com>, <speechsc@ietf.org>
References: <EDD694D47377D7119C8400D0B77FD331016AFA75@nhmail2.brooktrout.com>
Subject: Re: [Speechsc] START-OF-SPEECH in DTMF-only mode -> summary(?)
Date: Thu, 7 Jul 2005 14:14:48 +0100
MIME-Version: 1.0
Content-Type: text/plain; format=flowed; charset="iso-8859-1";
	reply-type=original
Content-Transfer-Encoding: 7bit
X-Priority: 3
X-MSMail-Priority: Normal
X-Mailer: Microsoft Outlook Express 6.00.2900.2180
X-MimeOLE: Produced By Microsoft MimeOLE V6.00.2900.2180
X-Spam-Score: 0.0 (/)
X-Scan-Signature: 07d4bcb4600b627a0786c2557bc62e06
Content-Transfer-Encoding: 7bit
Cc: 
X-BeenThere: speechsc@ietf.org
X-Mailman-Version: 2.1.5
Precedence: list
List-Id: Speech Services Control Working Group <speechsc.ietf.org>
List-Unsubscribe: <https://www1.ietf.org/mailman/listinfo/speechsc>,
	<mailto:speechsc-request@ietf.org?subject=unsubscribe>
List-Post: <mailto:speechsc@ietf.org>
List-Help: <mailto:speechsc-request@ietf.org?subject=help>
List-Subscribe: <https://www1.ietf.org/mailman/listinfo/speechsc>,
	<mailto:speechsc-request@ietf.org?subject=subscribe>
Sender: speechsc-bounces@ietf.org
Errors-To: speechsc-bounces@ietf.org

I consent!

----- Original Message ----- 
From: "Eric Burger" <eburger@brooktrout.com>
To: <speechsc@ietf.org>
Sent: Thursday, July 07, 2005 1:32 PM
Subject: RE: [Speechsc] START-OF-SPEECH in DTMF-only mode -> summary(?)


> Can we declare consensus? 
> 
>> -----Original Message-----
>> From: speechsc-bounces@ietf.org [mailto:speechsc-bounces@ietf.org] 
>> Sent: Wednesday, July 06, 2005 10:45 AM
>> To: Dave Burke
>> Cc: speechsc@ietf.org; Shanmugham, Saravanan; Klaus Reifenrath
>> Subject: Re: [Speechsc] START-OF-SPEECH in DTMF-only mode -> 
>> summary(?)
>> 
>> I think we're getting close. I though about snipping out some 
>> pieces to cut down the text, but I realized the context is 
>> still needed. See inline.
>> 
>> On Jul 6, 2005, at 10:29 AM, Dave Burke wrote:
>> 
>> > Inline.
>> >
>> > Dave
>> >
>> > ----- Original Message ----- From: "David R Oran" <oran@cisco.com>
>> > To: "Dave Burke" <david.burke@voxpilot.com>
>> > Cc: <speechsc@ietf.org>; "Shanmugham, Saravanan" <sarvi@cisco.com>; 
>> > "Klaus Reifenrath" <Klaus.Reifenrath@Scansoft.com>
>> > Sent: Wednesday, July 06, 2005 1:07 PM
>> > Subject: Re: [Speechsc] START-OF-SPEECH in DTMF-only mode -> summary
>> > (?)
>> >
>> >
>> >
>> >>
>> >> On Jul 6, 2005, at 5:31 AM, Dave Burke wrote:
>> >>
>> >>
>> >>> + Attempting to summarise:
>> >>>
>> >>> 1. START-OF-SPEECH is useful for the client to know when 
>> to stop   
>> >>> playing prompts the in non-optimised case
>> >>> 2. START-OF-SPEECH is useful for the client to calculate the  
>> >>> bargin  time (e.g. VoiceXML 2.1 <mark>)
>> >>> 3. In the optimised case, a bargin automatically stops prompt   
>> >>> playing (assuming prompts barginable)
>> >>> 4. Because of the previous point, the question of what input  
>> >>> type  caused bargin is different and less important to 
>> what input  
>> >>> type(s)  the recogniser is listening for
>> >>>
>> >> I'm not sure I follow point 4. Could you elaborate?
>> >>
>> >
>> > DB> Adding a parameter to the START-OF-SPEECH event would 
>> certainly  
>> > allow the client to ignore (i.e. let prompts continue playing) the  
>> > event if the event type is not of interest (e.g. the client would  
>> > ignore speech start events when it is interested only in  DTMF  
>> > start events). This _only_ works for the non-optimised case,  
>> > however. For the optimised case, assuming START-OF-SPEECH 
>> coincides  
>> > with the bargin signal to the speechsynth, prompts will stop  
>> > playing for inputs that the client might not be interested 
>> in (e.g.  
>> > a speech input will stop prompts playing even if the client 
>> is only  
>> > interested in DTMF).
>> >
>> OK, now I get it. There's a need for the client to both handle the  
>> non-optimized case itself, and influence or at least have a 
>> clue what  
>> the server is going to do in the optimized case.
>> 
>> >
>> >>
>> >>
>> >>> 5. Currently, MRCPv2 has no way of indicating what input type(s)  
>> >>> a recogniser is listening for
>> >>>
>> >>>
>> >> Do you mean exactly this, or do you mean "for the client to  
>> >> indicate  to the resource what input types it should look for"?
>> >>
>> >
>> > DB> Yes exactly - apologies for not being clear.
>> >
>> >
>> >>
>> >>
>> >>> + Why implement 5?
>> >>>
>> >>> i. Noisy case: Need DTMF-only recognition (and may only have a  
>> >>> speechrecog)
>> >>>
>> >> I'm having difficulty following the logic of why noise would   
>> >> necessarily trigger START-OF-SPEECH if you were listening for  
>> >> speech  but not DTMF. I suppose you can use a more forgiving  
>> >> discriminator if  all you need to tell is if you're getting DTMF,  
>> >> but I've had a number  of real-world cases where wind noise was  
>> >> detected as DTMF, and  there's always the ambiguity when you have  
>> >> Captain Crunch on the  phone. In either case in the non-optimized  
>> >> case it's the client who  gets to decide whether an event should  
>> >> be interpreted as barge-in or  not, so it seems an aesthetic  
>> >> protocol design decision whether the  client tells the server  
>> >> ahead of time what circumstances to generate  the START-OF-SPEECH  
>> >> event for, or whether the event gets generated  and the client  
>> >> decides based on what's in the event whether it should  be 
>> treated  
>> >> as barge- in.
>> >>
>> >> I suppose one could make the argument that because the spec  
>> >> implies  that the event can only be generated once per request  
>> >> that if a DTMF/ speech capable recognizer first hears 
>> enough noise  
>> >> to think it's  hearing speech and later hears DTMF, the client  
>> >> will declare barge-in  when the event comes and not when he DTMF  
>> >> actually gets heard.
>> >>
>> >> If that's deemed a problem, we can still handle that in the  
>> >> design  where the server just reports what it's hearing by  
>> >> allowing multiple  events to be generated during a single request.
>> >>
>> >> Between the approach just outlined above, and an approach where  
>> >> the client provides a filter for whether to generate the event or  
>> >> not, I  have a mild preference (based on aesthetics rather than  
>> >> some hard  engineering tradeoff) for the approach where 
>> the server  
>> >> just reports  what it's hearing.
>> >>
>> >>
>> >> Having had some useful exchanges on this topic, it also is  
>> >> becoming apparent to me that this event is poorly named, and we  
>> >> should  consider renaming it to "INTERESTING-INPUT-HEARD" or  
>> >> something akin  to that, because as others have pointed out, a  
>> >> DTMF-only recognizer  will never detect "start of speech".
>> >>
>> >> Another consideration to fold into the design choice is   
>> >> extensibility. Bear with me through a little gedankenexperiment.
>> >>
>> >> Suppose we want to define a new recognizer type, which I'll call  
>> >> the "name that tune" recognizer. The client plays music to the  
>> >> server and  the server recognizes musical notes. The grammar is a  
>> >> standard  musical notation, augmented with a semantic  
>> >> interpretation that  transforms the notes into the title of the  
>> >> tune and provides that as  an answer.
>> >>
>> >> First, there's no speech involved (or is there...hang on a  
>> >> minute). Second, in order to accommodate the "name that tune"  
>> >> recognizer, we'd have to extend both the client and the server to  
>> >> undetstand a  directive as to whether to recognize music or now,  
>> >> inaddition to what  the server already knows what to do based on  
>> >> the grammar. If you  follow my logic above, whether or not we do  
>> >> that, we have to extend  "start-of-speech" to say "I'm hearing  
>> >> music". So far fairly  straightforward, but let me now throw in  
>> >> the pathological twist.
>> >>
>> >> Suppose what I feed to a  combined music/speech recognizer is a  
>> >> work  in sprechstimme (spoken music), like the "Geographical  
>> >> Fugue" (aside:  this is a wonderful piece of music I highly  
>> >> recommend to anyone  interested in small ensemble singing). In  
>> >> this case, the tune could  be named by either doing speech or  
>> >> music recognition. Why is there  any need for the client to  
>> >> constrain the server as to which it tries  to do when it's 
>> already  
>> >> told the server what it wants through the  grammar?
>> >>
>> >> A few other comments below
>> >>
>> >>
>> >>> ii. Flexibility: Want speech-only recognition (because a second  
>> >>> recogniser is doing hotword on DTMF)
>> >>>
>> >>>
>> >> I don't see how flexibility is affected by this deisgn 
>> choice. If   
>> >> that's what you want, feed the speech-only recognizer a grammar   
>> >> without any DTMF rules.
>> >>
>> >>
>> >>> + How to implement 5?
>> >>>
>> >>> a. Implicitly:
>> >>>    - dtmfrecog: always DTMF-only recognition
>> >>>    - speechrecog: depends on active grammar type
>> >>>        > if a dtmf grammar is active then DTMF input is "on"
>> >>>        > if a speech grammar is active then speech input is "on"
>> >>>
>> >>> b. Explicitly:
>> >>>    - Add inputmodes header to RECOGNIZE
>> >>>
>> >>> Option a is David's "do what I mean case"; option b is 
>> the extra   
>> >>> dial for the client.
>> >>>
>> >>>
>> >> Actually, that's not the point I was making with "do what I  
>> >> mean",  but it's not essential to the discussion so let's move on.
>> >>
>> >>
>> >>> It is worth noting that VoiceXML uses option b. This allows one  
>> >>> to activate both speech grammars and DTMF grammars (and 
>> therefore  
>> >>> be informed of any errors in the grammars at activation 
>> time) but  
>> >>> independently turn on whichever input mode you like e.g. 
>> perhaps   
>> >>> start with "both" then change to "dtmf".
>> >>>
>> >>>
>> >> I'm not sure the VXML precedent is relevant here, because the   
>> >> application behind VXML is working a different part of the 
>> problem  
>> >> -  how to traverse a TUI dialog based on different parts of the  
>> >> input  space. In fact, I suspect that the VXML: choice was  
>> >> conditioned more  by limitations at the time it was 
>> specified than  
>> >> an underlying good  design choice. Clearly having to specify this  
>> >> in VXML make the job of  handling a TUI with nodes like "Say or  
>> >> press 5" harder rather than  easier.
>> >>
>> >
>> > DB> The VoiceXML edge-case is pretty weird so it's not a major  
>> > concern. My main concern is that the client can indicate, somehow,  
>> > what the input modes are.
>> >
>> >
>> >>
>> >> Summing up, while I don't feel strongly one way or the other, I  
>> >> have  a preference for handling this as follows:
>> >>
>> >> a) Rename "START-OF-SPEECH" to "INTERESTING-INPUT-RECEIVED" or   
>> >> something equivalent.
>> >> b) Include a parameter in the event saying what was interesting  
>> >> about  the input you received, with a registry of values which  
>> >> includes:
>> >>     - signal above noise floor
>> >>     - speech
>> >>     - dtmf
>> >>     - (possibly) music
>> >> c) allow the event to be generated multiple times during a request
>> >>
>> >>
>> >
>> > DB> I like these suggestions (START-OF-INPUT?). However, I don't  
>> > see how the problem of the optimised case is not solved by them. I  
>> > think the optimised case is fine if we have the following rules:
>> >
>> Yes, I hadn't thought through the optimized case as thoroughly as  
>> you. Your suggested method name is fine by me as well
>> 
>> > 1. START-OF-SPEECH (and optimised bargin) is only generated 
>> for the  
>> > input type that is being listened for
>> > 2. A speechrecog listens for DTMF if DTMF grammars are active,  
>> > speech if speech grammars are active, or speech and DTMF if both  
>> > grammar types are active.
>> >
>> Works for me.
>> 
>> >
>> >> Note that all of the above I'm saying with my technical 
>> hat on and  
>> >> my chair hat off.
>> >> Putting my chair hat on for a moment, we really need to get this  
>> >> spec  to last call, so at some point Eric or I is going to 
>> declare  
>> >> rough  consensus so we can move on.
>> >>
>> >
>> > DB> Agreed!
>> >
>> >
>> >>
>> >> Dave Oran.
>> >>
>> >>> Dave
>> >>>
>> >>>
>> >>>
>> >>>
>> >>>
>> >>>
>> >>>
>> >>>
>> >>>
>> >>>
>> >>>
>> >>>
>> >>> Sarvi makes a good point that adding the reason why the 
>> START-OF-  
>> >>> SPEECH occurred does not fix the optimised bargin case.
>> >>>
>> >>> dtmfrecog - listens for DTMF only
>> >>> speechrecog - listens for DTMF only, or speech only, or speech &  
>> >>> DTMF
>> >>>
>> >>>
>> >>>
>> >>>
>> >>>
>> >>>
>> >>>
>> >>>
>> >>> ----- Original Message ----- From: "Shanmugham, Saravanan"  
>> >>> <sarvi@cisco.com>
>> >>> To: "David R Oran" <oran@cisco.com>; "Dave Burke"  
>> >>> <david.burke@voxpilot.com>
>> >>> Cc: <speechsc@ietf.org>; "Klaus Reifenrath"  
>> >>> <Klaus.Reifenrath@Scansoft.com>
>> >>> Sent: Tuesday, July 05, 2005 8:51 PM
>> >>> Subject: RE: [Speechsc] START-OF-SPEECH in DTMF-only mode
>> >>>
>> >>>
>> >>> inline.
>> >>>
>> >>>     -----Original Message-----
>> >>>     From: David R Oran [mailto:oran@cisco.com]
>> >>>     Sent: Tuesday, July 05, 2005 11:08 AM
>> >>>     To: Dave Burke
>> >>>     Cc: Shanmugham, Saravanan; Klaus Reifenrath; speechsc@ietf.org
>> >>>     Subject: Re: [Speechsc] START-OF-SPEECH in DTMF-only mode
>> >>>
>> >>>
>> >>>     On Jul 5, 2005, at 12:54 PM, Dave Burke wrote:
>> >>>
>> >>>     > Inline.
>> >>>     >
>> >>>     > Dave
>> >>>     >
>> >>>     > ----- Original Message ----- From: "Shanmugham, Saravanan"
>> >>>     > <sarvi@cisco.com>
>> >>>     > To: "David R Oran" <oran@cisco.com>; "Klaus Reifenrath"
>> >>>     > <Klaus.Reifenrath@Scansoft.com>
>> >>>     > Cc: <speechsc@ietf.org>
>> >>>     > Sent: Tuesday, July 05, 2005 5:22 PM
>> >>>     > Subject: RE: [Speechsc] START-OF-SPEECH in DTMF-only mode
>> >>>     >
>> >>>     >
>> >>>     > I agree with Dave's analysis. The purpose of this event
>> >>>     was barge-in.
>> >>>     > And barge-in should happen for both DTMF and speech.
>> >>>     >
>> >>>     > Is there a case where you think it should not behave
>> >>>     this way. If soe,
>> >>>     > please provide a scenario where you think
>> >>>     >    1. Barge-in should happen for DTMF and not voice or
>> >>>     vice-versa.
>> >>>     >
>> >>>     > DB> You want to do a DTMF recognition only because it is  
>> >>> noisy.
>> >>>     > While waiting for DTMF input, the speechrecog resource
>> >>>     (or advanced
>> >>>     > dtmfrecog) generates a START-OF-SPEECH because it heard
>> >>>     some speech.
>> >>>     > The client does not want to stop prompt playing unless
>> >>>     DTMF was heard
>> >>>     > but it can't tell by the START-OF-SPEECH whether speech
>> >>>     or DTMF was
>> >>>     > heard. Similarly vice versa.
>> >>>     >
>> >>>     It's an interesting design question what part of the
>> >>>     policy resides at the client and what at the server, and
>> >>>     who makes the "final decision" about whether what was
>> >>>     heard was relevant to the control channel. Right now we
>> >>>     (IMO) have a weird partitioning in many cases where the
>> >>>     client basically says "do what I mean", but there are no
>> >>>     constraints of what the server actually does, and no
>> >>>     normalized basis for the client to figure out what to set
>> >>>     various magic numbers to (e.g. sensitivity).
>> >>>
>> >>>     In this case the only thing the client needs to decide is
>> >>>     whether to kill the prompt because the server thinks
>> >>>     something that would interfere with the feedback
>> >>>     ear/mouth/finger control happened. What this says to me is
>> >>>     that it isn't necessarily a good idea for the client to
>> >>>     have more knobs to control the server (especially if those
>> >>>     knows are just more value/policy input ungrounded in any
>> >>>     physics/ acoustics). On the other hand, having the server
>> >>>     tell the client more about what it thinks is going on is
>> >>>     probably valuable.
>> >>>
>> >>>     So, Coming to the point after this long rambling
>> >>>     introduction, I think it would in fact be useful for the
>> >>>     START-Of-SPEECH event to indicate some extra information,
>> >>>     for example:
>> >>>     a) I got something enough above the noise floor to qualify
>> >>>     for exceeding the "Sensisitvity" parameter you sent in on
>> >>>     the request but I really can't tell what it is (could be a
>> >>>     hippopatmus fart, or a siren in the background, or captain
>> >>>     crunch trying to whistle DTMF).
>> >>>     b) I think I'm hearing speech
>> >>>     c) I think I'm hearing DTMF
>> >>>
>> >>> Though I agree with your former part of your response. I am not  
>> >>> sure I
>> >>> agree with your proposed solution.
>> >>> The way I see this problem is that, it is more of what 
>> constitues a
>> >>> barge-in event. This boils down to whether it is speech, DTMF or  
>> >>> both.
>> >>> This is inturn boils down to what type of recognizer 
>> resource we are
>> >>> using, dtmf-recog, speech-recog, and speech-only-recog(we 
>> don't have
>> >>> this and I don't think we should add it, but think of this as a  
>> >>> place
>> >>> holder that explains the concept).
>> >>>
>> >>> A client knowing what type of barge-in happenned, does 
>> not impact  
>> >>> the
>> >>> barge-in operation itself as it may be too late(for the optimized
>> >>> barge-in case). It may have other use cases, and if we 
>> can identify
>> >>> them, I don't mind adding support for the START-OF-SPEECH event  
>> >>> to say
>> >>> what type of barge-in happenned. But that itself does not 
>> solve the
>> >>> original problem raised. Refer to my previous response.
>> >>>
>> >>> The solution lies in defining what what is a barge-in event.  
>> >>> That  boils
>> >>> down to what type of recognition is happenning, 
>> dtmf-only, speech- 
>> >>> dtmf
>> >>> or speech-only. We do not support speech-only as a resource  
>> >>> today, the
>> >>> question is do we need a header to force it.
>> >>>
>> >>> Sarvi
>> >>>
>> >>>
>> >>>     >    2. You would benefit from the client knowing 
>> what caused  
>> >>> the
>> >>>     > barge-in, DTMF Vs speech.
>> >>>     >
>> >>>     > DB> See previous comment. And previous e-mail: either add an
>> >>>     > inputmodes header (taking value speech, dtmf, both) to
>> >>>     the RECOGNIZE
>> >>>     > request or add a header to the START-OF-SPEECH event
>> >>>     indicating DTMF
>> >>>     > or speech.
>> >>>     >
>> >>>     I'm leaning in your direction on this latter point - as
>> >>>     should be evident from what I wrote above.
>> >>>
>> >>>     > Sarvi
>> >>>     >
>> >>>     >     -----Original Message-----
>> >>>     >     From: speechsc-bounces@ietf.org
>> >>>     >     [mailto:speechsc-bounces@ietf.org] On Behalf Of 
>> David R  
>> >>> Oran
>> >>>     >     Sent: Tuesday, July 05, 2005 5:24 AM
>> >>>     >     To: Klaus Reifenrath
>> >>>     >     Cc: 'speechsc@ietf.org'
>> >>>     >     Subject: Re: [Speechsc] START-OF-SPEECH in 
>> DTMF-only mode
>> >>>     >
>> >>>     >
>> >>>     >     On Jul 5, 2005, at 3:46 AM, Reifenrath, Klaus wrote:
>> >>>     >
>> >>>     >     > The current spec is not clear when 
>> START-OF-SPEECH need
>> >>>     >     to be send in
>> >>>     >     > the following scenarios:
>> >>>     >     > A) The client requested a DTMF Recognizer. Is the
>> >>>     >     START-OF-SPEECH
>> >>>     >     > event send to the client also if speech was detected?
>> >>>     >     I suspect so, since one of the prime purposes 
>> is to enable
>> >>>     >     client- mediated barge-in handling. However, if the
>> >>>     >     recognizer is in fact only capable of recognizing DTMF
>> >>>     >     then it may in fact not report anythin unless it's using
>> >>>     >     some primitive thresholding machinery, like a SN  
>> >>> threshold.
>> >>>     >     > B) The client requested a Speech Recognizer, but only
>> >>>     >     activated DTMF
>> >>>     >     > grammars. Is the START-OF-SPEECH event send to the
>> >>>     >     client also if
>> >>>     >     > speech was detected?
>> >>>     >     Again, I'd say yes, for the same reason as above.
>> >>>     >     > I think in both cases START-OF-SPEECH should only
>> >>>     be send after
>> >>>     >     > detecting a DTMF digit (see Figure 12 of VoiceXML
>> >>>     2.0: http://
>> >>>     >     > www.w3.org/TR/voicexml20/#dmlATiming).
>> >>>     >     We seem to have reached different conclusions. I'd be
>> >>>     >     interested in why you think my analysis above is wrong.
>> >>>     >
>> >>>     >     Dave.
>> >>>     >
>> >>>     >     > Klaus
>> >>>     >     >
>> >>>     >     > _______________________________________________
>> >>>     >     > Speechsc mailing list
>> >>>     >     > Speechsc@ietf.org
>> >>>     >     > https://www1.ietf.org/mailman/listinfo/speechsc
>> >>>     >     >
>> >>>     >
>> >>>     >     _______________________________________________
>> >>>     >     Speechsc mailing list
>> >>>     >     Speechsc@ietf.org
>> >>>     >     https://www1.ietf.org/mailman/listinfo/speechsc
>> >>>     >
>> >>>     >
>> >>>     > _______________________________________________
>> >>>     > Speechsc mailing list
>> >>>     > Speechsc@ietf.org
>> >>>     > https://www1.ietf.org/mailman/listinfo/speechsc
>> >>>     >
>> >>>
>> >>>
>> >>> _______________________________________________
>> >>> Speechsc mailing list
>> >>> Speechsc@ietf.org
>> >>> https://www1.ietf.org/mailman/listinfo/speechsc
>> >>>
>> >>>
>> >>
>> >> _______________________________________________
>> >> Speechsc mailing list
>> >> Speechsc@ietf.org
>> >> https://www1.ietf.org/mailman/listinfo/speechsc
>> >
>> 
>> _______________________________________________
>> Speechsc mailing list
>> Speechsc@ietf.org
>> https://www1.ietf.org/mailman/listinfo/speechsc
>> 
> 
> 
> _______________________________________________
> Speechsc mailing list
> Speechsc@ietf.org
> https://www1.ietf.org/mailman/listinfo/speechsc
>

_______________________________________________
Speechsc mailing list
Speechsc@ietf.org
https://www1.ietf.org/mailman/listinfo/speechsc



From speechsc-bounces@ietf.org Thu Jul 07 09:22:16 2005
Received: from localhost.localdomain ([127.0.0.1] helo=megatron.ietf.org)
	by megatron.ietf.org with esmtp (Exim 4.32)
	id 1DqWKV-0001mP-I7; Thu, 07 Jul 2005 09:22:16 -0400
Received: from odin.ietf.org ([132.151.1.176] helo=ietf.org)
	by megatron.ietf.org with esmtp (Exim 4.32) id 1DqWKS-0001kD-HV
	for speechsc@megatron.ietf.org; Thu, 07 Jul 2005 09:22:13 -0400
Received: from ietf-mx.ietf.org (ietf-mx [132.151.6.1])
	by ietf.org (8.9.1a/8.9.1a) with ESMTP id JAA05416
	for <speechsc@ietf.org>; Thu, 7 Jul 2005 09:22:10 -0400 (EDT)
Received: from letter.nuance.com ([207.107.210.132])
	by ietf-mx.ietf.org with esmtp (Exim 4.43) id 1DqWlb-0007Mc-Kf
	for speechsc@ietf.org; Thu, 07 Jul 2005 09:50:18 -0400
Received: from postcard.nuance.com ([10.3.6.20]:49291)
	by letter.nuance.com with esmtp id 1DqWKA-0002eG-MD;
	Thu, 07 Jul 2005 06:21:54 -0700
Received: from mtb1exch01.nuance.com ([10.3.2.6]) by postcard.nuance.com with
	Microsoft SMTPSVC(6.0.3790.0); Thu, 7 Jul 2005 09:21:31 -0400
X-MimeOLE: Produced By Microsoft Exchange V6.5.7226.0
Content-class: urn:content-classes:message
MIME-Version: 1.0
Content-Type: text/plain;
	charset="us-ascii"
Content-Transfer-Encoding: quoted-printable
Subject: RE: [Speechsc] START-OF-SPEECH in DTMF-only mode -> summary(?)
Date: Thu, 7 Jul 2005 09:21:48 -0400
Message-ID: <7DE7C4EF3B7C8B4B82955191378290D802ED3F88@mtb1exch01.nuance.com>
Thread-Topic: [Speechsc] START-OF-SPEECH in DTMF-only mode -> summary(?)
Thread-Index: AcWC9PY0LPkvxBeXQQOVCo/nDyd3egAAaF2g
From: "Pierre Forgues" <forgues@nuance.com>
To: "Klaus Reifenrath" <Klaus.Reifenrath@Scansoft.com>,
	"Eric Burger" <eburger@brooktrout.com>, <speechsc@ietf.org>
X-OriginalArrivalTime: 07 Jul 2005 13:21:31.0325 (UTC)
	FILETIME=[CA0B96D0:01C582F6]
X-FromHost: postcard.nuance.com [10.3.6.20]:49291
Lines: 607
X-Spam-Score: 0.0 (/)
X-Scan-Signature: 9668dd9718e12afaf579fddf1143437a
Content-Transfer-Encoding: quoted-printable
Cc: 
X-BeenThere: speechsc@ietf.org
X-Mailman-Version: 2.1.5
Precedence: list
List-Id: Speech Services Control Working Group <speechsc.ietf.org>
List-Unsubscribe: <https://www1.ietf.org/mailman/listinfo/speechsc>,
	<mailto:speechsc-request@ietf.org?subject=unsubscribe>
List-Post: <mailto:speechsc@ietf.org>
List-Help: <mailto:speechsc-request@ietf.org?subject=help>
List-Subscribe: <https://www1.ietf.org/mailman/listinfo/speechsc>,
	<mailto:speechsc-request@ietf.org?subject=subscribe>
Sender: speechsc-bounces@ietf.org
Errors-To: speechsc-bounces@ietf.org

I agree with Klaus's position. =20

Moreover, when the SOS event is generated, it is possible exact nature
of the barge-in method is not known at that time.  It is critical to
notify the user as soon as possible to stop the prompt or this impacts
recognition accuracy.  In some cases, the resource might not even have
exact determination at that ninstant in time whether this is DTMF or
speech.

/Pierre

-----Original Message-----
From: speechsc-bounces@ietf.org [mailto:speechsc-bounces@ietf.org] On
Behalf Of Klaus Reifenrath
Sent: Thursday, July 07, 2005 9:04 AM
To: 'Eric Burger'; speechsc@ietf.org
Subject: RE: [Speechsc] START-OF-SPEECH in DTMF-only mode -> summary(?)

I) START-XYZ for the new event is misleading if we allow the event to be
generated multiple times.
Do we really need to introduce a new (optional) event? Why should a DTMF
only resource do speech detection?=20

II)
  1. START-OF-SPEECH (and optimised bargin) is only generated for the =20
  input type that is being listened for
  2. A speechrecog listens for DTMF if DTMF grammars are active, =20
  speech if speech grammars are active, or speech and DTMF if both =20
  grammar types are active.
I personally like this, BUT doesn't it conflict with VoiceXML 2.0
section
4.1.5.1 (bargein type speech):
"The prompt will be stopped as soon as speech or DTMF input is detected.
The
prompt is stopped irrespective of whether or not the input matches a
grammar
and irrespective of which grammars are active." =20

Klaus

-----Original Message-----
From: Eric Burger [mailto:eburger@brooktrout.com]=20
Sent: Donnerstag, 7. Juli 2005 14:32
To: speechsc@ietf.org
Subject: RE: [Speechsc] START-OF-SPEECH in DTMF-only mode -> summary(?)

Can we declare consensus?=20

> -----Original Message-----
> From: speechsc-bounces@ietf.org [mailto:speechsc-bounces@ietf.org]
> Sent: Wednesday, July 06, 2005 10:45 AM
> To: Dave Burke
> Cc: speechsc@ietf.org; Shanmugham, Saravanan; Klaus Reifenrath
> Subject: Re: [Speechsc] START-OF-SPEECH in DTMF-only mode ->
> summary(?)
>=20
> I think we're getting close. I though about snipping out some pieces=20
> to cut down the text, but I realized the context is still needed. See=20
> inline.
>=20
> On Jul 6, 2005, at 10:29 AM, Dave Burke wrote:
>=20
> > Inline.
> >
> > Dave
> >
> > ----- Original Message ----- From: "David R Oran" <oran@cisco.com>
> > To: "Dave Burke" <david.burke@voxpilot.com>
> > Cc: <speechsc@ietf.org>; "Shanmugham, Saravanan" <sarvi@cisco.com>;=20
> > "Klaus Reifenrath" <Klaus.Reifenrath@Scansoft.com>
> > Sent: Wednesday, July 06, 2005 1:07 PM
> > Subject: Re: [Speechsc] START-OF-SPEECH in DTMF-only mode -> summary
> > (?)
> >
> >
> >
> >>
> >> On Jul 6, 2005, at 5:31 AM, Dave Burke wrote:
> >>
> >>
> >>> + Attempting to summarise:
> >>>
> >>> 1. START-OF-SPEECH is useful for the client to know when
> to stop  =20
> >>> playing prompts the in non-optimised case 2. START-OF-SPEECH is=20
> >>> useful for the client to calculate the bargin  time (e.g. VoiceXML

> >>> 2.1 <mark>)
> >>> 3. In the optimised case, a bargin automatically stops prompt  =20
> >>> playing (assuming prompts barginable) 4. Because of the previous=20
> >>> point, the question of what input type  caused bargin is different

> >>> and less important to
> what input
> >>> type(s)  the recogniser is listening for
> >>>
> >> I'm not sure I follow point 4. Could you elaborate?
> >>
> >
> > DB> Adding a parameter to the START-OF-SPEECH event would
> certainly
> > allow the client to ignore (i.e. let prompts continue playing) the=20
> > event if the event type is not of interest (e.g. the client would=20
> > ignore speech start events when it is interested only in  DTMF start

> > events). This _only_ works for the non-optimised case, however. For=20
> > the optimised case, assuming START-OF-SPEECH
> coincides
> > with the bargin signal to the speechsynth, prompts will stop playing

> > for inputs that the client might not be interested
> in (e.g. =20
> > a speech input will stop prompts playing even if the client
> is only
> > interested in DTMF).
> >
> OK, now I get it. There's a need for the client to both handle the=20
> non-optimized case itself, and influence or at least have a clue what=20
> the server is going to do in the optimized case.
>=20
> >
> >>
> >>
> >>> 5. Currently, MRCPv2 has no way of indicating what input type(s) a

> >>> recogniser is listening for
> >>>
> >>>
> >> Do you mean exactly this, or do you mean "for the client to=20
> >> indicate  to the resource what input types it should look for"?
> >>
> >
> > DB> Yes exactly - apologies for not being clear.
> >
> >
> >>
> >>
> >>> + Why implement 5?
> >>>
> >>> i. Noisy case: Need DTMF-only recognition (and may only have a
> >>> speechrecog)
> >>>
> >> I'm having difficulty following the logic of why noise would  =20
> >> necessarily trigger START-OF-SPEECH if you were listening for=20
> >> speech  but not DTMF. I suppose you can use a more forgiving=20
> >> discriminator if  all you need to tell is if you're getting DTMF,=20
> >> but I've had a number  of real-world cases where wind noise was=20
> >> detected as DTMF, and  there's always the ambiguity when you have=20
> >> Captain Crunch on the  phone. In either case in the non-optimized=20
> >> case it's the client who  gets to decide whether an event should be

> >> interpreted as barge-in or  not, so it seems an aesthetic protocol=20
> >> design decision whether the  client tells the server ahead of time=20
> >> what circumstances to generate  the START-OF-SPEECH event for, or=20
> >> whether the event gets generated  and the client decides based on=20
> >> what's in the event whether it should  be
> treated
> >> as barge- in.
> >>
> >> I suppose one could make the argument that because the spec implies

> >> that the event can only be generated once per request that if a=20
> >> DTMF/ speech capable recognizer first hears
> enough noise
> >> to think it's  hearing speech and later hears DTMF, the client will

> >> declare barge-in  when the event comes and not when he DTMF=20
> >> actually gets heard.
> >>
> >> If that's deemed a problem, we can still handle that in the design

> >> where the server just reports what it's hearing by allowing=20
> >> multiple  events to be generated during a single request.
> >>
> >> Between the approach just outlined above, and an approach where the

> >> client provides a filter for whether to generate the event or not,=20
> >> I  have a mild preference (based on aesthetics rather than some=20
> >> hard  engineering tradeoff) for the approach where
> the server
> >> just reports  what it's hearing.
> >>
> >>
> >> Having had some useful exchanges on this topic, it also is becoming

> >> apparent to me that this event is poorly named, and we should =20
> >> consider renaming it to "INTERESTING-INPUT-HEARD" or something akin

> >> to that, because as others have pointed out, a DTMF-only recognizer

> >> will never detect "start of speech".
> >>
> >> Another consideration to fold into the design choice is  =20
> >> extensibility. Bear with me through a little gedankenexperiment.
> >>
> >> Suppose we want to define a new recognizer type, which I'll call=20
> >> the "name that tune" recognizer. The client plays music to the=20
> >> server and  the server recognizes musical notes. The grammar is a=20
> >> standard  musical notation, augmented with a semantic=20
> >> interpretation that  transforms the notes into the title of the=20
> >> tune and provides that as  an answer.
> >>
> >> First, there's no speech involved (or is there...hang on a minute).

> >> Second, in order to accommodate the "name that tune"
> >> recognizer, we'd have to extend both the client and the server to=20
> >> undetstand a  directive as to whether to recognize music or now,=20
> >> inaddition to what  the server already knows what to do based on=20
> >> the grammar. If you  follow my logic above, whether or not we do=20
> >> that, we have to extend  "start-of-speech" to say "I'm hearing=20
> >> music". So far fairly  straightforward, but let me now throw in the

> >> pathological twist.
> >>
> >> Suppose what I feed to a  combined music/speech recognizer is a=20
> >> work  in sprechstimme (spoken music), like the "Geographical Fugue"

> >> (aside:  this is a wonderful piece of music I highly recommend to=20
> >> anyone  interested in small ensemble singing). In this case, the=20
> >> tune could  be named by either doing speech or music recognition.=20
> >> Why is there  any need for the client to constrain the server as to

> >> which it tries  to do when it's
> already
> >> told the server what it wants through the  grammar?
> >>
> >> A few other comments below
> >>
> >>
> >>> ii. Flexibility: Want speech-only recognition (because a second=20
> >>> recogniser is doing hotword on DTMF)
> >>>
> >>>
> >> I don't see how flexibility is affected by this deisgn
> choice. If  =20
> >> that's what you want, feed the speech-only recognizer a grammar  =20
> >> without any DTMF rules.
> >>
> >>
> >>> + How to implement 5?
> >>>
> >>> a. Implicitly:
> >>>    - dtmfrecog: always DTMF-only recognition
> >>>    - speechrecog: depends on active grammar type
> >>>        > if a dtmf grammar is active then DTMF input is "on"
> >>>        > if a speech grammar is active then speech input is "on"
> >>>
> >>> b. Explicitly:
> >>>    - Add inputmodes header to RECOGNIZE
> >>>
> >>> Option a is David's "do what I mean case"; option b is
> the extra  =20
> >>> dial for the client.
> >>>
> >>>
> >> Actually, that's not the point I was making with "do what I mean",

> >> but it's not essential to the discussion so let's move on.
> >>
> >>
> >>> It is worth noting that VoiceXML uses option b. This allows one to

> >>> activate both speech grammars and DTMF grammars (and
> therefore
> >>> be informed of any errors in the grammars at activation
> time) but
> >>> independently turn on whichever input mode you like e.g.=20
> perhaps  =20
> >>> start with "both" then change to "dtmf".
> >>>
> >>>
> >> I'm not sure the VXML precedent is relevant here, because the  =20
> >> application behind VXML is working a different part of the
> problem
> >> -  how to traverse a TUI dialog based on different parts of the=20
> >> input  space. In fact, I suspect that the VXML: choice was=20
> >> conditioned more  by limitations at the time it was
> specified than
> >> an underlying good  design choice. Clearly having to specify this=20
> >> in VXML make the job of  handling a TUI with nodes like "Say or=20
> >> press 5" harder rather than  easier.
> >>
> >
> > DB> The VoiceXML edge-case is pretty weird so it's not a major
> > concern. My main concern is that the client can indicate, somehow,=20
> > what the input modes are.
> >
> >
> >>
> >> Summing up, while I don't feel strongly one way or the other, I=20
> >> have  a preference for handling this as follows:
> >>
> >> a) Rename "START-OF-SPEECH" to "INTERESTING-INPUT-RECEIVED" or  =20
> >> something equivalent.
> >> b) Include a parameter in the event saying what was interesting=20
> >> about  the input you received, with a registry of values which
> >> includes:
> >>     - signal above noise floor
> >>     - speech
> >>     - dtmf
> >>     - (possibly) music
> >> c) allow the event to be generated multiple times during a request
> >>
> >>
> >
> > DB> I like these suggestions (START-OF-INPUT?). However, I don't
> > see how the problem of the optimised case is not solved by them. I=20
> > think the optimised case is fine if we have the following rules:
> >
> Yes, I hadn't thought through the optimized case as thoroughly as you.

> Your suggested method name is fine by me as well
>=20
> > 1. START-OF-SPEECH (and optimised bargin) is only generated
> for the
> > input type that is being listened for 2. A speechrecog listens for=20
> > DTMF if DTMF grammars are active, speech if speech grammars are=20
> > active, or speech and DTMF if both grammar types are active.
> >
> Works for me.
>=20
> >
> >> Note that all of the above I'm saying with my technical
> hat on and
> >> my chair hat off.
> >> Putting my chair hat on for a moment, we really need to get this=20
> >> spec  to last call, so at some point Eric or I is going to
> declare
> >> rough  consensus so we can move on.
> >>
> >
> > DB> Agreed!
> >
> >
> >>
> >> Dave Oran.
> >>
> >>> Dave
> >>>
> >>>
> >>>
> >>>
> >>>
> >>>
> >>>
> >>>
> >>>
> >>>
> >>>
> >>>
> >>> Sarvi makes a good point that adding the reason why the
> START-OF-
> >>> SPEECH occurred does not fix the optimised bargin case.
> >>>
> >>> dtmfrecog - listens for DTMF only
> >>> speechrecog - listens for DTMF only, or speech only, or speech &=20
> >>> DTMF
> >>>
> >>>
> >>>
> >>>
> >>>
> >>>
> >>>
> >>>
> >>> ----- Original Message ----- From: "Shanmugham, Saravanan" =20
> >>> <sarvi@cisco.com>
> >>> To: "David R Oran" <oran@cisco.com>; "Dave Burke" =20
> >>> <david.burke@voxpilot.com>
> >>> Cc: <speechsc@ietf.org>; "Klaus Reifenrath" =20
> >>> <Klaus.Reifenrath@Scansoft.com>
> >>> Sent: Tuesday, July 05, 2005 8:51 PM
> >>> Subject: RE: [Speechsc] START-OF-SPEECH in DTMF-only mode
> >>>
> >>>
> >>> inline.
> >>>
> >>>     -----Original Message-----
> >>>     From: David R Oran [mailto:oran@cisco.com]
> >>>     Sent: Tuesday, July 05, 2005 11:08 AM
> >>>     To: Dave Burke
> >>>     Cc: Shanmugham, Saravanan; Klaus Reifenrath; speechsc@ietf.org
> >>>     Subject: Re: [Speechsc] START-OF-SPEECH in DTMF-only mode
> >>>
> >>>
> >>>     On Jul 5, 2005, at 12:54 PM, Dave Burke wrote:
> >>>
> >>>     > Inline.
> >>>     >
> >>>     > Dave
> >>>     >
> >>>     > ----- Original Message ----- From: "Shanmugham, Saravanan"
> >>>     > <sarvi@cisco.com>
> >>>     > To: "David R Oran" <oran@cisco.com>; "Klaus Reifenrath"
> >>>     > <Klaus.Reifenrath@Scansoft.com>
> >>>     > Cc: <speechsc@ietf.org>
> >>>     > Sent: Tuesday, July 05, 2005 5:22 PM
> >>>     > Subject: RE: [Speechsc] START-OF-SPEECH in DTMF-only mode
> >>>     >
> >>>     >
> >>>     > I agree with Dave's analysis. The purpose of this event
> >>>     was barge-in.
> >>>     > And barge-in should happen for both DTMF and speech.
> >>>     >
> >>>     > Is there a case where you think it should not behave
> >>>     this way. If soe,
> >>>     > please provide a scenario where you think
> >>>     >    1. Barge-in should happen for DTMF and not voice or
> >>>     vice-versa.
> >>>     >
> >>>     > DB> You want to do a DTMF recognition only because it is=20
> >>> noisy.
> >>>     > While waiting for DTMF input, the speechrecog resource
> >>>     (or advanced
> >>>     > dtmfrecog) generates a START-OF-SPEECH because it heard
> >>>     some speech.
> >>>     > The client does not want to stop prompt playing unless
> >>>     DTMF was heard
> >>>     > but it can't tell by the START-OF-SPEECH whether speech
> >>>     or DTMF was
> >>>     > heard. Similarly vice versa.
> >>>     >
> >>>     It's an interesting design question what part of the
> >>>     policy resides at the client and what at the server, and
> >>>     who makes the "final decision" about whether what was
> >>>     heard was relevant to the control channel. Right now we
> >>>     (IMO) have a weird partitioning in many cases where the
> >>>     client basically says "do what I mean", but there are no
> >>>     constraints of what the server actually does, and no
> >>>     normalized basis for the client to figure out what to set
> >>>     various magic numbers to (e.g. sensitivity).
> >>>
> >>>     In this case the only thing the client needs to decide is
> >>>     whether to kill the prompt because the server thinks
> >>>     something that would interfere with the feedback
> >>>     ear/mouth/finger control happened. What this says to me is
> >>>     that it isn't necessarily a good idea for the client to
> >>>     have more knobs to control the server (especially if those
> >>>     knows are just more value/policy input ungrounded in any
> >>>     physics/ acoustics). On the other hand, having the server
> >>>     tell the client more about what it thinks is going on is
> >>>     probably valuable.
> >>>
> >>>     So, Coming to the point after this long rambling
> >>>     introduction, I think it would in fact be useful for the
> >>>     START-Of-SPEECH event to indicate some extra information,
> >>>     for example:
> >>>     a) I got something enough above the noise floor to qualify
> >>>     for exceeding the "Sensisitvity" parameter you sent in on
> >>>     the request but I really can't tell what it is (could be a
> >>>     hippopatmus fart, or a siren in the background, or captain
> >>>     crunch trying to whistle DTMF).
> >>>     b) I think I'm hearing speech
> >>>     c) I think I'm hearing DTMF
> >>>
> >>> Though I agree with your former part of your response. I am not=20
> >>> sure I agree with your proposed solution.
> >>> The way I see this problem is that, it is more of what
> constitues a
> >>> barge-in event. This boils down to whether it is speech, DTMF or=20
> >>> both.
> >>> This is inturn boils down to what type of recognizer
> resource we are
> >>> using, dtmf-recog, speech-recog, and speech-only-recog(we
> don't have
> >>> this and I don't think we should add it, but think of this as a=20
> >>> place holder that explains the concept).
> >>>
> >>> A client knowing what type of barge-in happenned, does
> not impact
> >>> the
> >>> barge-in operation itself as it may be too late(for the optimized=20
> >>> barge-in case). It may have other use cases, and if we
> can identify
> >>> them, I don't mind adding support for the START-OF-SPEECH event to

> >>> say what type of barge-in happenned. But that itself does not
> solve the
> >>> original problem raised. Refer to my previous response.
> >>>
> >>> The solution lies in defining what what is a barge-in event. =20
> >>> That  boils
> >>> down to what type of recognition is happenning,
> dtmf-only, speech-
> >>> dtmf
> >>> or speech-only. We do not support speech-only as a resource today,

> >>> the question is do we need a header to force it.
> >>>
> >>> Sarvi
> >>>
> >>>
> >>>     >    2. You would benefit from the client knowing=20
> what caused
> >>> the
> >>>     > barge-in, DTMF Vs speech.
> >>>     >
> >>>     > DB> See previous comment. And previous e-mail: either add an
> >>>     > inputmodes header (taking value speech, dtmf, both) to
> >>>     the RECOGNIZE
> >>>     > request or add a header to the START-OF-SPEECH event
> >>>     indicating DTMF
> >>>     > or speech.
> >>>     >
> >>>     I'm leaning in your direction on this latter point - as
> >>>     should be evident from what I wrote above.
> >>>
> >>>     > Sarvi
> >>>     >
> >>>     >     -----Original Message-----
> >>>     >     From: speechsc-bounces@ietf.org
> >>>     >     [mailto:speechsc-bounces@ietf.org] On Behalf Of=20
> David R
> >>> Oran
> >>>     >     Sent: Tuesday, July 05, 2005 5:24 AM
> >>>     >     To: Klaus Reifenrath
> >>>     >     Cc: 'speechsc@ietf.org'
> >>>     >     Subject: Re: [Speechsc] START-OF-SPEECH in=20
> DTMF-only mode
> >>>     >
> >>>     >
> >>>     >     On Jul 5, 2005, at 3:46 AM, Reifenrath, Klaus wrote:
> >>>     >
> >>>     >     > The current spec is not clear when=20
> START-OF-SPEECH need
> >>>     >     to be send in
> >>>     >     > the following scenarios:
> >>>     >     > A) The client requested a DTMF Recognizer. Is the
> >>>     >     START-OF-SPEECH
> >>>     >     > event send to the client also if speech was detected?
> >>>     >     I suspect so, since one of the prime purposes=20
> is to enable
> >>>     >     client- mediated barge-in handling. However, if the
> >>>     >     recognizer is in fact only capable of recognizing DTMF
> >>>     >     then it may in fact not report anythin unless it's using
> >>>     >     some primitive thresholding machinery, like a SN =20
> >>> threshold.
> >>>     >     > B) The client requested a Speech Recognizer, but only
> >>>     >     activated DTMF
> >>>     >     > grammars. Is the START-OF-SPEECH event send to the
> >>>     >     client also if
> >>>     >     > speech was detected?
> >>>     >     Again, I'd say yes, for the same reason as above.
> >>>     >     > I think in both cases START-OF-SPEECH should only
> >>>     be send after
> >>>     >     > detecting a DTMF digit (see Figure 12 of VoiceXML
> >>>     2.0: http://
> >>>     >     > www.w3.org/TR/voicexml20/#dmlATiming).
> >>>     >     We seem to have reached different conclusions. I'd be
> >>>     >     interested in why you think my analysis above is wrong.
> >>>     >
> >>>     >     Dave.
> >>>     >
> >>>     >     > Klaus
> >>>     >     >
> >>>     >     > _______________________________________________
> >>>     >     > Speechsc mailing list
> >>>     >     > Speechsc@ietf.org
> >>>     >     > https://www1.ietf.org/mailman/listinfo/speechsc
> >>>     >     >
> >>>     >
> >>>     >     _______________________________________________
> >>>     >     Speechsc mailing list
> >>>     >     Speechsc@ietf.org
> >>>     >     https://www1.ietf.org/mailman/listinfo/speechsc
> >>>     >
> >>>     >
> >>>     > _______________________________________________
> >>>     > Speechsc mailing list
> >>>     > Speechsc@ietf.org
> >>>     > https://www1.ietf.org/mailman/listinfo/speechsc
> >>>     >
> >>>
> >>>
> >>> _______________________________________________
> >>> Speechsc mailing list
> >>> Speechsc@ietf.org
> >>> https://www1.ietf.org/mailman/listinfo/speechsc
> >>>
> >>>
> >>
> >> _______________________________________________
> >> Speechsc mailing list
> >> Speechsc@ietf.org
> >> https://www1.ietf.org/mailman/listinfo/speechsc
> >
>=20
> _______________________________________________
> Speechsc mailing list
> Speechsc@ietf.org
> https://www1.ietf.org/mailman/listinfo/speechsc
>=20


_______________________________________________
Speechsc mailing list
Speechsc@ietf.org
https://www1.ietf.org/mailman/listinfo/speechsc

_______________________________________________
Speechsc mailing list
Speechsc@ietf.org
https://www1.ietf.org/mailman/listinfo/speechsc

=20
 =20


_______________________________________________
Speechsc mailing list
Speechsc@ietf.org
https://www1.ietf.org/mailman/listinfo/speechsc



From speechsc-bounces@ietf.org Thu Jul 07 12:52:27 2005
Received: from localhost.localdomain ([127.0.0.1] helo=megatron.ietf.org)
	by megatron.ietf.org with esmtp (Exim 4.32)
	id 1DqZbu-0000Qi-Sr; Thu, 07 Jul 2005 12:52:26 -0400
Received: from odin.ietf.org ([132.151.1.176] helo=ietf.org)
	by megatron.ietf.org with esmtp (Exim 4.32) id 1DqZbs-0000Pj-2d
	for speechsc@megatron.ietf.org; Thu, 07 Jul 2005 12:52:24 -0400
Received: from ietf-mx.ietf.org (ietf-mx [132.151.6.1])
	by ietf.org (8.9.1a/8.9.1a) with ESMTP id MAA25112
	for <speechsc@ietf.org>; Thu, 7 Jul 2005 12:52:20 -0400 (EDT)
Received: from sj-iport-5.cisco.com ([171.68.10.87])
	by ietf-mx.ietf.org with esmtp (Exim 4.43) id 1Dqa33-0008Eb-44
	for speechsc@ietf.org; Thu, 07 Jul 2005 13:20:32 -0400
Received: from sj-core-3.cisco.com (171.68.223.137)
	by sj-iport-5.cisco.com with ESMTP; 07 Jul 2005 09:52:12 -0700
X-IronPort-AV: i="3.93,270,1115017200"; 
	d="scan'208"; a="196919891:sNHT34476908"
Received: from vtg-um-e2k6.sj21ad.cisco.com (vtg-um-e2k6.cisco.com
	[171.70.93.77])
	by sj-core-3.cisco.com (8.12.10/8.12.6) with ESMTP id j67Gq96p006458;
	Thu, 7 Jul 2005 09:52:09 -0700 (PDT)
Content-class: urn:content-classes:message
MIME-Version: 1.0
Content-Type: text/plain;
	charset="us-ascii"
Content-Transfer-Encoding: quoted-printable
X-MimeOLE: Produced By Microsoft Exchange V6.5.7226.0
Subject: RE: [Speechsc] START-OF-SPEECH in DTMF-only mode -> summary(?)
Date: Thu, 7 Jul 2005 09:52:08 -0700
Message-ID: <03772D1EC8DE624A863058C75874A75C1128E5@vtg-um-e2k6.sj21ad.cisco.com>
Thread-Topic: [Speechsc] START-OF-SPEECH in DTMF-only mode -> summary(?)
Thread-Index: AcWC8i/wtcNwTItNQBekCl5mGebKswAIMqJA
From: "Shanmugham, Saravanan" <sarvi@cisco.com>
To: "Eric Burger" <eburger@brooktrout.com>, <speechsc@ietf.org>
X-Spam-Score: 0.0 (/)
X-Scan-Signature: b4c10eaa27436d806c79842272125a2a
Content-Transfer-Encoding: quoted-printable
Cc: 
X-BeenThere: speechsc@ietf.org
X-Mailman-Version: 2.1.5
Precedence: list
List-Id: Speech Services Control Working Group <speechsc.ietf.org>
List-Unsubscribe: <https://www1.ietf.org/mailman/listinfo/speechsc>,
	<mailto:speechsc-request@ietf.org?subject=unsubscribe>
List-Post: <mailto:speechsc@ietf.org>
List-Help: <mailto:speechsc-request@ietf.org?subject=help>
List-Subscribe: <https://www1.ietf.org/mailman/listinfo/speechsc>,
	<mailto:speechsc-request@ietf.org?subject=subscribe>
Sender: speechsc-bounces@ietf.org
Errors-To: speechsc-bounces@ietf.org

I don't think generating multiple START-OF-SPEECH events is a solution.=20
We still haven't addressed, what constitutes a barge-in, for the
optimised case. That should also be the single point when a single
START-OF-SPEECH(or whathever else you want to name it) should be
generated.

That point, in my opinion should be
   1. For "dtmf-recog" resources should be the beginning of a DTMF key
press.
   2. For "speech-recog" resources should be the beginning of a DTMF key
press or the beginning of speech. This is should be irrespective of what
type of grammar is being used. Coz even numbers only grammars can still
be spoken and hence cannot be assumed to be a cue for DTMF only
recognition.
   3. For "speech-only-recog" resources(which are not defined today) the
time to barge-in is the beginning of speech. I don't see a need for such
a resource today. But I am mentioning this for completeness.

Sarvi
=20

     -----Original Message-----
     From: speechsc-bounces@ietf.org=20
     [mailto:speechsc-bounces@ietf.org] On Behalf Of Eric Burger
     Sent: Thursday, July 07, 2005 5:32 AM
     To: speechsc@ietf.org
     Subject: RE: [Speechsc] START-OF-SPEECH in DTMF-only mode=20
     -> summary(?)
    =20
     Can we declare consensus?=20
    =20
     > -----Original Message-----
     > From: speechsc-bounces@ietf.org=20
     [mailto:speechsc-bounces@ietf.org]
     > Sent: Wednesday, July 06, 2005 10:45 AM
     > To: Dave Burke
     > Cc: speechsc@ietf.org; Shanmugham, Saravanan; Klaus Reifenrath
     > Subject: Re: [Speechsc] START-OF-SPEECH in DTMF-only mode ->
     > summary(?)
     >=20
     > I think we're getting close. I though about snipping out=20
     some pieces=20
     > to cut down the text, but I realized the context is=20
     still needed. See=20
     > inline.
     >=20
     > On Jul 6, 2005, at 10:29 AM, Dave Burke wrote:
     >=20
     > > Inline.
     > >
     > > Dave
     > >
     > > ----- Original Message ----- From: "David R Oran"=20
     <oran@cisco.com>
     > > To: "Dave Burke" <david.burke@voxpilot.com>
     > > Cc: <speechsc@ietf.org>; "Shanmugham, Saravanan"=20
     <sarvi@cisco.com>;=20
     > > "Klaus Reifenrath" <Klaus.Reifenrath@Scansoft.com>
     > > Sent: Wednesday, July 06, 2005 1:07 PM
     > > Subject: Re: [Speechsc] START-OF-SPEECH in DTMF-only=20
     mode -> summary
     > > (?)
     > >
     > >
     > >
     > >>
     > >> On Jul 6, 2005, at 5:31 AM, Dave Burke wrote:
     > >>
     > >>
     > >>> + Attempting to summarise:
     > >>>
     > >>> 1. START-OF-SPEECH is useful for the client to know when
     > to stop  =20
     > >>> playing prompts the in non-optimised case 2.=20
     START-OF-SPEECH is=20
     > >>> useful for the client to calculate the bargin  time=20
     (e.g. VoiceXML=20
     > >>> 2.1 <mark>)
     > >>> 3. In the optimised case, a bargin automatically=20
     stops prompt  =20
     > >>> playing (assuming prompts barginable) 4. Because of=20
     the previous=20
     > >>> point, the question of what input type  caused=20
     bargin is different=20
     > >>> and less important to
     > what input
     > >>> type(s)  the recogniser is listening for
     > >>>
     > >> I'm not sure I follow point 4. Could you elaborate?
     > >>
     > >
     > > DB> Adding a parameter to the START-OF-SPEECH event would
     > certainly
     > > allow the client to ignore (i.e. let prompts continue=20
     playing) the=20
     > > event if the event type is not of interest (e.g. the=20
     client would=20
     > > ignore speech start events when it is interested only=20
     in  DTMF start=20
     > > events). This _only_ works for the non-optimised case,=20
     however. For=20
     > > the optimised case, assuming START-OF-SPEECH
     > coincides
     > > with the bargin signal to the speechsynth, prompts=20
     will stop playing=20
     > > for inputs that the client might not be interested
     > in (e.g. =20
     > > a speech input will stop prompts playing even if the client
     > is only
     > > interested in DTMF).
     > >
     > OK, now I get it. There's a need for the client to both=20
     handle the=20
     > non-optimized case itself, and influence or at least=20
     have a clue what=20
     > the server is going to do in the optimized case.
     >=20
     > >
     > >>
     > >>
     > >>> 5. Currently, MRCPv2 has no way of indicating what=20
     input type(s) a=20
     > >>> recogniser is listening for
     > >>>
     > >>>
     > >> Do you mean exactly this, or do you mean "for the client to=20
     > >> indicate  to the resource what input types it should=20
     look for"?
     > >>
     > >
     > > DB> Yes exactly - apologies for not being clear.
     > >
     > >
     > >>
     > >>
     > >>> + Why implement 5?
     > >>>
     > >>> i. Noisy case: Need DTMF-only recognition (and may=20
     only have a
     > >>> speechrecog)
     > >>>
     > >> I'm having difficulty following the logic of why=20
     noise would  =20
     > >> necessarily trigger START-OF-SPEECH if you were listening for=20
     > >> speech  but not DTMF. I suppose you can use a more forgiving=20
     > >> discriminator if  all you need to tell is if you're=20
     getting DTMF,=20
     > >> but I've had a number  of real-world cases where wind=20
     noise was=20
     > >> detected as DTMF, and  there's always the ambiguity=20
     when you have=20
     > >> Captain Crunch on the  phone. In either case in the=20
     non-optimized=20
     > >> case it's the client who  gets to decide whether an=20
     event should be=20
     > >> interpreted as barge-in or  not, so it seems an=20
     aesthetic protocol=20
     > >> design decision whether the  client tells the server=20
     ahead of time=20
     > >> what circumstances to generate  the START-OF-SPEECH=20
     event for, or=20
     > >> whether the event gets generated  and the client=20
     decides based on=20
     > >> what's in the event whether it should  be
     > treated
     > >> as barge- in.
     > >>
     > >> I suppose one could make the argument that because=20
     the spec implies =20
     > >> that the event can only be generated once per request=20
     that if a=20
     > >> DTMF/ speech capable recognizer first hears
     > enough noise
     > >> to think it's  hearing speech and later hears DTMF,=20
     the client will=20
     > >> declare barge-in  when the event comes and not when he DTMF=20
     > >> actually gets heard.
     > >>
     > >> If that's deemed a problem, we can still handle that=20
     in the design =20
     > >> where the server just reports what it's hearing by allowing=20
     > >> multiple  events to be generated during a single request.
     > >>
     > >> Between the approach just outlined above, and an=20
     approach where the=20
     > >> client provides a filter for whether to generate the=20
     event or not,=20
     > >> I  have a mild preference (based on aesthetics rather=20
     than some=20
     > >> hard  engineering tradeoff) for the approach where
     > the server
     > >> just reports  what it's hearing.
     > >>
     > >>
     > >> Having had some useful exchanges on this topic, it=20
     also is becoming=20
     > >> apparent to me that this event is poorly named, and=20
     we should =20
     > >> consider renaming it to "INTERESTING-INPUT-HEARD" or=20
     something akin =20
     > >> to that, because as others have pointed out, a=20
     DTMF-only recognizer =20
     > >> will never detect "start of speech".
     > >>
     > >> Another consideration to fold into the design choice is  =20
     > >> extensibility. Bear with me through a little=20
     gedankenexperiment.
     > >>
     > >> Suppose we want to define a new recognizer type,=20
     which I'll call=20
     > >> the "name that tune" recognizer. The client plays=20
     music to the=20
     > >> server and  the server recognizes musical notes. The=20
     grammar is a=20
     > >> standard  musical notation, augmented with a semantic=20
     > >> interpretation that  transforms the notes into the=20
     title of the=20
     > >> tune and provides that as  an answer.
     > >>
     > >> First, there's no speech involved (or is there...hang=20
     on a minute).=20
     > >> Second, in order to accommodate the "name that tune"
     > >> recognizer, we'd have to extend both the client and=20
     the server to=20
     > >> undetstand a  directive as to whether to recognize=20
     music or now,=20
     > >> inaddition to what  the server already knows what to=20
     do based on=20
     > >> the grammar. If you  follow my logic above, whether=20
     or not we do=20
     > >> that, we have to extend  "start-of-speech" to say=20
     "I'm hearing=20
     > >> music". So far fairly  straightforward, but let me=20
     now throw in the=20
     > >> pathological twist.
     > >>
     > >> Suppose what I feed to a  combined music/speech=20
     recognizer is a=20
     > >> work  in sprechstimme (spoken music), like the=20
     "Geographical Fugue"=20
     > >> (aside:  this is a wonderful piece of music I highly=20
     recommend to=20
     > >> anyone  interested in small ensemble singing). In=20
     this case, the=20
     > >> tune could  be named by either doing speech or music=20
     recognition.=20
     > >> Why is there  any need for the client to constrain=20
     the server as to=20
     > >> which it tries  to do when it's
     > already
     > >> told the server what it wants through the  grammar?
     > >>
     > >> A few other comments below
     > >>
     > >>
     > >>> ii. Flexibility: Want speech-only recognition=20
     (because a second=20
     > >>> recogniser is doing hotword on DTMF)
     > >>>
     > >>>
     > >> I don't see how flexibility is affected by this deisgn
     > choice. If  =20
     > >> that's what you want, feed the speech-only recognizer=20
     a grammar  =20
     > >> without any DTMF rules.
     > >>
     > >>
     > >>> + How to implement 5?
     > >>>
     > >>> a. Implicitly:
     > >>>    - dtmfrecog: always DTMF-only recognition
     > >>>    - speechrecog: depends on active grammar type
     > >>>        > if a dtmf grammar is active then DTMF input is "on"
     > >>>        > if a speech grammar is active then speech=20
     input is "on"
     > >>>
     > >>> b. Explicitly:
     > >>>    - Add inputmodes header to RECOGNIZE
     > >>>
     > >>> Option a is David's "do what I mean case"; option b is
     > the extra  =20
     > >>> dial for the client.
     > >>>
     > >>>
     > >> Actually, that's not the point I was making with "do=20
     what I mean", =20
     > >> but it's not essential to the discussion so let's move on.
     > >>
     > >>
     > >>> It is worth noting that VoiceXML uses option b. This=20
     allows one to=20
     > >>> activate both speech grammars and DTMF grammars (and
     > therefore
     > >>> be informed of any errors in the grammars at activation
     > time) but
     > >>> independently turn on whichever input mode you like e.g.=20
     > perhaps  =20
     > >>> start with "both" then change to "dtmf".
     > >>>
     > >>>
     > >> I'm not sure the VXML precedent is relevant here,=20
     because the  =20
     > >> application behind VXML is working a different part of the
     > problem
     > >> -  how to traverse a TUI dialog based on different=20
     parts of the=20
     > >> input  space. In fact, I suspect that the VXML: choice was=20
     > >> conditioned more  by limitations at the time it was
     > specified than
     > >> an underlying good  design choice. Clearly having to=20
     specify this=20
     > >> in VXML make the job of  handling a TUI with nodes=20
     like "Say or=20
     > >> press 5" harder rather than  easier.
     > >>
     > >
     > > DB> The VoiceXML edge-case is pretty weird so it's not a major
     > > concern. My main concern is that the client can=20
     indicate, somehow,=20
     > > what the input modes are.
     > >
     > >
     > >>
     > >> Summing up, while I don't feel strongly one way or=20
     the other, I=20
     > >> have  a preference for handling this as follows:
     > >>
     > >> a) Rename "START-OF-SPEECH" to=20
     "INTERESTING-INPUT-RECEIVED" or  =20
     > >> something equivalent.
     > >> b) Include a parameter in the event saying what was=20
     interesting=20
     > >> about  the input you received, with a registry of values which
     > >> includes:
     > >>     - signal above noise floor
     > >>     - speech
     > >>     - dtmf
     > >>     - (possibly) music
     > >> c) allow the event to be generated multiple times=20
     during a request
     > >>
     > >>
     > >
     > > DB> I like these suggestions (START-OF-INPUT?).=20
     However, I don't
     > > see how the problem of the optimised case is not=20
     solved by them. I=20
     > > think the optimised case is fine if we have the=20
     following rules:
     > >
     > Yes, I hadn't thought through the optimized case as=20
     thoroughly as you.=20
     > Your suggested method name is fine by me as well
     >=20
     > > 1. START-OF-SPEECH (and optimised bargin) is only generated
     > for the
     > > input type that is being listened for 2. A speechrecog=20
     listens for=20
     > > DTMF if DTMF grammars are active, speech if speech=20
     grammars are=20
     > > active, or speech and DTMF if both grammar types are active.
     > >
     > Works for me.
     >=20
     > >
     > >> Note that all of the above I'm saying with my technical
     > hat on and
     > >> my chair hat off.
     > >> Putting my chair hat on for a moment, we really need=20
     to get this=20
     > >> spec  to last call, so at some point Eric or I is going to
     > declare
     > >> rough  consensus so we can move on.
     > >>
     > >
     > > DB> Agreed!
     > >
     > >
     > >>
     > >> Dave Oran.
     > >>
     > >>> Dave
     > >>>
     > >>>
     > >>>
     > >>>
     > >>>
     > >>>
     > >>>
     > >>>
     > >>>
     > >>>
     > >>>
     > >>>
     > >>> Sarvi makes a good point that adding the reason why the
     > START-OF-
     > >>> SPEECH occurred does not fix the optimised bargin case.
     > >>>
     > >>> dtmfrecog - listens for DTMF only
     > >>> speechrecog - listens for DTMF only, or speech only,=20
     or speech &=20
     > >>> DTMF
     > >>>
     > >>>
     > >>>
     > >>>
     > >>>
     > >>>
     > >>>
     > >>>
     > >>> ----- Original Message ----- From: "Shanmugham, Saravanan" =20
     > >>> <sarvi@cisco.com>
     > >>> To: "David R Oran" <oran@cisco.com>; "Dave Burke" =20
     > >>> <david.burke@voxpilot.com>
     > >>> Cc: <speechsc@ietf.org>; "Klaus Reifenrath" =20
     > >>> <Klaus.Reifenrath@Scansoft.com>
     > >>> Sent: Tuesday, July 05, 2005 8:51 PM
     > >>> Subject: RE: [Speechsc] START-OF-SPEECH in DTMF-only mode
     > >>>
     > >>>
     > >>> inline.
     > >>>
     > >>>     -----Original Message-----
     > >>>     From: David R Oran [mailto:oran@cisco.com]
     > >>>     Sent: Tuesday, July 05, 2005 11:08 AM
     > >>>     To: Dave Burke
     > >>>     Cc: Shanmugham, Saravanan; Klaus Reifenrath;=20
     speechsc@ietf.org
     > >>>     Subject: Re: [Speechsc] START-OF-SPEECH in DTMF-only mode
     > >>>
     > >>>
     > >>>     On Jul 5, 2005, at 12:54 PM, Dave Burke wrote:
     > >>>
     > >>>     > Inline.
     > >>>     >
     > >>>     > Dave
     > >>>     >
     > >>>     > ----- Original Message ----- From:=20
     "Shanmugham, Saravanan"
     > >>>     > <sarvi@cisco.com>
     > >>>     > To: "David R Oran" <oran@cisco.com>; "Klaus Reifenrath"
     > >>>     > <Klaus.Reifenrath@Scansoft.com>
     > >>>     > Cc: <speechsc@ietf.org>
     > >>>     > Sent: Tuesday, July 05, 2005 5:22 PM
     > >>>     > Subject: RE: [Speechsc] START-OF-SPEECH in=20
     DTMF-only mode
     > >>>     >
     > >>>     >
     > >>>     > I agree with Dave's analysis. The purpose of this event
     > >>>     was barge-in.
     > >>>     > And barge-in should happen for both DTMF and speech.
     > >>>     >
     > >>>     > Is there a case where you think it should not behave
     > >>>     this way. If soe,
     > >>>     > please provide a scenario where you think
     > >>>     >    1. Barge-in should happen for DTMF and not voice or
     > >>>     vice-versa.
     > >>>     >
     > >>>     > DB> You want to do a DTMF recognition only=20
     because it is=20
     > >>> noisy.
     > >>>     > While waiting for DTMF input, the speechrecog resource
     > >>>     (or advanced
     > >>>     > dtmfrecog) generates a START-OF-SPEECH because it heard
     > >>>     some speech.
     > >>>     > The client does not want to stop prompt playing unless
     > >>>     DTMF was heard
     > >>>     > but it can't tell by the START-OF-SPEECH whether speech
     > >>>     or DTMF was
     > >>>     > heard. Similarly vice versa.
     > >>>     >
     > >>>     It's an interesting design question what part of the
     > >>>     policy resides at the client and what at the server, and
     > >>>     who makes the "final decision" about whether what was
     > >>>     heard was relevant to the control channel. Right now we
     > >>>     (IMO) have a weird partitioning in many cases where the
     > >>>     client basically says "do what I mean", but there are no
     > >>>     constraints of what the server actually does, and no
     > >>>     normalized basis for the client to figure out what to set
     > >>>     various magic numbers to (e.g. sensitivity).
     > >>>
     > >>>     In this case the only thing the client needs to decide is
     > >>>     whether to kill the prompt because the server thinks
     > >>>     something that would interfere with the feedback
     > >>>     ear/mouth/finger control happened. What this=20
     says to me is
     > >>>     that it isn't necessarily a good idea for the client to
     > >>>     have more knobs to control the server=20
     (especially if those
     > >>>     knows are just more value/policy input ungrounded in any
     > >>>     physics/ acoustics). On the other hand, having the server
     > >>>     tell the client more about what it thinks is going on is
     > >>>     probably valuable.
     > >>>
     > >>>     So, Coming to the point after this long rambling
     > >>>     introduction, I think it would in fact be useful for the
     > >>>     START-Of-SPEECH event to indicate some extra information,
     > >>>     for example:
     > >>>     a) I got something enough above the noise floor=20
     to qualify
     > >>>     for exceeding the "Sensisitvity" parameter you sent in on
     > >>>     the request but I really can't tell what it is=20
     (could be a
     > >>>     hippopatmus fart, or a siren in the background,=20
     or captain
     > >>>     crunch trying to whistle DTMF).
     > >>>     b) I think I'm hearing speech
     > >>>     c) I think I'm hearing DTMF
     > >>>
     > >>> Though I agree with your former part of your=20
     response. I am not=20
     > >>> sure I agree with your proposed solution.
     > >>> The way I see this problem is that, it is more of what
     > constitues a
     > >>> barge-in event. This boils down to whether it is=20
     speech, DTMF or=20
     > >>> both.
     > >>> This is inturn boils down to what type of recognizer
     > resource we are
     > >>> using, dtmf-recog, speech-recog, and speech-only-recog(we
     > don't have
     > >>> this and I don't think we should add it, but think=20
     of this as a=20
     > >>> place holder that explains the concept).
     > >>>
     > >>> A client knowing what type of barge-in happenned, does
     > not impact
     > >>> the
     > >>> barge-in operation itself as it may be too late(for=20
     the optimized=20
     > >>> barge-in case). It may have other use cases, and if we
     > can identify
     > >>> them, I don't mind adding support for the=20
     START-OF-SPEECH event to=20
     > >>> say what type of barge-in happenned. But that itself does not
     > solve the
     > >>> original problem raised. Refer to my previous response.
     > >>>
     > >>> The solution lies in defining what what is a=20
     barge-in event. =20
     > >>> That  boils
     > >>> down to what type of recognition is happenning,
     > dtmf-only, speech-
     > >>> dtmf
     > >>> or speech-only. We do not support speech-only as a=20
     resource today,=20
     > >>> the question is do we need a header to force it.
     > >>>
     > >>> Sarvi
     > >>>
     > >>>
     > >>>     >    2. You would benefit from the client knowing=20
     > what caused
     > >>> the
     > >>>     > barge-in, DTMF Vs speech.
     > >>>     >
     > >>>     > DB> See previous comment. And previous e-mail:=20
     either add an
     > >>>     > inputmodes header (taking value speech, dtmf, both) to
     > >>>     the RECOGNIZE
     > >>>     > request or add a header to the START-OF-SPEECH event
     > >>>     indicating DTMF
     > >>>     > or speech.
     > >>>     >
     > >>>     I'm leaning in your direction on this latter point - as
     > >>>     should be evident from what I wrote above.
     > >>>
     > >>>     > Sarvi
     > >>>     >
     > >>>     >     -----Original Message-----
     > >>>     >     From: speechsc-bounces@ietf.org
     > >>>     >     [mailto:speechsc-bounces@ietf.org] On Behalf Of=20
     > David R
     > >>> Oran
     > >>>     >     Sent: Tuesday, July 05, 2005 5:24 AM
     > >>>     >     To: Klaus Reifenrath
     > >>>     >     Cc: 'speechsc@ietf.org'
     > >>>     >     Subject: Re: [Speechsc] START-OF-SPEECH in=20
     > DTMF-only mode
     > >>>     >
     > >>>     >
     > >>>     >     On Jul 5, 2005, at 3:46 AM, Reifenrath,=20
     Klaus wrote:
     > >>>     >
     > >>>     >     > The current spec is not clear when=20
     > START-OF-SPEECH need
     > >>>     >     to be send in
     > >>>     >     > the following scenarios:
     > >>>     >     > A) The client requested a DTMF Recognizer. Is the
     > >>>     >     START-OF-SPEECH
     > >>>     >     > event send to the client also if speech=20
     was detected?
     > >>>     >     I suspect so, since one of the prime purposes=20
     > is to enable
     > >>>     >     client- mediated barge-in handling. However, if the
     > >>>     >     recognizer is in fact only capable of=20
     recognizing DTMF
     > >>>     >     then it may in fact not report anythin=20
     unless it's using
     > >>>     >     some primitive thresholding machinery, like a SN =20
     > >>> threshold.
     > >>>     >     > B) The client requested a Speech=20
     Recognizer, but only
     > >>>     >     activated DTMF
     > >>>     >     > grammars. Is the START-OF-SPEECH event=20
     send to the
     > >>>     >     client also if
     > >>>     >     > speech was detected?
     > >>>     >     Again, I'd say yes, for the same reason as above.
     > >>>     >     > I think in both cases START-OF-SPEECH should only
     > >>>     be send after
     > >>>     >     > detecting a DTMF digit (see Figure 12 of VoiceXML
     > >>>     2.0: http://
     > >>>     >     > www.w3.org/TR/voicexml20/#dmlATiming).
     > >>>     >     We seem to have reached different=20
     conclusions. I'd be
     > >>>     >     interested in why you think my analysis=20
     above is wrong.
     > >>>     >
     > >>>     >     Dave.
     > >>>     >
     > >>>     >     > Klaus
     > >>>     >     >
     > >>>     >     > _______________________________________________
     > >>>     >     > Speechsc mailing list
     > >>>     >     > Speechsc@ietf.org
     > >>>     >     > https://www1.ietf.org/mailman/listinfo/speechsc
     > >>>     >     >
     > >>>     >
     > >>>     >     _______________________________________________
     > >>>     >     Speechsc mailing list
     > >>>     >     Speechsc@ietf.org
     > >>>     >     https://www1.ietf.org/mailman/listinfo/speechsc
     > >>>     >
     > >>>     >
     > >>>     > _______________________________________________
     > >>>     > Speechsc mailing list
     > >>>     > Speechsc@ietf.org
     > >>>     > https://www1.ietf.org/mailman/listinfo/speechsc
     > >>>     >
     > >>>
     > >>>
     > >>> _______________________________________________
     > >>> Speechsc mailing list
     > >>> Speechsc@ietf.org
     > >>> https://www1.ietf.org/mailman/listinfo/speechsc
     > >>>
     > >>>
     > >>
     > >> _______________________________________________
     > >> Speechsc mailing list
     > >> Speechsc@ietf.org
     > >> https://www1.ietf.org/mailman/listinfo/speechsc
     > >
     >=20
     > _______________________________________________
     > Speechsc mailing list
     > Speechsc@ietf.org
     > https://www1.ietf.org/mailman/listinfo/speechsc
     >=20
    =20
    =20
     _______________________________________________
     Speechsc mailing list
     Speechsc@ietf.org
     https://www1.ietf.org/mailman/listinfo/speechsc
    =20

_______________________________________________
Speechsc mailing list
Speechsc@ietf.org
https://www1.ietf.org/mailman/listinfo/speechsc



From speechsc-bounces@ietf.org Fri Jul 08 07:00:20 2005
Received: from localhost.localdomain ([127.0.0.1] helo=megatron.ietf.org)
	by megatron.ietf.org with esmtp (Exim 4.32)
	id 1Dqqah-0001JS-Te; Fri, 08 Jul 2005 07:00:19 -0400
Received: from odin.ietf.org ([132.151.1.176] helo=ietf.org)
	by megatron.ietf.org with esmtp (Exim 4.32) id 1Dqqaf-00017K-9i
	for speechsc@megatron.ietf.org; Fri, 08 Jul 2005 07:00:17 -0400
Received: from ietf-mx.ietf.org (ietf-mx [132.151.6.1])
	by ietf.org (8.9.1a/8.9.1a) with ESMTP id HAA20194
	for <speechsc@ietf.org>; Fri, 8 Jul 2005 07:00:14 -0400 (EDT)
Received: from fw01.db01.voxpilot.com ([212.17.54.82] helo=mail.voxpilot.com)
	by ietf-mx.ietf.org with esmtp (Exim 4.43) id 1Dqr1s-0003ql-B8
	for speechsc@ietf.org; Fri, 08 Jul 2005 07:28:34 -0400
Received: from daburkewxp (unknown [213.233.148.254])
	by mail.voxpilot.com (Postfix) with ESMTP
	id A4FF6214041; Fri,  8 Jul 2005 10:59:48 +0000 (GMT)
Message-ID: <025201c583ac$28ef4470$cb00000a@db01.voxpilot.com>
From: "Dave Burke" <david.burke@voxpilot.com>
To: "Shanmugham, Saravanan" <sarvi@cisco.com>,
	"Eric Burger" <eburger@brooktrout.com>, <speechsc@ietf.org>
References: <03772D1EC8DE624A863058C75874A75C1128E5@vtg-um-e2k6.sj21ad.cisco.com>
Subject: Re: [Speechsc] START-OF-SPEECH in DTMF-only mode -> summary(?)
Date: Fri, 8 Jul 2005 11:59:44 +0100
MIME-Version: 1.0
Content-Type: text/plain; format=flowed; charset="iso-8859-1";
	reply-type=original
Content-Transfer-Encoding: 7bit
X-Priority: 3
X-MSMail-Priority: Normal
X-Mailer: Microsoft Outlook Express 6.00.2900.2180
X-MimeOLE: Produced By Microsoft MimeOLE V6.00.2900.2180
X-Spam-Score: 0.0 (/)
X-Scan-Signature: 320ddee61cb2d39f489e9191ae3fdc8c
Content-Transfer-Encoding: 7bit
Cc: 
X-BeenThere: speechsc@ietf.org
X-Mailman-Version: 2.1.5
Precedence: list
List-Id: Speech Services Control Working Group <speechsc.ietf.org>
List-Unsubscribe: <https://www1.ietf.org/mailman/listinfo/speechsc>,
	<mailto:speechsc-request@ietf.org?subject=unsubscribe>
List-Post: <mailto:speechsc@ietf.org>
List-Help: <mailto:speechsc-request@ietf.org?subject=help>
List-Subscribe: <https://www1.ietf.org/mailman/listinfo/speechsc>,
	<mailto:speechsc-request@ietf.org?subject=subscribe>
Sender: speechsc-bounces@ietf.org
Errors-To: speechsc-bounces@ietf.org

I agree with your opinion on what constitutes a barge-in. Based on Klaus' 
e-mail, I am more concerned that using the grammar type to decide what 
constitutes a barge-in is going to make VoiceXML implementations difficult.

It is easy to map VoiceXML application selected inputmodes to MRCP 
resources:
a. inputmodes="dtmf" -> Use a dtmfrecog
b. inputmodes="speech" -> Use a "speech-only-recog"
c. inputmodes="both" -> Use a speechrecog (or a combination of 
"speech-only-recog" + dtmfrecog)

The only problem is (b). While not as important as being able to do 
DTMF-only recognition, I believe we DO need to support speech-only 
recongition so as (a) to avoid unnecessary limitations in VUI design, and 
(b) to facilitate VoiceXML implementations. I think we should add a header 
to RECOGNIZE so the client is able to always _explicitly_ set the 
inputmodes.

In retrospect, I don't like the idea of multiple START-OF-SPEECH events 
being generated from the same media resource. This is because one assumes 
that the START-OF-SPEECH should be of the same type as the hypothesis 
returned in the RECOGNITION-COMPETE message - most implementations, on 
hearing one input mode type disable the recogniser of the other type. I do 
like David's idea of renaming START-OF-SPEECH to something like 
START-OF-INPUT and carrying a type header because it is neater and more 
extensible.

So in summary, I propose we modify the spec to:

1. Clarify what constitutes a barge-in for a dtmfrecog and speechrecog as 
per Sarvi's e-mail (and in agreement with Klaus' for dtmfrecog).

2. Specifiy an InputModes header to RECOGNIZE (defaults to "both", can also 
be "speech" or "DTMF"). Setting to speech for a speechrecog results in the 
hypothesised "speech-only-recog". Setting to DTMF for a speechrecog is 
equivalent to using a dtmfrecog. Edge cases: Setting to speech for a 
dtmfrecog will result in a noinput as will setting to DTMF for a speechrecog 
which does not support DTMF.

3. Change START-OF-SPEECH to START-OF-INPUT. Clarify it occurs once and 
coincides with barge-in.

4. Add a header of InputType to START-OF-INPUT. Current specified values are 
"dtmf" or "speech".

Dave



----- Original Message ----- 
From: "Shanmugham, Saravanan" <sarvi@cisco.com>
To: "Eric Burger" <eburger@brooktrout.com>; <speechsc@ietf.org>
Sent: Thursday, July 07, 2005 5:52 PM
Subject: RE: [Speechsc] START-OF-SPEECH in DTMF-only mode -> summary(?)


I don't think generating multiple START-OF-SPEECH events is a solution.
We still haven't addressed, what constitutes a barge-in, for the
optimised case. That should also be the single point when a single
START-OF-SPEECH(or whathever else you want to name it) should be
generated.

That point, in my opinion should be
   1. For "dtmf-recog" resources should be the beginning of a DTMF key
press.
   2. For "speech-recog" resources should be the beginning of a DTMF key
press or the beginning of speech. This is should be irrespective of what
type of grammar is being used. Coz even numbers only grammars can still
be spoken and hence cannot be assumed to be a cue for DTMF only
recognition.
   3. For "speech-only-recog" resources(which are not defined today) the
time to barge-in is the beginning of speech. I don't see a need for such
a resource today. But I am mentioning this for completeness.

Sarvi


     -----Original Message-----
     From: speechsc-bounces@ietf.org
     [mailto:speechsc-bounces@ietf.org] On Behalf Of Eric Burger
     Sent: Thursday, July 07, 2005 5:32 AM
     To: speechsc@ietf.org
     Subject: RE: [Speechsc] START-OF-SPEECH in DTMF-only mode
     -> summary(?)

     Can we declare consensus?

     > -----Original Message-----
     > From: speechsc-bounces@ietf.org
     [mailto:speechsc-bounces@ietf.org]
     > Sent: Wednesday, July 06, 2005 10:45 AM
     > To: Dave Burke
     > Cc: speechsc@ietf.org; Shanmugham, Saravanan; Klaus Reifenrath
     > Subject: Re: [Speechsc] START-OF-SPEECH in DTMF-only mode ->
     > summary(?)
     >
     > I think we're getting close. I though about snipping out
     some pieces
     > to cut down the text, but I realized the context is
     still needed. See
     > inline.
     >
     > On Jul 6, 2005, at 10:29 AM, Dave Burke wrote:
     >
     > > Inline.
     > >
     > > Dave
     > >
     > > ----- Original Message ----- From: "David R Oran"
     <oran@cisco.com>
     > > To: "Dave Burke" <david.burke@voxpilot.com>
     > > Cc: <speechsc@ietf.org>; "Shanmugham, Saravanan"
     <sarvi@cisco.com>;
     > > "Klaus Reifenrath" <Klaus.Reifenrath@Scansoft.com>
     > > Sent: Wednesday, July 06, 2005 1:07 PM
     > > Subject: Re: [Speechsc] START-OF-SPEECH in DTMF-only
     mode -> summary
     > > (?)
     > >
     > >
     > >
     > >>
     > >> On Jul 6, 2005, at 5:31 AM, Dave Burke wrote:
     > >>
     > >>
     > >>> + Attempting to summarise:
     > >>>
     > >>> 1. START-OF-SPEECH is useful for the client to know when
     > to stop
     > >>> playing prompts the in non-optimised case 2.
     START-OF-SPEECH is
     > >>> useful for the client to calculate the bargin  time
     (e.g. VoiceXML
     > >>> 2.1 <mark>)
     > >>> 3. In the optimised case, a bargin automatically
     stops prompt
     > >>> playing (assuming prompts barginable) 4. Because of
     the previous
     > >>> point, the question of what input type  caused
     bargin is different
     > >>> and less important to
     > what input
     > >>> type(s)  the recogniser is listening for
     > >>>
     > >> I'm not sure I follow point 4. Could you elaborate?
     > >>
     > >
     > > DB> Adding a parameter to the START-OF-SPEECH event would
     > certainly
     > > allow the client to ignore (i.e. let prompts continue
     playing) the
     > > event if the event type is not of interest (e.g. the
     client would
     > > ignore speech start events when it is interested only
     in  DTMF start
     > > events). This _only_ works for the non-optimised case,
     however. For
     > > the optimised case, assuming START-OF-SPEECH
     > coincides
     > > with the bargin signal to the speechsynth, prompts
     will stop playing
     > > for inputs that the client might not be interested
     > in (e.g.
     > > a speech input will stop prompts playing even if the client
     > is only
     > > interested in DTMF).
     > >
     > OK, now I get it. There's a need for the client to both
     handle the
     > non-optimized case itself, and influence or at least
     have a clue what
     > the server is going to do in the optimized case.
     >
     > >
     > >>
     > >>
     > >>> 5. Currently, MRCPv2 has no way of indicating what
     input type(s) a
     > >>> recogniser is listening for
     > >>>
     > >>>
     > >> Do you mean exactly this, or do you mean "for the client to
     > >> indicate  to the resource what input types it should
     look for"?
     > >>
     > >
     > > DB> Yes exactly - apologies for not being clear.
     > >
     > >
     > >>
     > >>
     > >>> + Why implement 5?
     > >>>
     > >>> i. Noisy case: Need DTMF-only recognition (and may
     only have a
     > >>> speechrecog)
     > >>>
     > >> I'm having difficulty following the logic of why
     noise would
     > >> necessarily trigger START-OF-SPEECH if you were listening for
     > >> speech  but not DTMF. I suppose you can use a more forgiving
     > >> discriminator if  all you need to tell is if you're
     getting DTMF,
     > >> but I've had a number  of real-world cases where wind
     noise was
     > >> detected as DTMF, and  there's always the ambiguity
     when you have
     > >> Captain Crunch on the  phone. In either case in the
     non-optimized
     > >> case it's the client who  gets to decide whether an
     event should be
     > >> interpreted as barge-in or  not, so it seems an
     aesthetic protocol
     > >> design decision whether the  client tells the server
     ahead of time
     > >> what circumstances to generate  the START-OF-SPEECH
     event for, or
     > >> whether the event gets generated  and the client
     decides based on
     > >> what's in the event whether it should  be
     > treated
     > >> as barge- in.
     > >>
     > >> I suppose one could make the argument that because
     the spec implies
     > >> that the event can only be generated once per request
     that if a
     > >> DTMF/ speech capable recognizer first hears
     > enough noise
     > >> to think it's  hearing speech and later hears DTMF,
     the client will
     > >> declare barge-in  when the event comes and not when he DTMF
     > >> actually gets heard.
     > >>
     > >> If that's deemed a problem, we can still handle that
     in the design
     > >> where the server just reports what it's hearing by allowing
     > >> multiple  events to be generated during a single request.
     > >>
     > >> Between the approach just outlined above, and an
     approach where the
     > >> client provides a filter for whether to generate the
     event or not,
     > >> I  have a mild preference (based on aesthetics rather
     than some
     > >> hard  engineering tradeoff) for the approach where
     > the server
     > >> just reports  what it's hearing.
     > >>
     > >>
     > >> Having had some useful exchanges on this topic, it
     also is becoming
     > >> apparent to me that this event is poorly named, and
     we should
     > >> consider renaming it to "INTERESTING-INPUT-HEARD" or
     something akin
     > >> to that, because as others have pointed out, a
     DTMF-only recognizer
     > >> will never detect "start of speech".
     > >>
     > >> Another consideration to fold into the design choice is
     > >> extensibility. Bear with me through a little
     gedankenexperiment.
     > >>
     > >> Suppose we want to define a new recognizer type,
     which I'll call
     > >> the "name that tune" recognizer. The client plays
     music to the
     > >> server and  the server recognizes musical notes. The
     grammar is a
     > >> standard  musical notation, augmented with a semantic
     > >> interpretation that  transforms the notes into the
     title of the
     > >> tune and provides that as  an answer.
     > >>
     > >> First, there's no speech involved (or is there...hang
     on a minute).
     > >> Second, in order to accommodate the "name that tune"
     > >> recognizer, we'd have to extend both the client and
     the server to
     > >> undetstand a  directive as to whether to recognize
     music or now,
     > >> inaddition to what  the server already knows what to
     do based on
     > >> the grammar. If you  follow my logic above, whether
     or not we do
     > >> that, we have to extend  "start-of-speech" to say
     "I'm hearing
     > >> music". So far fairly  straightforward, but let me
     now throw in the
     > >> pathological twist.
     > >>
     > >> Suppose what I feed to a  combined music/speech
     recognizer is a
     > >> work  in sprechstimme (spoken music), like the
     "Geographical Fugue"
     > >> (aside:  this is a wonderful piece of music I highly
     recommend to
     > >> anyone  interested in small ensemble singing). In
     this case, the
     > >> tune could  be named by either doing speech or music
     recognition.
     > >> Why is there  any need for the client to constrain
     the server as to
     > >> which it tries  to do when it's
     > already
     > >> told the server what it wants through the  grammar?
     > >>
     > >> A few other comments below
     > >>
     > >>
     > >>> ii. Flexibility: Want speech-only recognition
     (because a second
     > >>> recogniser is doing hotword on DTMF)
     > >>>
     > >>>
     > >> I don't see how flexibility is affected by this deisgn
     > choice. If
     > >> that's what you want, feed the speech-only recognizer
     a grammar
     > >> without any DTMF rules.
     > >>
     > >>
     > >>> + How to implement 5?
     > >>>
     > >>> a. Implicitly:
     > >>>    - dtmfrecog: always DTMF-only recognition
     > >>>    - speechrecog: depends on active grammar type
     > >>>        > if a dtmf grammar is active then DTMF input is "on"
     > >>>        > if a speech grammar is active then speech
     input is "on"
     > >>>
     > >>> b. Explicitly:
     > >>>    - Add inputmodes header to RECOGNIZE
     > >>>
     > >>> Option a is David's "do what I mean case"; option b is
     > the extra
     > >>> dial for the client.
     > >>>
     > >>>
     > >> Actually, that's not the point I was making with "do
     what I mean",
     > >> but it's not essential to the discussion so let's move on.
     > >>
     > >>
     > >>> It is worth noting that VoiceXML uses option b. This
     allows one to
     > >>> activate both speech grammars and DTMF grammars (and
     > therefore
     > >>> be informed of any errors in the grammars at activation
     > time) but
     > >>> independently turn on whichever input mode you like e.g.
     > perhaps
     > >>> start with "both" then change to "dtmf".
     > >>>
     > >>>
     > >> I'm not sure the VXML precedent is relevant here,
     because the
     > >> application behind VXML is working a different part of the
     > problem
     > >> -  how to traverse a TUI dialog based on different
     parts of the
     > >> input  space. In fact, I suspect that the VXML: choice was
     > >> conditioned more  by limitations at the time it was
     > specified than
     > >> an underlying good  design choice. Clearly having to
     specify this
     > >> in VXML make the job of  handling a TUI with nodes
     like "Say or
     > >> press 5" harder rather than  easier.
     > >>
     > >
     > > DB> The VoiceXML edge-case is pretty weird so it's not a major
     > > concern. My main concern is that the client can
     indicate, somehow,
     > > what the input modes are.
     > >
     > >
     > >>
     > >> Summing up, while I don't feel strongly one way or
     the other, I
     > >> have  a preference for handling this as follows:
     > >>
     > >> a) Rename "START-OF-SPEECH" to
     "INTERESTING-INPUT-RECEIVED" or
     > >> something equivalent.
     > >> b) Include a parameter in the event saying what was
     interesting
     > >> about  the input you received, with a registry of values which
     > >> includes:
     > >>     - signal above noise floor
     > >>     - speech
     > >>     - dtmf
     > >>     - (possibly) music
     > >> c) allow the event to be generated multiple times
     during a request
     > >>
     > >>
     > >
     > > DB> I like these suggestions (START-OF-INPUT?).
     However, I don't
     > > see how the problem of the optimised case is not
     solved by them. I
     > > think the optimised case is fine if we have the
     following rules:
     > >
     > Yes, I hadn't thought through the optimized case as
     thoroughly as you.
     > Your suggested method name is fine by me as well
     >
     > > 1. START-OF-SPEECH (and optimised bargin) is only generated
     > for the
     > > input type that is being listened for 2. A speechrecog
     listens for
     > > DTMF if DTMF grammars are active, speech if speech
     grammars are
     > > active, or speech and DTMF if both grammar types are active.
     > >
     > Works for me.
     >
     > >
     > >> Note that all of the above I'm saying with my technical
     > hat on and
     > >> my chair hat off.
     > >> Putting my chair hat on for a moment, we really need
     to get this
     > >> spec  to last call, so at some point Eric or I is going to
     > declare
     > >> rough  consensus so we can move on.
     > >>
     > >
     > > DB> Agreed!
     > >
     > >
     > >>
     > >> Dave Oran.
     > >>
     > >>> Dave
     > >>>
     > >>>
     > >>>
     > >>>
     > >>>
     > >>>
     > >>>
     > >>>
     > >>>
     > >>>
     > >>>
     > >>>
     > >>> Sarvi makes a good point that adding the reason why the
     > START-OF-
     > >>> SPEECH occurred does not fix the optimised bargin case.
     > >>>
     > >>> dtmfrecog - listens for DTMF only
     > >>> speechrecog - listens for DTMF only, or speech only,
     or speech &
     > >>> DTMF
     > >>>
     > >>>
     > >>>
     > >>>
     > >>>
     > >>>
     > >>>
     > >>>
     > >>> ----- Original Message ----- From: "Shanmugham, Saravanan"
     > >>> <sarvi@cisco.com>
     > >>> To: "David R Oran" <oran@cisco.com>; "Dave Burke"
     > >>> <david.burke@voxpilot.com>
     > >>> Cc: <speechsc@ietf.org>; "Klaus Reifenrath"
     > >>> <Klaus.Reifenrath@Scansoft.com>
     > >>> Sent: Tuesday, July 05, 2005 8:51 PM
     > >>> Subject: RE: [Speechsc] START-OF-SPEECH in DTMF-only mode
     > >>>
     > >>>
     > >>> inline.
     > >>>
     > >>>     -----Original Message-----
     > >>>     From: David R Oran [mailto:oran@cisco.com]
     > >>>     Sent: Tuesday, July 05, 2005 11:08 AM
     > >>>     To: Dave Burke
     > >>>     Cc: Shanmugham, Saravanan; Klaus Reifenrath;
     speechsc@ietf.org
     > >>>     Subject: Re: [Speechsc] START-OF-SPEECH in DTMF-only mode
     > >>>
     > >>>
     > >>>     On Jul 5, 2005, at 12:54 PM, Dave Burke wrote:
     > >>>
     > >>>     > Inline.
     > >>>     >
     > >>>     > Dave
     > >>>     >
     > >>>     > ----- Original Message ----- From:
     "Shanmugham, Saravanan"
     > >>>     > <sarvi@cisco.com>
     > >>>     > To: "David R Oran" <oran@cisco.com>; "Klaus Reifenrath"
     > >>>     > <Klaus.Reifenrath@Scansoft.com>
     > >>>     > Cc: <speechsc@ietf.org>
     > >>>     > Sent: Tuesday, July 05, 2005 5:22 PM
     > >>>     > Subject: RE: [Speechsc] START-OF-SPEECH in
     DTMF-only mode
     > >>>     >
     > >>>     >
     > >>>     > I agree with Dave's analysis. The purpose of this event
     > >>>     was barge-in.
     > >>>     > And barge-in should happen for both DTMF and speech.
     > >>>     >
     > >>>     > Is there a case where you think it should not behave
     > >>>     this way. If soe,
     > >>>     > please provide a scenario where you think
     > >>>     >    1. Barge-in should happen for DTMF and not voice or
     > >>>     vice-versa.
     > >>>     >
     > >>>     > DB> You want to do a DTMF recognition only
     because it is
     > >>> noisy.
     > >>>     > While waiting for DTMF input, the speechrecog resource
     > >>>     (or advanced
     > >>>     > dtmfrecog) generates a START-OF-SPEECH because it heard
     > >>>     some speech.
     > >>>     > The client does not want to stop prompt playing unless
     > >>>     DTMF was heard
     > >>>     > but it can't tell by the START-OF-SPEECH whether speech
     > >>>     or DTMF was
     > >>>     > heard. Similarly vice versa.
     > >>>     >
     > >>>     It's an interesting design question what part of the
     > >>>     policy resides at the client and what at the server, and
     > >>>     who makes the "final decision" about whether what was
     > >>>     heard was relevant to the control channel. Right now we
     > >>>     (IMO) have a weird partitioning in many cases where the
     > >>>     client basically says "do what I mean", but there are no
     > >>>     constraints of what the server actually does, and no
     > >>>     normalized basis for the client to figure out what to set
     > >>>     various magic numbers to (e.g. sensitivity).
     > >>>
     > >>>     In this case the only thing the client needs to decide is
     > >>>     whether to kill the prompt because the server thinks
     > >>>     something that would interfere with the feedback
     > >>>     ear/mouth/finger control happened. What this
     says to me is
     > >>>     that it isn't necessarily a good idea for the client to
     > >>>     have more knobs to control the server
     (especially if those
     > >>>     knows are just more value/policy input ungrounded in any
     > >>>     physics/ acoustics). On the other hand, having the server
     > >>>     tell the client more about what it thinks is going on is
     > >>>     probably valuable.
     > >>>
     > >>>     So, Coming to the point after this long rambling
     > >>>     introduction, I think it would in fact be useful for the
     > >>>     START-Of-SPEECH event to indicate some extra information,
     > >>>     for example:
     > >>>     a) I got something enough above the noise floor
     to qualify
     > >>>     for exceeding the "Sensisitvity" parameter you sent in on
     > >>>     the request but I really can't tell what it is
     (could be a
     > >>>     hippopatmus fart, or a siren in the background,
     or captain
     > >>>     crunch trying to whistle DTMF).
     > >>>     b) I think I'm hearing speech
     > >>>     c) I think I'm hearing DTMF
     > >>>
     > >>> Though I agree with your former part of your
     response. I am not
     > >>> sure I agree with your proposed solution.
     > >>> The way I see this problem is that, it is more of what
     > constitues a
     > >>> barge-in event. This boils down to whether it is
     speech, DTMF or
     > >>> both.
     > >>> This is inturn boils down to what type of recognizer
     > resource we are
     > >>> using, dtmf-recog, speech-recog, and speech-only-recog(we
     > don't have
     > >>> this and I don't think we should add it, but think
     of this as a
     > >>> place holder that explains the concept).
     > >>>
     > >>> A client knowing what type of barge-in happenned, does
     > not impact
     > >>> the
     > >>> barge-in operation itself as it may be too late(for
     the optimized
     > >>> barge-in case). It may have other use cases, and if we
     > can identify
     > >>> them, I don't mind adding support for the
     START-OF-SPEECH event to
     > >>> say what type of barge-in happenned. But that itself does not
     > solve the
     > >>> original problem raised. Refer to my previous response.
     > >>>
     > >>> The solution lies in defining what what is a
     barge-in event.
     > >>> That  boils
     > >>> down to what type of recognition is happenning,
     > dtmf-only, speech-
     > >>> dtmf
     > >>> or speech-only. We do not support speech-only as a
     resource today,
     > >>> the question is do we need a header to force it.
     > >>>
     > >>> Sarvi
     > >>>
     > >>>
     > >>>     >    2. You would benefit from the client knowing
     > what caused
     > >>> the
     > >>>     > barge-in, DTMF Vs speech.
     > >>>     >
     > >>>     > DB> See previous comment. And previous e-mail:
     either add an
     > >>>     > inputmodes header (taking value speech, dtmf, both) to
     > >>>     the RECOGNIZE
     > >>>     > request or add a header to the START-OF-SPEECH event
     > >>>     indicating DTMF
     > >>>     > or speech.
     > >>>     >
     > >>>     I'm leaning in your direction on this latter point - as
     > >>>     should be evident from what I wrote above.
     > >>>
     > >>>     > Sarvi
     > >>>     >
     > >>>     >     -----Original Message-----
     > >>>     >     From: speechsc-bounces@ietf.org
     > >>>     >     [mailto:speechsc-bounces@ietf.org] On Behalf Of
     > David R
     > >>> Oran
     > >>>     >     Sent: Tuesday, July 05, 2005 5:24 AM
     > >>>     >     To: Klaus Reifenrath
     > >>>     >     Cc: 'speechsc@ietf.org'
     > >>>     >     Subject: Re: [Speechsc] START-OF-SPEECH in
     > DTMF-only mode
     > >>>     >
     > >>>     >
     > >>>     >     On Jul 5, 2005, at 3:46 AM, Reifenrath,
     Klaus wrote:
     > >>>     >
     > >>>     >     > The current spec is not clear when
     > START-OF-SPEECH need
     > >>>     >     to be send in
     > >>>     >     > the following scenarios:
     > >>>     >     > A) The client requested a DTMF Recognizer. Is the
     > >>>     >     START-OF-SPEECH
     > >>>     >     > event send to the client also if speech
     was detected?
     > >>>     >     I suspect so, since one of the prime purposes
     > is to enable
     > >>>     >     client- mediated barge-in handling. However, if the
     > >>>     >     recognizer is in fact only capable of
     recognizing DTMF
     > >>>     >     then it may in fact not report anythin
     unless it's using
     > >>>     >     some primitive thresholding machinery, like a SN
     > >>> threshold.
     > >>>     >     > B) The client requested a Speech
     Recognizer, but only
     > >>>     >     activated DTMF
     > >>>     >     > grammars. Is the START-OF-SPEECH event
     send to the
     > >>>     >     client also if
     > >>>     >     > speech was detected?
     > >>>     >     Again, I'd say yes, for the same reason as above.
     > >>>     >     > I think in both cases START-OF-SPEECH should only
     > >>>     be send after
     > >>>     >     > detecting a DTMF digit (see Figure 12 of VoiceXML
     > >>>     2.0: http://
     > >>>     >     > www.w3.org/TR/voicexml20/#dmlATiming).
     > >>>     >     We seem to have reached different
     conclusions. I'd be
     > >>>     >     interested in why you think my analysis
     above is wrong.
     > >>>     >
     > >>>     >     Dave.
     > >>>     >
     > >>>     >     > Klaus
     > >>>     >     >
     > >>>     >     > _______________________________________________
     > >>>     >     > Speechsc mailing list
     > >>>     >     > Speechsc@ietf.org
     > >>>     >     > https://www1.ietf.org/mailman/listinfo/speechsc
     > >>>     >     >
     > >>>     >
     > >>>     >     _______________________________________________
     > >>>     >     Speechsc mailing list
     > >>>     >     Speechsc@ietf.org
     > >>>     >     https://www1.ietf.org/mailman/listinfo/speechsc
     > >>>     >
     > >>>     >
     > >>>     > _______________________________________________
     > >>>     > Speechsc mailing list
     > >>>     > Speechsc@ietf.org
     > >>>     > https://www1.ietf.org/mailman/listinfo/speechsc
     > >>>     >
     > >>>
     > >>>
     > >>> _______________________________________________
     > >>> Speechsc mailing list
     > >>> Speechsc@ietf.org
     > >>> https://www1.ietf.org/mailman/listinfo/speechsc
     > >>>
     > >>>
     > >>
     > >> _______________________________________________
     > >> Speechsc mailing list
     > >> Speechsc@ietf.org
     > >> https://www1.ietf.org/mailman/listinfo/speechsc
     > >
     >
     > _______________________________________________
     > Speechsc mailing list
     > Speechsc@ietf.org
     > https://www1.ietf.org/mailman/listinfo/speechsc
     >


     _______________________________________________
     Speechsc mailing list
     Speechsc@ietf.org
     https://www1.ietf.org/mailman/listinfo/speechsc


_______________________________________________
Speechsc mailing list
Speechsc@ietf.org
https://www1.ietf.org/mailman/listinfo/speechsc


_______________________________________________
Speechsc mailing list
Speechsc@ietf.org
https://www1.ietf.org/mailman/listinfo/speechsc



From speechsc-bounces@ietf.org Fri Jul 08 10:05:48 2005
Received: from localhost.localdomain ([127.0.0.1] helo=megatron.ietf.org)
	by megatron.ietf.org with esmtp (Exim 4.32)
	id 1DqtUB-0007xz-S8; Fri, 08 Jul 2005 10:05:47 -0400
Received: from odin.ietf.org ([132.151.1.176] helo=ietf.org)
	by megatron.ietf.org with esmtp (Exim 4.32) id 1DqtU7-0007w6-CR
	for speechsc@megatron.ietf.org; Fri, 08 Jul 2005 10:05:46 -0400
Received: from ietf-mx.ietf.org (ietf-mx [132.151.6.1])
	by ietf.org (8.9.1a/8.9.1a) with ESMTP id KAA03979
	for <speechsc@ietf.org>; Fri, 8 Jul 2005 10:05:41 -0400 (EDT)
Received: from sj-iport-3-in.cisco.com ([171.71.176.72]
	helo=sj-iport-3.cisco.com) by ietf-mx.ietf.org with esmtp (Exim 4.43)
	id 1DqtvU-0001tV-EX
	for speechsc@ietf.org; Fri, 08 Jul 2005 10:34:02 -0400
Received: from sj-core-2.cisco.com (171.71.177.254)
	by sj-iport-3.cisco.com with ESMTP; 08 Jul 2005 07:05:31 -0700
X-IronPort-AV: i="3.93,274,1115017200"; 
	d="scan'208"; a="294991134:sNHT81343176"
Received: from imail.cisco.com (imail.cisco.com [128.107.200.91])
	by sj-core-2.cisco.com (8.12.10/8.12.6) with ESMTP id j68E5Ood012549;
	Fri, 8 Jul 2005 07:05:24 -0700 (PDT)
Received: from [10.32.245.152] (stealth-10-32-245-152.cisco.com
	[10.32.245.152])
	by imail.cisco.com (8.12.11/8.12.10) with SMTP id j68E4OmT026025;
	Fri, 8 Jul 2005 07:04:24 -0700
In-Reply-To: <025201c583ac$28ef4470$cb00000a@db01.voxpilot.com>
References: <03772D1EC8DE624A863058C75874A75C1128E5@vtg-um-e2k6.sj21ad.cisco.com>
	<025201c583ac$28ef4470$cb00000a@db01.voxpilot.com>
Mime-Version: 1.0 (Apple Message framework v730)
X-Priority: 3
Content-Type: text/plain; charset=US-ASCII; delsp=yes; format=flowed
Message-Id: <C477D12D-D2C0-46B4-B83B-E8D71BEF522C@cisco.com>
Content-Transfer-Encoding: 7bit
From: David R Oran <oran@cisco.com>
Subject: Re: [Speechsc] START-OF-SPEECH in DTMF-only mode -> summary(?)
Date: Fri, 8 Jul 2005 10:05:22 -0400
To: "Dave Burke" <david.burke@voxpilot.com>
X-Mailer: Apple Mail (2.730)
IIM-SIG: v:"1.1"; h:"imail.cisco.com"; d:"cisco.com"; z:"home"; m:"krs";
	t:"1120831467.108940"; x:"432200"; a:"rsa-sha1"; b:"nofws:21674";
	e:"Iw=="; n:"sQYarK2E51MdcTiUqeif3F7cWdxIfoCiXhdfb9vD5ee/j0jXL15gbFxF2p"
	"XIweAblu0N6XAgK7k+wrbr7bQDJaCDqOmzqpRUBjIRQAXQ7NzadpmR3pUL6wxaRUtW+c43sl9jC"
	"50Qg1sXHpPjt8Y+Y16ioyQAQAdSunM4YhevURc=";
	s:"XNDJD+Jkmethd1qw49ZcikYXxjqh710j1r0Wb1afpxVXxQNXkeY0yoBY/+OrB1Il47KimDNq"
	"xBfDSn50YTfRoPpJU282ABH4qDWJAtTjqpCcMrNpdgHK0MUTn2/XIlXWYDE0wc0NA35eU0KOjgg"
	"/1+F8AP1dHuE966V5oBfI4B4="; c:"From: David R Oran <oran@cisco.com>";
	c:"Subject: Re: [Speechsc] START-OF-SPEECH in DTMF-only mode ->
	summary" "(?)"; c:"Date: Fri, 8 Jul 2005 10:05:22 -0400"
IIM-VERIFY: s:"y"; v:"y"; r:"60"; h:"imail.cisco.com";
	c:"message from imail.cisco.com verified; "
X-Spam-Score: 0.0 (/)
X-Scan-Signature: 71427903e4ce43cf4879c36ef9c04bc1
Content-Transfer-Encoding: 7bit
Cc: speechsc@ietf.org, "Shanmugham, Saravanan" <sarvi@cisco.com>,
	Eric Burger <eburger@brooktrout.com>
X-BeenThere: speechsc@ietf.org
X-Mailman-Version: 2.1.5
Precedence: list
List-Id: Speech Services Control Working Group <speechsc.ietf.org>
List-Unsubscribe: <https://www1.ietf.org/mailman/listinfo/speechsc>,
	<mailto:speechsc-request@ietf.org?subject=unsubscribe>
List-Post: <mailto:speechsc@ietf.org>
List-Help: <mailto:speechsc-request@ietf.org?subject=help>
List-Subscribe: <https://www1.ietf.org/mailman/listinfo/speechsc>,
	<mailto:speechsc-request@ietf.org?subject=subscribe>
Sender: speechsc-bounces@ietf.org
Errors-To: speechsc-bounces@ietf.org


On Jul 8, 2005, at 6:59 AM, Dave Burke wrote:

> I agree with your opinion on what constitutes a barge-in. Based on  
> Klaus' e-mail, I am more concerned that using the grammar type to  
> decide what constitutes a barge-in is going to make VoiceXML  
> implementations difficult.
>
> It is easy to map VoiceXML application selected inputmodes to MRCP  
> resources:
> a. inputmodes="dtmf" -> Use a dtmfrecog
> b. inputmodes="speech" -> Use a "speech-only-recog"
> c. inputmodes="both" -> Use a speechrecog (or a combination of  
> "speech-only-recog" + dtmfrecog)
>
> The only problem is (b). While not as important as being able to do  
> DTMF-only recognition, I believe we DO need to support speech-only  
> recongition so as (a) to avoid unnecessary limitations in VUI  
> design, and (b) to facilitate VoiceXML implementations. I think we  
> should add a header to RECOGNIZE so the client is able to always  
> _explicitly_ set the inputmodes.
>
I'm not sure I buy this, since what the recognizer is looking for and  
what constitutes an input that should be considered a potential barge- 
in strike me as independent. I'm similarly not persuaded that we need  
the flexibility to set the input mode independently of the grammar,  
since it leads to all sorts of inconsistent states (e.g. speech-only  
grammar with an input-mode of dtmf). On the other hand I can see the  
need for the client to specify what sorts of input out to generate  
the start-of-input (nee start-of-speech) event on.

> In retrospect, I don't like the idea of multiple START-OF-SPEECH  
> events being generated from the same media resource. This is  
> because one assumes that the START-OF-SPEECH should be of the same  
> type as the hypothesis returned in the RECOGNITION-COMPETE message  
> - most implementations, on hearing one input mode type disable the  
> recogniser of the other type.
I'm not sure I follow this logic, but I'm not wedded to the idea of  
allowing multiple events. It was trying to solve the problem of  
ambiguity around what the recognizer was hearing and the idea that  
the client might care. If you don't think the client will ever care,  
then we don't need the capability.


> I do like David's idea of renaming START-OF-SPEECH to something  
> like START-OF-INPUT and carrying a type header because it is neater  
> and more extensible.
>
> So in summary, I propose we modify the spec to:
>
> 1. Clarify what constitutes a barge-in for a dtmfrecog and  
> speechrecog as per Sarvi's e-mail (and in agreement with Klaus' for  
> dtmfrecog).
>
Hmmm, ok, but I think the issue is actually clarifying what the  
recognizer declares as "interesting input", which it may decide also  
constitutes a barge-in in the optimized case, and in either case  
reports to the client that it heard.

> 2. Specifiy an InputModes header to RECOGNIZE (defaults to "both",  
> can also be "speech" or "DTMF"). Setting to speech for a  
> speechrecog results in the hypothesised "speech-only-recog".  
> Setting to DTMF for a speechrecog is equivalent to using a  
> dtmfrecog. Edge cases: Setting to speech for a dtmfrecog will  
> result in a noinput as will setting to DTMF for a speechrecog which  
> does not support DTMF.
>
I'm ok with having a header for the client to tell the server when it  
would like the start-of-input event to be generated and what the  
client considers to be the "interesting input" that the recognizer  
should use to do the discrimination, and possibly do barge-in  
processing in the optimized case.

I'm less ok with the proposed domain of values, since it will have  
extensibility problems. Especially problematical is have a code point  
of "both" since that will be ambiguous if we even define a "music"  
recognizer, or a "gesture" recognizer using video input. Here's my  
counter-proposal:

Create a header on the recognizer resource requests called "start- 
input-on:". Define a registry of values, with the following three  
values initially defined:
     - dtmf
     - speech
     - music

The semantics would be that the resource is to generate a "start-of- 
input" event, apply the defined grammars to what follows, and do any  
local optimized barge-in processing if ANY of the enumerated input  
types is detected. That way you can say
     "start-input-on: dtmf" if you want to just  dtmf,
     "start-input-on: speech" if you want just speech (dtmf input  
would be ignored even if the grammar supported dtmf)
     "start-input- on: speech, dtmf" if you wanted both
etc.

> 3. Change START-OF-SPEECH to START-OF-INPUT. Clarify it occurs once  
> and coincides with barge-in.
>
Ok for the "only once", but I'd like to tighten it up to talk about  
more than just barge-in, as I suggested above.

> 4. Add a header of InputType to START-OF-INPUT. Current specified  
> values are "dtmf" or "speech".
>
Ok, with slight modification. Say that the syntax of the "input-type"  
header is a single-valued subset of the registered value(s) that are  
defined for the start-input-on: header.

Comments?

Dave O. (technical hat on, chair hat off).

> Dave
>
>
>
> ----- Original Message ----- From: "Shanmugham, Saravanan"  
> <sarvi@cisco.com>
> To: "Eric Burger" <eburger@brooktrout.com>; <speechsc@ietf.org>
> Sent: Thursday, July 07, 2005 5:52 PM
> Subject: RE: [Speechsc] START-OF-SPEECH in DTMF-only mode -> summary 
> (?)
>
>
> I don't think generating multiple START-OF-SPEECH events is a  
> solution.
> We still haven't addressed, what constitutes a barge-in, for the
> optimised case. That should also be the single point when a single
> START-OF-SPEECH(or whathever else you want to name it) should be
> generated.
>
> That point, in my opinion should be
>   1. For "dtmf-recog" resources should be the beginning of a DTMF key
> press.
>   2. For "speech-recog" resources should be the beginning of a DTMF  
> key
> press or the beginning of speech. This is should be irrespective of  
> what
> type of grammar is being used. Coz even numbers only grammars can  
> still
> be spoken and hence cannot be assumed to be a cue for DTMF only
> recognition.
>   3. For "speech-only-recog" resources(which are not defined today)  
> the
> time to barge-in is the beginning of speech. I don't see a need for  
> such
> a resource today. But I am mentioning this for completeness.
>
> Sarvi
>
>
>     -----Original Message-----
>     From: speechsc-bounces@ietf.org
>     [mailto:speechsc-bounces@ietf.org] On Behalf Of Eric Burger
>     Sent: Thursday, July 07, 2005 5:32 AM
>     To: speechsc@ietf.org
>     Subject: RE: [Speechsc] START-OF-SPEECH in DTMF-only mode
>     -> summary(?)
>
>     Can we declare consensus?
>
>     > -----Original Message-----
>     > From: speechsc-bounces@ietf.org
>     [mailto:speechsc-bounces@ietf.org]
>     > Sent: Wednesday, July 06, 2005 10:45 AM
>     > To: Dave Burke
>     > Cc: speechsc@ietf.org; Shanmugham, Saravanan; Klaus Reifenrath
>     > Subject: Re: [Speechsc] START-OF-SPEECH in DTMF-only mode ->
>     > summary(?)
>     >
>     > I think we're getting close. I though about snipping out
>     some pieces
>     > to cut down the text, but I realized the context is
>     still needed. See
>     > inline.
>     >
>     > On Jul 6, 2005, at 10:29 AM, Dave Burke wrote:
>     >
>     > > Inline.
>     > >
>     > > Dave
>     > >
>     > > ----- Original Message ----- From: "David R Oran"
>     <oran@cisco.com>
>     > > To: "Dave Burke" <david.burke@voxpilot.com>
>     > > Cc: <speechsc@ietf.org>; "Shanmugham, Saravanan"
>     <sarvi@cisco.com>;
>     > > "Klaus Reifenrath" <Klaus.Reifenrath@Scansoft.com>
>     > > Sent: Wednesday, July 06, 2005 1:07 PM
>     > > Subject: Re: [Speechsc] START-OF-SPEECH in DTMF-only
>     mode -> summary
>     > > (?)
>     > >
>     > >
>     > >
>     > >>
>     > >> On Jul 6, 2005, at 5:31 AM, Dave Burke wrote:
>     > >>
>     > >>
>     > >>> + Attempting to summarise:
>     > >>>
>     > >>> 1. START-OF-SPEECH is useful for the client to know when
>     > to stop
>     > >>> playing prompts the in non-optimised case 2.
>     START-OF-SPEECH is
>     > >>> useful for the client to calculate the bargin  time
>     (e.g. VoiceXML
>     > >>> 2.1 <mark>)
>     > >>> 3. In the optimised case, a bargin automatically
>     stops prompt
>     > >>> playing (assuming prompts barginable) 4. Because of
>     the previous
>     > >>> point, the question of what input type  caused
>     bargin is different
>     > >>> and less important to
>     > what input
>     > >>> type(s)  the recogniser is listening for
>     > >>>
>     > >> I'm not sure I follow point 4. Could you elaborate?
>     > >>
>     > >
>     > > DB> Adding a parameter to the START-OF-SPEECH event would
>     > certainly
>     > > allow the client to ignore (i.e. let prompts continue
>     playing) the
>     > > event if the event type is not of interest (e.g. the
>     client would
>     > > ignore speech start events when it is interested only
>     in  DTMF start
>     > > events). This _only_ works for the non-optimised case,
>     however. For
>     > > the optimised case, assuming START-OF-SPEECH
>     > coincides
>     > > with the bargin signal to the speechsynth, prompts
>     will stop playing
>     > > for inputs that the client might not be interested
>     > in (e.g.
>     > > a speech input will stop prompts playing even if the client
>     > is only
>     > > interested in DTMF).
>     > >
>     > OK, now I get it. There's a need for the client to both
>     handle the
>     > non-optimized case itself, and influence or at least
>     have a clue what
>     > the server is going to do in the optimized case.
>     >
>     > >
>     > >>
>     > >>
>     > >>> 5. Currently, MRCPv2 has no way of indicating what
>     input type(s) a
>     > >>> recogniser is listening for
>     > >>>
>     > >>>
>     > >> Do you mean exactly this, or do you mean "for the client to
>     > >> indicate  to the resource what input types it should
>     look for"?
>     > >>
>     > >
>     > > DB> Yes exactly - apologies for not being clear.
>     > >
>     > >
>     > >>
>     > >>
>     > >>> + Why implement 5?
>     > >>>
>     > >>> i. Noisy case: Need DTMF-only recognition (and may
>     only have a
>     > >>> speechrecog)
>     > >>>
>     > >> I'm having difficulty following the logic of why
>     noise would
>     > >> necessarily trigger START-OF-SPEECH if you were listening for
>     > >> speech  but not DTMF. I suppose you can use a more forgiving
>     > >> discriminator if  all you need to tell is if you're
>     getting DTMF,
>     > >> but I've had a number  of real-world cases where wind
>     noise was
>     > >> detected as DTMF, and  there's always the ambiguity
>     when you have
>     > >> Captain Crunch on the  phone. In either case in the
>     non-optimized
>     > >> case it's the client who  gets to decide whether an
>     event should be
>     > >> interpreted as barge-in or  not, so it seems an
>     aesthetic protocol
>     > >> design decision whether the  client tells the server
>     ahead of time
>     > >> what circumstances to generate  the START-OF-SPEECH
>     event for, or
>     > >> whether the event gets generated  and the client
>     decides based on
>     > >> what's in the event whether it should  be
>     > treated
>     > >> as barge- in.
>     > >>
>     > >> I suppose one could make the argument that because
>     the spec implies
>     > >> that the event can only be generated once per request
>     that if a
>     > >> DTMF/ speech capable recognizer first hears
>     > enough noise
>     > >> to think it's  hearing speech and later hears DTMF,
>     the client will
>     > >> declare barge-in  when the event comes and not when he DTMF
>     > >> actually gets heard.
>     > >>
>     > >> If that's deemed a problem, we can still handle that
>     in the design
>     > >> where the server just reports what it's hearing by allowing
>     > >> multiple  events to be generated during a single request.
>     > >>
>     > >> Between the approach just outlined above, and an
>     approach where the
>     > >> client provides a filter for whether to generate the
>     event or not,
>     > >> I  have a mild preference (based on aesthetics rather
>     than some
>     > >> hard  engineering tradeoff) for the approach where
>     > the server
>     > >> just reports  what it's hearing.
>     > >>
>     > >>
>     > >> Having had some useful exchanges on this topic, it
>     also is becoming
>     > >> apparent to me that this event is poorly named, and
>     we should
>     > >> consider renaming it to "INTERESTING-INPUT-HEARD" or
>     something akin
>     > >> to that, because as others have pointed out, a
>     DTMF-only recognizer
>     > >> will never detect "start of speech".
>     > >>
>     > >> Another consideration to fold into the design choice is
>     > >> extensibility. Bear with me through a little
>     gedankenexperiment.
>     > >>
>     > >> Suppose we want to define a new recognizer type,
>     which I'll call
>     > >> the "name that tune" recognizer. The client plays
>     music to the
>     > >> server and  the server recognizes musical notes. The
>     grammar is a
>     > >> standard  musical notation, augmented with a semantic
>     > >> interpretation that  transforms the notes into the
>     title of the
>     > >> tune and provides that as  an answer.
>     > >>
>     > >> First, there's no speech involved (or is there...hang
>     on a minute).
>     > >> Second, in order to accommodate the "name that tune"
>     > >> recognizer, we'd have to extend both the client and
>     the server to
>     > >> undetstand a  directive as to whether to recognize
>     music or now,
>     > >> inaddition to what  the server already knows what to
>     do based on
>     > >> the grammar. If you  follow my logic above, whether
>     or not we do
>     > >> that, we have to extend  "start-of-speech" to say
>     "I'm hearing
>     > >> music". So far fairly  straightforward, but let me
>     now throw in the
>     > >> pathological twist.
>     > >>
>     > >> Suppose what I feed to a  combined music/speech
>     recognizer is a
>     > >> work  in sprechstimme (spoken music), like the
>     "Geographical Fugue"
>     > >> (aside:  this is a wonderful piece of music I highly
>     recommend to
>     > >> anyone  interested in small ensemble singing). In
>     this case, the
>     > >> tune could  be named by either doing speech or music
>     recognition.
>     > >> Why is there  any need for the client to constrain
>     the server as to
>     > >> which it tries  to do when it's
>     > already
>     > >> told the server what it wants through the  grammar?
>     > >>
>     > >> A few other comments below
>     > >>
>     > >>
>     > >>> ii. Flexibility: Want speech-only recognition
>     (because a second
>     > >>> recogniser is doing hotword on DTMF)
>     > >>>
>     > >>>
>     > >> I don't see how flexibility is affected by this deisgn
>     > choice. If
>     > >> that's what you want, feed the speech-only recognizer
>     a grammar
>     > >> without any DTMF rules.
>     > >>
>     > >>
>     > >>> + How to implement 5?
>     > >>>
>     > >>> a. Implicitly:
>     > >>>    - dtmfrecog: always DTMF-only recognition
>     > >>>    - speechrecog: depends on active grammar type
>     > >>>        > if a dtmf grammar is active then DTMF input is "on"
>     > >>>        > if a speech grammar is active then speech
>     input is "on"
>     > >>>
>     > >>> b. Explicitly:
>     > >>>    - Add inputmodes header to RECOGNIZE
>     > >>>
>     > >>> Option a is David's "do what I mean case"; option b is
>     > the extra
>     > >>> dial for the client.
>     > >>>
>     > >>>
>     > >> Actually, that's not the point I was making with "do
>     what I mean",
>     > >> but it's not essential to the discussion so let's move on.
>     > >>
>     > >>
>     > >>> It is worth noting that VoiceXML uses option b. This
>     allows one to
>     > >>> activate both speech grammars and DTMF grammars (and
>     > therefore
>     > >>> be informed of any errors in the grammars at activation
>     > time) but
>     > >>> independently turn on whichever input mode you like e.g.
>     > perhaps
>     > >>> start with "both" then change to "dtmf".
>     > >>>
>     > >>>
>     > >> I'm not sure the VXML precedent is relevant here,
>     because the
>     > >> application behind VXML is working a different part of the
>     > problem
>     > >> -  how to traverse a TUI dialog based on different
>     parts of the
>     > >> input  space. In fact, I suspect that the VXML: choice was
>     > >> conditioned more  by limitations at the time it was
>     > specified than
>     > >> an underlying good  design choice. Clearly having to
>     specify this
>     > >> in VXML make the job of  handling a TUI with nodes
>     like "Say or
>     > >> press 5" harder rather than  easier.
>     > >>
>     > >
>     > > DB> The VoiceXML edge-case is pretty weird so it's not a major
>     > > concern. My main concern is that the client can
>     indicate, somehow,
>     > > what the input modes are.
>     > >
>     > >
>     > >>
>     > >> Summing up, while I don't feel strongly one way or
>     the other, I
>     > >> have  a preference for handling this as follows:
>     > >>
>     > >> a) Rename "START-OF-SPEECH" to
>     "INTERESTING-INPUT-RECEIVED" or
>     > >> something equivalent.
>     > >> b) Include a parameter in the event saying what was
>     interesting
>     > >> about  the input you received, with a registry of values  
> which
>     > >> includes:
>     > >>     - signal above noise floor
>     > >>     - speech
>     > >>     - dtmf
>     > >>     - (possibly) music
>     > >> c) allow the event to be generated multiple times
>     during a request
>     > >>
>     > >>
>     > >
>     > > DB> I like these suggestions (START-OF-INPUT?).
>     However, I don't
>     > > see how the problem of the optimised case is not
>     solved by them. I
>     > > think the optimised case is fine if we have the
>     following rules:
>     > >
>     > Yes, I hadn't thought through the optimized case as
>     thoroughly as you.
>     > Your suggested method name is fine by me as well
>     >
>     > > 1. START-OF-SPEECH (and optimised bargin) is only generated
>     > for the
>     > > input type that is being listened for 2. A speechrecog
>     listens for
>     > > DTMF if DTMF grammars are active, speech if speech
>     grammars are
>     > > active, or speech and DTMF if both grammar types are active.
>     > >
>     > Works for me.
>     >
>     > >
>     > >> Note that all of the above I'm saying with my technical
>     > hat on and
>     > >> my chair hat off.
>     > >> Putting my chair hat on for a moment, we really need
>     to get this
>     > >> spec  to last call, so at some point Eric or I is going to
>     > declare
>     > >> rough  consensus so we can move on.
>     > >>
>     > >
>     > > DB> Agreed!
>     > >
>     > >
>     > >>
>     > >> Dave Oran.
>     > >>
>     > >>> Dave
>     > >>>
>     > >>>
>     > >>>
>     > >>>
>     > >>>
>     > >>>
>     > >>>
>     > >>>
>     > >>>
>     > >>>
>     > >>>
>     > >>>
>     > >>> Sarvi makes a good point that adding the reason why the
>     > START-OF-
>     > >>> SPEECH occurred does not fix the optimised bargin case.
>     > >>>
>     > >>> dtmfrecog - listens for DTMF only
>     > >>> speechrecog - listens for DTMF only, or speech only,
>     or speech &
>     > >>> DTMF
>     > >>>
>     > >>>
>     > >>>
>     > >>>
>     > >>>
>     > >>>
>     > >>>
>     > >>>
>     > >>> ----- Original Message ----- From: "Shanmugham, Saravanan"
>     > >>> <sarvi@cisco.com>
>     > >>> To: "David R Oran" <oran@cisco.com>; "Dave Burke"
>     > >>> <david.burke@voxpilot.com>
>     > >>> Cc: <speechsc@ietf.org>; "Klaus Reifenrath"
>     > >>> <Klaus.Reifenrath@Scansoft.com>
>     > >>> Sent: Tuesday, July 05, 2005 8:51 PM
>     > >>> Subject: RE: [Speechsc] START-OF-SPEECH in DTMF-only mode
>     > >>>
>     > >>>
>     > >>> inline.
>     > >>>
>     > >>>     -----Original Message-----
>     > >>>     From: David R Oran [mailto:oran@cisco.com]
>     > >>>     Sent: Tuesday, July 05, 2005 11:08 AM
>     > >>>     To: Dave Burke
>     > >>>     Cc: Shanmugham, Saravanan; Klaus Reifenrath;
>     speechsc@ietf.org
>     > >>>     Subject: Re: [Speechsc] START-OF-SPEECH in DTMF-only  
> mode
>     > >>>
>     > >>>
>     > >>>     On Jul 5, 2005, at 12:54 PM, Dave Burke wrote:
>     > >>>
>     > >>>     > Inline.
>     > >>>     >
>     > >>>     > Dave
>     > >>>     >
>     > >>>     > ----- Original Message ----- From:
>     "Shanmugham, Saravanan"
>     > >>>     > <sarvi@cisco.com>
>     > >>>     > To: "David R Oran" <oran@cisco.com>; "Klaus  
> Reifenrath"
>     > >>>     > <Klaus.Reifenrath@Scansoft.com>
>     > >>>     > Cc: <speechsc@ietf.org>
>     > >>>     > Sent: Tuesday, July 05, 2005 5:22 PM
>     > >>>     > Subject: RE: [Speechsc] START-OF-SPEECH in
>     DTMF-only mode
>     > >>>     >
>     > >>>     >
>     > >>>     > I agree with Dave's analysis. The purpose of this  
> event
>     > >>>     was barge-in.
>     > >>>     > And barge-in should happen for both DTMF and speech.
>     > >>>     >
>     > >>>     > Is there a case where you think it should not behave
>     > >>>     this way. If soe,
>     > >>>     > please provide a scenario where you think
>     > >>>     >    1. Barge-in should happen for DTMF and not voice or
>     > >>>     vice-versa.
>     > >>>     >
>     > >>>     > DB> You want to do a DTMF recognition only
>     because it is
>     > >>> noisy.
>     > >>>     > While waiting for DTMF input, the speechrecog resource
>     > >>>     (or advanced
>     > >>>     > dtmfrecog) generates a START-OF-SPEECH because it  
> heard
>     > >>>     some speech.
>     > >>>     > The client does not want to stop prompt playing unless
>     > >>>     DTMF was heard
>     > >>>     > but it can't tell by the START-OF-SPEECH whether  
> speech
>     > >>>     or DTMF was
>     > >>>     > heard. Similarly vice versa.
>     > >>>     >
>     > >>>     It's an interesting design question what part of the
>     > >>>     policy resides at the client and what at the server, and
>     > >>>     who makes the "final decision" about whether what was
>     > >>>     heard was relevant to the control channel. Right now we
>     > >>>     (IMO) have a weird partitioning in many cases where the
>     > >>>     client basically says "do what I mean", but there are no
>     > >>>     constraints of what the server actually does, and no
>     > >>>     normalized basis for the client to figure out what to  
> set
>     > >>>     various magic numbers to (e.g. sensitivity).
>     > >>>
>     > >>>     In this case the only thing the client needs to  
> decide is
>     > >>>     whether to kill the prompt because the server thinks
>     > >>>     something that would interfere with the feedback
>     > >>>     ear/mouth/finger control happened. What this
>     says to me is
>     > >>>     that it isn't necessarily a good idea for the client to
>     > >>>     have more knobs to control the server
>     (especially if those
>     > >>>     knows are just more value/policy input ungrounded in any
>     > >>>     physics/ acoustics). On the other hand, having the  
> server
>     > >>>     tell the client more about what it thinks is going on is
>     > >>>     probably valuable.
>     > >>>
>     > >>>     So, Coming to the point after this long rambling
>     > >>>     introduction, I think it would in fact be useful for the
>     > >>>     START-Of-SPEECH event to indicate some extra  
> information,
>     > >>>     for example:
>     > >>>     a) I got something enough above the noise floor
>     to qualify
>     > >>>     for exceeding the "Sensisitvity" parameter you sent  
> in on
>     > >>>     the request but I really can't tell what it is
>     (could be a
>     > >>>     hippopatmus fart, or a siren in the background,
>     or captain
>     > >>>     crunch trying to whistle DTMF).
>     > >>>     b) I think I'm hearing speech
>     > >>>     c) I think I'm hearing DTMF
>     > >>>
>     > >>> Though I agree with your former part of your
>     response. I am not
>     > >>> sure I agree with your proposed solution.
>     > >>> The way I see this problem is that, it is more of what
>     > constitues a
>     > >>> barge-in event. This boils down to whether it is
>     speech, DTMF or
>     > >>> both.
>     > >>> This is inturn boils down to what type of recognizer
>     > resource we are
>     > >>> using, dtmf-recog, speech-recog, and speech-only-recog(we
>     > don't have
>     > >>> this and I don't think we should add it, but think
>     of this as a
>     > >>> place holder that explains the concept).
>     > >>>
>     > >>> A client knowing what type of barge-in happenned, does
>     > not impact
>     > >>> the
>     > >>> barge-in operation itself as it may be too late(for
>     the optimized
>     > >>> barge-in case). It may have other use cases, and if we
>     > can identify
>     > >>> them, I don't mind adding support for the
>     START-OF-SPEECH event to
>     > >>> say what type of barge-in happenned. But that itself does  
> not
>     > solve the
>     > >>> original problem raised. Refer to my previous response.
>     > >>>
>     > >>> The solution lies in defining what what is a
>     barge-in event.
>     > >>> That  boils
>     > >>> down to what type of recognition is happenning,
>     > dtmf-only, speech-
>     > >>> dtmf
>     > >>> or speech-only. We do not support speech-only as a
>     resource today,
>     > >>> the question is do we need a header to force it.
>     > >>>
>     > >>> Sarvi
>     > >>>
>     > >>>
>     > >>>     >    2. You would benefit from the client knowing
>     > what caused
>     > >>> the
>     > >>>     > barge-in, DTMF Vs speech.
>     > >>>     >
>     > >>>     > DB> See previous comment. And previous e-mail:
>     either add an
>     > >>>     > inputmodes header (taking value speech, dtmf, both) to
>     > >>>     the RECOGNIZE
>     > >>>     > request or add a header to the START-OF-SPEECH event
>     > >>>     indicating DTMF
>     > >>>     > or speech.
>     > >>>     >
>     > >>>     I'm leaning in your direction on this latter point - as
>     > >>>     should be evident from what I wrote above.
>     > >>>
>     > >>>     > Sarvi
>     > >>>     >
>     > >>>     >     -----Original Message-----
>     > >>>     >     From: speechsc-bounces@ietf.org
>     > >>>     >     [mailto:speechsc-bounces@ietf.org] On Behalf Of
>     > David R
>     > >>> Oran
>     > >>>     >     Sent: Tuesday, July 05, 2005 5:24 AM
>     > >>>     >     To: Klaus Reifenrath
>     > >>>     >     Cc: 'speechsc@ietf.org'
>     > >>>     >     Subject: Re: [Speechsc] START-OF-SPEECH in
>     > DTMF-only mode
>     > >>>     >
>     > >>>     >
>     > >>>     >     On Jul 5, 2005, at 3:46 AM, Reifenrath,
>     Klaus wrote:
>     > >>>     >
>     > >>>     >     > The current spec is not clear when
>     > START-OF-SPEECH need
>     > >>>     >     to be send in
>     > >>>     >     > the following scenarios:
>     > >>>     >     > A) The client requested a DTMF Recognizer. Is  
> the
>     > >>>     >     START-OF-SPEECH
>     > >>>     >     > event send to the client also if speech
>     was detected?
>     > >>>     >     I suspect so, since one of the prime purposes
>     > is to enable
>     > >>>     >     client- mediated barge-in handling. However, if  
> the
>     > >>>     >     recognizer is in fact only capable of
>     recognizing DTMF
>     > >>>     >     then it may in fact not report anythin
>     unless it's using
>     > >>>     >     some primitive thresholding machinery, like a SN
>     > >>> threshold.
>     > >>>     >     > B) The client requested a Speech
>     Recognizer, but only
>     > >>>     >     activated DTMF
>     > >>>     >     > grammars. Is the START-OF-SPEECH event
>     send to the
>     > >>>     >     client also if
>     > >>>     >     > speech was detected?
>     > >>>     >     Again, I'd say yes, for the same reason as above.
>     > >>>     >     > I think in both cases START-OF-SPEECH should  
> only
>     > >>>     be send after
>     > >>>     >     > detecting a DTMF digit (see Figure 12 of  
> VoiceXML
>     > >>>     2.0: http://
>     > >>>     >     > www.w3.org/TR/voicexml20/#dmlATiming).
>     > >>>     >     We seem to have reached different
>     conclusions. I'd be
>     > >>>     >     interested in why you think my analysis
>     above is wrong.
>     > >>>     >
>     > >>>     >     Dave.
>     > >>>     >
>     > >>>     >     > Klaus
>     > >>>     >     >
>     > >>>     >     > _______________________________________________
>     > >>>     >     > Speechsc mailing list
>     > >>>     >     > Speechsc@ietf.org
>     > >>>     >     > https://www1.ietf.org/mailman/listinfo/speechsc
>     > >>>     >     >
>     > >>>     >
>     > >>>     >     _______________________________________________
>     > >>>     >     Speechsc mailing list
>     > >>>     >     Speechsc@ietf.org
>     > >>>     >     https://www1.ietf.org/mailman/listinfo/speechsc
>     > >>>     >
>     > >>>     >
>     > >>>     > _______________________________________________
>     > >>>     > Speechsc mailing list
>     > >>>     > Speechsc@ietf.org
>     > >>>     > https://www1.ietf.org/mailman/listinfo/speechsc
>     > >>>     >
>     > >>>
>     > >>>
>     > >>> _______________________________________________
>     > >>> Speechsc mailing list
>     > >>> Speechsc@ietf.org
>     > >>> https://www1.ietf.org/mailman/listinfo/speechsc
>     > >>>
>     > >>>
>     > >>
>     > >> _______________________________________________
>     > >> Speechsc mailing list
>     > >> Speechsc@ietf.org
>     > >> https://www1.ietf.org/mailman/listinfo/speechsc
>     > >
>     >
>     > _______________________________________________
>     > Speechsc mailing list
>     > Speechsc@ietf.org
>     > https://www1.ietf.org/mailman/listinfo/speechsc
>     >
>
>
>     _______________________________________________
>     Speechsc mailing list
>     Speechsc@ietf.org
>     https://www1.ietf.org/mailman/listinfo/speechsc
>
>
> _______________________________________________
> Speechsc mailing list
> Speechsc@ietf.org
> https://www1.ietf.org/mailman/listinfo/speechsc
>
>
> _______________________________________________
> Speechsc mailing list
> Speechsc@ietf.org
> https://www1.ietf.org/mailman/listinfo/speechsc
>

_______________________________________________
Speechsc mailing list
Speechsc@ietf.org
https://www1.ietf.org/mailman/listinfo/speechsc



From speechsc-bounces@ietf.org Fri Jul 08 11:36:54 2005
Received: from localhost.localdomain ([127.0.0.1] helo=megatron.ietf.org)
	by megatron.ietf.org with esmtp (Exim 4.32)
	id 1DquuM-0004ge-9C; Fri, 08 Jul 2005 11:36:54 -0400
Received: from odin.ietf.org ([132.151.1.176] helo=ietf.org)
	by megatron.ietf.org with esmtp (Exim 4.32) id 1DquuJ-0004e9-9B
	for speechsc@megatron.ietf.org; Fri, 08 Jul 2005 11:36:52 -0400
Received: from ietf-mx.ietf.org (ietf-mx [132.151.6.1])
	by ietf.org (8.9.1a/8.9.1a) with ESMTP id LAA21941
	for <speechsc@ietf.org>; Fri, 8 Jul 2005 11:36:49 -0400 (EDT)
Received: from letter.nuance.com ([207.107.210.132])
	by ietf-mx.ietf.org with esmtp (Exim 4.43) id 1DqvLi-0001e8-Bs
	for speechsc@ietf.org; Fri, 08 Jul 2005 12:05:11 -0400
Received: from postcard.nuance.com ([10.3.6.20]:9775)
	by letter.nuance.com with esmtp id 1Dquto-0005xj-AA;
	Fri, 08 Jul 2005 08:36:20 -0700
Received: from mtb1exch01.nuance.com ([10.3.2.6]) by postcard.nuance.com with
	Microsoft SMTPSVC(6.0.3790.0); Fri, 8 Jul 2005 11:36:19 -0400
X-MimeOLE: Produced By Microsoft Exchange V6.5.7226.0
Content-class: urn:content-classes:message
MIME-Version: 1.0
Content-Type: text/plain;
	charset="us-ascii"
Content-Transfer-Encoding: quoted-printable
Subject: RE: [Speechsc] START-OF-SPEECH in DTMF-only mode -> summary(?)
Date: Fri, 8 Jul 2005 11:36:13 -0400
Message-ID: <7DE7C4EF3B7C8B4B82955191378290D802ED4176@mtb1exch01.nuance.com>
Thread-Topic: [Speechsc] START-OF-SPEECH in DTMF-only mode -> summary(?)
Thread-Index: AcWDxnsqacO68WIQRraPdgMIBhokzAACzz6A
From: "Pierre Forgues" <forgues@nuance.com>
To: "David R Oran" <oran@cisco.com>, "Dave Burke" <david.burke@voxpilot.com>
X-OriginalArrivalTime: 08 Jul 2005 15:36:19.0052 (UTC)
	FILETIME=[C91E8AC0:01C583D2]
X-FromHost: postcard.nuance.com [10.3.6.20]:9775
Lines: 860
X-Spam-Score: 0.0 (/)
X-Scan-Signature: d7d3f294c642eb492c2d44336474db48
Content-Transfer-Encoding: quoted-printable
Cc: speechsc@ietf.org, "Shanmugham, Saravanan" <sarvi@cisco.com>,
	Eric Burger <eburger@brooktrout.com>
X-BeenThere: speechsc@ietf.org
X-Mailman-Version: 2.1.5
Precedence: list
List-Id: Speech Services Control Working Group <speechsc.ietf.org>
List-Unsubscribe: <https://www1.ietf.org/mailman/listinfo/speechsc>,
	<mailto:speechsc-request@ietf.org?subject=unsubscribe>
List-Post: <mailto:speechsc@ietf.org>
List-Help: <mailto:speechsc-request@ietf.org?subject=help>
List-Subscribe: <https://www1.ietf.org/mailman/listinfo/speechsc>,
	<mailto:speechsc-request@ietf.org?subject=subscribe>
Sender: speechsc-bounces@ietf.org
Errors-To: speechsc-bounces@ietf.org



-----Original Message-----
From: speechsc-bounces@ietf.org [mailto:speechsc-bounces@ietf.org] On
Behalf Of David R Oran
Sent: Friday, July 08, 2005 10:05 AM
To: Dave Burke
Cc: speechsc@ietf.org; Shanmugham, Saravanan; Eric Burger
Subject: Re: [Speechsc] START-OF-SPEECH in DTMF-only mode -> summary(?)


On Jul 8, 2005, at 6:59 AM, Dave Burke wrote:

> I agree with your opinion on what constitutes a barge-in. Based on =20
> Klaus' e-mail, I am more concerned that using the grammar type to =20
> decide what constitutes a barge-in is going to make VoiceXML =20
> implementations difficult.
>
> It is easy to map VoiceXML application selected inputmodes to MRCP =20
> resources:
> a. inputmodes=3D"dtmf" -> Use a dtmfrecog
> b. inputmodes=3D"speech" -> Use a "speech-only-recog"
> c. inputmodes=3D"both" -> Use a speechrecog (or a combination of =20
> "speech-only-recog" + dtmfrecog)
>
> The only problem is (b). While not as important as being able to do =20
> DTMF-only recognition, I believe we DO need to support speech-only =20
> recongition so as (a) to avoid unnecessary limitations in VUI =20
> design, and (b) to facilitate VoiceXML implementations. I think we =20
> should add a header to RECOGNIZE so the client is able to always =20
> _explicitly_ set the inputmodes.
>
I'm not sure I buy this, since what the recognizer is looking for and =20
what constitutes an input that should be considered a potential barge-=20
in strike me as independent. I'm similarly not persuaded that we need =20
the flexibility to set the input mode independently of the grammar, =20
since it leads to all sorts of inconsistent states (e.g. speech-only =20
grammar with an input-mode of dtmf). On the other hand I can see the =20
need for the client to specify what sorts of input out to generate =20
the start-of-input (nee start-of-speech) event on.

Pmf> The mechanism for a client to specify the type of input is using
grammars.  These can have DTMF and/or speech requirements.  If you add
the complexity of an independent header to specify the input mode then
you will need to document the behavior for inconsistencies.

> In retrospect, I don't like the idea of multiple START-OF-SPEECH =20
> events being generated from the same media resource. This is =20
> because one assumes that the START-OF-SPEECH should be of the same =20
> type as the hypothesis returned in the RECOGNITION-COMPETE message =20
> - most implementations, on hearing one input mode type disable the =20
> recogniser of the other type.
I'm not sure I follow this logic, but I'm not wedded to the idea of =20
allowing multiple events. It was trying to solve the problem of =20
ambiguity around what the recognizer was hearing and the idea that =20
the client might care. If you don't think the client will ever care, =20
then we don't need the capability.
Pmf> I agree we should not have multiple SOS events.  Too much chatter.


> I do like David's idea of renaming START-OF-SPEECH to something =20
> like START-OF-INPUT and carrying a type header because it is neater =20
> and more extensible.
>
> So in summary, I propose we modify the spec to:
>
> 1. Clarify what constitutes a barge-in for a dtmfrecog and =20
> speechrecog as per Sarvi's e-mail (and in agreement with Klaus' for =20
> dtmfrecog).
>
Hmmm, ok, but I think the issue is actually clarifying what the =20
recognizer declares as "interesting input", which it may decide also =20
constitutes a barge-in in the optimized case, and in either case =20
reports to the client that it heard.
Pmf> I'm not going to argue changing the name of the event, but my
preference would be to keep existing method names unless there is a
clear reason for changing - which I do not see here.

> 2. Specifiy an InputModes header to RECOGNIZE (defaults to "both", =20
> can also be "speech" or "DTMF"). Setting to speech for a =20
> speechrecog results in the hypothesised "speech-only-recog". =20
> Setting to DTMF for a speechrecog is equivalent to using a =20
> dtmfrecog. Edge cases: Setting to speech for a dtmfrecog will =20
> result in a noinput as will setting to DTMF for a speechrecog which =20
> does not support DTMF.
>
I'm ok with having a header for the client to tell the server when it =20
would like the start-of-input event to be generated and what the =20
client considers to be the "interesting input" that the recognizer =20
should use to do the discrimination, and possibly do barge-in =20
processing in the optimized case.

I'm less ok with the proposed domain of values, since it will have =20
extensibility problems. Especially problematical is have a code point =20
of "both" since that will be ambiguous if we even define a "music" =20
recognizer, or a "gesture" recognizer using video input. Here's my =20
counter-proposal:

Create a header on the recognizer resource requests called "start-=20
input-on:". Define a registry of values, with the following three =20
values initially defined:
     - dtmf
     - speech
     - music

The semantics would be that the resource is to generate a "start-of-=20
input" event, apply the defined grammars to what follows, and do any =20
local optimized barge-in processing if ANY of the enumerated input =20
types is detected. That way you can say
     "start-input-on: dtmf" if you want to just  dtmf,
     "start-input-on: speech" if you want just speech (dtmf input =20
would be ignored even if the grammar supported dtmf)
     "start-input- on: speech, dtmf" if you wanted both
etc.

Pmf> Doesn't this come back to the previous point of having two
independent ways to specify input modes?  Through grammars and this
proposed new header?  If you add this then you will have to document
behavior on inconsistencies.  Anyway, if you proceed then I would
definitely make this optional

> 3. Change START-OF-SPEECH to START-OF-INPUT. Clarify it occurs once =20
> and coincides with barge-in.
>
Ok for the "only once", but I'd like to tighten it up to talk about =20
more than just barge-in, as I suggested above.

> 4. Add a header of InputType to START-OF-INPUT. Current specified =20
> values are "dtmf" or "speech".
>
Ok, with slight modification. Say that the syntax of the "input-type" =20
header is a single-valued subset of the registered value(s) that are =20
defined for the start-input-on: header.

Comments?

Dave O. (technical hat on, chair hat off).

> Dave
>
>
>
> ----- Original Message ----- From: "Shanmugham, Saravanan" =20
> <sarvi@cisco.com>
> To: "Eric Burger" <eburger@brooktrout.com>; <speechsc@ietf.org>
> Sent: Thursday, July 07, 2005 5:52 PM
> Subject: RE: [Speechsc] START-OF-SPEECH in DTMF-only mode -> summary=20
> (?)
>
>
> I don't think generating multiple START-OF-SPEECH events is a =20
> solution.
> We still haven't addressed, what constitutes a barge-in, for the
> optimised case. That should also be the single point when a single
> START-OF-SPEECH(or whathever else you want to name it) should be
> generated.
>
> That point, in my opinion should be
>   1. For "dtmf-recog" resources should be the beginning of a DTMF key
> press.
>   2. For "speech-recog" resources should be the beginning of a DTMF =20
> key
> press or the beginning of speech. This is should be irrespective of =20
> what
> type of grammar is being used. Coz even numbers only grammars can =20
> still
> be spoken and hence cannot be assumed to be a cue for DTMF only
> recognition.
>   3. For "speech-only-recog" resources(which are not defined today) =20
> the
> time to barge-in is the beginning of speech. I don't see a need for =20
> such
> a resource today. But I am mentioning this for completeness.
>
> Sarvi
>
>
>     -----Original Message-----
>     From: speechsc-bounces@ietf.org
>     [mailto:speechsc-bounces@ietf.org] On Behalf Of Eric Burger
>     Sent: Thursday, July 07, 2005 5:32 AM
>     To: speechsc@ietf.org
>     Subject: RE: [Speechsc] START-OF-SPEECH in DTMF-only mode
>     -> summary(?)
>
>     Can we declare consensus?
>
>     > -----Original Message-----
>     > From: speechsc-bounces@ietf.org
>     [mailto:speechsc-bounces@ietf.org]
>     > Sent: Wednesday, July 06, 2005 10:45 AM
>     > To: Dave Burke
>     > Cc: speechsc@ietf.org; Shanmugham, Saravanan; Klaus Reifenrath
>     > Subject: Re: [Speechsc] START-OF-SPEECH in DTMF-only mode ->
>     > summary(?)
>     >
>     > I think we're getting close. I though about snipping out
>     some pieces
>     > to cut down the text, but I realized the context is
>     still needed. See
>     > inline.
>     >
>     > On Jul 6, 2005, at 10:29 AM, Dave Burke wrote:
>     >
>     > > Inline.
>     > >
>     > > Dave
>     > >
>     > > ----- Original Message ----- From: "David R Oran"
>     <oran@cisco.com>
>     > > To: "Dave Burke" <david.burke@voxpilot.com>
>     > > Cc: <speechsc@ietf.org>; "Shanmugham, Saravanan"
>     <sarvi@cisco.com>;
>     > > "Klaus Reifenrath" <Klaus.Reifenrath@Scansoft.com>
>     > > Sent: Wednesday, July 06, 2005 1:07 PM
>     > > Subject: Re: [Speechsc] START-OF-SPEECH in DTMF-only
>     mode -> summary
>     > > (?)
>     > >
>     > >
>     > >
>     > >>
>     > >> On Jul 6, 2005, at 5:31 AM, Dave Burke wrote:
>     > >>
>     > >>
>     > >>> + Attempting to summarise:
>     > >>>
>     > >>> 1. START-OF-SPEECH is useful for the client to know when
>     > to stop
>     > >>> playing prompts the in non-optimised case 2.
>     START-OF-SPEECH is
>     > >>> useful for the client to calculate the bargin  time
>     (e.g. VoiceXML
>     > >>> 2.1 <mark>)
>     > >>> 3. In the optimised case, a bargin automatically
>     stops prompt
>     > >>> playing (assuming prompts barginable) 4. Because of
>     the previous
>     > >>> point, the question of what input type  caused
>     bargin is different
>     > >>> and less important to
>     > what input
>     > >>> type(s)  the recogniser is listening for
>     > >>>
>     > >> I'm not sure I follow point 4. Could you elaborate?
>     > >>
>     > >
>     > > DB> Adding a parameter to the START-OF-SPEECH event would
>     > certainly
>     > > allow the client to ignore (i.e. let prompts continue
>     playing) the
>     > > event if the event type is not of interest (e.g. the
>     client would
>     > > ignore speech start events when it is interested only
>     in  DTMF start
>     > > events). This _only_ works for the non-optimised case,
>     however. For
>     > > the optimised case, assuming START-OF-SPEECH
>     > coincides
>     > > with the bargin signal to the speechsynth, prompts
>     will stop playing
>     > > for inputs that the client might not be interested
>     > in (e.g.
>     > > a speech input will stop prompts playing even if the client
>     > is only
>     > > interested in DTMF).
>     > >
>     > OK, now I get it. There's a need for the client to both
>     handle the
>     > non-optimized case itself, and influence or at least
>     have a clue what
>     > the server is going to do in the optimized case.
>     >
>     > >
>     > >>
>     > >>
>     > >>> 5. Currently, MRCPv2 has no way of indicating what
>     input type(s) a
>     > >>> recogniser is listening for
>     > >>>
>     > >>>
>     > >> Do you mean exactly this, or do you mean "for the client to
>     > >> indicate  to the resource what input types it should
>     look for"?
>     > >>
>     > >
>     > > DB> Yes exactly - apologies for not being clear.
>     > >
>     > >
>     > >>
>     > >>
>     > >>> + Why implement 5?
>     > >>>
>     > >>> i. Noisy case: Need DTMF-only recognition (and may
>     only have a
>     > >>> speechrecog)
>     > >>>
>     > >> I'm having difficulty following the logic of why
>     noise would
>     > >> necessarily trigger START-OF-SPEECH if you were listening for
>     > >> speech  but not DTMF. I suppose you can use a more forgiving
>     > >> discriminator if  all you need to tell is if you're
>     getting DTMF,
>     > >> but I've had a number  of real-world cases where wind
>     noise was
>     > >> detected as DTMF, and  there's always the ambiguity
>     when you have
>     > >> Captain Crunch on the  phone. In either case in the
>     non-optimized
>     > >> case it's the client who  gets to decide whether an
>     event should be
>     > >> interpreted as barge-in or  not, so it seems an
>     aesthetic protocol
>     > >> design decision whether the  client tells the server
>     ahead of time
>     > >> what circumstances to generate  the START-OF-SPEECH
>     event for, or
>     > >> whether the event gets generated  and the client
>     decides based on
>     > >> what's in the event whether it should  be
>     > treated
>     > >> as barge- in.
>     > >>
>     > >> I suppose one could make the argument that because
>     the spec implies
>     > >> that the event can only be generated once per request
>     that if a
>     > >> DTMF/ speech capable recognizer first hears
>     > enough noise
>     > >> to think it's  hearing speech and later hears DTMF,
>     the client will
>     > >> declare barge-in  when the event comes and not when he DTMF
>     > >> actually gets heard.
>     > >>
>     > >> If that's deemed a problem, we can still handle that
>     in the design
>     > >> where the server just reports what it's hearing by allowing
>     > >> multiple  events to be generated during a single request.
>     > >>
>     > >> Between the approach just outlined above, and an
>     approach where the
>     > >> client provides a filter for whether to generate the
>     event or not,
>     > >> I  have a mild preference (based on aesthetics rather
>     than some
>     > >> hard  engineering tradeoff) for the approach where
>     > the server
>     > >> just reports  what it's hearing.
>     > >>
>     > >>
>     > >> Having had some useful exchanges on this topic, it
>     also is becoming
>     > >> apparent to me that this event is poorly named, and
>     we should
>     > >> consider renaming it to "INTERESTING-INPUT-HEARD" or
>     something akin
>     > >> to that, because as others have pointed out, a
>     DTMF-only recognizer
>     > >> will never detect "start of speech".
>     > >>
>     > >> Another consideration to fold into the design choice is
>     > >> extensibility. Bear with me through a little
>     gedankenexperiment.
>     > >>
>     > >> Suppose we want to define a new recognizer type,
>     which I'll call
>     > >> the "name that tune" recognizer. The client plays
>     music to the
>     > >> server and  the server recognizes musical notes. The
>     grammar is a
>     > >> standard  musical notation, augmented with a semantic
>     > >> interpretation that  transforms the notes into the
>     title of the
>     > >> tune and provides that as  an answer.
>     > >>
>     > >> First, there's no speech involved (or is there...hang
>     on a minute).
>     > >> Second, in order to accommodate the "name that tune"
>     > >> recognizer, we'd have to extend both the client and
>     the server to
>     > >> undetstand a  directive as to whether to recognize
>     music or now,
>     > >> inaddition to what  the server already knows what to
>     do based on
>     > >> the grammar. If you  follow my logic above, whether
>     or not we do
>     > >> that, we have to extend  "start-of-speech" to say
>     "I'm hearing
>     > >> music". So far fairly  straightforward, but let me
>     now throw in the
>     > >> pathological twist.
>     > >>
>     > >> Suppose what I feed to a  combined music/speech
>     recognizer is a
>     > >> work  in sprechstimme (spoken music), like the
>     "Geographical Fugue"
>     > >> (aside:  this is a wonderful piece of music I highly
>     recommend to
>     > >> anyone  interested in small ensemble singing). In
>     this case, the
>     > >> tune could  be named by either doing speech or music
>     recognition.
>     > >> Why is there  any need for the client to constrain
>     the server as to
>     > >> which it tries  to do when it's
>     > already
>     > >> told the server what it wants through the  grammar?
>     > >>
>     > >> A few other comments below
>     > >>
>     > >>
>     > >>> ii. Flexibility: Want speech-only recognition
>     (because a second
>     > >>> recogniser is doing hotword on DTMF)
>     > >>>
>     > >>>
>     > >> I don't see how flexibility is affected by this deisgn
>     > choice. If
>     > >> that's what you want, feed the speech-only recognizer
>     a grammar
>     > >> without any DTMF rules.
>     > >>
>     > >>
>     > >>> + How to implement 5?
>     > >>>
>     > >>> a. Implicitly:
>     > >>>    - dtmfrecog: always DTMF-only recognition
>     > >>>    - speechrecog: depends on active grammar type
>     > >>>        > if a dtmf grammar is active then DTMF input is "on"
>     > >>>        > if a speech grammar is active then speech
>     input is "on"
>     > >>>
>     > >>> b. Explicitly:
>     > >>>    - Add inputmodes header to RECOGNIZE
>     > >>>
>     > >>> Option a is David's "do what I mean case"; option b is
>     > the extra
>     > >>> dial for the client.
>     > >>>
>     > >>>
>     > >> Actually, that's not the point I was making with "do
>     what I mean",
>     > >> but it's not essential to the discussion so let's move on.
>     > >>
>     > >>
>     > >>> It is worth noting that VoiceXML uses option b. This
>     allows one to
>     > >>> activate both speech grammars and DTMF grammars (and
>     > therefore
>     > >>> be informed of any errors in the grammars at activation
>     > time) but
>     > >>> independently turn on whichever input mode you like e.g.
>     > perhaps
>     > >>> start with "both" then change to "dtmf".
>     > >>>
>     > >>>
>     > >> I'm not sure the VXML precedent is relevant here,
>     because the
>     > >> application behind VXML is working a different part of the
>     > problem
>     > >> -  how to traverse a TUI dialog based on different
>     parts of the
>     > >> input  space. In fact, I suspect that the VXML: choice was
>     > >> conditioned more  by limitations at the time it was
>     > specified than
>     > >> an underlying good  design choice. Clearly having to
>     specify this
>     > >> in VXML make the job of  handling a TUI with nodes
>     like "Say or
>     > >> press 5" harder rather than  easier.
>     > >>
>     > >
>     > > DB> The VoiceXML edge-case is pretty weird so it's not a major
>     > > concern. My main concern is that the client can
>     indicate, somehow,
>     > > what the input modes are.
>     > >
>     > >
>     > >>
>     > >> Summing up, while I don't feel strongly one way or
>     the other, I
>     > >> have  a preference for handling this as follows:
>     > >>
>     > >> a) Rename "START-OF-SPEECH" to
>     "INTERESTING-INPUT-RECEIVED" or
>     > >> something equivalent.
>     > >> b) Include a parameter in the event saying what was
>     interesting
>     > >> about  the input you received, with a registry of values =20
> which
>     > >> includes:
>     > >>     - signal above noise floor
>     > >>     - speech
>     > >>     - dtmf
>     > >>     - (possibly) music
>     > >> c) allow the event to be generated multiple times
>     during a request
>     > >>
>     > >>
>     > >
>     > > DB> I like these suggestions (START-OF-INPUT?).
>     However, I don't
>     > > see how the problem of the optimised case is not
>     solved by them. I
>     > > think the optimised case is fine if we have the
>     following rules:
>     > >
>     > Yes, I hadn't thought through the optimized case as
>     thoroughly as you.
>     > Your suggested method name is fine by me as well
>     >
>     > > 1. START-OF-SPEECH (and optimised bargin) is only generated
>     > for the
>     > > input type that is being listened for 2. A speechrecog
>     listens for
>     > > DTMF if DTMF grammars are active, speech if speech
>     grammars are
>     > > active, or speech and DTMF if both grammar types are active.
>     > >
>     > Works for me.
>     >
>     > >
>     > >> Note that all of the above I'm saying with my technical
>     > hat on and
>     > >> my chair hat off.
>     > >> Putting my chair hat on for a moment, we really need
>     to get this
>     > >> spec  to last call, so at some point Eric or I is going to
>     > declare
>     > >> rough  consensus so we can move on.
>     > >>
>     > >
>     > > DB> Agreed!
>     > >
>     > >
>     > >>
>     > >> Dave Oran.
>     > >>
>     > >>> Dave
>     > >>>
>     > >>>
>     > >>>
>     > >>>
>     > >>>
>     > >>>
>     > >>>
>     > >>>
>     > >>>
>     > >>>
>     > >>>
>     > >>>
>     > >>> Sarvi makes a good point that adding the reason why the
>     > START-OF-
>     > >>> SPEECH occurred does not fix the optimised bargin case.
>     > >>>
>     > >>> dtmfrecog - listens for DTMF only
>     > >>> speechrecog - listens for DTMF only, or speech only,
>     or speech &
>     > >>> DTMF
>     > >>>
>     > >>>
>     > >>>
>     > >>>
>     > >>>
>     > >>>
>     > >>>
>     > >>>
>     > >>> ----- Original Message ----- From: "Shanmugham, Saravanan"
>     > >>> <sarvi@cisco.com>
>     > >>> To: "David R Oran" <oran@cisco.com>; "Dave Burke"
>     > >>> <david.burke@voxpilot.com>
>     > >>> Cc: <speechsc@ietf.org>; "Klaus Reifenrath"
>     > >>> <Klaus.Reifenrath@Scansoft.com>
>     > >>> Sent: Tuesday, July 05, 2005 8:51 PM
>     > >>> Subject: RE: [Speechsc] START-OF-SPEECH in DTMF-only mode
>     > >>>
>     > >>>
>     > >>> inline.
>     > >>>
>     > >>>     -----Original Message-----
>     > >>>     From: David R Oran [mailto:oran@cisco.com]
>     > >>>     Sent: Tuesday, July 05, 2005 11:08 AM
>     > >>>     To: Dave Burke
>     > >>>     Cc: Shanmugham, Saravanan; Klaus Reifenrath;
>     speechsc@ietf.org
>     > >>>     Subject: Re: [Speechsc] START-OF-SPEECH in DTMF-only =20
> mode
>     > >>>
>     > >>>
>     > >>>     On Jul 5, 2005, at 12:54 PM, Dave Burke wrote:
>     > >>>
>     > >>>     > Inline.
>     > >>>     >
>     > >>>     > Dave
>     > >>>     >
>     > >>>     > ----- Original Message ----- From:
>     "Shanmugham, Saravanan"
>     > >>>     > <sarvi@cisco.com>
>     > >>>     > To: "David R Oran" <oran@cisco.com>; "Klaus =20
> Reifenrath"
>     > >>>     > <Klaus.Reifenrath@Scansoft.com>
>     > >>>     > Cc: <speechsc@ietf.org>
>     > >>>     > Sent: Tuesday, July 05, 2005 5:22 PM
>     > >>>     > Subject: RE: [Speechsc] START-OF-SPEECH in
>     DTMF-only mode
>     > >>>     >
>     > >>>     >
>     > >>>     > I agree with Dave's analysis. The purpose of this =20
> event
>     > >>>     was barge-in.
>     > >>>     > And barge-in should happen for both DTMF and speech.
>     > >>>     >
>     > >>>     > Is there a case where you think it should not behave
>     > >>>     this way. If soe,
>     > >>>     > please provide a scenario where you think
>     > >>>     >    1. Barge-in should happen for DTMF and not voice or
>     > >>>     vice-versa.
>     > >>>     >
>     > >>>     > DB> You want to do a DTMF recognition only
>     because it is
>     > >>> noisy.
>     > >>>     > While waiting for DTMF input, the speechrecog resource
>     > >>>     (or advanced
>     > >>>     > dtmfrecog) generates a START-OF-SPEECH because it =20
> heard
>     > >>>     some speech.
>     > >>>     > The client does not want to stop prompt playing unless
>     > >>>     DTMF was heard
>     > >>>     > but it can't tell by the START-OF-SPEECH whether =20
> speech
>     > >>>     or DTMF was
>     > >>>     > heard. Similarly vice versa.
>     > >>>     >
>     > >>>     It's an interesting design question what part of the
>     > >>>     policy resides at the client and what at the server, and
>     > >>>     who makes the "final decision" about whether what was
>     > >>>     heard was relevant to the control channel. Right now we
>     > >>>     (IMO) have a weird partitioning in many cases where the
>     > >>>     client basically says "do what I mean", but there are no
>     > >>>     constraints of what the server actually does, and no
>     > >>>     normalized basis for the client to figure out what to =20
> set
>     > >>>     various magic numbers to (e.g. sensitivity).
>     > >>>
>     > >>>     In this case the only thing the client needs to =20
> decide is
>     > >>>     whether to kill the prompt because the server thinks
>     > >>>     something that would interfere with the feedback
>     > >>>     ear/mouth/finger control happened. What this
>     says to me is
>     > >>>     that it isn't necessarily a good idea for the client to
>     > >>>     have more knobs to control the server
>     (especially if those
>     > >>>     knows are just more value/policy input ungrounded in any
>     > >>>     physics/ acoustics). On the other hand, having the =20
> server
>     > >>>     tell the client more about what it thinks is going on is
>     > >>>     probably valuable.
>     > >>>
>     > >>>     So, Coming to the point after this long rambling
>     > >>>     introduction, I think it would in fact be useful for the
>     > >>>     START-Of-SPEECH event to indicate some extra =20
> information,
>     > >>>     for example:
>     > >>>     a) I got something enough above the noise floor
>     to qualify
>     > >>>     for exceeding the "Sensisitvity" parameter you sent =20
> in on
>     > >>>     the request but I really can't tell what it is
>     (could be a
>     > >>>     hippopatmus fart, or a siren in the background,
>     or captain
>     > >>>     crunch trying to whistle DTMF).
>     > >>>     b) I think I'm hearing speech
>     > >>>     c) I think I'm hearing DTMF
>     > >>>
>     > >>> Though I agree with your former part of your
>     response. I am not
>     > >>> sure I agree with your proposed solution.
>     > >>> The way I see this problem is that, it is more of what
>     > constitues a
>     > >>> barge-in event. This boils down to whether it is
>     speech, DTMF or
>     > >>> both.
>     > >>> This is inturn boils down to what type of recognizer
>     > resource we are
>     > >>> using, dtmf-recog, speech-recog, and speech-only-recog(we
>     > don't have
>     > >>> this and I don't think we should add it, but think
>     of this as a
>     > >>> place holder that explains the concept).
>     > >>>
>     > >>> A client knowing what type of barge-in happenned, does
>     > not impact
>     > >>> the
>     > >>> barge-in operation itself as it may be too late(for
>     the optimized
>     > >>> barge-in case). It may have other use cases, and if we
>     > can identify
>     > >>> them, I don't mind adding support for the
>     START-OF-SPEECH event to
>     > >>> say what type of barge-in happenned. But that itself does =20
> not
>     > solve the
>     > >>> original problem raised. Refer to my previous response.
>     > >>>
>     > >>> The solution lies in defining what what is a
>     barge-in event.
>     > >>> That  boils
>     > >>> down to what type of recognition is happenning,
>     > dtmf-only, speech-
>     > >>> dtmf
>     > >>> or speech-only. We do not support speech-only as a
>     resource today,
>     > >>> the question is do we need a header to force it.
>     > >>>
>     > >>> Sarvi
>     > >>>
>     > >>>
>     > >>>     >    2. You would benefit from the client knowing
>     > what caused
>     > >>> the
>     > >>>     > barge-in, DTMF Vs speech.
>     > >>>     >
>     > >>>     > DB> See previous comment. And previous e-mail:
>     either add an
>     > >>>     > inputmodes header (taking value speech, dtmf, both) to
>     > >>>     the RECOGNIZE
>     > >>>     > request or add a header to the START-OF-SPEECH event
>     > >>>     indicating DTMF
>     > >>>     > or speech.
>     > >>>     >
>     > >>>     I'm leaning in your direction on this latter point - as
>     > >>>     should be evident from what I wrote above.
>     > >>>
>     > >>>     > Sarvi
>     > >>>     >
>     > >>>     >     -----Original Message-----
>     > >>>     >     From: speechsc-bounces@ietf.org
>     > >>>     >     [mailto:speechsc-bounces@ietf.org] On Behalf Of
>     > David R
>     > >>> Oran
>     > >>>     >     Sent: Tuesday, July 05, 2005 5:24 AM
>     > >>>     >     To: Klaus Reifenrath
>     > >>>     >     Cc: 'speechsc@ietf.org'
>     > >>>     >     Subject: Re: [Speechsc] START-OF-SPEECH in
>     > DTMF-only mode
>     > >>>     >
>     > >>>     >
>     > >>>     >     On Jul 5, 2005, at 3:46 AM, Reifenrath,
>     Klaus wrote:
>     > >>>     >
>     > >>>     >     > The current spec is not clear when
>     > START-OF-SPEECH need
>     > >>>     >     to be send in
>     > >>>     >     > the following scenarios:
>     > >>>     >     > A) The client requested a DTMF Recognizer. Is =20
> the
>     > >>>     >     START-OF-SPEECH
>     > >>>     >     > event send to the client also if speech
>     was detected?
>     > >>>     >     I suspect so, since one of the prime purposes
>     > is to enable
>     > >>>     >     client- mediated barge-in handling. However, if =20
> the
>     > >>>     >     recognizer is in fact only capable of
>     recognizing DTMF
>     > >>>     >     then it may in fact not report anythin
>     unless it's using
>     > >>>     >     some primitive thresholding machinery, like a SN
>     > >>> threshold.
>     > >>>     >     > B) The client requested a Speech
>     Recognizer, but only
>     > >>>     >     activated DTMF
>     > >>>     >     > grammars. Is the START-OF-SPEECH event
>     send to the
>     > >>>     >     client also if
>     > >>>     >     > speech was detected?
>     > >>>     >     Again, I'd say yes, for the same reason as above.
>     > >>>     >     > I think in both cases START-OF-SPEECH should =20
> only
>     > >>>     be send after
>     > >>>     >     > detecting a DTMF digit (see Figure 12 of =20
> VoiceXML
>     > >>>     2.0: http://
>     > >>>     >     > www.w3.org/TR/voicexml20/#dmlATiming).
>     > >>>     >     We seem to have reached different
>     conclusions. I'd be
>     > >>>     >     interested in why you think my analysis
>     above is wrong.
>     > >>>     >
>     > >>>     >     Dave.
>     > >>>     >
>     > >>>     >     > Klaus
>     > >>>     >     >
>     > >>>     >     > _______________________________________________
>     > >>>     >     > Speechsc mailing list
>     > >>>     >     > Speechsc@ietf.org
>     > >>>     >     > https://www1.ietf.org/mailman/listinfo/speechsc
>     > >>>     >     >
>     > >>>     >
>     > >>>     >     _______________________________________________
>     > >>>     >     Speechsc mailing list
>     > >>>     >     Speechsc@ietf.org
>     > >>>     >     https://www1.ietf.org/mailman/listinfo/speechsc
>     > >>>     >
>     > >>>     >
>     > >>>     > _______________________________________________
>     > >>>     > Speechsc mailing list
>     > >>>     > Speechsc@ietf.org
>     > >>>     > https://www1.ietf.org/mailman/listinfo/speechsc
>     > >>>     >
>     > >>>
>     > >>>
>     > >>> _______________________________________________
>     > >>> Speechsc mailing list
>     > >>> Speechsc@ietf.org
>     > >>> https://www1.ietf.org/mailman/listinfo/speechsc
>     > >>>
>     > >>>
>     > >>
>     > >> _______________________________________________
>     > >> Speechsc mailing list
>     > >> Speechsc@ietf.org
>     > >> https://www1.ietf.org/mailman/listinfo/speechsc
>     > >
>     >
>     > _______________________________________________
>     > Speechsc mailing list
>     > Speechsc@ietf.org
>     > https://www1.ietf.org/mailman/listinfo/speechsc
>     >
>
>
>     _______________________________________________
>     Speechsc mailing list
>     Speechsc@ietf.org
>     https://www1.ietf.org/mailman/listinfo/speechsc
>
>
> _______________________________________________
> Speechsc mailing list
> Speechsc@ietf.org
> https://www1.ietf.org/mailman/listinfo/speechsc
>
>
> _______________________________________________
> Speechsc mailing list
> Speechsc@ietf.org
> https://www1.ietf.org/mailman/listinfo/speechsc
>

_______________________________________________
Speechsc mailing list
Speechsc@ietf.org
https://www1.ietf.org/mailman/listinfo/speechsc

=20
 =20


_______________________________________________
Speechsc mailing list
Speechsc@ietf.org
https://www1.ietf.org/mailman/listinfo/speechsc



From speechsc-bounces@ietf.org Mon Jul 11 09:15:52 2005
Received: from localhost.localdomain ([127.0.0.1] helo=megatron.ietf.org)
	by megatron.ietf.org with esmtp (Exim 4.32)
	id 1Dry8W-0006R9-P2; Mon, 11 Jul 2005 09:15:52 -0400
Received: from odin.ietf.org ([132.151.1.176] helo=ietf.org)
	by megatron.ietf.org with esmtp (Exim 4.32) id 1Dry8U-0006R4-TT
	for speechsc@megatron.ietf.org; Mon, 11 Jul 2005 09:15:51 -0400
Received: from ietf-mx.ietf.org (ietf-mx [132.151.6.1])
	by ietf.org (8.9.1a/8.9.1a) with ESMTP id JAA08805
	for <speechsc@ietf.org>; Mon, 11 Jul 2005 09:15:49 -0400 (EDT)
Received: from sj-iport-3-in.cisco.com ([171.71.176.72]
	helo=sj-iport-3.cisco.com) by ietf-mx.ietf.org with esmtp (Exim 4.43)
	id 1DryaU-0006Sb-BZ
	for speechsc@ietf.org; Mon, 11 Jul 2005 09:44:47 -0400
Received: from sj-core-2.cisco.com (171.71.177.254)
	by sj-iport-3.cisco.com with ESMTP; 11 Jul 2005 06:15:40 -0700
X-IronPort-AV: i="3.93,278,1115017200"; 
	d="scan'208"; a="300454095:sNHT115176380"
Received: from imail.cisco.com (imail.cisco.com [128.107.200.91])
	by sj-core-2.cisco.com (8.12.10/8.12.6) with ESMTP id j6BDFXod006408;
	Mon, 11 Jul 2005 06:15:33 -0700 (PDT)
Received: from [10.32.245.152] (stealth-10-32-245-152.cisco.com
	[10.32.245.152])
	by imail.cisco.com (8.12.11/8.12.10) with SMTP id j6BDEKVk012406;
	Mon, 11 Jul 2005 06:14:22 -0700
In-Reply-To: <7DE7C4EF3B7C8B4B82955191378290D802ED4176@mtb1exch01.nuance.com>
References: <7DE7C4EF3B7C8B4B82955191378290D802ED4176@mtb1exch01.nuance.com>
Mime-Version: 1.0 (Apple Message framework v730)
Content-Type: text/plain; charset=US-ASCII; delsp=yes; format=flowed
Message-Id: <C7673C4C-0624-47A1-B1DF-115FB284785A@cisco.com>
Content-Transfer-Encoding: 7bit
From: David R Oran <oran@cisco.com>
Subject: Re: [Speechsc] START-OF-SPEECH in DTMF-only mode -> summary(?)
Date: Mon, 11 Jul 2005 09:15:28 -0400
To: Pierre Forgues <forgues@nuance.com>
X-Mailer: Apple Mail (2.730)
IIM-SIG: v:"1.1"; h:"imail.cisco.com"; d:"cisco.com"; z:"home"; m:"krs";
	t:"1121087667.278746"; x:"432200"; a:"rsa-sha1"; b:"nofws:26543";
	e:"Iw=="; n:"sQYarK2E51MdcTiUqeif3F7cWdxIfoCiXhdfb9vD5ee/j0jXL15gbFxF2p"
	"XIweAblu0N6XAgK7k+wrbr7bQDJaCDqOmzqpRUBjIRQAXQ7NzadpmR3pUL6wxaRUtW+c43sl9jC"
	"50Qg1sXHpPjt8Y+Y16ioyQAQAdSunM4YhevURc=";
	s:"GIPgroEN6XyHEg1veEn/1JVI/gWjSRG59m1lXCWE0lDfm0CHATv9J/U2+0yIECo7PkcSVXng"
	"3RykcSOZ+f6rEgFSmcbyV4LKn+Pa4sBlHIYuKAjWy8AwxAn9N+52ij6idlm7IB77k1IZaxZN36h"
	"onKa/9SoySoCjChrIR4TSmAU="; c:"From: David R Oran <oran@cisco.com>";
	c:"Subject: Re: [Speechsc] START-OF-SPEECH in DTMF-only mode ->
	summary" "(?)"; c:"Date: Mon, 11 Jul 2005 09:15:28 -0400"
IIM-VERIFY: s:"y"; v:"y"; r:"60"; h:"imail.cisco.com";
	c:"message from imail.cisco.com verified; "
X-Spam-Score: 0.0 (/)
X-Scan-Signature: 0f5efb848eea823449b646bdc843df76
Content-Transfer-Encoding: 7bit
Cc: speechsc@ietf.org, "Shanmugham, Saravanan" <sarvi@cisco.com>,
	Eric Burger <eburger@brooktrout.com>, Dave Burke <david.burke@voxpilot.com>
X-BeenThere: speechsc@ietf.org
X-Mailman-Version: 2.1.5
Precedence: list
List-Id: Speech Services Control Working Group <speechsc.ietf.org>
List-Unsubscribe: <https://www1.ietf.org/mailman/listinfo/speechsc>,
	<mailto:speechsc-request@ietf.org?subject=unsubscribe>
List-Post: <mailto:speechsc@ietf.org>
List-Help: <mailto:speechsc-request@ietf.org?subject=help>
List-Subscribe: <https://www1.ietf.org/mailman/listinfo/speechsc>,
	<mailto:speechsc-request@ietf.org?subject=subscribe>
Sender: speechsc-bounces@ietf.org
Errors-To: speechsc-bounces@ietf.org


On Jul 8, 2005, at 11:36 AM, Pierre Forgues wrote:

>
>
> -----Original Message-----
> From: speechsc-bounces@ietf.org [mailto:speechsc-bounces@ietf.org] On
> Behalf Of David R Oran
> Sent: Friday, July 08, 2005 10:05 AM
> To: Dave Burke
> Cc: speechsc@ietf.org; Shanmugham, Saravanan; Eric Burger
> Subject: Re: [Speechsc] START-OF-SPEECH in DTMF-only mode -> summary 
> (?)
>
>
> On Jul 8, 2005, at 6:59 AM, Dave Burke wrote:
>
>
>> I agree with your opinion on what constitutes a barge-in. Based on
>> Klaus' e-mail, I am more concerned that using the grammar type to
>> decide what constitutes a barge-in is going to make VoiceXML
>> implementations difficult.
>>
>> It is easy to map VoiceXML application selected inputmodes to MRCP
>> resources:
>> a. inputmodes="dtmf" -> Use a dtmfrecog
>> b. inputmodes="speech" -> Use a "speech-only-recog"
>> c. inputmodes="both" -> Use a speechrecog (or a combination of
>> "speech-only-recog" + dtmfrecog)
>>
>> The only problem is (b). While not as important as being able to do
>> DTMF-only recognition, I believe we DO need to support speech-only
>> recongition so as (a) to avoid unnecessary limitations in VUI
>> design, and (b) to facilitate VoiceXML implementations. I think we
>> should add a header to RECOGNIZE so the client is able to always
>> _explicitly_ set the inputmodes.
>>
>>
> I'm not sure I buy this, since what the recognizer is looking for and
> what constitutes an input that should be considered a potential barge-
> in strike me as independent. I'm similarly not persuaded that we need
> the flexibility to set the input mode independently of the grammar,
> since it leads to all sorts of inconsistent states (e.g. speech-only
> grammar with an input-mode of dtmf). On the other hand I can see the
> need for the client to specify what sorts of input out to generate
> the start-of-input (nee start-of-speech) event on.
>
> Pmf> The mechanism for a client to specify the type of input is using
> grammars.  These can have DTMF and/or speech requirements.  If you add
> the complexity of an independent header to specify the input mode then
> you will need to document the behavior for inconsistencies.
>
Right. I generally agree with Pierre, which is why I proposed a  
compromise position where the client can use a header to specify what  
types of input should be considered the start of something  
potentially interesting from a barge-in point of view, as opposed to  
a directive to the server for what to recognize.

I don't think this is crucial functionality, but it does make for  
more consistent behavior between the optimized and non-optimized  
barge-in cases. In fact, I can see cases where the client doesn't  
want any barge-in processing to happen. Imagine an application called  
"mark the musical phrases", where the client listens to playout of a  
musical score, and hits various DTMF buttons to indicate markers  
(e.g. end-of-phrase, end-of-theme, key-change) while listening. For  
this application you want to make sure that barge-in doesn't stop the  
music!

>
>> In retrospect, I don't like the idea of multiple START-OF-SPEECH
>> events being generated from the same media resource. This is
>> because one assumes that the START-OF-SPEECH should be of the same
>> type as the hypothesis returned in the RECOGNITION-COMPETE message
>> - most implementations, on hearing one input mode type disable the
>> recogniser of the other type.
>>
> I'm not sure I follow this logic, but I'm not wedded to the idea of
> allowing multiple events. It was trying to solve the problem of
> ambiguity around what the recognizer was hearing and the idea that
> the client might care. If you don't think the client will ever care,
> then we don't need the capability.
> Pmf> I agree we should not have multiple SOS events.  Too much  
> chatter.
>
I'm not concerned about the "chatter" since this is a low-bandwidth  
control channel even with multiple events, but as I said it's a small  
point and I'm happy to concede.

>
>
>> I do like David's idea of renaming START-OF-SPEECH to something
>> like START-OF-INPUT and carrying a type header because it is neater
>> and more extensible.
>>
>> So in summary, I propose we modify the spec to:
>>
>> 1. Clarify what constitutes a barge-in for a dtmfrecog and
>> speechrecog as per Sarvi's e-mail (and in agreement with Klaus' for
>> dtmfrecog).
>>
>>
> Hmmm, ok, but I think the issue is actually clarifying what the
> recognizer declares as "interesting input", which it may decide also
> constitutes a barge-in in the optimized case, and in either case
> reports to the client that it heard.
> Pmf> I'm not going to argue changing the name of the event, but my
> preference would be to keep existing method names unless there is a
> clear reason for changing - which I do not see here.
>
I do see a clear reason, since the thing that you see the start of  
may not be speech. I like (I think it was Dave's suggestion) "start- 
of-input:" since it mirrors the other method names.
>
>> 2. Specifiy an InputModes header to RECOGNIZE (defaults to "both",
>> can also be "speech" or "DTMF"). Setting to speech for a
>> speechrecog results in the hypothesised "speech-only-recog".
>> Setting to DTMF for a speechrecog is equivalent to using a
>> dtmfrecog. Edge cases: Setting to speech for a dtmfrecog will
>> result in a noinput as will setting to DTMF for a speechrecog which
>> does not support DTMF.
>>
>>
> I'm ok with having a header for the client to tell the server when it
> would like the start-of-input event to be generated and what the
> client considers to be the "interesting input" that the recognizer
> should use to do the discrimination, and possibly do barge-in
> processing in the optimized case.
>
> I'm less ok with the proposed domain of values, since it will have
> extensibility problems. Especially problematical is have a code point
> of "both" since that will be ambiguous if we even define a "music"
> recognizer, or a "gesture" recognizer using video input. Here's my
> counter-proposal:
>
> Create a header on the recognizer resource requests called "start-
> input-on:". Define a registry of values, with the following three
> values initially defined:
>      - dtmf
>      - speech
>      - music
>
> The semantics would be that the resource is to generate a "start-of-
> input" event, apply the defined grammars to what follows, and do any
> local optimized barge-in processing if ANY of the enumerated input
> types is detected. That way you can say
>      "start-input-on: dtmf" if you want to just  dtmf,
>      "start-input-on: speech" if you want just speech (dtmf input
> would be ignored even if the grammar supported dtmf)
>      "start-input- on: speech, dtmf" if you wanted both
> etc.
>
> Pmf> Doesn't this come back to the previous point of having two
> independent ways to specify input modes?  Through grammars and this
> proposed new header?  If you add this then you will have to document
> behavior on inconsistencies.  Anyway, if you proceed then I would
> definitely make this optional
>
I hope I explained the intent above. This header would not define the  
input types that the recognize would process, but rather the input  
types the client wants it to consider potential barge-in and hence  
generate the "start-of-input:" event for and do any local barge-in  
optimizations for.


Dave.

>
>> 3. Change START-OF-SPEECH to START-OF-INPUT. Clarify it occurs once
>> and coincides with barge-in.
>>
>>
> Ok for the "only once", but I'd like to tighten it up to talk about
> more than just barge-in, as I suggested above.
>
>
>> 4. Add a header of InputType to START-OF-INPUT. Current specified
>> values are "dtmf" or "speech".
>>
>>
> Ok, with slight modification. Say that the syntax of the "input-type"
> header is a single-valued subset of the registered value(s) that are
> defined for the start-input-on: header.
>
> Comments?
>
> Dave O. (technical hat on, chair hat off).
>
>
>> Dave
>>
>>
>>
>> ----- Original Message ----- From: "Shanmugham, Saravanan"
>> <sarvi@cisco.com>
>> To: "Eric Burger" <eburger@brooktrout.com>; <speechsc@ietf.org>
>> Sent: Thursday, July 07, 2005 5:52 PM
>> Subject: RE: [Speechsc] START-OF-SPEECH in DTMF-only mode -> summary
>> (?)
>>
>>
>> I don't think generating multiple START-OF-SPEECH events is a
>> solution.
>> We still haven't addressed, what constitutes a barge-in, for the
>> optimised case. That should also be the single point when a single
>> START-OF-SPEECH(or whathever else you want to name it) should be
>> generated.
>>
>> That point, in my opinion should be
>>   1. For "dtmf-recog" resources should be the beginning of a DTMF key
>> press.
>>   2. For "speech-recog" resources should be the beginning of a DTMF
>> key
>> press or the beginning of speech. This is should be irrespective of
>> what
>> type of grammar is being used. Coz even numbers only grammars can
>> still
>> be spoken and hence cannot be assumed to be a cue for DTMF only
>> recognition.
>>   3. For "speech-only-recog" resources(which are not defined today)
>> the
>> time to barge-in is the beginning of speech. I don't see a need for
>> such
>> a resource today. But I am mentioning this for completeness.
>>
>> Sarvi
>>
>>
>>     -----Original Message-----
>>     From: speechsc-bounces@ietf.org
>>     [mailto:speechsc-bounces@ietf.org] On Behalf Of Eric Burger
>>     Sent: Thursday, July 07, 2005 5:32 AM
>>     To: speechsc@ietf.org
>>     Subject: RE: [Speechsc] START-OF-SPEECH in DTMF-only mode
>>     -> summary(?)
>>
>>     Can we declare consensus?
>>
>>
>>> -----Original Message-----
>>> From: speechsc-bounces@ietf.org
>>>
>>     [mailto:speechsc-bounces@ietf.org]
>>
>>> Sent: Wednesday, July 06, 2005 10:45 AM
>>> To: Dave Burke
>>> Cc: speechsc@ietf.org; Shanmugham, Saravanan; Klaus Reifenrath
>>> Subject: Re: [Speechsc] START-OF-SPEECH in DTMF-only mode ->
>>> summary(?)
>>>
>>> I think we're getting close. I though about snipping out
>>>
>>     some pieces
>>
>>> to cut down the text, but I realized the context is
>>>
>>     still needed. See
>>
>>> inline.
>>>
>>> On Jul 6, 2005, at 10:29 AM, Dave Burke wrote:
>>>
>>>
>>>> Inline.
>>>>
>>>> Dave
>>>>
>>>> ----- Original Message ----- From: "David R Oran"
>>>>
>>     <oran@cisco.com>
>>
>>>> To: "Dave Burke" <david.burke@voxpilot.com>
>>>> Cc: <speechsc@ietf.org>; "Shanmugham, Saravanan"
>>>>
>>     <sarvi@cisco.com>;
>>
>>>> "Klaus Reifenrath" <Klaus.Reifenrath@Scansoft.com>
>>>> Sent: Wednesday, July 06, 2005 1:07 PM
>>>> Subject: Re: [Speechsc] START-OF-SPEECH in DTMF-only
>>>>
>>     mode -> summary
>>
>>>> (?)
>>>>
>>>>
>>>>
>>>>
>>>>>
>>>>> On Jul 6, 2005, at 5:31 AM, Dave Burke wrote:
>>>>>
>>>>>
>>>>>
>>>>>> + Attempting to summarise:
>>>>>>
>>>>>> 1. START-OF-SPEECH is useful for the client to know when
>>>>>>
>>> to stop
>>>
>>>>>> playing prompts the in non-optimised case 2.
>>>>>>
>>     START-OF-SPEECH is
>>
>>>>>> useful for the client to calculate the bargin  time
>>>>>>
>>     (e.g. VoiceXML
>>
>>>>>> 2.1 <mark>)
>>>>>> 3. In the optimised case, a bargin automatically
>>>>>>
>>     stops prompt
>>
>>>>>> playing (assuming prompts barginable) 4. Because of
>>>>>>
>>     the previous
>>
>>>>>> point, the question of what input type  caused
>>>>>>
>>     bargin is different
>>
>>>>>> and less important to
>>>>>>
>>> what input
>>>
>>>>>> type(s)  the recogniser is listening for
>>>>>>
>>>>>>
>>>>> I'm not sure I follow point 4. Could you elaborate?
>>>>>
>>>>>
>>>>
>>>> DB> Adding a parameter to the START-OF-SPEECH event would
>>>>
>>> certainly
>>>
>>>> allow the client to ignore (i.e. let prompts continue
>>>>
>>     playing) the
>>
>>>> event if the event type is not of interest (e.g. the
>>>>
>>     client would
>>
>>>> ignore speech start events when it is interested only
>>>>
>>     in  DTMF start
>>
>>>> events). This _only_ works for the non-optimised case,
>>>>
>>     however. For
>>
>>>> the optimised case, assuming START-OF-SPEECH
>>>>
>>> coincides
>>>
>>>> with the bargin signal to the speechsynth, prompts
>>>>
>>     will stop playing
>>
>>>> for inputs that the client might not be interested
>>>>
>>> in (e.g.
>>>
>>>> a speech input will stop prompts playing even if the client
>>>>
>>> is only
>>>
>>>> interested in DTMF).
>>>>
>>>>
>>> OK, now I get it. There's a need for the client to both
>>>
>>     handle the
>>
>>> non-optimized case itself, and influence or at least
>>>
>>     have a clue what
>>
>>> the server is going to do in the optimized case.
>>>
>>>
>>>>
>>>>
>>>>>
>>>>>
>>>>>
>>>>>> 5. Currently, MRCPv2 has no way of indicating what
>>>>>>
>>     input type(s) a
>>
>>>>>> recogniser is listening for
>>>>>>
>>>>>>
>>>>>>
>>>>> Do you mean exactly this, or do you mean "for the client to
>>>>> indicate  to the resource what input types it should
>>>>>
>>     look for"?
>>
>>>>>
>>>>>
>>>>
>>>> DB> Yes exactly - apologies for not being clear.
>>>>
>>>>
>>>>
>>>>>
>>>>>
>>>>>
>>>>>> + Why implement 5?
>>>>>>
>>>>>> i. Noisy case: Need DTMF-only recognition (and may
>>>>>>
>>     only have a
>>
>>>>>> speechrecog)
>>>>>>
>>>>>>
>>>>> I'm having difficulty following the logic of why
>>>>>
>>     noise would
>>
>>>>> necessarily trigger START-OF-SPEECH if you were listening for
>>>>> speech  but not DTMF. I suppose you can use a more forgiving
>>>>> discriminator if  all you need to tell is if you're
>>>>>
>>     getting DTMF,
>>
>>>>> but I've had a number  of real-world cases where wind
>>>>>
>>     noise was
>>
>>>>> detected as DTMF, and  there's always the ambiguity
>>>>>
>>     when you have
>>
>>>>> Captain Crunch on the  phone. In either case in the
>>>>>
>>     non-optimized
>>
>>>>> case it's the client who  gets to decide whether an
>>>>>
>>     event should be
>>
>>>>> interpreted as barge-in or  not, so it seems an
>>>>>
>>     aesthetic protocol
>>
>>>>> design decision whether the  client tells the server
>>>>>
>>     ahead of time
>>
>>>>> what circumstances to generate  the START-OF-SPEECH
>>>>>
>>     event for, or
>>
>>>>> whether the event gets generated  and the client
>>>>>
>>     decides based on
>>
>>>>> what's in the event whether it should  be
>>>>>
>>> treated
>>>
>>>>> as barge- in.
>>>>>
>>>>> I suppose one could make the argument that because
>>>>>
>>     the spec implies
>>
>>>>> that the event can only be generated once per request
>>>>>
>>     that if a
>>
>>>>> DTMF/ speech capable recognizer first hears
>>>>>
>>> enough noise
>>>
>>>>> to think it's  hearing speech and later hears DTMF,
>>>>>
>>     the client will
>>
>>>>> declare barge-in  when the event comes and not when he DTMF
>>>>> actually gets heard.
>>>>>
>>>>> If that's deemed a problem, we can still handle that
>>>>>
>>     in the design
>>
>>>>> where the server just reports what it's hearing by allowing
>>>>> multiple  events to be generated during a single request.
>>>>>
>>>>> Between the approach just outlined above, and an
>>>>>
>>     approach where the
>>
>>>>> client provides a filter for whether to generate the
>>>>>
>>     event or not,
>>
>>>>> I  have a mild preference (based on aesthetics rather
>>>>>
>>     than some
>>
>>>>> hard  engineering tradeoff) for the approach where
>>>>>
>>> the server
>>>
>>>>> just reports  what it's hearing.
>>>>>
>>>>>
>>>>> Having had some useful exchanges on this topic, it
>>>>>
>>     also is becoming
>>
>>>>> apparent to me that this event is poorly named, and
>>>>>
>>     we should
>>
>>>>> consider renaming it to "INTERESTING-INPUT-HEARD" or
>>>>>
>>     something akin
>>
>>>>> to that, because as others have pointed out, a
>>>>>
>>     DTMF-only recognizer
>>
>>>>> will never detect "start of speech".
>>>>>
>>>>> Another consideration to fold into the design choice is
>>>>> extensibility. Bear with me through a little
>>>>>
>>     gedankenexperiment.
>>
>>>>>
>>>>> Suppose we want to define a new recognizer type,
>>>>>
>>     which I'll call
>>
>>>>> the "name that tune" recognizer. The client plays
>>>>>
>>     music to the
>>
>>>>> server and  the server recognizes musical notes. The
>>>>>
>>     grammar is a
>>
>>>>> standard  musical notation, augmented with a semantic
>>>>> interpretation that  transforms the notes into the
>>>>>
>>     title of the
>>
>>>>> tune and provides that as  an answer.
>>>>>
>>>>> First, there's no speech involved (or is there...hang
>>>>>
>>     on a minute).
>>
>>>>> Second, in order to accommodate the "name that tune"
>>>>> recognizer, we'd have to extend both the client and
>>>>>
>>     the server to
>>
>>>>> undetstand a  directive as to whether to recognize
>>>>>
>>     music or now,
>>
>>>>> inaddition to what  the server already knows what to
>>>>>
>>     do based on
>>
>>>>> the grammar. If you  follow my logic above, whether
>>>>>
>>     or not we do
>>
>>>>> that, we have to extend  "start-of-speech" to say
>>>>>
>>     "I'm hearing
>>
>>>>> music". So far fairly  straightforward, but let me
>>>>>
>>     now throw in the
>>
>>>>> pathological twist.
>>>>>
>>>>> Suppose what I feed to a  combined music/speech
>>>>>
>>     recognizer is a
>>
>>>>> work  in sprechstimme (spoken music), like the
>>>>>
>>     "Geographical Fugue"
>>
>>>>> (aside:  this is a wonderful piece of music I highly
>>>>>
>>     recommend to
>>
>>>>> anyone  interested in small ensemble singing). In
>>>>>
>>     this case, the
>>
>>>>> tune could  be named by either doing speech or music
>>>>>
>>     recognition.
>>
>>>>> Why is there  any need for the client to constrain
>>>>>
>>     the server as to
>>
>>>>> which it tries  to do when it's
>>>>>
>>> already
>>>
>>>>> told the server what it wants through the  grammar?
>>>>>
>>>>> A few other comments below
>>>>>
>>>>>
>>>>>
>>>>>> ii. Flexibility: Want speech-only recognition
>>>>>>
>>     (because a second
>>
>>>>>> recogniser is doing hotword on DTMF)
>>>>>>
>>>>>>
>>>>>>
>>>>> I don't see how flexibility is affected by this deisgn
>>>>>
>>> choice. If
>>>
>>>>> that's what you want, feed the speech-only recognizer
>>>>>
>>     a grammar
>>
>>>>> without any DTMF rules.
>>>>>
>>>>>
>>>>>
>>>>>> + How to implement 5?
>>>>>>
>>>>>> a. Implicitly:
>>>>>>    - dtmfrecog: always DTMF-only recognition
>>>>>>    - speechrecog: depends on active grammar type
>>>>>>
>>>>>>> if a dtmf grammar is active then DTMF input is "on"
>>>>>>> if a speech grammar is active then speech
>>>>>>>
>>     input is "on"
>>
>>>>>>
>>>>>> b. Explicitly:
>>>>>>    - Add inputmodes header to RECOGNIZE
>>>>>>
>>>>>> Option a is David's "do what I mean case"; option b is
>>>>>>
>>> the extra
>>>
>>>>>> dial for the client.
>>>>>>
>>>>>>
>>>>>>
>>>>> Actually, that's not the point I was making with "do
>>>>>
>>     what I mean",
>>
>>>>> but it's not essential to the discussion so let's move on.
>>>>>
>>>>>
>>>>>
>>>>>> It is worth noting that VoiceXML uses option b. This
>>>>>>
>>     allows one to
>>
>>>>>> activate both speech grammars and DTMF grammars (and
>>>>>>
>>> therefore
>>>
>>>>>> be informed of any errors in the grammars at activation
>>>>>>
>>> time) but
>>>
>>>>>> independently turn on whichever input mode you like e.g.
>>>>>>
>>> perhaps
>>>
>>>>>> start with "both" then change to "dtmf".
>>>>>>
>>>>>>
>>>>>>
>>>>> I'm not sure the VXML precedent is relevant here,
>>>>>
>>     because the
>>
>>>>> application behind VXML is working a different part of the
>>>>>
>>> problem
>>>
>>>>> -  how to traverse a TUI dialog based on different
>>>>>
>>     parts of the
>>
>>>>> input  space. In fact, I suspect that the VXML: choice was
>>>>> conditioned more  by limitations at the time it was
>>>>>
>>> specified than
>>>
>>>>> an underlying good  design choice. Clearly having to
>>>>>
>>     specify this
>>
>>>>> in VXML make the job of  handling a TUI with nodes
>>>>>
>>     like "Say or
>>
>>>>> press 5" harder rather than  easier.
>>>>>
>>>>>
>>>>
>>>> DB> The VoiceXML edge-case is pretty weird so it's not a major
>>>> concern. My main concern is that the client can
>>>>
>>     indicate, somehow,
>>
>>>> what the input modes are.
>>>>
>>>>
>>>>
>>>>>
>>>>> Summing up, while I don't feel strongly one way or
>>>>>
>>     the other, I
>>
>>>>> have  a preference for handling this as follows:
>>>>>
>>>>> a) Rename "START-OF-SPEECH" to
>>>>>
>>     "INTERESTING-INPUT-RECEIVED" or
>>
>>>>> something equivalent.
>>>>> b) Include a parameter in the event saying what was
>>>>>
>>     interesting
>>
>>>>> about  the input you received, with a registry of values
>>>>>
>> which
>>
>>>>> includes:
>>>>>     - signal above noise floor
>>>>>     - speech
>>>>>     - dtmf
>>>>>     - (possibly) music
>>>>> c) allow the event to be generated multiple times
>>>>>
>>     during a request
>>
>>>>>
>>>>>
>>>>>
>>>>
>>>> DB> I like these suggestions (START-OF-INPUT?).
>>>>
>>     However, I don't
>>
>>>> see how the problem of the optimised case is not
>>>>
>>     solved by them. I
>>
>>>> think the optimised case is fine if we have the
>>>>
>>     following rules:
>>
>>>>
>>>>
>>> Yes, I hadn't thought through the optimized case as
>>>
>>     thoroughly as you.
>>
>>> Your suggested method name is fine by me as well
>>>
>>>
>>>> 1. START-OF-SPEECH (and optimised bargin) is only generated
>>>>
>>> for the
>>>
>>>> input type that is being listened for 2. A speechrecog
>>>>
>>     listens for
>>
>>>> DTMF if DTMF grammars are active, speech if speech
>>>>
>>     grammars are
>>
>>>> active, or speech and DTMF if both grammar types are active.
>>>>
>>>>
>>> Works for me.
>>>
>>>
>>>>
>>>>
>>>>> Note that all of the above I'm saying with my technical
>>>>>
>>> hat on and
>>>
>>>>> my chair hat off.
>>>>> Putting my chair hat on for a moment, we really need
>>>>>
>>     to get this
>>
>>>>> spec  to last call, so at some point Eric or I is going to
>>>>>
>>> declare
>>>
>>>>> rough  consensus so we can move on.
>>>>>
>>>>>
>>>>
>>>> DB> Agreed!
>>>>
>>>>
>>>>
>>>>>
>>>>> Dave Oran.
>>>>>
>>>>>
>>>>>> Dave
>>>>>>
>>>>>>
>>>>>>
>>>>>>
>>>>>>
>>>>>>
>>>>>>
>>>>>>
>>>>>>
>>>>>>
>>>>>>
>>>>>>
>>>>>> Sarvi makes a good point that adding the reason why the
>>>>>>
>>> START-OF-
>>>
>>>>>> SPEECH occurred does not fix the optimised bargin case.
>>>>>>
>>>>>> dtmfrecog - listens for DTMF only
>>>>>> speechrecog - listens for DTMF only, or speech only,
>>>>>>
>>     or speech &
>>
>>>>>> DTMF
>>>>>>
>>>>>>
>>>>>>
>>>>>>
>>>>>>
>>>>>>
>>>>>>
>>>>>>
>>>>>> ----- Original Message ----- From: "Shanmugham, Saravanan"
>>>>>> <sarvi@cisco.com>
>>>>>> To: "David R Oran" <oran@cisco.com>; "Dave Burke"
>>>>>> <david.burke@voxpilot.com>
>>>>>> Cc: <speechsc@ietf.org>; "Klaus Reifenrath"
>>>>>> <Klaus.Reifenrath@Scansoft.com>
>>>>>> Sent: Tuesday, July 05, 2005 8:51 PM
>>>>>> Subject: RE: [Speechsc] START-OF-SPEECH in DTMF-only mode
>>>>>>
>>>>>>
>>>>>> inline.
>>>>>>
>>>>>>     -----Original Message-----
>>>>>>     From: David R Oran [mailto:oran@cisco.com]
>>>>>>     Sent: Tuesday, July 05, 2005 11:08 AM
>>>>>>     To: Dave Burke
>>>>>>     Cc: Shanmugham, Saravanan; Klaus Reifenrath;
>>>>>>
>>     speechsc@ietf.org
>>
>>>>>>     Subject: Re: [Speechsc] START-OF-SPEECH in DTMF-only
>>>>>>
>> mode
>>
>>>>>>
>>>>>>
>>>>>>     On Jul 5, 2005, at 12:54 PM, Dave Burke wrote:
>>>>>>
>>>>>>
>>>>>>> Inline.
>>>>>>>
>>>>>>> Dave
>>>>>>>
>>>>>>> ----- Original Message ----- From:
>>>>>>>
>>     "Shanmugham, Saravanan"
>>
>>>>>>> <sarvi@cisco.com>
>>>>>>> To: "David R Oran" <oran@cisco.com>; "Klaus
>>>>>>>
>> Reifenrath"
>>
>>>>>>> <Klaus.Reifenrath@Scansoft.com>
>>>>>>> Cc: <speechsc@ietf.org>
>>>>>>> Sent: Tuesday, July 05, 2005 5:22 PM
>>>>>>> Subject: RE: [Speechsc] START-OF-SPEECH in
>>>>>>>
>>     DTMF-only mode
>>
>>>>>>>
>>>>>>>
>>>>>>> I agree with Dave's analysis. The purpose of this
>>>>>>>
>> event
>>
>>>>>>     was barge-in.
>>>>>>
>>>>>>> And barge-in should happen for both DTMF and speech.
>>>>>>>
>>>>>>> Is there a case where you think it should not behave
>>>>>>>
>>>>>>     this way. If soe,
>>>>>>
>>>>>>> please provide a scenario where you think
>>>>>>>    1. Barge-in should happen for DTMF and not voice or
>>>>>>>
>>>>>>     vice-versa.
>>>>>>
>>>>>>>
>>>>>>> DB> You want to do a DTMF recognition only
>>>>>>>
>>     because it is
>>
>>>>>> noisy.
>>>>>>
>>>>>>> While waiting for DTMF input, the speechrecog resource
>>>>>>>
>>>>>>     (or advanced
>>>>>>
>>>>>>> dtmfrecog) generates a START-OF-SPEECH because it
>>>>>>>
>> heard
>>
>>>>>>     some speech.
>>>>>>
>>>>>>> The client does not want to stop prompt playing unless
>>>>>>>
>>>>>>     DTMF was heard
>>>>>>
>>>>>>> but it can't tell by the START-OF-SPEECH whether
>>>>>>>
>> speech
>>
>>>>>>     or DTMF was
>>>>>>
>>>>>>> heard. Similarly vice versa.
>>>>>>>
>>>>>>>
>>>>>>     It's an interesting design question what part of the
>>>>>>     policy resides at the client and what at the server, and
>>>>>>     who makes the "final decision" about whether what was
>>>>>>     heard was relevant to the control channel. Right now we
>>>>>>     (IMO) have a weird partitioning in many cases where the
>>>>>>     client basically says "do what I mean", but there are no
>>>>>>     constraints of what the server actually does, and no
>>>>>>     normalized basis for the client to figure out what to
>>>>>>
>> set
>>
>>>>>>     various magic numbers to (e.g. sensitivity).
>>>>>>
>>>>>>     In this case the only thing the client needs to
>>>>>>
>> decide is
>>
>>>>>>     whether to kill the prompt because the server thinks
>>>>>>     something that would interfere with the feedback
>>>>>>     ear/mouth/finger control happened. What this
>>>>>>
>>     says to me is
>>
>>>>>>     that it isn't necessarily a good idea for the client to
>>>>>>     have more knobs to control the server
>>>>>>
>>     (especially if those
>>
>>>>>>     knows are just more value/policy input ungrounded in any
>>>>>>     physics/ acoustics). On the other hand, having the
>>>>>>
>> server
>>
>>>>>>     tell the client more about what it thinks is going on is
>>>>>>     probably valuable.
>>>>>>
>>>>>>     So, Coming to the point after this long rambling
>>>>>>     introduction, I think it would in fact be useful for the
>>>>>>     START-Of-SPEECH event to indicate some extra
>>>>>>
>> information,
>>
>>>>>>     for example:
>>>>>>     a) I got something enough above the noise floor
>>>>>>
>>     to qualify
>>
>>>>>>     for exceeding the "Sensisitvity" parameter you sent
>>>>>>
>> in on
>>
>>>>>>     the request but I really can't tell what it is
>>>>>>
>>     (could be a
>>
>>>>>>     hippopatmus fart, or a siren in the background,
>>>>>>
>>     or captain
>>
>>>>>>     crunch trying to whistle DTMF).
>>>>>>     b) I think I'm hearing speech
>>>>>>     c) I think I'm hearing DTMF
>>>>>>
>>>>>> Though I agree with your former part of your
>>>>>>
>>     response. I am not
>>
>>>>>> sure I agree with your proposed solution.
>>>>>> The way I see this problem is that, it is more of what
>>>>>>
>>> constitues a
>>>
>>>>>> barge-in event. This boils down to whether it is
>>>>>>
>>     speech, DTMF or
>>
>>>>>> both.
>>>>>> This is inturn boils down to what type of recognizer
>>>>>>
>>> resource we are
>>>
>>>>>> using, dtmf-recog, speech-recog, and speech-only-recog(we
>>>>>>
>>> don't have
>>>
>>>>>> this and I don't think we should add it, but think
>>>>>>
>>     of this as a
>>
>>>>>> place holder that explains the concept).
>>>>>>
>>>>>> A client knowing what type of barge-in happenned, does
>>>>>>
>>> not impact
>>>
>>>>>> the
>>>>>> barge-in operation itself as it may be too late(for
>>>>>>
>>     the optimized
>>
>>>>>> barge-in case). It may have other use cases, and if we
>>>>>>
>>> can identify
>>>
>>>>>> them, I don't mind adding support for the
>>>>>>
>>     START-OF-SPEECH event to
>>
>>>>>> say what type of barge-in happenned. But that itself does
>>>>>>
>> not
>>
>>> solve the
>>>
>>>>>> original problem raised. Refer to my previous response.
>>>>>>
>>>>>> The solution lies in defining what what is a
>>>>>>
>>     barge-in event.
>>
>>>>>> That  boils
>>>>>> down to what type of recognition is happenning,
>>>>>>
>>> dtmf-only, speech-
>>>
>>>>>> dtmf
>>>>>> or speech-only. We do not support speech-only as a
>>>>>>
>>     resource today,
>>
>>>>>> the question is do we need a header to force it.
>>>>>>
>>>>>> Sarvi
>>>>>>
>>>>>>
>>>>>>
>>>>>>>    2. You would benefit from the client knowing
>>>>>>>
>>> what caused
>>>
>>>>>> the
>>>>>>
>>>>>>> barge-in, DTMF Vs speech.
>>>>>>>
>>>>>>> DB> See previous comment. And previous e-mail:
>>>>>>>
>>     either add an
>>
>>>>>>> inputmodes header (taking value speech, dtmf, both) to
>>>>>>>
>>>>>>     the RECOGNIZE
>>>>>>
>>>>>>> request or add a header to the START-OF-SPEECH event
>>>>>>>
>>>>>>     indicating DTMF
>>>>>>
>>>>>>> or speech.
>>>>>>>
>>>>>>>
>>>>>>     I'm leaning in your direction on this latter point - as
>>>>>>     should be evident from what I wrote above.
>>>>>>
>>>>>>
>>>>>>> Sarvi
>>>>>>>
>>>>>>>     -----Original Message-----
>>>>>>>     From: speechsc-bounces@ietf.org
>>>>>>>     [mailto:speechsc-bounces@ietf.org] On Behalf Of
>>>>>>>
>>> David R
>>>
>>>>>> Oran
>>>>>>
>>>>>>>     Sent: Tuesday, July 05, 2005 5:24 AM
>>>>>>>     To: Klaus Reifenrath
>>>>>>>     Cc: 'speechsc@ietf.org'
>>>>>>>     Subject: Re: [Speechsc] START-OF-SPEECH in
>>>>>>>
>>> DTMF-only mode
>>>
>>>>>>>
>>>>>>>
>>>>>>>     On Jul 5, 2005, at 3:46 AM, Reifenrath,
>>>>>>>
>>     Klaus wrote:
>>
>>>>>>>
>>>>>>>
>>>>>>>> The current spec is not clear when
>>>>>>>>
>>> START-OF-SPEECH need
>>>
>>>>>>>     to be send in
>>>>>>>
>>>>>>>> the following scenarios:
>>>>>>>> A) The client requested a DTMF Recognizer. Is
>>>>>>>>
>> the
>>
>>>>>>>     START-OF-SPEECH
>>>>>>>
>>>>>>>> event send to the client also if speech
>>>>>>>>
>>     was detected?
>>
>>>>>>>     I suspect so, since one of the prime purposes
>>>>>>>
>>> is to enable
>>>
>>>>>>>     client- mediated barge-in handling. However, if
>>>>>>>
>> the
>>
>>>>>>>     recognizer is in fact only capable of
>>>>>>>
>>     recognizing DTMF
>>
>>>>>>>     then it may in fact not report anythin
>>>>>>>
>>     unless it's using
>>
>>>>>>>     some primitive thresholding machinery, like a SN
>>>>>>>
>>>>>> threshold.
>>>>>>
>>>>>>>> B) The client requested a Speech
>>>>>>>>
>>     Recognizer, but only
>>
>>>>>>>     activated DTMF
>>>>>>>
>>>>>>>> grammars. Is the START-OF-SPEECH event
>>>>>>>>
>>     send to the
>>
>>>>>>>     client also if
>>>>>>>
>>>>>>>> speech was detected?
>>>>>>>>
>>>>>>>     Again, I'd say yes, for the same reason as above.
>>>>>>>
>>>>>>>> I think in both cases START-OF-SPEECH should
>>>>>>>>
>> only
>>
>>>>>>     be send after
>>>>>>
>>>>>>>> detecting a DTMF digit (see Figure 12 of
>>>>>>>>
>> VoiceXML
>>
>>>>>>     2.0: http://
>>>>>>
>>>>>>>> www.w3.org/TR/voicexml20/#dmlATiming).
>>>>>>>>
>>>>>>>     We seem to have reached different
>>>>>>>
>>     conclusions. I'd be
>>
>>>>>>>     interested in why you think my analysis
>>>>>>>
>>     above is wrong.
>>
>>>>>>>
>>>>>>>     Dave.
>>>>>>>
>>>>>>>
>>>>>>>> Klaus
>>>>>>>>
>>>>>>>> _______________________________________________
>>>>>>>> Speechsc mailing list
>>>>>>>> Speechsc@ietf.org
>>>>>>>> https://www1.ietf.org/mailman/listinfo/speechsc
>>>>>>>>
>>>>>>>>
>>>>>>>
>>>>>>>     _______________________________________________
>>>>>>>     Speechsc mailing list
>>>>>>>     Speechsc@ietf.org
>>>>>>>     https://www1.ietf.org/mailman/listinfo/speechsc
>>>>>>>
>>>>>>>
>>>>>>> _______________________________________________
>>>>>>> Speechsc mailing list
>>>>>>> Speechsc@ietf.org
>>>>>>> https://www1.ietf.org/mailman/listinfo/speechsc
>>>>>>>
>>>>>>>
>>>>>>
>>>>>>
>>>>>> _______________________________________________
>>>>>> Speechsc mailing list
>>>>>> Speechsc@ietf.org
>>>>>> https://www1.ietf.org/mailman/listinfo/speechsc
>>>>>>
>>>>>>
>>>>>>
>>>>>
>>>>> _______________________________________________
>>>>> Speechsc mailing list
>>>>> Speechsc@ietf.org
>>>>> https://www1.ietf.org/mailman/listinfo/speechsc
>>>>>
>>>>
>>>>
>>>
>>> _______________________________________________
>>> Speechsc mailing list
>>> Speechsc@ietf.org
>>> https://www1.ietf.org/mailman/listinfo/speechsc
>>>
>>>
>>
>>
>>     _______________________________________________
>>     Speechsc mailing list
>>     Speechsc@ietf.org
>>     https://www1.ietf.org/mailman/listinfo/speechsc
>>
>>
>> _______________________________________________
>> Speechsc mailing list
>> Speechsc@ietf.org
>> https://www1.ietf.org/mailman/listinfo/speechsc
>>
>>
>> _______________________________________________
>> Speechsc mailing list
>> Speechsc@ietf.org
>> https://www1.ietf.org/mailman/listinfo/speechsc
>>
>>
>
> _______________________________________________
> Speechsc mailing list
> Speechsc@ietf.org
> https://www1.ietf.org/mailman/listinfo/speechsc
>
>
>
>

_______________________________________________
Speechsc mailing list
Speechsc@ietf.org
https://www1.ietf.org/mailman/listinfo/speechsc



From speechsc-bounces@ietf.org Mon Jul 11 17:22:04 2005
Received: from localhost.localdomain ([127.0.0.1] helo=megatron.ietf.org)
	by megatron.ietf.org with esmtp (Exim 4.32)
	id 1Ds5j2-0003UF-G2; Mon, 11 Jul 2005 17:22:04 -0400
Received: from odin.ietf.org ([132.151.1.176] helo=ietf.org)
	by megatron.ietf.org with esmtp (Exim 4.32) id 1Ds5j1-0003Sk-Kv
	for speechsc@megatron.ietf.org; Mon, 11 Jul 2005 17:22:03 -0400
Received: from ietf-mx.ietf.org (ietf-mx [132.151.6.1])
	by ietf.org (8.9.1a/8.9.1a) with ESMTP id RAA17498
	for <speechsc@ietf.org>; Mon, 11 Jul 2005 17:22:00 -0400 (EDT)
Received: from fw01.db01.voxpilot.com ([212.17.54.82] helo=mail.voxpilot.com)
	by ietf-mx.ietf.org with esmtp (Exim 4.43) id 1Ds6B3-0001hB-7g
	for speechsc@ietf.org; Mon, 11 Jul 2005 17:51:04 -0400
Received: from daburkewxp (ppp-155.net-207.magic.fr [62.210.222.155])
	by mail.voxpilot.com (Postfix) with ESMTP
	id 1847C214041; Mon, 11 Jul 2005 21:21:27 +0000 (GMT)
Message-ID: <028f01c5865e$7f92c4a0$038ae9d5@db01.voxpilot.com>
From: "Dave Burke" <david.burke@voxpilot.com>
To: "David R Oran" <oran@cisco.com>, "Pierre Forgues" <forgues@nuance.com>
References: <7DE7C4EF3B7C8B4B82955191378290D802ED4176@mtb1exch01.nuance.com>
	<C7673C4C-0624-47A1-B1DF-115FB284785A@cisco.com>
Subject: Re: [Speechsc] START-OF-SPEECH in DTMF-only mode -> summary(?)
Date: Mon, 11 Jul 2005 22:21:25 +0100
MIME-Version: 1.0
Content-Type: text/plain; format=flowed; charset="iso-8859-1";
	reply-type=response
Content-Transfer-Encoding: 7bit
X-Priority: 3
X-MSMail-Priority: Normal
X-Mailer: Microsoft Outlook Express 6.00.2900.2180
X-MimeOLE: Produced By Microsoft MimeOLE V6.00.2900.2180
X-Spam-Score: 0.0 (/)
X-Scan-Signature: fcb81a552e93de47ff406e53e992da89
Content-Transfer-Encoding: 7bit
Cc: speechsc@ietf.org, "Shanmugham, Saravanan" <sarvi@cisco.com>,
	Eric Burger <eburger@brooktrout.com>
X-BeenThere: speechsc@ietf.org
X-Mailman-Version: 2.1.5
Precedence: list
List-Id: Speech Services Control Working Group <speechsc.ietf.org>
List-Unsubscribe: <https://www1.ietf.org/mailman/listinfo/speechsc>,
	<mailto:speechsc-request@ietf.org?subject=unsubscribe>
List-Post: <mailto:speechsc@ietf.org>
List-Help: <mailto:speechsc-request@ietf.org?subject=help>
List-Subscribe: <https://www1.ietf.org/mailman/listinfo/speechsc>,
	<mailto:speechsc-request@ietf.org?subject=subscribe>
Sender: speechsc-bounces@ietf.org
Errors-To: speechsc-bounces@ietf.org

I don't think it is possible in practice to separate what input type is 
considered potential barge-in and what input type is being recognised 
because a recognition hypothesis is always generated when a barge-in occurs. 
For example, imagine a client invoked a speechrecog and asked it to only 
consider DTMF for barge-in. Once speech is detected, although no barge-in 
would happen, a RECOGNITION-COMPLETE will result after 
Speech-Incomplete-Timeout milliseconds of silence thus ending the 
recognition state machine with a 'nomatch'. The converse example applies for 
a speech-only recognition (DTMF-Interdigit-Timeout replaces the 
Speech-Incomplete-Timeout).

Although I prefer the idea of an InputModes header in RECOGNIZE (it's the 
cleanest solution for VoiceXML implementors), I am willing to compromise on 
using the grammar type to determine the input types (in what follows I use 
input types to mean both what is considered potential barge-in and what type 
of input is to be recognised). The main issue with this approach is 
incompatibility with VoiceXML (recall VoiceXML grammar activation is 
independent of what inputmodes are set).

This incompatibility results in side-effects that, in my opinion, are 
inconsequential for real applications:
    a. speech barge-in (noise) and nomatch will not happen for 
inputmodes="both" when no speech grammars are activated
    b. dtmf barge-in and nomatch will not happen for inputmodes="both" when 
no dtmf grammars are activated
    c. grammars with errors will not be detected if the grammar type is not 
compatible with the inputmodes property
A workaround for the really conscientious VoiceXML platform is to activate a 
dummy grammar (maybe with the NULL rule) of type speech (dtmf) when 
inputmodes="both" is set but no speech (dtmf) grammar is activated in the 
application.

-----------------

Summarising the proposed changes again  (refined a little and to avoid 
trawling though this massive thread!):

1. For a speechrecog resource, the types of the activated grammars determine 
the input types the recogniser processes and considers for potential 
barge-in.

2. Clarify a dtmfrecog only processes DTMF and hence can only generate 
barge-in / START-OF-SPEECH for DTMF inputs

Nice-to-have:

3. Change START-OF-SPEECH to START-OF-INPUT

4. Add a header of InputType to START-OF-INPUT. Current specified values are 
"dtmf" or "speech".

-----------------

Dave

----- Original Message ----- 
From: "David R Oran" <oran@cisco.com>
To: "Pierre Forgues" <forgues@nuance.com>
Cc: <speechsc@ietf.org>; "Shanmugham, Saravanan" <sarvi@cisco.com>; "Eric 
Burger" <eburger@brooktrout.com>; "Dave Burke" <david.burke@voxpilot.com>
Sent: Monday, July 11, 2005 2:15 PM
Subject: Re: [Speechsc] START-OF-SPEECH in DTMF-only mode -> summary(?)


>
> On Jul 8, 2005, at 11:36 AM, Pierre Forgues wrote:
>
>>
>>
>> -----Original Message-----
>> From: speechsc-bounces@ietf.org [mailto:speechsc-bounces@ietf.org] On
>> Behalf Of David R Oran
>> Sent: Friday, July 08, 2005 10:05 AM
>> To: Dave Burke
>> Cc: speechsc@ietf.org; Shanmugham, Saravanan; Eric Burger
>> Subject: Re: [Speechsc] START-OF-SPEECH in DTMF-only mode -> summary (?)
>>
>>
>> On Jul 8, 2005, at 6:59 AM, Dave Burke wrote:
>>
>>
>>> I agree with your opinion on what constitutes a barge-in. Based on
>>> Klaus' e-mail, I am more concerned that using the grammar type to
>>> decide what constitutes a barge-in is going to make VoiceXML
>>> implementations difficult.
>>>
>>> It is easy to map VoiceXML application selected inputmodes to MRCP
>>> resources:
>>> a. inputmodes="dtmf" -> Use a dtmfrecog
>>> b. inputmodes="speech" -> Use a "speech-only-recog"
>>> c. inputmodes="both" -> Use a speechrecog (or a combination of
>>> "speech-only-recog" + dtmfrecog)
>>>
>>> The only problem is (b). While not as important as being able to do
>>> DTMF-only recognition, I believe we DO need to support speech-only
>>> recongition so as (a) to avoid unnecessary limitations in VUI
>>> design, and (b) to facilitate VoiceXML implementations. I think we
>>> should add a header to RECOGNIZE so the client is able to always
>>> _explicitly_ set the inputmodes.
>>>
>>>
>> I'm not sure I buy this, since what the recognizer is looking for and
>> what constitutes an input that should be considered a potential barge-
>> in strike me as independent. I'm similarly not persuaded that we need
>> the flexibility to set the input mode independently of the grammar,
>> since it leads to all sorts of inconsistent states (e.g. speech-only
>> grammar with an input-mode of dtmf). On the other hand I can see the
>> need for the client to specify what sorts of input out to generate
>> the start-of-input (nee start-of-speech) event on.
>>
>> Pmf> The mechanism for a client to specify the type of input is using
>> grammars.  These can have DTMF and/or speech requirements.  If you add
>> the complexity of an independent header to specify the input mode then
>> you will need to document the behavior for inconsistencies.
>>
> Right. I generally agree with Pierre, which is why I proposed a 
> compromise position where the client can use a header to specify what 
> types of input should be considered the start of something  potentially 
> interesting from a barge-in point of view, as opposed to  a directive to 
> the server for what to recognize.
>
> I don't think this is crucial functionality, but it does make for  more 
> consistent behavior between the optimized and non-optimized  barge-in 
> cases. In fact, I can see cases where the client doesn't  want any 
> barge-in processing to happen. Imagine an application called  "mark the 
> musical phrases", where the client listens to playout of a  musical score, 
> and hits various DTMF buttons to indicate markers  (e.g. end-of-phrase, 
> end-of-theme, key-change) while listening. For  this application you want 
> to make sure that barge-in doesn't stop the  music!
>
>>
>>> In retrospect, I don't like the idea of multiple START-OF-SPEECH
>>> events being generated from the same media resource. This is
>>> because one assumes that the START-OF-SPEECH should be of the same
>>> type as the hypothesis returned in the RECOGNITION-COMPETE message
>>> - most implementations, on hearing one input mode type disable the
>>> recogniser of the other type.
>>>
>> I'm not sure I follow this logic, but I'm not wedded to the idea of
>> allowing multiple events. It was trying to solve the problem of
>> ambiguity around what the recognizer was hearing and the idea that
>> the client might care. If you don't think the client will ever care,
>> then we don't need the capability.
>> Pmf> I agree we should not have multiple SOS events.  Too much  chatter.
>>
> I'm not concerned about the "chatter" since this is a low-bandwidth 
> control channel even with multiple events, but as I said it's a small 
> point and I'm happy to concede.
>
>>
>>
>>> I do like David's idea of renaming START-OF-SPEECH to something
>>> like START-OF-INPUT and carrying a type header because it is neater
>>> and more extensible.
>>>
>>> So in summary, I propose we modify the spec to:
>>>
>>> 1. Clarify what constitutes a barge-in for a dtmfrecog and
>>> speechrecog as per Sarvi's e-mail (and in agreement with Klaus' for
>>> dtmfrecog).
>>>
>>>
>> Hmmm, ok, but I think the issue is actually clarifying what the
>> recognizer declares as "interesting input", which it may decide also
>> constitutes a barge-in in the optimized case, and in either case
>> reports to the client that it heard.
>> Pmf> I'm not going to argue changing the name of the event, but my
>> preference would be to keep existing method names unless there is a
>> clear reason for changing - which I do not see here.
>>
> I do see a clear reason, since the thing that you see the start of  may 
> not be speech. I like (I think it was Dave's suggestion) "start- 
> of-input:" since it mirrors the other method names.
>>
>>> 2. Specifiy an InputModes header to RECOGNIZE (defaults to "both",
>>> can also be "speech" or "DTMF"). Setting to speech for a
>>> speechrecog results in the hypothesised "speech-only-recog".
>>> Setting to DTMF for a speechrecog is equivalent to using a
>>> dtmfrecog. Edge cases: Setting to speech for a dtmfrecog will
>>> result in a noinput as will setting to DTMF for a speechrecog which
>>> does not support DTMF.
>>>
>>>
>> I'm ok with having a header for the client to tell the server when it
>> would like the start-of-input event to be generated and what the
>> client considers to be the "interesting input" that the recognizer
>> should use to do the discrimination, and possibly do barge-in
>> processing in the optimized case.
>>
>> I'm less ok with the proposed domain of values, since it will have
>> extensibility problems. Especially problematical is have a code point
>> of "both" since that will be ambiguous if we even define a "music"
>> recognizer, or a "gesture" recognizer using video input. Here's my
>> counter-proposal:
>>
>> Create a header on the recognizer resource requests called "start-
>> input-on:". Define a registry of values, with the following three
>> values initially defined:
>>      - dtmf
>>      - speech
>>      - music
>>
>> The semantics would be that the resource is to generate a "start-of-
>> input" event, apply the defined grammars to what follows, and do any
>> local optimized barge-in processing if ANY of the enumerated input
>> types is detected. That way you can say
>>      "start-input-on: dtmf" if you want to just  dtmf,
>>      "start-input-on: speech" if you want just speech (dtmf input
>> would be ignored even if the grammar supported dtmf)
>>      "start-input- on: speech, dtmf" if you wanted both
>> etc.
>>
>> Pmf> Doesn't this come back to the previous point of having two
>> independent ways to specify input modes?  Through grammars and this
>> proposed new header?  If you add this then you will have to document
>> behavior on inconsistencies.  Anyway, if you proceed then I would
>> definitely make this optional
>>
> I hope I explained the intent above. This header would not define the 
> input types that the recognize would process, but rather the input  types 
> the client wants it to consider potential barge-in and hence  generate the 
> "start-of-input:" event for and do any local barge-in  optimizations for.
>
>
> Dave.
>
>>
>>> 3. Change START-OF-SPEECH to START-OF-INPUT. Clarify it occurs once
>>> and coincides with barge-in.
>>>
>>>
>> Ok for the "only once", but I'd like to tighten it up to talk about
>> more than just barge-in, as I suggested above.
>>
>>
>>> 4. Add a header of InputType to START-OF-INPUT. Current specified
>>> values are "dtmf" or "speech".
>>>
>>>
>> Ok, with slight modification. Say that the syntax of the "input-type"
>> header is a single-valued subset of the registered value(s) that are
>> defined for the start-input-on: header.
>>
>> Comments?
>>
>> Dave O. (technical hat on, chair hat off).
>>
>>
>>> Dave
>>>
>>>
>>>
>>> ----- Original Message ----- From: "Shanmugham, Saravanan"
>>> <sarvi@cisco.com>
>>> To: "Eric Burger" <eburger@brooktrout.com>; <speechsc@ietf.org>
>>> Sent: Thursday, July 07, 2005 5:52 PM
>>> Subject: RE: [Speechsc] START-OF-SPEECH in DTMF-only mode -> summary
>>> (?)
>>>
>>>
>>> I don't think generating multiple START-OF-SPEECH events is a
>>> solution.
>>> We still haven't addressed, what constitutes a barge-in, for the
>>> optimised case. That should also be the single point when a single
>>> START-OF-SPEECH(or whathever else you want to name it) should be
>>> generated.
>>>
>>> That point, in my opinion should be
>>>   1. For "dtmf-recog" resources should be the beginning of a DTMF key
>>> press.
>>>   2. For "speech-recog" resources should be the beginning of a DTMF
>>> key
>>> press or the beginning of speech. This is should be irrespective of
>>> what
>>> type of grammar is being used. Coz even numbers only grammars can
>>> still
>>> be spoken and hence cannot be assumed to be a cue for DTMF only
>>> recognition.
>>>   3. For "speech-only-recog" resources(which are not defined today)
>>> the
>>> time to barge-in is the beginning of speech. I don't see a need for
>>> such
>>> a resource today. But I am mentioning this for completeness.
>>>
>>> Sarvi
>>>
>>>
>>>     -----Original Message-----
>>>     From: speechsc-bounces@ietf.org
>>>     [mailto:speechsc-bounces@ietf.org] On Behalf Of Eric Burger
>>>     Sent: Thursday, July 07, 2005 5:32 AM
>>>     To: speechsc@ietf.org
>>>     Subject: RE: [Speechsc] START-OF-SPEECH in DTMF-only mode
>>>     -> summary(?)
>>>
>>>     Can we declare consensus?
>>>
>>>
>>>> -----Original Message-----
>>>> From: speechsc-bounces@ietf.org
>>>>
>>>     [mailto:speechsc-bounces@ietf.org]
>>>
>>>> Sent: Wednesday, July 06, 2005 10:45 AM
>>>> To: Dave Burke
>>>> Cc: speechsc@ietf.org; Shanmugham, Saravanan; Klaus Reifenrath
>>>> Subject: Re: [Speechsc] START-OF-SPEECH in DTMF-only mode ->
>>>> summary(?)
>>>>
>>>> I think we're getting close. I though about snipping out
>>>>
>>>     some pieces
>>>
>>>> to cut down the text, but I realized the context is
>>>>
>>>     still needed. See
>>>
>>>> inline.
>>>>
>>>> On Jul 6, 2005, at 10:29 AM, Dave Burke wrote:
>>>>
>>>>
>>>>> Inline.
>>>>>
>>>>> Dave
>>>>>
>>>>> ----- Original Message ----- From: "David R Oran"
>>>>>
>>>     <oran@cisco.com>
>>>
>>>>> To: "Dave Burke" <david.burke@voxpilot.com>
>>>>> Cc: <speechsc@ietf.org>; "Shanmugham, Saravanan"
>>>>>
>>>     <sarvi@cisco.com>;
>>>
>>>>> "Klaus Reifenrath" <Klaus.Reifenrath@Scansoft.com>
>>>>> Sent: Wednesday, July 06, 2005 1:07 PM
>>>>> Subject: Re: [Speechsc] START-OF-SPEECH in DTMF-only
>>>>>
>>>     mode -> summary
>>>
>>>>> (?)
>>>>>
>>>>>
>>>>>
>>>>>
>>>>>>
>>>>>> On Jul 6, 2005, at 5:31 AM, Dave Burke wrote:
>>>>>>
>>>>>>
>>>>>>
>>>>>>> + Attempting to summarise:
>>>>>>>
>>>>>>> 1. START-OF-SPEECH is useful for the client to know when
>>>>>>>
>>>> to stop
>>>>
>>>>>>> playing prompts the in non-optimised case 2.
>>>>>>>
>>>     START-OF-SPEECH is
>>>
>>>>>>> useful for the client to calculate the bargin  time
>>>>>>>
>>>     (e.g. VoiceXML
>>>
>>>>>>> 2.1 <mark>)
>>>>>>> 3. In the optimised case, a bargin automatically
>>>>>>>
>>>     stops prompt
>>>
>>>>>>> playing (assuming prompts barginable) 4. Because of
>>>>>>>
>>>     the previous
>>>
>>>>>>> point, the question of what input type  caused
>>>>>>>
>>>     bargin is different
>>>
>>>>>>> and less important to
>>>>>>>
>>>> what input
>>>>
>>>>>>> type(s)  the recogniser is listening for
>>>>>>>
>>>>>>>
>>>>>> I'm not sure I follow point 4. Could you elaborate?
>>>>>>
>>>>>>
>>>>>
>>>>> DB> Adding a parameter to the START-OF-SPEECH event would
>>>>>
>>>> certainly
>>>>
>>>>> allow the client to ignore (i.e. let prompts continue
>>>>>
>>>     playing) the
>>>
>>>>> event if the event type is not of interest (e.g. the
>>>>>
>>>     client would
>>>
>>>>> ignore speech start events when it is interested only
>>>>>
>>>     in  DTMF start
>>>
>>>>> events). This _only_ works for the non-optimised case,
>>>>>
>>>     however. For
>>>
>>>>> the optimised case, assuming START-OF-SPEECH
>>>>>
>>>> coincides
>>>>
>>>>> with the bargin signal to the speechsynth, prompts
>>>>>
>>>     will stop playing
>>>
>>>>> for inputs that the client might not be interested
>>>>>
>>>> in (e.g.
>>>>
>>>>> a speech input will stop prompts playing even if the client
>>>>>
>>>> is only
>>>>
>>>>> interested in DTMF).
>>>>>
>>>>>
>>>> OK, now I get it. There's a need for the client to both
>>>>
>>>     handle the
>>>
>>>> non-optimized case itself, and influence or at least
>>>>
>>>     have a clue what
>>>
>>>> the server is going to do in the optimized case.
>>>>
>>>>
>>>>>
>>>>>
>>>>>>
>>>>>>
>>>>>>
>>>>>>> 5. Currently, MRCPv2 has no way of indicating what
>>>>>>>
>>>     input type(s) a
>>>
>>>>>>> recogniser is listening for
>>>>>>>
>>>>>>>
>>>>>>>
>>>>>> Do you mean exactly this, or do you mean "for the client to
>>>>>> indicate  to the resource what input types it should
>>>>>>
>>>     look for"?
>>>
>>>>>>
>>>>>>
>>>>>
>>>>> DB> Yes exactly - apologies for not being clear.
>>>>>
>>>>>
>>>>>
>>>>>>
>>>>>>
>>>>>>
>>>>>>> + Why implement 5?
>>>>>>>
>>>>>>> i. Noisy case: Need DTMF-only recognition (and may
>>>>>>>
>>>     only have a
>>>
>>>>>>> speechrecog)
>>>>>>>
>>>>>>>
>>>>>> I'm having difficulty following the logic of why
>>>>>>
>>>     noise would
>>>
>>>>>> necessarily trigger START-OF-SPEECH if you were listening for
>>>>>> speech  but not DTMF. I suppose you can use a more forgiving
>>>>>> discriminator if  all you need to tell is if you're
>>>>>>
>>>     getting DTMF,
>>>
>>>>>> but I've had a number  of real-world cases where wind
>>>>>>
>>>     noise was
>>>
>>>>>> detected as DTMF, and  there's always the ambiguity
>>>>>>
>>>     when you have
>>>
>>>>>> Captain Crunch on the  phone. In either case in the
>>>>>>
>>>     non-optimized
>>>
>>>>>> case it's the client who  gets to decide whether an
>>>>>>
>>>     event should be
>>>
>>>>>> interpreted as barge-in or  not, so it seems an
>>>>>>
>>>     aesthetic protocol
>>>
>>>>>> design decision whether the  client tells the server
>>>>>>
>>>     ahead of time
>>>
>>>>>> what circumstances to generate  the START-OF-SPEECH
>>>>>>
>>>     event for, or
>>>
>>>>>> whether the event gets generated  and the client
>>>>>>
>>>     decides based on
>>>
>>>>>> what's in the event whether it should  be
>>>>>>
>>>> treated
>>>>
>>>>>> as barge- in.
>>>>>>
>>>>>> I suppose one could make the argument that because
>>>>>>
>>>     the spec implies
>>>
>>>>>> that the event can only be generated once per request
>>>>>>
>>>     that if a
>>>
>>>>>> DTMF/ speech capable recognizer first hears
>>>>>>
>>>> enough noise
>>>>
>>>>>> to think it's  hearing speech and later hears DTMF,
>>>>>>
>>>     the client will
>>>
>>>>>> declare barge-in  when the event comes and not when he DTMF
>>>>>> actually gets heard.
>>>>>>
>>>>>> If that's deemed a problem, we can still handle that
>>>>>>
>>>     in the design
>>>
>>>>>> where the server just reports what it's hearing by allowing
>>>>>> multiple  events to be generated during a single request.
>>>>>>
>>>>>> Between the approach just outlined above, and an
>>>>>>
>>>     approach where the
>>>
>>>>>> client provides a filter for whether to generate the
>>>>>>
>>>     event or not,
>>>
>>>>>> I  have a mild preference (based on aesthetics rather
>>>>>>
>>>     than some
>>>
>>>>>> hard  engineering tradeoff) for the approach where
>>>>>>
>>>> the server
>>>>
>>>>>> just reports  what it's hearing.
>>>>>>
>>>>>>
>>>>>> Having had some useful exchanges on this topic, it
>>>>>>
>>>     also is becoming
>>>
>>>>>> apparent to me that this event is poorly named, and
>>>>>>
>>>     we should
>>>
>>>>>> consider renaming it to "INTERESTING-INPUT-HEARD" or
>>>>>>
>>>     something akin
>>>
>>>>>> to that, because as others have pointed out, a
>>>>>>
>>>     DTMF-only recognizer
>>>
>>>>>> will never detect "start of speech".
>>>>>>
>>>>>> Another consideration to fold into the design choice is
>>>>>> extensibility. Bear with me through a little
>>>>>>
>>>     gedankenexperiment.
>>>
>>>>>>
>>>>>> Suppose we want to define a new recognizer type,
>>>>>>
>>>     which I'll call
>>>
>>>>>> the "name that tune" recognizer. The client plays
>>>>>>
>>>     music to the
>>>
>>>>>> server and  the server recognizes musical notes. The
>>>>>>
>>>     grammar is a
>>>
>>>>>> standard  musical notation, augmented with a semantic
>>>>>> interpretation that  transforms the notes into the
>>>>>>
>>>     title of the
>>>
>>>>>> tune and provides that as  an answer.
>>>>>>
>>>>>> First, there's no speech involved (or is there...hang
>>>>>>
>>>     on a minute).
>>>
>>>>>> Second, in order to accommodate the "name that tune"
>>>>>> recognizer, we'd have to extend both the client and
>>>>>>
>>>     the server to
>>>
>>>>>> undetstand a  directive as to whether to recognize
>>>>>>
>>>     music or now,
>>>
>>>>>> inaddition to what  the server already knows what to
>>>>>>
>>>     do based on
>>>
>>>>>> the grammar. If you  follow my logic above, whether
>>>>>>
>>>     or not we do
>>>
>>>>>> that, we have to extend  "start-of-speech" to say
>>>>>>
>>>     "I'm hearing
>>>
>>>>>> music". So far fairly  straightforward, but let me
>>>>>>
>>>     now throw in the
>>>
>>>>>> pathological twist.
>>>>>>
>>>>>> Suppose what I feed to a  combined music/speech
>>>>>>
>>>     recognizer is a
>>>
>>>>>> work  in sprechstimme (spoken music), like the
>>>>>>
>>>     "Geographical Fugue"
>>>
>>>>>> (aside:  this is a wonderful piece of music I highly
>>>>>>
>>>     recommend to
>>>
>>>>>> anyone  interested in small ensemble singing). In
>>>>>>
>>>     this case, the
>>>
>>>>>> tune could  be named by either doing speech or music
>>>>>>
>>>     recognition.
>>>
>>>>>> Why is there  any need for the client to constrain
>>>>>>
>>>     the server as to
>>>
>>>>>> which it tries  to do when it's
>>>>>>
>>>> already
>>>>
>>>>>> told the server what it wants through the  grammar?
>>>>>>
>>>>>> A few other comments below
>>>>>>
>>>>>>
>>>>>>
>>>>>>> ii. Flexibility: Want speech-only recognition
>>>>>>>
>>>     (because a second
>>>
>>>>>>> recogniser is doing hotword on DTMF)
>>>>>>>
>>>>>>>
>>>>>>>
>>>>>> I don't see how flexibility is affected by this deisgn
>>>>>>
>>>> choice. If
>>>>
>>>>>> that's what you want, feed the speech-only recognizer
>>>>>>
>>>     a grammar
>>>
>>>>>> without any DTMF rules.
>>>>>>
>>>>>>
>>>>>>
>>>>>>> + How to implement 5?
>>>>>>>
>>>>>>> a. Implicitly:
>>>>>>>    - dtmfrecog: always DTMF-only recognition
>>>>>>>    - speechrecog: depends on active grammar type
>>>>>>>
>>>>>>>> if a dtmf grammar is active then DTMF input is "on"
>>>>>>>> if a speech grammar is active then speech
>>>>>>>>
>>>     input is "on"
>>>
>>>>>>>
>>>>>>> b. Explicitly:
>>>>>>>    - Add inputmodes header to RECOGNIZE
>>>>>>>
>>>>>>> Option a is David's "do what I mean case"; option b is
>>>>>>>
>>>> the extra
>>>>
>>>>>>> dial for the client.
>>>>>>>
>>>>>>>
>>>>>>>
>>>>>> Actually, that's not the point I was making with "do
>>>>>>
>>>     what I mean",
>>>
>>>>>> but it's not essential to the discussion so let's move on.
>>>>>>
>>>>>>
>>>>>>
>>>>>>> It is worth noting that VoiceXML uses option b. This
>>>>>>>
>>>     allows one to
>>>
>>>>>>> activate both speech grammars and DTMF grammars (and
>>>>>>>
>>>> therefore
>>>>
>>>>>>> be informed of any errors in the grammars at activation
>>>>>>>
>>>> time) but
>>>>
>>>>>>> independently turn on whichever input mode you like e.g.
>>>>>>>
>>>> perhaps
>>>>
>>>>>>> start with "both" then change to "dtmf".
>>>>>>>
>>>>>>>
>>>>>>>
>>>>>> I'm not sure the VXML precedent is relevant here,
>>>>>>
>>>     because the
>>>
>>>>>> application behind VXML is working a different part of the
>>>>>>
>>>> problem
>>>>
>>>>>> -  how to traverse a TUI dialog based on different
>>>>>>
>>>     parts of the
>>>
>>>>>> input  space. In fact, I suspect that the VXML: choice was
>>>>>> conditioned more  by limitations at the time it was
>>>>>>
>>>> specified than
>>>>
>>>>>> an underlying good  design choice. Clearly having to
>>>>>>
>>>     specify this
>>>
>>>>>> in VXML make the job of  handling a TUI with nodes
>>>>>>
>>>     like "Say or
>>>
>>>>>> press 5" harder rather than  easier.
>>>>>>
>>>>>>
>>>>>
>>>>> DB> The VoiceXML edge-case is pretty weird so it's not a major
>>>>> concern. My main concern is that the client can
>>>>>
>>>     indicate, somehow,
>>>
>>>>> what the input modes are.
>>>>>
>>>>>
>>>>>
>>>>>>
>>>>>> Summing up, while I don't feel strongly one way or
>>>>>>
>>>     the other, I
>>>
>>>>>> have  a preference for handling this as follows:
>>>>>>
>>>>>> a) Rename "START-OF-SPEECH" to
>>>>>>
>>>     "INTERESTING-INPUT-RECEIVED" or
>>>
>>>>>> something equivalent.
>>>>>> b) Include a parameter in the event saying what was
>>>>>>
>>>     interesting
>>>
>>>>>> about  the input you received, with a registry of values
>>>>>>
>>> which
>>>
>>>>>> includes:
>>>>>>     - signal above noise floor
>>>>>>     - speech
>>>>>>     - dtmf
>>>>>>     - (possibly) music
>>>>>> c) allow the event to be generated multiple times
>>>>>>
>>>     during a request
>>>
>>>>>>
>>>>>>
>>>>>>
>>>>>
>>>>> DB> I like these suggestions (START-OF-INPUT?).
>>>>>
>>>     However, I don't
>>>
>>>>> see how the problem of the optimised case is not
>>>>>
>>>     solved by them. I
>>>
>>>>> think the optimised case is fine if we have the
>>>>>
>>>     following rules:
>>>
>>>>>
>>>>>
>>>> Yes, I hadn't thought through the optimized case as
>>>>
>>>     thoroughly as you.
>>>
>>>> Your suggested method name is fine by me as well
>>>>
>>>>
>>>>> 1. START-OF-SPEECH (and optimised bargin) is only generated
>>>>>
>>>> for the
>>>>
>>>>> input type that is being listened for 2. A speechrecog
>>>>>
>>>     listens for
>>>
>>>>> DTMF if DTMF grammars are active, speech if speech
>>>>>
>>>     grammars are
>>>
>>>>> active, or speech and DTMF if both grammar types are active.
>>>>>
>>>>>
>>>> Works for me.
>>>>
>>>>
>>>>>
>>>>>
>>>>>> Note that all of the above I'm saying with my technical
>>>>>>
>>>> hat on and
>>>>
>>>>>> my chair hat off.
>>>>>> Putting my chair hat on for a moment, we really need
>>>>>>
>>>     to get this
>>>
>>>>>> spec  to last call, so at some point Eric or I is going to
>>>>>>
>>>> declare
>>>>
>>>>>> rough  consensus so we can move on.
>>>>>>
>>>>>>
>>>>>
>>>>> DB> Agreed!
>>>>>
>>>>>
>>>>>
>>>>>>
>>>>>> Dave Oran.
>>>>>>
>>>>>>
>>>>>>> Dave
>>>>>>>
>>>>>>>
>>>>>>>
>>>>>>>
>>>>>>>
>>>>>>>
>>>>>>>
>>>>>>>
>>>>>>>
>>>>>>>
>>>>>>>
>>>>>>>
>>>>>>> Sarvi makes a good point that adding the reason why the
>>>>>>>
>>>> START-OF-
>>>>
>>>>>>> SPEECH occurred does not fix the optimised bargin case.
>>>>>>>
>>>>>>> dtmfrecog - listens for DTMF only
>>>>>>> speechrecog - listens for DTMF only, or speech only,
>>>>>>>
>>>     or speech &
>>>
>>>>>>> DTMF
>>>>>>>
>>>>>>>
>>>>>>>
>>>>>>>
>>>>>>>
>>>>>>>
>>>>>>>
>>>>>>>
>>>>>>> ----- Original Message ----- From: "Shanmugham, Saravanan"
>>>>>>> <sarvi@cisco.com>
>>>>>>> To: "David R Oran" <oran@cisco.com>; "Dave Burke"
>>>>>>> <david.burke@voxpilot.com>
>>>>>>> Cc: <speechsc@ietf.org>; "Klaus Reifenrath"
>>>>>>> <Klaus.Reifenrath@Scansoft.com>
>>>>>>> Sent: Tuesday, July 05, 2005 8:51 PM
>>>>>>> Subject: RE: [Speechsc] START-OF-SPEECH in DTMF-only mode
>>>>>>>
>>>>>>>
>>>>>>> inline.
>>>>>>>
>>>>>>>     -----Original Message-----
>>>>>>>     From: David R Oran [mailto:oran@cisco.com]
>>>>>>>     Sent: Tuesday, July 05, 2005 11:08 AM
>>>>>>>     To: Dave Burke
>>>>>>>     Cc: Shanmugham, Saravanan; Klaus Reifenrath;
>>>>>>>
>>>     speechsc@ietf.org
>>>
>>>>>>>     Subject: Re: [Speechsc] START-OF-SPEECH in DTMF-only
>>>>>>>
>>> mode
>>>
>>>>>>>
>>>>>>>
>>>>>>>     On Jul 5, 2005, at 12:54 PM, Dave Burke wrote:
>>>>>>>
>>>>>>>
>>>>>>>> Inline.
>>>>>>>>
>>>>>>>> Dave
>>>>>>>>
>>>>>>>> ----- Original Message ----- From:
>>>>>>>>
>>>     "Shanmugham, Saravanan"
>>>
>>>>>>>> <sarvi@cisco.com>
>>>>>>>> To: "David R Oran" <oran@cisco.com>; "Klaus
>>>>>>>>
>>> Reifenrath"
>>>
>>>>>>>> <Klaus.Reifenrath@Scansoft.com>
>>>>>>>> Cc: <speechsc@ietf.org>
>>>>>>>> Sent: Tuesday, July 05, 2005 5:22 PM
>>>>>>>> Subject: RE: [Speechsc] START-OF-SPEECH in
>>>>>>>>
>>>     DTMF-only mode
>>>
>>>>>>>>
>>>>>>>>
>>>>>>>> I agree with Dave's analysis. The purpose of this
>>>>>>>>
>>> event
>>>
>>>>>>>     was barge-in.
>>>>>>>
>>>>>>>> And barge-in should happen for both DTMF and speech.
>>>>>>>>
>>>>>>>> Is there a case where you think it should not behave
>>>>>>>>
>>>>>>>     this way. If soe,
>>>>>>>
>>>>>>>> please provide a scenario where you think
>>>>>>>>    1. Barge-in should happen for DTMF and not voice or
>>>>>>>>
>>>>>>>     vice-versa.
>>>>>>>
>>>>>>>>
>>>>>>>> DB> You want to do a DTMF recognition only
>>>>>>>>
>>>     because it is
>>>
>>>>>>> noisy.
>>>>>>>
>>>>>>>> While waiting for DTMF input, the speechrecog resource
>>>>>>>>
>>>>>>>     (or advanced
>>>>>>>
>>>>>>>> dtmfrecog) generates a START-OF-SPEECH because it
>>>>>>>>
>>> heard
>>>
>>>>>>>     some speech.
>>>>>>>
>>>>>>>> The client does not want to stop prompt playing unless
>>>>>>>>
>>>>>>>     DTMF was heard
>>>>>>>
>>>>>>>> but it can't tell by the START-OF-SPEECH whether
>>>>>>>>
>>> speech
>>>
>>>>>>>     or DTMF was
>>>>>>>
>>>>>>>> heard. Similarly vice versa.
>>>>>>>>
>>>>>>>>
>>>>>>>     It's an interesting design question what part of the
>>>>>>>     policy resides at the client and what at the server, and
>>>>>>>     who makes the "final decision" about whether what was
>>>>>>>     heard was relevant to the control channel. Right now we
>>>>>>>     (IMO) have a weird partitioning in many cases where the
>>>>>>>     client basically says "do what I mean", but there are no
>>>>>>>     constraints of what the server actually does, and no
>>>>>>>     normalized basis for the client to figure out what to
>>>>>>>
>>> set
>>>
>>>>>>>     various magic numbers to (e.g. sensitivity).
>>>>>>>
>>>>>>>     In this case the only thing the client needs to
>>>>>>>
>>> decide is
>>>
>>>>>>>     whether to kill the prompt because the server thinks
>>>>>>>     something that would interfere with the feedback
>>>>>>>     ear/mouth/finger control happened. What this
>>>>>>>
>>>     says to me is
>>>
>>>>>>>     that it isn't necessarily a good idea for the client to
>>>>>>>     have more knobs to control the server
>>>>>>>
>>>     (especially if those
>>>
>>>>>>>     knows are just more value/policy input ungrounded in any
>>>>>>>     physics/ acoustics). On the other hand, having the
>>>>>>>
>>> server
>>>
>>>>>>>     tell the client more about what it thinks is going on is
>>>>>>>     probably valuable.
>>>>>>>
>>>>>>>     So, Coming to the point after this long rambling
>>>>>>>     introduction, I think it would in fact be useful for the
>>>>>>>     START-Of-SPEECH event to indicate some extra
>>>>>>>
>>> information,
>>>
>>>>>>>     for example:
>>>>>>>     a) I got something enough above the noise floor
>>>>>>>
>>>     to qualify
>>>
>>>>>>>     for exceeding the "Sensisitvity" parameter you sent
>>>>>>>
>>> in on
>>>
>>>>>>>     the request but I really can't tell what it is
>>>>>>>
>>>     (could be a
>>>
>>>>>>>     hippopatmus fart, or a siren in the background,
>>>>>>>
>>>     or captain
>>>
>>>>>>>     crunch trying to whistle DTMF).
>>>>>>>     b) I think I'm hearing speech
>>>>>>>     c) I think I'm hearing DTMF
>>>>>>>
>>>>>>> Though I agree with your former part of your
>>>>>>>
>>>     response. I am not
>>>
>>>>>>> sure I agree with your proposed solution.
>>>>>>> The way I see this problem is that, it is more of what
>>>>>>>
>>>> constitues a
>>>>
>>>>>>> barge-in event. This boils down to whether it is
>>>>>>>
>>>     speech, DTMF or
>>>
>>>>>>> both.
>>>>>>> This is inturn boils down to what type of recognizer
>>>>>>>
>>>> resource we are
>>>>
>>>>>>> using, dtmf-recog, speech-recog, and speech-only-recog(we
>>>>>>>
>>>> don't have
>>>>
>>>>>>> this and I don't think we should add it, but think
>>>>>>>
>>>     of this as a
>>>
>>>>>>> place holder that explains the concept).
>>>>>>>
>>>>>>> A client knowing what type of barge-in happenned, does
>>>>>>>
>>>> not impact
>>>>
>>>>>>> the
>>>>>>> barge-in operation itself as it may be too late(for
>>>>>>>
>>>     the optimized
>>>
>>>>>>> barge-in case). It may have other use cases, and if we
>>>>>>>
>>>> can identify
>>>>
>>>>>>> them, I don't mind adding support for the
>>>>>>>
>>>     START-OF-SPEECH event to
>>>
>>>>>>> say what type of barge-in happenned. But that itself does
>>>>>>>
>>> not
>>>
>>>> solve the
>>>>
>>>>>>> original problem raised. Refer to my previous response.
>>>>>>>
>>>>>>> The solution lies in defining what what is a
>>>>>>>
>>>     barge-in event.
>>>
>>>>>>> That  boils
>>>>>>> down to what type of recognition is happenning,
>>>>>>>
>>>> dtmf-only, speech-
>>>>
>>>>>>> dtmf
>>>>>>> or speech-only. We do not support speech-only as a
>>>>>>>
>>>     resource today,
>>>
>>>>>>> the question is do we need a header to force it.
>>>>>>>
>>>>>>> Sarvi
>>>>>>>
>>>>>>>
>>>>>>>
>>>>>>>>    2. You would benefit from the client knowing
>>>>>>>>
>>>> what caused
>>>>
>>>>>>> the
>>>>>>>
>>>>>>>> barge-in, DTMF Vs speech.
>>>>>>>>
>>>>>>>> DB> See previous comment. And previous e-mail:
>>>>>>>>
>>>     either add an
>>>
>>>>>>>> inputmodes header (taking value speech, dtmf, both) to
>>>>>>>>
>>>>>>>     the RECOGNIZE
>>>>>>>
>>>>>>>> request or add a header to the START-OF-SPEECH event
>>>>>>>>
>>>>>>>     indicating DTMF
>>>>>>>
>>>>>>>> or speech.
>>>>>>>>
>>>>>>>>
>>>>>>>     I'm leaning in your direction on this latter point - as
>>>>>>>     should be evident from what I wrote above.
>>>>>>>
>>>>>>>
>>>>>>>> Sarvi
>>>>>>>>
>>>>>>>>     -----Original Message-----
>>>>>>>>     From: speechsc-bounces@ietf.org
>>>>>>>>     [mailto:speechsc-bounces@ietf.org] On Behalf Of
>>>>>>>>
>>>> David R
>>>>
>>>>>>> Oran
>>>>>>>
>>>>>>>>     Sent: Tuesday, July 05, 2005 5:24 AM
>>>>>>>>     To: Klaus Reifenrath
>>>>>>>>     Cc: 'speechsc@ietf.org'
>>>>>>>>     Subject: Re: [Speechsc] START-OF-SPEECH in
>>>>>>>>
>>>> DTMF-only mode
>>>>
>>>>>>>>
>>>>>>>>
>>>>>>>>     On Jul 5, 2005, at 3:46 AM, Reifenrath,
>>>>>>>>
>>>     Klaus wrote:
>>>
>>>>>>>>
>>>>>>>>
>>>>>>>>> The current spec is not clear when
>>>>>>>>>
>>>> START-OF-SPEECH need
>>>>
>>>>>>>>     to be send in
>>>>>>>>
>>>>>>>>> the following scenarios:
>>>>>>>>> A) The client requested a DTMF Recognizer. Is
>>>>>>>>>
>>> the
>>>
>>>>>>>>     START-OF-SPEECH
>>>>>>>>
>>>>>>>>> event send to the client also if speech
>>>>>>>>>
>>>     was detected?
>>>
>>>>>>>>     I suspect so, since one of the prime purposes
>>>>>>>>
>>>> is to enable
>>>>
>>>>>>>>     client- mediated barge-in handling. However, if
>>>>>>>>
>>> the
>>>
>>>>>>>>     recognizer is in fact only capable of
>>>>>>>>
>>>     recognizing DTMF
>>>
>>>>>>>>     then it may in fact not report anythin
>>>>>>>>
>>>     unless it's using
>>>
>>>>>>>>     some primitive thresholding machinery, like a SN
>>>>>>>>
>>>>>>> threshold.
>>>>>>>
>>>>>>>>> B) The client requested a Speech
>>>>>>>>>
>>>     Recognizer, but only
>>>
>>>>>>>>     activated DTMF
>>>>>>>>
>>>>>>>>> grammars. Is the START-OF-SPEECH event
>>>>>>>>>
>>>     send to the
>>>
>>>>>>>>     client also if
>>>>>>>>
>>>>>>>>> speech was detected?
>>>>>>>>>
>>>>>>>>     Again, I'd say yes, for the same reason as above.
>>>>>>>>
>>>>>>>>> I think in both cases START-OF-SPEECH should
>>>>>>>>>
>>> only
>>>
>>>>>>>     be send after
>>>>>>>
>>>>>>>>> detecting a DTMF digit (see Figure 12 of
>>>>>>>>>
>>> VoiceXML
>>>
>>>>>>>     2.0: http://
>>>>>>>
>>>>>>>>> www.w3.org/TR/voicexml20/#dmlATiming).
>>>>>>>>>
>>>>>>>>     We seem to have reached different
>>>>>>>>
>>>     conclusions. I'd be
>>>
>>>>>>>>     interested in why you think my analysis
>>>>>>>>
>>>     above is wrong.
>>>
>>>>>>>>
>>>>>>>>     Dave.
>>>>>>>>
>>>>>>>>
>>>>>>>>> Klaus
>>>>>>>>>
>>>>>>>>> _______________________________________________
>>>>>>>>> Speechsc mailing list
>>>>>>>>> Speechsc@ietf.org
>>>>>>>>> https://www1.ietf.org/mailman/listinfo/speechsc
>>>>>>>>>
>>>>>>>>>
>>>>>>>>
>>>>>>>>     _______________________________________________
>>>>>>>>     Speechsc mailing list
>>>>>>>>     Speechsc@ietf.org
>>>>>>>>     https://www1.ietf.org/mailman/listinfo/speechsc
>>>>>>>>
>>>>>>>>
>>>>>>>> _______________________________________________
>>>>>>>> Speechsc mailing list
>>>>>>>> Speechsc@ietf.org
>>>>>>>> https://www1.ietf.org/mailman/listinfo/speechsc
>>>>>>>>
>>>>>>>>
>>>>>>>
>>>>>>>
>>>>>>> _______________________________________________
>>>>>>> Speechsc mailing list
>>>>>>> Speechsc@ietf.org
>>>>>>> https://www1.ietf.org/mailman/listinfo/speechsc
>>>>>>>
>>>>>>>
>>>>>>>
>>>>>>
>>>>>> _______________________________________________
>>>>>> Speechsc mailing list
>>>>>> Speechsc@ietf.org
>>>>>> https://www1.ietf.org/mailman/listinfo/speechsc
>>>>>>
>>>>>
>>>>>
>>>>
>>>> _______________________________________________
>>>> Speechsc mailing list
>>>> Speechsc@ietf.org
>>>> https://www1.ietf.org/mailman/listinfo/speechsc
>>>>
>>>>
>>>
>>>
>>>     _______________________________________________
>>>     Speechsc mailing list
>>>     Speechsc@ietf.org
>>>     https://www1.ietf.org/mailman/listinfo/speechsc
>>>
>>>
>>> _______________________________________________
>>> Speechsc mailing list
>>> Speechsc@ietf.org
>>> https://www1.ietf.org/mailman/listinfo/speechsc
>>>
>>>
>>> _______________________________________________
>>> Speechsc mailing list
>>> Speechsc@ietf.org
>>> https://www1.ietf.org/mailman/listinfo/speechsc
>>>
>>>
>>
>> _______________________________________________
>> Speechsc mailing list
>> Speechsc@ietf.org
>> https://www1.ietf.org/mailman/listinfo/speechsc
>>
>>
>>
>>
>
> _______________________________________________
> Speechsc mailing list
> Speechsc@ietf.org
> https://www1.ietf.org/mailman/listinfo/speechsc
> 


_______________________________________________
Speechsc mailing list
Speechsc@ietf.org
https://www1.ietf.org/mailman/listinfo/speechsc



From speechsc-bounces@ietf.org Wed Jul 13 09:20:45 2005
Received: from localhost.localdomain ([127.0.0.1] helo=megatron.ietf.org)
	by megatron.ietf.org with esmtp (Exim 4.32)
	id 1DshAL-0001Sa-2S; Wed, 13 Jul 2005 09:20:45 -0400
Received: from odin.ietf.org ([132.151.1.176] helo=ietf.org)
	by megatron.ietf.org with esmtp (Exim 4.32) id 1DshAI-0001R2-Kz
	for speechsc@megatron.ietf.org; Wed, 13 Jul 2005 09:20:43 -0400
Received: from ietf-mx.ietf.org (ietf-mx [132.151.6.1])
	by ietf.org (8.9.1a/8.9.1a) with ESMTP id JAA05492
	for <speechsc@ietf.org>; Wed, 13 Jul 2005 09:20:41 -0400 (EDT)
Received: from sj-iport-2-in.cisco.com ([171.71.176.71]
	helo=sj-iport-2.cisco.com) by ietf-mx.ietf.org with esmtp (Exim 4.43)
	id 1Dshce-0008Gx-PE
	for speechsc@ietf.org; Wed, 13 Jul 2005 09:50:04 -0400
Received: from sj-core-1.cisco.com (171.71.177.237)
	by sj-iport-2.cisco.com with ESMTP; 13 Jul 2005 06:20:29 -0700
Received: from imail.cisco.com (imail.cisco.com [128.107.200.91])
	by sj-core-1.cisco.com (8.12.10/8.12.6) with ESMTP id j6DDKSvM021404;
	Wed, 13 Jul 2005 06:20:28 -0700 (PDT)
Received: from [10.32.245.152] (stealth-10-32-245-152.cisco.com
	[10.32.245.152])
	by imail.cisco.com (8.12.11/8.12.10) with SMTP id j6DDJ4op030379;
	Wed, 13 Jul 2005 06:19:05 -0700
In-Reply-To: <028f01c5865e$7f92c4a0$038ae9d5@db01.voxpilot.com>
References: <7DE7C4EF3B7C8B4B82955191378290D802ED4176@mtb1exch01.nuance.com>
	<C7673C4C-0624-47A1-B1DF-115FB284785A@cisco.com>
	<028f01c5865e$7f92c4a0$038ae9d5@db01.voxpilot.com>
Mime-Version: 1.0 (Apple Message framework v733)
X-Priority: 3
Content-Type: text/plain; charset=US-ASCII; delsp=yes; format=flowed
Message-Id: <9B4BE6A2-82ED-4C8B-BAE5-006FB0C2DBBB@cisco.com>
Content-Transfer-Encoding: 7bit
From: David R Oran <oran@cisco.com>
Subject: Re: [Speechsc] START-OF-SPEECH in DTMF-only mode -> summary(?)
Date: Wed, 13 Jul 2005 09:20:22 -0400
To: Dave Burke <david.burke@voxpilot.com>
X-Mailer: Apple Mail (2.733)
IIM-SIG: v:"1.1"; h:"imail.cisco.com"; d:"cisco.com"; z:"home"; m:"krs";
	t:"1121260748.645718"; x:"432200"; a:"rsa-sha1"; b:"nofws:35539";
	e:"Iw=="; n:"sQYarK2E51MdcTiUqeif3F7cWdxIfoCiXhdfb9vD5ee/j0jXL15gbFxF2p"
	"XIweAblu0N6XAgK7k+wrbr7bQDJaCDqOmzqpRUBjIRQAXQ7NzadpmR3pUL6wxaRUtW+c43sl9jC"
	"50Qg1sXHpPjt8Y+Y16ioyQAQAdSunM4YhevURc=";
	s:"H1pKFgSW9OHBcRmdPyivVypXUwqQXJ4P88NKi7sEQhv6uDzxMbxe4ExezsIpKAduH3MVq+Ld"
	"NVKoJ1e1aG/opAL3CC0jKbx+2Z6+/HTXHxhyCfUdSX4oNKR3SCLivb5cdMLW5S/R1bjZY+Ns4po"
	"TKiHTQU6Z3eyvYw21iRQq+s8="; c:"From: David R Oran <oran@cisco.com>";
	c:"Subject: Re: [Speechsc] START-OF-SPEECH in DTMF-only mode ->
	summary" "(?)"; c:"Date: Wed, 13 Jul 2005 09:20:22 -0400"
IIM-VERIFY: s:"y"; v:"y"; r:"60"; h:"imail.cisco.com";
	c:"message from imail.cisco.com verified; "
X-Spam-Score: 0.0 (/)
X-Scan-Signature: aceaed67ecf223b6febbd7c5b72f7957
Content-Transfer-Encoding: 7bit
Cc: speechsc@ietf.org, Pierre Forgues <forgues@nuance.com>,
	Eric Burger <eburger@brooktrout.com>, "Shanmugham,
	Saravanan" <sarvi@cisco.com>
X-BeenThere: speechsc@ietf.org
X-Mailman-Version: 2.1.5
Precedence: list
List-Id: Speech Services Control Working Group <speechsc.ietf.org>
List-Unsubscribe: <https://www1.ietf.org/mailman/listinfo/speechsc>,
	<mailto:speechsc-request@ietf.org?subject=unsubscribe>
List-Post: <mailto:speechsc@ietf.org>
List-Help: <mailto:speechsc-request@ietf.org?subject=help>
List-Subscribe: <https://www1.ietf.org/mailman/listinfo/speechsc>,
	<mailto:speechsc-request@ietf.org?subject=subscribe>
Sender: speechsc-bounces@ietf.org
Errors-To: speechsc-bounces@ietf.org


On Jul 11, 2005, at 5:21 PM, Dave Burke wrote:

> I don't think it is possible in practice to separate what input  
> type is considered potential barge-in and what input type is being  
> recognised because a recognition hypothesis is always generated  
> when a barge-in occurs.
True, but what I'm trying to tease apart is whether to declare barge- 
in from the point of view of stopping output from simply detecting  
input for the purposes of cranking up the non-signal-processing parts  
of the recognition engine. Did you not find my "mark the phrases"  
application example compelling?

> For example, imagine a client invoked a speechrecog and asked it to  
> only consider DTMF for barge-in. Once speech is detected, although  
> no barge-in would happen, a RECOGNITION-COMPLETE will result after  
> Speech-Incomplete-Timeout milliseconds of silence thus ending the  
> recognition state machine with a 'nomatch'.
Sure, but of course the timeout could be very long... I must be  
dense, because I see this relationship as tenuous rather than  
strongly coupled.

> The converse example applies for a speech-only recognition (DTMF- 
> Interdigit-Timeout replaces the Speech-Incomplete-Timeout).
>
> Although I prefer the idea of an InputModes header in RECOGNIZE  
> (it's the cleanest solution for VoiceXML implementors), I am  
> willing to compromise on using the grammar type to determine the  
> input types (in what follows I use input types to mean both what is  
> considered potential barge-in and what type of input is to be  
> recognised). The main issue with this approach is incompatibility  
> with VoiceXML (recall VoiceXML grammar activation is independent of  
> what inputmodes are set).
>
And voice XML does not have any explicit control over barge-in  
either. We could of course back off not have clients involved at all  
in barge-in processing, but I think that would be a mistake.

> This incompatibility results in side-effects that, in my opinion,  
> are inconsequential for real applications:
>    a. speech barge-in (noise) and nomatch will not happen for  
> inputmodes="both" when no speech grammars are activated
>    b. dtmf barge-in and nomatch will not happen for  
> inputmodes="both" when no dtmf grammars are activated
>    c. grammars with errors will not be detected if the grammar type  
> is not compatible with the inputmodes property
> A workaround for the really conscientious VoiceXML platform is to  
> activate a dummy grammar (maybe with the NULL rule) of type speech  
> (dtmf) when inputmodes="both" is set but no speech (dtmf) grammar  
> is activated in the application.
>
I agree with this assessment. I'll also point out that the  
flexibility will allow a class of VXML applications that can't be  
done today (e.g. ones that don't stop prompts on hearing input.

> -----------------
>
> Summarising the proposed changes again  (refined a little and to  
> avoid trawling though this massive thread!):
>
> 1. For a speechrecog resource, the types of the activated grammars  
> determine the input types the recogniser processes and considers  
> for potential barge-in.
>
Good. Agree.

> 2. Clarify a dtmfrecog only processes DTMF and hence can only  
> generate barge-in / START-OF-SPEECH for DTMF inputs
>
I'm ok with this even though I'd rather keep input detection and  
recognition behavior orthogonal to accommodate smarter server  
implementations and to give better consistency between the server- 
optimized and client-intermediated cases. In particular, I do really  
feel like we need a way for the client to tell a server to NOT try to  
outguess it by doing optimized barge in processing since the  
application might in fact not want the output stopped.


> Nice-to-have:
>
> 3. Change START-OF-SPEECH to START-OF-INPUT
>
I'm in favor of this.

> 4. Add a header of InputType to START-OF-INPUT. Current specified  
> values are "dtmf" or "speech".
>
I think we really need this, and would press again that we make this  
a list of values and future-proof it with the other values we suspect  
will be quite useful (e.g. music, gesture).

Dave. (chair hat sort-of-on since we're trying to get final consensus  
here).


> -----------------
>
> Dave
>
> ----- Original Message ----- From: "David R Oran" <oran@cisco.com>
> To: "Pierre Forgues" <forgues@nuance.com>
> Cc: <speechsc@ietf.org>; "Shanmugham, Saravanan" <sarvi@cisco.com>;  
> "Eric Burger" <eburger@brooktrout.com>; "Dave Burke"  
> <david.burke@voxpilot.com>
> Sent: Monday, July 11, 2005 2:15 PM
> Subject: Re: [Speechsc] START-OF-SPEECH in DTMF-only mode -> summary 
> (?)
>
>
>
>>
>> On Jul 8, 2005, at 11:36 AM, Pierre Forgues wrote:
>>
>>
>>>
>>>
>>> -----Original Message-----
>>> From: speechsc-bounces@ietf.org [mailto:speechsc- 
>>> bounces@ietf.org] On
>>> Behalf Of David R Oran
>>> Sent: Friday, July 08, 2005 10:05 AM
>>> To: Dave Burke
>>> Cc: speechsc@ietf.org; Shanmugham, Saravanan; Eric Burger
>>> Subject: Re: [Speechsc] START-OF-SPEECH in DTMF-only mode ->  
>>> summary (?)
>>>
>>>
>>> On Jul 8, 2005, at 6:59 AM, Dave Burke wrote:
>>>
>>>
>>>
>>>> I agree with your opinion on what constitutes a barge-in. Based on
>>>> Klaus' e-mail, I am more concerned that using the grammar type to
>>>> decide what constitutes a barge-in is going to make VoiceXML
>>>> implementations difficult.
>>>>
>>>> It is easy to map VoiceXML application selected inputmodes to MRCP
>>>> resources:
>>>> a. inputmodes="dtmf" -> Use a dtmfrecog
>>>> b. inputmodes="speech" -> Use a "speech-only-recog"
>>>> c. inputmodes="both" -> Use a speechrecog (or a combination of
>>>> "speech-only-recog" + dtmfrecog)
>>>>
>>>> The only problem is (b). While not as important as being able to do
>>>> DTMF-only recognition, I believe we DO need to support speech-only
>>>> recongition so as (a) to avoid unnecessary limitations in VUI
>>>> design, and (b) to facilitate VoiceXML implementations. I think we
>>>> should add a header to RECOGNIZE so the client is able to always
>>>> _explicitly_ set the inputmodes.
>>>>
>>>>
>>>>
>>> I'm not sure I buy this, since what the recognizer is looking for  
>>> and
>>> what constitutes an input that should be considered a potential  
>>> barge-
>>> in strike me as independent. I'm similarly not persuaded that we  
>>> need
>>> the flexibility to set the input mode independently of the grammar,
>>> since it leads to all sorts of inconsistent states (e.g. speech-only
>>> grammar with an input-mode of dtmf). On the other hand I can see the
>>> need for the client to specify what sorts of input out to generate
>>> the start-of-input (nee start-of-speech) event on.
>>>
>>> Pmf> The mechanism for a client to specify the type of input is  
>>> using
>>> grammars.  These can have DTMF and/or speech requirements.  If  
>>> you add
>>> the complexity of an independent header to specify the input mode  
>>> then
>>> you will need to document the behavior for inconsistencies.
>>>
>>>
>> Right. I generally agree with Pierre, which is why I proposed a  
>> compromise position where the client can use a header to specify  
>> what types of input should be considered the start of something   
>> potentially interesting from a barge-in point of view, as opposed  
>> to  a directive to the server for what to recognize.
>>
>> I don't think this is crucial functionality, but it does make for   
>> more consistent behavior between the optimized and non-optimized   
>> barge-in cases. In fact, I can see cases where the client doesn't   
>> want any barge-in processing to happen. Imagine an application  
>> called  "mark the musical phrases", where the client listens to  
>> playout of a  musical score, and hits various DTMF buttons to  
>> indicate markers  (e.g. end-of-phrase, end-of-theme, key-change)  
>> while listening. For  this application you want to make sure that  
>> barge-in doesn't stop the  music!
>>
>>
>>>
>>>
>>>> In retrospect, I don't like the idea of multiple START-OF-SPEECH
>>>> events being generated from the same media resource. This is
>>>> because one assumes that the START-OF-SPEECH should be of the same
>>>> type as the hypothesis returned in the RECOGNITION-COMPETE message
>>>> - most implementations, on hearing one input mode type disable the
>>>> recogniser of the other type.
>>>>
>>>>
>>> I'm not sure I follow this logic, but I'm not wedded to the idea of
>>> allowing multiple events. It was trying to solve the problem of
>>> ambiguity around what the recognizer was hearing and the idea that
>>> the client might care. If you don't think the client will ever care,
>>> then we don't need the capability.
>>> Pmf> I agree we should not have multiple SOS events.  Too much   
>>> chatter.
>>>
>>>
>> I'm not concerned about the "chatter" since this is a low- 
>> bandwidth control channel even with multiple events, but as I said  
>> it's a small point and I'm happy to concede.
>>
>>
>>>
>>>
>>>
>>>> I do like David's idea of renaming START-OF-SPEECH to something
>>>> like START-OF-INPUT and carrying a type header because it is neater
>>>> and more extensible.
>>>>
>>>> So in summary, I propose we modify the spec to:
>>>>
>>>> 1. Clarify what constitutes a barge-in for a dtmfrecog and
>>>> speechrecog as per Sarvi's e-mail (and in agreement with Klaus' for
>>>> dtmfrecog).
>>>>
>>>>
>>>>
>>> Hmmm, ok, but I think the issue is actually clarifying what the
>>> recognizer declares as "interesting input", which it may decide also
>>> constitutes a barge-in in the optimized case, and in either case
>>> reports to the client that it heard.
>>> Pmf> I'm not going to argue changing the name of the event, but my
>>> preference would be to keep existing method names unless there is a
>>> clear reason for changing - which I do not see here.
>>>
>>>
>> I do see a clear reason, since the thing that you see the start  
>> of  may not be speech. I like (I think it was Dave's suggestion)  
>> "start- of-input:" since it mirrors the other method names.
>>
>>>
>>>
>>>> 2. Specifiy an InputModes header to RECOGNIZE (defaults to "both",
>>>> can also be "speech" or "DTMF"). Setting to speech for a
>>>> speechrecog results in the hypothesised "speech-only-recog".
>>>> Setting to DTMF for a speechrecog is equivalent to using a
>>>> dtmfrecog. Edge cases: Setting to speech for a dtmfrecog will
>>>> result in a noinput as will setting to DTMF for a speechrecog which
>>>> does not support DTMF.
>>>>
>>>>
>>>>
>>> I'm ok with having a header for the client to tell the server  
>>> when it
>>> would like the start-of-input event to be generated and what the
>>> client considers to be the "interesting input" that the recognizer
>>> should use to do the discrimination, and possibly do barge-in
>>> processing in the optimized case.
>>>
>>> I'm less ok with the proposed domain of values, since it will have
>>> extensibility problems. Especially problematical is have a code  
>>> point
>>> of "both" since that will be ambiguous if we even define a "music"
>>> recognizer, or a "gesture" recognizer using video input. Here's my
>>> counter-proposal:
>>>
>>> Create a header on the recognizer resource requests called "start-
>>> input-on:". Define a registry of values, with the following three
>>> values initially defined:
>>>      - dtmf
>>>      - speech
>>>      - music
>>>
>>> The semantics would be that the resource is to generate a "start-of-
>>> input" event, apply the defined grammars to what follows, and do any
>>> local optimized barge-in processing if ANY of the enumerated input
>>> types is detected. That way you can say
>>>      "start-input-on: dtmf" if you want to just  dtmf,
>>>      "start-input-on: speech" if you want just speech (dtmf input
>>> would be ignored even if the grammar supported dtmf)
>>>      "start-input- on: speech, dtmf" if you wanted both
>>> etc.
>>>
>>> Pmf> Doesn't this come back to the previous point of having two
>>> independent ways to specify input modes?  Through grammars and this
>>> proposed new header?  If you add this then you will have to document
>>> behavior on inconsistencies.  Anyway, if you proceed then I would
>>> definitely make this optional
>>>
>>>
>> I hope I explained the intent above. This header would not define  
>> the input types that the recognize would process, but rather the  
>> input  types the client wants it to consider potential barge-in  
>> and hence  generate the "start-of-input:" event for and do any  
>> local barge-in  optimizations for.
>>
>>
>> Dave.
>>
>>
>>>
>>>
>>>> 3. Change START-OF-SPEECH to START-OF-INPUT. Clarify it occurs once
>>>> and coincides with barge-in.
>>>>
>>>>
>>>>
>>> Ok for the "only once", but I'd like to tighten it up to talk about
>>> more than just barge-in, as I suggested above.
>>>
>>>
>>>
>>>> 4. Add a header of InputType to START-OF-INPUT. Current specified
>>>> values are "dtmf" or "speech".
>>>>
>>>>
>>>>
>>> Ok, with slight modification. Say that the syntax of the "input- 
>>> type"
>>> header is a single-valued subset of the registered value(s) that are
>>> defined for the start-input-on: header.
>>>
>>> Comments?
>>>
>>> Dave O. (technical hat on, chair hat off).
>>>
>>>
>>>
>>>> Dave
>>>>
>>>>
>>>>
>>>> ----- Original Message ----- From: "Shanmugham, Saravanan"
>>>> <sarvi@cisco.com>
>>>> To: "Eric Burger" <eburger@brooktrout.com>; <speechsc@ietf.org>
>>>> Sent: Thursday, July 07, 2005 5:52 PM
>>>> Subject: RE: [Speechsc] START-OF-SPEECH in DTMF-only mode ->  
>>>> summary
>>>> (?)
>>>>
>>>>
>>>> I don't think generating multiple START-OF-SPEECH events is a
>>>> solution.
>>>> We still haven't addressed, what constitutes a barge-in, for the
>>>> optimised case. That should also be the single point when a single
>>>> START-OF-SPEECH(or whathever else you want to name it) should be
>>>> generated.
>>>>
>>>> That point, in my opinion should be
>>>>   1. For "dtmf-recog" resources should be the beginning of a  
>>>> DTMF key
>>>> press.
>>>>   2. For "speech-recog" resources should be the beginning of a DTMF
>>>> key
>>>> press or the beginning of speech. This is should be irrespective of
>>>> what
>>>> type of grammar is being used. Coz even numbers only grammars can
>>>> still
>>>> be spoken and hence cannot be assumed to be a cue for DTMF only
>>>> recognition.
>>>>   3. For "speech-only-recog" resources(which are not defined today)
>>>> the
>>>> time to barge-in is the beginning of speech. I don't see a need for
>>>> such
>>>> a resource today. But I am mentioning this for completeness.
>>>>
>>>> Sarvi
>>>>
>>>>
>>>>     -----Original Message-----
>>>>     From: speechsc-bounces@ietf.org
>>>>     [mailto:speechsc-bounces@ietf.org] On Behalf Of Eric Burger
>>>>     Sent: Thursday, July 07, 2005 5:32 AM
>>>>     To: speechsc@ietf.org
>>>>     Subject: RE: [Speechsc] START-OF-SPEECH in DTMF-only mode
>>>>     -> summary(?)
>>>>
>>>>     Can we declare consensus?
>>>>
>>>>
>>>>
>>>>> -----Original Message-----
>>>>> From: speechsc-bounces@ietf.org
>>>>>
>>>>>
>>>>     [mailto:speechsc-bounces@ietf.org]
>>>>
>>>>
>>>>> Sent: Wednesday, July 06, 2005 10:45 AM
>>>>> To: Dave Burke
>>>>> Cc: speechsc@ietf.org; Shanmugham, Saravanan; Klaus Reifenrath
>>>>> Subject: Re: [Speechsc] START-OF-SPEECH in DTMF-only mode ->
>>>>> summary(?)
>>>>>
>>>>> I think we're getting close. I though about snipping out
>>>>>
>>>>>
>>>>     some pieces
>>>>
>>>>
>>>>> to cut down the text, but I realized the context is
>>>>>
>>>>>
>>>>     still needed. See
>>>>
>>>>
>>>>> inline.
>>>>>
>>>>> On Jul 6, 2005, at 10:29 AM, Dave Burke wrote:
>>>>>
>>>>>
>>>>>
>>>>>> Inline.
>>>>>>
>>>>>> Dave
>>>>>>
>>>>>> ----- Original Message ----- From: "David R Oran"
>>>>>>
>>>>>>
>>>>     <oran@cisco.com>
>>>>
>>>>
>>>>>> To: "Dave Burke" <david.burke@voxpilot.com>
>>>>>> Cc: <speechsc@ietf.org>; "Shanmugham, Saravanan"
>>>>>>
>>>>>>
>>>>     <sarvi@cisco.com>;
>>>>
>>>>
>>>>>> "Klaus Reifenrath" <Klaus.Reifenrath@Scansoft.com>
>>>>>> Sent: Wednesday, July 06, 2005 1:07 PM
>>>>>> Subject: Re: [Speechsc] START-OF-SPEECH in DTMF-only
>>>>>>
>>>>>>
>>>>     mode -> summary
>>>>
>>>>
>>>>>> (?)
>>>>>>
>>>>>>
>>>>>>
>>>>>>
>>>>>>
>>>>>>>
>>>>>>> On Jul 6, 2005, at 5:31 AM, Dave Burke wrote:
>>>>>>>
>>>>>>>
>>>>>>>
>>>>>>>
>>>>>>>> + Attempting to summarise:
>>>>>>>>
>>>>>>>> 1. START-OF-SPEECH is useful for the client to know when
>>>>>>>>
>>>>>>>>
>>>>> to stop
>>>>>
>>>>>
>>>>>>>> playing prompts the in non-optimised case 2.
>>>>>>>>
>>>>>>>>
>>>>     START-OF-SPEECH is
>>>>
>>>>
>>>>>>>> useful for the client to calculate the bargin  time
>>>>>>>>
>>>>>>>>
>>>>     (e.g. VoiceXML
>>>>
>>>>
>>>>>>>> 2.1 <mark>)
>>>>>>>> 3. In the optimised case, a bargin automatically
>>>>>>>>
>>>>>>>>
>>>>     stops prompt
>>>>
>>>>
>>>>>>>> playing (assuming prompts barginable) 4. Because of
>>>>>>>>
>>>>>>>>
>>>>     the previous
>>>>
>>>>
>>>>>>>> point, the question of what input type  caused
>>>>>>>>
>>>>>>>>
>>>>     bargin is different
>>>>
>>>>
>>>>>>>> and less important to
>>>>>>>>
>>>>>>>>
>>>>> what input
>>>>>
>>>>>
>>>>>>>> type(s)  the recogniser is listening for
>>>>>>>>
>>>>>>>>
>>>>>>>>
>>>>>>> I'm not sure I follow point 4. Could you elaborate?
>>>>>>>
>>>>>>>
>>>>>>>
>>>>>>
>>>>>> DB> Adding a parameter to the START-OF-SPEECH event would
>>>>>>
>>>>>>
>>>>> certainly
>>>>>
>>>>>
>>>>>> allow the client to ignore (i.e. let prompts continue
>>>>>>
>>>>>>
>>>>     playing) the
>>>>
>>>>
>>>>>> event if the event type is not of interest (e.g. the
>>>>>>
>>>>>>
>>>>     client would
>>>>
>>>>
>>>>>> ignore speech start events when it is interested only
>>>>>>
>>>>>>
>>>>     in  DTMF start
>>>>
>>>>
>>>>>> events). This _only_ works for the non-optimised case,
>>>>>>
>>>>>>
>>>>     however. For
>>>>
>>>>
>>>>>> the optimised case, assuming START-OF-SPEECH
>>>>>>
>>>>>>
>>>>> coincides
>>>>>
>>>>>
>>>>>> with the bargin signal to the speechsynth, prompts
>>>>>>
>>>>>>
>>>>     will stop playing
>>>>
>>>>
>>>>>> for inputs that the client might not be interested
>>>>>>
>>>>>>
>>>>> in (e.g.
>>>>>
>>>>>
>>>>>> a speech input will stop prompts playing even if the client
>>>>>>
>>>>>>
>>>>> is only
>>>>>
>>>>>
>>>>>> interested in DTMF).
>>>>>>
>>>>>>
>>>>>>
>>>>> OK, now I get it. There's a need for the client to both
>>>>>
>>>>>
>>>>     handle the
>>>>
>>>>
>>>>> non-optimized case itself, and influence or at least
>>>>>
>>>>>
>>>>     have a clue what
>>>>
>>>>
>>>>> the server is going to do in the optimized case.
>>>>>
>>>>>
>>>>>
>>>>>>
>>>>>>
>>>>>>
>>>>>>>
>>>>>>>
>>>>>>>
>>>>>>>
>>>>>>>> 5. Currently, MRCPv2 has no way of indicating what
>>>>>>>>
>>>>>>>>
>>>>     input type(s) a
>>>>
>>>>
>>>>>>>> recogniser is listening for
>>>>>>>>
>>>>>>>>
>>>>>>>>
>>>>>>>>
>>>>>>> Do you mean exactly this, or do you mean "for the client to
>>>>>>> indicate  to the resource what input types it should
>>>>>>>
>>>>>>>
>>>>     look for"?
>>>>
>>>>
>>>>>>>
>>>>>>>
>>>>>>>
>>>>>>
>>>>>> DB> Yes exactly - apologies for not being clear.
>>>>>>
>>>>>>
>>>>>>
>>>>>>
>>>>>>>
>>>>>>>
>>>>>>>
>>>>>>>
>>>>>>>> + Why implement 5?
>>>>>>>>
>>>>>>>> i. Noisy case: Need DTMF-only recognition (and may
>>>>>>>>
>>>>>>>>
>>>>     only have a
>>>>
>>>>
>>>>>>>> speechrecog)
>>>>>>>>
>>>>>>>>
>>>>>>>>
>>>>>>> I'm having difficulty following the logic of why
>>>>>>>
>>>>>>>
>>>>     noise would
>>>>
>>>>
>>>>>>> necessarily trigger START-OF-SPEECH if you were listening for
>>>>>>> speech  but not DTMF. I suppose you can use a more forgiving
>>>>>>> discriminator if  all you need to tell is if you're
>>>>>>>
>>>>>>>
>>>>     getting DTMF,
>>>>
>>>>
>>>>>>> but I've had a number  of real-world cases where wind
>>>>>>>
>>>>>>>
>>>>     noise was
>>>>
>>>>
>>>>>>> detected as DTMF, and  there's always the ambiguity
>>>>>>>
>>>>>>>
>>>>     when you have
>>>>
>>>>
>>>>>>> Captain Crunch on the  phone. In either case in the
>>>>>>>
>>>>>>>
>>>>     non-optimized
>>>>
>>>>
>>>>>>> case it's the client who  gets to decide whether an
>>>>>>>
>>>>>>>
>>>>     event should be
>>>>
>>>>
>>>>>>> interpreted as barge-in or  not, so it seems an
>>>>>>>
>>>>>>>
>>>>     aesthetic protocol
>>>>
>>>>
>>>>>>> design decision whether the  client tells the server
>>>>>>>
>>>>>>>
>>>>     ahead of time
>>>>
>>>>
>>>>>>> what circumstances to generate  the START-OF-SPEECH
>>>>>>>
>>>>>>>
>>>>     event for, or
>>>>
>>>>
>>>>>>> whether the event gets generated  and the client
>>>>>>>
>>>>>>>
>>>>     decides based on
>>>>
>>>>
>>>>>>> what's in the event whether it should  be
>>>>>>>
>>>>>>>
>>>>> treated
>>>>>
>>>>>
>>>>>>> as barge- in.
>>>>>>>
>>>>>>> I suppose one could make the argument that because
>>>>>>>
>>>>>>>
>>>>     the spec implies
>>>>
>>>>
>>>>>>> that the event can only be generated once per request
>>>>>>>
>>>>>>>
>>>>     that if a
>>>>
>>>>
>>>>>>> DTMF/ speech capable recognizer first hears
>>>>>>>
>>>>>>>
>>>>> enough noise
>>>>>
>>>>>
>>>>>>> to think it's  hearing speech and later hears DTMF,
>>>>>>>
>>>>>>>
>>>>     the client will
>>>>
>>>>
>>>>>>> declare barge-in  when the event comes and not when he DTMF
>>>>>>> actually gets heard.
>>>>>>>
>>>>>>> If that's deemed a problem, we can still handle that
>>>>>>>
>>>>>>>
>>>>     in the design
>>>>
>>>>
>>>>>>> where the server just reports what it's hearing by allowing
>>>>>>> multiple  events to be generated during a single request.
>>>>>>>
>>>>>>> Between the approach just outlined above, and an
>>>>>>>
>>>>>>>
>>>>     approach where the
>>>>
>>>>
>>>>>>> client provides a filter for whether to generate the
>>>>>>>
>>>>>>>
>>>>     event or not,
>>>>
>>>>
>>>>>>> I  have a mild preference (based on aesthetics rather
>>>>>>>
>>>>>>>
>>>>     than some
>>>>
>>>>
>>>>>>> hard  engineering tradeoff) for the approach where
>>>>>>>
>>>>>>>
>>>>> the server
>>>>>
>>>>>
>>>>>>> just reports  what it's hearing.
>>>>>>>
>>>>>>>
>>>>>>> Having had some useful exchanges on this topic, it
>>>>>>>
>>>>>>>
>>>>     also is becoming
>>>>
>>>>
>>>>>>> apparent to me that this event is poorly named, and
>>>>>>>
>>>>>>>
>>>>     we should
>>>>
>>>>
>>>>>>> consider renaming it to "INTERESTING-INPUT-HEARD" or
>>>>>>>
>>>>>>>
>>>>     something akin
>>>>
>>>>
>>>>>>> to that, because as others have pointed out, a
>>>>>>>
>>>>>>>
>>>>     DTMF-only recognizer
>>>>
>>>>
>>>>>>> will never detect "start of speech".
>>>>>>>
>>>>>>> Another consideration to fold into the design choice is
>>>>>>> extensibility. Bear with me through a little
>>>>>>>
>>>>>>>
>>>>     gedankenexperiment.
>>>>
>>>>
>>>>>>>
>>>>>>> Suppose we want to define a new recognizer type,
>>>>>>>
>>>>>>>
>>>>     which I'll call
>>>>
>>>>
>>>>>>> the "name that tune" recognizer. The client plays
>>>>>>>
>>>>>>>
>>>>     music to the
>>>>
>>>>
>>>>>>> server and  the server recognizes musical notes. The
>>>>>>>
>>>>>>>
>>>>     grammar is a
>>>>
>>>>
>>>>>>> standard  musical notation, augmented with a semantic
>>>>>>> interpretation that  transforms the notes into the
>>>>>>>
>>>>>>>
>>>>     title of the
>>>>
>>>>
>>>>>>> tune and provides that as  an answer.
>>>>>>>
>>>>>>> First, there's no speech involved (or is there...hang
>>>>>>>
>>>>>>>
>>>>     on a minute).
>>>>
>>>>
>>>>>>> Second, in order to accommodate the "name that tune"
>>>>>>> recognizer, we'd have to extend both the client and
>>>>>>>
>>>>>>>
>>>>     the server to
>>>>
>>>>
>>>>>>> undetstand a  directive as to whether to recognize
>>>>>>>
>>>>>>>
>>>>     music or now,
>>>>
>>>>
>>>>>>> inaddition to what  the server already knows what to
>>>>>>>
>>>>>>>
>>>>     do based on
>>>>
>>>>
>>>>>>> the grammar. If you  follow my logic above, whether
>>>>>>>
>>>>>>>
>>>>     or not we do
>>>>
>>>>
>>>>>>> that, we have to extend  "start-of-speech" to say
>>>>>>>
>>>>>>>
>>>>     "I'm hearing
>>>>
>>>>
>>>>>>> music". So far fairly  straightforward, but let me
>>>>>>>
>>>>>>>
>>>>     now throw in the
>>>>
>>>>
>>>>>>> pathological twist.
>>>>>>>
>>>>>>> Suppose what I feed to a  combined music/speech
>>>>>>>
>>>>>>>
>>>>     recognizer is a
>>>>
>>>>
>>>>>>> work  in sprechstimme (spoken music), like the
>>>>>>>
>>>>>>>
>>>>     "Geographical Fugue"
>>>>
>>>>
>>>>>>> (aside:  this is a wonderful piece of music I highly
>>>>>>>
>>>>>>>
>>>>     recommend to
>>>>
>>>>
>>>>>>> anyone  interested in small ensemble singing). In
>>>>>>>
>>>>>>>
>>>>     this case, the
>>>>
>>>>
>>>>>>> tune could  be named by either doing speech or music
>>>>>>>
>>>>>>>
>>>>     recognition.
>>>>
>>>>
>>>>>>> Why is there  any need for the client to constrain
>>>>>>>
>>>>>>>
>>>>     the server as to
>>>>
>>>>
>>>>>>> which it tries  to do when it's
>>>>>>>
>>>>>>>
>>>>> already
>>>>>
>>>>>
>>>>>>> told the server what it wants through the  grammar?
>>>>>>>
>>>>>>> A few other comments below
>>>>>>>
>>>>>>>
>>>>>>>
>>>>>>>
>>>>>>>> ii. Flexibility: Want speech-only recognition
>>>>>>>>
>>>>>>>>
>>>>     (because a second
>>>>
>>>>
>>>>>>>> recogniser is doing hotword on DTMF)
>>>>>>>>
>>>>>>>>
>>>>>>>>
>>>>>>>>
>>>>>>> I don't see how flexibility is affected by this deisgn
>>>>>>>
>>>>>>>
>>>>> choice. If
>>>>>
>>>>>
>>>>>>> that's what you want, feed the speech-only recognizer
>>>>>>>
>>>>>>>
>>>>     a grammar
>>>>
>>>>
>>>>>>> without any DTMF rules.
>>>>>>>
>>>>>>>
>>>>>>>
>>>>>>>
>>>>>>>> + How to implement 5?
>>>>>>>>
>>>>>>>> a. Implicitly:
>>>>>>>>    - dtmfrecog: always DTMF-only recognition
>>>>>>>>    - speechrecog: depends on active grammar type
>>>>>>>>
>>>>>>>>
>>>>>>>>> if a dtmf grammar is active then DTMF input is "on"
>>>>>>>>> if a speech grammar is active then speech
>>>>>>>>>
>>>>>>>>>
>>>>     input is "on"
>>>>
>>>>
>>>>>>>>
>>>>>>>> b. Explicitly:
>>>>>>>>    - Add inputmodes header to RECOGNIZE
>>>>>>>>
>>>>>>>> Option a is David's "do what I mean case"; option b is
>>>>>>>>
>>>>>>>>
>>>>> the extra
>>>>>
>>>>>
>>>>>>>> dial for the client.
>>>>>>>>
>>>>>>>>
>>>>>>>>
>>>>>>>>
>>>>>>> Actually, that's not the point I was making with "do
>>>>>>>
>>>>>>>
>>>>     what I mean",
>>>>
>>>>
>>>>>>> but it's not essential to the discussion so let's move on.
>>>>>>>
>>>>>>>
>>>>>>>
>>>>>>>
>>>>>>>> It is worth noting that VoiceXML uses option b. This
>>>>>>>>
>>>>>>>>
>>>>     allows one to
>>>>
>>>>
>>>>>>>> activate both speech grammars and DTMF grammars (and
>>>>>>>>
>>>>>>>>
>>>>> therefore
>>>>>
>>>>>
>>>>>>>> be informed of any errors in the grammars at activation
>>>>>>>>
>>>>>>>>
>>>>> time) but
>>>>>
>>>>>
>>>>>>>> independently turn on whichever input mode you like e.g.
>>>>>>>>
>>>>>>>>
>>>>> perhaps
>>>>>
>>>>>
>>>>>>>> start with "both" then change to "dtmf".
>>>>>>>>
>>>>>>>>
>>>>>>>>
>>>>>>>>
>>>>>>> I'm not sure the VXML precedent is relevant here,
>>>>>>>
>>>>>>>
>>>>     because the
>>>>
>>>>
>>>>>>> application behind VXML is working a different part of the
>>>>>>>
>>>>>>>
>>>>> problem
>>>>>
>>>>>
>>>>>>> -  how to traverse a TUI dialog based on different
>>>>>>>
>>>>>>>
>>>>     parts of the
>>>>
>>>>
>>>>>>> input  space. In fact, I suspect that the VXML: choice was
>>>>>>> conditioned more  by limitations at the time it was
>>>>>>>
>>>>>>>
>>>>> specified than
>>>>>
>>>>>
>>>>>>> an underlying good  design choice. Clearly having to
>>>>>>>
>>>>>>>
>>>>     specify this
>>>>
>>>>
>>>>>>> in VXML make the job of  handling a TUI with nodes
>>>>>>>
>>>>>>>
>>>>     like "Say or
>>>>
>>>>
>>>>>>> press 5" harder rather than  easier.
>>>>>>>
>>>>>>>
>>>>>>>
>>>>>>
>>>>>> DB> The VoiceXML edge-case is pretty weird so it's not a major
>>>>>> concern. My main concern is that the client can
>>>>>>
>>>>>>
>>>>     indicate, somehow,
>>>>
>>>>
>>>>>> what the input modes are.
>>>>>>
>>>>>>
>>>>>>
>>>>>>
>>>>>>>
>>>>>>> Summing up, while I don't feel strongly one way or
>>>>>>>
>>>>>>>
>>>>     the other, I
>>>>
>>>>
>>>>>>> have  a preference for handling this as follows:
>>>>>>>
>>>>>>> a) Rename "START-OF-SPEECH" to
>>>>>>>
>>>>>>>
>>>>     "INTERESTING-INPUT-RECEIVED" or
>>>>
>>>>
>>>>>>> something equivalent.
>>>>>>> b) Include a parameter in the event saying what was
>>>>>>>
>>>>>>>
>>>>     interesting
>>>>
>>>>
>>>>>>> about  the input you received, with a registry of values
>>>>>>>
>>>>>>>
>>>> which
>>>>
>>>>
>>>>>>> includes:
>>>>>>>     - signal above noise floor
>>>>>>>     - speech
>>>>>>>     - dtmf
>>>>>>>     - (possibly) music
>>>>>>> c) allow the event to be generated multiple times
>>>>>>>
>>>>>>>
>>>>     during a request
>>>>
>>>>
>>>>>>>
>>>>>>>
>>>>>>>
>>>>>>>
>>>>>>
>>>>>> DB> I like these suggestions (START-OF-INPUT?).
>>>>>>
>>>>>>
>>>>     However, I don't
>>>>
>>>>
>>>>>> see how the problem of the optimised case is not
>>>>>>
>>>>>>
>>>>     solved by them. I
>>>>
>>>>
>>>>>> think the optimised case is fine if we have the
>>>>>>
>>>>>>
>>>>     following rules:
>>>>
>>>>
>>>>>>
>>>>>>
>>>>>>
>>>>> Yes, I hadn't thought through the optimized case as
>>>>>
>>>>>
>>>>     thoroughly as you.
>>>>
>>>>
>>>>> Your suggested method name is fine by me as well
>>>>>
>>>>>
>>>>>
>>>>>> 1. START-OF-SPEECH (and optimised bargin) is only generated
>>>>>>
>>>>>>
>>>>> for the
>>>>>
>>>>>
>>>>>> input type that is being listened for 2. A speechrecog
>>>>>>
>>>>>>
>>>>     listens for
>>>>
>>>>
>>>>>> DTMF if DTMF grammars are active, speech if speech
>>>>>>
>>>>>>
>>>>     grammars are
>>>>
>>>>
>>>>>> active, or speech and DTMF if both grammar types are active.
>>>>>>
>>>>>>
>>>>>>
>>>>> Works for me.
>>>>>
>>>>>
>>>>>
>>>>>>
>>>>>>
>>>>>>
>>>>>>> Note that all of the above I'm saying with my technical
>>>>>>>
>>>>>>>
>>>>> hat on and
>>>>>
>>>>>
>>>>>>> my chair hat off.
>>>>>>> Putting my chair hat on for a moment, we really need
>>>>>>>
>>>>>>>
>>>>     to get this
>>>>
>>>>
>>>>>>> spec  to last call, so at some point Eric or I is going to
>>>>>>>
>>>>>>>
>>>>> declare
>>>>>
>>>>>
>>>>>>> rough  consensus so we can move on.
>>>>>>>
>>>>>>>
>>>>>>>
>>>>>>
>>>>>> DB> Agreed!
>>>>>>
>>>>>>
>>>>>>
>>>>>>
>>>>>>>
>>>>>>> Dave Oran.
>>>>>>>
>>>>>>>
>>>>>>>
>>>>>>>> Dave
>>>>>>>>
>>>>>>>>
>>>>>>>>
>>>>>>>>
>>>>>>>>
>>>>>>>>
>>>>>>>>
>>>>>>>>
>>>>>>>>
>>>>>>>>
>>>>>>>>
>>>>>>>>
>>>>>>>> Sarvi makes a good point that adding the reason why the
>>>>>>>>
>>>>>>>>
>>>>> START-OF-
>>>>>
>>>>>
>>>>>>>> SPEECH occurred does not fix the optimised bargin case.
>>>>>>>>
>>>>>>>> dtmfrecog - listens for DTMF only
>>>>>>>> speechrecog - listens for DTMF only, or speech only,
>>>>>>>>
>>>>>>>>
>>>>     or speech &
>>>>
>>>>
>>>>>>>> DTMF
>>>>>>>>
>>>>>>>>
>>>>>>>>
>>>>>>>>
>>>>>>>>
>>>>>>>>
>>>>>>>>
>>>>>>>>
>>>>>>>> ----- Original Message ----- From: "Shanmugham, Saravanan"
>>>>>>>> <sarvi@cisco.com>
>>>>>>>> To: "David R Oran" <oran@cisco.com>; "Dave Burke"
>>>>>>>> <david.burke@voxpilot.com>
>>>>>>>> Cc: <speechsc@ietf.org>; "Klaus Reifenrath"
>>>>>>>> <Klaus.Reifenrath@Scansoft.com>
>>>>>>>> Sent: Tuesday, July 05, 2005 8:51 PM
>>>>>>>> Subject: RE: [Speechsc] START-OF-SPEECH in DTMF-only mode
>>>>>>>>
>>>>>>>>
>>>>>>>> inline.
>>>>>>>>
>>>>>>>>     -----Original Message-----
>>>>>>>>     From: David R Oran [mailto:oran@cisco.com]
>>>>>>>>     Sent: Tuesday, July 05, 2005 11:08 AM
>>>>>>>>     To: Dave Burke
>>>>>>>>     Cc: Shanmugham, Saravanan; Klaus Reifenrath;
>>>>>>>>
>>>>>>>>
>>>>     speechsc@ietf.org
>>>>
>>>>
>>>>>>>>     Subject: Re: [Speechsc] START-OF-SPEECH in DTMF-only
>>>>>>>>
>>>>>>>>
>>>> mode
>>>>
>>>>
>>>>>>>>
>>>>>>>>
>>>>>>>>     On Jul 5, 2005, at 12:54 PM, Dave Burke wrote:
>>>>>>>>
>>>>>>>>
>>>>>>>>
>>>>>>>>> Inline.
>>>>>>>>>
>>>>>>>>> Dave
>>>>>>>>>
>>>>>>>>> ----- Original Message ----- From:
>>>>>>>>>
>>>>>>>>>
>>>>     "Shanmugham, Saravanan"
>>>>
>>>>
>>>>>>>>> <sarvi@cisco.com>
>>>>>>>>> To: "David R Oran" <oran@cisco.com>; "Klaus
>>>>>>>>>
>>>>>>>>>
>>>> Reifenrath"
>>>>
>>>>
>>>>>>>>> <Klaus.Reifenrath@Scansoft.com>
>>>>>>>>> Cc: <speechsc@ietf.org>
>>>>>>>>> Sent: Tuesday, July 05, 2005 5:22 PM
>>>>>>>>> Subject: RE: [Speechsc] START-OF-SPEECH in
>>>>>>>>>
>>>>>>>>>
>>>>     DTMF-only mode
>>>>
>>>>
>>>>>>>>>
>>>>>>>>>
>>>>>>>>> I agree with Dave's analysis. The purpose of this
>>>>>>>>>
>>>>>>>>>
>>>> event
>>>>
>>>>
>>>>>>>>     was barge-in.
>>>>>>>>
>>>>>>>>
>>>>>>>>> And barge-in should happen for both DTMF and speech.
>>>>>>>>>
>>>>>>>>> Is there a case where you think it should not behave
>>>>>>>>>
>>>>>>>>>
>>>>>>>>     this way. If soe,
>>>>>>>>
>>>>>>>>
>>>>>>>>> please provide a scenario where you think
>>>>>>>>>    1. Barge-in should happen for DTMF and not voice or
>>>>>>>>>
>>>>>>>>>
>>>>>>>>     vice-versa.
>>>>>>>>
>>>>>>>>
>>>>>>>>>
>>>>>>>>> DB> You want to do a DTMF recognition only
>>>>>>>>>
>>>>>>>>>
>>>>     because it is
>>>>
>>>>
>>>>>>>> noisy.
>>>>>>>>
>>>>>>>>
>>>>>>>>> While waiting for DTMF input, the speechrecog resource
>>>>>>>>>
>>>>>>>>>
>>>>>>>>     (or advanced
>>>>>>>>
>>>>>>>>
>>>>>>>>> dtmfrecog) generates a START-OF-SPEECH because it
>>>>>>>>>
>>>>>>>>>
>>>> heard
>>>>
>>>>
>>>>>>>>     some speech.
>>>>>>>>
>>>>>>>>
>>>>>>>>> The client does not want to stop prompt playing unless
>>>>>>>>>
>>>>>>>>>
>>>>>>>>     DTMF was heard
>>>>>>>>
>>>>>>>>
>>>>>>>>> but it can't tell by the START-OF-SPEECH whether
>>>>>>>>>
>>>>>>>>>
>>>> speech
>>>>
>>>>
>>>>>>>>     or DTMF was
>>>>>>>>
>>>>>>>>
>>>>>>>>> heard. Similarly vice versa.
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>
>>>>>>>>     It's an interesting design question what part of the
>>>>>>>>     policy resides at the client and what at the server, and
>>>>>>>>     who makes the "final decision" about whether what was
>>>>>>>>     heard was relevant to the control channel. Right now we
>>>>>>>>     (IMO) have a weird partitioning in many cases where the
>>>>>>>>     client basically says "do what I mean", but there are no
>>>>>>>>     constraints of what the server actually does, and no
>>>>>>>>     normalized basis for the client to figure out what to
>>>>>>>>
>>>>>>>>
>>>> set
>>>>
>>>>
>>>>>>>>     various magic numbers to (e.g. sensitivity).
>>>>>>>>
>>>>>>>>     In this case the only thing the client needs to
>>>>>>>>
>>>>>>>>
>>>> decide is
>>>>
>>>>
>>>>>>>>     whether to kill the prompt because the server thinks
>>>>>>>>     something that would interfere with the feedback
>>>>>>>>     ear/mouth/finger control happened. What this
>>>>>>>>
>>>>>>>>
>>>>     says to me is
>>>>
>>>>
>>>>>>>>     that it isn't necessarily a good idea for the client to
>>>>>>>>     have more knobs to control the server
>>>>>>>>
>>>>>>>>
>>>>     (especially if those
>>>>
>>>>
>>>>>>>>     knows are just more value/policy input ungrounded in any
>>>>>>>>     physics/ acoustics). On the other hand, having the
>>>>>>>>
>>>>>>>>
>>>> server
>>>>
>>>>
>>>>>>>>     tell the client more about what it thinks is going on is
>>>>>>>>     probably valuable.
>>>>>>>>
>>>>>>>>     So, Coming to the point after this long rambling
>>>>>>>>     introduction, I think it would in fact be useful for the
>>>>>>>>     START-Of-SPEECH event to indicate some extra
>>>>>>>>
>>>>>>>>
>>>> information,
>>>>
>>>>
>>>>>>>>     for example:
>>>>>>>>     a) I got something enough above the noise floor
>>>>>>>>
>>>>>>>>
>>>>     to qualify
>>>>
>>>>
>>>>>>>>     for exceeding the "Sensisitvity" parameter you sent
>>>>>>>>
>>>>>>>>
>>>> in on
>>>>
>>>>
>>>>>>>>     the request but I really can't tell what it is
>>>>>>>>
>>>>>>>>
>>>>     (could be a
>>>>
>>>>
>>>>>>>>     hippopatmus fart, or a siren in the background,
>>>>>>>>
>>>>>>>>
>>>>     or captain
>>>>
>>>>
>>>>>>>>     crunch trying to whistle DTMF).
>>>>>>>>     b) I think I'm hearing speech
>>>>>>>>     c) I think I'm hearing DTMF
>>>>>>>>
>>>>>>>> Though I agree with your former part of your
>>>>>>>>
>>>>>>>>
>>>>     response. I am not
>>>>
>>>>
>>>>>>>> sure I agree with your proposed solution.
>>>>>>>> The way I see this problem is that, it is more of what
>>>>>>>>
>>>>>>>>
>>>>> constitues a
>>>>>
>>>>>
>>>>>>>> barge-in event. This boils down to whether it is
>>>>>>>>
>>>>>>>>
>>>>     speech, DTMF or
>>>>
>>>>
>>>>>>>> both.
>>>>>>>> This is inturn boils down to what type of recognizer
>>>>>>>>
>>>>>>>>
>>>>> resource we are
>>>>>
>>>>>
>>>>>>>> using, dtmf-recog, speech-recog, and speech-only-recog(we
>>>>>>>>
>>>>>>>>
>>>>> don't have
>>>>>
>>>>>
>>>>>>>> this and I don't think we should add it, but think
>>>>>>>>
>>>>>>>>
>>>>     of this as a
>>>>
>>>>
>>>>>>>> place holder that explains the concept).
>>>>>>>>
>>>>>>>> A client knowing what type of barge-in happenned, does
>>>>>>>>
>>>>>>>>
>>>>> not impact
>>>>>
>>>>>
>>>>>>>> the
>>>>>>>> barge-in operation itself as it may be too late(for
>>>>>>>>
>>>>>>>>
>>>>     the optimized
>>>>
>>>>
>>>>>>>> barge-in case). It may have other use cases, and if we
>>>>>>>>
>>>>>>>>
>>>>> can identify
>>>>>
>>>>>
>>>>>>>> them, I don't mind adding support for the
>>>>>>>>
>>>>>>>>
>>>>     START-OF-SPEECH event to
>>>>
>>>>
>>>>>>>> say what type of barge-in happenned. But that itself does
>>>>>>>>
>>>>>>>>
>>>> not
>>>>
>>>>
>>>>> solve the
>>>>>
>>>>>
>>>>>>>> original problem raised. Refer to my previous response.
>>>>>>>>
>>>>>>>> The solution lies in defining what what is a
>>>>>>>>
>>>>>>>>
>>>>     barge-in event.
>>>>
>>>>
>>>>>>>> That  boils
>>>>>>>> down to what type of recognition is happenning,
>>>>>>>>
>>>>>>>>
>>>>> dtmf-only, speech-
>>>>>
>>>>>
>>>>>>>> dtmf
>>>>>>>> or speech-only. We do not support speech-only as a
>>>>>>>>
>>>>>>>>
>>>>     resource today,
>>>>
>>>>
>>>>>>>> the question is do we need a header to force it.
>>>>>>>>
>>>>>>>> Sarvi
>>>>>>>>
>>>>>>>>
>>>>>>>>
>>>>>>>>
>>>>>>>>>    2. You would benefit from the client knowing
>>>>>>>>>
>>>>>>>>>
>>>>> what caused
>>>>>
>>>>>
>>>>>>>> the
>>>>>>>>
>>>>>>>>
>>>>>>>>> barge-in, DTMF Vs speech.
>>>>>>>>>
>>>>>>>>> DB> See previous comment. And previous e-mail:
>>>>>>>>>
>>>>>>>>>
>>>>     either add an
>>>>
>>>>
>>>>>>>>> inputmodes header (taking value speech, dtmf, both) to
>>>>>>>>>
>>>>>>>>>
>>>>>>>>     the RECOGNIZE
>>>>>>>>
>>>>>>>>
>>>>>>>>> request or add a header to the START-OF-SPEECH event
>>>>>>>>>
>>>>>>>>>
>>>>>>>>     indicating DTMF
>>>>>>>>
>>>>>>>>
>>>>>>>>> or speech.
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>
>>>>>>>>     I'm leaning in your direction on this latter point - as
>>>>>>>>     should be evident from what I wrote above.
>>>>>>>>
>>>>>>>>
>>>>>>>>
>>>>>>>>> Sarvi
>>>>>>>>>
>>>>>>>>>     -----Original Message-----
>>>>>>>>>     From: speechsc-bounces@ietf.org
>>>>>>>>>     [mailto:speechsc-bounces@ietf.org] On Behalf Of
>>>>>>>>>
>>>>>>>>>
>>>>> David R
>>>>>
>>>>>
>>>>>>>> Oran
>>>>>>>>
>>>>>>>>
>>>>>>>>>     Sent: Tuesday, July 05, 2005 5:24 AM
>>>>>>>>>     To: Klaus Reifenrath
>>>>>>>>>     Cc: 'speechsc@ietf.org'
>>>>>>>>>     Subject: Re: [Speechsc] START-OF-SPEECH in
>>>>>>>>>
>>>>>>>>>
>>>>> DTMF-only mode
>>>>>
>>>>>
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>     On Jul 5, 2005, at 3:46 AM, Reifenrath,
>>>>>>>>>
>>>>>>>>>
>>>>     Klaus wrote:
>>>>
>>>>
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>> The current spec is not clear when
>>>>>>>>>>
>>>>>>>>>>
>>>>> START-OF-SPEECH need
>>>>>
>>>>>
>>>>>>>>>     to be send in
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>> the following scenarios:
>>>>>>>>>> A) The client requested a DTMF Recognizer. Is
>>>>>>>>>>
>>>>>>>>>>
>>>> the
>>>>
>>>>
>>>>>>>>>     START-OF-SPEECH
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>> event send to the client also if speech
>>>>>>>>>>
>>>>>>>>>>
>>>>     was detected?
>>>>
>>>>
>>>>>>>>>     I suspect so, since one of the prime purposes
>>>>>>>>>
>>>>>>>>>
>>>>> is to enable
>>>>>
>>>>>
>>>>>>>>>     client- mediated barge-in handling. However, if
>>>>>>>>>
>>>>>>>>>
>>>> the
>>>>
>>>>
>>>>>>>>>     recognizer is in fact only capable of
>>>>>>>>>
>>>>>>>>>
>>>>     recognizing DTMF
>>>>
>>>>
>>>>>>>>>     then it may in fact not report anythin
>>>>>>>>>
>>>>>>>>>
>>>>     unless it's using
>>>>
>>>>
>>>>>>>>>     some primitive thresholding machinery, like a SN
>>>>>>>>>
>>>>>>>>>
>>>>>>>> threshold.
>>>>>>>>
>>>>>>>>
>>>>>>>>>> B) The client requested a Speech
>>>>>>>>>>
>>>>>>>>>>
>>>>     Recognizer, but only
>>>>
>>>>
>>>>>>>>>     activated DTMF
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>> grammars. Is the START-OF-SPEECH event
>>>>>>>>>>
>>>>>>>>>>
>>>>     send to the
>>>>
>>>>
>>>>>>>>>     client also if
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>> speech was detected?
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>     Again, I'd say yes, for the same reason as above.
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>> I think in both cases START-OF-SPEECH should
>>>>>>>>>>
>>>>>>>>>>
>>>> only
>>>>
>>>>
>>>>>>>>     be send after
>>>>>>>>
>>>>>>>>
>>>>>>>>>> detecting a DTMF digit (see Figure 12 of
>>>>>>>>>>
>>>>>>>>>>
>>>> VoiceXML
>>>>
>>>>
>>>>>>>>     2.0: http://
>>>>>>>>
>>>>>>>>
>>>>>>>>>> www.w3.org/TR/voicexml20/#dmlATiming).
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>     We seem to have reached different
>>>>>>>>>
>>>>>>>>>
>>>>     conclusions. I'd be
>>>>
>>>>
>>>>>>>>>     interested in why you think my analysis
>>>>>>>>>
>>>>>>>>>
>>>>     above is wrong.
>>>>
>>>>
>>>>>>>>>
>>>>>>>>>     Dave.
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>> Klaus
>>>>>>>>>>
>>>>>>>>>> _______________________________________________
>>>>>>>>>> Speechsc mailing list
>>>>>>>>>> Speechsc@ietf.org
>>>>>>>>>> https://www1.ietf.org/mailman/listinfo/speechsc
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>
>>>>>>>>>     _______________________________________________
>>>>>>>>>     Speechsc mailing list
>>>>>>>>>     Speechsc@ietf.org
>>>>>>>>>     https://www1.ietf.org/mailman/listinfo/speechsc
>>>>>>>>>
>>>>>>>>>
>>>>>>>>> _______________________________________________
>>>>>>>>> Speechsc mailing list
>>>>>>>>> Speechsc@ietf.org
>>>>>>>>> https://www1.ietf.org/mailman/listinfo/speechsc
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>
>>>>>>>>
>>>>>>>>
>>>>>>>> _______________________________________________
>>>>>>>> Speechsc mailing list
>>>>>>>> Speechsc@ietf.org
>>>>>>>> https://www1.ietf.org/mailman/listinfo/speechsc
>>>>>>>>
>>>>>>>>
>>>>>>>>
>>>>>>>>
>>>>>>>
>>>>>>> _______________________________________________
>>>>>>> Speechsc mailing list
>>>>>>> Speechsc@ietf.org
>>>>>>> https://www1.ietf.org/mailman/listinfo/speechsc
>>>>>>>
>>>>>>>
>>>>>>
>>>>>>
>>>>>>
>>>>>
>>>>> _______________________________________________
>>>>> Speechsc mailing list
>>>>> Speechsc@ietf.org
>>>>> https://www1.ietf.org/mailman/listinfo/speechsc
>>>>>
>>>>>
>>>>>
>>>>
>>>>
>>>>     _______________________________________________
>>>>     Speechsc mailing list
>>>>     Speechsc@ietf.org
>>>>     https://www1.ietf.org/mailman/listinfo/speechsc
>>>>
>>>>
>>>> _______________________________________________
>>>> Speechsc mailing list
>>>> Speechsc@ietf.org
>>>> https://www1.ietf.org/mailman/listinfo/speechsc
>>>>
>>>>
>>>> _______________________________________________
>>>> Speechsc mailing list
>>>> Speechsc@ietf.org
>>>> https://www1.ietf.org/mailman/listinfo/speechsc
>>>>
>>>>
>>>>
>>>
>>> _______________________________________________
>>> Speechsc mailing list
>>> Speechsc@ietf.org
>>> https://www1.ietf.org/mailman/listinfo/speechsc
>>>
>>>
>>>
>>>
>>>
>>
>> _______________________________________________
>> Speechsc mailing list
>> Speechsc@ietf.org
>> https://www1.ietf.org/mailman/listinfo/speechsc
>

_______________________________________________
Speechsc mailing list
Speechsc@ietf.org
https://www1.ietf.org/mailman/listinfo/speechsc



From speechsc-bounces@ietf.org Wed Jul 13 11:08:01 2005
Received: from localhost.localdomain ([127.0.0.1] helo=megatron.ietf.org)
	by megatron.ietf.org with esmtp (Exim 4.32)
	id 1Dsiq9-00048k-Bo; Wed, 13 Jul 2005 11:08:01 -0400
Received: from odin.ietf.org ([132.151.1.176] helo=ietf.org)
	by megatron.ietf.org with esmtp (Exim 4.32) id 1Dsiq2-00047L-Vs
	for speechsc@megatron.ietf.org; Wed, 13 Jul 2005 11:08:00 -0400
Received: from ietf-mx.ietf.org (ietf-mx [132.151.6.1])
	by ietf.org (8.9.1a/8.9.1a) with ESMTP id LAA16838
	for <speechsc@ietf.org>; Wed, 13 Jul 2005 11:07:52 -0400 (EDT)
Received: from fw01.db01.voxpilot.com ([212.17.54.82] helo=mail.voxpilot.com)
	by ietf-mx.ietf.org with esmtp (Exim 4.43) id 1DsjIO-0004mR-4G
	for speechsc@ietf.org; Wed, 13 Jul 2005 11:37:17 -0400
Received: from daburkewxp (67.5-14-84.ripe.coltfrance.com [84.14.5.67])
	by mail.voxpilot.com (Postfix) with ESMTP
	id 3FFB1214041; Wed, 13 Jul 2005 15:07:22 +0000 (GMT)
Message-ID: <020a01c587c4$f2eb3dd0$6901a8c0@db01.voxpilot.com>
From: "Dave Burke" <david.burke@voxpilot.com>
To: "David R Oran" <oran@cisco.com>
References: <7DE7C4EF3B7C8B4B82955191378290D802ED4176@mtb1exch01.nuance.com>
	<C7673C4C-0624-47A1-B1DF-115FB284785A@cisco.com>
	<028f01c5865e$7f92c4a0$038ae9d5@db01.voxpilot.com>
	<9B4BE6A2-82ED-4C8B-BAE5-006FB0C2DBBB@cisco.com>
Subject: Re: [Speechsc] START-OF-SPEECH in DTMF-only mode -> summary(?)
Date: Wed, 13 Jul 2005 17:07:19 +0100
MIME-Version: 1.0
Content-Type: text/plain; format=flowed; charset="iso-8859-1";
	reply-type=response
Content-Transfer-Encoding: 7bit
X-Priority: 3
X-MSMail-Priority: Normal
X-Mailer: Microsoft Outlook Express 6.00.2900.2180
X-MimeOLE: Produced By Microsoft MimeOLE V6.00.2900.2180
X-Spam-Score: 0.0 (/)
X-Scan-Signature: 8c313f1b0f40db66833fce0a5a529476
Content-Transfer-Encoding: 7bit
Cc: speechsc@ietf.org, Pierre Forgues <forgues@nuance.com>,
	Eric Burger <eburger@brooktrout.com>, "Shanmugham,
	Saravanan" <sarvi@cisco.com>
X-BeenThere: speechsc@ietf.org
X-Mailman-Version: 2.1.5
Precedence: list
List-Id: Speech Services Control Working Group <speechsc.ietf.org>
List-Unsubscribe: <https://www1.ietf.org/mailman/listinfo/speechsc>,
	<mailto:speechsc-request@ietf.org?subject=unsubscribe>
List-Post: <mailto:speechsc@ietf.org>
List-Help: <mailto:speechsc-request@ietf.org?subject=help>
List-Subscribe: <https://www1.ietf.org/mailman/listinfo/speechsc>,
	<mailto:speechsc-request@ietf.org?subject=subscribe>
Sender: speechsc-bounces@ietf.org
Errors-To: speechsc-bounces@ietf.org

Inline

Dave

----- Original Message ----- 
From: "David R Oran" <oran@cisco.com>
To: "Dave Burke" <david.burke@voxpilot.com>
Cc: "Pierre Forgues" <forgues@nuance.com>; <speechsc@ietf.org>; "Shanmugham, 
Saravanan" <sarvi@cisco.com>; "Eric Burger" <eburger@brooktrout.com>
Sent: Wednesday, July 13, 2005 2:20 PM
Subject: Re: [Speechsc] START-OF-SPEECH in DTMF-only mode -> summary(?)


>
> On Jul 11, 2005, at 5:21 PM, Dave Burke wrote:
>
>> I don't think it is possible in practice to separate what input  type is 
>> considered potential barge-in and what input type is being  recognised 
>> because a recognition hypothesis is always generated  when a barge-in 
>> occurs.
> True, but what I'm trying to tease apart is whether to declare barge- in 
> from the point of view of stopping output from simply detecting  input for 
> the purposes of cranking up the non-signal-processing parts  of the 
> recognition engine. Did you not find my "mark the phrases"  application 
> example compelling?

DB> Indeed it's a nice application (albeit producing disturbing memories of 
musical theory exams from my distant childhood!)... Can't this application 
be created by setting Kill-On-Barge-In to false on the SPEAK request that 
queued the music?

>
>> For example, imagine a client invoked a speechrecog and asked it to  only 
>> consider DTMF for barge-in. Once speech is detected, although  no 
>> barge-in would happen, a RECOGNITION-COMPLETE will result after 
>> Speech-Incomplete-Timeout milliseconds of silence thus ending the 
>> recognition state machine with a 'nomatch'.
> Sure, but of course the timeout could be very long... I must be  dense, 
> because I see this relationship as tenuous rather than  strongly coupled.

DB> But the recognition will terminate with nomatch at some point if it 
hears noise/speech first regardless of a good DTMF sequence (I say first 
only out of pragmatism because that's what modern day recognizers do - the 
first input type recognized results in turning off the other input type for 
the remainder of the recognition). The recognition will need to be restarted 
but then some DTMF is lost, timers have to be reset etc.

>
>> The converse example applies for a speech-only recognition (DTMF- 
>> Interdigit-Timeout replaces the Speech-Incomplete-Timeout).
>>
>> Although I prefer the idea of an InputModes header in RECOGNIZE  (it's 
>> the cleanest solution for VoiceXML implementors), I am  willing to 
>> compromise on using the grammar type to determine the  input types (in 
>> what follows I use input types to mean both what is  considered potential 
>> barge-in and what type of input is to be  recognised). The main issue 
>> with this approach is incompatibility  with VoiceXML (recall VoiceXML 
>> grammar activation is independent of  what inputmodes are set).
>>
> And voice XML does not have any explicit control over barge-in  either. We 
> could of course back off not have clients involved at all  in barge-in 
> processing, but I think that would be a mistake.

DB> I'm not sure I understand this (unless you mean VoiceXML doesn't 
decouple barge-in input type from recognition input type in which case I 
agree). VoiceXML allows you to switch on/off barge-in by setting the 
barge-in attribute on <prompt> (maps to MRCP's Kill-On-Barge-In). It allows 
you to simultaneously set the input type for recognition and barge-in. 
VoiceXML does have a bargeintype attribute although this is just to turn on 
hotword...

>
>> This incompatibility results in side-effects that, in my opinion,  are 
>> inconsequential for real applications:
>>    a. speech barge-in (noise) and nomatch will not happen for 
>> inputmodes="both" when no speech grammars are activated
>>    b. dtmf barge-in and nomatch will not happen for  inputmodes="both" 
>> when no dtmf grammars are activated
>>    c. grammars with errors will not be detected if the grammar type  is 
>> not compatible with the inputmodes property
>> A workaround for the really conscientious VoiceXML platform is to 
>> activate a dummy grammar (maybe with the NULL rule) of type speech 
>> (dtmf) when inputmodes="both" is set but no speech (dtmf) grammar  is 
>> activated in the application.
>>
> I agree with this assessment. I'll also point out that the  flexibility 
> will allow a class of VXML applications that can't be  done today (e.g. 
> ones that don't stop prompts on hearing input.

DB> This is possible already with Kill-On-Barge-On false

>
>> -----------------
>>
>> Summarising the proposed changes again  (refined a little and to  avoid 
>> trawling though this massive thread!):
>>
>> 1. For a speechrecog resource, the types of the activated grammars 
>> determine the input types the recogniser processes and considers  for 
>> potential barge-in.
>>
> Good. Agree.
>
>> 2. Clarify a dtmfrecog only processes DTMF and hence can only  generate 
>> barge-in / START-OF-SPEECH for DTMF inputs
>>
> I'm ok with this even though I'd rather keep input detection and 
> recognition behavior orthogonal to accommodate smarter server 
> implementations and to give better consistency between the server- 
> optimized and client-intermediated cases. In particular, I do really  feel 
> like we need a way for the client to tell a server to NOT try to  outguess 
> it by doing optimized barge in processing since the  application might in 
> fact not want the output stopped.
>
>
>> Nice-to-have:
>>
>> 3. Change START-OF-SPEECH to START-OF-INPUT
>>
> I'm in favor of this.
>
>> 4. Add a header of InputType to START-OF-INPUT. Current specified  values 
>> are "dtmf" or "speech".
>>
> I think we really need this, and would press again that we make this  a 
> list of values and future-proof it with the other values we suspect  will 
> be quite useful (e.g. music, gesture).
>
> Dave. (chair hat sort-of-on since we're trying to get final consensus 
> here).
>
>
>> -----------------
>>
>> Dave
>>
>> ----- Original Message ----- From: "David R Oran" <oran@cisco.com>
>> To: "Pierre Forgues" <forgues@nuance.com>
>> Cc: <speechsc@ietf.org>; "Shanmugham, Saravanan" <sarvi@cisco.com>; 
>> "Eric Burger" <eburger@brooktrout.com>; "Dave Burke" 
>> <david.burke@voxpilot.com>
>> Sent: Monday, July 11, 2005 2:15 PM
>> Subject: Re: [Speechsc] START-OF-SPEECH in DTMF-only mode -> summary (?)
>>
>>
>>
>>>
>>> On Jul 8, 2005, at 11:36 AM, Pierre Forgues wrote:
>>>
>>>
>>>>
>>>>
>>>> -----Original Message-----
>>>> From: speechsc-bounces@ietf.org [mailto:speechsc- bounces@ietf.org] On
>>>> Behalf Of David R Oran
>>>> Sent: Friday, July 08, 2005 10:05 AM
>>>> To: Dave Burke
>>>> Cc: speechsc@ietf.org; Shanmugham, Saravanan; Eric Burger
>>>> Subject: Re: [Speechsc] START-OF-SPEECH in DTMF-only mode ->  summary 
>>>> (?)
>>>>
>>>>
>>>> On Jul 8, 2005, at 6:59 AM, Dave Burke wrote:
>>>>
>>>>
>>>>
>>>>> I agree with your opinion on what constitutes a barge-in. Based on
>>>>> Klaus' e-mail, I am more concerned that using the grammar type to
>>>>> decide what constitutes a barge-in is going to make VoiceXML
>>>>> implementations difficult.
>>>>>
>>>>> It is easy to map VoiceXML application selected inputmodes to MRCP
>>>>> resources:
>>>>> a. inputmodes="dtmf" -> Use a dtmfrecog
>>>>> b. inputmodes="speech" -> Use a "speech-only-recog"
>>>>> c. inputmodes="both" -> Use a speechrecog (or a combination of
>>>>> "speech-only-recog" + dtmfrecog)
>>>>>
>>>>> The only problem is (b). While not as important as being able to do
>>>>> DTMF-only recognition, I believe we DO need to support speech-only
>>>>> recongition so as (a) to avoid unnecessary limitations in VUI
>>>>> design, and (b) to facilitate VoiceXML implementations. I think we
>>>>> should add a header to RECOGNIZE so the client is able to always
>>>>> _explicitly_ set the inputmodes.
>>>>>
>>>>>
>>>>>
>>>> I'm not sure I buy this, since what the recognizer is looking for  and
>>>> what constitutes an input that should be considered a potential  barge-
>>>> in strike me as independent. I'm similarly not persuaded that we  need
>>>> the flexibility to set the input mode independently of the grammar,
>>>> since it leads to all sorts of inconsistent states (e.g. speech-only
>>>> grammar with an input-mode of dtmf). On the other hand I can see the
>>>> need for the client to specify what sorts of input out to generate
>>>> the start-of-input (nee start-of-speech) event on.
>>>>
>>>> Pmf> The mechanism for a client to specify the type of input is  using
>>>> grammars.  These can have DTMF and/or speech requirements.  If  you add
>>>> the complexity of an independent header to specify the input mode  then
>>>> you will need to document the behavior for inconsistencies.
>>>>
>>>>
>>> Right. I generally agree with Pierre, which is why I proposed a 
>>> compromise position where the client can use a header to specify  what 
>>> types of input should be considered the start of something   potentially 
>>> interesting from a barge-in point of view, as opposed  to  a directive 
>>> to the server for what to recognize.
>>>
>>> I don't think this is crucial functionality, but it does make for   more 
>>> consistent behavior between the optimized and non-optimized   barge-in 
>>> cases. In fact, I can see cases where the client doesn't   want any 
>>> barge-in processing to happen. Imagine an application  called  "mark the 
>>> musical phrases", where the client listens to  playout of a  musical 
>>> score, and hits various DTMF buttons to  indicate markers  (e.g. 
>>> end-of-phrase, end-of-theme, key-change)  while listening. For  this 
>>> application you want to make sure that  barge-in doesn't stop the 
>>> music!
>>>
>>>
>>>>
>>>>
>>>>> In retrospect, I don't like the idea of multiple START-OF-SPEECH
>>>>> events being generated from the same media resource. This is
>>>>> because one assumes that the START-OF-SPEECH should be of the same
>>>>> type as the hypothesis returned in the RECOGNITION-COMPETE message
>>>>> - most implementations, on hearing one input mode type disable the
>>>>> recogniser of the other type.
>>>>>
>>>>>
>>>> I'm not sure I follow this logic, but I'm not wedded to the idea of
>>>> allowing multiple events. It was trying to solve the problem of
>>>> ambiguity around what the recognizer was hearing and the idea that
>>>> the client might care. If you don't think the client will ever care,
>>>> then we don't need the capability.
>>>> Pmf> I agree we should not have multiple SOS events.  Too much 
>>>> chatter.
>>>>
>>>>
>>> I'm not concerned about the "chatter" since this is a low- bandwidth 
>>> control channel even with multiple events, but as I said  it's a small 
>>> point and I'm happy to concede.
>>>
>>>
>>>>
>>>>
>>>>
>>>>> I do like David's idea of renaming START-OF-SPEECH to something
>>>>> like START-OF-INPUT and carrying a type header because it is neater
>>>>> and more extensible.
>>>>>
>>>>> So in summary, I propose we modify the spec to:
>>>>>
>>>>> 1. Clarify what constitutes a barge-in for a dtmfrecog and
>>>>> speechrecog as per Sarvi's e-mail (and in agreement with Klaus' for
>>>>> dtmfrecog).
>>>>>
>>>>>
>>>>>
>>>> Hmmm, ok, but I think the issue is actually clarifying what the
>>>> recognizer declares as "interesting input", which it may decide also
>>>> constitutes a barge-in in the optimized case, and in either case
>>>> reports to the client that it heard.
>>>> Pmf> I'm not going to argue changing the name of the event, but my
>>>> preference would be to keep existing method names unless there is a
>>>> clear reason for changing - which I do not see here.
>>>>
>>>>
>>> I do see a clear reason, since the thing that you see the start  of  may 
>>> not be speech. I like (I think it was Dave's suggestion)  "start- 
>>> of-input:" since it mirrors the other method names.
>>>
>>>>
>>>>
>>>>> 2. Specifiy an InputModes header to RECOGNIZE (defaults to "both",
>>>>> can also be "speech" or "DTMF"). Setting to speech for a
>>>>> speechrecog results in the hypothesised "speech-only-recog".
>>>>> Setting to DTMF for a speechrecog is equivalent to using a
>>>>> dtmfrecog. Edge cases: Setting to speech for a dtmfrecog will
>>>>> result in a noinput as will setting to DTMF for a speechrecog which
>>>>> does not support DTMF.
>>>>>
>>>>>
>>>>>
>>>> I'm ok with having a header for the client to tell the server  when it
>>>> would like the start-of-input event to be generated and what the
>>>> client considers to be the "interesting input" that the recognizer
>>>> should use to do the discrimination, and possibly do barge-in
>>>> processing in the optimized case.
>>>>
>>>> I'm less ok with the proposed domain of values, since it will have
>>>> extensibility problems. Especially problematical is have a code  point
>>>> of "both" since that will be ambiguous if we even define a "music"
>>>> recognizer, or a "gesture" recognizer using video input. Here's my
>>>> counter-proposal:
>>>>
>>>> Create a header on the recognizer resource requests called "start-
>>>> input-on:". Define a registry of values, with the following three
>>>> values initially defined:
>>>>      - dtmf
>>>>      - speech
>>>>      - music
>>>>
>>>> The semantics would be that the resource is to generate a "start-of-
>>>> input" event, apply the defined grammars to what follows, and do any
>>>> local optimized barge-in processing if ANY of the enumerated input
>>>> types is detected. That way you can say
>>>>      "start-input-on: dtmf" if you want to just  dtmf,
>>>>      "start-input-on: speech" if you want just speech (dtmf input
>>>> would be ignored even if the grammar supported dtmf)
>>>>      "start-input- on: speech, dtmf" if you wanted both
>>>> etc.
>>>>
>>>> Pmf> Doesn't this come back to the previous point of having two
>>>> independent ways to specify input modes?  Through grammars and this
>>>> proposed new header?  If you add this then you will have to document
>>>> behavior on inconsistencies.  Anyway, if you proceed then I would
>>>> definitely make this optional
>>>>
>>>>
>>> I hope I explained the intent above. This header would not define  the 
>>> input types that the recognize would process, but rather the  input 
>>> types the client wants it to consider potential barge-in  and hence 
>>> generate the "start-of-input:" event for and do any  local barge-in 
>>> optimizations for.
>>>
>>>
>>> Dave.
>>>
>>>
>>>>
>>>>
>>>>> 3. Change START-OF-SPEECH to START-OF-INPUT. Clarify it occurs once
>>>>> and coincides with barge-in.
>>>>>
>>>>>
>>>>>
>>>> Ok for the "only once", but I'd like to tighten it up to talk about
>>>> more than just barge-in, as I suggested above.
>>>>
>>>>
>>>>
>>>>> 4. Add a header of InputType to START-OF-INPUT. Current specified
>>>>> values are "dtmf" or "speech".
>>>>>
>>>>>
>>>>>
>>>> Ok, with slight modification. Say that the syntax of the "input- type"
>>>> header is a single-valued subset of the registered value(s) that are
>>>> defined for the start-input-on: header.
>>>>
>>>> Comments?
>>>>
>>>> Dave O. (technical hat on, chair hat off).
>>>>
>>>>
>>>>
>>>>> Dave
>>>>>
>>>>>
>>>>>
>>>>> ----- Original Message ----- From: "Shanmugham, Saravanan"
>>>>> <sarvi@cisco.com>
>>>>> To: "Eric Burger" <eburger@brooktrout.com>; <speechsc@ietf.org>
>>>>> Sent: Thursday, July 07, 2005 5:52 PM
>>>>> Subject: RE: [Speechsc] START-OF-SPEECH in DTMF-only mode ->  summary
>>>>> (?)
>>>>>
>>>>>
>>>>> I don't think generating multiple START-OF-SPEECH events is a
>>>>> solution.
>>>>> We still haven't addressed, what constitutes a barge-in, for the
>>>>> optimised case. That should also be the single point when a single
>>>>> START-OF-SPEECH(or whathever else you want to name it) should be
>>>>> generated.
>>>>>
>>>>> That point, in my opinion should be
>>>>>   1. For "dtmf-recog" resources should be the beginning of a  DTMF key
>>>>> press.
>>>>>   2. For "speech-recog" resources should be the beginning of a DTMF
>>>>> key
>>>>> press or the beginning of speech. This is should be irrespective of
>>>>> what
>>>>> type of grammar is being used. Coz even numbers only grammars can
>>>>> still
>>>>> be spoken and hence cannot be assumed to be a cue for DTMF only
>>>>> recognition.
>>>>>   3. For "speech-only-recog" resources(which are not defined today)
>>>>> the
>>>>> time to barge-in is the beginning of speech. I don't see a need for
>>>>> such
>>>>> a resource today. But I am mentioning this for completeness.
>>>>>
>>>>> Sarvi
>>>>>
>>>>>
>>>>>     -----Original Message-----
>>>>>     From: speechsc-bounces@ietf.org
>>>>>     [mailto:speechsc-bounces@ietf.org] On Behalf Of Eric Burger
>>>>>     Sent: Thursday, July 07, 2005 5:32 AM
>>>>>     To: speechsc@ietf.org
>>>>>     Subject: RE: [Speechsc] START-OF-SPEECH in DTMF-only mode
>>>>>     -> summary(?)
>>>>>
>>>>>     Can we declare consensus?
>>>>>
>>>>>
>>>>>
>>>>>> -----Original Message-----
>>>>>> From: speechsc-bounces@ietf.org
>>>>>>
>>>>>>
>>>>>     [mailto:speechsc-bounces@ietf.org]
>>>>>
>>>>>
>>>>>> Sent: Wednesday, July 06, 2005 10:45 AM
>>>>>> To: Dave Burke
>>>>>> Cc: speechsc@ietf.org; Shanmugham, Saravanan; Klaus Reifenrath
>>>>>> Subject: Re: [Speechsc] START-OF-SPEECH in DTMF-only mode ->
>>>>>> summary(?)
>>>>>>
>>>>>> I think we're getting close. I though about snipping out
>>>>>>
>>>>>>
>>>>>     some pieces
>>>>>
>>>>>
>>>>>> to cut down the text, but I realized the context is
>>>>>>
>>>>>>
>>>>>     still needed. See
>>>>>
>>>>>
>>>>>> inline.
>>>>>>
>>>>>> On Jul 6, 2005, at 10:29 AM, Dave Burke wrote:
>>>>>>
>>>>>>
>>>>>>
>>>>>>> Inline.
>>>>>>>
>>>>>>> Dave
>>>>>>>
>>>>>>> ----- Original Message ----- From: "David R Oran"
>>>>>>>
>>>>>>>
>>>>>     <oran@cisco.com>
>>>>>
>>>>>
>>>>>>> To: "Dave Burke" <david.burke@voxpilot.com>
>>>>>>> Cc: <speechsc@ietf.org>; "Shanmugham, Saravanan"
>>>>>>>
>>>>>>>
>>>>>     <sarvi@cisco.com>;
>>>>>
>>>>>
>>>>>>> "Klaus Reifenrath" <Klaus.Reifenrath@Scansoft.com>
>>>>>>> Sent: Wednesday, July 06, 2005 1:07 PM
>>>>>>> Subject: Re: [Speechsc] START-OF-SPEECH in DTMF-only
>>>>>>>
>>>>>>>
>>>>>     mode -> summary
>>>>>
>>>>>
>>>>>>> (?)
>>>>>>>
>>>>>>>
>>>>>>>
>>>>>>>
>>>>>>>
>>>>>>>>
>>>>>>>> On Jul 6, 2005, at 5:31 AM, Dave Burke wrote:
>>>>>>>>
>>>>>>>>
>>>>>>>>
>>>>>>>>
>>>>>>>>> + Attempting to summarise:
>>>>>>>>>
>>>>>>>>> 1. START-OF-SPEECH is useful for the client to know when
>>>>>>>>>
>>>>>>>>>
>>>>>> to stop
>>>>>>
>>>>>>
>>>>>>>>> playing prompts the in non-optimised case 2.
>>>>>>>>>
>>>>>>>>>
>>>>>     START-OF-SPEECH is
>>>>>
>>>>>
>>>>>>>>> useful for the client to calculate the bargin  time
>>>>>>>>>
>>>>>>>>>
>>>>>     (e.g. VoiceXML
>>>>>
>>>>>
>>>>>>>>> 2.1 <mark>)
>>>>>>>>> 3. In the optimised case, a bargin automatically
>>>>>>>>>
>>>>>>>>>
>>>>>     stops prompt
>>>>>
>>>>>
>>>>>>>>> playing (assuming prompts barginable) 4. Because of
>>>>>>>>>
>>>>>>>>>
>>>>>     the previous
>>>>>
>>>>>
>>>>>>>>> point, the question of what input type  caused
>>>>>>>>>
>>>>>>>>>
>>>>>     bargin is different
>>>>>
>>>>>
>>>>>>>>> and less important to
>>>>>>>>>
>>>>>>>>>
>>>>>> what input
>>>>>>
>>>>>>
>>>>>>>>> type(s)  the recogniser is listening for
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>
>>>>>>>> I'm not sure I follow point 4. Could you elaborate?
>>>>>>>>
>>>>>>>>
>>>>>>>>
>>>>>>>
>>>>>>> DB> Adding a parameter to the START-OF-SPEECH event would
>>>>>>>
>>>>>>>
>>>>>> certainly
>>>>>>
>>>>>>
>>>>>>> allow the client to ignore (i.e. let prompts continue
>>>>>>>
>>>>>>>
>>>>>     playing) the
>>>>>
>>>>>
>>>>>>> event if the event type is not of interest (e.g. the
>>>>>>>
>>>>>>>
>>>>>     client would
>>>>>
>>>>>
>>>>>>> ignore speech start events when it is interested only
>>>>>>>
>>>>>>>
>>>>>     in  DTMF start
>>>>>
>>>>>
>>>>>>> events). This _only_ works for the non-optimised case,
>>>>>>>
>>>>>>>
>>>>>     however. For
>>>>>
>>>>>
>>>>>>> the optimised case, assuming START-OF-SPEECH
>>>>>>>
>>>>>>>
>>>>>> coincides
>>>>>>
>>>>>>
>>>>>>> with the bargin signal to the speechsynth, prompts
>>>>>>>
>>>>>>>
>>>>>     will stop playing
>>>>>
>>>>>
>>>>>>> for inputs that the client might not be interested
>>>>>>>
>>>>>>>
>>>>>> in (e.g.
>>>>>>
>>>>>>
>>>>>>> a speech input will stop prompts playing even if the client
>>>>>>>
>>>>>>>
>>>>>> is only
>>>>>>
>>>>>>
>>>>>>> interested in DTMF).
>>>>>>>
>>>>>>>
>>>>>>>
>>>>>> OK, now I get it. There's a need for the client to both
>>>>>>
>>>>>>
>>>>>     handle the
>>>>>
>>>>>
>>>>>> non-optimized case itself, and influence or at least
>>>>>>
>>>>>>
>>>>>     have a clue what
>>>>>
>>>>>
>>>>>> the server is going to do in the optimized case.
>>>>>>
>>>>>>
>>>>>>
>>>>>>>
>>>>>>>
>>>>>>>
>>>>>>>>
>>>>>>>>
>>>>>>>>
>>>>>>>>
>>>>>>>>> 5. Currently, MRCPv2 has no way of indicating what
>>>>>>>>>
>>>>>>>>>
>>>>>     input type(s) a
>>>>>
>>>>>
>>>>>>>>> recogniser is listening for
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>
>>>>>>>> Do you mean exactly this, or do you mean "for the client to
>>>>>>>> indicate  to the resource what input types it should
>>>>>>>>
>>>>>>>>
>>>>>     look for"?
>>>>>
>>>>>
>>>>>>>>
>>>>>>>>
>>>>>>>>
>>>>>>>
>>>>>>> DB> Yes exactly - apologies for not being clear.
>>>>>>>
>>>>>>>
>>>>>>>
>>>>>>>
>>>>>>>>
>>>>>>>>
>>>>>>>>
>>>>>>>>
>>>>>>>>> + Why implement 5?
>>>>>>>>>
>>>>>>>>> i. Noisy case: Need DTMF-only recognition (and may
>>>>>>>>>
>>>>>>>>>
>>>>>     only have a
>>>>>
>>>>>
>>>>>>>>> speechrecog)
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>
>>>>>>>> I'm having difficulty following the logic of why
>>>>>>>>
>>>>>>>>
>>>>>     noise would
>>>>>
>>>>>
>>>>>>>> necessarily trigger START-OF-SPEECH if you were listening for
>>>>>>>> speech  but not DTMF. I suppose you can use a more forgiving
>>>>>>>> discriminator if  all you need to tell is if you're
>>>>>>>>
>>>>>>>>
>>>>>     getting DTMF,
>>>>>
>>>>>
>>>>>>>> but I've had a number  of real-world cases where wind
>>>>>>>>
>>>>>>>>
>>>>>     noise was
>>>>>
>>>>>
>>>>>>>> detected as DTMF, and  there's always the ambiguity
>>>>>>>>
>>>>>>>>
>>>>>     when you have
>>>>>
>>>>>
>>>>>>>> Captain Crunch on the  phone. In either case in the
>>>>>>>>
>>>>>>>>
>>>>>     non-optimized
>>>>>
>>>>>
>>>>>>>> case it's the client who  gets to decide whether an
>>>>>>>>
>>>>>>>>
>>>>>     event should be
>>>>>
>>>>>
>>>>>>>> interpreted as barge-in or  not, so it seems an
>>>>>>>>
>>>>>>>>
>>>>>     aesthetic protocol
>>>>>
>>>>>
>>>>>>>> design decision whether the  client tells the server
>>>>>>>>
>>>>>>>>
>>>>>     ahead of time
>>>>>
>>>>>
>>>>>>>> what circumstances to generate  the START-OF-SPEECH
>>>>>>>>
>>>>>>>>
>>>>>     event for, or
>>>>>
>>>>>
>>>>>>>> whether the event gets generated  and the client
>>>>>>>>
>>>>>>>>
>>>>>     decides based on
>>>>>
>>>>>
>>>>>>>> what's in the event whether it should  be
>>>>>>>>
>>>>>>>>
>>>>>> treated
>>>>>>
>>>>>>
>>>>>>>> as barge- in.
>>>>>>>>
>>>>>>>> I suppose one could make the argument that because
>>>>>>>>
>>>>>>>>
>>>>>     the spec implies
>>>>>
>>>>>
>>>>>>>> that the event can only be generated once per request
>>>>>>>>
>>>>>>>>
>>>>>     that if a
>>>>>
>>>>>
>>>>>>>> DTMF/ speech capable recognizer first hears
>>>>>>>>
>>>>>>>>
>>>>>> enough noise
>>>>>>
>>>>>>
>>>>>>>> to think it's  hearing speech and later hears DTMF,
>>>>>>>>
>>>>>>>>
>>>>>     the client will
>>>>>
>>>>>
>>>>>>>> declare barge-in  when the event comes and not when he DTMF
>>>>>>>> actually gets heard.
>>>>>>>>
>>>>>>>> If that's deemed a problem, we can still handle that
>>>>>>>>
>>>>>>>>
>>>>>     in the design
>>>>>
>>>>>
>>>>>>>> where the server just reports what it's hearing by allowing
>>>>>>>> multiple  events to be generated during a single request.
>>>>>>>>
>>>>>>>> Between the approach just outlined above, and an
>>>>>>>>
>>>>>>>>
>>>>>     approach where the
>>>>>
>>>>>
>>>>>>>> client provides a filter for whether to generate the
>>>>>>>>
>>>>>>>>
>>>>>     event or not,
>>>>>
>>>>>
>>>>>>>> I  have a mild preference (based on aesthetics rather
>>>>>>>>
>>>>>>>>
>>>>>     than some
>>>>>
>>>>>
>>>>>>>> hard  engineering tradeoff) for the approach where
>>>>>>>>
>>>>>>>>
>>>>>> the server
>>>>>>
>>>>>>
>>>>>>>> just reports  what it's hearing.
>>>>>>>>
>>>>>>>>
>>>>>>>> Having had some useful exchanges on this topic, it
>>>>>>>>
>>>>>>>>
>>>>>     also is becoming
>>>>>
>>>>>
>>>>>>>> apparent to me that this event is poorly named, and
>>>>>>>>
>>>>>>>>
>>>>>     we should
>>>>>
>>>>>
>>>>>>>> consider renaming it to "INTERESTING-INPUT-HEARD" or
>>>>>>>>
>>>>>>>>
>>>>>     something akin
>>>>>
>>>>>
>>>>>>>> to that, because as others have pointed out, a
>>>>>>>>
>>>>>>>>
>>>>>     DTMF-only recognizer
>>>>>
>>>>>
>>>>>>>> will never detect "start of speech".
>>>>>>>>
>>>>>>>> Another consideration to fold into the design choice is
>>>>>>>> extensibility. Bear with me through a little
>>>>>>>>
>>>>>>>>
>>>>>     gedankenexperiment.
>>>>>
>>>>>
>>>>>>>>
>>>>>>>> Suppose we want to define a new recognizer type,
>>>>>>>>
>>>>>>>>
>>>>>     which I'll call
>>>>>
>>>>>
>>>>>>>> the "name that tune" recognizer. The client plays
>>>>>>>>
>>>>>>>>
>>>>>     music to the
>>>>>
>>>>>
>>>>>>>> server and  the server recognizes musical notes. The
>>>>>>>>
>>>>>>>>
>>>>>     grammar is a
>>>>>
>>>>>
>>>>>>>> standard  musical notation, augmented with a semantic
>>>>>>>> interpretation that  transforms the notes into the
>>>>>>>>
>>>>>>>>
>>>>>     title of the
>>>>>
>>>>>
>>>>>>>> tune and provides that as  an answer.
>>>>>>>>
>>>>>>>> First, there's no speech involved (or is there...hang
>>>>>>>>
>>>>>>>>
>>>>>     on a minute).
>>>>>
>>>>>
>>>>>>>> Second, in order to accommodate the "name that tune"
>>>>>>>> recognizer, we'd have to extend both the client and
>>>>>>>>
>>>>>>>>
>>>>>     the server to
>>>>>
>>>>>
>>>>>>>> undetstand a  directive as to whether to recognize
>>>>>>>>
>>>>>>>>
>>>>>     music or now,
>>>>>
>>>>>
>>>>>>>> inaddition to what  the server already knows what to
>>>>>>>>
>>>>>>>>
>>>>>     do based on
>>>>>
>>>>>
>>>>>>>> the grammar. If you  follow my logic above, whether
>>>>>>>>
>>>>>>>>
>>>>>     or not we do
>>>>>
>>>>>
>>>>>>>> that, we have to extend  "start-of-speech" to say
>>>>>>>>
>>>>>>>>
>>>>>     "I'm hearing
>>>>>
>>>>>
>>>>>>>> music". So far fairly  straightforward, but let me
>>>>>>>>
>>>>>>>>
>>>>>     now throw in the
>>>>>
>>>>>
>>>>>>>> pathological twist.
>>>>>>>>
>>>>>>>> Suppose what I feed to a  combined music/speech
>>>>>>>>
>>>>>>>>
>>>>>     recognizer is a
>>>>>
>>>>>
>>>>>>>> work  in sprechstimme (spoken music), like the
>>>>>>>>
>>>>>>>>
>>>>>     "Geographical Fugue"
>>>>>
>>>>>
>>>>>>>> (aside:  this is a wonderful piece of music I highly
>>>>>>>>
>>>>>>>>
>>>>>     recommend to
>>>>>
>>>>>
>>>>>>>> anyone  interested in small ensemble singing). In
>>>>>>>>
>>>>>>>>
>>>>>     this case, the
>>>>>
>>>>>
>>>>>>>> tune could  be named by either doing speech or music
>>>>>>>>
>>>>>>>>
>>>>>     recognition.
>>>>>
>>>>>
>>>>>>>> Why is there  any need for the client to constrain
>>>>>>>>
>>>>>>>>
>>>>>     the server as to
>>>>>
>>>>>
>>>>>>>> which it tries  to do when it's
>>>>>>>>
>>>>>>>>
>>>>>> already
>>>>>>
>>>>>>
>>>>>>>> told the server what it wants through the  grammar?
>>>>>>>>
>>>>>>>> A few other comments below
>>>>>>>>
>>>>>>>>
>>>>>>>>
>>>>>>>>
>>>>>>>>> ii. Flexibility: Want speech-only recognition
>>>>>>>>>
>>>>>>>>>
>>>>>     (because a second
>>>>>
>>>>>
>>>>>>>>> recogniser is doing hotword on DTMF)
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>
>>>>>>>> I don't see how flexibility is affected by this deisgn
>>>>>>>>
>>>>>>>>
>>>>>> choice. If
>>>>>>
>>>>>>
>>>>>>>> that's what you want, feed the speech-only recognizer
>>>>>>>>
>>>>>>>>
>>>>>     a grammar
>>>>>
>>>>>
>>>>>>>> without any DTMF rules.
>>>>>>>>
>>>>>>>>
>>>>>>>>
>>>>>>>>
>>>>>>>>> + How to implement 5?
>>>>>>>>>
>>>>>>>>> a. Implicitly:
>>>>>>>>>    - dtmfrecog: always DTMF-only recognition
>>>>>>>>>    - speechrecog: depends on active grammar type
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>> if a dtmf grammar is active then DTMF input is "on"
>>>>>>>>>> if a speech grammar is active then speech
>>>>>>>>>>
>>>>>>>>>>
>>>>>     input is "on"
>>>>>
>>>>>
>>>>>>>>>
>>>>>>>>> b. Explicitly:
>>>>>>>>>    - Add inputmodes header to RECOGNIZE
>>>>>>>>>
>>>>>>>>> Option a is David's "do what I mean case"; option b is
>>>>>>>>>
>>>>>>>>>
>>>>>> the extra
>>>>>>
>>>>>>
>>>>>>>>> dial for the client.
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>
>>>>>>>> Actually, that's not the point I was making with "do
>>>>>>>>
>>>>>>>>
>>>>>     what I mean",
>>>>>
>>>>>
>>>>>>>> but it's not essential to the discussion so let's move on.
>>>>>>>>
>>>>>>>>
>>>>>>>>
>>>>>>>>
>>>>>>>>> It is worth noting that VoiceXML uses option b. This
>>>>>>>>>
>>>>>>>>>
>>>>>     allows one to
>>>>>
>>>>>
>>>>>>>>> activate both speech grammars and DTMF grammars (and
>>>>>>>>>
>>>>>>>>>
>>>>>> therefore
>>>>>>
>>>>>>
>>>>>>>>> be informed of any errors in the grammars at activation
>>>>>>>>>
>>>>>>>>>
>>>>>> time) but
>>>>>>
>>>>>>
>>>>>>>>> independently turn on whichever input mode you like e.g.
>>>>>>>>>
>>>>>>>>>
>>>>>> perhaps
>>>>>>
>>>>>>
>>>>>>>>> start with "both" then change to "dtmf".
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>
>>>>>>>> I'm not sure the VXML precedent is relevant here,
>>>>>>>>
>>>>>>>>
>>>>>     because the
>>>>>
>>>>>
>>>>>>>> application behind VXML is working a different part of the
>>>>>>>>
>>>>>>>>
>>>>>> problem
>>>>>>
>>>>>>
>>>>>>>> -  how to traverse a TUI dialog based on different
>>>>>>>>
>>>>>>>>
>>>>>     parts of the
>>>>>
>>>>>
>>>>>>>> input  space. In fact, I suspect that the VXML: choice was
>>>>>>>> conditioned more  by limitations at the time it was
>>>>>>>>
>>>>>>>>
>>>>>> specified than
>>>>>>
>>>>>>
>>>>>>>> an underlying good  design choice. Clearly having to
>>>>>>>>
>>>>>>>>
>>>>>     specify this
>>>>>
>>>>>
>>>>>>>> in VXML make the job of  handling a TUI with nodes
>>>>>>>>
>>>>>>>>
>>>>>     like "Say or
>>>>>
>>>>>
>>>>>>>> press 5" harder rather than  easier.
>>>>>>>>
>>>>>>>>
>>>>>>>>
>>>>>>>
>>>>>>> DB> The VoiceXML edge-case is pretty weird so it's not a major
>>>>>>> concern. My main concern is that the client can
>>>>>>>
>>>>>>>
>>>>>     indicate, somehow,
>>>>>
>>>>>
>>>>>>> what the input modes are.
>>>>>>>
>>>>>>>
>>>>>>>
>>>>>>>
>>>>>>>>
>>>>>>>> Summing up, while I don't feel strongly one way or
>>>>>>>>
>>>>>>>>
>>>>>     the other, I
>>>>>
>>>>>
>>>>>>>> have  a preference for handling this as follows:
>>>>>>>>
>>>>>>>> a) Rename "START-OF-SPEECH" to
>>>>>>>>
>>>>>>>>
>>>>>     "INTERESTING-INPUT-RECEIVED" or
>>>>>
>>>>>
>>>>>>>> something equivalent.
>>>>>>>> b) Include a parameter in the event saying what was
>>>>>>>>
>>>>>>>>
>>>>>     interesting
>>>>>
>>>>>
>>>>>>>> about  the input you received, with a registry of values
>>>>>>>>
>>>>>>>>
>>>>> which
>>>>>
>>>>>
>>>>>>>> includes:
>>>>>>>>     - signal above noise floor
>>>>>>>>     - speech
>>>>>>>>     - dtmf
>>>>>>>>     - (possibly) music
>>>>>>>> c) allow the event to be generated multiple times
>>>>>>>>
>>>>>>>>
>>>>>     during a request
>>>>>
>>>>>
>>>>>>>>
>>>>>>>>
>>>>>>>>
>>>>>>>>
>>>>>>>
>>>>>>> DB> I like these suggestions (START-OF-INPUT?).
>>>>>>>
>>>>>>>
>>>>>     However, I don't
>>>>>
>>>>>
>>>>>>> see how the problem of the optimised case is not
>>>>>>>
>>>>>>>
>>>>>     solved by them. I
>>>>>
>>>>>
>>>>>>> think the optimised case is fine if we have the
>>>>>>>
>>>>>>>
>>>>>     following rules:
>>>>>
>>>>>
>>>>>>>
>>>>>>>
>>>>>>>
>>>>>> Yes, I hadn't thought through the optimized case as
>>>>>>
>>>>>>
>>>>>     thoroughly as you.
>>>>>
>>>>>
>>>>>> Your suggested method name is fine by me as well
>>>>>>
>>>>>>
>>>>>>
>>>>>>> 1. START-OF-SPEECH (and optimised bargin) is only generated
>>>>>>>
>>>>>>>
>>>>>> for the
>>>>>>
>>>>>>
>>>>>>> input type that is being listened for 2. A speechrecog
>>>>>>>
>>>>>>>
>>>>>     listens for
>>>>>
>>>>>
>>>>>>> DTMF if DTMF grammars are active, speech if speech
>>>>>>>
>>>>>>>
>>>>>     grammars are
>>>>>
>>>>>
>>>>>>> active, or speech and DTMF if both grammar types are active.
>>>>>>>
>>>>>>>
>>>>>>>
>>>>>> Works for me.
>>>>>>
>>>>>>
>>>>>>
>>>>>>>
>>>>>>>
>>>>>>>
>>>>>>>> Note that all of the above I'm saying with my technical
>>>>>>>>
>>>>>>>>
>>>>>> hat on and
>>>>>>
>>>>>>
>>>>>>>> my chair hat off.
>>>>>>>> Putting my chair hat on for a moment, we really need
>>>>>>>>
>>>>>>>>
>>>>>     to get this
>>>>>
>>>>>
>>>>>>>> spec  to last call, so at some point Eric or I is going to
>>>>>>>>
>>>>>>>>
>>>>>> declare
>>>>>>
>>>>>>
>>>>>>>> rough  consensus so we can move on.
>>>>>>>>
>>>>>>>>
>>>>>>>>
>>>>>>>
>>>>>>> DB> Agreed!
>>>>>>>
>>>>>>>
>>>>>>>
>>>>>>>
>>>>>>>>
>>>>>>>> Dave Oran.
>>>>>>>>
>>>>>>>>
>>>>>>>>
>>>>>>>>> Dave
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>
>>>>>>>>> Sarvi makes a good point that adding the reason why the
>>>>>>>>>
>>>>>>>>>
>>>>>> START-OF-
>>>>>>
>>>>>>
>>>>>>>>> SPEECH occurred does not fix the optimised bargin case.
>>>>>>>>>
>>>>>>>>> dtmfrecog - listens for DTMF only
>>>>>>>>> speechrecog - listens for DTMF only, or speech only,
>>>>>>>>>
>>>>>>>>>
>>>>>     or speech &
>>>>>
>>>>>
>>>>>>>>> DTMF
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>
>>>>>>>>> ----- Original Message ----- From: "Shanmugham, Saravanan"
>>>>>>>>> <sarvi@cisco.com>
>>>>>>>>> To: "David R Oran" <oran@cisco.com>; "Dave Burke"
>>>>>>>>> <david.burke@voxpilot.com>
>>>>>>>>> Cc: <speechsc@ietf.org>; "Klaus Reifenrath"
>>>>>>>>> <Klaus.Reifenrath@Scansoft.com>
>>>>>>>>> Sent: Tuesday, July 05, 2005 8:51 PM
>>>>>>>>> Subject: RE: [Speechsc] START-OF-SPEECH in DTMF-only mode
>>>>>>>>>
>>>>>>>>>
>>>>>>>>> inline.
>>>>>>>>>
>>>>>>>>>     -----Original Message-----
>>>>>>>>>     From: David R Oran [mailto:oran@cisco.com]
>>>>>>>>>     Sent: Tuesday, July 05, 2005 11:08 AM
>>>>>>>>>     To: Dave Burke
>>>>>>>>>     Cc: Shanmugham, Saravanan; Klaus Reifenrath;
>>>>>>>>>
>>>>>>>>>
>>>>>     speechsc@ietf.org
>>>>>
>>>>>
>>>>>>>>>     Subject: Re: [Speechsc] START-OF-SPEECH in DTMF-only
>>>>>>>>>
>>>>>>>>>
>>>>> mode
>>>>>
>>>>>
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>     On Jul 5, 2005, at 12:54 PM, Dave Burke wrote:
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>> Inline.
>>>>>>>>>>
>>>>>>>>>> Dave
>>>>>>>>>>
>>>>>>>>>> ----- Original Message ----- From:
>>>>>>>>>>
>>>>>>>>>>
>>>>>     "Shanmugham, Saravanan"
>>>>>
>>>>>
>>>>>>>>>> <sarvi@cisco.com>
>>>>>>>>>> To: "David R Oran" <oran@cisco.com>; "Klaus
>>>>>>>>>>
>>>>>>>>>>
>>>>> Reifenrath"
>>>>>
>>>>>
>>>>>>>>>> <Klaus.Reifenrath@Scansoft.com>
>>>>>>>>>> Cc: <speechsc@ietf.org>
>>>>>>>>>> Sent: Tuesday, July 05, 2005 5:22 PM
>>>>>>>>>> Subject: RE: [Speechsc] START-OF-SPEECH in
>>>>>>>>>>
>>>>>>>>>>
>>>>>     DTMF-only mode
>>>>>
>>>>>
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>> I agree with Dave's analysis. The purpose of this
>>>>>>>>>>
>>>>>>>>>>
>>>>> event
>>>>>
>>>>>
>>>>>>>>>     was barge-in.
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>> And barge-in should happen for both DTMF and speech.
>>>>>>>>>>
>>>>>>>>>> Is there a case where you think it should not behave
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>     this way. If soe,
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>> please provide a scenario where you think
>>>>>>>>>>    1. Barge-in should happen for DTMF and not voice or
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>     vice-versa.
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>> DB> You want to do a DTMF recognition only
>>>>>>>>>>
>>>>>>>>>>
>>>>>     because it is
>>>>>
>>>>>
>>>>>>>>> noisy.
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>> While waiting for DTMF input, the speechrecog resource
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>     (or advanced
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>> dtmfrecog) generates a START-OF-SPEECH because it
>>>>>>>>>>
>>>>>>>>>>
>>>>> heard
>>>>>
>>>>>
>>>>>>>>>     some speech.
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>> The client does not want to stop prompt playing unless
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>     DTMF was heard
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>> but it can't tell by the START-OF-SPEECH whether
>>>>>>>>>>
>>>>>>>>>>
>>>>> speech
>>>>>
>>>>>
>>>>>>>>>     or DTMF was
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>> heard. Similarly vice versa.
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>     It's an interesting design question what part of the
>>>>>>>>>     policy resides at the client and what at the server, and
>>>>>>>>>     who makes the "final decision" about whether what was
>>>>>>>>>     heard was relevant to the control channel. Right now we
>>>>>>>>>     (IMO) have a weird partitioning in many cases where the
>>>>>>>>>     client basically says "do what I mean", but there are no
>>>>>>>>>     constraints of what the server actually does, and no
>>>>>>>>>     normalized basis for the client to figure out what to
>>>>>>>>>
>>>>>>>>>
>>>>> set
>>>>>
>>>>>
>>>>>>>>>     various magic numbers to (e.g. sensitivity).
>>>>>>>>>
>>>>>>>>>     In this case the only thing the client needs to
>>>>>>>>>
>>>>>>>>>
>>>>> decide is
>>>>>
>>>>>
>>>>>>>>>     whether to kill the prompt because the server thinks
>>>>>>>>>     something that would interfere with the feedback
>>>>>>>>>     ear/mouth/finger control happened. What this
>>>>>>>>>
>>>>>>>>>
>>>>>     says to me is
>>>>>
>>>>>
>>>>>>>>>     that it isn't necessarily a good idea for the client to
>>>>>>>>>     have more knobs to control the server
>>>>>>>>>
>>>>>>>>>
>>>>>     (especially if those
>>>>>
>>>>>
>>>>>>>>>     knows are just more value/policy input ungrounded in any
>>>>>>>>>     physics/ acoustics). On the other hand, having the
>>>>>>>>>
>>>>>>>>>
>>>>> server
>>>>>
>>>>>
>>>>>>>>>     tell the client more about what it thinks is going on is
>>>>>>>>>     probably valuable.
>>>>>>>>>
>>>>>>>>>     So, Coming to the point after this long rambling
>>>>>>>>>     introduction, I think it would in fact be useful for the
>>>>>>>>>     START-Of-SPEECH event to indicate some extra
>>>>>>>>>
>>>>>>>>>
>>>>> information,
>>>>>
>>>>>
>>>>>>>>>     for example:
>>>>>>>>>     a) I got something enough above the noise floor
>>>>>>>>>
>>>>>>>>>
>>>>>     to qualify
>>>>>
>>>>>
>>>>>>>>>     for exceeding the "Sensisitvity" parameter you sent
>>>>>>>>>
>>>>>>>>>
>>>>> in on
>>>>>
>>>>>
>>>>>>>>>     the request but I really can't tell what it is
>>>>>>>>>
>>>>>>>>>
>>>>>     (could be a
>>>>>
>>>>>
>>>>>>>>>     hippopatmus fart, or a siren in the background,
>>>>>>>>>
>>>>>>>>>
>>>>>     or captain
>>>>>
>>>>>
>>>>>>>>>     crunch trying to whistle DTMF).
>>>>>>>>>     b) I think I'm hearing speech
>>>>>>>>>     c) I think I'm hearing DTMF
>>>>>>>>>
>>>>>>>>> Though I agree with your former part of your
>>>>>>>>>
>>>>>>>>>
>>>>>     response. I am not
>>>>>
>>>>>
>>>>>>>>> sure I agree with your proposed solution.
>>>>>>>>> The way I see this problem is that, it is more of what
>>>>>>>>>
>>>>>>>>>
>>>>>> constitues a
>>>>>>
>>>>>>
>>>>>>>>> barge-in event. This boils down to whether it is
>>>>>>>>>
>>>>>>>>>
>>>>>     speech, DTMF or
>>>>>
>>>>>
>>>>>>>>> both.
>>>>>>>>> This is inturn boils down to what type of recognizer
>>>>>>>>>
>>>>>>>>>
>>>>>> resource we are
>>>>>>
>>>>>>
>>>>>>>>> using, dtmf-recog, speech-recog, and speech-only-recog(we
>>>>>>>>>
>>>>>>>>>
>>>>>> don't have
>>>>>>
>>>>>>
>>>>>>>>> this and I don't think we should add it, but think
>>>>>>>>>
>>>>>>>>>
>>>>>     of this as a
>>>>>
>>>>>
>>>>>>>>> place holder that explains the concept).
>>>>>>>>>
>>>>>>>>> A client knowing what type of barge-in happenned, does
>>>>>>>>>
>>>>>>>>>
>>>>>> not impact
>>>>>>
>>>>>>
>>>>>>>>> the
>>>>>>>>> barge-in operation itself as it may be too late(for
>>>>>>>>>
>>>>>>>>>
>>>>>     the optimized
>>>>>
>>>>>
>>>>>>>>> barge-in case). It may have other use cases, and if we
>>>>>>>>>
>>>>>>>>>
>>>>>> can identify
>>>>>>
>>>>>>
>>>>>>>>> them, I don't mind adding support for the
>>>>>>>>>
>>>>>>>>>
>>>>>     START-OF-SPEECH event to
>>>>>
>>>>>
>>>>>>>>> say what type of barge-in happenned. But that itself does
>>>>>>>>>
>>>>>>>>>
>>>>> not
>>>>>
>>>>>
>>>>>> solve the
>>>>>>
>>>>>>
>>>>>>>>> original problem raised. Refer to my previous response.
>>>>>>>>>
>>>>>>>>> The solution lies in defining what what is a
>>>>>>>>>
>>>>>>>>>
>>>>>     barge-in event.
>>>>>
>>>>>
>>>>>>>>> That  boils
>>>>>>>>> down to what type of recognition is happenning,
>>>>>>>>>
>>>>>>>>>
>>>>>> dtmf-only, speech-
>>>>>>
>>>>>>
>>>>>>>>> dtmf
>>>>>>>>> or speech-only. We do not support speech-only as a
>>>>>>>>>
>>>>>>>>>
>>>>>     resource today,
>>>>>
>>>>>
>>>>>>>>> the question is do we need a header to force it.
>>>>>>>>>
>>>>>>>>> Sarvi
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>>    2. You would benefit from the client knowing
>>>>>>>>>>
>>>>>>>>>>
>>>>>> what caused
>>>>>>
>>>>>>
>>>>>>>>> the
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>> barge-in, DTMF Vs speech.
>>>>>>>>>>
>>>>>>>>>> DB> See previous comment. And previous e-mail:
>>>>>>>>>>
>>>>>>>>>>
>>>>>     either add an
>>>>>
>>>>>
>>>>>>>>>> inputmodes header (taking value speech, dtmf, both) to
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>     the RECOGNIZE
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>> request or add a header to the START-OF-SPEECH event
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>     indicating DTMF
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>> or speech.
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>     I'm leaning in your direction on this latter point - as
>>>>>>>>>     should be evident from what I wrote above.
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>> Sarvi
>>>>>>>>>>
>>>>>>>>>>     -----Original Message-----
>>>>>>>>>>     From: speechsc-bounces@ietf.org
>>>>>>>>>>     [mailto:speechsc-bounces@ietf.org] On Behalf Of
>>>>>>>>>>
>>>>>>>>>>
>>>>>> David R
>>>>>>
>>>>>>
>>>>>>>>> Oran
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>>     Sent: Tuesday, July 05, 2005 5:24 AM
>>>>>>>>>>     To: Klaus Reifenrath
>>>>>>>>>>     Cc: 'speechsc@ietf.org'
>>>>>>>>>>     Subject: Re: [Speechsc] START-OF-SPEECH in
>>>>>>>>>>
>>>>>>>>>>
>>>>>> DTMF-only mode
>>>>>>
>>>>>>
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>>     On Jul 5, 2005, at 3:46 AM, Reifenrath,
>>>>>>>>>>
>>>>>>>>>>
>>>>>     Klaus wrote:
>>>>>
>>>>>
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>>> The current spec is not clear when
>>>>>>>>>>>
>>>>>>>>>>>
>>>>>> START-OF-SPEECH need
>>>>>>
>>>>>>
>>>>>>>>>>     to be send in
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>>> the following scenarios:
>>>>>>>>>>> A) The client requested a DTMF Recognizer. Is
>>>>>>>>>>>
>>>>>>>>>>>
>>>>> the
>>>>>
>>>>>
>>>>>>>>>>     START-OF-SPEECH
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>>> event send to the client also if speech
>>>>>>>>>>>
>>>>>>>>>>>
>>>>>     was detected?
>>>>>
>>>>>
>>>>>>>>>>     I suspect so, since one of the prime purposes
>>>>>>>>>>
>>>>>>>>>>
>>>>>> is to enable
>>>>>>
>>>>>>
>>>>>>>>>>     client- mediated barge-in handling. However, if
>>>>>>>>>>
>>>>>>>>>>
>>>>> the
>>>>>
>>>>>
>>>>>>>>>>     recognizer is in fact only capable of
>>>>>>>>>>
>>>>>>>>>>
>>>>>     recognizing DTMF
>>>>>
>>>>>
>>>>>>>>>>     then it may in fact not report anythin
>>>>>>>>>>
>>>>>>>>>>
>>>>>     unless it's using
>>>>>
>>>>>
>>>>>>>>>>     some primitive thresholding machinery, like a SN
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>> threshold.
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>>> B) The client requested a Speech
>>>>>>>>>>>
>>>>>>>>>>>
>>>>>     Recognizer, but only
>>>>>
>>>>>
>>>>>>>>>>     activated DTMF
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>>> grammars. Is the START-OF-SPEECH event
>>>>>>>>>>>
>>>>>>>>>>>
>>>>>     send to the
>>>>>
>>>>>
>>>>>>>>>>     client also if
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>>> speech was detected?
>>>>>>>>>>>
>>>>>>>>>>>
>>>>>>>>>>     Again, I'd say yes, for the same reason as above.
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>>> I think in both cases START-OF-SPEECH should
>>>>>>>>>>>
>>>>>>>>>>>
>>>>> only
>>>>>
>>>>>
>>>>>>>>>     be send after
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>>> detecting a DTMF digit (see Figure 12 of
>>>>>>>>>>>
>>>>>>>>>>>
>>>>> VoiceXML
>>>>>
>>>>>
>>>>>>>>>     2.0: http://
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>>> www.w3.org/TR/voicexml20/#dmlATiming).
>>>>>>>>>>>
>>>>>>>>>>>
>>>>>>>>>>     We seem to have reached different
>>>>>>>>>>
>>>>>>>>>>
>>>>>     conclusions. I'd be
>>>>>
>>>>>
>>>>>>>>>>     interested in why you think my analysis
>>>>>>>>>>
>>>>>>>>>>
>>>>>     above is wrong.
>>>>>
>>>>>
>>>>>>>>>>
>>>>>>>>>>     Dave.
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>>> Klaus
>>>>>>>>>>>
>>>>>>>>>>> _______________________________________________
>>>>>>>>>>> Speechsc mailing list
>>>>>>>>>>> Speechsc@ietf.org
>>>>>>>>>>> https://www1.ietf.org/mailman/listinfo/speechsc
>>>>>>>>>>>
>>>>>>>>>>>
>>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>>     _______________________________________________
>>>>>>>>>>     Speechsc mailing list
>>>>>>>>>>     Speechsc@ietf.org
>>>>>>>>>>     https://www1.ietf.org/mailman/listinfo/speechsc
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>> _______________________________________________
>>>>>>>>>> Speechsc mailing list
>>>>>>>>>> Speechsc@ietf.org
>>>>>>>>>> https://www1.ietf.org/mailman/listinfo/speechsc
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>
>>>>>>>>>
>>>>>>>>> _______________________________________________
>>>>>>>>> Speechsc mailing list
>>>>>>>>> Speechsc@ietf.org
>>>>>>>>> https://www1.ietf.org/mailman/listinfo/speechsc
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>
>>>>>>>>
>>>>>>>> _______________________________________________
>>>>>>>> Speechsc mailing list
>>>>>>>> Speechsc@ietf.org
>>>>>>>> https://www1.ietf.org/mailman/listinfo/speechsc
>>>>>>>>
>>>>>>>>
>>>>>>>
>>>>>>>
>>>>>>>
>>>>>>
>>>>>> _______________________________________________
>>>>>> Speechsc mailing list
>>>>>> Speechsc@ietf.org
>>>>>> https://www1.ietf.org/mailman/listinfo/speechsc
>>>>>>
>>>>>>
>>>>>>
>>>>>
>>>>>
>>>>>     _______________________________________________
>>>>>     Speechsc mailing list
>>>>>     Speechsc@ietf.org
>>>>>     https://www1.ietf.org/mailman/listinfo/speechsc
>>>>>
>>>>>
>>>>> _______________________________________________
>>>>> Speechsc mailing list
>>>>> Speechsc@ietf.org
>>>>> https://www1.ietf.org/mailman/listinfo/speechsc
>>>>>
>>>>>
>>>>> _______________________________________________
>>>>> Speechsc mailing list
>>>>> Speechsc@ietf.org
>>>>> https://www1.ietf.org/mailman/listinfo/speechsc
>>>>>
>>>>>
>>>>>
>>>>
>>>> _______________________________________________
>>>> Speechsc mailing list
>>>> Speechsc@ietf.org
>>>> https://www1.ietf.org/mailman/listinfo/speechsc
>>>>
>>>>
>>>>
>>>>
>>>>
>>>
>>> _______________________________________________
>>> Speechsc mailing list
>>> Speechsc@ietf.org
>>> https://www1.ietf.org/mailman/listinfo/speechsc
>>
> 


_______________________________________________
Speechsc mailing list
Speechsc@ietf.org
https://www1.ietf.org/mailman/listinfo/speechsc



From speechsc-bounces@ietf.org Wed Jul 13 12:55:05 2005
Received: from localhost.localdomain ([127.0.0.1] helo=megatron.ietf.org)
	by megatron.ietf.org with esmtp (Exim 4.32)
	id 1DskVl-0007MK-3M; Wed, 13 Jul 2005 12:55:05 -0400
Received: from odin.ietf.org ([132.151.1.176] helo=ietf.org)
	by megatron.ietf.org with esmtp (Exim 4.32) id 1DskVh-0007Hr-T3
	for speechsc@megatron.ietf.org; Wed, 13 Jul 2005 12:55:03 -0400
Received: from ietf-mx.ietf.org (ietf-mx [132.151.6.1])
	by ietf.org (8.9.1a/8.9.1a) with ESMTP id MAA02977
	for <speechsc@ietf.org>; Wed, 13 Jul 2005 12:54:58 -0400 (EDT)
Received: from sj-iport-5.cisco.com ([171.68.10.87])
	by ietf-mx.ietf.org with esmtp (Exim 4.43) id 1Dsky5-00030A-Nf
	for speechsc@ietf.org; Wed, 13 Jul 2005 13:24:25 -0400
Received: from sj-core-2.cisco.com (171.71.177.254)
	by sj-iport-5.cisco.com with ESMTP; 13 Jul 2005 09:54:47 -0700
X-IronPort-AV: i="3.93,287,1115017200"; 
	d="scan'208"; a="198074652:sNHT51492284"
Received: from imail.cisco.com (imail.cisco.com [128.107.200.91])
	by sj-core-2.cisco.com (8.12.10/8.12.6) with ESMTP id j6DGshod017407;
	Wed, 13 Jul 2005 09:54:43 -0700 (PDT)
Received: from [10.32.245.152] (stealth-10-32-245-152.cisco.com
	[10.32.245.152])
	by imail.cisco.com (8.12.11/8.12.10) with SMTP id j6DGrLDt032337;
	Wed, 13 Jul 2005 09:53:22 -0700
In-Reply-To: <020a01c587c4$f2eb3dd0$6901a8c0@db01.voxpilot.com>
References: <7DE7C4EF3B7C8B4B82955191378290D802ED4176@mtb1exch01.nuance.com>
	<C7673C4C-0624-47A1-B1DF-115FB284785A@cisco.com>
	<028f01c5865e$7f92c4a0$038ae9d5@db01.voxpilot.com>
	<9B4BE6A2-82ED-4C8B-BAE5-006FB0C2DBBB@cisco.com>
	<020a01c587c4$f2eb3dd0$6901a8c0@db01.voxpilot.com>
Mime-Version: 1.0 (Apple Message framework v733)
X-Priority: 3
Content-Type: text/plain; charset=US-ASCII; delsp=yes; format=flowed
Message-Id: <B339DA52-6994-4E58-9DAC-9987FDCA23CD@cisco.com>
Content-Transfer-Encoding: 7bit
From: David R Oran <oran@cisco.com>
Subject: Re: [Speechsc] START-OF-SPEECH in DTMF-only mode -> summary(?)
Date: Wed, 13 Jul 2005 12:54:40 -0400
To: "Dave Burke" <david.burke@voxpilot.com>
X-Mailer: Apple Mail (2.733)
IIM-SIG: v:"1.1"; h:"imail.cisco.com"; d:"cisco.com"; z:"home"; m:"krs";
	t:"1121273605.899061"; x:"432200"; a:"rsa-sha1"; b:"nofws:44447";
	e:"Iw=="; n:"sQYarK2E51MdcTiUqeif3F7cWdxIfoCiXhdfb9vD5ee/j0jXL15gbFxF2p"
	"XIweAblu0N6XAgK7k+wrbr7bQDJaCDqOmzqpRUBjIRQAXQ7NzadpmR3pUL6wxaRUtW+c43sl9jC"
	"50Qg1sXHpPjt8Y+Y16ioyQAQAdSunM4YhevURc=";
	s:"rgt41Y5VFia3OmnOFJOLoo1JZzt4ZdQVG38cFzx3ndjirwNradmnlanUvTO/8i0TdfKyU34t"
	"9OIGMcAxMnZvtlvJ9fd2syjFLNz5wld1vetA3mjWT5f9Zss5BTCFIoyAtxtHeAKz/3RFonaxXFo"
	"rdZDu3FANY7NKPjyXkagRpKs="; c:"From: David R Oran <oran@cisco.com>";
	c:"Subject: Re: [Speechsc] START-OF-SPEECH in DTMF-only mode ->
	summary" "(?)"; c:"Date: Wed, 13 Jul 2005 12:54:40 -0400"
IIM-VERIFY: s:"y"; v:"y"; r:"60"; h:"imail.cisco.com";
	c:"message from imail.cisco.com verified; "
X-Spam-Score: 0.0 (/)
X-Scan-Signature: 02b0ecec2de8b7bc26112405dc3efaf0
Content-Transfer-Encoding: 7bit
Cc: speechsc@ietf.org, Pierre Forgues <forgues@nuance.com>,
	Eric Burger <eburger@brooktrout.com>, "Shanmugham,
	Saravanan" <sarvi@cisco.com>
X-BeenThere: speechsc@ietf.org
X-Mailman-Version: 2.1.5
Precedence: list
List-Id: Speech Services Control Working Group <speechsc.ietf.org>
List-Unsubscribe: <https://www1.ietf.org/mailman/listinfo/speechsc>,
	<mailto:speechsc-request@ietf.org?subject=unsubscribe>
List-Post: <mailto:speechsc@ietf.org>
List-Help: <mailto:speechsc-request@ietf.org?subject=help>
List-Subscribe: <https://www1.ietf.org/mailman/listinfo/speechsc>,
	<mailto:speechsc-request@ietf.org?subject=subscribe>
Sender: speechsc-bounces@ietf.org
Errors-To: speechsc-bounces@ietf.org


On Jul 13, 2005, at 12:07 PM, Dave Burke wrote:

> Inline
>
> Dave
>
> ----- Original Message ----- From: "David R Oran" <oran@cisco.com>
> To: "Dave Burke" <david.burke@voxpilot.com>
> Cc: "Pierre Forgues" <forgues@nuance.com>; <speechsc@ietf.org>;  
> "Shanmugham, Saravanan" <sarvi@cisco.com>; "Eric Burger"  
> <eburger@brooktrout.com>
> Sent: Wednesday, July 13, 2005 2:20 PM
> Subject: Re: [Speechsc] START-OF-SPEECH in DTMF-only mode -> summary 
> (?)
>
>
>
>>
>> On Jul 11, 2005, at 5:21 PM, Dave Burke wrote:
>>
>>
>>> I don't think it is possible in practice to separate what input   
>>> type is considered potential barge-in and what input type is  
>>> being  recognised because a recognition hypothesis is always  
>>> generated  when a barge-in occurs.
>>>
>> True, but what I'm trying to tease apart is whether to declare  
>> barge- in from the point of view of stopping output from simply  
>> detecting  input for the purposes of cranking up the non-signal- 
>> processing parts  of the recognition engine. Did you not find my  
>> "mark the phrases"  application example compelling?
>>
>
> DB> Indeed it's a nice application (albeit producing disturbing  
> memories of musical theory exams from my distant childhood!)...  
> Can't this application be created by setting Kill-On-Barge-In to  
> false on the SPEAK request that queued the music?
>
Ah, true if the output was from a SPEAK request. Not sure if we need  
to accommodate other cases.

>
>>
>>
>>> For example, imagine a client invoked a speechrecog and asked it  
>>> to  only consider DTMF for barge-in. Once speech is detected,  
>>> although  no barge-in would happen, a RECOGNITION-COMPLETE will  
>>> result after Speech-Incomplete-Timeout milliseconds of silence  
>>> thus ending the recognition state machine with a 'nomatch'.
>>>
>> Sure, but of course the timeout could be very long... I must be   
>> dense, because I see this relationship as tenuous rather than   
>> strongly coupled.
>>
>
> DB> But the recognition will terminate with nomatch at some point  
> if it hears noise/speech first regardless of a good DTMF sequence  
> (I say first only out of pragmatism because that's what modern day  
> recognizers do - the first input type recognized results in turning  
> off the other input type for the remainder of the recognition). The  
> recognition will need to be restarted but then some DTMF is lost,  
> timers have to be reset etc.
>
Yes, it's a tradeoff between pragmatic stuff and clean aesthetics and  
abstraction. I tend to favor the latter, all things being equal  
(which they rarely are).
>
>>
>>
>>> The converse example applies for a speech-only recognition (DTMF-  
>>> Interdigit-Timeout replaces the Speech-Incomplete-Timeout).
>>>
>>> Although I prefer the idea of an InputModes header in RECOGNIZE   
>>> (it's the cleanest solution for VoiceXML implementors), I am   
>>> willing to compromise on using the grammar type to determine the   
>>> input types (in what follows I use input types to mean both what  
>>> is  considered potential barge-in and what type of input is to  
>>> be  recognised). The main issue with this approach is  
>>> incompatibility  with VoiceXML (recall VoiceXML grammar  
>>> activation is independent of  what inputmodes are set).
>>>
>>>
>> And voice XML does not have any explicit control over barge-in   
>> either. We could of course back off not have clients involved at  
>> all  in barge-in processing, but I think that would be a mistake.
>>
>
> DB> I'm not sure I understand this (unless you mean VoiceXML  
> doesn't decouple barge-in input type from recognition input type in  
> which case I agree).
Yes, that's what I meant.

> VoiceXML allows you to switch on/off barge-in by setting the barge- 
> in attribute on <prompt> (maps to MRCP's Kill-On-Barge-In). It  
> allows you to simultaneously set the input type for recognition and  
> barge-in. VoiceXML does have a bargeintype attribute although this  
> is just to turn on hotword...
>
Which seems another mixing of fruit in the fruit salad :-)

>
>>
>>
>>> This incompatibility results in side-effects that, in my  
>>> opinion,  are inconsequential for real applications:
>>>    a. speech barge-in (noise) and nomatch will not happen for  
>>> inputmodes="both" when no speech grammars are activated
>>>    b. dtmf barge-in and nomatch will not happen for   
>>> inputmodes="both" when no dtmf grammars are activated
>>>    c. grammars with errors will not be detected if the grammar  
>>> type  is not compatible with the inputmodes property
>>> A workaround for the really conscientious VoiceXML platform is to  
>>> activate a dummy grammar (maybe with the NULL rule) of type  
>>> speech (dtmf) when inputmodes="both" is set but no speech (dtmf)  
>>> grammar  is activated in the application.
>>>
>>>
>> I agree with this assessment. I'll also point out that the   
>> flexibility will allow a class of VXML applications that can't be   
>> done today (e.g. ones that don't stop prompts on hearing input.
>>
>
> DB> This is possible already with Kill-On-Barge-On false
>
Ok, I concede the point. You're right.
>
>>
>>
>>> -----------------
>>>
>>> Summarising the proposed changes again  (refined a little and to   
>>> avoid trawling though this massive thread!):
>>>
>>> 1. For a speechrecog resource, the types of the activated  
>>> grammars determine the input types the recogniser processes and  
>>> considers  for potential barge-in.
>>>
>>>
>> Good. Agree.
>>
>>
>>> 2. Clarify a dtmfrecog only processes DTMF and hence can only   
>>> generate barge-in / START-OF-SPEECH for DTMF inputs
>>>
>>>
>> I'm ok with this even though I'd rather keep input detection and  
>> recognition behavior orthogonal to accommodate smarter server  
>> implementations and to give better consistency between the server-  
>> optimized and client-intermediated cases. In particular, I do  
>> really  feel like we need a way for the client to tell a server to  
>> NOT try to  outguess it by doing optimized barge in processing  
>> since the  application might in fact not want the output stopped.
>>
>>
>>
>>> Nice-to-have:
>>>
>>> 3. Change START-OF-SPEECH to START-OF-INPUT
>>>
>>>
>> I'm in favor of this.
>>
>>
>>> 4. Add a header of InputType to START-OF-INPUT. Current  
>>> specified  values are "dtmf" or "speech".
>>>
>>>
>> I think we really need this, and would press again that we make  
>> this  a list of values and future-proof it with the other values  
>> we suspect  will be quite useful (e.g. music, gesture).
>>
>> Dave. (chair hat sort-of-on since we're trying to get final  
>> consensus here).
>>
>>
>>
>>> -----------------
>>>
>>> Dave
>>>
>>> ----- Original Message ----- From: "David R Oran" <oran@cisco.com>
>>> To: "Pierre Forgues" <forgues@nuance.com>
>>> Cc: <speechsc@ietf.org>; "Shanmugham, Saravanan"  
>>> <sarvi@cisco.com>; "Eric Burger" <eburger@brooktrout.com>; "Dave  
>>> Burke" <david.burke@voxpilot.com>
>>> Sent: Monday, July 11, 2005 2:15 PM
>>> Subject: Re: [Speechsc] START-OF-SPEECH in DTMF-only mode ->  
>>> summary (?)
>>>
>>>
>>>
>>>
>>>>
>>>> On Jul 8, 2005, at 11:36 AM, Pierre Forgues wrote:
>>>>
>>>>
>>>>
>>>>>
>>>>>
>>>>> -----Original Message-----
>>>>> From: speechsc-bounces@ietf.org [mailto:speechsc-  
>>>>> bounces@ietf.org] On
>>>>> Behalf Of David R Oran
>>>>> Sent: Friday, July 08, 2005 10:05 AM
>>>>> To: Dave Burke
>>>>> Cc: speechsc@ietf.org; Shanmugham, Saravanan; Eric Burger
>>>>> Subject: Re: [Speechsc] START-OF-SPEECH in DTMF-only mode ->   
>>>>> summary (?)
>>>>>
>>>>>
>>>>> On Jul 8, 2005, at 6:59 AM, Dave Burke wrote:
>>>>>
>>>>>
>>>>>
>>>>>
>>>>>> I agree with your opinion on what constitutes a barge-in.  
>>>>>> Based on
>>>>>> Klaus' e-mail, I am more concerned that using the grammar type to
>>>>>> decide what constitutes a barge-in is going to make VoiceXML
>>>>>> implementations difficult.
>>>>>>
>>>>>> It is easy to map VoiceXML application selected inputmodes to  
>>>>>> MRCP
>>>>>> resources:
>>>>>> a. inputmodes="dtmf" -> Use a dtmfrecog
>>>>>> b. inputmodes="speech" -> Use a "speech-only-recog"
>>>>>> c. inputmodes="both" -> Use a speechrecog (or a combination of
>>>>>> "speech-only-recog" + dtmfrecog)
>>>>>>
>>>>>> The only problem is (b). While not as important as being able  
>>>>>> to do
>>>>>> DTMF-only recognition, I believe we DO need to support speech- 
>>>>>> only
>>>>>> recongition so as (a) to avoid unnecessary limitations in VUI
>>>>>> design, and (b) to facilitate VoiceXML implementations. I  
>>>>>> think we
>>>>>> should add a header to RECOGNIZE so the client is able to always
>>>>>> _explicitly_ set the inputmodes.
>>>>>>
>>>>>>
>>>>>>
>>>>>>
>>>>> I'm not sure I buy this, since what the recognizer is looking  
>>>>> for  and
>>>>> what constitutes an input that should be considered a  
>>>>> potential  barge-
>>>>> in strike me as independent. I'm similarly not persuaded that  
>>>>> we  need
>>>>> the flexibility to set the input mode independently of the  
>>>>> grammar,
>>>>> since it leads to all sorts of inconsistent states (e.g. speech- 
>>>>> only
>>>>> grammar with an input-mode of dtmf). On the other hand I can  
>>>>> see the
>>>>> need for the client to specify what sorts of input out to generate
>>>>> the start-of-input (nee start-of-speech) event on.
>>>>>
>>>>> Pmf> The mechanism for a client to specify the type of input  
>>>>> is  using
>>>>> grammars.  These can have DTMF and/or speech requirements.  If   
>>>>> you add
>>>>> the complexity of an independent header to specify the input  
>>>>> mode  then
>>>>> you will need to document the behavior for inconsistencies.
>>>>>
>>>>>
>>>>>
>>>> Right. I generally agree with Pierre, which is why I proposed a  
>>>> compromise position where the client can use a header to  
>>>> specify  what types of input should be considered the start of  
>>>> something   potentially interesting from a barge-in point of  
>>>> view, as opposed  to  a directive to the server for what to  
>>>> recognize.
>>>>
>>>> I don't think this is crucial functionality, but it does make  
>>>> for   more consistent behavior between the optimized and non- 
>>>> optimized   barge-in cases. In fact, I can see cases where the  
>>>> client doesn't   want any barge-in processing to happen. Imagine  
>>>> an application  called  "mark the musical phrases", where the  
>>>> client listens to  playout of a  musical score, and hits various  
>>>> DTMF buttons to  indicate markers  (e.g. end-of-phrase, end-of- 
>>>> theme, key-change)  while listening. For  this application you  
>>>> want to make sure that  barge-in doesn't stop the music!
>>>>
>>>>
>>>>
>>>>>
>>>>>
>>>>>
>>>>>> In retrospect, I don't like the idea of multiple START-OF-SPEECH
>>>>>> events being generated from the same media resource. This is
>>>>>> because one assumes that the START-OF-SPEECH should be of the  
>>>>>> same
>>>>>> type as the hypothesis returned in the RECOGNITION-COMPETE  
>>>>>> message
>>>>>> - most implementations, on hearing one input mode type disable  
>>>>>> the
>>>>>> recogniser of the other type.
>>>>>>
>>>>>>
>>>>>>
>>>>> I'm not sure I follow this logic, but I'm not wedded to the  
>>>>> idea of
>>>>> allowing multiple events. It was trying to solve the problem of
>>>>> ambiguity around what the recognizer was hearing and the idea that
>>>>> the client might care. If you don't think the client will ever  
>>>>> care,
>>>>> then we don't need the capability.
>>>>> Pmf> I agree we should not have multiple SOS events.  Too much  
>>>>> chatter.
>>>>>
>>>>>
>>>>>
>>>> I'm not concerned about the "chatter" since this is a low-  
>>>> bandwidth control channel even with multiple events, but as I  
>>>> said  it's a small point and I'm happy to concede.
>>>>
>>>>
>>>>
>>>>>
>>>>>
>>>>>
>>>>>
>>>>>> I do like David's idea of renaming START-OF-SPEECH to something
>>>>>> like START-OF-INPUT and carrying a type header because it is  
>>>>>> neater
>>>>>> and more extensible.
>>>>>>
>>>>>> So in summary, I propose we modify the spec to:
>>>>>>
>>>>>> 1. Clarify what constitutes a barge-in for a dtmfrecog and
>>>>>> speechrecog as per Sarvi's e-mail (and in agreement with  
>>>>>> Klaus' for
>>>>>> dtmfrecog).
>>>>>>
>>>>>>
>>>>>>
>>>>>>
>>>>> Hmmm, ok, but I think the issue is actually clarifying what the
>>>>> recognizer declares as "interesting input", which it may decide  
>>>>> also
>>>>> constitutes a barge-in in the optimized case, and in either case
>>>>> reports to the client that it heard.
>>>>> Pmf> I'm not going to argue changing the name of the event, but my
>>>>> preference would be to keep existing method names unless there  
>>>>> is a
>>>>> clear reason for changing - which I do not see here.
>>>>>
>>>>>
>>>>>
>>>> I do see a clear reason, since the thing that you see the start   
>>>> of  may not be speech. I like (I think it was Dave's  
>>>> suggestion)  "start- of-input:" since it mirrors the other  
>>>> method names.
>>>>
>>>>
>>>>>
>>>>>
>>>>>
>>>>>> 2. Specifiy an InputModes header to RECOGNIZE (defaults to  
>>>>>> "both",
>>>>>> can also be "speech" or "DTMF"). Setting to speech for a
>>>>>> speechrecog results in the hypothesised "speech-only-recog".
>>>>>> Setting to DTMF for a speechrecog is equivalent to using a
>>>>>> dtmfrecog. Edge cases: Setting to speech for a dtmfrecog will
>>>>>> result in a noinput as will setting to DTMF for a speechrecog  
>>>>>> which
>>>>>> does not support DTMF.
>>>>>>
>>>>>>
>>>>>>
>>>>>>
>>>>> I'm ok with having a header for the client to tell the server   
>>>>> when it
>>>>> would like the start-of-input event to be generated and what the
>>>>> client considers to be the "interesting input" that the recognizer
>>>>> should use to do the discrimination, and possibly do barge-in
>>>>> processing in the optimized case.
>>>>>
>>>>> I'm less ok with the proposed domain of values, since it will have
>>>>> extensibility problems. Especially problematical is have a  
>>>>> code  point
>>>>> of "both" since that will be ambiguous if we even define a "music"
>>>>> recognizer, or a "gesture" recognizer using video input. Here's my
>>>>> counter-proposal:
>>>>>
>>>>> Create a header on the recognizer resource requests called "start-
>>>>> input-on:". Define a registry of values, with the following three
>>>>> values initially defined:
>>>>>      - dtmf
>>>>>      - speech
>>>>>      - music
>>>>>
>>>>> The semantics would be that the resource is to generate a  
>>>>> "start-of-
>>>>> input" event, apply the defined grammars to what follows, and  
>>>>> do any
>>>>> local optimized barge-in processing if ANY of the enumerated input
>>>>> types is detected. That way you can say
>>>>>      "start-input-on: dtmf" if you want to just  dtmf,
>>>>>      "start-input-on: speech" if you want just speech (dtmf input
>>>>> would be ignored even if the grammar supported dtmf)
>>>>>      "start-input- on: speech, dtmf" if you wanted both
>>>>> etc.
>>>>>
>>>>> Pmf> Doesn't this come back to the previous point of having two
>>>>> independent ways to specify input modes?  Through grammars and  
>>>>> this
>>>>> proposed new header?  If you add this then you will have to  
>>>>> document
>>>>> behavior on inconsistencies.  Anyway, if you proceed then I would
>>>>> definitely make this optional
>>>>>
>>>>>
>>>>>
>>>> I hope I explained the intent above. This header would not  
>>>> define  the input types that the recognize would process, but  
>>>> rather the  input types the client wants it to consider  
>>>> potential barge-in  and hence generate the "start-of-input:"  
>>>> event for and do any  local barge-in optimizations for.
>>>>
>>>>
>>>> Dave.
>>>>
>>>>
>>>>
>>>>>
>>>>>
>>>>>
>>>>>> 3. Change START-OF-SPEECH to START-OF-INPUT. Clarify it occurs  
>>>>>> once
>>>>>> and coincides with barge-in.
>>>>>>
>>>>>>
>>>>>>
>>>>>>
>>>>> Ok for the "only once", but I'd like to tighten it up to talk  
>>>>> about
>>>>> more than just barge-in, as I suggested above.
>>>>>
>>>>>
>>>>>
>>>>>
>>>>>> 4. Add a header of InputType to START-OF-INPUT. Current specified
>>>>>> values are "dtmf" or "speech".
>>>>>>
>>>>>>
>>>>>>
>>>>>>
>>>>> Ok, with slight modification. Say that the syntax of the  
>>>>> "input- type"
>>>>> header is a single-valued subset of the registered value(s)  
>>>>> that are
>>>>> defined for the start-input-on: header.
>>>>>
>>>>> Comments?
>>>>>
>>>>> Dave O. (technical hat on, chair hat off).
>>>>>
>>>>>
>>>>>
>>>>>
>>>>>> Dave
>>>>>>
>>>>>>
>>>>>>
>>>>>> ----- Original Message ----- From: "Shanmugham, Saravanan"
>>>>>> <sarvi@cisco.com>
>>>>>> To: "Eric Burger" <eburger@brooktrout.com>; <speechsc@ietf.org>
>>>>>> Sent: Thursday, July 07, 2005 5:52 PM
>>>>>> Subject: RE: [Speechsc] START-OF-SPEECH in DTMF-only mode ->   
>>>>>> summary
>>>>>> (?)
>>>>>>
>>>>>>
>>>>>> I don't think generating multiple START-OF-SPEECH events is a
>>>>>> solution.
>>>>>> We still haven't addressed, what constitutes a barge-in, for the
>>>>>> optimised case. That should also be the single point when a  
>>>>>> single
>>>>>> START-OF-SPEECH(or whathever else you want to name it) should be
>>>>>> generated.
>>>>>>
>>>>>> That point, in my opinion should be
>>>>>>   1. For "dtmf-recog" resources should be the beginning of a   
>>>>>> DTMF key
>>>>>> press.
>>>>>>   2. For "speech-recog" resources should be the beginning of a  
>>>>>> DTMF
>>>>>> key
>>>>>> press or the beginning of speech. This is should be  
>>>>>> irrespective of
>>>>>> what
>>>>>> type of grammar is being used. Coz even numbers only grammars can
>>>>>> still
>>>>>> be spoken and hence cannot be assumed to be a cue for DTMF only
>>>>>> recognition.
>>>>>>   3. For "speech-only-recog" resources(which are not defined  
>>>>>> today)
>>>>>> the
>>>>>> time to barge-in is the beginning of speech. I don't see a  
>>>>>> need for
>>>>>> such
>>>>>> a resource today. But I am mentioning this for completeness.
>>>>>>
>>>>>> Sarvi
>>>>>>
>>>>>>
>>>>>>     -----Original Message-----
>>>>>>     From: speechsc-bounces@ietf.org
>>>>>>     [mailto:speechsc-bounces@ietf.org] On Behalf Of Eric Burger
>>>>>>     Sent: Thursday, July 07, 2005 5:32 AM
>>>>>>     To: speechsc@ietf.org
>>>>>>     Subject: RE: [Speechsc] START-OF-SPEECH in DTMF-only mode
>>>>>>     -> summary(?)
>>>>>>
>>>>>>     Can we declare consensus?
>>>>>>
>>>>>>
>>>>>>
>>>>>>
>>>>>>> -----Original Message-----
>>>>>>> From: speechsc-bounces@ietf.org
>>>>>>>
>>>>>>>
>>>>>>>
>>>>>>     [mailto:speechsc-bounces@ietf.org]
>>>>>>
>>>>>>
>>>>>>
>>>>>>> Sent: Wednesday, July 06, 2005 10:45 AM
>>>>>>> To: Dave Burke
>>>>>>> Cc: speechsc@ietf.org; Shanmugham, Saravanan; Klaus Reifenrath
>>>>>>> Subject: Re: [Speechsc] START-OF-SPEECH in DTMF-only mode ->
>>>>>>> summary(?)
>>>>>>>
>>>>>>> I think we're getting close. I though about snipping out
>>>>>>>
>>>>>>>
>>>>>>>
>>>>>>     some pieces
>>>>>>
>>>>>>
>>>>>>
>>>>>>> to cut down the text, but I realized the context is
>>>>>>>
>>>>>>>
>>>>>>>
>>>>>>     still needed. See
>>>>>>
>>>>>>
>>>>>>
>>>>>>> inline.
>>>>>>>
>>>>>>> On Jul 6, 2005, at 10:29 AM, Dave Burke wrote:
>>>>>>>
>>>>>>>
>>>>>>>
>>>>>>>
>>>>>>>> Inline.
>>>>>>>>
>>>>>>>> Dave
>>>>>>>>
>>>>>>>> ----- Original Message ----- From: "David R Oran"
>>>>>>>>
>>>>>>>>
>>>>>>>>
>>>>>>     <oran@cisco.com>
>>>>>>
>>>>>>
>>>>>>
>>>>>>>> To: "Dave Burke" <david.burke@voxpilot.com>
>>>>>>>> Cc: <speechsc@ietf.org>; "Shanmugham, Saravanan"
>>>>>>>>
>>>>>>>>
>>>>>>>>
>>>>>>     <sarvi@cisco.com>;
>>>>>>
>>>>>>
>>>>>>
>>>>>>>> "Klaus Reifenrath" <Klaus.Reifenrath@Scansoft.com>
>>>>>>>> Sent: Wednesday, July 06, 2005 1:07 PM
>>>>>>>> Subject: Re: [Speechsc] START-OF-SPEECH in DTMF-only
>>>>>>>>
>>>>>>>>
>>>>>>>>
>>>>>>     mode -> summary
>>>>>>
>>>>>>
>>>>>>
>>>>>>>> (?)
>>>>>>>>
>>>>>>>>
>>>>>>>>
>>>>>>>>
>>>>>>>>
>>>>>>>>
>>>>>>>>>
>>>>>>>>> On Jul 6, 2005, at 5:31 AM, Dave Burke wrote:
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>> + Attempting to summarise:
>>>>>>>>>>
>>>>>>>>>> 1. START-OF-SPEECH is useful for the client to know when
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>>
>>>>>>> to stop
>>>>>>>
>>>>>>>
>>>>>>>
>>>>>>>>>> playing prompts the in non-optimised case 2.
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>>
>>>>>>     START-OF-SPEECH is
>>>>>>
>>>>>>
>>>>>>
>>>>>>>>>> useful for the client to calculate the bargin  time
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>>
>>>>>>     (e.g. VoiceXML
>>>>>>
>>>>>>
>>>>>>
>>>>>>>>>> 2.1 <mark>)
>>>>>>>>>> 3. In the optimised case, a bargin automatically
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>>
>>>>>>     stops prompt
>>>>>>
>>>>>>
>>>>>>
>>>>>>>>>> playing (assuming prompts barginable) 4. Because of
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>>
>>>>>>     the previous
>>>>>>
>>>>>>
>>>>>>
>>>>>>>>>> point, the question of what input type  caused
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>>
>>>>>>     bargin is different
>>>>>>
>>>>>>
>>>>>>
>>>>>>>>>> and less important to
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>>
>>>>>>> what input
>>>>>>>
>>>>>>>
>>>>>>>
>>>>>>>>>> type(s)  the recogniser is listening for
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>> I'm not sure I follow point 4. Could you elaborate?
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>
>>>>>>>>
>>>>>>>> DB> Adding a parameter to the START-OF-SPEECH event would
>>>>>>>>
>>>>>>>>
>>>>>>>>
>>>>>>> certainly
>>>>>>>
>>>>>>>
>>>>>>>
>>>>>>>> allow the client to ignore (i.e. let prompts continue
>>>>>>>>
>>>>>>>>
>>>>>>>>
>>>>>>     playing) the
>>>>>>
>>>>>>
>>>>>>
>>>>>>>> event if the event type is not of interest (e.g. the
>>>>>>>>
>>>>>>>>
>>>>>>>>
>>>>>>     client would
>>>>>>
>>>>>>
>>>>>>
>>>>>>>> ignore speech start events when it is interested only
>>>>>>>>
>>>>>>>>
>>>>>>>>
>>>>>>     in  DTMF start
>>>>>>
>>>>>>
>>>>>>
>>>>>>>> events). This _only_ works for the non-optimised case,
>>>>>>>>
>>>>>>>>
>>>>>>>>
>>>>>>     however. For
>>>>>>
>>>>>>
>>>>>>
>>>>>>>> the optimised case, assuming START-OF-SPEECH
>>>>>>>>
>>>>>>>>
>>>>>>>>
>>>>>>> coincides
>>>>>>>
>>>>>>>
>>>>>>>
>>>>>>>> with the bargin signal to the speechsynth, prompts
>>>>>>>>
>>>>>>>>
>>>>>>>>
>>>>>>     will stop playing
>>>>>>
>>>>>>
>>>>>>
>>>>>>>> for inputs that the client might not be interested
>>>>>>>>
>>>>>>>>
>>>>>>>>
>>>>>>> in (e.g.
>>>>>>>
>>>>>>>
>>>>>>>
>>>>>>>> a speech input will stop prompts playing even if the client
>>>>>>>>
>>>>>>>>
>>>>>>>>
>>>>>>> is only
>>>>>>>
>>>>>>>
>>>>>>>
>>>>>>>> interested in DTMF).
>>>>>>>>
>>>>>>>>
>>>>>>>>
>>>>>>>>
>>>>>>> OK, now I get it. There's a need for the client to both
>>>>>>>
>>>>>>>
>>>>>>>
>>>>>>     handle the
>>>>>>
>>>>>>
>>>>>>
>>>>>>> non-optimized case itself, and influence or at least
>>>>>>>
>>>>>>>
>>>>>>>
>>>>>>     have a clue what
>>>>>>
>>>>>>
>>>>>>
>>>>>>> the server is going to do in the optimized case.
>>>>>>>
>>>>>>>
>>>>>>>
>>>>>>>
>>>>>>>>
>>>>>>>>
>>>>>>>>
>>>>>>>>
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>> 5. Currently, MRCPv2 has no way of indicating what
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>>
>>>>>>     input type(s) a
>>>>>>
>>>>>>
>>>>>>
>>>>>>>>>> recogniser is listening for
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>> Do you mean exactly this, or do you mean "for the client to
>>>>>>>>> indicate  to the resource what input types it should
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>
>>>>>>     look for"?
>>>>>>
>>>>>>
>>>>>>
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>
>>>>>>>>
>>>>>>>> DB> Yes exactly - apologies for not being clear.
>>>>>>>>
>>>>>>>>
>>>>>>>>
>>>>>>>>
>>>>>>>>
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>> + Why implement 5?
>>>>>>>>>>
>>>>>>>>>> i. Noisy case: Need DTMF-only recognition (and may
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>>
>>>>>>     only have a
>>>>>>
>>>>>>
>>>>>>
>>>>>>>>>> speechrecog)
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>> I'm having difficulty following the logic of why
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>
>>>>>>     noise would
>>>>>>
>>>>>>
>>>>>>
>>>>>>>>> necessarily trigger START-OF-SPEECH if you were listening for
>>>>>>>>> speech  but not DTMF. I suppose you can use a more forgiving
>>>>>>>>> discriminator if  all you need to tell is if you're
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>
>>>>>>     getting DTMF,
>>>>>>
>>>>>>
>>>>>>
>>>>>>>>> but I've had a number  of real-world cases where wind
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>
>>>>>>     noise was
>>>>>>
>>>>>>
>>>>>>
>>>>>>>>> detected as DTMF, and  there's always the ambiguity
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>
>>>>>>     when you have
>>>>>>
>>>>>>
>>>>>>
>>>>>>>>> Captain Crunch on the  phone. In either case in the
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>
>>>>>>     non-optimized
>>>>>>
>>>>>>
>>>>>>
>>>>>>>>> case it's the client who  gets to decide whether an
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>
>>>>>>     event should be
>>>>>>
>>>>>>
>>>>>>
>>>>>>>>> interpreted as barge-in or  not, so it seems an
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>
>>>>>>     aesthetic protocol
>>>>>>
>>>>>>
>>>>>>
>>>>>>>>> design decision whether the  client tells the server
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>
>>>>>>     ahead of time
>>>>>>
>>>>>>
>>>>>>
>>>>>>>>> what circumstances to generate  the START-OF-SPEECH
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>
>>>>>>     event for, or
>>>>>>
>>>>>>
>>>>>>
>>>>>>>>> whether the event gets generated  and the client
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>
>>>>>>     decides based on
>>>>>>
>>>>>>
>>>>>>
>>>>>>>>> what's in the event whether it should  be
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>
>>>>>>> treated
>>>>>>>
>>>>>>>
>>>>>>>
>>>>>>>>> as barge- in.
>>>>>>>>>
>>>>>>>>> I suppose one could make the argument that because
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>
>>>>>>     the spec implies
>>>>>>
>>>>>>
>>>>>>
>>>>>>>>> that the event can only be generated once per request
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>
>>>>>>     that if a
>>>>>>
>>>>>>
>>>>>>
>>>>>>>>> DTMF/ speech capable recognizer first hears
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>
>>>>>>> enough noise
>>>>>>>
>>>>>>>
>>>>>>>
>>>>>>>>> to think it's  hearing speech and later hears DTMF,
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>
>>>>>>     the client will
>>>>>>
>>>>>>
>>>>>>
>>>>>>>>> declare barge-in  when the event comes and not when he DTMF
>>>>>>>>> actually gets heard.
>>>>>>>>>
>>>>>>>>> If that's deemed a problem, we can still handle that
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>
>>>>>>     in the design
>>>>>>
>>>>>>
>>>>>>
>>>>>>>>> where the server just reports what it's hearing by allowing
>>>>>>>>> multiple  events to be generated during a single request.
>>>>>>>>>
>>>>>>>>> Between the approach just outlined above, and an
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>
>>>>>>     approach where the
>>>>>>
>>>>>>
>>>>>>
>>>>>>>>> client provides a filter for whether to generate the
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>
>>>>>>     event or not,
>>>>>>
>>>>>>
>>>>>>
>>>>>>>>> I  have a mild preference (based on aesthetics rather
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>
>>>>>>     than some
>>>>>>
>>>>>>
>>>>>>
>>>>>>>>> hard  engineering tradeoff) for the approach where
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>
>>>>>>> the server
>>>>>>>
>>>>>>>
>>>>>>>
>>>>>>>>> just reports  what it's hearing.
>>>>>>>>>
>>>>>>>>>
>>>>>>>>> Having had some useful exchanges on this topic, it
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>
>>>>>>     also is becoming
>>>>>>
>>>>>>
>>>>>>
>>>>>>>>> apparent to me that this event is poorly named, and
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>
>>>>>>     we should
>>>>>>
>>>>>>
>>>>>>
>>>>>>>>> consider renaming it to "INTERESTING-INPUT-HEARD" or
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>
>>>>>>     something akin
>>>>>>
>>>>>>
>>>>>>
>>>>>>>>> to that, because as others have pointed out, a
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>
>>>>>>     DTMF-only recognizer
>>>>>>
>>>>>>
>>>>>>
>>>>>>>>> will never detect "start of speech".
>>>>>>>>>
>>>>>>>>> Another consideration to fold into the design choice is
>>>>>>>>> extensibility. Bear with me through a little
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>
>>>>>>     gedankenexperiment.
>>>>>>
>>>>>>
>>>>>>
>>>>>>>>>
>>>>>>>>> Suppose we want to define a new recognizer type,
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>
>>>>>>     which I'll call
>>>>>>
>>>>>>
>>>>>>
>>>>>>>>> the "name that tune" recognizer. The client plays
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>
>>>>>>     music to the
>>>>>>
>>>>>>
>>>>>>
>>>>>>>>> server and  the server recognizes musical notes. The
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>
>>>>>>     grammar is a
>>>>>>
>>>>>>
>>>>>>
>>>>>>>>> standard  musical notation, augmented with a semantic
>>>>>>>>> interpretation that  transforms the notes into the
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>
>>>>>>     title of the
>>>>>>
>>>>>>
>>>>>>
>>>>>>>>> tune and provides that as  an answer.
>>>>>>>>>
>>>>>>>>> First, there's no speech involved (or is there...hang
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>
>>>>>>     on a minute).
>>>>>>
>>>>>>
>>>>>>
>>>>>>>>> Second, in order to accommodate the "name that tune"
>>>>>>>>> recognizer, we'd have to extend both the client and
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>
>>>>>>     the server to
>>>>>>
>>>>>>
>>>>>>
>>>>>>>>> undetstand a  directive as to whether to recognize
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>
>>>>>>     music or now,
>>>>>>
>>>>>>
>>>>>>
>>>>>>>>> inaddition to what  the server already knows what to
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>
>>>>>>     do based on
>>>>>>
>>>>>>
>>>>>>
>>>>>>>>> the grammar. If you  follow my logic above, whether
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>
>>>>>>     or not we do
>>>>>>
>>>>>>
>>>>>>
>>>>>>>>> that, we have to extend  "start-of-speech" to say
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>
>>>>>>     "I'm hearing
>>>>>>
>>>>>>
>>>>>>
>>>>>>>>> music". So far fairly  straightforward, but let me
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>
>>>>>>     now throw in the
>>>>>>
>>>>>>
>>>>>>
>>>>>>>>> pathological twist.
>>>>>>>>>
>>>>>>>>> Suppose what I feed to a  combined music/speech
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>
>>>>>>     recognizer is a
>>>>>>
>>>>>>
>>>>>>
>>>>>>>>> work  in sprechstimme (spoken music), like the
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>
>>>>>>     "Geographical Fugue"
>>>>>>
>>>>>>
>>>>>>
>>>>>>>>> (aside:  this is a wonderful piece of music I highly
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>
>>>>>>     recommend to
>>>>>>
>>>>>>
>>>>>>
>>>>>>>>> anyone  interested in small ensemble singing). In
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>
>>>>>>     this case, the
>>>>>>
>>>>>>
>>>>>>
>>>>>>>>> tune could  be named by either doing speech or music
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>
>>>>>>     recognition.
>>>>>>
>>>>>>
>>>>>>
>>>>>>>>> Why is there  any need for the client to constrain
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>
>>>>>>     the server as to
>>>>>>
>>>>>>
>>>>>>
>>>>>>>>> which it tries  to do when it's
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>
>>>>>>> already
>>>>>>>
>>>>>>>
>>>>>>>
>>>>>>>>> told the server what it wants through the  grammar?
>>>>>>>>>
>>>>>>>>> A few other comments below
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>> ii. Flexibility: Want speech-only recognition
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>>
>>>>>>     (because a second
>>>>>>
>>>>>>
>>>>>>
>>>>>>>>>> recogniser is doing hotword on DTMF)
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>> I don't see how flexibility is affected by this deisgn
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>
>>>>>>> choice. If
>>>>>>>
>>>>>>>
>>>>>>>
>>>>>>>>> that's what you want, feed the speech-only recognizer
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>
>>>>>>     a grammar
>>>>>>
>>>>>>
>>>>>>
>>>>>>>>> without any DTMF rules.
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>> + How to implement 5?
>>>>>>>>>>
>>>>>>>>>> a. Implicitly:
>>>>>>>>>>    - dtmfrecog: always DTMF-only recognition
>>>>>>>>>>    - speechrecog: depends on active grammar type
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>>> if a dtmf grammar is active then DTMF input is "on"
>>>>>>>>>>> if a speech grammar is active then speech
>>>>>>>>>>>
>>>>>>>>>>>
>>>>>>>>>>>
>>>>>>     input is "on"
>>>>>>
>>>>>>
>>>>>>
>>>>>>>>>>
>>>>>>>>>> b. Explicitly:
>>>>>>>>>>    - Add inputmodes header to RECOGNIZE
>>>>>>>>>>
>>>>>>>>>> Option a is David's "do what I mean case"; option b is
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>>
>>>>>>> the extra
>>>>>>>
>>>>>>>
>>>>>>>
>>>>>>>>>> dial for the client.
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>> Actually, that's not the point I was making with "do
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>
>>>>>>     what I mean",
>>>>>>
>>>>>>
>>>>>>
>>>>>>>>> but it's not essential to the discussion so let's move on.
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>> It is worth noting that VoiceXML uses option b. This
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>>
>>>>>>     allows one to
>>>>>>
>>>>>>
>>>>>>
>>>>>>>>>> activate both speech grammars and DTMF grammars (and
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>>
>>>>>>> therefore
>>>>>>>
>>>>>>>
>>>>>>>
>>>>>>>>>> be informed of any errors in the grammars at activation
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>>
>>>>>>> time) but
>>>>>>>
>>>>>>>
>>>>>>>
>>>>>>>>>> independently turn on whichever input mode you like e.g.
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>>
>>>>>>> perhaps
>>>>>>>
>>>>>>>
>>>>>>>
>>>>>>>>>> start with "both" then change to "dtmf".
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>> I'm not sure the VXML precedent is relevant here,
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>
>>>>>>     because the
>>>>>>
>>>>>>
>>>>>>
>>>>>>>>> application behind VXML is working a different part of the
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>
>>>>>>> problem
>>>>>>>
>>>>>>>
>>>>>>>
>>>>>>>>> -  how to traverse a TUI dialog based on different
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>
>>>>>>     parts of the
>>>>>>
>>>>>>
>>>>>>
>>>>>>>>> input  space. In fact, I suspect that the VXML: choice was
>>>>>>>>> conditioned more  by limitations at the time it was
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>
>>>>>>> specified than
>>>>>>>
>>>>>>>
>>>>>>>
>>>>>>>>> an underlying good  design choice. Clearly having to
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>
>>>>>>     specify this
>>>>>>
>>>>>>
>>>>>>
>>>>>>>>> in VXML make the job of  handling a TUI with nodes
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>
>>>>>>     like "Say or
>>>>>>
>>>>>>
>>>>>>
>>>>>>>>> press 5" harder rather than  easier.
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>
>>>>>>>>
>>>>>>>> DB> The VoiceXML edge-case is pretty weird so it's not a major
>>>>>>>> concern. My main concern is that the client can
>>>>>>>>
>>>>>>>>
>>>>>>>>
>>>>>>     indicate, somehow,
>>>>>>
>>>>>>
>>>>>>
>>>>>>>> what the input modes are.
>>>>>>>>
>>>>>>>>
>>>>>>>>
>>>>>>>>
>>>>>>>>
>>>>>>>>>
>>>>>>>>> Summing up, while I don't feel strongly one way or
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>
>>>>>>     the other, I
>>>>>>
>>>>>>
>>>>>>
>>>>>>>>> have  a preference for handling this as follows:
>>>>>>>>>
>>>>>>>>> a) Rename "START-OF-SPEECH" to
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>
>>>>>>     "INTERESTING-INPUT-RECEIVED" or
>>>>>>
>>>>>>
>>>>>>
>>>>>>>>> something equivalent.
>>>>>>>>> b) Include a parameter in the event saying what was
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>
>>>>>>     interesting
>>>>>>
>>>>>>
>>>>>>
>>>>>>>>> about  the input you received, with a registry of values
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>
>>>>>> which
>>>>>>
>>>>>>
>>>>>>
>>>>>>>>> includes:
>>>>>>>>>     - signal above noise floor
>>>>>>>>>     - speech
>>>>>>>>>     - dtmf
>>>>>>>>>     - (possibly) music
>>>>>>>>> c) allow the event to be generated multiple times
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>
>>>>>>     during a request
>>>>>>
>>>>>>
>>>>>>
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>
>>>>>>>>
>>>>>>>> DB> I like these suggestions (START-OF-INPUT?).
>>>>>>>>
>>>>>>>>
>>>>>>>>
>>>>>>     However, I don't
>>>>>>
>>>>>>
>>>>>>
>>>>>>>> see how the problem of the optimised case is not
>>>>>>>>
>>>>>>>>
>>>>>>>>
>>>>>>     solved by them. I
>>>>>>
>>>>>>
>>>>>>
>>>>>>>> think the optimised case is fine if we have the
>>>>>>>>
>>>>>>>>
>>>>>>>>
>>>>>>     following rules:
>>>>>>
>>>>>>
>>>>>>
>>>>>>>>
>>>>>>>>
>>>>>>>>
>>>>>>>>
>>>>>>> Yes, I hadn't thought through the optimized case as
>>>>>>>
>>>>>>>
>>>>>>>
>>>>>>     thoroughly as you.
>>>>>>
>>>>>>
>>>>>>
>>>>>>> Your suggested method name is fine by me as well
>>>>>>>
>>>>>>>
>>>>>>>
>>>>>>>
>>>>>>>> 1. START-OF-SPEECH (and optimised bargin) is only generated
>>>>>>>>
>>>>>>>>
>>>>>>>>
>>>>>>> for the
>>>>>>>
>>>>>>>
>>>>>>>
>>>>>>>> input type that is being listened for 2. A speechrecog
>>>>>>>>
>>>>>>>>
>>>>>>>>
>>>>>>     listens for
>>>>>>
>>>>>>
>>>>>>
>>>>>>>> DTMF if DTMF grammars are active, speech if speech
>>>>>>>>
>>>>>>>>
>>>>>>>>
>>>>>>     grammars are
>>>>>>
>>>>>>
>>>>>>
>>>>>>>> active, or speech and DTMF if both grammar types are active.
>>>>>>>>
>>>>>>>>
>>>>>>>>
>>>>>>>>
>>>>>>> Works for me.
>>>>>>>
>>>>>>>
>>>>>>>
>>>>>>>
>>>>>>>>
>>>>>>>>
>>>>>>>>
>>>>>>>>
>>>>>>>>> Note that all of the above I'm saying with my technical
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>
>>>>>>> hat on and
>>>>>>>
>>>>>>>
>>>>>>>
>>>>>>>>> my chair hat off.
>>>>>>>>> Putting my chair hat on for a moment, we really need
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>
>>>>>>     to get this
>>>>>>
>>>>>>
>>>>>>
>>>>>>>>> spec  to last call, so at some point Eric or I is going to
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>
>>>>>>> declare
>>>>>>>
>>>>>>>
>>>>>>>
>>>>>>>>> rough  consensus so we can move on.
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>
>>>>>>>>
>>>>>>>> DB> Agreed!
>>>>>>>>
>>>>>>>>
>>>>>>>>
>>>>>>>>
>>>>>>>>
>>>>>>>>>
>>>>>>>>> Dave Oran.
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>> Dave
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>> Sarvi makes a good point that adding the reason why the
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>>
>>>>>>> START-OF-
>>>>>>>
>>>>>>>
>>>>>>>
>>>>>>>>>> SPEECH occurred does not fix the optimised bargin case.
>>>>>>>>>>
>>>>>>>>>> dtmfrecog - listens for DTMF only
>>>>>>>>>> speechrecog - listens for DTMF only, or speech only,
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>>
>>>>>>     or speech &
>>>>>>
>>>>>>
>>>>>>
>>>>>>>>>> DTMF
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>> ----- Original Message ----- From: "Shanmugham, Saravanan"
>>>>>>>>>> <sarvi@cisco.com>
>>>>>>>>>> To: "David R Oran" <oran@cisco.com>; "Dave Burke"
>>>>>>>>>> <david.burke@voxpilot.com>
>>>>>>>>>> Cc: <speechsc@ietf.org>; "Klaus Reifenrath"
>>>>>>>>>> <Klaus.Reifenrath@Scansoft.com>
>>>>>>>>>> Sent: Tuesday, July 05, 2005 8:51 PM
>>>>>>>>>> Subject: RE: [Speechsc] START-OF-SPEECH in DTMF-only mode
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>> inline.
>>>>>>>>>>
>>>>>>>>>>     -----Original Message-----
>>>>>>>>>>     From: David R Oran [mailto:oran@cisco.com]
>>>>>>>>>>     Sent: Tuesday, July 05, 2005 11:08 AM
>>>>>>>>>>     To: Dave Burke
>>>>>>>>>>     Cc: Shanmugham, Saravanan; Klaus Reifenrath;
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>>
>>>>>>     speechsc@ietf.org
>>>>>>
>>>>>>
>>>>>>
>>>>>>>>>>     Subject: Re: [Speechsc] START-OF-SPEECH in DTMF-only
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>>
>>>>>> mode
>>>>>>
>>>>>>
>>>>>>
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>>     On Jul 5, 2005, at 12:54 PM, Dave Burke wrote:
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>>> Inline.
>>>>>>>>>>>
>>>>>>>>>>> Dave
>>>>>>>>>>>
>>>>>>>>>>> ----- Original Message ----- From:
>>>>>>>>>>>
>>>>>>>>>>>
>>>>>>>>>>>
>>>>>>     "Shanmugham, Saravanan"
>>>>>>
>>>>>>
>>>>>>
>>>>>>>>>>> <sarvi@cisco.com>
>>>>>>>>>>> To: "David R Oran" <oran@cisco.com>; "Klaus
>>>>>>>>>>>
>>>>>>>>>>>
>>>>>>>>>>>
>>>>>> Reifenrath"
>>>>>>
>>>>>>
>>>>>>
>>>>>>>>>>> <Klaus.Reifenrath@Scansoft.com>
>>>>>>>>>>> Cc: <speechsc@ietf.org>
>>>>>>>>>>> Sent: Tuesday, July 05, 2005 5:22 PM
>>>>>>>>>>> Subject: RE: [Speechsc] START-OF-SPEECH in
>>>>>>>>>>>
>>>>>>>>>>>
>>>>>>>>>>>
>>>>>>     DTMF-only mode
>>>>>>
>>>>>>
>>>>>>
>>>>>>>>>>>
>>>>>>>>>>>
>>>>>>>>>>> I agree with Dave's analysis. The purpose of this
>>>>>>>>>>>
>>>>>>>>>>>
>>>>>>>>>>>
>>>>>> event
>>>>>>
>>>>>>
>>>>>>
>>>>>>>>>>     was barge-in.
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>>> And barge-in should happen for both DTMF and speech.
>>>>>>>>>>>
>>>>>>>>>>> Is there a case where you think it should not behave
>>>>>>>>>>>
>>>>>>>>>>>
>>>>>>>>>>>
>>>>>>>>>>     this way. If soe,
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>>> please provide a scenario where you think
>>>>>>>>>>>    1. Barge-in should happen for DTMF and not voice or
>>>>>>>>>>>
>>>>>>>>>>>
>>>>>>>>>>>
>>>>>>>>>>     vice-versa.
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>>>
>>>>>>>>>>> DB> You want to do a DTMF recognition only
>>>>>>>>>>>
>>>>>>>>>>>
>>>>>>>>>>>
>>>>>>     because it is
>>>>>>
>>>>>>
>>>>>>
>>>>>>>>>> noisy.
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>>> While waiting for DTMF input, the speechrecog resource
>>>>>>>>>>>
>>>>>>>>>>>
>>>>>>>>>>>
>>>>>>>>>>     (or advanced
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>>> dtmfrecog) generates a START-OF-SPEECH because it
>>>>>>>>>>>
>>>>>>>>>>>
>>>>>>>>>>>
>>>>>> heard
>>>>>>
>>>>>>
>>>>>>
>>>>>>>>>>     some speech.
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>>> The client does not want to stop prompt playing unless
>>>>>>>>>>>
>>>>>>>>>>>
>>>>>>>>>>>
>>>>>>>>>>     DTMF was heard
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>>> but it can't tell by the START-OF-SPEECH whether
>>>>>>>>>>>
>>>>>>>>>>>
>>>>>>>>>>>
>>>>>> speech
>>>>>>
>>>>>>
>>>>>>
>>>>>>>>>>     or DTMF was
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>>> heard. Similarly vice versa.
>>>>>>>>>>>
>>>>>>>>>>>
>>>>>>>>>>>
>>>>>>>>>>>
>>>>>>>>>>     It's an interesting design question what part of the
>>>>>>>>>>     policy resides at the client and what at the server, and
>>>>>>>>>>     who makes the "final decision" about whether what was
>>>>>>>>>>     heard was relevant to the control channel. Right now we
>>>>>>>>>>     (IMO) have a weird partitioning in many cases where the
>>>>>>>>>>     client basically says "do what I mean", but there are no
>>>>>>>>>>     constraints of what the server actually does, and no
>>>>>>>>>>     normalized basis for the client to figure out what to
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>>
>>>>>> set
>>>>>>
>>>>>>
>>>>>>
>>>>>>>>>>     various magic numbers to (e.g. sensitivity).
>>>>>>>>>>
>>>>>>>>>>     In this case the only thing the client needs to
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>>
>>>>>> decide is
>>>>>>
>>>>>>
>>>>>>
>>>>>>>>>>     whether to kill the prompt because the server thinks
>>>>>>>>>>     something that would interfere with the feedback
>>>>>>>>>>     ear/mouth/finger control happened. What this
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>>
>>>>>>     says to me is
>>>>>>
>>>>>>
>>>>>>
>>>>>>>>>>     that it isn't necessarily a good idea for the client to
>>>>>>>>>>     have more knobs to control the server
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>>
>>>>>>     (especially if those
>>>>>>
>>>>>>
>>>>>>
>>>>>>>>>>     knows are just more value/policy input ungrounded in any
>>>>>>>>>>     physics/ acoustics). On the other hand, having the
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>>
>>>>>> server
>>>>>>
>>>>>>
>>>>>>
>>>>>>>>>>     tell the client more about what it thinks is going on is
>>>>>>>>>>     probably valuable.
>>>>>>>>>>
>>>>>>>>>>     So, Coming to the point after this long rambling
>>>>>>>>>>     introduction, I think it would in fact be useful for the
>>>>>>>>>>     START-Of-SPEECH event to indicate some extra
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>>
>>>>>> information,
>>>>>>
>>>>>>
>>>>>>
>>>>>>>>>>     for example:
>>>>>>>>>>     a) I got something enough above the noise floor
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>>
>>>>>>     to qualify
>>>>>>
>>>>>>
>>>>>>
>>>>>>>>>>     for exceeding the "Sensisitvity" parameter you sent
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>>
>>>>>> in on
>>>>>>
>>>>>>
>>>>>>
>>>>>>>>>>     the request but I really can't tell what it is
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>>
>>>>>>     (could be a
>>>>>>
>>>>>>
>>>>>>
>>>>>>>>>>     hippopatmus fart, or a siren in the background,
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>>
>>>>>>     or captain
>>>>>>
>>>>>>
>>>>>>
>>>>>>>>>>     crunch trying to whistle DTMF).
>>>>>>>>>>     b) I think I'm hearing speech
>>>>>>>>>>     c) I think I'm hearing DTMF
>>>>>>>>>>
>>>>>>>>>> Though I agree with your former part of your
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>>
>>>>>>     response. I am not
>>>>>>
>>>>>>
>>>>>>
>>>>>>>>>> sure I agree with your proposed solution.
>>>>>>>>>> The way I see this problem is that, it is more of what
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>>
>>>>>>> constitues a
>>>>>>>
>>>>>>>
>>>>>>>
>>>>>>>>>> barge-in event. This boils down to whether it is
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>>
>>>>>>     speech, DTMF or
>>>>>>
>>>>>>
>>>>>>
>>>>>>>>>> both.
>>>>>>>>>> This is inturn boils down to what type of recognizer
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>>
>>>>>>> resource we are
>>>>>>>
>>>>>>>
>>>>>>>
>>>>>>>>>> using, dtmf-recog, speech-recog, and speech-only-recog(we
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>>
>>>>>>> don't have
>>>>>>>
>>>>>>>
>>>>>>>
>>>>>>>>>> this and I don't think we should add it, but think
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>>
>>>>>>     of this as a
>>>>>>
>>>>>>
>>>>>>
>>>>>>>>>> place holder that explains the concept).
>>>>>>>>>>
>>>>>>>>>> A client knowing what type of barge-in happenned, does
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>>
>>>>>>> not impact
>>>>>>>
>>>>>>>
>>>>>>>
>>>>>>>>>> the
>>>>>>>>>> barge-in operation itself as it may be too late(for
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>>
>>>>>>     the optimized
>>>>>>
>>>>>>
>>>>>>
>>>>>>>>>> barge-in case). It may have other use cases, and if we
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>>
>>>>>>> can identify
>>>>>>>
>>>>>>>
>>>>>>>
>>>>>>>>>> them, I don't mind adding support for the
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>>
>>>>>>     START-OF-SPEECH event to
>>>>>>
>>>>>>
>>>>>>
>>>>>>>>>> say what type of barge-in happenned. But that itself does
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>>
>>>>>> not
>>>>>>
>>>>>>
>>>>>>
>>>>>>> solve the
>>>>>>>
>>>>>>>
>>>>>>>
>>>>>>>>>> original problem raised. Refer to my previous response.
>>>>>>>>>>
>>>>>>>>>> The solution lies in defining what what is a
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>>
>>>>>>     barge-in event.
>>>>>>
>>>>>>
>>>>>>
>>>>>>>>>> That  boils
>>>>>>>>>> down to what type of recognition is happenning,
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>>
>>>>>>> dtmf-only, speech-
>>>>>>>
>>>>>>>
>>>>>>>
>>>>>>>>>> dtmf
>>>>>>>>>> or speech-only. We do not support speech-only as a
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>>
>>>>>>     resource today,
>>>>>>
>>>>>>
>>>>>>
>>>>>>>>>> the question is do we need a header to force it.
>>>>>>>>>>
>>>>>>>>>> Sarvi
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>>>    2. You would benefit from the client knowing
>>>>>>>>>>>
>>>>>>>>>>>
>>>>>>>>>>>
>>>>>>> what caused
>>>>>>>
>>>>>>>
>>>>>>>
>>>>>>>>>> the
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>>> barge-in, DTMF Vs speech.
>>>>>>>>>>>
>>>>>>>>>>> DB> See previous comment. And previous e-mail:
>>>>>>>>>>>
>>>>>>>>>>>
>>>>>>>>>>>
>>>>>>     either add an
>>>>>>
>>>>>>
>>>>>>
>>>>>>>>>>> inputmodes header (taking value speech, dtmf, both) to
>>>>>>>>>>>
>>>>>>>>>>>
>>>>>>>>>>>
>>>>>>>>>>     the RECOGNIZE
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>>> request or add a header to the START-OF-SPEECH event
>>>>>>>>>>>
>>>>>>>>>>>
>>>>>>>>>>>
>>>>>>>>>>     indicating DTMF
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>>> or speech.
>>>>>>>>>>>
>>>>>>>>>>>
>>>>>>>>>>>
>>>>>>>>>>>
>>>>>>>>>>     I'm leaning in your direction on this latter point - as
>>>>>>>>>>     should be evident from what I wrote above.
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>>> Sarvi
>>>>>>>>>>>
>>>>>>>>>>>     -----Original Message-----
>>>>>>>>>>>     From: speechsc-bounces@ietf.org
>>>>>>>>>>>     [mailto:speechsc-bounces@ietf.org] On Behalf Of
>>>>>>>>>>>
>>>>>>>>>>>
>>>>>>>>>>>
>>>>>>> David R
>>>>>>>
>>>>>>>
>>>>>>>
>>>>>>>>>> Oran
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>>>     Sent: Tuesday, July 05, 2005 5:24 AM
>>>>>>>>>>>     To: Klaus Reifenrath
>>>>>>>>>>>     Cc: 'speechsc@ietf.org'
>>>>>>>>>>>     Subject: Re: [Speechsc] START-OF-SPEECH in
>>>>>>>>>>>
>>>>>>>>>>>
>>>>>>>>>>>
>>>>>>> DTMF-only mode
>>>>>>>
>>>>>>>
>>>>>>>
>>>>>>>>>>>
>>>>>>>>>>>
>>>>>>>>>>>     On Jul 5, 2005, at 3:46 AM, Reifenrath,
>>>>>>>>>>>
>>>>>>>>>>>
>>>>>>>>>>>
>>>>>>     Klaus wrote:
>>>>>>
>>>>>>
>>>>>>
>>>>>>>>>>>
>>>>>>>>>>>
>>>>>>>>>>>
>>>>>>>>>>>
>>>>>>>>>>>> The current spec is not clear when
>>>>>>>>>>>>
>>>>>>>>>>>>
>>>>>>>>>>>>
>>>>>>> START-OF-SPEECH need
>>>>>>>
>>>>>>>
>>>>>>>
>>>>>>>>>>>     to be send in
>>>>>>>>>>>
>>>>>>>>>>>
>>>>>>>>>>>
>>>>>>>>>>>> the following scenarios:
>>>>>>>>>>>> A) The client requested a DTMF Recognizer. Is
>>>>>>>>>>>>
>>>>>>>>>>>>
>>>>>>>>>>>>
>>>>>> the
>>>>>>
>>>>>>
>>>>>>
>>>>>>>>>>>     START-OF-SPEECH
>>>>>>>>>>>
>>>>>>>>>>>
>>>>>>>>>>>
>>>>>>>>>>>> event send to the client also if speech
>>>>>>>>>>>>
>>>>>>>>>>>>
>>>>>>>>>>>>
>>>>>>     was detected?
>>>>>>
>>>>>>
>>>>>>
>>>>>>>>>>>     I suspect so, since one of the prime purposes
>>>>>>>>>>>
>>>>>>>>>>>
>>>>>>>>>>>
>>>>>>> is to enable
>>>>>>>
>>>>>>>
>>>>>>>
>>>>>>>>>>>     client- mediated barge-in handling. However, if
>>>>>>>>>>>
>>>>>>>>>>>
>>>>>>>>>>>
>>>>>> the
>>>>>>
>>>>>>
>>>>>>
>>>>>>>>>>>     recognizer is in fact only capable of
>>>>>>>>>>>
>>>>>>>>>>>
>>>>>>>>>>>
>>>>>>     recognizing DTMF
>>>>>>
>>>>>>
>>>>>>
>>>>>>>>>>>     then it may in fact not report anythin
>>>>>>>>>>>
>>>>>>>>>>>
>>>>>>>>>>>
>>>>>>     unless it's using
>>>>>>
>>>>>>
>>>>>>
>>>>>>>>>>>     some primitive thresholding machinery, like a SN
>>>>>>>>>>>
>>>>>>>>>>>
>>>>>>>>>>>
>>>>>>>>>> threshold.
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>>>> B) The client requested a Speech
>>>>>>>>>>>>
>>>>>>>>>>>>
>>>>>>>>>>>>
>>>>>>     Recognizer, but only
>>>>>>
>>>>>>
>>>>>>
>>>>>>>>>>>     activated DTMF
>>>>>>>>>>>
>>>>>>>>>>>
>>>>>>>>>>>
>>>>>>>>>>>> grammars. Is the START-OF-SPEECH event
>>>>>>>>>>>>
>>>>>>>>>>>>
>>>>>>>>>>>>
>>>>>>     send to the
>>>>>>
>>>>>>
>>>>>>
>>>>>>>>>>>     client also if
>>>>>>>>>>>
>>>>>>>>>>>
>>>>>>>>>>>
>>>>>>>>>>>> speech was detected?
>>>>>>>>>>>>
>>>>>>>>>>>>
>>>>>>>>>>>>
>>>>>>>>>>>     Again, I'd say yes, for the same reason as above.
>>>>>>>>>>>
>>>>>>>>>>>
>>>>>>>>>>>
>>>>>>>>>>>> I think in both cases START-OF-SPEECH should
>>>>>>>>>>>>
>>>>>>>>>>>>
>>>>>>>>>>>>
>>>>>> only
>>>>>>
>>>>>>
>>>>>>
>>>>>>>>>>     be send after
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>>>> detecting a DTMF digit (see Figure 12 of
>>>>>>>>>>>>
>>>>>>>>>>>>
>>>>>>>>>>>>
>>>>>> VoiceXML
>>>>>>
>>>>>>
>>>>>>
>>>>>>>>>>     2.0: http://
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>>>> www.w3.org/TR/voicexml20/#dmlATiming).
>>>>>>>>>>>>
>>>>>>>>>>>>
>>>>>>>>>>>>
>>>>>>>>>>>     We seem to have reached different
>>>>>>>>>>>
>>>>>>>>>>>
>>>>>>>>>>>
>>>>>>     conclusions. I'd be
>>>>>>
>>>>>>
>>>>>>
>>>>>>>>>>>     interested in why you think my analysis
>>>>>>>>>>>
>>>>>>>>>>>
>>>>>>>>>>>
>>>>>>     above is wrong.
>>>>>>
>>>>>>
>>>>>>
>>>>>>>>>>>
>>>>>>>>>>>     Dave.
>>>>>>>>>>>
>>>>>>>>>>>
>>>>>>>>>>>
>>>>>>>>>>>
>>>>>>>>>>>> Klaus
>>>>>>>>>>>>
>>>>>>>>>>>> _______________________________________________
>>>>>>>>>>>> Speechsc mailing list
>>>>>>>>>>>> Speechsc@ietf.org
>>>>>>>>>>>> https://www1.ietf.org/mailman/listinfo/speechsc
>>>>>>>>>>>>
>>>>>>>>>>>>
>>>>>>>>>>>>
>>>>>>>>>>>>
>>>>>>>>>>>
>>>>>>>>>>>     _______________________________________________
>>>>>>>>>>>     Speechsc mailing list
>>>>>>>>>>>     Speechsc@ietf.org
>>>>>>>>>>>     https://www1.ietf.org/mailman/listinfo/speechsc
>>>>>>>>>>>
>>>>>>>>>>>
>>>>>>>>>>> _______________________________________________
>>>>>>>>>>> Speechsc mailing list
>>>>>>>>>>> Speechsc@ietf.org
>>>>>>>>>>> https://www1.ietf.org/mailman/listinfo/speechsc
>>>>>>>>>>>
>>>>>>>>>>>
>>>>>>>>>>>
>>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>> _______________________________________________
>>>>>>>>>> Speechsc mailing list
>>>>>>>>>> Speechsc@ietf.org
>>>>>>>>>> https://www1.ietf.org/mailman/listinfo/speechsc
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>
>>>>>>>>> _______________________________________________
>>>>>>>>> Speechsc mailing list
>>>>>>>>> Speechsc@ietf.org
>>>>>>>>> https://www1.ietf.org/mailman/listinfo/speechsc
>>>>>>>>>
>>>>>>>>>
>>>>>>>>>
>>>>>>>>
>>>>>>>>
>>>>>>>>
>>>>>>>>
>>>>>>>
>>>>>>> _______________________________________________
>>>>>>> Speechsc mailing list
>>>>>>> Speechsc@ietf.org
>>>>>>> https://www1.ietf.org/mailman/listinfo/speechsc
>>>>>>>
>>>>>>>
>>>>>>>
>>>>>>>
>>>>>>
>>>>>>
>>>>>>     _______________________________________________
>>>>>>     Speechsc mailing list
>>>>>>     Speechsc@ietf.org
>>>>>>     https://www1.ietf.org/mailman/listinfo/speechsc
>>>>>>
>>>>>>
>>>>>> _______________________________________________
>>>>>> Speechsc mailing list
>>>>>> Speechsc@ietf.org
>>>>>> https://www1.ietf.org/mailman/listinfo/speechsc
>>>>>>
>>>>>>
>>>>>> _______________________________________________
>>>>>> Speechsc mailing list
>>>>>> Speechsc@ietf.org
>>>>>> https://www1.ietf.org/mailman/listinfo/speechsc
>>>>>>
>>>>>>
>>>>>>
>>>>>>
>>>>>
>>>>> _______________________________________________
>>>>> Speechsc mailing list
>>>>> Speechsc@ietf.org
>>>>> https://www1.ietf.org/mailman/listinfo/speechsc
>>>>>
>>>>>
>>>>>
>>>>>
>>>>>
>>>>>
>>>>
>>>> _______________________________________________
>>>> Speechsc mailing list
>>>> Speechsc@ietf.org
>>>> https://www1.ietf.org/mailman/listinfo/speechsc
>>>>
>>>
>

_______________________________________________
Speechsc mailing list
Speechsc@ietf.org
https://www1.ietf.org/mailman/listinfo/speechsc



From speechsc-bounces@ietf.org Fri Jul 15 07:42:00 2005
Received: from localhost.localdomain ([127.0.0.1] helo=megatron.ietf.org)
	by megatron.ietf.org with esmtp (Exim 4.32)
	id 1DtOZs-00076b-AZ; Fri, 15 Jul 2005 07:42:00 -0400
Received: from odin.ietf.org ([132.151.1.176] helo=ietf.org)
	by megatron.ietf.org with esmtp (Exim 4.32) id 1DtOZq-00076U-MZ
	for speechsc@megatron.ietf.org; Fri, 15 Jul 2005 07:41:58 -0400
Received: from ietf-mx.ietf.org (ietf-mx [132.151.6.1])
	by ietf.org (8.9.1a/8.9.1a) with ESMTP id HAA27339
	for <speechsc@ietf.org>; Fri, 15 Jul 2005 07:41:56 -0400 (EDT)
Received: from [195.222.227.20] (helo=gromit.ibp.de ident=Debian-exim)
	by ietf-mx.ietf.org with esmtp (Exim 4.43) id 1DtP2d-00058c-41
	for speechsc@ietf.org; Fri, 15 Jul 2005 08:11:44 -0400
Received: from [195.222.227.22] (helo=[195.222.227.22])
	by gromit.ibp.de with esmtp (Exim 4.50) id 1DtOPV-00059r-6W
	for speechsc@ietf.org; Fri, 15 Jul 2005 13:31:17 +0200
Message-ID: <42D7A17B.6000902@ibp.de>
Date: Fri, 15 Jul 2005 13:43:55 +0200
From: Claudia Daboul <claudia@ibp.de>
User-Agent: Mozilla Thunderbird 0.8 (Windows/20040913)
X-Accept-Language: en-us, en
MIME-Version: 1.0
To: speechsc@ietf.org
X-Enigmail-Version: 0.86.1.0
X-Enigmail-Supports: pgp-inline, pgp-mime
Content-Type: text/plain; charset=ISO-8859-1; format=flowed
Content-Transfer-Encoding: 7bit
X-Spam-Score: 0.0 (/)
X-Spam-Score: 0.0 (/)
X-Scan-Signature: b19722fc8d3865b147c75ae2495625f2
Content-Transfer-Encoding: 7bit
Subject: [Speechsc] RECOGNIZE completion-cause 003 and 008
X-BeenThere: speechsc@ietf.org
X-Mailman-Version: 2.1.5
Precedence: list
List-Id: Speech Services Control Working Group <speechsc.ietf.org>
List-Unsubscribe: <https://www1.ietf.org/mailman/listinfo/speechsc>,
	<mailto:speechsc-request@ietf.org?subject=unsubscribe>
List-Post: <mailto:speechsc@ietf.org>
List-Help: <mailto:speechsc-request@ietf.org?subject=help>
List-Subscribe: <https://www1.ietf.org/mailman/listinfo/speechsc>,
	<mailto:speechsc-request@ietf.org?subject=subscribe>
Sender: speechsc-bounces@ietf.org
Errors-To: speechsc-bounces@ietf.org

The spec defines the following completion-causes for the RECOGNIZE request

003          recognition-timeout
                                    RECOGNIZE completed without a match
                                    due to a recognition-timeout
and

008           too-much-speech-timeout
                                    RECOGNIZE request terminated because
                                    speech was too long.

The recognition timeout parameter is explained as follows:

Recognition Timeout

    When recognition is started and there is no match for a certain
    period of time, the recognizer can send a RECOGNITION-COMPLETE event
    to the client and terminate the recognition operation. It is the
    timer that is started when START-OF-SPEECH event is generated by the
    resource and specifies the maximum duration of the utterance. When
    this timer expires the recognition request would complete with a
    status code of "008 too-much-speech-timeout". The recognition-
    timeout header field sets this timeout value. The value is in
    milliseconds. The value for this field ranges from 0 to MAXTIMEOUT,
    where MAXTIMEOUT is platform specific. The default value is 10
    seconds. This header field MAY occur in RECOGNIZE, SET-PARAMS or
    GET-PARAMS.


I am wondering if there is any difference between completion-causes 003 
and 008. Can anybody clarify this?

Claudia

_______________________________________________
Speechsc mailing list
Speechsc@ietf.org
https://www1.ietf.org/mailman/listinfo/speechsc



From speechsc-bounces@ietf.org Fri Jul 15 08:13:40 2005
Received: from localhost.localdomain ([127.0.0.1] helo=megatron.ietf.org)
	by megatron.ietf.org with esmtp (Exim 4.32)
	id 1DtP4W-0001FQ-Bj; Fri, 15 Jul 2005 08:13:40 -0400
Received: from odin.ietf.org ([132.151.1.176] helo=ietf.org)
	by megatron.ietf.org with esmtp (Exim 4.32) id 1DtP4V-0001FL-QU
	for speechsc@megatron.ietf.org; Fri, 15 Jul 2005 08:13:39 -0400
Received: from ietf-mx.ietf.org (ietf-mx [132.151.6.1])
	by ietf.org (8.9.1a/8.9.1a) with ESMTP id IAA29358
	for <speechsc@ietf.org>; Fri, 15 Jul 2005 08:13:38 -0400 (EDT)
Received: from sj-iport-2-in.cisco.com ([171.71.176.71]
	helo=sj-iport-2.cisco.com) by ietf-mx.ietf.org with esmtp (Exim 4.43)
	id 1DtPXJ-0006Db-9G
	for speechsc@ietf.org; Fri, 15 Jul 2005 08:43:26 -0400
Received: from sj-core-2.cisco.com (171.71.177.254)
	by sj-iport-2.cisco.com with ESMTP; 15 Jul 2005 05:13:38 -0700
Received: from imail.cisco.com (imail.cisco.com [128.107.200.91])
	by sj-core-2.cisco.com (8.12.10/8.12.6) with ESMTP id j6FCDWod008786;
	Fri, 15 Jul 2005 05:13:32 -0700 (PDT)
Received: from [10.32.245.152] (stealth-10-32-245-152.cisco.com
	[10.32.245.152])
	by imail.cisco.com (8.12.11/8.12.10) with SMTP id j6FCC9VU015929;
	Fri, 15 Jul 2005 05:12:09 -0700
In-Reply-To: <42D7A17B.6000902@ibp.de>
References: <42D7A17B.6000902@ibp.de>
Mime-Version: 1.0 (Apple Message framework v733)
Content-Type: text/plain; charset=US-ASCII; delsp=yes; format=flowed
Message-Id: <F68CC13F-5C6D-4337-B8E4-D3F3E5C3AD8D@cisco.com>
Content-Transfer-Encoding: 7bit
From: David R Oran <oran@cisco.com>
Subject: Re: [Speechsc] RECOGNIZE completion-cause 003 and 008
Date: Fri, 15 Jul 2005 08:13:34 -0400
To: Claudia Daboul <claudia@ibp.de>
X-Mailer: Apple Mail (2.733)
IIM-SIG: v:"1.1"; h:"imail.cisco.com"; d:"cisco.com"; z:"home"; m:"krs";
	t:"1121429530.332169"; x:"432200"; a:"rsa-sha1"; b:"nofws:1441";
	e:"Iw=="; n:"sQYarK2E51MdcTiUqeif3F7cWdxIfoCiXhdfb9vD5ee/j0jXL15gbFxF2p"
	"XIweAblu0N6XAgK7k+wrbr7bQDJaCDqOmzqpRUBjIRQAXQ7NzadpmR3pUL6wxaRUtW+c43sl9jC"
	"50Qg1sXHpPjt8Y+Y16ioyQAQAdSunM4YhevURc=";
	s:"eV0XNhOJEIreaZ3OfVPOKjZAPxE9WOiCDoTKsDiS77ofDrEGu12N17bjXAhSveuonU+jwsuE"
	"tttj9vObFforpgqe7WuaRv7lQlHOSv7Q2Vl6ye56CU9+H6q/27dUeqP7b5YO1VcwtMPS+X3adDQ"
	"QxjXIpPsyIjaVyKRzP+EJ6RA="; c:"From: David R Oran <oran@cisco.com>";
	c:"Subject: Re: [Speechsc] RECOGNIZE completion-cause 003 and 008";
	c:"Date: Fri, 15 Jul 2005 08:13:34 -0400"
IIM-VERIFY: s:"y"; v:"y"; r:"60"; h:"imail.cisco.com";
	c:"message from imail.cisco.com verified; "
X-Spam-Score: 0.0 (/)
X-Scan-Signature: 8b431ad66d60be2d47c7bfeb879db82c
Content-Transfer-Encoding: 7bit
Cc: speechsc@ietf.org
X-BeenThere: speechsc@ietf.org
X-Mailman-Version: 2.1.5
Precedence: list
List-Id: Speech Services Control Working Group <speechsc.ietf.org>
List-Unsubscribe: <https://www1.ietf.org/mailman/listinfo/speechsc>,
	<mailto:speechsc-request@ietf.org?subject=unsubscribe>
List-Post: <mailto:speechsc@ietf.org>
List-Help: <mailto:speechsc-request@ietf.org?subject=help>
List-Subscribe: <https://www1.ietf.org/mailman/listinfo/speechsc>,
	<mailto:speechsc-request@ietf.org?subject=subscribe>
Sender: speechsc-bounces@ietf.org
Errors-To: speechsc-bounces@ietf.org


On Jul 15, 2005, at 7:43 AM, Claudia Daboul wrote:

> The spec defines the following completion-causes for the RECOGNIZE  
> request
>
> 003          recognition-timeout
>                                    RECOGNIZE completed without a match
>                                    due to a recognition-timeout
> and
>
> 008           too-much-speech-timeout
>                                    RECOGNIZE request terminated  
> because
>                                    speech was too long.
>
> The recognition timeout parameter is explained as follows:
>
> Recognition Timeout
>
>    When recognition is started and there is no match for a certain
>    period of time, the recognizer can send a RECOGNITION-COMPLETE  
> event
>    to the client and terminate the recognition operation. It is the
>    timer that is started when START-OF-SPEECH event is generated by  
> the
>    resource and specifies the maximum duration of the utterance. When
>    this timer expires the recognition request would complete with a
>    status code of "008 too-much-speech-timeout". The recognition-
>    timeout header field sets this timeout value. The value is in
>    milliseconds. The value for this field ranges from 0 to MAXTIMEOUT,
>    where MAXTIMEOUT is platform specific. The default value is 10
>    seconds. This header field MAY occur in RECOGNIZE, SET-PARAMS or
>    GET-PARAMS.
>
>
> I am wondering if there is any difference between completion-causes  
> 003 and 008. Can anybody clarify this?
>
I'll take a whack. 003 says that the recognizer didn't see anything  
that matched an active grammar before the timer ran out. 008 says the  
recognizer got an utterance that matched the grammar but the user  
kept babbling on and didn't stop so the match is probably bogus.

> Claudia
>
> _______________________________________________
> Speechsc mailing list
> Speechsc@ietf.org
> https://www1.ietf.org/mailman/listinfo/speechsc
>

_______________________________________________
Speechsc mailing list
Speechsc@ietf.org
https://www1.ietf.org/mailman/listinfo/speechsc



From speechsc-bounces@ietf.org Fri Jul 15 09:36:49 2005
Received: from localhost.localdomain ([127.0.0.1] helo=megatron.ietf.org)
	by megatron.ietf.org with esmtp (Exim 4.32)
	id 1DtQMy-0006pu-V4; Fri, 15 Jul 2005 09:36:48 -0400
Received: from odin.ietf.org ([132.151.1.176] helo=ietf.org)
	by megatron.ietf.org with esmtp (Exim 4.32) id 1DtQMx-0006pp-1W
	for speechsc@megatron.ietf.org; Fri, 15 Jul 2005 09:36:47 -0400
Received: from ietf-mx.ietf.org (ietf-mx [132.151.6.1])
	by ietf.org (8.9.1a/8.9.1a) with ESMTP id JAA05544
	for <speechsc@ietf.org>; Fri, 15 Jul 2005 09:36:44 -0400 (EDT)
Received: from dns1.tilab.com ([163.162.42.4])
	by ietf-mx.ietf.org with esmtp (Exim 4.43) id 1DtQpk-0000fL-Ky
	for speechsc@ietf.org; Fri, 15 Jul 2005 10:06:34 -0400
Received: from iowa2k01a.cselt.it ([163.162.242.201])
	by dns1.cselt.it (PMDF V6.0-025 #38895)
	with ESMTP id <0IJO0004R8BQRT@dns1.cselt.it> for speechsc@ietf.org; Fri,
	15 Jul 2005 15:33:26 +0200 (MEST)
Received: from EXC01C.cselt.it ([163.162.4.218]) by iowa2k01a.cselt.it with
	Microsoft SMTPSVC(6.0.3790.211); Fri, 15 Jul 2005 15:39:46 +0200
Date: Fri, 15 Jul 2005 15:36:18 +0200
From: Baggia Paolo <Paolo.Baggia@LOQUENDO.COM>
Subject: RE: [Speechsc] RECOGNIZE completion-cause 003 and 008
To: David R Oran <oran@cisco.com>, Claudia Daboul <claudia@ibp.de>
Message-id: <E5880434292FCB448F00BDAEE44A60D07DC48F@EXC01C.cselt.it>
MIME-version: 1.0
X-MIMEOLE: Produced By Microsoft MimeOLE V6.00.3790.326
Content-type: text/plain; charset=iso-8859-1
Content-transfer-encoding: quoted-printable
Importance: normal
Priority: normal
Thread-Topic: [Speechsc] RECOGNIZE completion-cause 003 and 008
thread-index: AcWJNwd7g2p5E4lzRGiAF+Nh2RpYOAACswtg
Content-Class: urn:content-classes:message
X-OriginalArrivalTime: 15 Jul 2005 13:39:46.0312 (UTC)
	FILETIME=[AA035880:01C58942]
X-Spam-Score: 0.0 (/)
X-Scan-Signature: 6e922792024732fb1bb6f346e63517e4
Content-Transfer-Encoding: quoted-printable
Cc: speechsc@ietf.org, Baggia Paolo <Paolo.Baggia@LOQUENDO.COM>
X-BeenThere: speechsc@ietf.org
X-Mailman-Version: 2.1.5
Precedence: list
List-Id: Speech Services Control Working Group <speechsc.ietf.org>
List-Unsubscribe: <https://www1.ietf.org/mailman/listinfo/speechsc>,
	<mailto:speechsc-request@ietf.org?subject=unsubscribe>
List-Post: <mailto:speechsc@ietf.org>
List-Help: <mailto:speechsc-request@ietf.org?subject=help>
List-Subscribe: <https://www1.ietf.org/mailman/listinfo/speechsc>,
	<mailto:speechsc-request@ietf.org?subject=subscribe>
Sender: speechsc-bounces@ietf.org
Errors-To: speechsc-bounces@ietf.org

Dear Claudia,

I completely agree with David.

The "008" is related to what VoiceXML 2.0 call:
"maxspeechtimeout" that is defined at:
http://www.w3.org/TR/2004/REC-voicexml20-20040316/#dml6.3.2 as:

"The maximum duration of user speech. If this time elapsed before=20
the user stops speaking, the event "maxspeechtimeout" is thrown.=20
The value is a Time Designation (see Section 6.5).=20
The default duration is platform-dependent."

The "003" is the "timeout" for generating "noinput" event.
The user didn't speak at all.

Paolo Baggia, Loquendo.

-----Original Message-----
From: speechsc-bounces@ietf.org [mailto:speechsc-bounces@ietf.org]On
Behalf Of David R Oran
Sent: Friday, July 15, 2005 2:14 PM
To: Claudia Daboul
Cc: speechsc@ietf.org
Subject: Re: [Speechsc] RECOGNIZE completion-cause 003 and 008



On Jul 15, 2005, at 7:43 AM, Claudia Daboul wrote:

> The spec defines the following completion-causes for the RECOGNIZE =20
> request
>
> 003          recognition-timeout
>                                    RECOGNIZE completed without a match
>                                    due to a recognition-timeout
> and
>
> 008           too-much-speech-timeout
>                                    RECOGNIZE request terminated =20
> because
>                                    speech was too long.
>
> The recognition timeout parameter is explained as follows:
>
> Recognition Timeout
>
>    When recognition is started and there is no match for a certain
>    period of time, the recognizer can send a RECOGNITION-COMPLETE =20
> event
>    to the client and terminate the recognition operation. It is the
>    timer that is started when START-OF-SPEECH event is generated by =20
> the
>    resource and specifies the maximum duration of the utterance. When
>    this timer expires the recognition request would complete with a
>    status code of "008 too-much-speech-timeout". The recognition-
>    timeout header field sets this timeout value. The value is in
>    milliseconds. The value for this field ranges from 0 to MAXTIMEOUT,
>    where MAXTIMEOUT is platform specific. The default value is 10
>    seconds. This header field MAY occur in RECOGNIZE, SET-PARAMS or
>    GET-PARAMS.
>
>
> I am wondering if there is any difference between completion-causes =20
> 003 and 008. Can anybody clarify this?
>
I'll take a whack. 003 says that the recognizer didn't see anything =20
that matched an active grammar before the timer ran out. 008 says the =20
recognizer got an utterance that matched the grammar but the user =20
kept babbling on and didn't stop so the match is probably bogus.

> Claudia
>
> _______________________________________________
> Speechsc mailing list
> Speechsc@ietf.org
> https://www1.ietf.org/mailman/listinfo/speechsc
>

_______________________________________________
Speechsc mailing list
Speechsc@ietf.org
https://www1.ietf.org/mailman/listinfo/speechsc


Gruppo Telecom Italia - Direzione e coordinamento di Telecom Italia =
S.p.A.

=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=
=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=
=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D
CONFIDENTIALITY NOTICE
This message and its attachments are addressed solely to the persons
above and may contain confidential information. If you have received
the message in error, be informed that any use of the content hereof
is prohibited. Please return it immediately to the sender and delete
the message. Should you have any questions, please send an e_mail to=20
MailAdmin@tilab.com. Thank you
=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=
=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=
=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D

_______________________________________________
Speechsc mailing list
Speechsc@ietf.org
https://www1.ietf.org/mailman/listinfo/speechsc



From speechsc-bounces@ietf.org Fri Jul 15 16:30:22 2005
Received: from localhost.localdomain ([127.0.0.1] helo=megatron.ietf.org)
	by megatron.ietf.org with esmtp (Exim 4.32)
	id 1DtWpC-00086m-IZ; Fri, 15 Jul 2005 16:30:22 -0400
Received: from odin.ietf.org ([132.151.1.176] helo=ietf.org)
	by megatron.ietf.org with esmtp (Exim 4.32) id 1DtWpA-00086D-UT
	for speechsc@megatron.ietf.org; Fri, 15 Jul 2005 16:30:20 -0400
Received: from ietf-mx.ietf.org (ietf-mx [132.151.6.1])
	by ietf.org (8.9.1a/8.9.1a) with ESMTP id QAA27784
	for <speechsc@ietf.org>; Fri, 15 Jul 2005 16:30:17 -0400 (EDT)
Received: from [195.222.227.20] (helo=gromit.ibp.de ident=Debian-exim)
	by ietf-mx.ietf.org with esmtp (Exim 4.43) id 1DtXI0-0005df-Gx
	for speechsc@ietf.org; Fri, 15 Jul 2005 17:00:09 -0400
Received: from [195.222.227.24] (helo=[195.222.227.24])
	by gromit.ibp.de with esmtpsa (TLS-1.0:DHE_RSA_AES_256_CBC_SHA:32)
	(Exim 4.50) id 1DtWf0-0006dc-4N; Fri, 15 Jul 2005 22:19:50 +0200
Message-ID: <42D81CF3.2020008@ibp.de>
Date: Fri, 15 Jul 2005 22:30:43 +0200
From: Claudia Daboul <claudia@ibp.de>
User-Agent: Mozilla Thunderbird 1.0 (Macintosh/20041206)
X-Accept-Language: en-us, en
MIME-Version: 1.0
To: David R Oran <oran@cisco.com>
Subject: Re: [Speechsc] RECOGNIZE completion-cause 003 and 008
References: <42D7A17B.6000902@ibp.de>
	<F68CC13F-5C6D-4337-B8E4-D3F3E5C3AD8D@cisco.com>
In-Reply-To: <F68CC13F-5C6D-4337-B8E4-D3F3E5C3AD8D@cisco.com>
Content-Type: text/plain; charset=ISO-8859-1; format=flowed
Content-Transfer-Encoding: 7bit
X-Spam-Score: 0.0 (/)
X-Spam-Score: 0.0 (/)
X-Scan-Signature: 41c17b4b16d1eedaa8395c26e9a251c4
Content-Transfer-Encoding: 7bit
Cc: speechsc@ietf.org
X-BeenThere: speechsc@ietf.org
X-Mailman-Version: 2.1.5
Precedence: list
List-Id: Speech Services Control Working Group <speechsc.ietf.org>
List-Unsubscribe: <https://www1.ietf.org/mailman/listinfo/speechsc>,
	<mailto:speechsc-request@ietf.org?subject=unsubscribe>
List-Post: <mailto:speechsc@ietf.org>
List-Help: <mailto:speechsc-request@ietf.org?subject=help>
List-Subscribe: <https://www1.ietf.org/mailman/listinfo/speechsc>,
	<mailto:speechsc-request@ietf.org?subject=subscribe>
Sender: speechsc-bounces@ietf.org
Errors-To: speechsc-bounces@ietf.org

 > I'll take a whack. 003 says that the recognizer didn't see anything
 > that matched an active grammar before the timer ran out. 008 says the
 > recognizer got an utterance that matched the grammar but the user  kept
 > babbling on and didn't stop so the match is probably bogus.

Do you mean that for example when the grammar understands "yes" or "no" 
and the caller says "yes, but blah blah blah..." I would get 008, but if 
he just says "blah blah blah..." I would get 003?

I think in both cases it would make more sense to return 001 for 
no-match. More generally if the recognizer has determined that the 
utterance could not have matched the grammar, even if the 
recognition-timeout would have been higher, it should return the 
no-match code.
Only if this can't be determined, either because the recognizer doesn't 
process the cut off utterance, or because it determined that the 
utterance did indeed match the grammar until it was cut off, the 003 (or 
008) should be returned.


David R Oran wrote:
> 
> On Jul 15, 2005, at 7:43 AM, Claudia Daboul wrote:
> 
>> The spec defines the following completion-causes for the RECOGNIZE  
>> request
>>
>> 003          recognition-timeout
>>                                    RECOGNIZE completed without a match
>>                                    due to a recognition-timeout
>> and
>>
>> 008           too-much-speech-timeout
>>                                    RECOGNIZE request terminated  because
>>                                    speech was too long.
>>
>> The recognition timeout parameter is explained as follows:
>>
>> Recognition Timeout
>>
>>    When recognition is started and there is no match for a certain
>>    period of time, the recognizer can send a RECOGNITION-COMPLETE  event
>>    to the client and terminate the recognition operation. It is the
>>    timer that is started when START-OF-SPEECH event is generated by  the
>>    resource and specifies the maximum duration of the utterance. When
>>    this timer expires the recognition request would complete with a
>>    status code of "008 too-much-speech-timeout". The recognition-
>>    timeout header field sets this timeout value. The value is in
>>    milliseconds. The value for this field ranges from 0 to MAXTIMEOUT,
>>    where MAXTIMEOUT is platform specific. The default value is 10
>>    seconds. This header field MAY occur in RECOGNIZE, SET-PARAMS or
>>    GET-PARAMS.
>>
>>
>> I am wondering if there is any difference between completion-causes  
>> 003 and 008. Can anybody clarify this?
>>
> I'll take a whack. 003 says that the recognizer didn't see anything  
> that matched an active grammar before the timer ran out. 008 says the  
> recognizer got an utterance that matched the grammar but the user  kept 
> babbling on and didn't stop so the match is probably bogus.
> 
>> Claudia
>>
>> _______________________________________________
>> Speechsc mailing list
>> Speechsc@ietf.org
>> https://www1.ietf.org/mailman/listinfo/speechsc
>>

_______________________________________________
Speechsc mailing list
Speechsc@ietf.org
https://www1.ietf.org/mailman/listinfo/speechsc



From speechsc-bounces@ietf.org Fri Jul 15 16:39:55 2005
Received: from localhost.localdomain ([127.0.0.1] helo=megatron.ietf.org)
	by megatron.ietf.org with esmtp (Exim 4.32)
	id 1DtWyR-0000r9-Hx; Fri, 15 Jul 2005 16:39:55 -0400
Received: from odin.ietf.org ([132.151.1.176] helo=ietf.org)
	by megatron.ietf.org with esmtp (Exim 4.32) id 1DtWyQ-0000ql-10
	for speechsc@megatron.ietf.org; Fri, 15 Jul 2005 16:39:54 -0400
Received: from ietf-mx.ietf.org (ietf-mx [132.151.6.1])
	by ietf.org (8.9.1a/8.9.1a) with ESMTP id QAA02206
	for <speechsc@ietf.org>; Fri, 15 Jul 2005 16:39:50 -0400 (EDT)
Received: from [195.222.227.20] (helo=gromit.ibp.de ident=Debian-exim)
	by ietf-mx.ietf.org with esmtp (Exim 4.43) id 1DtXRH-0007BH-RY
	for speechsc@ietf.org; Fri, 15 Jul 2005 17:09:44 -0400
Received: from [195.222.227.24] (helo=[195.222.227.24])
	by gromit.ibp.de with esmtpsa (TLS-1.0:DHE_RSA_AES_256_CBC_SHA:32)
	(Exim 4.50) id 1DtWoK-0006eV-DF; Fri, 15 Jul 2005 22:29:29 +0200
Message-ID: <42D81F35.7000200@ibp.de>
Date: Fri, 15 Jul 2005 22:40:21 +0200
From: Claudia Daboul <claudia@ibp.de>
User-Agent: Mozilla Thunderbird 1.0 (Macintosh/20041206)
X-Accept-Language: en-us, en
MIME-Version: 1.0
To: Baggia Paolo <Paolo.Baggia@LOQUENDO.COM>
Subject: Re: [Speechsc] RECOGNIZE completion-cause 003 and 008
References: <E5880434292FCB448F00BDAEE44A60D07DC48F@EXC01C.cselt.it>
In-Reply-To: <E5880434292FCB448F00BDAEE44A60D07DC48F@EXC01C.cselt.it>
Content-Type: text/plain; charset=ISO-8859-1; format=flowed
Content-Transfer-Encoding: 7bit
X-Spam-Score: 0.0 (/)
X-Spam-Score: 0.0 (/)
X-Scan-Signature: c83ccb5cc10e751496398f1233ca9c3a
Content-Transfer-Encoding: 7bit
Cc: speechsc@ietf.org, David R Oran <oran@cisco.com>
X-BeenThere: speechsc@ietf.org
X-Mailman-Version: 2.1.5
Precedence: list
List-Id: Speech Services Control Working Group <speechsc.ietf.org>
List-Unsubscribe: <https://www1.ietf.org/mailman/listinfo/speechsc>,
	<mailto:speechsc-request@ietf.org?subject=unsubscribe>
List-Post: <mailto:speechsc@ietf.org>
List-Help: <mailto:speechsc-request@ietf.org?subject=help>
List-Subscribe: <https://www1.ietf.org/mailman/listinfo/speechsc>,
	<mailto:speechsc-request@ietf.org?subject=subscribe>
Sender: speechsc-bounces@ietf.org
Errors-To: speechsc-bounces@ietf.org

 > The "003" is the "timeout" for generating "noinput" event.
 > The user didn't speak at all.

No, in the MRCP v2 spec the no-input-timeout has code "002".

I agree that the recognition-timeout parameter in MRCP v2 corresponds to 
the maxspeechtimeout in VoiceXML 2.0, but in VoiceXML there is only one 
event, namely the maxspeechtimeout event, which can be returned when 
this timeout has elapsed. In contrast it seems that in MRCP there are 
two different possible return codes corresponding to this same timeout.

Baggia Paolo wrote:
> Dear Claudia,
> 
> I completely agree with David.
> 
> The "008" is related to what VoiceXML 2.0 call:
> "maxspeechtimeout" that is defined at:
> http://www.w3.org/TR/2004/REC-voicexml20-20040316/#dml6.3.2 as:
> 
> "The maximum duration of user speech. If this time elapsed before 
> the user stops speaking, the event "maxspeechtimeout" is thrown. 
> The value is a Time Designation (see Section 6.5). 
> The default duration is platform-dependent."
> 
> The "003" is the "timeout" for generating "noinput" event.
> The user didn't speak at all.
> 
> Paolo Baggia, Loquendo.
> 
> -----Original Message-----
> From: speechsc-bounces@ietf.org [mailto:speechsc-bounces@ietf.org]On
> Behalf Of David R Oran
> Sent: Friday, July 15, 2005 2:14 PM
> To: Claudia Daboul
> Cc: speechsc@ietf.org
> Subject: Re: [Speechsc] RECOGNIZE completion-cause 003 and 008
> 
> 
> 
> On Jul 15, 2005, at 7:43 AM, Claudia Daboul wrote:
> 
> 
>>The spec defines the following completion-causes for the RECOGNIZE  
>>request
>>
>>003          recognition-timeout
>>                                   RECOGNIZE completed without a match
>>                                   due to a recognition-timeout
>>and
>>
>>008           too-much-speech-timeout
>>                                   RECOGNIZE request terminated  
>>because
>>                                   speech was too long.
>>
>>The recognition timeout parameter is explained as follows:
>>
>>Recognition Timeout
>>
>>   When recognition is started and there is no match for a certain
>>   period of time, the recognizer can send a RECOGNITION-COMPLETE  
>>event
>>   to the client and terminate the recognition operation. It is the
>>   timer that is started when START-OF-SPEECH event is generated by  
>>the
>>   resource and specifies the maximum duration of the utterance. When
>>   this timer expires the recognition request would complete with a
>>   status code of "008 too-much-speech-timeout". The recognition-
>>   timeout header field sets this timeout value. The value is in
>>   milliseconds. The value for this field ranges from 0 to MAXTIMEOUT,
>>   where MAXTIMEOUT is platform specific. The default value is 10
>>   seconds. This header field MAY occur in RECOGNIZE, SET-PARAMS or
>>   GET-PARAMS.
>>
>>
>>I am wondering if there is any difference between completion-causes  
>>003 and 008. Can anybody clarify this?
>>
> 
> I'll take a whack. 003 says that the recognizer didn't see anything  
> that matched an active grammar before the timer ran out. 008 says the  
> recognizer got an utterance that matched the grammar but the user  
> kept babbling on and didn't stop so the match is probably bogus.
> 
> 
>>Claudia
>>
>>_______________________________________________
>>Speechsc mailing list
>>Speechsc@ietf.org
>>https://www1.ietf.org/mailman/listinfo/speechsc
>>
> 
> 
> _______________________________________________
> Speechsc mailing list
> Speechsc@ietf.org
> https://www1.ietf.org/mailman/listinfo/speechsc
> 
> 
> Gruppo Telecom Italia - Direzione e coordinamento di Telecom Italia S.p.A.
> 
> ====================================================================
> CONFIDENTIALITY NOTICE
> This message and its attachments are addressed solely to the persons
> above and may contain confidential information. If you have received
> the message in error, be informed that any use of the content hereof
> is prohibited. Please return it immediately to the sender and delete
> the message. Should you have any questions, please send an e_mail to 
> MailAdmin@tilab.com. Thank you
> ====================================================================

_______________________________________________
Speechsc mailing list
Speechsc@ietf.org
https://www1.ietf.org/mailman/listinfo/speechsc



From speechsc-bounces@ietf.org Fri Jul 15 19:30:21 2005
Received: from localhost.localdomain ([127.0.0.1] helo=megatron.ietf.org)
	by megatron.ietf.org with esmtp (Exim 4.32)
	id 1DtZdN-0007Vu-5P; Fri, 15 Jul 2005 19:30:21 -0400
Received: from odin.ietf.org ([132.151.1.176] helo=ietf.org)
	by megatron.ietf.org with esmtp (Exim 4.32) id 1DtZdL-0007Vp-Fi
	for speechsc@megatron.ietf.org; Fri, 15 Jul 2005 19:30:19 -0400
Received: from ietf-mx.ietf.org (ietf-mx [132.151.6.1])
	by ietf.org (8.9.1a/8.9.1a) with ESMTP id TAA06871
	for <speechsc@ietf.org>; Fri, 15 Jul 2005 19:30:15 -0400 (EDT)
Received: from sj-iport-5.cisco.com ([171.68.10.87])
	by ietf-mx.ietf.org with esmtp (Exim 4.43) id 1Dta6D-00036h-Ls
	for speechsc@ietf.org; Fri, 15 Jul 2005 20:00:12 -0400
Received: from sj-core-3.cisco.com (171.68.223.137)
	by sj-iport-5.cisco.com with ESMTP; 15 Jul 2005 16:30:07 -0700
X-IronPort-AV: i="3.93,294,1115017200"; 
	d="scan'208"; a="198631267:sNHT28685304"
Received: from vtg-um-e2k6.sj21ad.cisco.com (vtg-um-e2k6.cisco.com
	[171.70.93.77])
	by sj-core-3.cisco.com (8.12.10/8.12.6) with ESMTP id j6FNU56p000708;
	Fri, 15 Jul 2005 16:30:05 -0700 (PDT)
Content-class: urn:content-classes:message
MIME-Version: 1.0
Content-Type: text/plain;
	charset="us-ascii"
Content-Transfer-Encoding: quoted-printable
X-MimeOLE: Produced By Microsoft Exchange V6.5.7226.0
Subject: RE: [Speechsc] RECOGNIZE completion-cause 003 and 008
Date: Fri, 15 Jul 2005 16:30:04 -0700
Message-ID: <03772D1EC8DE624A863058C75874A75C1E008D@vtg-um-e2k6.sj21ad.cisco.com>
Thread-Topic: [Speechsc] RECOGNIZE completion-cause 003 and 008
Thread-Index: AcWJfJK+Q2TMksjVTdqxLJZi3ZhTxgAF6Bvw
From: "Shanmugham, Saravanan" <sarvi@cisco.com>
To: "Claudia Daboul" <claudia@ibp.de>, "David R Oran" <oran@cisco.com>
X-Spam-Score: 0.0 (/)
X-Scan-Signature: 5ebbf074524e58e662bc8209a6235027
Content-Transfer-Encoding: quoted-printable
Cc: speechsc@ietf.org
X-BeenThere: speechsc@ietf.org
X-Mailman-Version: 2.1.5
Precedence: list
List-Id: Speech Services Control Working Group <speechsc.ietf.org>
List-Unsubscribe: <https://www1.ietf.org/mailman/listinfo/speechsc>,
	<mailto:speechsc-request@ietf.org?subject=unsubscribe>
List-Post: <mailto:speechsc@ietf.org>
List-Help: <mailto:speechsc-request@ietf.org?subject=help>
List-Subscribe: <https://www1.ietf.org/mailman/listinfo/speechsc>,
	<mailto:speechsc-request@ietf.org?subject=subscribe>
Sender: speechsc-bounces@ietf.org
Errors-To: speechsc-bounces@ietf.org



No-Match - means there was no match at all
Recognition-timeout - means there was a partial match but the user
stopped speaking in the middle and a mid speech timeout happenned.=20
Too-much-speech-timeout means the user continued to speak for too long
but everything that was spoken until the timeout does match a grammar.

Sarvi

     -----Original Message-----
     From: speechsc-bounces@ietf.org=20
     [mailto:speechsc-bounces@ietf.org] On Behalf Of Claudia Daboul
     Sent: Friday, July 15, 2005 1:31 PM
     To: David R Oran
     Cc: speechsc@ietf.org
     Subject: Re: [Speechsc] RECOGNIZE completion-cause 003 and 008
    =20
      > I'll take a whack. 003 says that the recognizer didn't=20
     see anything  > that matched an active grammar before the=20
     timer ran out. 008 says the  > recognizer got an utterance=20
     that matched the grammar but the user  kept  > babbling on=20
     and didn't stop so the match is probably bogus.
    =20
     Do you mean that for example when the grammar understands=20
     "yes" or "no"=20
     and the caller says "yes, but blah blah blah..." I would=20
     get 008, but if he just says "blah blah blah..." I would get 003?
    =20
     I think in both cases it would make more sense to return=20
     001 for no-match. More generally if the recognizer has=20
     determined that the utterance could not have matched the=20
     grammar, even if the recognition-timeout would have been=20
     higher, it should return the no-match code.
     Only if this can't be determined, either because the=20
     recognizer doesn't process the cut off utterance, or=20
     because it determined that the utterance did indeed match=20
     the grammar until it was cut off, the 003 (or
     008) should be returned.
    =20
    =20
     David R Oran wrote:
     >=20
     > On Jul 15, 2005, at 7:43 AM, Claudia Daboul wrote:
     >=20
     >> The spec defines the following completion-causes for=20
     the RECOGNIZE=20
     >> request
     >>
     >> 003          recognition-timeout
     >>                                    RECOGNIZE completed=20
     without a match
     >>                                    due to a=20
     recognition-timeout and
     >>
     >> 008           too-much-speech-timeout
     >>                                    RECOGNIZE request=20
     terminated  because
     >>                                    speech was too long.
     >>
     >> The recognition timeout parameter is explained as follows:
     >>
     >> Recognition Timeout
     >>
     >>    When recognition is started and there is no match=20
     for a certain
     >>    period of time, the recognizer can send a=20
     RECOGNITION-COMPLETE  event
     >>    to the client and terminate the recognition=20
     operation. It is the
     >>    timer that is started when START-OF-SPEECH event is=20
     generated by  the
     >>    resource and specifies the maximum duration of the=20
     utterance. When
     >>    this timer expires the recognition request would=20
     complete with a
     >>    status code of "008 too-much-speech-timeout". The=20
     recognition-
     >>    timeout header field sets this timeout value. The value is in
     >>    milliseconds. The value for this field ranges from 0=20
     to MAXTIMEOUT,
     >>    where MAXTIMEOUT is platform specific. The default=20
     value is 10
     >>    seconds. This header field MAY occur in RECOGNIZE,=20
     SET-PARAMS or
     >>    GET-PARAMS.
     >>
     >>
     >> I am wondering if there is any difference between=20
     completion-causes
     >> 003 and 008. Can anybody clarify this?
     >>
     > I'll take a whack. 003 says that the recognizer didn't=20
     see anything=20
     > that matched an active grammar before the timer ran out.=20
     008 says the=20
     > recognizer got an utterance that matched the grammar but=20
     the user =20
     > kept babbling on and didn't stop so the match is probably bogus.
     >=20
     >> Claudia
     >>
     >> _______________________________________________
     >> Speechsc mailing list
     >> Speechsc@ietf.org
     >> https://www1.ietf.org/mailman/listinfo/speechsc
     >>
    =20
     _______________________________________________
     Speechsc mailing list
     Speechsc@ietf.org
     https://www1.ietf.org/mailman/listinfo/speechsc
    =20

_______________________________________________
Speechsc mailing list
Speechsc@ietf.org
https://www1.ietf.org/mailman/listinfo/speechsc



From speechsc-bounces@ietf.org Mon Jul 18 07:23:39 2005
Received: from localhost.localdomain ([127.0.0.1] helo=megatron.ietf.org)
	by megatron.ietf.org with esmtp (Exim 4.32)
	id 1DuTil-0000i3-Mt; Mon, 18 Jul 2005 07:23:39 -0400
Received: from odin.ietf.org ([132.151.1.176] helo=ietf.org)
	by megatron.ietf.org with esmtp (Exim 4.32) id 1DuTik-0000hp-2N
	for speechsc@megatron.ietf.org; Mon, 18 Jul 2005 07:23:39 -0400
Received: from ietf-mx.ietf.org (ietf-mx [132.151.6.1])
	by ietf.org (8.9.1a/8.9.1a) with ESMTP id HAA01738
	for <speechsc@ietf.org>; Mon, 18 Jul 2005 07:23:37 -0400 (EDT)
Received: from pb-exchcon2.scansoft.com ([199.4.160.64])
	by ietf-mx.ietf.org with esmtp (Exim 4.43) id 1DuUC7-0000P7-TT
	for speechsc@ietf.org; Mon, 18 Jul 2005 07:54:02 -0400
Received: by pb-exchcon2.pb.scansoft.com with Internet Mail Service
	(5.5.2658.27) id <PADWBRWJ>; Mon, 18 Jul 2005 07:23:19 -0400
Message-ID: <BBF29C9B95E52E4DB5C29A0ACC94E83B016AA0F4@ac-exch1.eu.scansoft.com>
From: "Reifenrath, Klaus" <Klaus.Reifenrath@Scansoft.com>
To: "'claudia@ibp.de'" <claudia@ibp.de>
Subject: Re: [Speechsc] RECOGNIZE completion-cause 003 and 008
Date: Mon, 18 Jul 2005 07:11:18 -0400
MIME-Version: 1.0
X-Mailer: Internet Mail Service (5.5.2658.27)
Content-Type: text/plain
X-Spam-Score: 0.0 (/)
X-Scan-Signature: 3d7f2f6612d734db849efa86ea692407
Cc: "'speechsc@ietf.org'" <speechsc@ietf.org>, "NUAN-Forgues,
	Pierre" <forgues@nuance.com>, "'oran@cisco.com'" <oran@cisco.com>,
	"'Paolo.Baggia@LOQUENDO.COM'" <Paolo.Baggia@LOQUENDO.COM>
X-BeenThere: speechsc@ietf.org
X-Mailman-Version: 2.1.5
Precedence: list
List-Id: Speech Services Control Working Group <speechsc.ietf.org>
List-Unsubscribe: <https://www1.ietf.org/mailman/listinfo/speechsc>,
	<mailto:speechsc-request@ietf.org?subject=unsubscribe>
List-Post: <mailto:speechsc@ietf.org>
List-Help: <mailto:speechsc-request@ietf.org?subject=help>
List-Subscribe: <https://www1.ietf.org/mailman/listinfo/speechsc>,
	<mailto:speechsc-request@ietf.org?subject=subscribe>
Sender: speechsc-bounces@ietf.org
Errors-To: speechsc-bounces@ietf.org

Hi Claudia,

below is an old e-mail thread on this topic.

Regards,
Klaus

-----Original Message-----
From: Pierre Forgues [mailto:forgues@nuance.com] 
Sent: Mittwoch, 28. April 2004 14:32
To: Reifenrath, Klaus
Cc: speechsc@ietf.org
Subject: RE: [Speechsc] Recognition Timeout

OK, agreed!  This would leave the choice to the recognition engine.  The
second option would be used if there are worthwhile results to return.  

Practically speaking, we have not seen this as a common case.
Pierre

-----Original Message-----
From: speechsc-admin@ietf.org [mailto:speechsc-admin@ietf.org] On Behalf Of
Reifenrath, Klaus
Sent: 28 avril, 2004 04:46
To: Pierre Forgues
Cc: speechsc@ietf.org
Subject: RE: [Speechsc] Recognition Timeout

Hi Pierre!

Why should the available result be discarded right away? This would only
elongate the call.
MRCP should allow two ways of handling this situation:
1) immediately return with a 008 "too-much-speech"
2) return 00x "success-maxtime" plus a recognition result In 2) the
recognizer treats the situation, when the maximum length of speech has been
reached, as if it had detected end-of-speech.

Klaus

-----Original Message-----
From: Pierre Forgues [mailto:forgues@nuance.com]
Sent: Dienstag, 27. April 2004 20:49
To: Reifenrath, Klaus; speechsc@ietf.org
Subject: RE: [Speechsc] Recognition Timeout


We return immediately with a 008 "too-much-speech-timeout".  How about just
clarifying the documentation to reflect this?

Pierre

-----Original Message-----
From: Reifenrath, Klaus [mailto:Klaus.Reifenrath@Scansoft.com]
Sent: 27 avril, 2004 04:13
To: speechsc@ietf.org
Cc: Pierre Forgues
Subject: RE: [Speechsc] Recognition Timeout

Hi group!

In case the callers utterance exceeds the maximum duration (given by the
recognition-timeout) and the recognizer has a complete match, but the caller
did not stop speaking (speech-complete-timeout did not elapse) the current
spec allows the following alternatives
1) continue recognition until speech-complete-timeout indicates that the
caller stopped speaking;
2) return completion cause "success" and a recognition result (as if
speech-complete-timeout did elapse). 
In the former case the recognition timeout is only applied to utterances
that have no match. But this seems not to go with "maxspeechtimeout" of
VoiceXML. 
In the later case the client does not get informed that only part of the
callers utterance was recognized.
 
I suggest to return a recognition result and a new completion cause
"success-maxtime" (compare recorder completion causes in 10.4). Either
"too-much-speech-timeout" or "recognition-timeout" would become obsolete. 

Klaus

-----Original Message-----
From: Pierre Forgues [mailto:forgues@nuance.com]
Sent: Montag, 19. April 2004 19:04
To: Reifenrath, Klaus
Cc: speechsc@ietf.org
Subject: RE: [Speechsc] Recognition Timeout


Klaus,

In our case, when this timer expires, we never have a match.  The 003
recognition-timeout and 008 too-much-speech-timeout are equivalent and
ambiguous it seems.  In this case, we return a 008 too-much-speech-timeout
since in our case we don't produce a match.

Pierre

-----Original Message-----
From: Reifenrath, Klaus [mailto:Klaus.Reifenrath@Scansoft.com]
Sent: 19 avril, 2004 11:50
To: Pierre Forgues
Cc: speechsc@ietf.org
Subject: RE: [Speechsc] Recognition Timeout

Hi Pierre,

this is inline with my understanding of the spec.

If the timer expires and there is no match, completion code 003
(recognition-timeout) should be returned, right?
If the timer expires and there is a match, completion code 008
(too-much-speech-timeout) along with a recognition result is returned? 

Klaus


-----Original Message-----
From: Pierre Forgues [mailto:forgues@nuance.com]
Sent: Montag, 19. April 2004 16:36
To: Reifenrath, Klaus; speechsc@ietf.org
Subject: RE: [Speechsc] Recognition Timeout


Hi Klaus,

The definition is indeed a bit ambiguous.  It is the timer that is started
when START-OF-SPEECH event is generated by the server.  It is indeed used to
specify the maximum duration of an utterance.  

Pierre
-----Original Message-----
From: speechsc-admin@ietf.org [mailto:speechsc-admin@ietf.org] On Behalf Of
Reifenrath, Klaus
Sent: 19 avril, 2004 07:51
To: 'speechsc@ietf.org'
Subject: [Speechsc] Recognition Timeout

The definition of the recognition-timeout header field could be more precise
about when the timer is started. Is the timer started when START-OF-SPEECH
is generated by the server or when RECOGNIZE/RECOGNITION-START-TIMERS is
send by the client? In the former case the parameter specifies the maximum
period of time for the utterance and maps nicely to the property
maxspeechtimeout in VoiceXML or the attribute babbletimeout in SALT. In the
later case it would be equivalent to the attribute maxtimeout in SALT.

Klaus

_______________________________________________
Speechsc mailing list
Speechsc@ietf.org
https://www1.ietf.org/mailman/listinfo/speechsc

_______________________________________________
Speechsc mailing list
Speechsc@ietf.org
https://www1.ietf.org/mailman/listinfo/speechsc

_______________________________________________
Speechsc mailing list
Speechsc@ietf.org
https://www1.ietf.org/mailman/listinfo/speechsc



From speechsc-bounces@ietf.org Tue Jul 19 06:45:28 2005
Received: from localhost.localdomain ([127.0.0.1] helo=megatron.ietf.org)
	by megatron.ietf.org with esmtp (Exim 4.32)
	id 1DupbM-0002Sc-Mb; Tue, 19 Jul 2005 06:45:28 -0400
Received: from odin.ietf.org ([132.151.1.176] helo=ietf.org)
	by megatron.ietf.org with esmtp (Exim 4.32) id 1DupbK-0002SM-Go
	for speechsc@megatron.ietf.org; Tue, 19 Jul 2005 06:45:28 -0400
Received: from ietf-mx.ietf.org (ietf-mx [132.151.6.1])
	by ietf.org (8.9.1a/8.9.1a) with ESMTP id GAA00028
	for <speechsc@ietf.org>; Tue, 19 Jul 2005 06:45:23 -0400 (EDT)
Received: from [195.222.227.20] (helo=gromit.ibp.de ident=Debian-exim)
	by ietf-mx.ietf.org with esmtp (Exim 4.43) id 1Duq4u-0008Pc-JY
	for speechsc@ietf.org; Tue, 19 Jul 2005 07:16:02 -0400
Received: from [195.222.227.22] (helo=[195.222.227.22])
	by gromit.ibp.de with esmtp (Exim 4.50)
	id 1DupQh-0002dU-Bc; Tue, 19 Jul 2005 12:34:28 +0200
Message-ID: <42DCDA3A.4080904@ibp.de>
Date: Tue, 19 Jul 2005 12:47:22 +0200
From: Claudia Daboul <claudia@ibp.de>
User-Agent: Mozilla Thunderbird 0.8 (Windows/20040913)
X-Accept-Language: en-us, en
MIME-Version: 1.0
To: "Shanmugham, Saravanan" <sarvi@cisco.com>,
	"Reifenrath, Klaus" <Klaus.Reifenrath@Scansoft.com>
Subject: Re: [Speechsc] RECOGNIZE completion-cause 003 and 008
References: <03772D1EC8DE624A863058C75874A75C1E008D@vtg-um-e2k6.sj21ad.cisco.com>
In-Reply-To: <03772D1EC8DE624A863058C75874A75C1E008D@vtg-um-e2k6.sj21ad.cisco.com>
X-Enigmail-Version: 0.86.1.0
X-Enigmail-Supports: pgp-inline, pgp-mime
Content-Type: text/plain; charset=ISO-8859-1; format=flowed
Content-Transfer-Encoding: 7bit
X-Spam-Score: 0.0 (/)
X-Spam-Score: 0.0 (/)
X-Scan-Signature: 36b1f8810cb91289d885dc8ab4fc8172
Content-Transfer-Encoding: 7bit
Cc: speechsc@ietf.org, David R Oran <oran@cisco.com>
X-BeenThere: speechsc@ietf.org
X-Mailman-Version: 2.1.5
Precedence: list
List-Id: Speech Services Control Working Group <speechsc.ietf.org>
List-Unsubscribe: <https://www1.ietf.org/mailman/listinfo/speechsc>,
	<mailto:speechsc-request@ietf.org?subject=unsubscribe>
List-Post: <mailto:speechsc@ietf.org>
List-Help: <mailto:speechsc-request@ietf.org?subject=help>
List-Subscribe: <https://www1.ietf.org/mailman/listinfo/speechsc>,
	<mailto:speechsc-request@ietf.org?subject=subscribe>
Sender: speechsc-bounces@ietf.org
Errors-To: speechsc-bounces@ietf.org

 > Recognition-timeout - means there was a partial match but the user
 > stopped speaking in the middle and a mid speech timeout happenned.

But in this case it would be the speech-incomplete-timeout that would 
apply and according to the definition of that header-field a nomatch 
event should be returned when it expires and there is only a partial 
(incomplete) match of the grammar.

 > Too-much-speech-timeout means the user continued to speak for too long
 > but everything that was spoken until the timeout does match a grammar.
 >
I think this would be what Klaus calls the "success-maxtime" event, 
(which is not yet part of the spec) and which would come together with a 
valid recognition result, or did I get that wrong Klaus?



Shanmugham, Saravanan wrote:
> 
> No-Match - means there was no match at all
> Recognition-timeout - means there was a partial match but the user
> stopped speaking in the middle and a mid speech timeout happenned. 
> Too-much-speech-timeout means the user continued to speak for too long
> but everything that was spoken until the timeout does match a grammar.
> 
> Sarvi
> 
>      -----Original Message-----
>      From: speechsc-bounces@ietf.org 
>      [mailto:speechsc-bounces@ietf.org] On Behalf Of Claudia Daboul
>      Sent: Friday, July 15, 2005 1:31 PM
>      To: David R Oran
>      Cc: speechsc@ietf.org
>      Subject: Re: [Speechsc] RECOGNIZE completion-cause 003 and 008
>      
>       > I'll take a whack. 003 says that the recognizer didn't 
>      see anything  > that matched an active grammar before the 
>      timer ran out. 008 says the  > recognizer got an utterance 
>      that matched the grammar but the user  kept  > babbling on 
>      and didn't stop so the match is probably bogus.
>      
>      Do you mean that for example when the grammar understands 
>      "yes" or "no" 
>      and the caller says "yes, but blah blah blah..." I would 
>      get 008, but if he just says "blah blah blah..." I would get 003?
>      
>      I think in both cases it would make more sense to return 
>      001 for no-match. More generally if the recognizer has 
>      determined that the utterance could not have matched the 
>      grammar, even if the recognition-timeout would have been 
>      higher, it should return the no-match code.
>      Only if this can't be determined, either because the 
>      recognizer doesn't process the cut off utterance, or 
>      because it determined that the utterance did indeed match 
>      the grammar until it was cut off, the 003 (or
>      008) should be returned.
>      
>      
>      David R Oran wrote:
>      > 
>      > On Jul 15, 2005, at 7:43 AM, Claudia Daboul wrote:
>      > 
>      >> The spec defines the following completion-causes for 
>      the RECOGNIZE 
>      >> request
>      >>
>      >> 003          recognition-timeout
>      >>                                    RECOGNIZE completed 
>      without a match
>      >>                                    due to a 
>      recognition-timeout and
>      >>
>      >> 008           too-much-speech-timeout
>      >>                                    RECOGNIZE request 
>      terminated  because
>      >>                                    speech was too long.
>      >>
>      >> The recognition timeout parameter is explained as follows:
>      >>
>      >> Recognition Timeout
>      >>
>      >>    When recognition is started and there is no match 
>      for a certain
>      >>    period of time, the recognizer can send a 
>      RECOGNITION-COMPLETE  event
>      >>    to the client and terminate the recognition 
>      operation. It is the
>      >>    timer that is started when START-OF-SPEECH event is 
>      generated by  the
>      >>    resource and specifies the maximum duration of the 
>      utterance. When
>      >>    this timer expires the recognition request would 
>      complete with a
>      >>    status code of "008 too-much-speech-timeout". The 
>      recognition-
>      >>    timeout header field sets this timeout value. The value is in
>      >>    milliseconds. The value for this field ranges from 0 
>      to MAXTIMEOUT,
>      >>    where MAXTIMEOUT is platform specific. The default 
>      value is 10
>      >>    seconds. This header field MAY occur in RECOGNIZE, 
>      SET-PARAMS or
>      >>    GET-PARAMS.
>      >>
>      >>
>      >> I am wondering if there is any difference between 
>      completion-causes
>      >> 003 and 008. Can anybody clarify this?
>      >>
>      > I'll take a whack. 003 says that the recognizer didn't 
>      see anything 
>      > that matched an active grammar before the timer ran out. 
>      008 says the 
>      > recognizer got an utterance that matched the grammar but 
>      the user  
>      > kept babbling on and didn't stop so the match is probably bogus.
>      > 
>      >> Claudia
>      >>
>      >> _______________________________________________
>      >> Speechsc mailing list
>      >> Speechsc@ietf.org
>      >> https://www1.ietf.org/mailman/listinfo/speechsc
>      >>
>      
>      _______________________________________________
>      Speechsc mailing list
>      Speechsc@ietf.org
>      https://www1.ietf.org/mailman/listinfo/speechsc
>      

-- 
Dr. Claudia Daboul
IBP UK (Limited)
Luetkensallee 19
22041 Hamburg
Tel.: +49 (40) 3172541
Fax.: +49 (40) 3172547
E-mail: claudia@ibp.de

_______________________________________________
Speechsc mailing list
Speechsc@ietf.org
https://www1.ietf.org/mailman/listinfo/speechsc



From speechsc-bounces@ietf.org Sun Jul 24 21:10:29 2005
Received: from localhost.localdomain ([127.0.0.1] helo=megatron.ietf.org)
	by megatron.ietf.org with esmtp (Exim 4.32)
	id 1DwrUD-0003QZ-6j; Sun, 24 Jul 2005 21:10:29 -0400
Received: from odin.ietf.org ([132.151.1.176] helo=ietf.org)
	by megatron.ietf.org with esmtp (Exim 4.32) id 1DwrUB-0003Px-Oy
	for speechsc@megatron.ietf.org; Sun, 24 Jul 2005 21:10:27 -0400
Received: from ietf-mx.ietf.org (ietf-mx [132.151.6.1])
	by ietf.org (8.9.1a/8.9.1a) with ESMTP id VAA09134
	for <speechsc@ietf.org>; Sun, 24 Jul 2005 21:10:25 -0400 (EDT)
Received: from salvelinus.brooktrout.com ([204.176.205.6])
	by ietf-mx.ietf.org with esmtp (Exim 4.43) id 1Dwryv-0003si-Oz
	for speechsc@ietf.org; Sun, 24 Jul 2005 21:42:15 -0400
Received: from ATLANTIS.Brooktrout.com (oceans11.brooktrout.com
	[204.176.75.121])
	by salvelinus.brooktrout.com (8.12.5/8.12.5) with ESMTP id
	j6P15gvN008211
	for <speechsc@ietf.org>; Sun, 24 Jul 2005 21:05:42 -0400 (EDT)
Content-class: urn:content-classes:message
MIME-Version: 1.0
Content-Type: text/plain;
	charset="US-ASCII"
Content-Transfer-Encoding: quoted-printable
X-MimeOLE: Produced By Microsoft Exchange V6.5.7226.0
Date: Sun, 24 Jul 2005 21:05:42 -0400
Message-ID: <330A23D8336C0346B5C1A5BB196666476E8FA1@ATLANTIS.Brooktrout.com>
Thread-Topic: Status of audio streaming for IETF 63.
Thread-Index: AcWHRx3ywkAfDMRLQbaXQ4JNMtn3UgJRVDHA
From: "Eric Burger" <eburger@brooktrout.com>
To: <speechsc@ietf.org>
X-Spam-Score: 0.0 (/)
X-Scan-Signature: 5a9a1bd6c2d06a21d748b7d0070ddcb8
Content-Transfer-Encoding: quoted-printable
Subject: [Speechsc] FW: Status of audio streaming for IETF 63.
X-BeenThere: speechsc@ietf.org
X-Mailman-Version: 2.1.5
Precedence: list
List-Id: Speech Services Control Working Group <speechsc.ietf.org>
List-Unsubscribe: <https://www1.ietf.org/mailman/listinfo/speechsc>,
	<mailto:speechsc-request@ietf.org?subject=unsubscribe>
List-Post: <mailto:speechsc@ietf.org>
List-Help: <mailto:speechsc-request@ietf.org?subject=help>
List-Subscribe: <https://www1.ietf.org/mailman/listinfo/speechsc>,
	<mailto:speechsc-request@ietf.org?subject=subscribe>
Sender: speechsc-bounces@ietf.org
Errors-To: speechsc-bounces@ietf.org



-----Original Message-----
From: ietf-bounces@ietf.org [mailto:joelja@darkwing.uoregon.edu]=20
Sent: Tuesday, July 12, 2005 9:05 PM
To: ietf@ietf.org
Subject: Status of audio streaming for IETF 63.

The new streaming effort continues for IETF 63. All eight parallel
tracks=20
as well as the plenaries will be covered. It is our hope that this
effort=20
will continue to provide useful timely and accessible access to the=20
proceedings of the IETF as they happen.

An internet draft (revised since IETF 62) desscribing what the project=20
intends to accomplish, and the efforts up to this point is available at:

http://www.ietf.org/internet-drafts/draft-jaeggli-ietftv-ng-01.txt

Streams are to be delivered as 64Kb/s unicast-http-streamed mp3 audio, a

popular and relativly standard way to deliver internet radio. Most
platforms=20
should have a client immediatly available (windows media player,
quicktime,=20
real, winamp, vlc, mplayer, zinf, etc) capable of playing back the
stream.

Streaming begins July 31. See the page located at:

http://videolab.uoregon.edu/events/ietf/ietf63.html

which will be updated as additional information including the final=20
schedule, becomes available.

To test your client against a currently active stream of the same type
prior to=20
the the meeting, visit the webpage for instructions.

regards
joelja


--=20
------------------------------------------------------------------------
--
Joel Jaeggli  	       Unix Consulting
joelja@darkwing.uoregon.edu
GPG Key Fingerprint:     5C6E 0104 BAF0 40B0 5BD3 C38B F000 35AB B67F
56B2


_______________________________________________
Ietf mailing list
Ietf@ietf.org
https://www1.ietf.org/mailman/listinfo/ietf

_______________________________________________
Speechsc mailing list
Speechsc@ietf.org
https://www1.ietf.org/mailman/listinfo/speechsc



From speechsc-bounces@ietf.org Mon Jul 25 01:40:43 2005
Received: from localhost.localdomain ([127.0.0.1] helo=megatron.ietf.org)
	by megatron.ietf.org with esmtp (Exim 4.32)
	id 1Dwvhj-0008Nb-NX; Mon, 25 Jul 2005 01:40:43 -0400
Received: from odin.ietf.org ([132.151.1.176] helo=ietf.org)
	by megatron.ietf.org with esmtp (Exim 4.32) id 1Dwvhh-0008NT-2b
	for speechsc@megatron.ietf.org; Mon, 25 Jul 2005 01:40:41 -0400
Received: from ietf-mx.ietf.org (ietf-mx [132.151.6.1])
	by ietf.org (8.9.1a/8.9.1a) with ESMTP id BAA26403
	for <speechsc@ietf.org>; Mon, 25 Jul 2005 01:40:40 -0400 (EDT)
Received: from sj-iport-5.cisco.com ([171.68.10.87])
	by ietf-mx.ietf.org with esmtp (Exim 4.43) id 1DwwCT-0003Hf-KV
	for speechsc@ietf.org; Mon, 25 Jul 2005 02:12:30 -0400
Received: from sj-core-3.cisco.com (171.68.223.137)
	by sj-iport-5.cisco.com with ESMTP; 24 Jul 2005 22:40:30 -0700
X-IronPort-AV: i="3.95,139,1120460400"; 
	d="scan'208"; a="200300030:sNHT32779092"
Received: from vtg-um-e2k6.sj21ad.cisco.com (vtg-um-e2k6.cisco.com
	[171.70.93.77])
	by sj-core-3.cisco.com (8.12.10/8.12.6) with ESMTP id j6P5eQ6p019736;
	Sun, 24 Jul 2005 22:40:27 -0700 (PDT)
Content-class: urn:content-classes:message
MIME-Version: 1.0
Content-Type: text/plain;
	charset="us-ascii"
Content-Transfer-Encoding: quoted-printable
X-MimeOLE: Produced By Microsoft Exchange V6.5.7226.0
Subject: RE: [Speechsc] RECOGNIZE completion-cause 003 and 008
Date: Sun, 24 Jul 2005 22:40:26 -0700
Message-ID: <03772D1EC8DE624A863058C75874A75C1E06B8@vtg-um-e2k6.sj21ad.cisco.com>
Thread-Topic: [Speechsc] RECOGNIZE completion-cause 003 and 008
Thread-Index: AcWMTzth0TjduA1uRtii3wcH9YybrgEipXOQ
From: "Shanmugham, Saravanan" <sarvi@cisco.com>
To: "Claudia Daboul" <claudia@ibp.de>,
	"Reifenrath, Klaus" <Klaus.Reifenrath@Scansoft.com>
X-Spam-Score: 0.0 (/)
X-Scan-Signature: 2bf730a014b318fd3efd65b39b48818c
Content-Transfer-Encoding: quoted-printable
Cc: speechsc@ietf.org, David R Oran <oran@cisco.com>
X-BeenThere: speechsc@ietf.org
X-Mailman-Version: 2.1.5
Precedence: list
List-Id: Speech Services Control Working Group <speechsc.ietf.org>
List-Unsubscribe: <https://www1.ietf.org/mailman/listinfo/speechsc>,
	<mailto:speechsc-request@ietf.org?subject=unsubscribe>
List-Post: <mailto:speechsc@ietf.org>
List-Help: <mailto:speechsc-request@ietf.org?subject=help>
List-Subscribe: <https://www1.ietf.org/mailman/listinfo/speechsc>,
	<mailto:speechsc-request@ietf.org?subject=subscribe>
Sender: speechsc-bounces@ietf.org
Errors-To: speechsc-bounces@ietf.org

I think see the problem.


I have made the following changes to the Recognizer Completion-Cause.

Changed speech-incomplete-timeout to "success-maxtime" as suggested by
Klaus, and meaning that there was a full match but the user continue to
speak and the recongition timeout expired.

Added a new Completion-Cause "013 partial-match" which is returned if
"Recognition timeout expired before there was a full match. But whatever
was spoken till that point was a partial match to one or more grammars."

I then updated Recognition-Timeout to address another problem relating
to timeouts of hotword RECOGNIZE operations.  "RECOGNIZE in hotword mode
completed without a match due to a recognition-timeout".

Thanks,
Sarvi

=20

     -----Original Message-----
     From: Claudia Daboul [mailto:claudia@ibp.de]=20
     Sent: Tuesday, July 19, 2005 3:47 AM
     To: Shanmugham, Saravanan; Reifenrath, Klaus
     Cc: David R Oran; speechsc@ietf.org
     Subject: Re: [Speechsc] RECOGNIZE completion-cause 003 and 008
    =20
      > Recognition-timeout - means there was a partial match=20
     but the user  > stopped speaking in the middle and a mid=20
     speech timeout happenned.
    =20
     But in this case it would be the speech-incomplete-timeout=20
     that would apply and according to the definition of that=20
     header-field a nomatch event should be returned when it=20
     expires and there is only a partial
     (incomplete) match of the grammar.
    =20
      > Too-much-speech-timeout means the user continued to=20
     speak for too long  > but everything that was spoken until=20
     the timeout does match a grammar.
      >
     I think this would be what Klaus calls the=20
     "success-maxtime" event, (which is not yet part of the=20
     spec) and which would come together with a valid=20
     recognition result, or did I get that wrong Klaus?
    =20
    =20
    =20
     Shanmugham, Saravanan wrote:
     >=20
     > No-Match - means there was no match at all=20
     Recognition-timeout - means=20
     > there was a partial match but the user stopped speaking=20
     in the middle=20
     > and a mid speech timeout happenned.
     > Too-much-speech-timeout means the user continued to=20
     speak for too long=20
     > but everything that was spoken until the timeout does=20
     match a grammar.
     >=20
     > Sarvi
     >=20
     >      -----Original Message-----
     >      From: speechsc-bounces@ietf.org=20
     >      [mailto:speechsc-bounces@ietf.org] On Behalf Of=20
     Claudia Daboul
     >      Sent: Friday, July 15, 2005 1:31 PM
     >      To: David R Oran
     >      Cc: speechsc@ietf.org
     >      Subject: Re: [Speechsc] RECOGNIZE completion-cause=20
     003 and 008
     >     =20
     >       > I'll take a whack. 003 says that the recognizer didn't=20
     >      see anything  > that matched an active grammar before the=20
     >      timer ran out. 008 says the  > recognizer got an utterance=20
     >      that matched the grammar but the user  kept  > babbling on=20
     >      and didn't stop so the match is probably bogus.
     >     =20
     >      Do you mean that for example when the grammar understands=20
     >      "yes" or "no"=20
     >      and the caller says "yes, but blah blah blah..." I would=20
     >      get 008, but if he just says "blah blah blah..." I=20
     would get 003?
     >     =20
     >      I think in both cases it would make more sense to return=20
     >      001 for no-match. More generally if the recognizer has=20
     >      determined that the utterance could not have matched the=20
     >      grammar, even if the recognition-timeout would have been=20
     >      higher, it should return the no-match code.
     >      Only if this can't be determined, either because the=20
     >      recognizer doesn't process the cut off utterance, or=20
     >      because it determined that the utterance did indeed match=20
     >      the grammar until it was cut off, the 003 (or
     >      008) should be returned.
     >     =20
     >     =20
     >      David R Oran wrote:
     >      >=20
     >      > On Jul 15, 2005, at 7:43 AM, Claudia Daboul wrote:
     >      >=20
     >      >> The spec defines the following completion-causes for=20
     >      the RECOGNIZE=20
     >      >> request
     >      >>
     >      >> 003          recognition-timeout
     >      >>                                    RECOGNIZE completed=20
     >      without a match
     >      >>                                    due to a=20
     >      recognition-timeout and
     >      >>
     >      >> 008           too-much-speech-timeout
     >      >>                                    RECOGNIZE request=20
     >      terminated  because
     >      >>                                    speech was too long.
     >      >>
     >      >> The recognition timeout parameter is explained=20
     as follows:
     >      >>
     >      >> Recognition Timeout
     >      >>
     >      >>    When recognition is started and there is no match=20
     >      for a certain
     >      >>    period of time, the recognizer can send a=20
     >      RECOGNITION-COMPLETE  event
     >      >>    to the client and terminate the recognition=20
     >      operation. It is the
     >      >>    timer that is started when START-OF-SPEECH event is=20
     >      generated by  the
     >      >>    resource and specifies the maximum duration of the=20
     >      utterance. When
     >      >>    this timer expires the recognition request would=20
     >      complete with a
     >      >>    status code of "008 too-much-speech-timeout". The=20
     >      recognition-
     >      >>    timeout header field sets this timeout value.=20
     The value is in
     >      >>    milliseconds. The value for this field ranges from 0=20
     >      to MAXTIMEOUT,
     >      >>    where MAXTIMEOUT is platform specific. The default=20
     >      value is 10
     >      >>    seconds. This header field MAY occur in RECOGNIZE,=20
     >      SET-PARAMS or
     >      >>    GET-PARAMS.
     >      >>
     >      >>
     >      >> I am wondering if there is any difference between=20
     >      completion-causes
     >      >> 003 and 008. Can anybody clarify this?
     >      >>
     >      > I'll take a whack. 003 says that the recognizer didn't=20
     >      see anything=20
     >      > that matched an active grammar before the timer ran out.=20
     >      008 says the=20
     >      > recognizer got an utterance that matched the grammar but=20
     >      the user =20
     >      > kept babbling on and didn't stop so the match is=20
     probably bogus.
     >      >=20
     >      >> Claudia
     >      >>
     >      >> _______________________________________________
     >      >> Speechsc mailing list
     >      >> Speechsc@ietf.org
     >      >> https://www1.ietf.org/mailman/listinfo/speechsc
     >      >>
     >     =20
     >      _______________________________________________
     >      Speechsc mailing list
     >      Speechsc@ietf.org
     >      https://www1.ietf.org/mailman/listinfo/speechsc
     >     =20
    =20
     --
     Dr. Claudia Daboul
     IBP UK (Limited)
     Luetkensallee 19
     22041 Hamburg
     Tel.: +49 (40) 3172541
     Fax.: +49 (40) 3172547
     E-mail: claudia@ibp.de
    =20

_______________________________________________
Speechsc mailing list
Speechsc@ietf.org
https://www1.ietf.org/mailman/listinfo/speechsc



From speechsc-bounces@ietf.org Mon Jul 25 02:06:06 2005
Received: from localhost.localdomain ([127.0.0.1] helo=megatron.ietf.org)
	by megatron.ietf.org with esmtp (Exim 4.32)
	id 1Dww6H-0003Wu-Ux; Mon, 25 Jul 2005 02:06:05 -0400
Received: from odin.ietf.org ([132.151.1.176] helo=ietf.org)
	by megatron.ietf.org with esmtp (Exim 4.32) id 1Dww6B-0003V3-Bx
	for speechsc@megatron.ietf.org; Mon, 25 Jul 2005 02:06:04 -0400
Received: from ietf-mx.ietf.org (ietf-mx [132.151.6.1])
	by ietf.org (8.9.1a/8.9.1a) with ESMTP id CAA03600
	for <speechsc@ietf.org>; Mon, 25 Jul 2005 02:05:58 -0400 (EDT)
Received: from sj-iport-2-in.cisco.com ([171.71.176.71]
	helo=sj-iport-2.cisco.com) by ietf-mx.ietf.org with esmtp (Exim 4.43)
	id 1Dwway-0003wk-1E
	for speechsc@ietf.org; Mon, 25 Jul 2005 02:37:49 -0400
Received: from sj-core-2.cisco.com (171.71.177.254)
	by sj-iport-2.cisco.com with ESMTP; 24 Jul 2005 23:05:49 -0700
Received: from vtg-um-e2k6.sj21ad.cisco.com (vtg-um-e2k6.cisco.com
	[171.70.93.77])
	by sj-core-2.cisco.com (8.12.10/8.12.6) with ESMTP id j6P65kul013574;
	Sun, 24 Jul 2005 23:05:46 -0700 (PDT)
Content-class: urn:content-classes:message
MIME-Version: 1.0
Content-Type: text/plain;
	charset="us-ascii"
Content-Transfer-Encoding: quoted-printable
X-MimeOLE: Produced By Microsoft Exchange V6.5.7226.0
Subject: RE: [Speechsc] MRCPv2 Hotword Mode Recognition clarification
Date: Sun, 24 Jul 2005 23:05:48 -0700
Message-ID: <03772D1EC8DE624A863058C75874A75C1E06B9@vtg-um-e2k6.sj21ad.cisco.com>
Thread-Topic: [Speechsc] MRCPv2 Hotword Mode Recognition clarification
Thread-Index: AcVQ1vOUjRYYHTsPQwODCx3lYkmtthABsijA
From: "Shanmugham, Saravanan" <sarvi@cisco.com>
To: "Brett Gavagni" <gavagni@us.ibm.com>
X-Spam-Score: 0.0 (/)
X-Scan-Signature: 9cc83ac38bbbabacbf00f656311dd8d8
Content-Transfer-Encoding: quoted-printable
Cc: speechsc@ietf.org
X-BeenThere: speechsc@ietf.org
X-Mailman-Version: 2.1.5
Precedence: list
List-Id: Speech Services Control Working Group <speechsc.ietf.org>
List-Unsubscribe: <https://www1.ietf.org/mailman/listinfo/speechsc>,
	<mailto:speechsc-request@ietf.org?subject=unsubscribe>
List-Post: <mailto:speechsc@ietf.org>
List-Help: <mailto:speechsc-request@ietf.org?subject=help>
List-Subscribe: <https://www1.ietf.org/mailman/listinfo/speechsc>,
	<mailto:speechsc-request@ietf.org?subject=subscribe>
Sender: speechsc-bounces@ietf.org
Errors-To: speechsc-bounces@ietf.org

Addressed this with the following.

Allowed No-Input and Recognition timeouts to happen for Hotword mode
recognition as well.
The could still disabled starting the timers during RECOGNIZE with usual
Start-Input-Timers=3D false header.

START-OF-SPEECH will not be generated for hoteword mode, but the
Recogntion-Timeout will be started when the user does begin speaking. No
change here, just a clarification on when the recognition timeout is
started.

"When recognition is started and there is no match for a certain period
of time, the recognizer can send a RECOGNITION-COMPLETE event to the
client and terminate the recognition operation. For regular recognition,
this timer is started when a START-OF-SPEECH event is generated by the
resource. When this timer expires and there is a partial match to a
grammar the recognition request completes with a status code of "008
success-maxtime". For a hotword recognition mode, this timer is started
when the user begins speaking or when the START-INPUT-TIMERS method is
received. Note that for Hotword mode recognition the START-OF-SPEECH
event is not generated. When this timer expires in the hotword mode the
recognition request completes with a status code of "003
recognition-timeout". The recognition-timeout header allows the client
to set this timeout value. The value is in milliseconds. The value for
this header ranges from 0 to an implementation specific maximum value.
The default value is 10 seconds. This header MAY occur in RECOGNIZE,
SET-PARAMS or GET-PARAMS."


Sarvi

=20

     -----Original Message-----
     From: Brett Gavagni [mailto:gavagni@us.ibm.com]=20
     Sent: Wednesday, May 04, 2005 11:26 AM
     To: Shanmugham, Saravanan
     Cc: Reifenrath, Klaus; speechsc@ietf.org; Thomas Gal
     Subject: RE: [Speechsc] MRCPv2 Hotword Mode Recognition=20
     clarification
    =20
     Hi,
    =20
     VoiceXML vendors utilizing MRCP for speech resources are=20
     dependent on
     "Recognition-Mode: hotword" to support VoiceXML's=20
     "bargeintype=3Dhotword".=20
    =20
     It would be consistent to have a MRCP server terminate a=20
     recognition request (RECOGNITION-COMPLETE) in the case of=20
     a timer that is triggered for a "no-input" or=20
     "recognition-timeout" in either recognition modes.=20
    =20
     It would appear inconsistent for a MRCP server to provide=20
     timer capability for a "Recognition-Mode: normal" mode=20
     ("non-bargein" or a "speech bargein"), and require a=20
     VoiceXML vendor to implement their own timers for
     "Recognition-Mode: hotword".=20
    =20
     Preferred Proposal:
     In the case where a MRCP client does not want any timers=20
     enabled,  the client can specify a long value for the=20
     specific timer to be disabled (ie.=20
      -1 or some other token).=20
    =20
     Alternative Proposal:
     Clarification on the behavior of the recognition timers=20
     for both recognition modes would be greatly appreciated,=20
     including the clarification of the Recognizer State=20
     Machine requirement of generating an error response on=20
     START-INPUT-TIMERS when "Recognition-Mode: hotword".
    =20
     Thanks,
    =20
     Brett Gavagni
     WebSphere Voice Server Development
     http://www-306.ibm.com/software/pervasive/voice_server/
     gavagni@us.ibm.com
    =20
    =20
    =20
    =20
     "Shanmugham, Saravanan" <sarvi@cisco.com>
     05/04/2005 12:33 PM
    =20
     To
     Brett Gavagni/West Palm Beach/IBM@IBMUS, "Thomas Gal"=20
     <tgal@nsi.edu> cc <speechsc@ietf.org>, "Reifenrath, Klaus"=20
     <Klaus.Reifenrath@Scansoft.com> Subject
     RE: [Speechsc] MRCPv2 Hotword Mode Recognition clarification
    =20
    =20
    =20
    =20
    =20
    =20
     Could you propose exact text as to how you like to modify=20
     and clarify the use of this field. If that works for most=20
     others, I will update the specfication.
    =20
     Thanks,
     Sarvi=20
    =20
          -----Original Message-----
          From: speechsc-bounces@ietf.org=20
          [mailto:speechsc-bounces@ietf.org] On Behalf Of Brett Gavagni
          Sent: Friday, April 29, 2005 12:40 PM
          To: Thomas Gal
          Cc: speechsc@ietf.org; 'Reifenrath, Klaus'
          Subject: RE: [Speechsc] MRCPv2 Hotword Mode Recognition=20
          clarification
     =20
          I'll re-iterate the objection; if a client wants the=20
          recognition to last for a timeout, then it should be able=20
          to specify this condition.=20
     =20
          The current wording in the specification doesn't clearly=20
          indicate that a client request to START-INPUT-TIMERS is=20
          invalid for Recognition-Mode:=20
          hotword.
     =20
          The Recognizer State Machine should be clearly documented=20
          for both distinct recognition mode if the interpretation=20
          is expected to be inconsistent.
     =20
          Inconsistency usually add to the challenge of implementing=20
          a specification and facilitating interoperability.
     =20
          Thanks,
     =20
          Brett Gavagni
          WebSphere Voice Server Development
          http://www-306.ibm.com/software/pervasive/voice_server/
          gavagni@us.ibm.com
     =20
     =20
     =20
     =20
          "Thomas Gal" <tgal@nsi.edu>
          04/29/2005 02:58 PM
     =20
          To
          Brett Gavagni/West Palm Beach/IBM@IBMUS,=20
     "'Reifenrath, Klaus'"=20
          <Klaus.Reifenrath@Scansoft.com>
          cc
          <speechsc@ietf.org>
          Subject
          RE: [Speechsc] MRCPv2 Hotword Mode Recognition clarification
     =20
     =20
     =20
     =20
     =20
     =20
                           Hotword mode recognition is specifically=20
          the case where there's no timeout and you want to be=20
          actively listening. I don't really see any problem. But to=20
          be pragmatic......
     =20
                           Obviously in the case where there is a=20
          limit to the time frame in which the client will want that=20
          recognition to occur, they can then send a CANCEL which as=20
          you point out may leave holes if there are faulty clients.
          This doesn't prevent the server from having it's own set=20
          of limits for such things, which should obviously be set=20
          in the context of the application.=20
          Now
          I wouldn't have a problem with a special hotword mode=20
          timeout parameter (or just honoring the other timeout=20
          values) to help in that case, but I also don't see why=20
          that can't just be an implementation dependant parameter=20
          that doesn't need to be specified or investigated in this=20
          protocol. Chances are any services/products sold using=20
          this protocol will be based on per-port licensing or some=20
          other metric which will assure that it's in the best=20
          interest of the user to not be wasting recognition=20
          resources, as speech recognition is hardware intensive=20
          already. The fact remains that having any sort of real=20
          timeout value which limits the recognition scope to=20
          something short of the complete "speech transaction" (lets=20
          say phone call or
          whatever)
          makes it become just a regular old recognition and=20
          specifically NOT HOTWORD mode recognition as people define it.
     =20
          Personally I agree with Klaus that timeout values should=20
          be errors in the context of Hotword recognitions.
     =20
          -Tom
     =20
     =20
          -----Original Message-----
          From: speechsc-bounces@ietf.org=20
          [mailto:speechsc-bounces@ietf.org] On Behalf Of Brett Gavagni
          Sent: Friday, April 29, 2005 11:02 AM
          To: Reifenrath, Klaus
          Cc: speechsc@ietf.org
          Subject: RE: [Speechsc] MRCPv2 Hotword Mode Recognition=20
          clarification
     =20
          The following statement is a bit convoluted in the current=20
          draft specification w.r.t. hotword mode.
     =20
          "It does not timeout nor generate a no-match and will=20
     complete=20
             only for a successful match of grammar."=20
     =20
          The current draft specification doesn't detail that the=20
          timeout headers are invalid for a RECOGNIZE request with=20
          Recognition-Mode: hotword.
     =20
          What would justify a server not honoring timeouts if a=20
          client has the ability to specify timeout values?=20
     =20
          The hotword statement listed above also raises concerns in=20
          terms of a server providing a robust implementation. An=20
          implementing server would have to entirely rely on a=20
          client to terminate a recognition request where=20
     =20
          a hotword isn't matched. The complete reliance on a client=20
          has a whole separate set of issues for a server to=20
          facilitate redundancy.
     =20
          Thanks,
     =20
          Brett Gavagni
          WebSphere Voice Server Development
          http://www-306.ibm.com/software/pervasive/voice_server/
          gavagni@us.ibm.com
     =20
     =20
     =20
     =20
          "Reifenrath, Klaus" <Klaus.Reifenrath@Scansoft.com>
          04/28/2005 09:59 AM
     =20
          To
          Brett Gavagni/West Palm Beach/IBM@IBMUS
          cc
          speechsc@ietf.org
          Subject
          RE: [Speechsc] MRCPv2 Hotword Mode Recognition clarification
     =20
     =20
     =20
     =20
     =20
     =20
          I think the server should return with status code 402=20
          (method not valid in=20
     =20
          this state).
     =20
          Klaus
     =20
          From: Brett Gavagni [mailto:gavagni@us.ibm.com]
          Sent: Donnerstag, 28. April 2005 15:39
          To: speechsc@ietf.org
          Subject: [Speechsc] MRCPv2 Hotword Mode Recognition=20
     clarification
     =20
     =20
          Hi,=20
     =20
          There's room for interpretation in the following section=20
          of the current=20
          draft of MRCPv2.=20
     =20
          p53 of the draft-ietf-speechsc-mrcpv2-06.txt=20
          Hotword Mode Recognition=20
             Hotword mode is where the recognizer looks for a=20
          specific speech=20
             grammar or dtmf sequence and ignores speech or DTMF=20
          that does not=20
             match. It does not timeout nor generate a no-match and=20
          will complete=20
             only for a successful match of grammar.=20
     =20
          What response should a server generate for a=20
          START-INPUT-TIMERS request,=20
          when the current session has a recognition in progress in=20
          hotword mode?=20
     =20
          What would justify a server not honoring timeouts if a=20
          client has the=20
          ability to specify timeout values?=20
     =20
          Thanks,
     =20
          Brett Gavagni=20
          WebSphere Voice Server Development=20
          http://www-306.ibm.com/software/pervasive/voice_server/
          gavagni@us.ibm.com
     =20
     =20
          _______________________________________________
          Speechsc mailing list
          Speechsc@ietf.org
          https://www1.ietf.org/mailman/listinfo/speechsc
     =20
     =20
     =20
     =20
          _______________________________________________
          Speechsc mailing list
          Speechsc@ietf.org
          https://www1.ietf.org/mailman/listinfo/speechsc
     =20
    =20

_______________________________________________
Speechsc mailing list
Speechsc@ietf.org
https://www1.ietf.org/mailman/listinfo/speechsc



From speechsc-bounces@ietf.org Mon Jul 25 10:00:52 2005
Received: from localhost.localdomain ([127.0.0.1] helo=megatron.ietf.org)
	by megatron.ietf.org with esmtp (Exim 4.32)
	id 1Dx3Vk-0000Tr-S7; Mon, 25 Jul 2005 10:00:52 -0400
Received: from odin.ietf.org ([132.151.1.176] helo=ietf.org)
	by megatron.ietf.org with esmtp (Exim 4.32) id 1Dx3Vj-0000Mu-1B
	for speechsc@megatron.ietf.org; Mon, 25 Jul 2005 10:00:51 -0400
Received: from ietf-mx.ietf.org (ietf-mx [132.151.6.1])
	by ietf.org (8.9.1a/8.9.1a) with ESMTP id KAA21419
	for <speechsc@ietf.org>; Mon, 25 Jul 2005 10:00:47 -0400 (EDT)
Received: from e31.co.us.ibm.com ([32.97.110.129])
	by ietf-mx.ietf.org with esmtp (Exim 4.43) id 1Dx40U-0003Qe-Vx
	for speechsc@ietf.org; Mon, 25 Jul 2005 10:32:42 -0400
Received: from d03relay04.boulder.ibm.com (d03relay04.boulder.ibm.com
	[9.17.195.106])
	by e31.co.us.ibm.com (8.12.10/8.12.9) with ESMTP id j6PE0UWQ400176
	for <speechsc@ietf.org>; Mon, 25 Jul 2005 10:00:31 -0400
Received: from d03av01.boulder.ibm.com (d03av01.boulder.ibm.com [9.17.195.167])
	by d03relay04.boulder.ibm.com (8.12.10/NCO/VERS6.7) with ESMTP id
	j6PE0WNT076336
	for <speechsc@ietf.org>; Mon, 25 Jul 2005 08:00:32 -0600
Received: from d03av01.boulder.ibm.com (loopback [127.0.0.1])
	by d03av01.boulder.ibm.com (8.12.11/8.13.3) with ESMTP id
	j6PE0UVd011034
	for <speechsc@ietf.org>; Mon, 25 Jul 2005 08:00:30 -0600
Received: from d03nm119.boulder.ibm.com (d03nm119.boulder.ibm.com
	[9.17.195.145])
	by d03av01.boulder.ibm.com (8.12.11/8.12.11) with ESMTP id
	j6PE0UQX011026
	for <speechsc@ietf.org>; Mon, 25 Jul 2005 08:00:30 -0600
In-Reply-To: <03772D1EC8DE624A863058C75874A75C1E06B9@vtg-um-e2k6.sj21ad.cisco.com>
To: "Shanmugham, Saravanan" <sarvi@cisco.com>
MIME-Version: 1.0
Subject: RE: [Speechsc] MRCPv2 Hotword Mode Recognition clarification
X-Mailer: Lotus Notes Release 6.0.2CF1 June 9, 2003
Message-ID: <OFE7EE2630.D6FF33D7-ON87257049.004CCA9C-85257049.004CF331@us.ibm.com>
From: Brett Gavagni <gavagni@us.ibm.com>
Date: Mon, 25 Jul 2005 10:00:31 -0400
X-MIMETrack: Serialize by Router on D03NM119/03/M/IBM(Release 6.5.4|March 27,
	2005) at 07/25/2005 08:00:31,
	Serialize complete at 07/25/2005 08:00:31
Content-Type: text/plain; charset="US-ASCII"
X-Spam-Score: 0.0 (/)
X-Scan-Signature: 79bb66f827e54e9d5c5c7f1f9d645608
Cc: speechsc@ietf.org
X-BeenThere: speechsc@ietf.org
X-Mailman-Version: 2.1.5
Precedence: list
List-Id: Speech Services Control Working Group <speechsc.ietf.org>
List-Unsubscribe: <https://www1.ietf.org/mailman/listinfo/speechsc>,
	<mailto:speechsc-request@ietf.org?subject=unsubscribe>
List-Post: <mailto:speechsc@ietf.org>
List-Help: <mailto:speechsc-request@ietf.org?subject=help>
List-Subscribe: <https://www1.ietf.org/mailman/listinfo/speechsc>,
	<mailto:speechsc-request@ietf.org?subject=subscribe>
Sender: speechsc-bounces@ietf.org
Errors-To: speechsc-bounces@ietf.org

Looks great!

Thanks,

Brett Gavagni 
WebSphere Voice Server Development 
http://www-306.ibm.com/software/pervasive/voice_server/
gavagni@us.ibm.com




"Shanmugham, Saravanan" <sarvi@cisco.com> 
07/25/2005 02:05 AM

To
Brett Gavagni/West Palm Beach/IBM@IBMUS
cc
<speechsc@ietf.org>
Subject
RE: [Speechsc] MRCPv2 Hotword Mode Recognition clarification






Addressed this with the following.

Allowed No-Input and Recognition timeouts to happen for Hotword mode
recognition as well.
The could still disabled starting the timers during RECOGNIZE with usual
Start-Input-Timers= false header.

START-OF-SPEECH will not be generated for hoteword mode, but the
Recogntion-Timeout will be started when the user does begin speaking. No
change here, just a clarification on when the recognition timeout is
started.

"When recognition is started and there is no match for a certain period
of time, the recognizer can send a RECOGNITION-COMPLETE event to the
client and terminate the recognition operation. For regular recognition,
this timer is started when a START-OF-SPEECH event is generated by the
resource. When this timer expires and there is a partial match to a
grammar the recognition request completes with a status code of "008
success-maxtime". For a hotword recognition mode, this timer is started
when the user begins speaking or when the START-INPUT-TIMERS method is
received. Note that for Hotword mode recognition the START-OF-SPEECH
event is not generated. When this timer expires in the hotword mode the
recognition request completes with a status code of "003
recognition-timeout". The recognition-timeout header allows the client
to set this timeout value. The value is in milliseconds. The value for
this header ranges from 0 to an implementation specific maximum value.
The default value is 10 seconds. This header MAY occur in RECOGNIZE,
SET-PARAMS or GET-PARAMS."


Sarvi

 

     -----Original Message-----
     From: Brett Gavagni [mailto:gavagni@us.ibm.com] 
     Sent: Wednesday, May 04, 2005 11:26 AM
     To: Shanmugham, Saravanan
     Cc: Reifenrath, Klaus; speechsc@ietf.org; Thomas Gal
     Subject: RE: [Speechsc] MRCPv2 Hotword Mode Recognition 
     clarification
 
     Hi,
 
     VoiceXML vendors utilizing MRCP for speech resources are 
     dependent on
     "Recognition-Mode: hotword" to support VoiceXML's 
     "bargeintype=hotword". 
 
     It would be consistent to have a MRCP server terminate a 
     recognition request (RECOGNITION-COMPLETE) in the case of 
     a timer that is triggered for a "no-input" or 
     "recognition-timeout" in either recognition modes. 
 
     It would appear inconsistent for a MRCP server to provide 
     timer capability for a "Recognition-Mode: normal" mode 
     ("non-bargein" or a "speech bargein"), and require a 
     VoiceXML vendor to implement their own timers for
     "Recognition-Mode: hotword". 
 
     Preferred Proposal:
     In the case where a MRCP client does not want any timers 
     enabled,  the client can specify a long value for the 
     specific timer to be disabled (ie. 
      -1 or some other token). 
 
     Alternative Proposal:
     Clarification on the behavior of the recognition timers 
     for both recognition modes would be greatly appreciated, 
     including the clarification of the Recognizer State 
     Machine requirement of generating an error response on 
     START-INPUT-TIMERS when "Recognition-Mode: hotword".
 
     Thanks,
 
     Brett Gavagni
     WebSphere Voice Server Development
     http://www-306.ibm.com/software/pervasive/voice_server/
     gavagni@us.ibm.com
 
 
 
 
     "Shanmugham, Saravanan" <sarvi@cisco.com>
     05/04/2005 12:33 PM
 
     To
     Brett Gavagni/West Palm Beach/IBM@IBMUS, "Thomas Gal" 
     <tgal@nsi.edu> cc <speechsc@ietf.org>, "Reifenrath, Klaus" 
     <Klaus.Reifenrath@Scansoft.com> Subject
     RE: [Speechsc] MRCPv2 Hotword Mode Recognition clarification
 
 
 
 
 
 
     Could you propose exact text as to how you like to modify 
     and clarify the use of this field. If that works for most 
     others, I will update the specfication.
 
     Thanks,
     Sarvi 
 
          -----Original Message-----
          From: speechsc-bounces@ietf.org 
          [mailto:speechsc-bounces@ietf.org] On Behalf Of Brett Gavagni
          Sent: Friday, April 29, 2005 12:40 PM
          To: Thomas Gal
          Cc: speechsc@ietf.org; 'Reifenrath, Klaus'
          Subject: RE: [Speechsc] MRCPv2 Hotword Mode Recognition 
          clarification
 
          I'll re-iterate the objection; if a client wants the 
          recognition to last for a timeout, then it should be able 
          to specify this condition. 
 
          The current wording in the specification doesn't clearly 
          indicate that a client request to START-INPUT-TIMERS is 
          invalid for Recognition-Mode: 
          hotword.
 
          The Recognizer State Machine should be clearly documented 
          for both distinct recognition mode if the interpretation 
          is expected to be inconsistent.
 
          Inconsistency usually add to the challenge of implementing 
          a specification and facilitating interoperability.
 
          Thanks,
 
          Brett Gavagni
          WebSphere Voice Server Development
          http://www-306.ibm.com/software/pervasive/voice_server/
          gavagni@us.ibm.com
 
 
 
 
          "Thomas Gal" <tgal@nsi.edu>
          04/29/2005 02:58 PM
 
          To
          Brett Gavagni/West Palm Beach/IBM@IBMUS, 
     "'Reifenrath, Klaus'" 
          <Klaus.Reifenrath@Scansoft.com>
          cc
          <speechsc@ietf.org>
          Subject
          RE: [Speechsc] MRCPv2 Hotword Mode Recognition clarification
 
 
 
 
 
 
                           Hotword mode recognition is specifically 
          the case where there's no timeout and you want to be 
          actively listening. I don't really see any problem. But to 
          be pragmatic......
 
                           Obviously in the case where there is a 
          limit to the time frame in which the client will want that 
          recognition to occur, they can then send a CANCEL which as 
          you point out may leave holes if there are faulty clients.
          This doesn't prevent the server from having it's own set 
          of limits for such things, which should obviously be set 
          in the context of the application. 
          Now
          I wouldn't have a problem with a special hotword mode 
          timeout parameter (or just honoring the other timeout 
          values) to help in that case, but I also don't see why 
          that can't just be an implementation dependant parameter 
          that doesn't need to be specified or investigated in this 
          protocol. Chances are any services/products sold using 
          this protocol will be based on per-port licensing or some 
          other metric which will assure that it's in the best 
          interest of the user to not be wasting recognition 
          resources, as speech recognition is hardware intensive 
          already. The fact remains that having any sort of real 
          timeout value which limits the recognition scope to 
          something short of the complete "speech transaction" (lets 
          say phone call or
          whatever)
          makes it become just a regular old recognition and 
          specifically NOT HOTWORD mode recognition as people define it.
 
          Personally I agree with Klaus that timeout values should 
          be errors in the context of Hotword recognitions.
 
          -Tom
 
 
          -----Original Message-----
          From: speechsc-bounces@ietf.org 
          [mailto:speechsc-bounces@ietf.org] On Behalf Of Brett Gavagni
          Sent: Friday, April 29, 2005 11:02 AM
          To: Reifenrath, Klaus
          Cc: speechsc@ietf.org
          Subject: RE: [Speechsc] MRCPv2 Hotword Mode Recognition 
          clarification
 
          The following statement is a bit convoluted in the current 
          draft specification w.r.t. hotword mode.
 
          "It does not timeout nor generate a no-match and will 
     complete 
             only for a successful match of grammar." 
 
          The current draft specification doesn't detail that the 
          timeout headers are invalid for a RECOGNIZE request with 
          Recognition-Mode: hotword.
 
          What would justify a server not honoring timeouts if a 
          client has the ability to specify timeout values? 
 
          The hotword statement listed above also raises concerns in 
          terms of a server providing a robust implementation. An 
          implementing server would have to entirely rely on a 
          client to terminate a recognition request where 
 
          a hotword isn't matched. The complete reliance on a client 
          has a whole separate set of issues for a server to 
          facilitate redundancy.
 
          Thanks,
 
          Brett Gavagni
          WebSphere Voice Server Development
          http://www-306.ibm.com/software/pervasive/voice_server/
          gavagni@us.ibm.com
 
 
 
 
          "Reifenrath, Klaus" <Klaus.Reifenrath@Scansoft.com>
          04/28/2005 09:59 AM
 
          To
          Brett Gavagni/West Palm Beach/IBM@IBMUS
          cc
          speechsc@ietf.org
          Subject
          RE: [Speechsc] MRCPv2 Hotword Mode Recognition clarification
 
 
 
 
 
 
          I think the server should return with status code 402 
          (method not valid in 
 
          this state).
 
          Klaus
 
          From: Brett Gavagni [mailto:gavagni@us.ibm.com]
          Sent: Donnerstag, 28. April 2005 15:39
          To: speechsc@ietf.org
          Subject: [Speechsc] MRCPv2 Hotword Mode Recognition 
     clarification
 
 
          Hi, 
 
          There's room for interpretation in the following section 
          of the current 
          draft of MRCPv2. 
 
          p53 of the draft-ietf-speechsc-mrcpv2-06.txt 
          Hotword Mode Recognition 
             Hotword mode is where the recognizer looks for a 
          specific speech 
             grammar or dtmf sequence and ignores speech or DTMF 
          that does not 
             match. It does not timeout nor generate a no-match and 
          will complete 
             only for a successful match of grammar. 
 
          What response should a server generate for a 
          START-INPUT-TIMERS request, 
          when the current session has a recognition in progress in 
          hotword mode? 
 
          What would justify a server not honoring timeouts if a 
          client has the 
          ability to specify timeout values? 
 
          Thanks,
 
          Brett Gavagni 
          WebSphere Voice Server Development 
          http://www-306.ibm.com/software/pervasive/voice_server/
          gavagni@us.ibm.com
 
 
          _______________________________________________
          Speechsc mailing list
          Speechsc@ietf.org
          https://www1.ietf.org/mailman/listinfo/speechsc
 
 
 
 
          _______________________________________________
          Speechsc mailing list
          Speechsc@ietf.org
          https://www1.ietf.org/mailman/listinfo/speechsc
 
 



_______________________________________________
Speechsc mailing list
Speechsc@ietf.org
https://www1.ietf.org/mailman/listinfo/speechsc



From speechsc-bounces@ietf.org Mon Jul 25 18:19:53 2005
Received: from localhost.localdomain ([127.0.0.1] helo=megatron.ietf.org)
	by megatron.ietf.org with esmtp (Exim 4.32)
	id 1DxBIf-0008UH-FE; Mon, 25 Jul 2005 18:19:53 -0400
Received: from odin.ietf.org ([132.151.1.176] helo=ietf.org)
	by megatron.ietf.org with esmtp (Exim 4.32) id 1DxBId-0008UC-Vh
	for speechsc@megatron.ietf.org; Mon, 25 Jul 2005 18:19:52 -0400
Received: from ietf-mx.ietf.org (ietf-mx [132.151.6.1])
	by ietf.org (8.9.1a/8.9.1a) with ESMTP id SAA29171
	for <speechsc@ietf.org>; Mon, 25 Jul 2005 18:19:47 -0400 (EDT)
Received: from e5.ny.us.ibm.com ([32.97.182.145])
	by ietf-mx.ietf.org with esmtp (Exim 4.43) id 1DxBnW-0002Ip-Vs
	for speechsc@ietf.org; Mon, 25 Jul 2005 18:51:48 -0400
Received: from d01relay04.pok.ibm.com (d01relay04.pok.ibm.com [9.56.227.236])
	by e5.ny.us.ibm.com (8.12.11/8.12.11) with ESMTP id j6PMJZmN023646
	for <speechsc@ietf.org>; Mon, 25 Jul 2005 18:19:35 -0400
Received: from d01av03.pok.ibm.com (d01av03.pok.ibm.com [9.56.224.217])
	by d01relay04.pok.ibm.com (8.12.10/NCO/VERS6.7) with ESMTP id
	j6PMJZFm205000
	for <speechsc@ietf.org>; Mon, 25 Jul 2005 18:19:35 -0400
Received: from d01av03.pok.ibm.com (loopback [127.0.0.1])
	by d01av03.pok.ibm.com (8.12.11/8.13.3) with ESMTP id j6PMJZmM024355
	for <speechsc@ietf.org>; Mon, 25 Jul 2005 18:19:35 -0400
Received: from d27mc602.rchland.ibm.com (d27mc602.rchland.ibm.com
	[9.10.229.36])
	by d01av03.pok.ibm.com (8.12.11/8.12.11) with ESMTP id j6PMJYRQ024320
	for <speechsc@ietf.org>; Mon, 25 Jul 2005 18:19:34 -0400
From: Can P Boyacigiller <can@us.ibm.com>
To: speechsc@ietf.org
Message-ID: <OFB30FF1D5.EBFD42C7-ON86257049.007AA37F-86257049.007AA37F@us.ibm.com>
Date: Mon, 25 Jul 2005 17:19:32 -0500
X-MIMETrack: Serialize by Router on d27mc602/27/M/IBM(Release 6.0.4|June 01,
	2004) at 07/25/2005 05:19:33 PM
MIME-Version: 1.0
Content-type: text/plain; charset=US-ASCII
X-Spam-Score: 0.0 (/)
X-Scan-Signature: 08e48e05374109708c00c6208b534009
Subject: [Speechsc] Can Paul has left the building
X-BeenThere: speechsc@ietf.org
X-Mailman-Version: 2.1.5
Precedence: list
List-Id: Speech Services Control Working Group <speechsc.ietf.org>
List-Unsubscribe: <https://www1.ietf.org/mailman/listinfo/speechsc>,
	<mailto:speechsc-request@ietf.org?subject=unsubscribe>
List-Post: <mailto:speechsc@ietf.org>
List-Help: <mailto:speechsc-request@ietf.org?subject=help>
List-Subscribe: <https://www1.ietf.org/mailman/listinfo/speechsc>,
	<mailto:speechsc-request@ietf.org?subject=subscribe>
Sender: speechsc-bounces@ietf.org
Errors-To: speechsc-bounces@ietf.org

I will be out of the office starting  07/25/2005 and will not return until
08/14/2005.

I will be out of the office at IBM Hursley UK for the next three weeks
returning on August 15th.  I will have no voice access, but will have
access to email during this time.   For any issues please contact Tom
Thacher (development manager) or my manager - Randy Freeman.


_______________________________________________
Speechsc mailing list
Speechsc@ietf.org
https://www1.ietf.org/mailman/listinfo/speechsc



From speechsc-bounces@ietf.org Tue Jul 26 11:28:26 2005
Received: from localhost.localdomain ([127.0.0.1] helo=megatron.ietf.org)
	by megatron.ietf.org with esmtp (Exim 4.32)
	id 1DxRM1-0003t4-RG; Tue, 26 Jul 2005 11:28:25 -0400
Received: from odin.ietf.org ([132.151.1.176] helo=ietf.org)
	by megatron.ietf.org with esmtp (Exim 4.32) id 1DxRM0-0003sb-6Q
	for speechsc@megatron.ietf.org; Tue, 26 Jul 2005 11:28:24 -0400
Received: from ietf-mx.ietf.org (ietf-mx [132.151.6.1])
	by ietf.org (8.9.1a/8.9.1a) with ESMTP id LAA27825
	for <speechsc@ietf.org>; Tue, 26 Jul 2005 11:28:19 -0400 (EDT)
Received: from e31.co.us.ibm.com ([32.97.110.129])
	by ietf-mx.ietf.org with esmtp (Exim 4.43) id 1DxRr1-00087x-Tr
	for speechsc@ietf.org; Tue, 26 Jul 2005 12:00:29 -0400
Received: from westrelay02.boulder.ibm.com (westrelay02.boulder.ibm.com
	[9.17.195.11])
	by e31.co.us.ibm.com (8.12.10/8.12.9) with ESMTP id j6QFS67K429062
	for <speechsc@ietf.org>; Tue, 26 Jul 2005 11:28:06 -0400
Received: from d03av03.boulder.ibm.com (d03av03.boulder.ibm.com [9.17.195.169])
	by westrelay02.boulder.ibm.com (8.12.10/NCO/VERS6.7) with ESMTP id
	j6QFS69v404738
	for <speechsc@ietf.org>; Tue, 26 Jul 2005 09:28:06 -0600
Received: from d03av03.boulder.ibm.com (loopback [127.0.0.1])
	by d03av03.boulder.ibm.com (8.12.11/8.13.3) with ESMTP id
	j6QFS6C0023014
	for <speechsc@ietf.org>; Tue, 26 Jul 2005 09:28:06 -0600
Received: from d03nm119.boulder.ibm.com (d03nm119.boulder.ibm.com
	[9.17.195.145])
	by d03av03.boulder.ibm.com (8.12.11/8.12.11) with ESMTP id
	j6QFS6YG023002
	for <speechsc@ietf.org>; Tue, 26 Jul 2005 09:28:06 -0600
To: speechsc@ietf.org
MIME-Version: 1.0
X-Mailer: Lotus Notes Release 6.0.2CF1 June 9, 2003
Message-ID: <OF84D8B908.97B7BE45-ON8725704A.005396DD-8525704A.0054F876@us.ibm.com>
From: Brett Gavagni <gavagni@us.ibm.com>
Date: Tue, 26 Jul 2005 11:28:07 -0400
X-MIMETrack: Serialize by Router on D03NM119/03/M/IBM(Release 6.5.4|March 27,
	2005) at 07/26/2005 09:28:09,
	Serialize complete at 07/26/2005 09:28:09
Content-Type: text/plain; charset="US-ASCII"
X-Spam-Score: 0.0 (/)
X-Scan-Signature: 79899194edc4f33a41f49410777972f8
Subject: [Speechsc] Personal-Grammar-URI clarification
X-BeenThere: speechsc@ietf.org
X-Mailman-Version: 2.1.5
Precedence: list
List-Id: Speech Services Control Working Group <speechsc.ietf.org>
List-Unsubscribe: <https://www1.ietf.org/mailman/listinfo/speechsc>,
	<mailto:speechsc-request@ietf.org?subject=unsubscribe>
List-Post: <mailto:speechsc@ietf.org>
List-Help: <mailto:speechsc-request@ietf.org?subject=help>
List-Subscribe: <https://www1.ietf.org/mailman/listinfo/speechsc>,
	<mailto:speechsc-request@ietf.org?subject=subscribe>
Sender: speechsc-bounces@ietf.org
Errors-To: speechsc-bounces@ietf.org

Hi,

The wording throughout the current draft specification is a bit confusing 
w.r.t "Personal-Grammar-URI".

Is the expectation that the "Personal-Grammar-URI" value is merely a 
client specified handle to an external grammar in lieu of using a 
"Content-Id" value?

Is the expectation that the server persistently store these references and 
the backing grammars, and when and how would the release of these 
references occur?

Thanks,

Brett Gavagni 
WebSphere Voice Server Development 
http://www-306.ibm.com/software/pervasive/voice_server/ 
gavagni@us.ibm.com


_______________________________________________
Speechsc mailing list
Speechsc@ietf.org
https://www1.ietf.org/mailman/listinfo/speechsc



From speechsc-bounces@ietf.org Wed Jul 27 18:54:10 2005
Received: from localhost.localdomain ([127.0.0.1] helo=megatron.ietf.org)
	by megatron.ietf.org with esmtp (Exim 4.32)
	id 1Dxumw-0000Dd-Iq; Wed, 27 Jul 2005 18:54:10 -0400
Received: from odin.ietf.org ([132.151.1.176] helo=ietf.org)
	by megatron.ietf.org with esmtp (Exim 4.32) id 1Dxumu-0000DT-V5
	for speechsc@megatron.ietf.org; Wed, 27 Jul 2005 18:54:09 -0400
Received: from ietf-mx.ietf.org (ietf-mx [132.151.6.1])
	by ietf.org (8.9.1a/8.9.1a) with ESMTP id SAA12266
	for <speechsc@ietf.org>; Wed, 27 Jul 2005 18:54:05 -0400 (EDT)
Received: from sj-iport-3-in.cisco.com ([171.71.176.72]
	helo=sj-iport-3.cisco.com) by ietf-mx.ietf.org with esmtp (Exim 4.43)
	id 1DxvIG-0006Hs-3q
	for speechsc@ietf.org; Wed, 27 Jul 2005 19:26:32 -0400
Received: from sj-core-5.cisco.com (171.71.177.238)
	by sj-iport-3.cisco.com with ESMTP; 27 Jul 2005 15:53:59 -0700
X-IronPort-AV: i="3.95,147,1120460400"; 
	d="scan'208,217"; a="326626949:sNHT55442302"
Received: from xbh-sjc-221.amer.cisco.com (xbh-sjc-221.cisco.com
	[128.107.191.63])
	by sj-core-5.cisco.com (8.12.10/8.12.6) with ESMTP id j6RMrwJR022858
	for <speechsc@ietf.org>; Wed, 27 Jul 2005 15:53:59 -0700 (PDT)
Received: from xmb-sjc-224.amer.cisco.com ([128.107.191.98]) by
	xbh-sjc-221.amer.cisco.com with Microsoft SMTPSVC(6.0.3790.211);
	Wed, 27 Jul 2005 15:53:58 -0700
X-MimeOLE: Produced By Microsoft Exchange V6.5.7226.0
Content-class: urn:content-classes:message
MIME-Version: 1.0
Subject: RE: [Speechsc] propose adding "media-type" header field
	inRECOGNIZEorSET-PARAMS/GET-PARAMS methods
Date: Wed, 27 Jul 2005 15:53:56 -0700
Message-ID: <AD8171548C328E46BD8D0FEE40D9968872A1CC@xmb-sjc-224.amer.cisco.com>
Thread-Topic: [Speechsc] propose adding "media-type" header field
	inRECOGNIZEorSET-PARAMS/GET-PARAMS methods
Thread-Index: AcVM1EeF6F3QJhBHRniGmUHADY5x4wAA+UewAPsQNVAQjcCwwAAABgpQ
From: "Jenny Yao \(jyao\)" <jyao@cisco.com>
To: <speechsc@ietf.org>
X-OriginalArrivalTime: 27 Jul 2005 22:53:58.0161 (UTC)
	FILETIME=[129DD810:01C592FE]
X-Spam-Score: 0.4 (/)
X-Scan-Signature: e06437eb72f6703f11713d345be8298a
X-BeenThere: speechsc@ietf.org
X-Mailman-Version: 2.1.5
Precedence: list
List-Id: Speech Services Control Working Group <speechsc.ietf.org>
List-Unsubscribe: <https://www1.ietf.org/mailman/listinfo/speechsc>,
	<mailto:speechsc-request@ietf.org?subject=unsubscribe>
List-Post: <mailto:speechsc@ietf.org>
List-Help: <mailto:speechsc-request@ietf.org?subject=help>
List-Subscribe: <https://www1.ietf.org/mailman/listinfo/speechsc>,
	<mailto:speechsc-request@ietf.org?subject=subscribe>
Content-Type: multipart/mixed; boundary="===============0856537083=="
Sender: speechsc-bounces@ietf.org
Errors-To: speechsc-bounces@ietf.org

This is a multi-part message in MIME format.

--===============0856537083==
Content-class: urn:content-classes:message
Content-Type: multipart/alternative;
	boundary="----_=_NextPart_001_01C592FE.127827D7"

This is a multi-part message in MIME format.

------_=_NextPart_001_01C592FE.127827D7
Content-Type: text/plain;
	charset="us-ascii"
Content-Transfer-Encoding: quoted-printable

There appears to be no opposition regarding the media-type header field
request that we sent earlier. So, can we request this feature be
included in the next v2 draft?=20
=20
In addition, VXML 2.1 needs to have the information about "duration" and
"size" of the recorded audio in the field "waveform-uri" in recog-only
header. We would like to propose the following fields be added to the
recog-only header:
=20
        waveform-duration - the duration of the audio file in
milliseconds
        waveform-size - the size of the audio file in bytes
=20
Thanks.
=20
Jenny

________________________________

From: speechsc-bounces@ietf.org [mailto:speechsc-bounces@ietf.org] On
Behalf Of Saravanan Shanmugham (sarvi)
Sent: Wednesday, May 04, 2005 9:27 AM
To: Thomas Gal; Jenny Yao (jyao); speechsc@ietf.org
Subject: RE: [Speechsc] propose adding "media-type" header field
inRECOGNIZEorSET-PARAMS/GET-PARAMS methods


I don't see why not, considering all this behaviour would apply only if
the "Save-Waveform" header is set to true and failure to caoture or save
the waveform does not necessarily stop the RECOGNIZE operation itself.=20
=20
Anyone opposed to doing this.=20
=20
Thanks,
Sarvi


________________________________

	From: speechsc-bounces@ietf.org
[mailto:speechsc-bounces@ietf.org] On Behalf Of Thomas Gal
	Sent: Friday, April 29, 2005 9:47 AM
	To: 'Jenny Yao (jyao)'; speechsc@ietf.org
	Subject: RE: [Speechsc] propose adding "media-type" header field
in RECOGNIZEorSET-PARAMS/GET-PARAMS methods
=09
=09

	Though this information could probably be inferred from the
filename extension, If we are going to follow/emulate the RECORD
methodology than it should also be available as a MIME body to the
RECOGNITION COMPLETE/STOP events as well. Otherwise I agree completely.

	=20

	-Tom

	=20

=09
________________________________


	From: speechsc-bounces@ietf.org
[mailto:speechsc-bounces@ietf.org] On Behalf Of Jenny Yao (jyao)
	Sent: Friday, April 29, 2005 8:58 AM
	To: speechsc@ietf.org
	Subject: [Speechsc] propose adding "media-type" header field in
RECOGNIZE orSET-PARAMS/GET-PARAMS methods

	=20

	=20

	=20

	VoiceXML 2.1 specifies "recording user utterances while
attempting recognition". This can be done through recog-only-header
"save-waveform" and "waveform-uri" in MRCP V1 and V2. VoiceXML 2.1 also
uses recordutterancetype property to specify the media format of the
result recording. However, there is no way in MRCP to pass the required
media format to the server. We propose using the "Media-Type" header
field, currently defined for the recording resource, in the RECOGNIZE
method or the SET-PARAMS/GET-PARAMS methods to specify a media type. The
Save-Waveform carries a URI pointing to the audio captured during
recognition. The captured audio SHOULD be saved with this media-type.

	=20

	Thanks.

	=20

	Jenny


------_=_NextPart_001_01C592FE.127827D7
Content-Type: text/html;
	charset="us-ascii"
Content-Transfer-Encoding: quoted-printable

<!DOCTYPE HTML PUBLIC "-//W3C//DTD HTML 4.0 Transitional//EN">
<HTML xmlns=3D"http://www.w3.org/TR/REC-html40" xmlns:v =3D=20
"urn:schemas-microsoft-com:vml" xmlns:o =3D=20
"urn:schemas-microsoft-com:office:office" xmlns:w =3D=20
"urn:schemas-microsoft-com:office:word"><HEAD>
<META http-equiv=3DContent-Type content=3D"text/html; =
charset=3Dus-ascii">
<META content=3D"MSHTML 6.00.2800.1505" name=3DGENERATOR><!--[if !mso]>
<STYLE>v\:* {
	BEHAVIOR: url(#default#VML)
}
o\:* {
	BEHAVIOR: url(#default#VML)
}
w\:* {
	BEHAVIOR: url(#default#VML)
}
.shape {
	BEHAVIOR: url(#default#VML)
}
</STYLE>
<![endif]-->
<STYLE>@font-face {
	font-family: Tahoma;
}
@page Section1 {size: 8.5in 11.0in; margin: 1.0in 1.25in 1.0in 1.25in; }
P.MsoNormal {
	FONT-SIZE: 12pt; MARGIN: 0in 0in 0pt; FONT-FAMILY: "Times New Roman"
}
LI.MsoNormal {
	FONT-SIZE: 12pt; MARGIN: 0in 0in 0pt; FONT-FAMILY: "Times New Roman"
}
DIV.MsoNormal {
	FONT-SIZE: 12pt; MARGIN: 0in 0in 0pt; FONT-FAMILY: "Times New Roman"
}
A:link {
	COLOR: blue; TEXT-DECORATION: underline
}
SPAN.MsoHyperlink {
	COLOR: blue; TEXT-DECORATION: underline
}
A:visited {
	COLOR: purple; TEXT-DECORATION: underline
}
SPAN.MsoHyperlinkFollowed {
	COLOR: purple; TEXT-DECORATION: underline
}
PRE {
	FONT-SIZE: 10pt; MARGIN: 0in 0in 0pt; FONT-FAMILY: "Courier New"
}
SPAN.EmailStyle18 {
	COLOR: navy; FONT-FAMILY: Arial; mso-style-type: personal-reply
}
DIV.Section1 {
	page: Section1
}
</STYLE>
</HEAD>
<BODY lang=3DEN-US vLink=3Dpurple link=3Dblue>
<DIV dir=3Dltr align=3Dleft><SPAN =
class=3D299473522-27072005></SPAN><FONT=20
face=3DArial><FONT color=3D#0000ff><FONT size=3D2><SPAN =
class=3D299473522-27072005>There=20
appears to be no opposition r</SPAN><SPAN =
class=3D299473522-27072005>egarding the=20
media-type header field request that&nbsp;we sent earlier. So, can we =
request=20
this feature be included in the next v2 draft?=20
</SPAN></FONT></FONT></FONT></DIV>
<DIV dir=3Dltr align=3Dleft><FONT face=3DArial><FONT =
color=3D#0000ff><FONT size=3D2><SPAN=20
class=3D299473522-27072005></SPAN></FONT></FONT></FONT>&nbsp;</DIV>
<DIV dir=3Dltr align=3Dleft><FONT face=3DArial><FONT =
color=3D#0000ff><FONT size=3D2><SPAN=20
class=3D299473522-27072005>In addition, VXML 2.1 needs to have the =
information=20
about "duration" and "size" of the&nbsp;recorded audio in the=20
field&nbsp;"waveform-uri" in recog-only header. We would like to propose =
the=20
following fields be added&nbsp;to the&nbsp;recog-only=20
header:</SPAN></FONT></FONT></FONT></DIV>
<DIV dir=3Dltr align=3Dleft><FONT face=3DArial><FONT =
color=3D#0000ff><FONT size=3D2><SPAN=20
class=3D299473522-27072005></SPAN></FONT></FONT></FONT>&nbsp;</DIV>
<DIV dir=3Dltr align=3Dleft><FONT face=3DArial><FONT =
color=3D#0000ff><FONT size=3D2><SPAN=20
class=3D299473522-27072005>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;=20
waveform-duration - the&nbsp;duration of the audio file in=20
milliseconds</SPAN></FONT></FONT></FONT></DIV>
<DIV dir=3Dltr align=3Dleft><FONT face=3DArial><FONT =
color=3D#0000ff><FONT size=3D2><SPAN=20
class=3D299473522-27072005>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbs=
p;waveform-size=20
- the size of the audio file in bytes</SPAN></FONT></FONT></FONT></DIV>
<DIV dir=3Dltr align=3Dleft><FONT face=3DArial><FONT =
color=3D#0000ff><FONT size=3D2><SPAN=20
class=3D299473522-27072005></SPAN></FONT></FONT></FONT>&nbsp;</DIV>
<DIV dir=3Dltr align=3Dleft><FONT face=3DArial><FONT =
color=3D#0000ff><FONT size=3D2><SPAN=20
class=3D299473522-27072005>Thanks.</SPAN></FONT></FONT></FONT></DIV>
<DIV><FONT face=3DArial><FONT color=3D#0000ff><FONT size=3D2><SPAN=20
class=3D299473522-27072005></SPAN></FONT></FONT></FONT>&nbsp;</DIV>
<DIV><FONT face=3DArial><FONT color=3D#0000ff><FONT size=3D2><SPAN=20
class=3D299473522-27072005>Jenny</SPAN></FONT></FONT></FONT></DIV>
<DIV dir=3Dltr align=3Dleft><BR></DIV>
<DIV class=3DOutlookMessageHeader lang=3Den-us dir=3Dltr align=3Dleft>
<HR tabIndex=3D-1>
<FONT face=3DTahoma size=3D2><B>From:</B> speechsc-bounces@ietf.org=20
[mailto:speechsc-bounces@ietf.org] <B>On Behalf Of </B>Saravanan =
Shanmugham=20
(sarvi)<BR><B>Sent:</B> Wednesday, May 04, 2005 9:27 AM<BR><B>To:</B> =
Thomas=20
Gal; Jenny Yao (jyao); speechsc@ietf.org<BR><B>Subject:</B> RE: =
[Speechsc]=20
propose adding "media-type" header field =
inRECOGNIZEorSET-PARAMS/GET-PARAMS=20
methods<BR></FONT><BR></DIV>
<DIV></DIV>
<DIV dir=3Dltr align=3Dleft><FONT face=3DArial color=3D#0000ff =
size=3D2><SPAN=20
class=3D062031516-04052005>I don't see why not, considering all this =
behaviour=20
would apply only if the "Save-Waveform" header is set to true and =
failure=20
to&nbsp;caoture or save the waveform&nbsp;does not necessarily stop the=20
RECOGNIZE operation itself.&nbsp;</SPAN></FONT></DIV>
<DIV dir=3Dltr align=3Dleft><FONT face=3DArial color=3D#0000ff =
size=3D2><SPAN=20
class=3D062031516-04052005></SPAN></FONT>&nbsp;</DIV>
<DIV dir=3Dltr align=3Dleft><FONT face=3DArial color=3D#0000ff =
size=3D2><SPAN=20
class=3D062031516-04052005>Anyone opposed to&nbsp;doing=20
this.&nbsp;</SPAN></FONT></DIV>
<DIV dir=3Dltr align=3Dleft><FONT face=3DArial color=3D#0000ff =
size=3D2><SPAN=20
class=3D062031516-04052005></SPAN></FONT>&nbsp;</DIV>
<DIV dir=3Dltr align=3Dleft><FONT face=3DArial color=3D#0000ff =
size=3D2><SPAN=20
class=3D062031516-04052005>Thanks,</SPAN></FONT></DIV>
<DIV dir=3Dltr align=3Dleft><FONT face=3DArial color=3D#0000ff =
size=3D2><SPAN=20
class=3D062031516-04052005>Sarvi</SPAN></FONT></DIV><BR>
<BLOCKQUOTE dir=3Dltr=20
style=3D"PADDING-LEFT: 5px; MARGIN-LEFT: 5px; BORDER-LEFT: #0000ff 2px =
solid; MARGIN-RIGHT: 0px">
  <DIV class=3DOutlookMessageHeader lang=3Den-us dir=3Dltr align=3Dleft>
  <HR tabIndex=3D-1>
  <FONT face=3DTahoma size=3D2><B>From:</B> speechsc-bounces@ietf.org=20
  [mailto:speechsc-bounces@ietf.org] <B>On Behalf Of </B>Thomas=20
  Gal<BR><B>Sent:</B> Friday, April 29, 2005 9:47 AM<BR><B>To:</B> =
'Jenny Yao=20
  (jyao)'; speechsc@ietf.org<BR><B>Subject:</B> RE: [Speechsc] propose =
adding=20
  "media-type" header field in RECOGNIZEorSET-PARAMS/GET-PARAMS=20
  methods<BR></FONT><BR></DIV>
  <DIV></DIV>
  <DIV class=3DSection1>
  <P class=3DMsoNormal><FONT face=3DArial color=3Dnavy size=3D2><SPAN=20
  style=3D"FONT-SIZE: 10pt; COLOR: navy; FONT-FAMILY: Arial">Though this =

  information could probably be inferred from the filename extension, If =
we are=20
  going to follow/emulate the RECORD methodology than it should also be=20
  available as a MIME body to the RECOGNITION COMPLETE/STOP events as =
well.=20
  Otherwise I agree completely.<o:p></o:p></SPAN></FONT></P>
  <P class=3DMsoNormal><FONT face=3DArial color=3Dnavy size=3D2><SPAN=20
  style=3D"FONT-SIZE: 10pt; COLOR: navy; FONT-FAMILY: =
Arial"><o:p>&nbsp;</o:p></SPAN></FONT></P>
  <P class=3DMsoNormal><FONT face=3DArial color=3Dnavy size=3D2><SPAN=20
  style=3D"FONT-SIZE: 10pt; COLOR: navy; FONT-FAMILY: =
Arial">-Tom<o:p></o:p></SPAN></FONT></P>
  <P class=3DMsoNormal><FONT face=3DArial color=3Dnavy size=3D2><SPAN=20
  style=3D"FONT-SIZE: 10pt; COLOR: navy; FONT-FAMILY: =
Arial"><o:p>&nbsp;</o:p></SPAN></FONT></P>
  <DIV>
  <DIV class=3DMsoNormal style=3D"TEXT-ALIGN: center" =
align=3Dcenter><FONT=20
  face=3D"Times New Roman" size=3D3><SPAN style=3D"FONT-SIZE: 12pt">
  <HR tabIndex=3D-1 align=3Dcenter width=3D"100%" SIZE=3D2>
  </SPAN></FONT></DIV>
  <P class=3DMsoNormal><B><FONT face=3DTahoma size=3D2><SPAN=20
  style=3D"FONT-WEIGHT: bold; FONT-SIZE: 10pt; FONT-FAMILY: =
Tahoma">From:</SPAN></FONT></B><FONT=20
  face=3DTahoma size=3D2><SPAN style=3D"FONT-SIZE: 10pt; FONT-FAMILY: =
Tahoma">=20
  speechsc-bounces@ietf.org [mailto:speechsc-bounces@ietf.org] <B><SPAN=20
  style=3D"FONT-WEIGHT: bold">On Behalf Of </SPAN></B>Jenny Yao =
(jyao)<BR><B><SPAN=20
  style=3D"FONT-WEIGHT: bold">Sent:</SPAN></B> Friday, April 29, 2005 =
8:58=20
  AM<BR><B><SPAN style=3D"FONT-WEIGHT: bold">To:</SPAN></B>=20
  speechsc@ietf.org<BR><B><SPAN style=3D"FONT-WEIGHT: =
bold">Subject:</SPAN></B>=20
  [Speechsc] propose adding "media-type" header field in RECOGNIZE=20
  orSET-PARAMS/GET-PARAMS methods</SPAN></FONT><o:p></o:p></P></DIV>
  <P class=3DMsoNormal><FONT face=3D"Times New Roman" size=3D3><SPAN=20
  style=3D"FONT-SIZE: 12pt"><o:p>&nbsp;</o:p></SPAN></FONT></P>
  <DIV>
  <DIV>
  <P class=3DMsoNormal><FONT face=3D"Times New Roman" size=3D2><SPAN=20
  style=3D"FONT-SIZE: 10pt">&nbsp;<o:p></o:p></SPAN></FONT></P></DIV>
  <DIV>
  <P class=3DMsoNormal><FONT face=3D"Times New Roman" size=3D2><SPAN=20
  style=3D"FONT-SIZE: 10pt">&nbsp;<o:p></o:p></SPAN></FONT></P></DIV>
  <DIV>
  <P class=3DMsoNormal><FONT face=3DArial size=3D2><SPAN=20
  style=3D"FONT-SIZE: 10pt; FONT-FAMILY: Arial">VoiceXML =
2.1&nbsp;specifies=20
  "recording user utterances while attempting recognition". This can be =
done=20
  through recog-only-header "save-waveform" and "waveform-uri" in MRCP =
V1 and=20
  V2. VoiceXML 2.1 also uses </SPAN></FONT><EM><I><FONT =
face=3DArial><SPAN=20
  style=3D"FONT-FAMILY: =
Arial">recordutterancetype</SPAN></FONT></I></EM><FONT=20
  face=3DArial><SPAN style=3D"FONT-FAMILY: Arial"> </SPAN></FONT><FONT =
face=3DArial=20
  size=3D2><SPAN style=3D"FONT-SIZE: 10pt; FONT-FAMILY: Arial">property =
to specify=20
  the media format of the result recording. However, there is =
no&nbsp;way=20
  in&nbsp;MRCP to pass the required media format to&nbsp;the =
server.&nbsp;We=20
  propose&nbsp;using the "Media-Type"&nbsp;header field, currently =
defined for=20
  the recording resource, in the RECOGNIZE method or the =
SET-PARAMS/GET-PARAMS=20
  methods to specify&nbsp;a media type.&nbsp;The Save-Waveform =
carries&nbsp;a=20
  URI pointing to the audio captured during recognition. The captured=20
  audio&nbsp;SHOULD be saved with this media-type.</SPAN></FONT><FONT=20
  size=3D2><SPAN style=3D"FONT-SIZE: =
10pt"><o:p></o:p></SPAN></FONT></P></DIV>
  <DIV>
  <P class=3DMsoNormal><FONT face=3D"Times New Roman" size=3D2><SPAN=20
  style=3D"FONT-SIZE: 10pt">&nbsp;<o:p></o:p></SPAN></FONT></P></DIV>
  <DIV>
  <P class=3DMsoNormal><FONT face=3DArial size=3D2><SPAN=20
  style=3D"FONT-SIZE: 10pt; FONT-FAMILY: =
Arial">Thanks.</SPAN></FONT><FONT=20
  size=3D2><SPAN style=3D"FONT-SIZE: =
10pt"><o:p></o:p></SPAN></FONT></P></DIV>
  <DIV>
  <P class=3DMsoNormal><FONT face=3D"Times New Roman" size=3D2><SPAN=20
  style=3D"FONT-SIZE: 10pt">&nbsp;<o:p></o:p></SPAN></FONT></P></DIV>
  <DIV>
  <P class=3DMsoNormal><FONT face=3DArial size=3D2><SPAN=20
  style=3D"FONT-SIZE: 10pt; FONT-FAMILY: Arial">Jenny</SPAN></FONT><FONT =

  size=3D2><SPAN=20
  style=3D"FONT-SIZE: =
10pt"><o:p></o:p></SPAN></FONT></P></DIV></DIV></DIV></BLOCKQUOTE></BODY>=
</HTML>

------_=_NextPart_001_01C592FE.127827D7--


--===============0856537083==
Content-Type: text/plain; charset="us-ascii"
MIME-Version: 1.0
Content-Transfer-Encoding: 7bit
Content-Disposition: inline

_______________________________________________
Speechsc mailing list
Speechsc@ietf.org
https://www1.ietf.org/mailman/listinfo/speechsc

--===============0856537083==--




From speechsc-bounces@ietf.org Thu Jul 28 05:27:23 2005
Received: from localhost.localdomain ([127.0.0.1] helo=megatron.ietf.org)
	by megatron.ietf.org with esmtp (Exim 4.32)
	id 1Dy4fj-0006JU-Eq; Thu, 28 Jul 2005 05:27:23 -0400
Received: from odin.ietf.org ([132.151.1.176] helo=ietf.org)
	by megatron.ietf.org with esmtp (Exim 4.32) id 1Dy4fh-0006Ic-A3
	for speechsc@megatron.ietf.org; Thu, 28 Jul 2005 05:27:21 -0400
Received: from ietf-mx.ietf.org (ietf-mx [132.151.6.1])
	by ietf.org (8.9.1a/8.9.1a) with ESMTP id FAA09372
	for <speechsc@ietf.org>; Thu, 28 Jul 2005 05:27:18 -0400 (EDT)
Received: from dns1.tilab.com ([163.162.42.4])
	by ietf-mx.ietf.org with esmtp (Exim 4.43) id 1Dy5B5-00065D-HX
	for speechsc@ietf.org; Thu, 28 Jul 2005 05:59:50 -0400
Received: from iowa2k01a.cselt.it ([163.162.242.201])
	by dns1.cselt.it (PMDF V6.0-025 #38895)
	with ESMTP id <0IKB00A37ZG9UG@dns1.cselt.it> for speechsc@ietf.org; Thu,
	28 Jul 2005 11:24:09 +0200 (MEST)
Received: from EXC01C.cselt.it ([163.162.4.218]) by iowa2k01a.cselt.it with
	Microsoft SMTPSVC(6.0.3790.211); Thu, 28 Jul 2005 11:30:35 +0200
Date: Thu, 28 Jul 2005 11:27:03 +0200
From: Baggia Paolo <Paolo.Baggia@LOQUENDO.COM>
Subject: RE: [Speechsc] propose adding "media-type" header
	fieldinRECOGNIZEorSET-PARAMS/GET-PARAMS methods
To: "Jenny Yao (jyao)" <jyao@cisco.com>, speechsc@ietf.org
Message-id: <E5880434292FCB448F00BDAEE44A60D07DC53F@EXC01C.cselt.it>
MIME-version: 1.0
X-MIMEOLE: Produced By Microsoft MimeOLE V6.00.3790.326
Importance: normal
Priority: normal
Thread-Topic: [Speechsc] propose adding "media-type" header
	fieldinRECOGNIZEorSET-PARAMS/GET-PARAMS methods
thread-index: AcVM1EeF6F3QJhBHRniGmUHADY5x4wAA+UewAPsQNVAQjcCwwAAABgpQABaCUWA=
Content-Class: urn:content-classes:message
X-OriginalArrivalTime: 28 Jul 2005 09:30:35.0296 (UTC)
	FILETIME=[01E21A00:01C59357]
X-Spam-Score: 0.4 (/)
X-Scan-Signature: 9d7e8d783239e9f0c425c823a9c950ff
Cc: Baggia Paolo <Paolo.Baggia@LOQUENDO.COM>
X-BeenThere: speechsc@ietf.org
X-Mailman-Version: 2.1.5
Precedence: list
List-Id: Speech Services Control Working Group <speechsc.ietf.org>
List-Unsubscribe: <https://www1.ietf.org/mailman/listinfo/speechsc>,
	<mailto:speechsc-request@ietf.org?subject=unsubscribe>
List-Post: <mailto:speechsc@ietf.org>
List-Help: <mailto:speechsc-request@ietf.org?subject=help>
List-Subscribe: <https://www1.ietf.org/mailman/listinfo/speechsc>,
	<mailto:speechsc-request@ietf.org?subject=subscribe>
Content-Type: multipart/mixed; boundary="===============1038945621=="
Sender: speechsc-bounces@ietf.org
Errors-To: speechsc-bounces@ietf.org

This is a multi-part message in MIME format.

--===============1038945621==
Content-type: multipart/alternative;
	boundary="----_=_NextPart_001_01C59356.83E579C9"
Content-transfer-encoding: 7bit
Content-Class: urn:content-classes:message

This is a multi-part message in MIME format.

------_=_NextPart_001_01C59356.83E579C9
Content-Type: text/plain;
	charset="iso-8859-1"
Content-Transfer-Encoding: quoted-printable

Jenny,
=20
As you well explain, these two new fields are essential for implementing =
VXML 2.1 on top of MRCPv2.
We think to add them is very important for MRCPv2, but there are two =
considerations we would=20
like to add.=20
=20
These two fields should at least be added to the "recorder-header", =
because also the recorder=20
needs to specify to the VoiceXML interpreter the size and duration of =
the recorded speech.
=20
Another option would be to add them to all the resources, because they =
may benefit of these fields
related to dump of audio ("save-waveform", "waveform-uri", =
"waveform-duration",  "waveform-size"):
- recognizer need them for VXML 2.1 support (need media-type, =
"waveform-duration" & "waveform-size"):
- recorder for both VXML 2.0 and 2.1 support (need"waveform-duration", & =
"waveform-size"):
- verifier possibly especially if the verification is done independet =
from recognition (if we think this
  is useful)
- sinthesizer would benefit for dumping the audio in a file, many =
platform support this feature too
  (need media-type and all the recording field headers)
=20
What do you think?
=20
Paolo.

-----Original Message-----
From: speechsc-bounces@ietf.org [mailto:speechsc-bounces@ietf.org]On =
Behalf Of Jenny Yao (jyao)
Sent: Thursday, July 28, 2005 12:54 AM
To: speechsc@ietf.org
Subject: RE: [Speechsc] propose adding "media-type" header =
fieldinRECOGNIZEorSET-PARAMS/GET-PARAMS methods


There appears to be no opposition regarding the media-type header field =
request that we sent earlier. So, can we request this feature be =
included in the next v2 draft?=20
=20
In addition, VXML 2.1 needs to have the information about "duration" and =
"size" of the recorded audio in the field "waveform-uri" in recog-only =
header. We would like to propose the following fields be added to the =
recog-only header:
=20
        waveform-duration - the duration of the audio file in =
milliseconds
        waveform-size - the size of the audio file in bytes
=20
Thanks.
=20
Jenny

  _____ =20

From: speechsc-bounces@ietf.org [mailto:speechsc-bounces@ietf.org] On =
Behalf Of Saravanan Shanmugham (sarvi)
Sent: Wednesday, May 04, 2005 9:27 AM
To: Thomas Gal; Jenny Yao (jyao); speechsc@ietf.org
Subject: RE: [Speechsc] propose adding "media-type" header field =
inRECOGNIZEorSET-PARAMS/GET-PARAMS methods


I don't see why not, considering all this behaviour would apply only if =
the "Save-Waveform" header is set to true and failure to caoture or save =
the waveform does not necessarily stop the RECOGNIZE operation itself.=20
=20
Anyone opposed to doing this.=20
=20
Thanks,
Sarvi


  _____ =20

From: speechsc-bounces@ietf.org [mailto:speechsc-bounces@ietf.org] On =
Behalf Of Thomas Gal
Sent: Friday, April 29, 2005 9:47 AM
To: 'Jenny Yao (jyao)'; speechsc@ietf.org
Subject: RE: [Speechsc] propose adding "media-type" header field in =
RECOGNIZEorSET-PARAMS/GET-PARAMS methods



Though this information could probably be inferred from the filename =
extension, If we are going to follow/emulate the RECORD methodology than =
it should also be available as a MIME body to the RECOGNITION =
COMPLETE/STOP events as well. Otherwise I agree completely.

=20

-Tom

=20


  _____ =20


From: speechsc-bounces@ietf.org [mailto:speechsc-bounces@ietf.org] On =
Behalf Of Jenny Yao (jyao)
Sent: Friday, April 29, 2005 8:58 AM
To: speechsc@ietf.org
Subject: [Speechsc] propose adding "media-type" header field in =
RECOGNIZE orSET-PARAMS/GET-PARAMS methods

=20

=20

=20

VoiceXML 2.1 specifies "recording user utterances while attempting =
recognition". This can be done through recog-only-header "save-waveform" =
and "waveform-uri" in MRCP V1 and V2. VoiceXML 2.1 also uses =
recordutterancetype property to specify the media format of the result =
recording. However, there is no way in MRCP to pass the required media =
format to the server. We propose using the "Media-Type" header field, =
currently defined for the recording resource, in the RECOGNIZE method or =
the SET-PARAMS/GET-PARAMS methods to specify a media type. The =
Save-Waveform carries a URI pointing to the audio captured during =
recognition. The captured audio SHOULD be saved with this media-type.

=20

Thanks.

=20

Jenny



Gruppo Telecom Italia - Direzione e coordinamento di Telecom Italia =
S.p.A.

=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=
=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=
=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D
CONFIDENTIALITY NOTICE
This message and its attachments are addressed solely to the persons
above and may contain confidential information. If you have received
the message in error, be informed that any use of the content hereof
is prohibited. Please return it immediately to the sender and delete
the message. Should you have any questions, please send an e_mail to=20
MailAdmin@tilab.com. Thank you
=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=
=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=
=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D

------_=_NextPart_001_01C59356.83E579C9
Content-Type: text/html;
	charset="iso-8859-1"
Content-Transfer-Encoding: quoted-printable

<!DOCTYPE HTML PUBLIC "-//W3C//DTD HTML 4.0 Transitional//EN">
<HTML xmlns=3D"http://www.w3.org/TR/REC-html40" xmlns:v =3D=20
"urn:schemas-microsoft-com:vml" xmlns:o =3D=20
"urn:schemas-microsoft-com:office:office" xmlns:w =3D=20
"urn:schemas-microsoft-com:office:word"><HEAD>
<META HTTP-EQUIV=3D"Content-Type" CONTENT=3D"text/html; =
charset=3Diso-8859-1">


<META content=3D"MSHTML 6.00.2800.1505" name=3DGENERATOR><!--[if !mso]>
<STYLE>v\:* {
	BEHAVIOR: url(#default#VML)
}
o\:* {
	BEHAVIOR: url(#default#VML)
}
w\:* {
	BEHAVIOR: url(#default#VML)
}
.shape {
	BEHAVIOR: url(#default#VML)
}
</STYLE>
<![endif]-->
<STYLE>@font-face {
	font-family: Tahoma;
}
@page Section1 {size: 8.5in 11.0in; margin: 1.0in 1.25in 1.0in 1.25in; }
P.MsoNormal {
	FONT-SIZE: 12pt; MARGIN: 0in 0in 0pt; FONT-FAMILY: "Times New Roman"
}
LI.MsoNormal {
	FONT-SIZE: 12pt; MARGIN: 0in 0in 0pt; FONT-FAMILY: "Times New Roman"
}
DIV.MsoNormal {
	FONT-SIZE: 12pt; MARGIN: 0in 0in 0pt; FONT-FAMILY: "Times New Roman"
}
A:link {
	COLOR: blue; TEXT-DECORATION: underline
}
SPAN.MsoHyperlink {
	COLOR: blue; TEXT-DECORATION: underline
}
A:visited {
	COLOR: purple; TEXT-DECORATION: underline
}
SPAN.MsoHyperlinkFollowed {
	COLOR: purple; TEXT-DECORATION: underline
}
PRE {
	FONT-SIZE: 10pt; MARGIN: 0in 0in 0pt; FONT-FAMILY: "Courier New"
}
SPAN.EmailStyle18 {
	COLOR: navy; FONT-FAMILY: Arial; mso-style-type: personal-reply
}
DIV.Section1 {
	page: Section1
}
</STYLE>
</HEAD>
<BODY lang=3DEN-US vLink=3Dpurple link=3Dblue>
<DIV>
<DIV><SPAN class=3D065225307-28072005><FONT face=3DArial color=3D#0000ff =

size=3D2>Jenny,</FONT></SPAN></DIV>
<DIV><SPAN class=3D065225307-28072005><FONT face=3DArial color=3D#0000ff =

size=3D2></FONT></SPAN>&nbsp;</DIV>
<DIV><SPAN class=3D065225307-28072005><FONT face=3DArial color=3D#0000ff =
size=3D2>As you=20
well explain, these two new fields are essential for implementing VXML =
2.1 on=20
top of MRCPv2.</FONT></SPAN></DIV>
<DIV><SPAN class=3D065225307-28072005><FONT face=3DArial color=3D#0000ff =
size=3D2>We=20
think to add them is very important for MRCPv2, but there&nbsp;<SPAN=20
class=3D505172009-28072005>are </SPAN>two considerations we would=20
</FONT></SPAN></DIV>
<DIV><FONT face=3DArial><FONT color=3D#0000ff><FONT size=3D2><SPAN=20
class=3D065225307-28072005>like to add. =
</SPAN></FONT></FONT></FONT></DIV>
<DIV><FONT face=3DArial><FONT color=3D#0000ff><FONT size=3D2><SPAN=20
class=3D065225307-28072005></SPAN></FONT></FONT></FONT>&nbsp;</DIV>
<DIV><FONT><FONT><FONT face=3DArial><FONT color=3D#0000ff><FONT =
size=3D2><SPAN=20
class=3D065225307-28072005>These&nbsp;<SPAN =
class=3D505172009-28072005>two=20
</SPAN>fields should&nbsp;at least</SPAN><SPAN=20
class=3D065225307-28072005>&nbsp;<SPAN class=3D505172009-28072005>be =
</SPAN>added to=20
the "recorder-header"<SPAN class=3D505172009-28072005>, </SPAN>because =
also the=20
recorder </SPAN></FONT></FONT></FONT></FONT></FONT></DIV>
<DIV><FONT face=3DArial><FONT color=3D#0000ff><FONT size=3D2><SPAN=20
class=3D065225307-28072005>need<SPAN class=3D505172009-28072005>s</SPAN> =
to specify=20
to the VoiceXML interpreter </SPAN><SPAN class=3D065225307-28072005>the =
size and=20
duration of the recorded speech.</SPAN></FONT></FONT></FONT></DIV>
<DIV><SPAN class=3D065225307-28072005><FONT face=3DArial color=3D#0000ff =

size=3D2></FONT></SPAN>&nbsp;</DIV>
<DIV><SPAN class=3D065225307-28072005><FONT face=3DArial color=3D#0000ff =

size=3D2>Another option would be to add them&nbsp;<SPAN=20
class=3D505172009-28072005>to&nbsp;</SPAN>all the&nbsp;<SPAN=20
class=3D505172009-28072005>resource</SPAN>s<SPAN =
class=3D505172009-28072005>,=20
because&nbsp;they may</SPAN>&nbsp;benefit of these =
fields</FONT></SPAN></DIV>
<DIV><SPAN class=3D065225307-28072005><FONT face=3DArial color=3D#0000ff =

size=3D2>related to dump of audio ("save-waveform", "waveform-uri",=20
"waveform-duration",&nbsp; "waveform-size"):</FONT></SPAN></DIV>
<DIV><SPAN class=3D065225307-28072005><FONT face=3DArial><FONT =
color=3D#0000ff><FONT=20
size=3D2>- recognizer need them for VXML 2.1 support<SPAN=20
class=3D505172009-28072005> (need media-type, "waveform-duration" &amp;=20
"waveform-size"):</SPAN></FONT></FONT></FONT></SPAN></DIV>
<DIV><SPAN class=3D065225307-28072005><FONT face=3DArial><FONT =
color=3D#0000ff><FONT=20
size=3D2>- recorder for both VXML 2.0 and 2.1 support<SPAN=20
class=3D505172009-28072005> (need"waveform-duration",=20
&amp;&nbsp;"waveform-size"):</SPAN></FONT></FONT></FONT></SPAN></DIV>
<DIV><FONT face=3DArial><FONT color=3D#0000ff><FONT size=3D2><SPAN=20
class=3D065225307-28072005>- verifier possibly especially if the =
verification is=20
done&nbsp;i</SPAN><SPAN class=3D065225307-28072005>ndependet from =
recognition<SPAN=20
class=3D505172009-28072005> (if we think=20
this</SPAN></SPAN></FONT></FONT></FONT></DIV>
<DIV><FONT face=3DArial><FONT color=3D#0000ff><FONT size=3D2><SPAN=20
class=3D065225307-28072005><SPAN class=3D505172009-28072005>&nbsp; is=20
useful)</SPAN></SPAN></FONT></FONT></FONT></DIV>
<DIV><SPAN class=3D065225307-28072005><FONT face=3DArial color=3D#0000ff =
size=3D2>-=20
sinthesizer would benefit for dumping the audio in a file<SPAN=20
class=3D505172009-28072005>, </SPAN>many platform support this feature=20
too</FONT></SPAN></DIV>
<DIV><SPAN class=3D065225307-28072005><SPAN =
class=3D505172009-28072005><FONT=20
face=3DArial color=3D#0000ff size=3D2>&nbsp; (need media-type and all =
the recording=20
field headers)</FONT></SPAN></SPAN></DIV>
<DIV><SPAN class=3D065225307-28072005><FONT face=3DArial color=3D#0000ff =

size=3D2></FONT></SPAN>&nbsp;</DIV>
<DIV><SPAN class=3D065225307-28072005><FONT face=3DArial color=3D#0000ff =
size=3D2>What=20
do you think?</FONT></SPAN></DIV>
<DIV><SPAN class=3D065225307-28072005><FONT face=3DArial color=3D#0000ff =

size=3D2></FONT></SPAN>&nbsp;</DIV>
<DIV><SPAN class=3D065225307-28072005><FONT face=3DArial color=3D#0000ff =

size=3D2>Paolo.</FONT></SPAN></DIV></DIV>
<BLOCKQUOTE dir=3Dltr style=3D"MARGIN-RIGHT: 0px">
  <DIV class=3DOutlookMessageHeader dir=3Dltr align=3Dleft><FONT =
face=3DTahoma=20
  size=3D2>-----Original Message-----<BR><B>From:</B> =
speechsc-bounces@ietf.org=20
  [mailto:speechsc-bounces@ietf.org]<B>On Behalf Of </B>Jenny Yao=20
  (jyao)<BR><B>Sent:</B> Thursday, July 28, 2005 12:54 AM<BR><B>To:</B>=20
  speechsc@ietf.org<BR><B>Subject:</B> RE: [Speechsc] propose adding=20
  "media-type" header fieldinRECOGNIZEorSET-PARAMS/GET-PARAMS=20
  methods<BR><BR></FONT></DIV>
  <DIV dir=3Dltr align=3Dleft><SPAN =
class=3D299473522-27072005></SPAN><FONT=20
  face=3DArial><FONT color=3D#0000ff><FONT size=3D2><SPAN=20
  class=3D299473522-27072005>There appears to be no opposition =
r</SPAN><SPAN=20
  class=3D299473522-27072005>egarding the media-type header field =
request=20
  that&nbsp;we sent earlier. So, can we request this feature be included =
in the=20
  next v2 draft? </SPAN></FONT></FONT></FONT></DIV>
  <DIV dir=3Dltr align=3Dleft><FONT face=3DArial><FONT =
color=3D#0000ff><FONT=20
  size=3D2><SPAN =
class=3D299473522-27072005></SPAN></FONT></FONT></FONT>&nbsp;</DIV>
  <DIV dir=3Dltr align=3Dleft><FONT face=3DArial><FONT =
color=3D#0000ff><FONT=20
  size=3D2><SPAN class=3D299473522-27072005>In addition, VXML 2.1 needs =
to have the=20
  information about "duration" and "size" of the&nbsp;recorded audio in =
the=20
  field&nbsp;"waveform-uri" in recog-only header. We would like to =
propose the=20
  following fields be added&nbsp;to the&nbsp;recog-only=20
  header:</SPAN></FONT></FONT></FONT></DIV>
  <DIV dir=3Dltr align=3Dleft><FONT face=3DArial><FONT =
color=3D#0000ff><FONT=20
  size=3D2><SPAN =
class=3D299473522-27072005></SPAN></FONT></FONT></FONT>&nbsp;</DIV>
  <DIV dir=3Dltr align=3Dleft><FONT face=3DArial><FONT =
color=3D#0000ff><FONT=20
  size=3D2><SPAN=20
  class=3D299473522-27072005>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;=20
  waveform-duration - the&nbsp;duration of the audio file in=20
  milliseconds</SPAN></FONT></FONT></FONT></DIV>
  <DIV dir=3Dltr align=3Dleft><FONT face=3DArial><FONT =
color=3D#0000ff><FONT=20
  size=3D2><SPAN=20
  =
class=3D299473522-27072005>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbs=
p;waveform-size=20
  - the size of the audio file in =
bytes</SPAN></FONT></FONT></FONT></DIV>
  <DIV dir=3Dltr align=3Dleft><FONT face=3DArial><FONT =
color=3D#0000ff><FONT=20
  size=3D2><SPAN =
class=3D299473522-27072005></SPAN></FONT></FONT></FONT>&nbsp;</DIV>
  <DIV dir=3Dltr align=3Dleft><FONT face=3DArial><FONT =
color=3D#0000ff><FONT=20
  size=3D2><SPAN=20
class=3D299473522-27072005>Thanks.</SPAN></FONT></FONT></FONT></DIV>
  <DIV><FONT face=3DArial><FONT color=3D#0000ff><FONT size=3D2><SPAN=20
  class=3D299473522-27072005></SPAN></FONT></FONT></FONT>&nbsp;</DIV>
  <DIV><FONT face=3DArial><FONT color=3D#0000ff><FONT size=3D2><SPAN=20
  class=3D299473522-27072005>Jenny</SPAN></FONT></FONT></FONT></DIV>
  <DIV dir=3Dltr align=3Dleft><BR></DIV>
  <DIV class=3DOutlookMessageHeader lang=3Den-us dir=3Dltr align=3Dleft>
  <HR tabIndex=3D-1>
  <FONT face=3DTahoma size=3D2><B>From:</B> speechsc-bounces@ietf.org=20
  [mailto:speechsc-bounces@ietf.org] <B>On Behalf Of </B>Saravanan =
Shanmugham=20
  (sarvi)<BR><B>Sent:</B> Wednesday, May 04, 2005 9:27 AM<BR><B>To:</B> =
Thomas=20
  Gal; Jenny Yao (jyao); speechsc@ietf.org<BR><B>Subject:</B> RE: =
[Speechsc]=20
  propose adding "media-type" header field =
inRECOGNIZEorSET-PARAMS/GET-PARAMS=20
  methods<BR></FONT><BR></DIV>
  <DIV></DIV>
  <DIV dir=3Dltr align=3Dleft><FONT face=3DArial color=3D#0000ff =
size=3D2><SPAN=20
  class=3D062031516-04052005>I don't see why not, considering all this =
behaviour=20
  would apply only if the "Save-Waveform" header is set to true and =
failure=20
  to&nbsp;caoture or save the waveform&nbsp;does not necessarily stop =
the=20
  RECOGNIZE operation itself.&nbsp;</SPAN></FONT></DIV>
  <DIV dir=3Dltr align=3Dleft><FONT face=3DArial color=3D#0000ff =
size=3D2><SPAN=20
  class=3D062031516-04052005></SPAN></FONT>&nbsp;</DIV>
  <DIV dir=3Dltr align=3Dleft><FONT face=3DArial color=3D#0000ff =
size=3D2><SPAN=20
  class=3D062031516-04052005>Anyone opposed to&nbsp;doing=20
  this.&nbsp;</SPAN></FONT></DIV>
  <DIV dir=3Dltr align=3Dleft><FONT face=3DArial color=3D#0000ff =
size=3D2><SPAN=20
  class=3D062031516-04052005></SPAN></FONT>&nbsp;</DIV>
  <DIV dir=3Dltr align=3Dleft><FONT face=3DArial color=3D#0000ff =
size=3D2><SPAN=20
  class=3D062031516-04052005>Thanks,</SPAN></FONT></DIV>
  <DIV dir=3Dltr align=3Dleft><FONT face=3DArial color=3D#0000ff =
size=3D2><SPAN=20
  class=3D062031516-04052005>Sarvi</SPAN></FONT></DIV><BR>
  <BLOCKQUOTE dir=3Dltr=20
  style=3D"PADDING-LEFT: 5px; MARGIN-LEFT: 5px; BORDER-LEFT: #0000ff 2px =
solid; MARGIN-RIGHT: 0px">
    <DIV class=3DOutlookMessageHeader lang=3Den-us dir=3Dltr =
align=3Dleft>
    <HR tabIndex=3D-1>
    <FONT face=3DTahoma size=3D2><B>From:</B> speechsc-bounces@ietf.org=20
    [mailto:speechsc-bounces@ietf.org] <B>On Behalf Of </B>Thomas=20
    Gal<BR><B>Sent:</B> Friday, April 29, 2005 9:47 AM<BR><B>To:</B> =
'Jenny Yao=20
    (jyao)'; speechsc@ietf.org<BR><B>Subject:</B> RE: [Speechsc] propose =
adding=20
    "media-type" header field in RECOGNIZEorSET-PARAMS/GET-PARAMS=20
    methods<BR></FONT><BR></DIV>
    <DIV></DIV>
    <DIV class=3DSection1>
    <P class=3DMsoNormal><FONT face=3DArial color=3Dnavy size=3D2><SPAN=20
    style=3D"FONT-SIZE: 10pt; COLOR: navy; FONT-FAMILY: Arial">Though =
this=20
    information could probably be inferred from the filename extension, =
If we=20
    are going to follow/emulate the RECORD methodology than it should =
also be=20
    available as a MIME body to the RECOGNITION COMPLETE/STOP events as =
well.=20
    Otherwise I agree completely.<o:p></o:p></SPAN></FONT></P>
    <P class=3DMsoNormal><FONT face=3DArial color=3Dnavy size=3D2><SPAN=20
    style=3D"FONT-SIZE: 10pt; COLOR: navy; FONT-FAMILY: =
Arial"><o:p>&nbsp;</o:p></SPAN></FONT></P>
    <P class=3DMsoNormal><FONT face=3DArial color=3Dnavy size=3D2><SPAN=20
    style=3D"FONT-SIZE: 10pt; COLOR: navy; FONT-FAMILY: =
Arial">-Tom<o:p></o:p></SPAN></FONT></P>
    <P class=3DMsoNormal><FONT face=3DArial color=3Dnavy size=3D2><SPAN=20
    style=3D"FONT-SIZE: 10pt; COLOR: navy; FONT-FAMILY: =
Arial"><o:p>&nbsp;</o:p></SPAN></FONT></P>
    <DIV>
    <DIV class=3DMsoNormal style=3D"TEXT-ALIGN: center" =
align=3Dcenter><FONT=20
    face=3D"Times New Roman" size=3D3><SPAN style=3D"FONT-SIZE: 12pt">
    <HR tabIndex=3D-1 align=3Dcenter width=3D"100%" SIZE=3D2>
    </SPAN></FONT></DIV>
    <P class=3DMsoNormal><B><FONT face=3DTahoma size=3D2><SPAN=20
    style=3D"FONT-WEIGHT: bold; FONT-SIZE: 10pt; FONT-FAMILY: =
Tahoma">From:</SPAN></FONT></B><FONT=20
    face=3DTahoma size=3D2><SPAN style=3D"FONT-SIZE: 10pt; FONT-FAMILY: =
Tahoma">=20
    speechsc-bounces@ietf.org [mailto:speechsc-bounces@ietf.org] =
<B><SPAN=20
    style=3D"FONT-WEIGHT: bold">On Behalf Of </SPAN></B>Jenny Yao=20
    (jyao)<BR><B><SPAN style=3D"FONT-WEIGHT: bold">Sent:</SPAN></B> =
Friday, April=20
    29, 2005 8:58 AM<BR><B><SPAN style=3D"FONT-WEIGHT: =
bold">To:</SPAN></B>=20
    speechsc@ietf.org<BR><B><SPAN style=3D"FONT-WEIGHT: =
bold">Subject:</SPAN></B>=20
    [Speechsc] propose adding "media-type" header field in RECOGNIZE=20
    orSET-PARAMS/GET-PARAMS methods</SPAN></FONT><o:p></o:p></P></DIV>
    <P class=3DMsoNormal><FONT face=3D"Times New Roman" size=3D3><SPAN=20
    style=3D"FONT-SIZE: 12pt"><o:p>&nbsp;</o:p></SPAN></FONT></P>
    <DIV>
    <DIV>
    <P class=3DMsoNormal><FONT face=3D"Times New Roman" size=3D2><SPAN=20
    style=3D"FONT-SIZE: 10pt">&nbsp;<o:p></o:p></SPAN></FONT></P></DIV>
    <DIV>
    <P class=3DMsoNormal><FONT face=3D"Times New Roman" size=3D2><SPAN=20
    style=3D"FONT-SIZE: 10pt">&nbsp;<o:p></o:p></SPAN></FONT></P></DIV>
    <DIV>
    <P class=3DMsoNormal><FONT face=3DArial size=3D2><SPAN=20
    style=3D"FONT-SIZE: 10pt; FONT-FAMILY: Arial">VoiceXML =
2.1&nbsp;specifies=20
    "recording user utterances while attempting recognition". This can =
be done=20
    through recog-only-header "save-waveform" and "waveform-uri" in MRCP =
V1 and=20
    V2. VoiceXML 2.1 also uses </SPAN></FONT><EM><I><FONT =
face=3DArial><SPAN=20
    style=3D"FONT-FAMILY: =
Arial">recordutterancetype</SPAN></FONT></I></EM><FONT=20
    face=3DArial><SPAN style=3D"FONT-FAMILY: Arial"> </SPAN></FONT><FONT =
face=3DArial=20
    size=3D2><SPAN style=3D"FONT-SIZE: 10pt; FONT-FAMILY: =
Arial">property to specify=20
    the media format of the result recording. However, there is =
no&nbsp;way=20
    in&nbsp;MRCP to pass the required media format to&nbsp;the =
server.&nbsp;We=20
    propose&nbsp;using the "Media-Type"&nbsp;header field, currently =
defined for=20
    the recording resource, in the RECOGNIZE method or the =
SET-PARAMS/GET-PARAMS=20
    methods to specify&nbsp;a media type.&nbsp;The Save-Waveform =
carries&nbsp;a=20
    URI pointing to the audio captured during recognition. The captured=20
    audio&nbsp;SHOULD be saved with this media-type.</SPAN></FONT><FONT=20
    size=3D2><SPAN style=3D"FONT-SIZE: =
10pt"><o:p></o:p></SPAN></FONT></P></DIV>
    <DIV>
    <P class=3DMsoNormal><FONT face=3D"Times New Roman" size=3D2><SPAN=20
    style=3D"FONT-SIZE: 10pt">&nbsp;<o:p></o:p></SPAN></FONT></P></DIV>
    <DIV>
    <P class=3DMsoNormal><FONT face=3DArial size=3D2><SPAN=20
    style=3D"FONT-SIZE: 10pt; FONT-FAMILY: =
Arial">Thanks.</SPAN></FONT><FONT=20
    size=3D2><SPAN style=3D"FONT-SIZE: =
10pt"><o:p></o:p></SPAN></FONT></P></DIV>
    <DIV>
    <P class=3DMsoNormal><FONT face=3D"Times New Roman" size=3D2><SPAN=20
    style=3D"FONT-SIZE: 10pt">&nbsp;<o:p></o:p></SPAN></FONT></P></DIV>
    <DIV>
    <P class=3DMsoNormal><FONT face=3DArial size=3D2><SPAN=20
    style=3D"FONT-SIZE: 10pt; FONT-FAMILY: =
Arial">Jenny</SPAN></FONT><FONT=20
    size=3D2><SPAN=20
    style=3D"FONT-SIZE: =
10pt"><o:p></o:p></SPAN></FONT></P></DIV></DIV></DIV></BLOCKQUOTE></BLOCK=
QUOTE><p></p><p> Gruppo Telecom Italia - Direzione e coordinamento di =
Telecom Italia =
S.p.A.<br><br>=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=
=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=
=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D<br>=
CONFIDENTIALITY NOTICE<br>This message and its attachments are addressed =
solely to the persons<br>above and may contain confidential information. =
If you have received<br>the message in error, be informed that any use =
of the content hereof<br>is prohibited. Please return it immediately to =
the sender and delete<br>the message. Should you have any questions, =
please send an e_mail to<br>MailAdmin@tilab.com. Thank =
you<br>=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=
=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=
=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D</BODY></H=
TML>

------_=_NextPart_001_01C59356.83E579C9--


--===============1038945621==
Content-Type: text/plain; charset="us-ascii"
MIME-Version: 1.0
Content-Transfer-Encoding: 7bit
Content-Disposition: inline

_______________________________________________
Speechsc mailing list
Speechsc@ietf.org
https://www1.ietf.org/mailman/listinfo/speechsc

--===============1038945621==--




From speechsc-bounces@ietf.org Fri Jul 29 10:31:03 2005
Received: from localhost.localdomain ([127.0.0.1] helo=megatron.ietf.org)
	by megatron.ietf.org with esmtp (Exim 4.32)
	id 1DyVt9-00038p-0X; Fri, 29 Jul 2005 10:31:03 -0400
Received: from odin.ietf.org ([132.151.1.176] helo=ietf.org)
	by megatron.ietf.org with esmtp (Exim 4.32) id 1DyVt6-00038a-3T
	for speechsc@megatron.ietf.org; Fri, 29 Jul 2005 10:31:01 -0400
Received: from ietf-mx.ietf.org (ietf-mx [132.151.6.1])
	by ietf.org (8.9.1a/8.9.1a) with ESMTP id KAA13089
	for <speechsc@ietf.org>; Fri, 29 Jul 2005 10:30:57 -0400 (EDT)
From: oran@cisco.com
Received: from sj-iport-2-in.cisco.com ([171.71.176.71]
	helo=sj-iport-2.cisco.com) by ietf-mx.ietf.org with esmtp (Exim 4.43)
	id 1DyWOj-0003Bp-HW
	for speechsc@ietf.org; Fri, 29 Jul 2005 11:03:44 -0400
Received: from sj-core-5.cisco.com (171.71.177.238)
	by sj-iport-2.cisco.com with ESMTP; 29 Jul 2005 07:30:47 -0700
Received: from imail.cisco.com (imail.cisco.com [128.107.200.91])
	by sj-core-5.cisco.com (8.12.10/8.12.6) with ESMTP id j6TEUlJL016976;
	Fri, 29 Jul 2005 07:30:47 -0700 (PDT)
Received: from [10.32.245.153] (stealth-10-32-245-153.cisco.com
	[10.32.245.153])
	by imail.cisco.com (8.12.11/8.12.10) with SMTP id j6TESR5b000681;
	Fri, 29 Jul 2005 07:28:27 -0700
In-Reply-To: <AD8171548C328E46BD8D0FEE40D9968872A1CC@xmb-sjc-224.amer.cisco.com>
References: <AD8171548C328E46BD8D0FEE40D9968872A1CC@xmb-sjc-224.amer.cisco.com>
Mime-Version: 1.0 (Apple Message framework v733)
Content-Type: text/plain; charset=US-ASCII; delsp=yes; format=flowed
Message-Id: <9E8FE0A0-882B-4BEB-8C4B-D36802DEEFF6@cisco.com>
Content-Transfer-Encoding: 7bit
Subject: Re: [Speechsc] propose adding "media-type" header field
	inRECOGNIZEorSET-PARAMS/GET-PARAMS methods
Date: Fri, 29 Jul 2005 00:23:08 -0400
To: Jenny Yao ((jyao)) <jyao@cisco.com>
X-Mailer: Apple Mail (2.733)
DKIM-Signature: a=rsa-sha1;  q=dns; l=2976; t=1122647308; x=1123079508;
	c=nowsp; s=nebraska;
	h=Subject:From:Sender:Date:Content-Type:Content-Transfer-Encoding;
	d=cisco.com; i=oran@cisco.com; 
	z=Subject:Re=3A=20[Speechsc]=20propose=20adding=20=22media-type=22=20header=20fiel
	d=20inRECOGNIZEorSET-PARAMS/GET-PARAMS=20methods|
	From:oran@cisco.com|
	Date:Fri,=2029=20Jul=202005=2000=3A23=3A08=20-0400|
	Content-Type:text/plain=3B=20charset=3DUS-ASCII=3B=20delsp=3Dyes=3B=20format=3Dflowed|
	Content-Transfer-Encoding:7bit;
	b=flnt7Kw8T2ucAivijtCIS7AZ17xxKb0dqevKACiDY4yncka9LCBPs21wQwjQGWmE5lQSsnnb
	myrj0YYdvBMAnBHaXDMPw+KY+xT9EyiY7p/BHVuzcmfUQIgWY+phq63nsErGUG6Bwl1bWshxOu9
	wxsE3byvD9VsnfvmewSpmEhg=
Authentication-Results: imail.cisco.com; header.From=oran@cisco.com;
	dkim=pass ( message from cisco.com verified; ); 
X-Spam-Score: 0.9 (/)
X-Scan-Signature: 287c806b254c6353fcb09ee0e53bbc5e
Content-Transfer-Encoding: 7bit
Cc: speechsc@ietf.org
X-BeenThere: speechsc@ietf.org
X-Mailman-Version: 2.1.5
Precedence: list
List-Id: Speech Services Control Working Group <speechsc.ietf.org>
List-Unsubscribe: <https://www1.ietf.org/mailman/listinfo/speechsc>,
	<mailto:speechsc-request@ietf.org?subject=unsubscribe>
List-Post: <mailto:speechsc@ietf.org>
List-Help: <mailto:speechsc-request@ietf.org?subject=help>
List-Subscribe: <https://www1.ietf.org/mailman/listinfo/speechsc>,
	<mailto:speechsc-request@ietf.org?subject=subscribe>
Sender: speechsc-bounces@ietf.org
Errors-To: speechsc-bounces@ietf.org


On Jul 27, 2005, at 6:53 PM, Jenny Yao ((jyao)) wrote:

> There appears to be no opposition regarding the media-type header  
> field request that we sent earlier. So, can we request this feature  
> be included in the next v2 draft?
>
Sarvi can confirm, but according to our internal author/chair issue  
tracker, this will appear in -07 (due out shortly).

>
> In addition, VXML 2.1 needs to have the information about  
> "duration" and "size" of the recorded audio in the field "waveform- 
> uri" in recog-only header. We would like to propose the following  
> fields be added to the recog-only header:
>
>         waveform-duration - the duration of the audio file in  
> milliseconds
>         waveform-size - the size of the audio file in bytes
>
The size is obtainable for content in the body through the content- 
length MIME header. Duration should be easily computable from the  
encoding type (e.g. coder) and the size.

If you mean input policy from the client to say the maximum size and  
duration of the returned data, that's another matter, and probably  
need discussion, since there can be complex interactions with other  
policy variables, like timeouts and success criteria.

Dave.

>
> Thanks.
>
> Jenny
>
> From: speechsc-bounces@ietf.org [mailto:speechsc-bounces@ietf.org]  
> On Behalf Of Saravanan Shanmugham (sarvi)
> Sent: Wednesday, May 04, 2005 9:27 AM
> To: Thomas Gal; Jenny Yao (jyao); speechsc@ietf.org
> Subject: RE: [Speechsc] propose adding "media-type" header field  
> inRECOGNIZEorSET-PARAMS/GET-PARAMS methods
>
> I don't see why not, considering all this behaviour would apply  
> only if the "Save-Waveform" header is set to true and failure to  
> caoture or save the waveform does not necessarily stop the  
> RECOGNIZE operation itself.
>
> Anyone opposed to doing this.
>
> Thanks,
> Sarvi
>
> From: speechsc-bounces@ietf.org [mailto:speechsc-bounces@ietf.org]  
> On Behalf Of Thomas Gal
> Sent: Friday, April 29, 2005 9:47 AM
> To: 'Jenny Yao (jyao)'; speechsc@ietf.org
> Subject: RE: [Speechsc] propose adding "media-type" header field in  
> RECOGNIZEorSET-PARAMS/GET-PARAMS methods
>
> Though this information could probably be inferred from the  
> filename extension, If we are going to follow/emulate the RECORD  
> methodology than it should also be available as a MIME body to the  
> RECOGNITION COMPLETE/STOP events as well. Otherwise I agree  
> completely.
>
>
>
> -Tom
>
>
>
> From: speechsc-bounces@ietf.org [mailto:speechsc-bounces@ietf.org]  
> On Behalf Of Jenny Yao (jyao)
> Sent: Friday, April 29, 2005 8:58 AM
> To: speechsc@ietf.org
> Subject: [Speechsc] propose adding "media-type" header field in  
> RECOGNIZE orSET-PARAMS/GET-PARAMS methods
>
>
>
>
>
>
>
> VoiceXML 2.1 specifies "recording user utterances while attempting  
> recognition". This can be done through recog-only-header "save- 
> waveform" and "waveform-uri" in MRCP V1 and V2. VoiceXML 2.1 also  
> uses recordutterancetype property to specify the media format of  
> the result recording. However, there is no way in MRCP to pass the  
> required media format to the server. We propose using the "Media- 
> Type" header field, currently defined for the recording resource,  
> in the RECOGNIZE method or the SET-PARAMS/GET-PARAMS methods to  
> specify a media type. The Save-Waveform carries a URI pointing to  
> the audio captured during recognition. The captured audio SHOULD be  
> saved with this media-type.
>
>
>
> Thanks.
>
>
>
> Jenny
>
> _______________________________________________
> Speechsc mailing list
> Speechsc@ietf.org
> https://www1.ietf.org/mailman/listinfo/speechsc
>

_______________________________________________
Speechsc mailing list
Speechsc@ietf.org
https://www1.ietf.org/mailman/listinfo/speechsc



From speechsc-bounces@ietf.org Fri Jul 29 11:21:10 2005
Received: from localhost.localdomain ([127.0.0.1] helo=megatron.ietf.org)
	by megatron.ietf.org with esmtp (Exim 4.32)
	id 1DyWfe-0001Kw-6l; Fri, 29 Jul 2005 11:21:10 -0400
Received: from odin.ietf.org ([132.151.1.176] helo=ietf.org)
	by megatron.ietf.org with esmtp (Exim 4.32) id 1DyWfc-0001Jk-O6
	for speechsc@megatron.ietf.org; Fri, 29 Jul 2005 11:21:08 -0400
Received: from ietf-mx.ietf.org (ietf-mx [132.151.6.1])
	by ietf.org (8.9.1a/8.9.1a) with ESMTP id LAA16727
	for <speechsc@ietf.org>; Fri, 29 Jul 2005 11:21:05 -0400 (EDT)
Received: from sj-iport-5.cisco.com ([171.68.10.87])
	by ietf-mx.ietf.org with esmtp (Exim 4.43) id 1DyXBJ-0004nK-AI
	for speechsc@ietf.org; Fri, 29 Jul 2005 11:53:53 -0400
Received: from sj-core-3.cisco.com (171.68.223.137)
	by sj-iport-5.cisco.com with ESMTP; 29 Jul 2005 08:21:00 -0700
X-IronPort-AV: i="3.95,153,1120460400"; 
	d="scan'208"; a="201453640:sNHT32637986"
Received: from vtg-um-e2k6.sj21ad.cisco.com (vtg-um-e2k6.cisco.com
	[171.70.93.77])
	by sj-core-3.cisco.com (8.12.10/8.12.6) with ESMTP id j6TFKu6p015797;
	Fri, 29 Jul 2005 08:20:57 -0700 (PDT)
Content-class: urn:content-classes:message
MIME-Version: 1.0
Content-Type: text/plain;
	charset="us-ascii"
Content-Transfer-Encoding: quoted-printable
X-MimeOLE: Produced By Microsoft Exchange V6.5.7226.0
Subject: RE: [Speechsc] propose adding "media-type" header
	fieldinRECOGNIZEorSET-PARAMS/GET-PARAMS methods
Date: Fri, 29 Jul 2005 08:20:55 -0700
Message-ID: <03772D1EC8DE624A863058C75874A75C1E0BA9@vtg-um-e2k6.sj21ad.cisco.com>
Thread-Topic: [Speechsc] propose adding "media-type" header
	fieldinRECOGNIZEorSET-PARAMS/GET-PARAMS methods
Thread-Index: AcWUSq55vxQ0QdnoSYe4AtOFkmSMcgABTj4Q
From: "Shanmugham, Saravanan" <sarvi@cisco.com>
To: <oran@cisco.com>, "Jenny Yao \(\(jyao\)\)" <jyao@cisco.com>
X-Spam-Score: 0.0 (/)
X-Scan-Signature: 20f22c03b5c66958bff5ef54fcda6e48
Content-Transfer-Encoding: quoted-printable
Cc: speechsc@ietf.org
X-BeenThere: speechsc@ietf.org
X-Mailman-Version: 2.1.5
Precedence: list
List-Id: Speech Services Control Working Group <speechsc.ietf.org>
List-Unsubscribe: <https://www1.ietf.org/mailman/listinfo/speechsc>,
	<mailto:speechsc-request@ietf.org?subject=unsubscribe>
List-Post: <mailto:speechsc@ietf.org>
List-Help: <mailto:speechsc-request@ietf.org?subject=help>
List-Subscribe: <https://www1.ietf.org/mailman/listinfo/speechsc>,
	<mailto:speechsc-request@ietf.org?subject=subscribe>
Sender: speechsc-bounces@ietf.org
Errors-To: speechsc-bounces@ietf.org

=20

     -----Original Message-----
     From: speechsc-bounces@ietf.org=20
     [mailto:speechsc-bounces@ietf.org] On Behalf Of oran@cisco.com
     Sent: Thursday, July 28, 2005 9:23 PM
     To: Jenny Yao ((jyao))
     Cc: speechsc@ietf.org
     Subject: Re: [Speechsc] propose adding "media-type" header=20
     fieldinRECOGNIZEorSET-PARAMS/GET-PARAMS methods
    =20
    =20
     On Jul 27, 2005, at 6:53 PM, Jenny Yao ((jyao)) wrote:
    =20
     > There appears to be no opposition regarding the=20
     media-type header=20
     > field request that we sent earlier. So, can we request=20
     this feature be=20
     > included in the next v2 draft?
     >
     Sarvi can confirm, but according to our internal=20
     author/chair issue tracker, this will appear in -07 (due=20
     out shortly).

Yes this covered.    =20
     >
     > In addition, VXML 2.1 needs to have the information about =20
     > "duration" and "size" of the recorded audio in the field=20
     "waveform-=20
     > uri" in recog-only header. We would like to propose the=20
     following =20
     > fields be added to the recog-only header:
     >
     >         waveform-duration - the duration of the audio file in =20
     > milliseconds
     >         waveform-size - the size of the audio file in bytes
     >
     The size is obtainable for content in the body through the=20
     content-=20
     length MIME header. Duration should be easily computable from the =20
     encoding type (e.g. coder) and the size.
    =20
     If you mean input policy from the client to say the=20
     maximum size and =20
     duration of the returned data, that's another matter, and=20
     probably =20
     need discussion, since there can be complex interactions=20
     with other =20
     policy variables, like timeouts and success criteria.
I don't believe this is policy piece sent from the client, but rather
information from the server. After a recognize the recorded
audio/waveform is available as a URI. The request, I believe was for the
server to send these 2 pieces of information in addition to the URI. We
could address this by allowing these fields to be sent in the same
header like the following.

Waveform-URI:<http://mediaserver.acme.com/recordin1234.wav>;size=3D743672=
;
duration=3D2300

Thx,
Sarvi    =20
     Dave.
    =20
     >
     > Thanks.
     >
     > Jenny
     >
     > From: speechsc-bounces@ietf.org=20
     [mailto:speechsc-bounces@ietf.org] =20
     > On Behalf Of Saravanan Shanmugham (sarvi)
     > Sent: Wednesday, May 04, 2005 9:27 AM
     > To: Thomas Gal; Jenny Yao (jyao); speechsc@ietf.org
     > Subject: RE: [Speechsc] propose adding "media-type"=20
     header field =20
     > inRECOGNIZEorSET-PARAMS/GET-PARAMS methods
     >
     > I don't see why not, considering all this behaviour would apply =20
     > only if the "Save-Waveform" header is set to true and=20
     failure to =20
     > caoture or save the waveform does not necessarily stop the =20
     > RECOGNIZE operation itself.
     >
     > Anyone opposed to doing this.
     >
     > Thanks,
     > Sarvi
     >
     > From: speechsc-bounces@ietf.org=20
     [mailto:speechsc-bounces@ietf.org] =20
     > On Behalf Of Thomas Gal
     > Sent: Friday, April 29, 2005 9:47 AM
     > To: 'Jenny Yao (jyao)'; speechsc@ietf.org
     > Subject: RE: [Speechsc] propose adding "media-type"=20
     header field in =20
     > RECOGNIZEorSET-PARAMS/GET-PARAMS methods
     >
     > Though this information could probably be inferred from the =20
     > filename extension, If we are going to follow/emulate=20
     the RECORD =20
     > methodology than it should also be available as a MIME=20
     body to the =20
     > RECOGNITION COMPLETE/STOP events as well. Otherwise I agree =20
     > completely.
     >
     >
     >
     > -Tom
     >
     >
     >
     > From: speechsc-bounces@ietf.org=20
     [mailto:speechsc-bounces@ietf.org] =20
     > On Behalf Of Jenny Yao (jyao)
     > Sent: Friday, April 29, 2005 8:58 AM
     > To: speechsc@ietf.org
     > Subject: [Speechsc] propose adding "media-type" header field in =20
     > RECOGNIZE orSET-PARAMS/GET-PARAMS methods
     >
     >
     >
     >
     >
     >
     >
     > VoiceXML 2.1 specifies "recording user utterances while=20
     attempting =20
     > recognition". This can be done through recog-only-header "save-=20
     > waveform" and "waveform-uri" in MRCP V1 and V2. VoiceXML=20
     2.1 also =20
     > uses recordutterancetype property to specify the media=20
     format of =20
     > the result recording. However, there is no way in MRCP=20
     to pass the =20
     > required media format to the server. We propose using=20
     the "Media-=20
     > Type" header field, currently defined for the recording=20
     resource, =20
     > in the RECOGNIZE method or the SET-PARAMS/GET-PARAMS methods to =20
     > specify a media type. The Save-Waveform carries a URI=20
     pointing to =20
     > the audio captured during recognition. The captured=20
     audio SHOULD be =20
     > saved with this media-type.
     >
     >
     >
     > Thanks.
     >
     >
     >
     > Jenny
     >
     > _______________________________________________
     > Speechsc mailing list
     > Speechsc@ietf.org
     > https://www1.ietf.org/mailman/listinfo/speechsc
     >
    =20
     _______________________________________________
     Speechsc mailing list
     Speechsc@ietf.org
     https://www1.ietf.org/mailman/listinfo/speechsc
    =20

_______________________________________________
Speechsc mailing list
Speechsc@ietf.org
https://www1.ietf.org/mailman/listinfo/speechsc



