TWIML AI
From Voice Agents to AI Avatars with Alexander Smola - #777
00:003910
00:00
This episode is brought to you by Blitzie, the Autonomous Software Development Platform
本期节目由 Blitzie 赞助,Autonomous Software Development Platform
00:05
built for enterprise scale.
专为 enterprise scale 打造。
00:07
Today's code bases have grown beyond human comprehension.
如今的 code base 已经大得超出人类的理解范围。
00:10
Millions of lines, decades of tech debt and complexity existing tools just can't fathom.
数百万行代码、几十年的 tech debt,还有现有工具根本摸不透的 complexity。
00:16
With Blitzie, thousands of specialized agents reverse engineer the code base, mapping
有了 Blitzie,成千上万个 specialized agents 会 reverse engineer 整个 code base,梳理
00:21
the architecture, dependencies, and business logic.
architecture、dependencies 和 business logic。
00:24
With that context, the platform then autonomously executes entire epochs, writing, validating
有了这些 context,平台就会自主执行整个 epochs,编写、validating
00:29
and testing the code for every project.
并 testing 每个项目的代码。
00:32
The result, Fortune 500 enterprises are able to modernize legacy systems and ship new features
结果是,Fortune 500 企业能够对 legacy systems 进行现代化改造,并发布新功能
00:38
five times faster.
速度能快五倍。
00:40
Want to try Blitzie on your code today?
想不想今天就在你的代码上试试 Blitzie?
00:42
Unlock 1 million lines of reverse engineering and 25k lines of code generation by visiting
通过访问以下链接,解锁 100 万行 reverse engineering 和 25k 行 code generation:
00:47
blitzie.com slash sandbox.
blitzie.com slash sandbox。
00:53
Voice AI has gotten very good over the past few years, but it still suffers from a bit
过去几年,Voice AI 已经变得非常出色,但它仍然有点
00:58
of an uncanny valley problem.
uncanny valley 问题。
01:01
These are remarkably sensitive to the subtle things that make a conversation feel natural
它们对让一段对话显得自然的细微之处异常敏感。
01:05
or not.
或者不是。
01:06
A little too much latency and interruption handled awkwardly, the wrong tone or emotional
Latency 稍微高一点,interruption 处理得生硬一点,语气或情绪回应不对,幻觉瞬间就破灭了。
01:10
response, and suddenly the illusion breaks.
而且随着 AI 系统不断进化,把我们的更多感官纳入进来,这个门槛只会越来越高——从文本和语音,走向能看见、也能被看见的系统,并在自然交互方面带来一整套全新的挑战。
01:14
And the bar only gets higher as AI systems evolve to incorporate more of our senses, moving
今天的嘉宾是 Alex Mola,Boson AI 的联合创始人兼 CEO,也是 Carnegie Mellon University 的教授。
01:19
beyond texts and voice to systems that can see and be seen and produces a whole new set
超越文本和语音,走向能看也能被看见的系统,这会带来一整套全新的
01:24
of challenges around natural interaction.
围绕 natural interaction 的挑战。
01:27
My guest today is Alex Mola, co-founder and CEO of Boson AI and a professor at Carnegie
今天我的嘉宾是 Alex Mola,Boson AI 的联合创始人兼 CEO,也是 Carnegie
01:33
Mellon University.
Mellon University 的教授。
01:34
Our conversation explores what it takes to build these systems, from the models and inference
我们的对话探讨了构建这些系统需要什么,从 models 和 inference
01:39
infrastructure needed to make voice work in real time, to emotional intelligence and ultimately
基础设施——让语音实时工作所需的——到 emotional intelligence,最终到
01:45
audiovisual avatars.
audiovisual avatars。
01:47
Here's Alex on where he sees all of this heading.
下面是 Alex 谈他如何看待这一切的走向。
01:50
I would argue that voice is an intermediate stepping stone.
我会说,语音是一个中间的过渡跳板。
01:54
We might think, well, what's next, it's clear that eventually we will have avatars.
我们可能会想,好吧,下一步是什么,很明显最终我们会有 avatars。
02:01
And so I think what this is converging to is that you'll be talking to an av agent that
所以我认为这一切正在收敛到:你会跟一个 AV agent 对话,它
02:07
looks and feels like a human and we're still very, very far away from making this really
看起来、感觉起来都像人类,而我们离真正实现这件事还非常非常远。
02:13
natural.
很自然。
02:15
So it's quite exciting, actually.
所以其实还挺让人兴奋的。
02:17
I'm Sam Charington and this is the Twomo AI podcast.
我是 Sam Charington,这里是 Twomo AI podcast。
02:21
For over a decade, I've been exploring the ideas and innovations shaping the future of
十多年来,我一直在探索那些塑造 AI 未来的想法和创新,
02:24
AI through conversations like this one that help you understand what's real, what's next
通过像这样的对话,帮你看清什么是真实的、下一步是什么,
02:30
and what matters.
以及什么才真正重要。
02:32
Let's jump in.
我们开始吧。
02:41
When I think about kind of the frontier of voice AI that most folks probably have access
当我想到大多数人可能都能接触到的那种 voice AI 前沿时
02:50
to, I'm imagining it's something like chat GPT advanced voice mode.
那,我猜它有点像 ChatGPT advanced voice mode。
02:57
Actually, I want you to react to this.
其实,我想让你对此回应一下。
02:59
Do you think that actually that's crap and there are much better systems and you should
你觉得那其实很烂,有更好的系统,而你应该
03:03
point me to X, Y and Z or are there better systems behind closed doors and labs and enterprises
给我指出 X、Y 和 Z,还是说在紧闭的门后、实验室和企业里有更好的系统
03:10
or what?
还是怎样?
03:11
But in general, I tend to think that they've gone through several iterations of it.
但总的来说,我倾向于认为他们已经对它迭代了好几轮。
03:16
It does keep getting better, but it's still very infuriating.
它确实一直在变好,但还是非常让人抓狂。
03:20
In my experience, it really only works in perfect conditions, meaning a silent room,
以我的经验,它真的只在完美条件下才管用,也就是在一个安静的房间里,
03:25
no background noise, being very conscientious about interrupting.
没有 background noise,对打断别人这件事非常谨慎。
03:29
If you interrupt either, you have to be committed to plowing forward or you have to stop and
如果你打断其中任何一个,你要么得下定决心一直往前推,要么就得停下来,
03:34
let the thing catch up to you.
让那个东西追上你。
03:36
You can't try to have a natural engagement.
你没法试着自然地互动。
03:39
It's very brittle, I think.
我觉得这非常脆弱。
03:42
You can definitely do better than that situation.
你肯定能做得比那种情况更好。
03:49
One situation from last week, so I was at Jeju in South Korea for the KDD conference
上周有个情况,我当时在 South Korea 的 Jeju 参加 KDD conference
03:57
and we were out at dinner for drinks in the bar and basically I was showing off our
然后我们出去吃晚饭,在酒吧里喝酒,基本上我是在炫耀我们的
04:06
system and just me foolishly say, hey, watch this, let's see what happens.
system 和我只是傻乎乎地说,嘿,看这个,看看会发生什么。
04:16
What I can confirm is that our model was able to handle people switching to Slovenia and
我能确认的是,我们的 model 能够处理人们切换到 Slovenia,然后
04:25
then somebody else to Hindi fairly well, even though the bar was very noisy.
另一个人切换到 Hindi 也处理得相当好,尽管酒吧里非常吵。
04:30
I would say a non-trivial amount of the credit definitely goes to the iOS team doing really
我会说,相当大一部分功劳肯定要归给 iOS team,他们在手机上做了非常
04:37
good noise cancellation and audio separation on the mobile phone, so I don't think that
好的 noise cancellation 和 audio separation,所以我不认为
04:44
this would have happened just by waving a microphone somewhere, but I think we are on
这事儿只是随便在哪儿挥个 microphone 就能发生,但我认为我们正在
04:53
our way and so basically having a single microphone will probably never really solve this.
走上正轨,所以基本上只有一个 microphone 很可能永远没法真正解决这个问题。
05:02
You need microphone arrays to really do proper noise cancellation, especially if you
你需要 microphone arrays 才能真正做好 noise cancellation,尤其是如果你
05:06
have many sources, but that's a solve problem, right?
有很多来源,但那是个可解决的问题,对吧?
05:09
I mean those things ship.
我是说,这些东西都已经出货了。
05:10
I mean, well, I've got a huge microphone array sitting in my laptop here.
我是说,嗯,我这台笔记本里就装着一个巨大的 microphone array。
05:15
I don't know if Codex was using it, but I've also had experiences with the relatively
我不知道 Codex 是不是在用它,但我也体验过相对
05:22
new Codex voice, it just that interaction just wasn't very fluid.
新的 Codex voice,只是那个交互就是不太流畅。
05:28
Yeah, okay, so that's, I mean, sometimes voice input is good if you're feeling too lazy
对,好吧,所以那,我是说,有时候 voice input 挺好,如果你懒得
05:37
to type.
打字。
05:38
I mean, typing way more precise for technical work, but for, so as in, you know, hey, let's
我是说,做技术工作的话,打字要精确得多,但为了,所以就是说,你知道,嘿,咱们
05:46
configure our storage array, well, no, I don't want to talk to you, I want to type because
配置我们的 storage array,呃,不,我不想跟你说话,我想打字,因为
05:51
here's a very specific serial number and a specific URL and I want to get this exactly right.
这里有一个非常具体的 serial number 和一个具体的 URL,我想把这件事弄得完全准确。
05:58
So that's where I still prefer text, but maybe this is old school, but for, you know, if you're
所以这就是为什么我还是更喜欢 text,但也许这有点老派,不过,你知道,如果你
06:05
out in a bar and you're asking for information advice or something, I think we're getting
在酒吧里,你在问信息、建议或者别的什么,我觉得我们正在
06:12
there.
接近了。
06:13
I will give it another year and this will actually become quite ubiquitous and I'm pretty
我会再给它一年,这真的会变得相当普遍,而且我挺
06:19
happy where we are at with our models.
满意我们的 models 现在所处的状态。
06:22
I think probably in a year, audio will become pretty much bulletproof, video feeds as in
我觉得可能再过一年,audio 会变得几乎无懈可击,video feeds 比如
06:31
avatars and so on are going to start making some appearances, probably a year and a half,
avatars 之类的会开始露面,大概一年半以后吧,
06:40
we'll see that quite widely deployed on robots, basically with a face fully animated.
我们会看到它相当广泛地部署在 robots 上,基本上就是一张完全动起来的脸。
06:47
So we're working actually with the startup on some of those things problems related to
所以其实我们正在和那家 startup 一起做一些事情,处理跟
06:53
that.
这个相关的问题。
06:54
It's a very talented team.
那是个非常有才华的团队。
06:57
So and there, I think the lead times are a little bit longer because if you want to do this
所以在那方面,我觉得 lead times 会稍微长一点,因为如果你想做这件事
07:02
in the, you know, and I made a tattoo or whatever, you actually need to make hardware and
在那个,你知道,我弄了个纹身还是什么的,你实际上得做 hardware,而且
07:09
hardware is hard, right, software is easy comparatively.
hardware 很难,对吧,相比之下 software 就简单多了。
07:16
When you think about the challenge of voice AI, how do you break it up in terms of, you
当你思考 voice AI 的挑战时,你会怎么拆解它——你知道,主要是 engineering 问题,还是主要是 research 问题?
07:27
know, primarily engineering problems, primarily research problems?
如果你试着那样做,结果不会好,你得同时做这两件事。
07:32
If you try that, it's not going to go well, you need to do both at the same time.
那让我给你举个简单的例子。
07:37
So let me give you a simple example.
所以咱们从,你知道,最基础的东西说起,也就是,从基本的作用打到你的 retina,到你的 cortex 真正对它做出处理,大约需要 150 毫秒。
07:41
So let's start with, you know, the very simple basics, namely it takes about 150 milliseconds
而且这个数字相当稳定,你其实可以把它当作一种 non-invasive
07:48
for, you know, basic effort on hitting your retina to your cortex actually doing something
因为,你知道,最基本的努力就是让打到你 retina 上的东西传到你 cortex,并真正做点什么
07:55
with it.
用它。
07:56
And that number is reasonably stable that you can actually use that as a non-invasive
而且那个数字相当稳定,你其实可以把它当作一种 non-invasive
08:00
diagnostic to find out whether you have a neurodegenerative disease, because in that
用来查你是否患有 neurodegenerative disease 的 diagnostic,因为在那种
08:05
case, the signal takes many detours and it takes longer.
情况下,signal 会绕很多弯路,花的时间也更长。
08:09
What people will do is they will flash a checkerboard pattern in front of you and then wait
人们会做的是,在你面前闪一个 checkerboard pattern,然后等
08:14
how long it takes for that stimulus to hit your cortex.
看那个 stimulus 要多久才能到达你的 cortex。
08:18
From the ears, it's actually a little bit shorter because it's a, well, half the way, right?
从耳朵这边走,其实会短一点,因为它,呃,只走了一半的路,对吧?
08:24
Now what that means is that basically humans kind of operate that around maybe six to ten
那这意味着,基本上人类大概是在六到十
08:31
hearts really when it comes to audiovisual perception.
hertz 左右运作,真的,就 audiovisual perception 来说。
08:35
I mean, of course we feel like the road is very fluid and I mean, we have our fancy 30 or 60
我的意思是,当然我们会觉得路非常流畅,我的意思是,我们有那些很炫的 30 或 60
08:41
hearts displays, but the brain kind of, you know, within that time frame, it's kind of okay.
hearts displays,但大脑呢,你知道,在那个时间范围内,其实还行。
08:48
So what we did correspondingly is make sure that our model is interoperable within about
所以我们相应地做的,就是确保我们的模型可互操作,大概在
08:55
a similar time frame.
一个相似的时间范围内。
08:57
So it does interruption handling within about 150 really seconds or so.
所以它处理 interruption handling,大概真的就在 150 秒左右。
09:04
Now the next thing is you need at least the current paradigm of how to deal with audio
接下来呢,你至少需要的是,当前处理 audio 的范式
09:12
and other things is at some point you convert everything into tokens.
以及其他东西的范式,就是到了某个时候,你把所有东西都转成 tokens。
09:17
And then those tokens will go to whatever LLM style backbone that you have, right?
然后这些 tokens 会进入你拥有的任何 LLM 风格 backbone,对吧?
09:25
I mean, the architectures may differ and maybe at some point we'll get to diffusion based models.
我是说,architectures 可能会不一样,也许到某个时候我们会走到 diffusion based models。
09:30
I would think it's probably going to be more like speculative decoding with diffusion and the more,
我会觉得,这大概率会更像带 diffusion 的 speculative decoding,以及更多,
09:39
you know, traditional sequential models for, you know, just overall the statistical modeling.
你知道,传统的 sequential models,用来,你知道,做整体的 statistical modeling。
09:45
But you know, that's a completely different argument and definitely one may have very
但你知道,那完全是个不同的论点,而且肯定有人对此可能会有非常
09:50
different opinions on that.
不同的看法。
09:54
But basically, you know, you need to turn the audio into tokens.
但基本上,你知道,你需要把 audio 转成 tokens。
09:59
Then your model does something and then you need to turn those tokens back into audio.
然后你的 model 做点什么,接着你需要把那些 tokens 再转回 audio。
10:05
Now you can get very high fidelity by having a very high token frequency.
现在你可以通过拥有非常高的 token frequency 来获得 very high fidelity。
10:12
The problem is a high token frequency means that you need to ingest,
问题是,高 token frequency 意味着你需要 ingest,
10:18
that's pre-fill and also generate many, many tokens per second.
那是在做 pre-fill,同时每秒生成很多很多 tokens。
10:22
And we all know that this is expensive.
而且我们都知道,这很贵。
10:24
So now you have this rather unpleasant dilemma where many tokens per second means
所以现在你就面临一个挺难受的两难境地:每秒很多 tokens 意味着
10:32
your model that cannot have too many parameters.
你的 model 就不能有太多 parameters。
10:34
Whereas if you have a smaller number of tokens per second, you can afford more parameters, right?
而如果每秒的 tokens 数量少一些,你就能负担更多 parameters,对吧?
10:39
Text is, you know, the ultimate compressed format in that sense, right?
文本,你知道,从这个意义上说就是终极的压缩格式,对吧?
10:43
It's about, you know, three to five tokens per second that humans want.
人类想要的,你知道,差不多就是每秒三到五个 tokens。
10:47
Press for audio, you can easily have, you know, 10 plus.
至于音频,你很容易就能有,你知道,10 以上。
10:52
And tokens per second also then means, you know, kind of, you know, the temporal granularity.
而 tokens per second 也就意味着,你知道,算是,你知道,temporal granularity。
10:58
Let's say I have 10 tokens per second.
假设我有 10 tokens per second。
11:00
Then that means that each token covers about 100 milliseconds.
那就意味着每个 token 覆盖大约 100 毫秒。
11:04
So this is why you actually get biology, user experience, engineering,
所以这就是为什么你实际上会牵涉到生物学、user experience、工程学,
11:10
because you need to make this cost effective.
因为你需要让这件事有成本效益。
11:12
And then science, namely, how do we actually represent this?
然后是科学,也就是,我们到底要怎么表示这个?
11:15
How do we send it into a model maybe doing something else more cleverly in the entire pipeline?
我们怎么把它送进一个 model,也许在整个 pipeline 里更聪明地做点别的事情?
11:22
So this is why all those things really need to come together.
所以这就是为什么所有这些事情真的需要结合在一起。
11:26
It may not necessarily be, you know, one engineer doing everything.
这不一定非得是,你知道,一个工程师把所有事都做了。
11:31
That would be pretty amazing if you could find somebody like that.
要是你能找到那样的人,那可就太厉害了。
11:35
I mean, there are very few people, but that's okay.
我是说,这样的人非常少,但没关系。
11:38
But, you know, it's a whole systems challenge.
但是,你知道,这是一个整个系统层面的挑战。
11:42
And then, of course, once you have all of that running, you need to also take care of
然后,当然,当你把这一切都跑起来之后,你还得处理
11:46
the engineering implementations, you probably need to buffer a little bit.
engineering implementations,你可能需要稍微 buffer 一下。
11:51
So remember when I mentioned you want to be interruptable, but you probably want to have
所以记得,我之前提到过,你希望自己是 interruptable 的,但你可能想要有
11:56
longer buffers.
更长的 buffers。
11:58
So now we are talking about essentially, you know, real-time type AV streaming.
所以现在,我们本质上聊的是,你知道,实时类型的 AV streaming。
12:03
Right?
对吧?
12:04
You also have a video feed.
你还有一路 video feed。
12:06
The video feed may come in at a different frame rate.
这个 video feed 可能会以不同的 frame rate 进来。
12:10
So yeah, basically, it's a really nice range of problems there.
所以对,基本上,这里面有一系列很不错的问题。
12:17
And for the video feed, I mean, we all know if you use WAN or some other models of flux,
而说到 video feed,我的意思是,我们都知道如果你用 WAN 或者 flux 的其他一些模型,
12:24
they will happily produce, you know, video segments of 10-ish seconds.
它们会很乐意生成,你知道,10 秒左右的 video segments。
12:29
But if you want to have a continuous feed that, you know, is visually consistent for an
但如果你想要一个 continuous feed,你知道,它在视觉上能保持一致,持续一段……
12:35
hour, you need to modify those models a little bit.
一小时,你需要稍微修改一下那些 models。
12:40
And if you then want to make those models effective such that you can have, you know, multiple
然后,如果你之后想让那些 models 有效,让你能有,你知道,多个
12:45
real-time factors, there is yet another design optimization to be made.
real-time factors,那就还有另一个 design optimization 要做。
12:50
The TLD are is, if you are doing video feeds for avatars, you know that it's a video feed
TLD are 就是,如果你是在给 avatars 做 video feeds,你知道它就是一条 video feed
12:58
for an avatar.
给一个 avatar 的。
12:59
So they are not going to be race cars driving in the background.
所以它们不会是在背景里飞驰的赛车。
13:04
In other words, most of the video feed is pre-boring.
换句话说,大部分 video feed 都是 pre-boring 的。
13:07
So, for instance, you could easily compress the av for this interview into a, well, fairly
所以,比如说,你可以很容易地把这次采访的 av compress 成一个,嗯,相当
13:15
effective stream.
有效的 stream。
13:17
Of course, if I start moving my hands like crazy, then that frame rate, then, you know,
当然,如果我开始疯狂地挥动双手,那这个 frame rate,然后,你知道,
13:21
the bit rate will immediately go up by a lot.
bit rate 就会立刻大幅上升。
13:24
But most humans don't do weird things like this.
但大多数人不会做这种奇怪的事。
13:28
And so it's a perfectly acceptable feed quality.
所以这完全是可接受的 feed quality。
13:32
And again, there is a trade-off between most beautiful quality and building something that
而且再说,要在最漂亮的质量和做出某种东西之间做 trade-off,而那个东西
13:38
actually people can afford.
实际上是人们能负担得起的。
13:39
And that's probably also the other point where we are maybe taking a slightly different
而这可能也是另一个点,我们或许在那里会采取稍微不同的
13:44
operating point from some of the, well, very large trophy models that are, you know, truly
operating point 来自一些,嗯,非常大的 trophy models,这些模型你知道,真的
13:51
and parameters just to do chat chat.
以及 parameters,只是为了聊聊天。
13:55
And which point in particular, the affordability or something about the trade-off or which?
那具体是哪个点,可负担性,还是关于 trade-off 的什么,还是哪个?
14:01
So the affordability, right?
所以是可负担性,对吧?
14:03
So basically, you can always make your models smarter by making it bigger.
所以基本上,你总是可以通过把 models 做大来让它们更聪明。
14:09
Absent of the real-time criteria that you mentioned with regards to the voice.
如果没有你提到的、关于语音的 real-time 标准。
14:12
There are some laws of physics that come in the play here.
这里有一些物理定律在起作用。
14:15
Yeah.
对。
14:16
And so the problem is basically, you know, how much compute do you need to stream, you
所以问题基本上就是,你知道,你需要多少 compute 才能把一段对话 stream 出去?
14:24
know, a conversation?
而如果你需要用,比如说,你知道,一整个 Blackwell server GPU 就为了单独一段对话,那这可能就不是最经济可行的模式了。
14:26
And if you need to use, let's say, you know, a full black well server GPU just for a single
我是说,这能做出非常惊艳的 demo,但你的客户负担不起。
14:33
conversation, then that may not be the most economically viable model.
而我觉得,这就是我们一开始先从价格入手,然后倒推,你知道,怎么才能做出人们真正负担得起的东西?
14:41
I mean, this produces gorgeous demos, but your customers can't afford it.
我们能不能退一步,让你稍微聊聊你这一路的发展轨迹,或者
14:49
And that's, I think, where we went in with the price first and then worked backwards
我觉得,这就是我们先从价格入手,然后再倒推的地方。
14:55
to, you know, how can you build something that actually people kind of afford?
到,你知道,你怎么才能做出人们实际上能勉强负担得起的东西?
14:59
Can we take a step back and maybe have you talk a little bit about your kind of arc or
我们能不能先退一步,也许让你稍微聊聊你的那种发展轨迹,或者
15:05
trajectory or path like you started the company in 2023?
发展轨迹或路径,比如你是在 2023 年创立公司的?
15:10
You were not initially focused on voice.
你一开始并没有专注于 voice。
15:12
You eventually shifted direction to voice when you started really focusing on voice kind
你最终转向了 voice,是在你开始真正专注于 voice 的时候。
15:20
of where did you start and what were the steps you took to kind of evolve to where you
大概你从哪里开始,以及你采取了哪些步骤,逐渐发展到
15:25
were today?
你今天所在的位置?
15:26
So we started off with text just like, I guess, others as well, maybe with a slightly
所以我们一开始是从 text 起步,就像我猜其他人也一样,可能稍微
15:32
stronger focus on AI for humans.
更侧重 AI for humans。
15:35
I mean, that's been with us since day one.
我的意思是,这一点从第一天起就一直伴随着我们。
15:39
And as mentioned, the one of the key issues was that, you know, the text interface felt
而且就像刚才提到的,其中一个关键问题是,你知道,text interface 总感觉
15:47
always a little bit awkward.
总是有点别扭。
15:50
And so we then started looking at, okay, what are good audio models.
于是我们接着开始看,好吧,什么样的 audio models 才算好。
15:54
We also realized that probably building your own LLM was not the smartest idea if you
我们还意识到,如果你
16:02
just wanted to have really good audio, but instead can you use the intelligence that's
只是想要特别好的 audio,那可能自己搭一个 LLM 并不是最聪明的做法,而是能不能利用那些
16:07
already baked into a high quality LLM and then make sure that it acquires effectively
已经内置在一个高质量 LLM 里的智能,然后确保它能有效地
16:15
one extra modality, namely in this case audio.
多掌握一种 modality,也就是在这个场景里的 audio。
16:19
Again, other people had done similar things, for instance, for video.
再者,其他人也做过类似的事情,比如在 video 方面。
16:24
So for instance, there are vision at least, so there are VLLM, so vision large language
所以比如说,至少有vision,所以有VLLM,所以vision large language
16:29
models.
models.
16:30
People are doing similar things for robotics and world models.
人们正在为robotics和world models做类似的事情。
16:34
So the overall pattern of using the intelligence that you kind of get for free from reasoning
所以整体模式是利用你从reasoning中免费获得的那种intelligence,
16:42
over large amounts of text, you then combine that with audio.
在大量text上,然后你把它和audio结合起来。
16:47
Now, one of the problems is if you, you know, teach this model on your modality and you
现在,其中一个问题是,如果你,你知道,在你的modality上教这个model,然后你
16:51
know, careful, it forgets everything that it knew before.
知道,小心,它会忘记之前知道的一切。
16:55
Probably the easiest way to imagine that is if you, and I've actually seen that with
可能最简单的想象方式就是,如果你,我实际上见过这种情况,
17:00
a friend, so she adopted a kid from Latin America.
有个朋友,她收养了一个来自 Latin America 的孩子。
17:07
So this was in Germany and she spoke only German to her and within a matter of months.
所以这件事发生在 Germany,她只对她说 German,结果没过几个月。
17:12
The girl had forgotten every single word of Spanish, right?
那个女孩已经把 Spanish 的每一个词都忘光了,对吧?
17:18
And the similar thing effectively happens if you take an LLM and you only get it to work
而类似的事情实际上也会发生:如果你拿一个 LLM,只让它去处理
17:24
with audio, then it will very quickly forget about all the reasoning and language itself.
audio,那它很快就会把所有的 reasoning 和语言本身都忘掉。
17:29
So you need to still maintain somewhat competent mid and post-training LLM pipeline while also
所以你需要仍然维持一个还算像样的 mid and post-training LLM pipeline,同时还要
17:38
having the same capabilities for audio.
在 audio 上具备同样的能力。
17:42
And you need to then also define tasks that nicely marry audio and text, so to ground things
然后你还需要定义一些能把 audio 和 text 很好地结合起来的任务,从而 ground 这些东西。
17:49
into each other.
互相转换。
17:50
I mean, just like if you have a multilingual LLM, at some point you need to make sure that,
我是说,就像你有一个多语言 LLM 的话,到了某个时候你就得确保,
17:57
you know, dog means shear or carne or wound.
你知道,dog 的意思是 shear,或者 carne,或者 wound。
18:02
And if you don't have that, then it becomes a little bit tricky for the model to reason
而如果你没有这个,那模型推理起来就会有点棘手
18:07
across languages and to get that strong generalization.
跨语言地,并获得那种很强的 generalization。
18:12
But anyway, so if we did this, we then, you know, released, I think, a fairly competent
但不管怎样,如果我们这么做了,那我们,你知道,发布了,我觉得,一个相当不错的
18:21
TTS model that was Hicks Audio V2 last year.
TTS 模型,就是去年的 Hicks Audio V2。
18:26
And I think we've significantly accelerated our release pipeline this year, so we've been
而且我觉得今年我们大幅加速了我们的 release pipeline,所以我们一直
18:34
putting out TTS and audio understanding in ASR models.
在 ASR models 里推出 TTS 和 audio understanding。
18:39
And so in case you wonder, what's the difference between, so TTS means text to speech.
所以,如果你好奇的话,这之间有什么区别,TTS 就是 text to speech。
18:43
So basically, you know, text in sound out, voice cloning, all of that.
基本上就是,你知道,text in、sound out,voice cloning,所有这些。
18:48
But what's the difference between audio understanding and ASR, so speech recognition, speech recognition
但 audio understanding 和 ASR 有什么区别呢,也就是 speech recognition,speech recognition
18:55
needs audio in text out.
需要 audio in、text out。
18:58
But the model doesn't really understand very much of what's going on in it, understands
但 model 其实并不太理解里面发生了什么,它能理解
19:02
a little bit, but these are fairly lightweight models, but they will not be able to resolve
一点点,但这些是相当 lightweight 的 models,不过它们没法分辨
19:10
whether to recognize speech means to recognize speech or whether it means to wreck a nice
到底 recognize speech 是指 recognize speech,还是说它是指 wreck a nice
19:19
speech, right?
speech,对吧?
19:20
They both sound the same.
它们俩听起来一模一样。
19:23
And obviously, one is nonsense, unless you are in some environmental conservation event
而且很明显,其中一个是没意义的,除非你是在某个环保活动上
19:32
where they probably mean the latter.
那种场合他们大概率指的是后者。
19:36
But you don't know, you know, the ASR isn't going to be able to handle that, but an audio
但你也说不准,你知道吧,ASR 根本处理不了这个,但一个 audio
19:41
understanding model that can reason over the audio plus maybe it takes prompt, plus maybe
understanding model,能对 audio 做 reasoning,再加上可能它会接收 prompt,再加上可能
19:49
other audio references can.
其他 audio references 也可以。
19:52
So these are the models that are much more similar to your, you know, favorite LLM, just
所以这些模型就更像你最喜欢的 LLM,你知道吧,只是
19:58
that they can now take different modalities as input.
它们现在可以把不同的 modalities 作为 input 了。
20:02
And you could then also consider having video as another input in addition to audio and
然后你还可以考虑,除了 audio 和
20:08
text.
text 之外,再把 video 作为另一个 input。
20:10
And you could have, you know, world model parameters or your robot or your self driving car and other
而且你可以有,你知道,world model parameters,或者你的 robot,或者你的 self driving car,以及其他
20:16
things also as inputs and then correspondingly all of that coming back out.
东西也作为 inputs,然后相应地,所有这些东西再出来。
20:21
So if you think about this is basically the tokens are like your bus on the back where
所以如果你想想,这基本上就是,tokens 就像是你背后的 bus,在那里
20:27
all the information gets sent through and then comes back out again.
所有信息都被送过去,然后再出来。
20:31
So it's, this is really the glue that ties everything together.
所以,这真的是把所有东西粘合在一起的粘合剂。
20:36
Okay, so we started releasing those models.
Okay,所以我们开始 release 那些 models 了。
20:40
And I think by now we have something that has very good latency.
而且我觉得到现在,我们已经有了 latency 非常好的东西。
20:45
Then we started having to really do performance tuning to make those models really interactive.
然后我们开始真的得做 performance tuning,才能让那些 models 真正 interactive。
20:55
Again, extra work.
又是额外的工作。
20:57
And I think by now we're in a decent position.
而且我觉得到现在,我们处在一个还不错的位置。
21:01
So I'm pretty happy and proud about what the team built.
所以我对团队做出来的东西挺开心,也挺自豪的。
21:04
And so along that path, what was the, when you think about kind of the significant technical
那么沿着这条路,那个,当你想到那些重大的 technical
21:10
challenges that you ran into that, you know, the team really had to, you know, go heads
challenges,你们遇到的那些,你知道,团队真的不得不,你知道,go heads
21:16
down and you think came up with a clever solution.
然后你觉得你想出了一个巧妙的解决方案。
21:21
Like, you know, talk about some of those key technical challenges.
就像,你知道的,聊聊其中一些关键的技术挑战。
21:25
So I think one of the things that are a meaningful differentiator is that we can process
所以我觉得,其中一个有意义的差异化因素是,我们能够处理
21:34
and have a lot of volume data.
并且拥有大量的 volume data。
21:38
And that's a meaningful note.
而且这是一个很有意义的点。
21:40
That's not the least training data set that you've collected or throughput or something
这不是你收集的最不重要的 training data set,也不是 throughput 或别的什么。
21:46
else.
不不,其实就是 training data set。
21:47
No, no, it's the training data set really.
不不,其实关键就是 training data set。
21:49
So we have in the order of 100 million hours of audio.
所以我们大概有一亿小时的 audio。
21:54
And it's about 200 human lifetimes.
这大约相当于 200 个人的一生。
22:00
If you live for 75 years in a very noisy environment, you would get about that amount of audio,
如果你在一个非常嘈杂的环境里活 75 年,你大概会接收到差不多这个量的 audio,
22:08
not necessarily spoken, but that amount of audio.
不一定是人说话,但就是这么多 audio。
22:13
But, you know, you can actually, you know, go forth and scrape and crawl a fair amount
但是,你知道,你其实可以,你知道,去网上 scrape 和 crawl 到相当多的这种数据。
22:17
of that data online.
但之后你需要 process 它,你需要 extract、tag、normalize、transcribe,会有
22:21
But then you need to process it, you need to extract, tag, normalize, transcribe, there's
很多额外的 processing 要做。
22:27
a lot of extra processing that's needed.
需要很多额外的 processing。
22:31
And having our own data center really helped us there.
而且拥有我们自己的 data center,这一点真的帮了我们大忙。
22:35
If you were to store this data on a new cloud, the storage bill would eat you alive and
如果你把这些数据存到一个新的 cloud 上,storage 账单会把你活活吃穷;而且,是啊,如果你用的是那些又大又好用的 cloud providers 中的一家,那就更麻烦了。
22:42
yeah, if you were on one of the big sweet cloud providers, it would become even more problematic.
所以至少直到最近 hard drives 又开始变得特别贵之前,这都还是一个非常好的局面。
22:52
So at least until recently when hard drives started becoming really expensive again, this
我是说,到现在 hard drives 的价格已经差不多是一年前的三倍了。
22:59
was a very nice situation.
所以,好吧,我们现在得稍微更谨慎一点了。
23:02
I mean, by now the hard drives are about three times the price of what they were a year
我是说,到现在为止,hard drives 的价格大约是它们一年
23:07
ago.
前的三倍。
23:08
So, okay, we'll have to be a little bit more prudent now.
所以,好吧,我们现在得稍微更谨慎一点了。
23:12
And so talk a little bit more about this data set, where how is it sourced?
那再稍微多聊聊这个 data set 吧,它是怎么来的?
23:19
Is it kind of internet style videos and maybe you extract audio from video, that kind
是那种互联网风格的 video,然后你可能从 video 里提取 audio,那种
23:26
of thing?
东西吗?
23:27
This and many other things, the one thing we didn't do is we did not spend unreasonable amounts
这个以及很多其他事情,我们唯一没做的一件事,就是没有花不合理的钱
23:34
of money on an annotation company.
去请一家 annotation 公司。
23:38
So every once in a while, it emails from a company saying, hey, we can annotate $10,000
所以时不时会有公司发邮件说,嘿,我们可以帮你 annotate 价值 $10,000
23:44
of audio for you and at that point, I'm like, okay, that's good for you.
的 audio,这时候我就会想,好吧,那对你们挺好。
23:53
We are at between 10, we're at 10,000 times that scale.
我们处在 10 之间,我们是在那个规模的 10,000 倍。
23:58
And so is that because the existing technology, like the existing tools are sufficient
所以这是因为现有技术,或者现有工具已经足够好,让你可以……这里面有很多复杂的地方,也投入了很多 engineering,我觉得这也是我们现有资产里很有意义的一部分。
24:08
enough that you can, there was a lot of intranscrime and this was, there's a lot of engineering
所以抱歉,我可能说得有点含糊,但没错,很多工作就是花在这上面了。
24:14
that went into this, that's I think part of, I think, what's a meaningful asset of what
但总的来说,就是你一开始会觉得,data 是不是越多越好。是的,data 越多越好。
24:23
we have.
而且是的,基本上,这么说吧,它就是从 internet 上获取的。
24:24
So biopologies for maybe being a little bit vague here, but yeah, that's where a lot of
所以,抱歉我可能在这里说得有点模糊,但 yeah,这就是很多……
24:31
work went into and that.
投入进去的工作,之类的。
24:33
But speaking in general, it is you start with, is more data is more better.
但总的来说,就是你一开始会觉得,data 越多越好。
24:38
And yes, it's, it's basically, let's put it this way, it's sourced on the internet.
对,它,它基本上,这么说吧,来源就是 internet。
24:43
And, you know, there are different types of data with different types of metadata that
而且,你知道,有不同类型的 data,也有你能获取到的不同类型的 metadata。
24:48
you can get.
然后你得是个好工程师,能识别出你能找到什么。
24:49
And then you need to be a good engineer and recognize what you can find.
这也让我想到一个问题,就是我们现有的工具,它们并不完美。
24:55
It prompts for me a question about the existing tools we have, it aren't perfect.
所以如果你是从 internet data 开始,用不完美的工具处理这些,你就会得到一个 noisy 的、相当 noisy 的 label set。
24:59
So if you're starting with internet data and processing those with imperfect tools, you
然后你拿它来做 training,就像,你知道,那样,那样会更好吗?
25:04
have a noisy, a fairly noisy label set.
还是更糟?
25:08
And you're using that for training, like, you know, is that, is that better?
然后你用它来做 training,就像,你知道,那,那会更好吗?
25:13
Is it worse?
还是会更差?
25:14
Is it great?
这很棒吗?
25:15
It's on unique challenges?
它是讲独特挑战的吗?
25:16
Like talk about the relationship between, you know, that, okay, so there's a couple of
就像聊聊,你知道,那个……之间的关系,好吧,所以有几
25:21
things.
件事。
25:22
First of all, I mean, we know that we, it's possible to extract meaningful information
首先,我是说,我们知道,我们,是有可能提取出有意义的信息的
25:29
even from noisy data.
哪怕是从 noisy data 里。
25:34
I mean, the simplest analogy is, let's say you have a really awful voltmeter and you want
我是说,最简单的类比就是,假设你有一个特别糟糕的 voltmeter,而你想
25:37
to find out what the voltage in your outlets in the houses, so you go around and plug it
弄清楚你房子里插座里的 voltage 是多少,于是你就到处跑,把它插上
25:42
in many times, you get the numbers out and you average in the end and you get something
很多时候,你把数字拿出来,最后做平均,就能得到某个结果
25:46
that's better than what an individual measurement of the voltmeter will do.
这比 voltmeter 单次测量给出的结果更好。
25:51
And that of course only works if your voltmeter is unbiased, if it has biased in that procedure
当然,这只有在你 voltmeter 是 unbiased 的情况下才成立;如果它是 biased 的,那这个做法
25:57
is useless.
就没用了。
25:59
That procedure has been around for ages, in the middle ages, there was something called
这个做法已经存在很久了;在中世纪,有一种东西叫
26:06
the foot rule where you estimate the foot by just having people walk out of church and
foot rule,就是让人们走出教堂来估计一英尺,然后
26:11
you grab the, the first 12 men, you send the two with the fouries and two with the longest
你抓住,呃,最先出来的 12 个男人,把两个脚最短的和两个脚最长的
26:17
feet away, probably due to deformity, take the other eight, take, average their length
排除掉,可能是因为畸形;然后取另外八个,取,平均他们的长度
26:23
and you've got a pretty good estimate of one foot, right?
那你就对一英尺有了一个相当不错的估计,对吧?
26:29
Now, okay, sorry for maybe giving a very food statistics explanation, but that's literally
现在,好吧,抱歉可能给了一个非常 food statistics 的解释,但这真的就是
26:35
where the foot rule and trim mean estimators come from, right?
foot rule 和 trim mean estimators 的来源,对吧?
26:39
So robust regression and estimation is half a million year old.
所以 robust regression 和 estimation 已经有五十万年的历史了。
26:46
Now in, with audio, right, I mean, you can, first of all, you know, it's not unreasonable
现在,在 audio 这块,对吧,我的意思是,你可以,首先,你懂的,这并非不合理
26:54
to assume that by having, you know, a lot of data, you can estimate a better model.
去假设,你懂的,通过拥有大量 data,你可以估计出更好的 model。
27:02
It's also not unreasonable to assume that once you have a better model, you can get better
同样并非不合理的是,假设一旦你有了更好的 model,你就能得到更好的
27:06
annotation, right?
annotation,对吧?
27:09
So basically all of the, you know, learning from weak supervision, all those ideas are
所以基本上,你知道,所有从 weak supervision 里学到的东西,所有这些想法都
27:16
applicable.
适用。
27:17
In some cases, you also have context, and the context helps you more.
有些情况下,你也有 context,而 context 会帮你更多。
27:22
And so, with that context of privilege or side information, again, you can do a little
所以,有了那种 privileged context 或者 side information,同样,你可以做得稍微
27:27
bit better.
好一点。
27:28
So, for instance, knowing that this podcast is between two people, and maybe I can even
所以,比如说,知道这个 podcast 是两个人之间的,也许我甚至可以
27:36
see that in, you know, the annotation or, you know, whatever, or having seen that other
在,你知道,annotation 里看到这一点,或者,你知道,随便什么,或者已经看到其他
27:44
podcasts on Twimble are usually Sam and one other person.
Twimble 上的 podcast 通常是 Sam 和另一个人。
27:49
So then you can go and use that to figure out if there are only two people.
那么你就可以用这个来判断是否只有两个人。
27:54
It's very easy to separate those speakers, and then, you know, you get larger amounts
分离那些 speakers 非常容易,然后,你知道,你就能得到更多
27:58
of, you know, single speaker audio.
的,你知道,single speaker audio。
28:02
It's also reasonably easy to tell, I guess, the two of our voices apart, which again
而且,我觉得,分辨我们俩的声音也相当容易,这又
28:09
helps you to then, you know, annotate a little bit better.
能帮你,然后,你知道,更好地 annotate 一点。
28:13
And so basically having a lot of this stuff allows you to build better models, and then
所以基本上,拥有大量这样的东西让你能构建更好的 models,然后
28:22
you get that flywheel.
你就能得到那个 flywheel。
28:25
I mean, if you think about it, people have used similar tricks for face recognition, where
我的意思是,如果你想想,人们在 face recognition 中用过类似的技巧,其中
28:31
usually matched impaired faces are not that easy to come by.
通常匹配好的受损人脸并不是那么容易找到。
28:36
But for instance, if you have a movie, you know that the actor may look very differently
但比如说,如果你有一部电影,你知道演员在整个电影中可能看起来非常不同,
28:41
throughout the movie, but it's still the same actor, and they're not that many actors,
但仍然是同一个演员,而且演员数量并不多,
28:44
so what you can do is it's not that hard to identify the same actor within the movie,
所以你能做的是,在电影中识别出同一个演员并不太难,
28:51
but then you can go and shuffle all the actors together.
但然后你可以把所有演员混在一起打乱。
28:54
And suddenly you get a much harder dataset where you have effectively labels that are
突然间你就得到了一个难得多得多的 dataset,而你的 labels 实际上
28:59
with very high likelihood, very good.
有非常高的概率,非常好。
29:03
So this is a lot of, you know, classical machine learning and statistics, and of course
所以这涉及大量的,你知道,经典的 machine learning 和 statistics,当然
29:08
it will help, right?
这会有帮助的,对吧?
29:10
So you do the right thing, and it's it's work, but it helps.
所以你做对的事,而且它、它确实费功夫,但它有帮助。
29:15
I mean, this is about, I think, as specific as I can be, but I think, I think, I think
我是说,这大概是,我觉得,我已经尽量说得具体了,但我觉得,我觉得,我觉得
29:22
most, most statisticians can do that, but you still need to do it at scale.
大多数、大多数统计学家都能做到,但你仍然需要 at scale 地去做。
29:31
And so, you know, you need, you need to invest computed scale, you need to invest into
所以,你知道,你需要,你需要投入 computed scale,你需要投入到
29:37
storage at scale for that.
storage at scale,为了这个。
29:40
So a non-trivial amount of our resources was invested in data, not just training.
所以,我们有相当一部分资源投在了 data 上,而不只是 training。
29:47
With regards to training, you mentioned earlier on that it was clear that you didn't want
关于 training,你之前提到过,很明显你不想
29:55
to build an LLM.
来构建一个 LLM。
29:57
Does that mean that your models are primarily, you know, fine-tunes or somehow based on
这是说你们的 models 主要是,你知道,fine-tunes,或者某种程度上基于
30:04
on other LLMs or how do you describe the model approach that you took for training, you
其他 LLM,还是说你会怎么描述你们用于 training 的 model 方法,你
30:12
know.
知道。
30:13
So it would be very nice if we could just go and take a model and fine-tune it, but, you
所以如果我们能直接拿一个 model 来 fine-tune 它,那就太好了,但是,你
30:20
know, modifying those models is fairly major modification, right?
知道,修改那些 models 其实是相当大的改动,对吧?
30:29
So there's a full pre-free mid and post-training stage like in LLM.
所以就像在 LLM 里一样,有一个完整的 pre-free mid 和 post-training 阶段。
30:35
And so what goes in is substantially different from what you get in the end, right?
所以输入进去的东西和最后得到的东西差别很大,对吧?
30:44
So they may share a non-trivial part of the same architecture, but there's a lot of training
所以他们可能在同一个 architecture 里共享了相当大一部分,但也会有很多 training
30:51
that happens even on the LLM side again in order to make sure that the models retain
即便在 LLM 这边,同样会发生,这是为了确保这些模型保留
30:56
the ability.
这种能力。
31:00
So this is basically whatever you can get on public models is just a good prior and then
所以基本上,你能从 public models 上拿到的任何东西都只是一个不错的 prior,然后
31:06
you optimize from there.
你就在那基础上 optimize。
31:08
Have you published or said like what your base model or models are or is that proprietary?
你们有发布过或者说过你们的 base model 或者这些模型是什么吗,还是说那是专有的?
31:15
It depends, so in some cases it's proprietary, so, for instance, one model that we released
这要看情况,所以有些情况下它是专有的,比如,我们去年发布的一个模型
31:22
last year was built on top of LLM and we acknowledged them appropriately in our model
就是构建在 LLM 之上的,我们也已经在我们的模型里向他们做了适当的致谢
31:28
card, right?
card,对吧?
31:30
So that basically some of those design decisions are a little bit variable for, you know,
所以基本上,其中一些设计决策会有一点点变数,取决于,你知道,
31:39
the, whether we just build an audio up model or an ASR model or whether it's an audio
那个,我们是直接构建一个 audio up model,还是一个 ASR model,或者它是不是一个
31:44
understanding model.
audio understanding model。
31:45
So you do need the full pipeline, but we can either, you know, build and train on top
所以你确实需要完整的 pipeline,但我们要么,你知道,可以在
31:56
of an existing model or we can build our own.
现有 model 的基础上构建和训练,要么我们可以自己构建。
32:01
But that's then often an economic decision and also, you know, which features you need.
但那通常是个经济上的决定,而且,你知道,还取决于你需要哪些 features。
32:06
So for instance, if you need particular performance in a particular language, then we may very
所以比如说,如果你在某个特定语言上需要特定的 performance,那我们可能非常
32:10
well decide to invest also significantly more still on the language ability there.
好吧,也决定还要在 language ability 那方面投入明显更多。
32:15
So it's, it, it, it's really depends on, on the use case.
所以这,这,这真的取决于,取决于 use case。
32:22
But you know, obviously, you know, if somebody gives you free steel, you don't build a steel
但你知道,很明显,你知道,如果有人白送你钢铁,你不会去建钢铁
32:28
mill, you build a car, right?
厂,你会去造车,对吧?
32:32
And you know, as long as, you know, those models are available, it makes sense to take
而且你知道,只要,你知道,那些 models 可用,去利用
32:40
advantage of it.
它就是合理的。
32:41
It will be economically foolish not to.
不这么做在经济上会很愚蠢。
32:45
But I would say the work that's required in building the audio models is somewhat commensurate
但我会说,构建 audio models 所需的工作量,某种程度上是相称的。
32:53
with the line with the language model itself, right?
跟 language model 本身那条线,对吧?
32:59
Maybe it's half or one third, but it's not one tenth.
也许是二分之一或三分之一,但不是十分之一。
33:04
And so the other thing is, for instance, if you want to have models that are able to then
所以另外一点是,比如说,如果你想要让 models 能够接着
33:11
have a higher amount of intelligence in the background, you need to, you know, specifically
在后台拥有更高程度的 intelligence,你就需要,你知道,特别
33:18
make sure that they can do this while there are still, you know, maintaining a conversation
确保它们能做到这一点,同时仍然,你知道,维持一场对话
33:24
in the foreground.
在前台。
33:26
And I mean, humans are pretty good at that.
而且我的意思是,人类在这方面挺擅长的。
33:28
So for instance, you may, you know, idly chat with somebody and in the back of your
所以比如说,你可能会,你知道,跟某人随便闲聊,而在你的……深处
33:37
head, you're thinking about something else like, well, do they add some coins to the
脑子里,你在想别的事情,比如,呃,他们要不要往停车计时器里投几个硬币,或者我真的得走了,又或者你其实在试图解决一个棘手的技术问题。
33:42
parking meter or I really need to go or you're actually trying to solve a difficult technical
所以,比如说,在考试里,你可能在拖延时间,而在脑海深处,你正在疯狂地思考解题方法,对吧?
33:50
problem.
所以那里面发生的事可能,你知道,相当多变,但人类很擅长这个。
33:51
So, for instance, in an exam, you may be stalling for time while in the back of your head,
机器仍在不断进步,而这实际上也是那里最令人兴奋的前沿之一。
33:59
you have feverishly thinking about the solution, right?
你会拼命地想着解决方案,对吧?
34:01
So what goes on there may be, you know, quite variable, but humans are pretty good at that.
所以那里发生的事可能,你知道,挺多变的,但人类挺擅长这个的。
34:11
Machines are still improving and that's actually also one of the really exciting frontiers
机器还在不断进步,而这其实也是一个非常令人兴奋的前沿领域。
34:17
there.
那里。
34:18
How do you design architectures?
你是怎么设计 architectures 的?
34:20
Talk a little bit more about that.
再多聊聊这个吧。
34:22
That's a great segue to the question that I wanted to ask, which is really around
这刚好能很自然地过渡到我一直想问的问题,其实就是关于
34:28
how you do that.
你是怎么做到的。
34:30
As you were talking about the challenges of audio AI earlier, one of the questions that
你前面聊到 audio AI 的挑战时,我心里冒出的一个问题
34:40
arose for me is, you know, are we able to do this all with a single model?
是,你知道,我们能不能只用一个 single model 就把这些全做了?
34:47
Does getting into hierarchical models, you know, your system one system two kind of bolt
是不是要开始用 hierarchical models,你知道,把你的 system one、system two 那种东西拼
34:52
it together to create a fluid experience?
在一起,创造出流畅的体验?
34:55
Does that ever make sense?
这真的说得通吗?
34:56
Like, how have you come to think about, you know, overall architecture for these types
就比如,你,你知道,是怎么开始考虑这些类型的
35:01
of systems?
系统的整体 architecture 的?
35:02
I mean, there are some systems that are just, you know, into into audio.
我是说,有些系统就是,你知道,专门做 audio 的。
35:07
And, you know, they have their place for something that's, you know, super responsive, fairly
而且,你知道,它们是有用武之地的,对于某种,你知道,响应超快、相当
35:13
small, fairly low latency, but also fairly dumb.
小、相当 low latency,但也相当笨的东西。
35:18
The problem is if you were to try and make them really smart, they would become an affordable
问题是,如果你想试着让它们变得非常聪明,它们就会在
35:24
in terms of compute costs, so you really don't want to do that.
compute costs 方面变得负担不起,所以你根本不想那么做。
35:28
Now, what you can do is you can have something that's maybe a little bit more of a two-stage
现在,你能做的是,你可以拥有一种可能稍微更偏向 two-stage 的
35:35
architecture where the understanding and reasoning happens in one end with the appropriate
architecture,其中理解和推理发生在一端,并带有适当的
35:41
tool calls being fired off in the back end, maybe then the answer's being received.
tool calls 在 back end 被触发,然后可能收到答案。
35:47
So it's really just like you would also do in, you know, multistrated programming, right,
所以这其实就像你在,你知道,multi-threaded programming 里也会做的那样,对吧,
35:53
where you have a main thread, and it may fire off other threads in parallel that may
你有一个 main thread,它可以 parallel 地触发其他 threads,这些 threads 可能会
35:58
do things, and then get the answer back such that you can get the nice data flow.
做一些事情,然后把答案拿回来,这样你就能获得漂亮的 data flow。
36:06
And I think a lot of tool calls, agentic programming, and so on, have actually been
而且我觉得很多 tool calls、agentic programming 等等,实际上已经
36:11
quite helpful there.
相当有帮助了。
36:14
The other thing is it also allows us to build models that, you know, you can instruct
另一件事是,它还让我们能构建一些 model,你知道,你可以去 instruct 它们,
36:20
just like it would, you know, language agent, you know, the prompts that you write for
就像你在 instruct 一个 language agent 一样,你知道,你为
36:26
our model are look and feel the same as if you were just, you know, instructing a regular
我们的 model 写的 prompts,看起来和感觉起来都跟你只是在,你知道,instruct 一个普通的
36:32
LLM, just that our model also talks.
LLM 一样,只不过我们的 model 还会说话。
36:37
So that makes it much, much easier for humans to work with that.
所以这让人类跟它打交道变得容易得多得多。
36:44
Of course, in the back, there is a lot of engineering going on deciding, you know,
当然,背后有大量工程在做决策,你知道,
36:49
going to do a tool call, when to answer it directly.
要不要去做一个 tool call,什么时候直接回答。
36:53
And then you also need to figure out, you know, when is the question easy enough that the
然后你还得弄清楚,你知道,问题什么时候足够简单,以至于这个
36:57
model can answer it, and when is it sufficiently hard that, you know, you probably need something
model 能回答它,那什么时候它才算足够难,你知道吧,难到你可能需要点别的东西?
37:04
else.
而且,这又取决于 model size,也取决于需要哪些信息。
37:05
And again, that depends on the model size, depends on the information that's required.
所以你需要一个 database query,需要跟 MCP server 交互,对吧。
37:11
So you need a database query, you need to talk to an MCP server, right.
而且,对,这其实就是我觉得我们花了很多时间的地方。
37:19
And yeah, that is actually something where we, I think, spent a lot of time.
我们之所以花这么多时间,是因为我们想确保这个 model 的成本是可接受的,对吧。
37:26
And the reason why we spent a lot of time is because we want to make sure that the model
所以你很容易就能搭出一个非常大的 model。
37:29
is affordable, right.
还算负担得起,对吧。
37:32
So you can easily build a very big model.
所以你可以很轻松地构建一个非常大的 model。
37:35
And that then uses a massive GPU.
然后那会用上一块巨大的 GPU。
37:39
And unfortunately, the user in the end pays for that massive GPU.
而不幸的是,最终是用户为那块巨大的 GPU 买单。
37:43
So right now, on benchmarks, we are better than, let's say, GPT and Gemini and GROC
所以现在,在 benchmarks 上,我们比,比如说,GPT、Gemini 和 GROC 更好
37:55
at a fraction of the cost.
成本却只有它们的一小部分。
37:59
So we are better than open eyes models at one tenths the cost, caveat with that on what
所以我们比 open eyes models 更好,成本只有十分之一,但有个前提,得看在哪些
38:06
metrics.
metrics。
38:08
So intelligence, you know, speech accuracy, all of the above.
所以 intelligence,你知道,speech accuracy,所有这些。
38:13
So this is, for instance, for, you know, big bench audio or complex function bench and
所以这比如说,是针对,你知道,big bench audio 或 complex function bench,以及
38:20
so on.
等等。
38:21
So there's a couple of corresponding benchmarks.
所以有几个对应的 benchmarks。
38:24
Now the little footnote is this is with thinking turned off in these models.
现在有个小脚注:这是在把这些 models 里的 thinking 关掉的情况下。
38:30
Now you might wonder, you know, why on earth would you turn off thinking?
现在你可能会想,你知道吧,你到底为什么会想把 thinking 关掉呢?
38:33
Well, because you don't want to have the big pause where the model thinks and then
嗯,因为你不想要那种大大的停顿——model 先思考,然后
38:40
it responds, right, because that feels very unnatural.
再回答,对吧,因为那感觉非常不自然。
38:45
So humans don't do that either.
所以人类也不会那样做。
38:48
I mean, they may say, hey, let me think.
我是说,他们可能会说,嘿,让我想想。
38:50
But that's now, you know, you need to correspondingly integrate that.
但那是现在,你知道,你得相应地把它整合进去。
38:55
I think that's what that was the actually the thought that led to kind of this hierarchical
我觉得那其实就是那个想法,促成了这种 hierarchical
38:59
thing.
的东西。
39:00
Like I'm envisioning a model whose primarily function, whose primary function is maintaining
就像我在设想一个 model,它的主要功能,它的主要功能是维持
39:06
the conversation and it might say, oh, that's a really interesting question.
对话,它可能会说,哦,这真是个很有意思的问题。
39:11
I'll have to think about that for a second and kind of like you alluded to kind of stalling,
我得先想一下这个,而且就像你提到的,有点像在拖延,
39:16
you know, while the tool is completing to retrieve the information.
你知道,等 tool 把信息 retrieve 完的时候。
39:22
So, so, so for instance, you know, if you test our, you know, you can test it out actually
所以,所以,所以比如说,你知道,如果你测试我们的,你知道,你其实可以测一测。
39:30
on our demo live afterwards.
之后在我们的 demo live 上。
39:32
So basically, you know, for the, you know, X live demo, this performs web search in the
所以基本上,你知道,对于这个,你知道,X live demo,它会在
39:42
background.
后台进行 web search。
39:44
And depending on how quickly gets the result back, it will just answer what it will actually
而根据它多快拿到结果,它就会直接回答;它实际会
39:50
tell you, hey, let me look for that.
告诉你的是:嘿,让我去找一下那个。
39:53
And the challenge is now to make this all feel very organic, such that, you know, the user
而现在的挑战是让这一切感觉非常自然,就是,你知道,让用户
39:59
doesn't feel any, you know, breakage of, you know, the entire interaction, does he really
感觉不到任何,你知道,整个互动的,你知道,断裂感,他是不是真的
40:05
want to maintain that illusion of, you know, a properly engaged other party.
想维持那种错觉,你知道,就是一个真正投入对话的另一方?
40:12
And so, you know, you don't want to break, you know, that sense that the model is able
所以,你知道,你不希望破坏那种感觉,你知道,就是 model 能在后台搜索、能做所有这些事情。
40:19
to search and do all of those things in the background.
那,那,那也就是,嗯,语音应用里相当一部分 engineering 发挥作用的地方。
40:25
That, that, that's where, yeah, fair amount of the engineering comes in for speech applications.
你想用 MCP servers 做的大多数事情,都能用现成的 MCP servers 吗?
40:33
Are you able to use off the shelf, MCP servers for most of the things that you might want
能,也不能。
40:40
to use them?
所以,是的,我们可以,但现在能同时启用的 servers 数量还有点有限。
40:41
Yes, and no.
是,也不是。
40:42
So, yes, we can, but the quantity of servers that enable at the same time right now is a
所以,是的,我们可以,但现在能同时启用的 server 数量有点有限。
40:51
little bit limited.
这非常有意思。
40:52
The next iteration is going to support significantly larger numbers of them at the same time.
下一个迭代会同时支持显著更多数量的它们。
41:00
Basically, this is all about, you know, context window sizes and so on.
基本上,这全都是关于,你知道,context window 的大小等等。
41:04
And again, you know, cost of inference.
而且,你知道,inference 的成本。
41:08
So there's a, there's a trade off because if you do quite a massive free fill, and then
所以有一个,有一个权衡,因为如果你做一个相当大规模的 free fill,然后
41:13
you have a large KV case that you need to log around with you, that of course, you know,
你就有一个很大的 KV case,需要一直带着走,那当然,你知道,
41:19
makes the token generation more expensive.
会让 token generation 变得更贵。
41:23
And this is really a little bit the trade off, but what you can do is you can then have
这其实有点像是权衡,但你能做的是,你可以接着有
41:28
tool calls and the tools themselves have MCP configured and so on.
tool calls,然后这些 tools 本身配置好 MCP 等等。
41:32
So there's plenty of nice ways how you can do this efficiently.
所以,有很多不错的方法可以让你高效地做到这一点。
41:38
Of course, you can also do things where you build, you know, you know, hundreds of billion
当然,你也可以做一些事情,比如构建一个,你知道,你知道,几千亿 parameters 的 front and audio model。
41:44
parameters front and audio model.
那能做出很惊艳的 demos,但一旦算上成本,就没那么惊艳了,你懂的,嗯。
41:47
And that makes for gorgeous demos, but not so gorgeous, you know, cost as soon as, yeah.
我觉得,这其实正是我们采取了一个稍微有点反主流策略的地方——我们先追求可负担性。
42:01
This is really, I think, where we took a slightly contrarian approach where we went for
这非常有意思。
42:07
affordability first.
这也跟我刚才问的有点不一样。
42:09
It's very interesting.
这也跟我刚才问的有点不太一样。
42:10
It's also a little different from what I was asking.
这也跟我刚才问的有点不一样。
42:15
And the experience that kind of, that was the motivation for my question is, yeah, as
然后那个经历,某种程度上,就是我提出这个问题的动机,是的,因为
42:23
I build agentic systems with just kind of off the shelf, MCP servers, internet, you know,
我用那种现成的、MCP servers、互联网、你知道、
42:31
database services, you know, personal data, they're slow, they're really slow.
数据库服务、你知道、个人数据来构建 agentic systems 的时候,它们很慢,它们真的非常慢。
42:39
And it's annoying even in text.
而且就算在文本里也很烦人。
42:41
It's unimaginable in a voice scenario.
在 voice 场景里简直无法想象。
42:44
And so really the question I was asking was almost, do you have to build everything
所以其实我当时问的问题几乎是,你是不是必须从
42:49
from the ground and from the ground up in order to make it work for voice like, I'm imagining
零开始、从底层开始构建一切,才能让它适用于 voice,就像,我在想
42:55
the way you would build a weather MCP server is a lot more efficient than, you know, in
你构建一个 weather MCP server 的方式会比,你知道,在
43:01
an ideal world.
一个理想的世界。
43:02
As a matter of fact, you can test out, you know, you know, it's basically search, news,
事实上,你可以测试一下,你知道,你知道,基本上就是搜索、新闻,
43:08
and weather.
还有天气。
43:09
And it feels very fluid in our application.
而且在我们应用里感觉非常流畅。
43:12
So you can go to, you know, buzz on the eye.
所以你可以去,你知道,buzz on the eye。
43:16
And presumably you didn't have to build those yourself, you're just using.
而且想必那些你不用自己搭,你只是在用。
43:19
So you use a good server for that.
所以你会为此用一个好的 server。
43:23
And being able to use these, I mean, it's, I think it spills a feature in a necessity.
而且能用这些,我的意思是,这,我觉得它仍然是一个 feature,也是一个刚需。
43:32
It's a necessity because you don't want to have to reinvent all the tooling again.
这是必须的,因为你不想又得把所有 tooling 重新发明一遍。
43:37
It's also a necessity because our customers don't want to have to relearn new techniques.
这也是必须的,因为我们的客户不想还得重新学新的 techniques。
43:46
And that, that's, that then also becomes a feature because it makes it easy to, or at
而且,这,这也就变成了一个 feature,因为它让……变得容易,或者说至少
43:52
least easier to integrate within an existing system, right?
更容易整合进现有系统里,对吧?
43:58
But, you know, if you have a service that takes a long time, I mean, humans are pretty
但是,你知道,如果你有一个耗时很长的 service,我的意思是,人类其实挺
44:04
good.
擅长的。
44:05
If, if I ask you a question that you have to look up, like, if I were to ask you, hey,
如果,如果我问你一个得去查一下的问题,比如,如果我问你,嘿,
44:10
when was the last time on Twimble?
你上一次上 Twimble 是什么时候?
44:14
And I don't think you remember the exact date.
而且我觉得你也不记得具体是哪一天了。
44:17
You would probably tell me something like, hey, Alex, let me look it up.
你大概会跟我说,嘿,Alex,让我查一下。
44:20
And then you'll, you know, open your laptop and do, you know, your search and maybe five
然后你会,你知道,打开你的笔记本电脑,然后,你知道,搜一下,可能五
44:25
for 10 seconds or maybe a minute or two later, you'll come back and say, oh, this was
秒或十秒钟,又或者一两分钟后,你会回来说,哦,那是在
44:31
at that and that date, right?
某某日期,对吧?
44:34
And it would feel totally normal for me, right?
而且对我来说,这感觉会完全正常,对吧?
44:37
But I would not be offended with you spending that time because you told me before, hey,
但我不会介意你花那段时间,因为你之前告诉过我,嘿,
44:43
this is going to take some time.
这要花点时间。
44:46
So humans are very tolerant to delays if they are told that they's a delay.
所以,只要提前告诉人类会有延迟,他们其实对延迟非常宽容。
44:53
My, my favorite examples are actually the boot screen from Apple.
我,我最喜欢的例子其实是 Apple 的 boot screen。
44:59
So Apple does something brilliantly sneaky there.
Apple 在这件事上做得特别聪明又鸡贼。
45:04
And I guess we all have seen the boot screens on an Apple device, which goes nicely, literally
我想我们都见过 Apple 设备上的 boot screen,它走得特别顺,真的,
45:12
and then typically before it's even completely done, then suddenly it's just to, okay, it's
然后通常还没完全走完,突然就——好吧,它已经
45:17
on, right?
开机了,对吧?
45:18
And we've also seen the infuriating boot screen on Microsoft device where it goes and goes
我们也见过 Microsoft 设备上那个让人抓狂的 boot screen,它一直转啊转
45:25
and goes.
转啊转。
45:26
And then it gets stuck at 99%.
然后它就卡在 99%。
45:29
And you wait for two minutes for the last percent to complete, right?
然后你得等两分钟,最后那百分之一才走完,对吧?
45:35
So you might wonder, you know, how does Apple do that, right?
所以你可能会想,你知道,Apple 是怎么做到这一点的,对吧?
45:40
They, you know, after all, you know, is there any secret magic?
他们,你知道,毕竟,你知道,这里面有什么秘密魔法吗?
45:43
No, actually, they lie to you.
不,其实他们是在骗你。
45:47
The Microsoft boot screen is the truth.
Microsoft 的 boot screen 才是真相。
45:51
What Apple does is very cleverly they measure the boot time that it took last time.
Apple 的做法非常巧妙,他们会测量上一次的 boot time。
45:59
And they add a tiny amount to it.
然后再往上加一点点。
46:02
And then they have a fake boot progress bar.
然后他们有一个假的 boot progress bar。
46:06
That's timed to go all the way to the end within the time that it took last time to boot.
它会设定成在上次 boot 所花的时间里,一路走到最后。
46:14
So in other words, you get this very nice progress bar that is, I can be, it means absolutely
所以换句话说,你会看到这个特别漂亮的 progress bar,也就是说,我的意思是,它完全没有任何意义。
46:21
nothing.
它只是,你知道,表示上次花了多长时间。
46:22
It just, you know, indicates how long it took last time.
你肯定对此完全没问题,对吧?
46:26
You must have perfectly fine with it, right?
所以我不知道现在是不是还是这个 implementation,但它很长时间以来都是这个 implementation。
46:28
So I don't know whether that is still the very implementation now, but it was the implementation
所以我不知道那现在是不是还是那个 implementation,但它曾经就是那个 implementation
46:33
for a long time.
很长一段时间里都是。
46:36
And this is a brilliant slate of hand where I can create a very good user experience without
而这真是一个绝妙的手法,让我可以打造非常好的用户体验,而不用
46:46
having to solve an impossible technical problem, right?
去解决一个不可能的技术问题,对吧?
46:50
The other thing that you also need to do is you need to start actually finding benchmarks
另一件你也需要做的事,是你得开始真正去找 benchmarks,
46:53
like what makes for a good user experience for interaction for voice, like, you know, is
比如,什么才算好的 voice 交互用户体验,就像,你知道,
47:00
the model properly interruptible?
这个 model 能被恰当地打断吗?
47:03
Can the model recover properly from that, right?
model 能从中恰当地恢复过来吗,对吧?
47:08
Does it know that, for instance, you know, in Japanese, it's very common to say, hey,
它知不知道,比如说,你知道,在日语里,人们很常说,嘿,
47:14
so there's no, so, you're back channeling the other person and saying, hey, I got it.
所以没有,所以,你是在给对方 back channeling,说,嘿,我懂了。
47:22
Yeah.
嗯。
47:23
That, right?
那个,对吧?
47:25
But if the agent stops for every one of those, that's going to be infuriating.
但如果 agent 每遇到一个这种情况都停下来,那真的会让人特别火大。
47:29
Exactly.
没错。
47:30
On the other hand, if that, if that person were to say, well, I don't understand, you want
另一方面,如果那个,如果那个人说,呃,我不明白,你是想让
47:38
the model to stop, right?
model 停下来,对吧?
47:40
So what that means is you need to make sure that the interruptability is really seen dependent.
所以这意味着,你得确保 interruptability 真的是 scene dependent。
47:48
So, for instance, we released a benchmark exactly on that that papers, I think we put
所以,比如说,我们发布了一个 benchmark,正好就是关于那个的,那些 paper,我想我们放了
47:55
up an archive a month ago, we looked at, you know, the level of productivity and interaction.
往上翻到一个月前的一期存档,我们看了,你知道,生产力和互动的水平。
48:01
So basically whether the model actually can pick up if there's something missing and whether
所以基本上,就是看 model 到底能不能察觉到是不是少了什么东西,以及
48:06
the model should also step in.
model 是不是也应该介入。
48:08
So basically, you do need to do extra user experience engineering and benchmarks and measurements
所以基本上,你确实需要做额外的 user experience engineering,还有 benchmarks 和 measurements
48:15
and optimization to really make the model work really well in this context.
以及 optimization,才能真正让 model 在这个 context 下跑得特别好。
48:21
So that's something that I think is quite different and new relative to what we had in
所以,我觉得这件事跟我们在
48:27
text.
text 里有的东西相比,是相当不同、相当新的。
48:28
And, you know, that's, I mean, to some extent, that's, you know, what AI for humans really
而且,你知道,那就是,我是说,在某种程度上,那就是,你知道,AI for humans 真正
48:35
means to optimize in a way that it's pleasant for humans.
也就是说,要以一种让人类觉得愉悦的方式去 optimize。
48:41
So this is a slightly different optimization track rather than models for code generation,
所以这是一条稍微不同的 optimization track,而不是用于 code generation 的模型,
48:47
right?
对吧?
48:48
They worry about task completion.
它们操心的是 task completion。
48:51
In our case, we worry about human happiness, you know, you still need task completion
但对我们来说,我们操心的是人类的幸福感,你知道,你还是需要 task completion
48:56
and all of that, but you also want to make it enjoyable.
以及所有这些,但你同时也想让它变得愉悦。
48:59
What was the name of that benchmark?
那个 benchmark 叫什么来着?
49:01
So, PROACT bench is the one for productivity and then there's another one also by the
所以,PROACT bench 是负责 productivity 的那个,然后还有一个也是由 the
49:07
same team.
同一个团队。
49:08
This is basically my Toronto team that has released it.
这基本上就是我 Toronto 团队发布的。
49:11
If you look at my blog, there is a fairly detailed analysis and discussion.
如果你看我的博客,里面有相当详细的分析和讨论。
49:16
So IH bench, which explains exactly, you know, the various metrics that we use.
所以 IH bench,它会确切解释,你知道,我们用的各种 metrics。
49:24
How you then also go and make this nicely reproducible till the R. I mean, we do use third
然后你还要怎么把它做得很 reproducible,till the R。我是说,我们确实用了 third
49:31
priority audio models in the defined benchmark.
priority audio models,在定义好的 benchmark 里。
49:36
So what happens is you have this, you actually need to measure appropriately all the interruptability
所以情况是,你有了这个之后,你实际上需要恰当地衡量所有的 interruptability
49:45
and, you know, responsiveness and whether, for instance, the audio matches what's being
以及,你知道,responsiveness,还有比如,audio 是否和正在被……的内容匹配
49:52
said.
说。
49:55
My favorite example is from the Hitchhiker's Guide to the Galaxy.
我最喜欢的例子来自 The Hitchhiker's Guide to the Galaxy。
49:59
I guess I'm showing my age here.
我猜这暴露了我的年龄。
50:01
There's the BBC series and they have this notoriously cheerful computer on a spaceship
BBC 剧集里有艘飞船,上面有台欢快得出名的电脑,
50:12
and there's a scene where the entire crew is about to die in two minutes because the
还有一幕,全体船员再过两分钟就要死了,因为那
50:19
spacecraft is going to crash into something and the computer very cheerfully announces
艘飞船马上要撞上什么东西,而那台电脑非常欢快地宣布
50:25
you will crash in two minutes, we will all die.
你们会在两分钟后坠毁,我们都会死。
50:31
And of course, this is for comedic relief and spoiler alert, no, they don't die in two
当然,这是为了喜剧效果,剧透预警,不,他们并没有在两分钟后死
50:35
minutes.
分钟。
50:36
Obviously that's an easy point that you were mentioning earlier.
显然,这是你之前提到的一个很简单的点。
50:40
Exactly.
没错。
50:41
So what I'm trying to say is that for voice and then also for video, you need to really
所以我想说的是,对于语音,然后还有视频,你真的需要
50:53
care about how humans feel rather than just doing text only and I mean, yeah, you want
在意人类的感受,而不是只做纯文本。而且我是说,对,你想要的是
50:59
models that are not as dumb as a brick, but there is a difference between IQ and EQ
别笨得像块砖头的模型,但 IQ 和 EQ 之间是有区别的,
51:04
and the latter is also important for humans.
而后者对人类来说也很重要。
51:08
I was going to ask there's quite a bit of research on EQ from a generic AI perspective,
我本来想问,从通用 AI 的角度来看,关于 EQ 其实有不少研究,
51:16
you know, probably text focused to, you know, but I'm imagining part of the kind of the
你知道,可能主要还是围绕文本,你知道,但我在想,那种
51:24
through line in our conversation is that things are just different for text, the challenges
贯穿我们对话的主线的一部分是,文本就是不一样,这些挑战
51:30
tend to be more systematic as opposed to we're just going to solve this model and put a text
往往更系统性,而不是说我们只要解决这个 model,然后放一个 text
51:35
front end.
front end。
51:36
There are a lot of moving pieces, I'm imagining that same thing is true for EQ, solving
有很多环节,我想 EQ 也是一样,从文本角度解决
51:44
EQ from a text perspective doesn't necessarily get you, you know, EQ voice AI.
EQ 并不一定能让你得到,你知道,EQ voice AI。
51:50
It will get you to parts of it, because so from a text transcript, for instance, for
它会让你实现其中一部分,因为比如从 text transcript 里,对于
51:56
interruptability, you can figure some things out, right?
interruptability,你能搞明白一些东西,对吧?
52:02
But then in some cases, it's also a matter of, you know, what's a good point of reference.
但然后,在某些情况下,这也是一个,你知道,什么才算好的参照标准的问题。
52:09
So let me give you examples of terrible points of reference.
所以让我给你举一些糟糕的参照标准的例子。
52:14
So you could think about, you know, maybe movies are a really great source of how humans
所以你可以想想,你知道,也许电影是一个很好的来源,说明人类
52:19
should interact with each other.
应该怎样彼此互动。
52:21
Well, take the romantic comedies and usually, you know, the, you know, the obsessed guy who
嗯,就拿浪漫喜剧来说,通常,你知道,那个,你知道,那个痴迷的家伙
52:32
in the end gets the girl essentially stalks her, right?
最后得到了女孩,本质上是在跟踪她,对吧?
52:37
If you did that in reality, the police would show up and lock you up, right?
如果你在现实中这么做,警察就会出现,把你关起来,对吧?
52:45
Likewise, in some other movies, people are quite liberal with their fists or kicks or
同样,在其他一些电影里,人们用拳头或脚相当随意,或者
52:51
whatever.
随便吧。
52:52
And in the end, you know, they will, you know, kiss and make up, so to say, or, you know,
到最后,你知道,他们会,你知道,可以说和好如初,或者,你知道,
52:56
be best friends again.
又重新成为最好的朋友。
52:58
Again, in reality, if you do this, you end up in jail.
但再说回来,现实中你要真这么做,最后会进监狱。
53:02
So it's very clear that a lot of social behavior in movies isn't quite so realistic.
所以很明显,电影里的很多社交行为并不那么真实。
53:10
I mean, there are more realistic things in terms of talk shows and so on.
我的意思是,就谈话节目之类的来说,有些东西还更真实一些。
53:14
And even there, I mean, I sincerely hope that nobody will train a mortal on Dr. Phil
而且即便是在那儿,我的意思是,我真心希望没人会用 Dr. Phil 来训练一个 model
53:23
and assume that this is normal human behavior.
然后还觉得这是正常的人类行为。
53:26
I mean, you know, it's entertaining TV, but to make it very clear that you do need to,
我是说,你知道吧,这节目挺有娱乐性的,但要把话说清楚,你确实需要,
53:36
you know, be more careful in terms of, you know, how interactions should really work.
你知道吧,更小心一点,去想想,你知道吧,互动到底应该怎么进行。
53:41
The good thing is that LMS by now have a decent theory of the mind, so that gets you somewhere.
好消息是,LMS 到现在对 theory of the mind 已经有了不错的理解,所以这能让你有点进展。
53:47
And then, yeah, I mean, we're now entering, you know, the exciting new world where we
然后,对,我是说,我们现在正进入,你知道吧,一个令人兴奋的新世界,在这里我们
53:53
can actually go and, you know, explore how humans really interact.
真的可以去,你知道吧,探索人类到底是怎么互动的。
53:59
We have to be very careful about it because, you know, ethics matter and you don't want
我们得对此非常小心,因为,你知道吧,伦理很重要,而且你不想
54:06
to, you know, experiment on humans, but, you know, this is, I think, a really exciting,
去,你知道吧,在人类身上做实验,但是,你知道吧,我觉得这真的是一个非常令人兴奋的,
54:13
at least not unless you do this appropriately.
至少,除非你以恰当的方式来做,否则不要这么做。
54:16
So I think it's a really exciting time of where this is going.
所以我觉得,现在这个发展方向真是一个特别令人兴奋的时刻。
54:19
And by that, you mean the technology is significant, you know, sufficiently far along enough
你说的这个意思,是指这项技术已经意义重大,你知道,已经发展得足够远了,
54:25
that we can start putting in front of real people, engaging their reactions and that kind
我们可以开始把它放到真人面前,激发他们的反应,那种
54:29
of thing, is that where you are going?
事情,这就是你说的方向吗?
54:31
That is effectively, I think, what will happen, where those models will get better.
实际上我觉得,基本上这就是接下来会发生的事,那些 models 会变得更好。
54:36
By learning how to interact with humans.
通过学习如何和人类互动。
54:41
And yeah, I mean, this is, this is, this is an exciting new world.
而且,对,我是说,这是一个,这是一个,这是一个令人兴奋的新世界。
54:48
Yeah, and that makes me think of, you know, there's a lot of conversation right now about
对,而且这让我想到,你知道,现在有很多讨论,都是关于
54:52
what cursive self improvement, right? And, you know, we, it's commonly discussed from the
什么,当然,self-improvement,对吧?而且,你知道,我们通常是从这个角度讨论的,你知道,就是我们用来构建工具或构建东西的 models,你知道,变得更好,你知道,我们就能构建更多东西,models,你知道,构建更好的 models,等等,等等。但是,是否存在某种程度,就像,那个,learning loop 可以变得有点连续,而且 models 可以,你知道,在对话中,你知道,学习什么对特定的人有效?就像那样,我认为这带来了很多东西,这种 self-improvement,就像 personalization。但这是我认为我们不太会看到的东西,你知道,我们今天在交互中看不到这一点。这正在开始,我会说。而且你有,你有许多
54:58
perspective of, you know, as the models that we use to build tools or to build things, you know,
从……的角度来看,你知道,随着我们用来构建工具或构建东西的 models,你知道,
55:03
get better, you know, we can build more things, the models, you know, build better models,
变得更好,你知道,我们就能构建更多东西,models,你知道,能构建更好的 models,
55:08
et cetera, et cetera. But is there an extent in which like the, the learning loop can become
等等,等等。但是不是存在某种程度,就是那个,那个 learning loop 可以变得
55:17
kind of continuous and the model can, you know, in a conversation, you know, learn what works for
某种连续,而且 model 可以,你知道,在一次对话中,你知道,学到什么对
55:25
a particular person? Like that, I think that brings in a lot of things, this kind of self-improvement,
某个特定的人有效?就像那样,我觉得那会带来很多东西,这种 self-improvement,
55:29
like personalization. But it's something that I don't think we see very, you know, we don't see
比如 personalization。但这是我觉得我们不太,你知道,我们不太看到的东西
55:37
that in interactions today. This is starting, I would say. And you have a, you have a number of
在今天的互动里就有这一点。我觉得,这还只是开始。而且你有,你有好几条路可以走。
55:43
avenues. The closest thing I think we see is like memory in, you know, you know, to traditional
我觉得我们目前看到最接近的,就是像 memory 加到传统 LLM 里,你知道,那种感觉。
55:49
LLM like. Okay, so, so, so let's actually go over a couple of pieces there. So the recursive
好,那,那,那我们就具体过几个点。比如 recursive self-improvement,如果你能,你知道,怎么说呢,完全在一台电脑上做,它就挺便宜的。
55:56
self-improvement, it's cheap if you can do it in, you know, so to say, fully on a computer
靠的就是,你知道,让它去互动,比如说,和一个有不错的 theory of the mind 的 LLM 互动。
56:06
by just, you know, have interacting, let's say, with an LLM that has a decent theory of the mind.
所以,比如说,我问一个 LLM,那,发生了什么?对,就像,它们是个 user simulator。
56:13
So, for instance, if I ask an LLM, well, what happened? Yeah, so like, and they're a user
所以,比如说,Nvidia 就做得很好,发布了一些 digital personas 和一些 scenarios 等等。我们就用那个。
56:20
simulator. So, for instance, Nvidia did a great job at releasing some digital personas and some
而且,我们也有关于它的专业论文。所以这也不是什么秘密。
56:25
scenarios and so on. And we use that. And that, we also have pro papers about it. So it's no secret.
scenarios 之类的。我们会用那个。而且,我们也有关于它的 pro papers。所以这不是什么秘密。
56:35
What that does is it also allows us to build models that work well for humans that are not always
这么做的作用是,它也让我们能构建出对并不总是千篇一律的人类也很好用的 models——因为比如你吃饭用右手,左手就背到身后,对吧?所以,或者说,如果在泰国有人看到你的脚底,那也不是什么好事,但基本上有很多文化上的奇怪之处,你知道,这些奇怪之处对生活在那个文化里的人真的很重要,而学会这些可能也能让我们更好地 personalize,针对,你知道,特定地区、特定文化,但也能针对那个文化里的个体;然后,你知道,在另一端,你知道,对于 agents 跟,你知道,那些软乎乎的人类互动时的行为改进。
56:41
cooperative, that are maybe a little bit abrasive, that are a little bit unpleasant to deal with,
合作的,可能有点冲,可能有点不太好打交道,
56:47
right? Because, or people who might might want to break the system, right? So, you know,
对吧?因为,或者是那些可能可能想搞坏系统的人,对吧?所以,你知道,
56:55
no human is perfect and people will try to have fun with those things. But okay, so there's the,
没有哪个人是完美的,人们会试着拿这些东西找乐子。但好吧,所以有那个,
57:04
you know, given that those models might not have a decent idea of how humans work,
你知道,考虑到那些 models 可能对人类是怎么运作的没有太好的概念,
57:09
you can use that for, you know, RSI to some extent. Another extent and direction, though, is
你可以把它用于,你知道,RSI 到某种程度上。不过另一个层面和方向,是
57:19
personalizing things. So, for instance, knowing that Alex likes, you know, equations and facts
personalizing 东西。比如说,知道 Alex 喜欢,你知道,方程和事实
57:27
and technical details and maybe, maybe has a little bit of rough edges on the social side or whatever,
和技术细节,而且可能,可能在社会方面有点毛糙,或者随便什么,
57:38
you know, these are things that I can use for personalization such that, you know, next time I
你知道,这些是我可以用来做 personalization 的东西,这样,你知道,下次我
57:44
may interact with Alex, I can probably, you know, personalize better for him, right? Or, you know,
可能会和 Alex 互动,我大概可以,你知道,更好地为他做 personalization,对吧?或者,你知道,
57:52
knowing which devices I have, what my preferences are, what my language background is and all of that,
知道我有哪些设备、我的偏好是什么、我的语言背景是什么,以及所有这些,
57:58
this all can be used to improve, you know, direct personalization by just having models
这些都可以用来改进,你知道,direct personalization,只要让 models
58:05
reason over it in a similar way. For instance, if you look at Hermes agent, which then, you know,
以类似的方式对它进行 reason over。比如,如果你看 Hermes agent,它接下来,你知道,
58:12
looks at past interactions and past things that it's done and the reasons and it
会查看过去的 interactions、它做过的事情以及原因,然后它
58:17
commits that into memory. So, think of this as a glorified CRM, but now for everybody.
把这些 commit 到 memory 里。所以,可以把它看成是一个升级版的 CRM,只不过现在是面向所有人。
58:24
But at the same time, you also want to learn across all the interactions. So, for instance,
但与此同时,你也想从所有 interactions 中学习。所以,比如说,
58:29
knowing that if I insult a user, the user will not take kindly to that. I mean, okay, sure,
知道如果我冒犯了一个用户,那个用户不会喜欢这样。我的意思是,好吧,当然,
58:37
this is a, it's a, it's an egregious example and of course, nobody would actually do that,
这是一个,这是一个,这是一个很过分的例子,当然,没人会真那么做,
58:42
but the point being, you know, learning such things across is something that you can then overall
但重点是,你知道,跨场景学习这些东西,是你之后整体上
58:48
use for an improved model. So, you basically get global improvements. So, it feels very similar to
可以用来打造一个改进后的 model。所以,你基本上会得到全局提升。所以,这感觉非常像
58:55
recommender systems just, you know, now in 2026, where you have an overall behavioral improvement,
recommender systems,只是,你知道,现在到了 2026 年,你会有整体行为上的提升,
59:03
but you also have personalized improvement, the latter, improving as the model gets to know more
但你也会有个性化提升,后者会随着 model 越来越了解
59:09
about, you know, let's say Sam or about Alex, but then overall, the model getting better at dealing
关于,你知道,比如说 Sam,或者关于 Alex,而提升,但整体上,model 越来越擅长处理
59:17
with those pesky humans or maybe dealing with those pesky humans in North America versus,
那些烦人的人类,或者也许是处理 North America 的那些烦人的人类,相比,
59:23
you know, there's certain appropriate behavior in some parts of the world that's perfectly
你知道,世界某些地方有某些得体的行为,那完全是
59:31
inappropriate elsewhere. So, like eating with your left hand in India is seriously frowned upon
在别的地方就不合适。所以,比如在 India 用左手吃饭是非常不受待见的,因为你是用右手吃饭,左手要背到身后,对吧?还有,如果在 Thailand 有人看到你的脚底,那也不是什么好事。但基本上,文化上有很多怪怪的地方,而且你知道,这些怪怪的地方对生活在那个文化里的人真的很重要,学习这些可能也能让你稍微更好地 personalize,针对特定地区、特定文化,但也能针对那个文化里的个体;然后,你知道,在另一头,就是为那些跟软趴趴的人类互动的 agents 做 behavioral improvements。
59:40
because you use the right hand for eating and the left hand goes behind your back, right? So,
因为你用右手吃饭,左手放到背后,对吧?所以,
59:48
that, or if somebody sees the bottom of your feet in Thailand, it's also not a good thing,
那个,或者如果有人在 Thailand 看到你的脚底,那也不是什么好事,
59:55
but basically there's a lot of cultural weirdness and, you know, weirdness that really
但基本上,有很多文化上的奇怪之处,而且,你知道,那些奇怪之处真的
60:05
matters for the people that, you know, live in that culture and learning that will allow probably
对那些人很重要,你知道,他们生活在那种文化里,而学会这些可能也会
60:12
also to personalize a little bit better for, you know, specific regions, specific cultures,
也能更好地 personalize 一点,你知道,针对特定地区、特定文化,
60:18
but then also for individuals within that culture and, you know, at the other end, you know,
但也能针对那个文化里的个体,而且,你知道,在另一端,你知道,
60:25
behavioral improvements for, you know, agents interacting with, you know, just those squishy humans.
行为上的改进,你知道,是为了 agents 和,你知道,就是那些软乎乎的人类互动。
60:30
So, that I think is a really exciting future and by now we will be able to do this at a scale
所以,我觉得那是一个非常令人兴奋的未来,而且到现在,我们能够在一个规模上做这件事
60:41
where we can actually build things rather than just for the theories, right? There's, you know,
在这个规模上,我们能够真正建造东西,而不仅仅是为了理论,对吧?你知道,有一些
60:48
glorious work by, so for instance, if you look at my mother's age, so she builds those robots that
出色的工作,比如,如果你看看我母亲的年纪,她建造了那些机器人,它们
60:55
interact with babies and get babies to move their legs and whatever. And there's a lot of
与婴儿互动,让婴儿动动腿之类的。而且有很多
61:03
engineering and science brilliance in building the stuff.
在建造这些东西时的工程和科学上的卓越才华。
61:06
Brilliant is cool, but brilliance is not repeatable and automatable.
才华很酷,但才华是不可重复和自动化的。
61:14
What I think the future is going to be is going to be systems that automatically learn the stuff
我认为未来将会是那些能自动学习这些东西的系统
61:22
such that in the future you don't need quite as much brilliance anymore and, you know,
这样在未来,你不再需要那么多的才华了,而且,你知道,
61:29
you can just let data speak. You know, for movie recommendations, this is by now a solve
你完全可以让 data 自己说话。你知道,对于 movie recommendations 来说,这到现在基本上已经是个解决得差不多的问题了,对吧?
61:37
problem mostly, right? But for overall human behavior interactions and improvement, I think,
但就整体的人类行为交互和改进来说,我觉得,连人类自己都不是特别擅长。
61:46
even humans are not particularly good at it. If I was trying to aggravate somebody and to,
如果我想激怒某个人,你知道,就是去骚扰他,很多人可能到某个时候就会失去冷静,然后你知道,以牙还牙地回应,对吧?不是所有人类都友善。
61:54
you know, just harass them, a lot of people might at some point lose their cool and, you know,
但如果是为了,你知道,让 agent 完成某个给定的 task,你其实希望 agent 保持冷静,并且把这件事处理好,对吧?
62:03
responding kind, right? It's not all humans are nice. But for the purpose of, you know,
而且我是说,这就像你,你知道,看学校老师,你知道,看他们对孩子多有一套,能让孩子守规矩,即便那些孩子……
62:12
achieving a given task for an agent, you actually want the agent to keep his cool and to
要让一个 agent 完成某个给定任务,其实你是希望这个 agent 保持冷静,并且
62:19
to handle this, right? And I mean, this is one thing when you, you know, watch, you know,
来处理这件事,对吧?我是说,这就像当你,你知道,看着,你知道,
62:25
school teachers and how great they are with kids and to get them to behave even though those kids
学校老师,看他们跟孩子相处得多好,能让孩子守规矩,哪怕那些孩子
62:31
are very uncooperative, maybe initially, right? There's a skill. Most humans don't have that.
一开始可能非常不配合,对吧?这是一种技能。大多数人都没有。
62:39
But, you know, we can learn those skills with, eventually, with those, with these
但是,你知道,我们最终能学会这些技能,通过那些,通过这些
62:47
authentic systems. I think this is an exciting world. Well, there's a bit of a paradox in that,
真实的系统。我觉得这是一个令人兴奋的世界。嗯,这里面有点悖论,
62:53
if most humans don't have them, they're not in the data. And therefore, we won't be able to
如果大多数人都没有这些技能,那它们就不在 data 里。因此,我们就没法
62:58
easily train models to follow those patterns. Well, you get the, this, this interact, I mean,
轻松训练 models 去遵循这些模式。嗯,你懂,这个,这个交互,我是说,
63:06
you get it from, from interactions, all right? And basically, you can think of each interaction as
你是从互动里得到的,对吧?基本上,你可以把每一次互动都看成
63:14
another new experiment, a new data point. And so, eventually, you know, by some randomness of
另一个新的实验、一个新的 data point。所以最终,你知道,靠某种
63:21
exploration and so on, you'll be able to learn what works and what doesn't work. I mean, the question
exploration 的随机性之类的,你就能学会什么行得通、什么行不通。我是说,问题
63:30
then is, you know, how you use data appropriately. But, for instance, if you need to have a
然后就是,你知道,怎么恰当地使用 data。但比如说,如果你需要和某个人
63:37
difficult conversation with somebody, I mean, the, you know, the fun example is the millennial
进行一场很难的对话,我的意思是,那个,你知道,好玩的例子就是千禧一代的
63:48
shit sandwich, right? Where you have praise, critique and praise. And by doing that, you know,
shit sandwich,对吧?也就是表扬、批评、再表扬。通过这样做,你知道,
63:57
you know, that it, you know, you can, you know, deliver, you know, the critique without
你知道,这样你,你知道,就能,你知道,给出,你知道,批评,而不会
64:04
people getting to upset, right? And this is a, this is a, this is a strategy of communication,
让别人太不高兴,对吧?而且这是一种,这是一种,这是一种沟通策略,
64:11
right? Or, you know, when you learn how to become a manager, you take all those training courses,
对吧?或者,你知道,当你学习怎么当 manager 的时候,你会参加所有那些培训课程,
64:17
how do you deal with people who are happy, who are unhappy, how do you, you know, basically
你怎么跟开心的人、不开心的人打交道,你怎么,你知道,基本上
64:27
deal with people in a way that, you know, they can be productive that they can enjoy.
以一种,你知道,他们能有产出、能享受其中的方式跟人打交道。
64:34
Media techniques. And, you know, these can be learned, they can be learned from instructions,
媒体技术。
64:41
you can take Chaudini's books and they probably contain some good instructions, but we can also
而且,你知道,这些都是可以学会的,可以从指令里学会,你可以拿 Chaudini 的书,它们可能包含一些不错的指导,但我们最终也可以从数据里学习。
64:47
then eventually go and learn from data. This is, this is, I think, really going to be quite the
这,这,我觉得,真的会是一场相当令人兴奋的革命,在接下来,也许两到三年里。
64:55
exciting revolution for the next, maybe two to three years. I think it's going to move fast.
我觉得它会发展得很快。
65:01
Well, Alex, it's been wonderful to catch up with you and get into a little bit of what you're seeing
嗯,Alex,很高兴能和你聊聊,了解一下你在 voice AI 方面看到和在做的一些事情。
65:08
and working on with regards to voice AI. Thanks. Thanks for having me. Awesome. Thank you.
谢谢。谢谢邀请我。太棒了。谢谢。