Latent Space

Underwriting Superintelligence: Backing Agents you can Sue — Rune Kvist, AIUC

2026-09-16 ·

00:00
00:03
Okay, we're in a studio with Rune from AIUC, AI Underwriting Company with trusty co-host Vibhu. Welcome. Thank you. Thanks for having me. Thank you.
好,我们现在在演播室里,和来自 AIUC,也就是 AI Underwriting Company 的 Rune,以及我们可靠的联合主持人 Vibhu 在一起。欢迎。谢谢。谢谢邀请我。谢谢。
00:11
What are you announcing today? We have raised $40 million by Ribbit Capital on First Money.
今天你们要宣布什么?我们已经在 First Money 上从 Ribbit Capital 融了 4000 万美元。
00:16
You first came to our attention when Nat and Dan invested in you guys. Is the story pretty much the same? Like, are you today what you thought you were back then?
我们最早注意到你们,是 Nat 和 Dan 投资你们的时候。故事基本还是那样吗?就是,你们今天还是当初以为的那个样子吗?
00:25
When we raised our seed round, we had a hypothesis that at some point, risk was going to hold down adoption. At that point in time, that felt kind of hypothetical.
当我们融 seed round 的时候,我们有一个假设:到了某个时候,风险会拖累 adoption。当时,这感觉还挺假设性的。
00:37
And I think that is now over. Clearly the moment is now, if measles and with fable is pretty obvious that literally the binding constraint on adoption is risk.
我觉得现在已经结束了。很明显,现在就是那个时刻,如果 measles 和 with fable 已经相当明显,那真的就是,限制 adoption 的 binding constraint 就是风险。
00:47
And so for us, it feels like this is a natural continuation of the same hypothesis, but where previously the speculation now, it feels like fact.
所以对我们来说,这感觉像是同一个假设的自然延续,只是以前还只是猜测,现在感觉已经是事实。
00:53
And let's get the list of the customers that you are highlighting this part of your series A. Totally. Yeah, so we are now working with folks like Cursor, Harvey, Loveable, 11 Labs.
那我们来列一下你们在这次 series A 中重点提到的客户吧。没问题。对,我们现在正在和 Cursor、Harvey、Loveable、11 Labs 这些团队合作。
01:04
Yeah, amazing. Congrats. Thank you.
对,太棒了。恭喜。谢谢。
01:07
So you were famously one of the first, the first hired and adopted for GTM and products. I'm just kind of curious, like, what was your path into AI? Just recap.
你就是最早的那批人之一,而且很有名,是第一个被招来负责 GTM 和产品、还真正上手的人。我挺好奇的,你进入 AI 的路径是什么?简单回顾一下。
01:18
Late 2021, I sold a company, my first company, and a tech company, I had a bit of time to think about what was next. I came across a scale in a little paper.
2021 年底,我卖掉了一家公司,我的第一家公司,也是一家科技公司,然后稍微有点时间想想下一步做什么。我偶然在一篇小论文里看到了 scaling。
01:27
And that just struck me like lightning. Files is like, this is a big idea. In short, the scale in those paper just says the bigger the model, the smarter the model.
那一下就像闪电击中了我。我当时就觉得,这是个了不起的想法。简单说,那些论文里的 scaling 就是在讲:model 越大,model 越聪明。
01:37
And this is the Kaplan one, not the Chinchilla one, exactly the Kaplan one. And the important thing that clicked for me there was, oh, now capital will understand us.
而且这是 Kaplan 那篇,不是 Chinchilla 那篇,就是 Kaplan 那篇。对我来说,那里最关键、让我一下子想通的是:哦,现在资本能听懂我们了。
01:49
If you put in more money, you get more money out. And so that will kick off a hype cycle. And so you'll actually kind of, you get a sense of predictable returns, which is in fact played out.
你投进去更多钱,就能产出更多钱。所以这会引爆一轮 hype cycle。这样你其实就能,有一种可预测回报的感觉,而这件事后来确实应验了。
02:00
And so I just pack my bags. I've never been to San Francisco. I just pack my back straight out here to find the people who have written that.
于是我就收拾行李。我从没去过 San Francisco。我直接收拾东西就来了这儿,想找到写下那篇论文的人。
02:07
And at the time, they had just started a small app called on topic. It was like 40 people at the time. So drank a bunch of coffee.
当时,他们刚创办了一个叫 on topic 的小 app。那时候大概有 40 个人。于是我就喝了一堆咖啡。
02:14
Until eventually got introduced to Daria. And at the time, they were wrestling with some of these questions. So like, should we deploy our models? Should we make revenue?
直到最后,有人把我介绍给了 Daria。当时他们正在纠结这些问题。比如,我们该不该 deploy 我们的 models?我们该不该创收?
02:23
Well, how should we engage with rest of the world? They're just broken off from opening eye. And it's been publicly reported that they were kind of concerned with how they were dealing with deployment.
那么,我们到底该怎么和外面的世界互动呢?
02:33
So they were wrestling with some of those questions. At that point, this is like early fog of war, like early 2022, the sexiest program at the time was like Jasper.
他们其实只是刚从 OpenAI 分出来而已。
02:41
There's nothing out there. So where is value going to crew? What are going to be the different parts of the stack or all open questions?
而且公开报道说,他们当时有点担心自己处理 deployment 的方式。
02:47
I want to highlight to people, you ask these questions because you have a PPE background. I actually was in Singapore in one of the sort of feeder programs for preppy people for PPE.
所以他们在纠结其中一些问题。
02:58
So I had a tutor. We learned philosophy and politics and economics. But I think you're kind of like machine learning people who read the neural skilling laws paper would not necessarily draw the same conclusions that you did.
那时候,这就像早期的战争迷雾,大概是 2022 年初,当时最火的项目就是 Jasper。
03:13
Whereas any capitalist would read that and go, holy shit.
外面什么都没有。
03:18
Correct.
那价值会积累到哪里?stack 的不同部分会是什么,还是说全都是开放问题?
03:20
Who tipped you onto that paper? Because there's not a paper that you normally read, right? Like in your circles?
我想跟大家强调一下,你会问这些问题,是因为你有 PPE 背景。
03:26
Yeah, I think I'd actually ever since AlphaGo had had some appreciation that AI was a big deal. But it kind of felt...
对,其实我觉得,从 AlphaGo 那会儿起,我就一直有点意识到 AI 是件大事了。但那种感觉又有点……
03:38
It raised all these kind of interesting philosophical questions, but it was kind of not clear from afar where exactly that would go.
它带来了各种挺有意思的哲学问题,但从远处看,又不太清楚这到底会走向哪里。
03:45
But it was obvious enough that it was like this is going to be a big thing if we find the kind of right mechanisms to kind of get the technocapital machine to work on this.
但有一点已经足够明显:如果我们能找到某种合适的机制,让 technocapital machine 在这件事上运转起来,这会是一件大事。
03:55
But it was just not clear. And so I think it was somewhere in which like that became obvious.
但就是还不清楚。所以我觉得,大概是在某个节点,这一点才变得明显起来。
04:00
And also it wasn't as obvious at the time than it is now, right? It was just like, wow, this is so interesting.
而且当时它也不像现在这么明显,对吧?当时就是,哇,这太有意思了。
04:06
But it still felt coming from kind of a philosophy and economics background.
但带着哲学和经济学的背景来看,还是会觉得……
04:10
It felt like if this turns out to be true, you're going to be wrestling with all of the big questions in society.
感觉就像,如果这成真了,你就得去跟社会上所有的大问题缠斗。
04:16
Everything you've learned about politics gets thrown out of the window.
你学过的关于政治的一切,都会被扔出窗外。
04:20
Everything you've learned about economics at least gets challenged.
你学到的所有经济学知识,至少都会受到挑战。
04:23
And so what felt interesting was to be a frontier that has just ramifications across everything.
所以有意思的是,它处在一个前沿,而影响会波及所有方面。
04:29
So that's why I thought it sorted out.
所以这就是为什么我觉得它说通了。
04:32
I mean, clearly really good insight.
我是说,这显然是非常好的洞见。
04:34
For people who don't know, the PPP program is like where prime ministers are born.
对不了解的人来说,PPP program 就像首相的诞生地。
04:38
So then you end up meeting Jared.
所以后来你就见到了 Jared。
04:40
Yep. First, Darrya.
对。先是 Darrya。
04:42
Yeah. Well, I mean, like, so did you get extra insights from talking with them that you didn't get from your original hypothesis?
是啊。呃,我是说,那你跟他们聊完之后,有没有得到一些你从最初的 hypothesis 里没有得到的额外洞见?
04:50
If you read the screen almost every day, there's like very vague sketch.
如果你几乎每天都看屏幕,那上面就像有一幅特别模糊的草图。
04:53
I was like, wow, this seems kind of important. There's some lines.
我当时就想,哇,这好像挺重要的。上面有一些线条。
04:55
And a child, this seems kind of important.
然后一个孩子,这好像挺重要的。
04:58
And what I think, it seemed like I thought more about than anyone was like, what are the implications of this?
而我觉得,我比任何人都更在想的一件事是,这到底会带来什么影响?
05:03
If you really play this out.
如果你真的把这件事推演下去。
05:05
And back then, they had kind of vision documents for what the world would look like in 2026.
而且当时,他们有一些类似愿景文档的东西,描绘 2026 年的世界会是什么样。
05:10
And they were kind of in vivid detail, playing out.
而且它们还挺生动细致地推演着。
05:13
How much compute is going to be needed?
到底需要多少 compute?
05:15
What does the Catholics kind of look like?
Catholics 大概看起来是什么样子?
05:17
What are going to be some of the kind of societal concerns, but also what is the amount of economic value coming out here?
会有哪些那种社会层面的担忧,但同时这里又能产生多少经济价值?
05:23
And so it kind of felt like they held a crystal ball that in hindsight, it turned out to just be dramatically correct.
所以感觉有点像是,他们拿着一个水晶球,事后回看,结果发现它就是惊人地准确。
05:31
And they weren't holding it like they were obviously correct.
而且他们拿着它的时候,并不是那种摆明了“我们肯定是对的”的样子。
05:34
They're just like, take this hypothesis really, really seriously.
他们就像是,真的、真的要认真对待这个 hypothesis。
05:37
Think it through.
把它想透。
05:38
And think it through.
然后再把它想透。
05:39
In the same way the kind of situation awareness that is now across the street.
同样地,那种现在就在街对面的 situation awareness。
05:43
Across the state.
全州各地。
05:44
Yeah.
是啊。
05:45
Oh my god, we're all in the same.
我的天,我们都在同一条船上。
05:47
All the same across the street.
街对面也都一样。
05:49
And that's now a couple of years old.
而这已经是好几年前的事了。
05:51
But also people keep referencing it.
但人们也一直在引用它。
05:53
These particular weeks with fabled and missiles.
这几个特定的星期,伴随着 fabled 和 missiles。
05:55
And it's like, wow.
然后就是,哇。
05:56
If you take this one idea seriously.
如果你认真对待这一个想法。
05:58
For the scale and the loss.
就 scale 和 loss 而言。
05:59
A lot of things fall into place.
很多事情就都说得通了。
06:01
And keep in mind at this point, this is the same team that had did GPT one, two, and three.
而且要记住,在那个时候,这还是做出过 GPT one、two 和 three 的同一个团队。
06:06
Which is also like, it's not just some experimentation.
这也就像是,这不只是某种实验。
06:10
Like this is a real model that we just scaled up.
就像,这是一个我们刚刚 scaled up 出来的真正模型。
06:13
And they had deemed conviction in, again, in this like big,
而且他们对此——再说一次——抱着一种坚定的信念,就是在这种很大的,
06:18
if you take a big blob of compute.
如果你有一大坨 compute。
06:20
They did, it just wants to learn.
他们确实做到了,它只是想学习。
06:23
And out of that will come smile on smile models.
而由此会诞生 smile on smile models。
06:25
And all the particulars were not clear.
而且所有细节都还不清楚。
06:26
Yeah, yeah.
对,对。
06:27
And all the implications were not clear.
而且所有影响都还不清楚。
06:28
But their deep conviction is like core thesis.
但他们深层的信念就像是核心论点。
06:32
And that was kind of dizzying.
而那有点让人眩晕。
06:35
It was both phenomenally interesting and exciting.
它既极其有趣,又令人兴奋。
06:38
And also very quickly you get to like the world we know today will no longer be.
而且很快你就会
06:43
Is this hypothesis hold?
这个 hypothesis 成立吗?
06:45
So I also felt like important in some kind of grandsons.
所以我也觉得自己在某种宏大计划里挺重要的。
06:48
What kind of shaped you there?
在那里,是什么塑造了你?
06:49
So that was your early 2022 not only had GPT one, two, three come out.
所以那是你 2022 年初的时候,不只是 GPT one, two, three 已经出来了。
06:53
But you know, the amazing co-founders of Anthropic that have never split up.
但你知道,Anthropic 那些很棒的联合创始人,他们从来没有分开过。
06:57
The only ones they actually had the conviction to leave opening.
他们实际上是唯一真正有决心离开 OpenAI 的人。
07:01
I start their lab.
并创办了他们自己的实验室。
07:02
You said they were about 40 people there.
你说过那里大概有40个人。
07:04
What was the time like there?
那段时光在那里是什么样的?
07:05
It was kind of remarkably like what it looks like on the outside today.
它有点惊人地像它今天在外界看起来的样子。
07:09
Extremely cohesive.
极度有凝聚力。
07:11
Extremely misreniant it.
极其有使命感。
07:14
And living in this tension between their two ideas, which is,
并且生活在他们两种想法之间的这种张力里,也就是,
07:19
AI could both go really well and really bad.
AI 既可能发展得非常好,也可能非常糟。
07:22
And we want to be part of building it.
而我们想成为构建它的一部分。
07:24
And that creates astounding amounts of tension.
而这会制造出惊人的张力。
07:27
And they were wrestling with this incentive challenge where they know they're kind of,
而他们当时一直在纠结这个激励挑战,他们知道自己有点,
07:32
there's a race that they're in where you might get forced to cut corners.
他们身处一场竞赛之中,在这场竞赛里,你可能被迫偷工减料。
07:35
But it also felt very important to them to be at the forefront of technology.
但对他们来说,站在技术最前沿也感觉非常重要。
07:40
And all of those ideas were just present at that time.
而且当时,所有这些想法都同时存在。
07:43
It kind of feels like that line has been just very, very clear.
某种程度上,感觉那条线一直都非常非常清晰。
07:48
And I think kind of loving more hate them, they have really stuck to their guns.
而且我觉得,不管你爱他们还是恨他们,他们真的坚持了自己的立场。
07:53
There's a core set of beliefs that they hold more deeply than most companies hold any beliefs.
他们有一套核心信念,坚守得比大多数公司对任何信念的坚守都更深。
07:57
Yeah.
对。
07:58
Fast forward to the day.
快进到那一天。
07:59
Yeah.
对。
08:00
What does that lead us to?
这会让我们得出什么结论?
08:01
Yeah, and the writing company.
对,还有那家写作公司。
08:02
What do you up to?
你在忙什么?
08:03
What motivated you to start this?
是什么促使你开始做这个的?
08:05
Yeah.
对。
08:06
AIC builds confidence and structure for Frontier AI through standards and insurance.
AIC 通过标准和保险,为 Frontier AI 建立信心和结构。
08:12
The link from a topic to building confidence and structure,
从一个 topic 到建立信心和结构之间的联系,
08:15
looking at the windows of the topic offices and seeing Waymo's driving by.
看着 topic 办公室的窗外,看到 Waymo 的车开过去。
08:19
Already back there in early 2022, Waymo's run some ways like AGI for cars,
早在 2022 年初,Waymo 在某些方面就已经像汽车界的 AGI 了,
08:23
like they were super human drivers, but you couldn't take one to the airport.
就像它们已经是超人类司机一样,但你没法坐一辆去机场。
08:27
And now, four and a bit years later, you still can't take away most of the airport.
而现在,四年多一点之后,你还是没法坐 Waymo 去机场。
08:31
Despite now everyone having kind of looked at the evidence and be like,
尽管现在大家多多少少都看了证据,然后会说,
08:34
they're better drivers than humans.
它们比人类司机开得更好。
08:35
So in that particular instance, what's clear is that the binding constraint on AI being useful
所以在那一个具体例子里,很清楚的是,限制 AI 是否有用的关键约束
08:40
is not capability, but instead liability or risk or trust.
不是能力,而是责任、风险或信任。
08:45
That problem is general.
这个问题很普遍。
08:49
The reason why right now, Fable is not open for access is not because it's not a good model.
现在 Fable 之所以不开放访问,并不是因为它不是一个好 model。
08:55
It's because it's a very good model.
而是因为它是一个非常好的 model。
08:57
It's just hard to make problems about what it will will not do.
只是很难保证它会做什么、不会做什么。
09:00
And this problem gets worse as AI gets better.
而且随着 AI 变得更好,这个问题会变得更糟。
09:04
Basically, more intelligent AI can be more autonomous.
基本上,更智能的 AI 可以更自主。
09:07
That's more valuable, but also the risk surface grows.
这更有价值,但 risk surface 也会变大。
09:10
And so what Waymo illustrates is that unless you build the confidence infrastructure
所以 Waymo 所说明的是,除非你建立起 confidence infrastructure
09:16
to make promises about AI or at least bring light to the risks,
来对 AI 做出承诺,或者至少让风险被看见,
09:20
you grind adoption to a halt.
否则 adoption 就会陷入停滞。
09:22
Governments, banks, hospitals, militaries need to have some sense of what AI will
政府、银行、医院、军方需要大致了解 AI 会
09:28
and will not do to be able to operate for them to incorporate it.
和不会做什么,才能运转,才能把 AI 纳入进来。
09:31
And that's the problem that we're trying to solve.
而这正是我们试图解决的问题。
09:35
Now, why stand as an insurance?
那么,为什么要以保险的形式来做?
09:38
If you trace this problem back through history,
如果你把这个问题追溯到历史里,
09:41
every technology wave has had some version of this problem.
每一波技术浪潮都出现过这个问题的某种版本。
09:44
So if you go back to like your 1900 electricity comes out,
所以你要是回到比如 1900 年代,电一出来,
09:47
cast burn down, sorry, houses burn down, lots of people die.
cast 被烧毁,抱歉,房子被烧毁,很多人死了。
09:51
1930s, cars are a big deal, kill lots of people.
1930 年代,汽车是件大事,害死了很多人。
09:55
50s, private nuclear energy is a big deal, poses big risks.
50 年代,私人核能是件大事,带来了很大风险。
09:59
In each of those instances, the market runs the head of regulation
在每一个例子里,市场都跑在了监管前面
10:04
to create confidence infrastructure because that's required to make go no-go decisions.
去建立信任基础设施,因为这是做 go/no-go 决策所必需的。
10:07
There's required for adoption and the market fundamentally wants adoption.
采用是有要求的,而市场从根本上就想要采用。
10:11
And in all of those instances, common blueprint emerges between standards and insurance.
在所有这些情况里,标准和保险之间都会浮现出一个共同的蓝图。
10:18
The reason is these two components is standards kind of provide the rules of the road
原因是,这两个组成部分里,标准有点像是在提供道路规则,
10:22
and they also specify like what are the tests that need to be run
而且它们还规定比如说需要跑哪些测试,
10:25
so we can get a sense of how high the risk is.
这样我们就能了解风险有多高。
10:28
So taking the case of cars, like a car crash, great everyone,
所以拿汽车来举例,比如车祸,好,各位,
10:31
they inform you insurance pricing today, they inform you purchasing decisions, etc.
它们会影响你现在的保险定价,也会影响你的购买决策,等等。
10:35
That's basically the risk framework.
这基本上就是 risk framework。
10:37
The insurers are important because they pick up the bill.
保险公司很重要,因为最后买单的是他们。
10:40
So they are the private institution that is most on the side of,
所以他们是那个最站在……一边的私营机构,
10:44
there's best incentivized to quantify the risks truthfully
他们最有动力去真实地量化风险
10:49
and then figure out all the ways to reduce the risk because that increases their profit.
然后找出所有降低风险的办法,因为那会增加他们的利润。
10:53
So they're basically, they help shape the incentives.
所以他们基本上,就是在帮助塑造激励机制。
10:56
And these two worked really well in unison.
而这两者配合得非常好。
10:59
Now, how does that show up as a company?
那么,这在一家公司身上是怎么体现出来的?
11:02
Well, one of the things that was obvious,
嗯,有一件事很明显,
11:05
even though it was starting to become a museum a couple years ago,
尽管几年前它就开始变成一个博物馆了,
11:08
was that frontier companies, some of our customers today,
是那些 frontier companies,也就是我们今天的一些客户,
11:11
like Chrysler, Sierra, 11 Labs, Harvey,
比如 Chrysler、Sierra、11 Labs、Harvey,
11:15
were going to have a very easy time selling a pilot to a bank.
会非常容易把 pilot 卖给银行。
11:20
The, like the, the damage itself itself is magic.
那个,就像那个,damage 本身本身就是魔法。
11:22
But bringing that through, if you want to do a wall-to-wall rollout at a bank or a hospital,
但真要把这个落地,如果你想在银行或医院做 wall-to-wall rollout,
11:27
you have to go through the risk process.
你就必须走 risk process。
11:29
These banks have no idea even which questions to ask,
这些银行甚至都不知道该问哪些问题,
11:32
let alone which answers are sufficient, let alone,
更别说哪些答案算是足够的了,更别说,
11:35
like how do they go and test whether these ages actually work the way they're supposed to?
比如他们到底怎么去测试这些 AIs 是不是真的按预期那样运作?
11:40
And so they had, there's problems like,
所以他们就有,比如这样的问题,
11:42
what can we say to earn the trust?
我们能说些什么来赢得信任?
11:45
And we think there's like a golden sentence that goes something like,
我们觉得,大概有这么一句金句,差不多是这样的,
11:49
hey, I hear you're really worried about hallucinations or jail breaks or whatever,
嘿,我听说你真的很担心 hallucinations 或者 jail breaks 之类的,
11:53
maybe we've had an independent third party,
也许我们已经有了一个独立第三方,
11:56
testers against the gold standard,
测试人员对照 gold standard 来测试,
11:58
with passive flying colors.
以被动的方式大获成功。
11:59
And as a vote of confidence,
而且,作为一张信任票,
12:01
the world's most conservative insurers have looked at the data
全球最保守的保险公司已经看过 data
12:04
and are willing to take some of the risk on to their balance sheet.
并且愿意把一部分风险放到自己的 balance sheet 上。
12:06
Yeah.
对。
12:07
So if something does go wrong.
所以,要是真出了什么问题,
12:08
There's money behind it, yeah.
这背后是有钱撑着的,对。
12:09
Exactly.
没错。
12:10
That's kind of like the link between all of them,
这有点像把它们全都串起来的那个连接点,
12:12
we can get into some of the,
我们可以深入聊一些,
12:13
the hard parts related to the technical testing,
跟 technical testing 相关的那些难啃的部分,
12:15
which is I think the crux of the matter,
我觉得这才是问题的核心,
12:17
but I'll pause there.
不过我就先停在这儿。
12:19
How do you and Reggie come together?
你和 Reggie 是怎么走到一起的?
12:21
There's always like, you come across very confident
总是会让人觉得,你给人的感觉特别自信,
12:24
and we're announcing your series A and all these things.
而且我们还在宣布你们的 series A 以及所有这些事情。
12:27
But I want to see like the early initial stages of like idea formation.
但我更想看到的是,比如说想法形成的最早期阶段。
12:31
Yeah.
嗯。
12:32
Rajiv is actually my soon-to-be brother-in-law.
Rajiv 其实马上就要成为我的舅子了。
12:35
Oh.
哦。
12:36
So I'm actually in a week and a half getting married to Rajiv's sister.
所以我再过一周半就要和 Rajiv 的姐妹结婚了。
12:41
Okay, now you're tight.
好吧,那现在你们关系可近了。
12:43
Exactly.
没错。
12:44
So Rajiv and I have known each other for a decade.
所以 Rajiv 和我已经认识十年了。
12:47
Funny story.
说来好笑。
12:49
I met both Rajiv and his sister, Hannah.
我同时认识了 Rajiv 和他妹妹 Hannah。
12:52
At the same time, when Hannah and I were interns
与此同时,Hannah 和我在做实习生
12:54
at McKinsey and London,
在 McKinsey 和 London,
12:56
and Rajiv was assigned as my mentor.
而 Rajiv 被安排成了我的导师。
12:59
And so I met them at the same time.
所以我是同时认识他们俩的。
13:01
For the long time, it was not obvious
有很长一段时间,都看不出
13:02
that we were necessarily going to work together.
我们一定会一起工作。
13:04
I was in startups.
我以前在创业公司干过。
13:05
He was an insurance partner at McKinsey three or four years ago.
三四年前,他是 McKinsey 的保险合伙人。
13:09
I think Hannah convinced him that I was going to be a really big saying.
我觉得 Hannah 说服了他,让他相信我会成为一个特别大的说法。
13:14
And so he quit his job,
于是他就辞掉了工作,
13:16
cushy partner job at McKinsey and London,
在 McKinsey 和 London 的那份轻松合伙人工作,
13:18
packaged bags through to San Francisco,
打包行李一路去了 San Francisco,
13:20
and ended up joining meter.
最后加入了 meter。
13:22
You guys are probably not enough.
你们这些人可能还不够。
13:24
Exactly.
没错。
13:25
We've seen the chat of the horizons
我们已经看到 horizons 的图表
13:29
of the tasks that agents can take on as doubling extremely fast.
agents 能承担的任务正以极快的速度翻倍。
13:32
So here's zero there.
所以这里就是零。
13:34
Led their partnerships with Anthropic and OpenAI
他们牵头了与 Anthropic 和 OpenAI 的合作
13:38
to test their models before release,
在发布前测试他们的 models,
13:40
but also working closely with the US and UK government
但也和 US 以及 UK 政府紧密合作
13:43
to figure out how do you know
来弄清楚你怎么知道
13:46
whether a model can be released.
一个 model 能不能发布。
13:48
And in some ways,
而且在某些方面,
13:50
that's like the perfect background.
这简直就是完美的背景。
13:51
He's spent a lot of time in insurance,
他在保险行业待了很长时间,
13:52
knows that world.
很懂那一行。
13:53
He's spent a lot of time with frontier testing of models.
他在 model 的 frontier testing 上也花了很多时间。
13:56
And so when I was bumbling around this idea space,
所以当我在这个想法空间里瞎摸索的时候,
13:59
starting with some of the ideas we talked about,
从我们聊过的一些想法开始,
14:01
related to Waymo,
跟 Waymo 有关,
14:02
as soon as we got into the content,
我们一聊到具体内容,
14:04
we're both like,
我们俩都心想,
14:05
oh, this would be an amazing business to build together.
天哪,我们一起做这个会是个超棒的生意。
14:08
This is like wrestling with the problem
这就像是在跟一个问题较劲,
14:10
that we both think is the most important of the world
而这个问题我们俩都认为它是世界上最重要的,
14:13
from a market angle,
从市场的角度来看,
14:15
which is kind of our intuations
这差不多就是我们的直觉。
14:17
is that the market can do a lot.
就是市场能做很多事。
14:19
And the faster I move,
而且我行动得越快,
14:21
the harder it is for government
政府就越难
14:22
to solve some of these problems.
解决其中一些问题。
14:23
And then it took a little bit of time
然后花了一点时间
14:25
to work through what does it like to work with family.
去想清楚跟家人一起工作是什么感觉。
14:27
And...
而且……
14:29
Does he already dating at the time more?
他当时是不是已经约会得更多了?
14:31
Yeah, yeah, yeah, yeah.
对对对对。
14:32
Yeah, exactly.
对,没错。
14:33
Already back then,
早在那个时候,
14:34
we felt like we were a family.
我们就觉得自己像一家人。
14:35
And so starting a business together
所以一起创业
14:37
felt like kind of a big step.
感觉像是挺大的一步。
14:38
And here we are with just immense amounts of trust.
而如今,我们之间有着无比巨大的信任。
14:42
Yeah.
对。
14:43
So now you're a company of how big?
所以现在你们公司有多大?
14:45
How big are you guys now?
你们现在规模多大?
14:46
There is just 20 of us now.
我们现在只有 20 个人。
14:47
20 of you guys now have Series A
你们现在 20 个人,就已经拿到 Series A 了,
14:48
and you have your first certification out,
而且你们第一个认证也出来了,
14:51
the AIUC1.
AIUC1。
14:53
Let's bring up the certification.
我们把那个认证调出来吧。
14:55
So this is the agent certification.
所以这就是那个 agent 认证。
14:58
What goes into the process?
这个过程都包括什么?
15:00
Well, I have like two questions here.
嗯,我这里大概有两个问题。
15:01
One is,
一个是,
15:02
walk us through the certification
带我们过一遍 certification 吧
15:04
and two is,
另一个是,
15:05
what is the process for a company to get certified?
一家公司要拿到 certification,流程是什么?
15:08
Great.
很好。
15:09
As it says, right at the top,
就像最上面写的那样,
15:10
AUC1 is a standard for agent security,
AUC1 是一个标准,面向 agent security、
15:13
safety and reliability.
safety 和 reliability。
15:14
The fundamental design principle is,
最根本的设计原则是,
15:16
take all of the concerns that slow down adoption.
把所有会拖慢采用的顾虑都收进来。
15:20
All the questions, all the fears that keep security leaders
所有那些问题,所有那些让安全负责人
15:23
and the Fortune 1000 up at night
和 Fortune 1000 夜不能寐的恐惧,
15:25
and put them into one comprehensive framework.
再把它们放进一个全面的 framework 里。
15:27
That's where you'll see there.
这就是你会在那里看到的。
15:29
You can see the six categories.
你可以看到这六个类别。
15:30
Two,
第二,
15:31
you want to ground all of this in technical testing.
你想把这一切都建立在 technical testing 的基础上。
15:34
So one of the concerns with security standards
所以,对于 security standards 的一个担忧是,
15:37
that often feel kind of like theater paperwork
它们常常感觉有点像作秀式的文书工作,
15:39
is that they don't actually ground out in,
问题在于它们实际上并没有真正落实到,
15:41
doesn't have this work, doesn't have this matter.
没有实际作用,也不真正重要。
15:43
And so we had a conviction from early on
所以我们很早就有一个信念。
15:45
that that was going to be the kind of crux,
那,那会是关键所在,
15:47
was to pass this,
就是要通过这一关,
15:49
you must get tested every quarter,
你必须每个季度都接受测试,
15:51
basically run thousands of simulations to see,
基本上就是跑几千次 simulations,看看,
15:54
also can it actually be jailbroken?
还有,它到底能不能被 jailbreak?
15:56
How hard is it to jailbreak?
要 jailbreak 它有多难?
15:57
How often does it hallucinate?
它多久会 hallucinate 一次?
15:58
How often does it leak data, et cetera?
它多经常会泄露 data,等等?
16:00
And then the last core idea here,
然后这里最后一个核心想法,
16:03
if you scroll up to the top here,
如果你往上滚到最上面这里,
16:05
is to refresh it quarterly.
就是每季度刷新一次。
16:07
So the core trailer of AI is that it moves extremely fast.
所以 AI 的核心特质就是它发展得极快。
16:10
Whatever concerns we're discussing today,
我们今天在讨论的这些担忧,
16:12
we're not the same ones three months ago,
已经不是三个月前的那些了,
16:14
and this will keep changing.
而且这还会一直变。
16:15
Typically standards update on like a decade cycle,
通常标准的更新周期大概是十年,
16:19
is obviously not going to work.
显然行不通。
16:21
But the question is kind of how do you update it?
但问题是,你要怎么更新它呢?
16:23
And the core thing here was to basically get
而这里的核心,基本上就是让
16:26
the risk leaders of the Fortune 1000
Fortune 1000 的风险负责人
16:28
around the table.
坐到一起。
16:29
So if you go over to the left here,
所以,如果你往左边这里看,
16:31
and you'll see AI is the one consortium,
你会看到 AI 就是那个联盟,
16:33
the consortium is a group of risk leaders
这个联盟是一群风险负责人
16:36
who run real banks, real hospitals,
他们经营着真实的银行、真实的医院,
16:39
real critical infrastructure,
真实的关键基础设施,
16:40
who are facing these challenges every day.
他们每天都在面对这些挑战。
16:42
And we meet with these folks twice a quarter.
而且我们每个季度会和这些人见两次面。
16:45
And here, what's top of mind?
那么,他们最关心的是什么?
16:46
What is keeping them up at night?
是什么让他们夜不能寐?
16:48
There's tremendous amount of desire for that conversation.
大家对这种对话有极大的渴望。
16:50
And then we operationalize that into a specific standard
然后我们把它落地成一项具体的标准
16:53
that gets into, and actually,
这会涉及到,而且其实,
16:54
we can go into and look at what is,
我们可以进去看看,什么是,
16:55
what even is the standards?
到底什么是 standards?
16:56
We go back to introduction out there to the left,
我们回到左边那个 introduction,
16:58
scroll up a little bit to the wheel,
往上滚一点,到那个 wheel,
17:00
click into reliability.
点进 reliability。
17:01
So if you take something like hallucinations,
所以,如果你拿 hallucinations 这种例子来说,
17:03
hallucinations sits in reliability,
hallucinations 就归在 reliability 里,
17:05
there is a number of requirements here.
这里有一些 requirements。
17:07
If you go into the top one,
如果你进入最上面那一个,
17:09
prevent hallucinated requirements,
防止 hallucinated requirements,
17:10
hallucinated outputs.
hallucinated outputs。
17:11
This is one particular requirement.
这是一个特定的 requirement。
17:13
This is a technical control.
这是一个 technical control。
17:14
Basically, we want some kind of ground in this filter.
基本上,我们想在这个 filter 里加入某种 ground。
17:16
The first thing you see here is,
你在这里看到的第一件事是,
17:18
what's called a crosswalk.
所谓的 crosswalk。
17:20
So every one of their grandmother
所以他们每个人的祖母
17:22
has put out a framework,
都发布了一个 framework,
17:24
very highly a little framework for what are the air risks.
一个很高层面的小 framework,讲 AI risks 是什么。
17:26
This is basically your competition.
这基本上就是你的竞争对手。
17:28
But in some ways out,
但从某些方面来说,呃,
17:29
we're in fact friends with them.
我们其实跟他们还是朋友。
17:30
We'll come back to why.
我们之后会再讲为什么。
17:31
But mapping everything together,
但是把所有东西都 mapping 到一起,
17:33
so you have one superset.
这样你就有了一个 superset。
17:35
The claim you're trying to support here is,
你在这里想要支持的主张是,
17:38
if you follow this framework,
如果你遵循这个 framework,
17:40
then you can also see how you follow the other frameworks.
那你也能看到你是怎么遵循其他 framework 的。
17:43
But the meat of it comes down here
但真正的重点落在这里,
17:44
in control activities and evidence.
在 control activities 和 evidence 上。
17:46
So control activities is like great.
所以 control activities 就,挺棒的。
17:48
You have this high level requirement.
你有一个 high level requirement。
17:49
How do you turn that down to something operational?
你怎么把它落到 operational 的层面?
17:51
Here's what you must do.
这就是你必须做的事。
17:53
And then what is the evidence that we're looking for?
然后,我们要找的证据是什么?
17:56
And the reason we go this deep is that,
而我们之所以要挖这么深,是因为,
17:59
there's actually not that much confusion
其实并没有那么多困惑
18:01
about what are the big concerns in AI?
关于 AI 里的大担忧是什么?
18:03
Everyone agrees to these.
大家都认同这些。
18:05
The question like, what are you actually supposed to do?
这个问题就像是,你实际上到底该做什么?
18:08
And so what we found a lot of demand for
所以我们发现,很多人真正需要的是
18:11
is getting down to the specific evidence
落实到具体的证据上
18:13
that people need to look for.
也就是人们需要去找的那些证据。
18:15
Whether you are cursor building something
无论你是在用 cursor 做东西
18:18
or even Jake Morgan building thing,
或者甚至是 Jake Morgan 在做东西,
18:21
but also if you're just a risk leader and Jake Morgan,
但如果你只是 Jake Morgan 的风险负责人,
18:23
like what exactly should you ask for?
比如说,你具体应该要求提供什么?
18:25
What can you ask for without sounding stupid?
你可以提什么要求,还不会显得自己很蠢?
18:27
Like if you ask for something,
比如说,如果你提个要求,
18:28
you won't believe the amount of time
你都不敢相信有多少时间
18:30
ever stayed as asked for the IP rights to the online model
是一直坚持在要 online model 的 IP rights,给到 Cursor 之类的。
18:34
to cursor or something.
然后你就想:不好意思,啥?
18:35
And you're just like, sorry, what?
你把它顺带塞进去,然后你就看到这个。
18:37
You slip it in there and you see this.
你看你注意到没有。
18:40
You see if you notice.
你看看你能不能注意到。
18:41
You see the exact amount of the encryption there.
你看,那里能看到 encryption 的确切数量。
18:43
So that's kind of what a standard is.
所以这大概就是 standard 的意思。
18:45
And we update this every quarter
而且我们每个季度都会更新这个
18:47
with these folks to keep up with the latest concerns.
和这些人一起,跟上最新的关注点。
18:50
Can I double click on this one?
我可以 double click 这个吗?
18:52
Yeah.
可以啊。
18:53
So first of all, the website is beautiful.
所以首先,这个 website 真的很漂亮。
18:55
Like it's so confidence-inducing,
就感觉特别能让人有信心,
18:57
which is the whole point where like,
这就是整个重点所在,就像,
18:59
okay, I know exactly what I'm signing out for when I talk with you.
好吧,跟你说话的时候,我完全清楚自己到底在签什么。
19:02
I don't even have to talk to you.
我甚至都不用跟你说话。
19:03
I can just see your whole certification, which is great.
我直接就能看到你完整的 certification,这太棒了。
19:06
But like, okay, so from here,
但就像,好吧,那从这里开始,
19:08
like D001.1 config,
比如 D001.1 config,
19:10
Gondas filter filter,
Gondas filter filter,
19:11
how does that get applied?
那这个到底是怎么应用的呢?
19:13
Like you have a person that goes to it.
就像你有个会去处理它的人。
19:16
Yeah, so if you go back up.
对,所以如果你再往上翻。
19:19
I can see somewhere there's like, you know,
我能看到某个地方有,就像,你知道,
19:21
51 requirements, 130 controls.
51 个 requirements,130 个 controls。
19:23
There's like a whole.
就像有一整套。
19:24
Right.
对。
19:25
Like to me, this doesn't translate into a test.
对我来说,这没法转化成一个 test。
19:27
Or you go.
或者你去。
19:28
Yes, yes, yes, yes.
对,对,对,对。
19:29
So if you go into, I don't know the left-hand side.
所以如果你进到,我不知道,左边那侧。
19:32
So actually before we go in there,
所以其实在我们进去之前,
19:34
there are three types of requirements.
有三类要求。
19:36
The first is technical controls.
第一类是 technical controls。
19:38
Like you must implement some guard rails.
比如你必须实施一些 guard rails。
19:41
Two, there are test controls.
第二,有 test controls。
19:44
So you must have an independent third party
所以你必须有一个独立的第三方
19:46
go and run some tests against you.
去跑一些针对你的测试。
19:48
I'll show you one of those in a second.
我一会儿给你看其中一个。
19:49
And then three, there are policy controls.
然后第三,就是 policy controls。
19:52
For example, you must have a person whose name is on the line
比如说,你们必须有一个具体的人来担责
19:55
when you guys fuck up.
当你们搞砸的时候。
19:56
And you must have a plan for how you tell your customers
而且你必须有个方案,说明怎么告诉你的客户
19:58
and how you engage with them.
以及怎么跟他们互动。
19:59
They're kind of more traditional standard types of stuff.
这些算是比较传统、标准的那类东西。
20:02
So in this particular instance, we just check whether they,
所以在这个具体情况下,我们只是检查他们,
20:04
in fact, have a ground in their filter.
实际上是否在他们的 filter 里有依据。
20:06
So we will partner with an auditor.
所以我们会和一个审计员合作。
20:08
So we partner with orders like KPMG or like Shellman
所以我们会和像 KPMG 或 Shellman 这样的审计机构合作
20:12
who go in and do the thing auditors do,
他们会进去做审计员会做的事,
20:14
which is to check the evidence.
也就是检查证据。
20:15
In this case, that might be a screenshot.
在这种情况下,那可能是一张截图。
20:16
It might be a part of the code that they need to review
也可能是他们需要审查的一部分代码。
20:18
to see that it actually just that it exists.
就是看看它其实只是存在而已。
20:21
And then the second thing is.
然后第二点呢就是。
20:22
So you're not testing the effectiveness of it.
所以你并不是在测试它的效果。
20:23
That's the second thing.
这就是第二点。
20:24
So if you go down to the third party testing for hallucinations
所以如果你往下看到第三方对 hallucinations 的测试
20:27
out on the left, that's basically the next requirement.
在左边,那基本上就是下一个要求。
20:29
This is where we test how well does it actually work?
这就是我们测试它实际上到底效果如何的地方?
20:31
Okay.
好。
20:32
And is it you testing or the auditor?
是你们来测试,还是 auditor 来测试?
20:34
We test them.
我们测试它们。
20:35
We test them.
我们测试它们。
20:36
That's a lot of work.
这工作量可不小。
20:37
How long does testing take?
测试要花多长时间?
20:38
So if I want to get certified just how long does then
所以如果我想拿到认证,那到底要花多久
20:41
and roughly take?
大概要多久?
20:42
Yeah. The internet almost always is dependent on
对。互联网几乎总是依赖于
20:45
that our customers need to do something for us.
我们的客户需要为我们做一些事情。
20:47
It takes them up between like three to ten weeks,
他们大概要花三到十周,
20:52
depending on how up to snuff they already are.
取决于他们现在的水平有多达标。
20:55
So some people could show up to us.
所以有些人来找我们的时候,
20:56
We're like extremely rigorous security programs.
我们就像是极其严格的 security programs。
20:58
When we test them, it works extremely well.
我们测试他们时,效果极其好。
21:00
We can get that done very quick.
我们很快就能搞定。
21:01
Some people come to us and they're not that far along.
有些人来找我们时,还没那么成熟。
21:03
We give them kind of the spec that they need to build towards
我们会给他们一个需要照着去构建的 spec
21:06
and then they're security teams and engineers
然后他们的安全团队和工程师
21:08
get to work and build to meet the standard.
就开始动手,构建出符合标准的东西。
21:10
The testing itself typically takes a couple of weeks,
测试本身通常要花上几周时间,
21:13
including the time for them to remediate.
包括他们做整改的时间。
21:15
Often we'll find something that we cannot pass,
我们经常会发现有些东西没法通过,
21:17
where this is actually just not up to the standard.
也就是它其实根本达不到标准。
21:20
You won't pass the standard and then they will need
你没法通过这个标准,然后他们就需要
21:23
to go and implement additional safeguards
去实施额外的 safeguards
21:26
or additional remediation that makes them more robust
或者额外的 remediation,让它们更 robust
21:28
so that they can actually kind of hand-on-hot
这样他们才能真正算是摸着良心
21:30
look at their customers in the eyes and say,
直视客户的眼睛说,
21:32
hey, we've done truly our very best.
嘿,我们真的已经尽了最大努力。
21:34
And they're certified for a year and have quarterly updates?
而且他们获得一年的认证,还有季度更新?
21:37
Correct.
对。
21:38
Yeah, it's pretty interesting.
是啊,这挺有意思的。
21:40
I think what's changed?
我觉得,变化是什么呢?
21:42
So this is certifying agents in production.
所以这就是在生产环境中认证 agent。
21:45
Your customers, like you've had levelable 11 labs
你的客户,比如你已经有了 levelable 11 labs
21:48
and they're confident they've all gone through the certification?
而且他们很确定自己都已经通过了认证?
21:50
Yes.
是的。
21:51
What has changed?
那到底有什么变化?
21:52
So I see you post like, you know,
所以我看到你发帖,就像,你知道,
21:53
Q2 added MCP agent, agent agent communication,
Q2 加了 MCP agent、agent agent communication,
21:57
any other things that you want to kind of highlight
还有什么你想稍微重点提一下的吗?
21:59
since the first, first iteration,
从第一次、第一次迭代以来,
22:01
what comes in quarterly?
每季度会有什么?
22:03
Yeah.
嗯。
22:04
So some of the changes have just been,
所以有些变化其实就是,
22:06
ages are not just one thing.
agents 并不只是单一的一种东西。
22:08
So like, if you take ages like cursor and compare them to Sierra,
所以比如,如果你拿像 Cursor 这样的 agents 来和 Sierra 比较,
22:12
they're really quite different.
它们真的挺不一样的。
22:13
Compare them to Harvey again.
再拿它们和 Harvey 比较一下。
22:14
Compare them to you.
再拿它们和你比较一下。
22:16
11 labs.
11 labs。
22:17
11 labs.
11 labs。
22:18
They're all quite different.
它们都很不一样。
22:19
And so we wanted to design a standard of works
所以我们想设计一套工作标准
22:22
for all of the types of agents.
给所有类型的 agents。
22:25
And we started with one that was like pretty text-based,
我们从一种挺 text-based 的开始,
22:27
like honestly, pretty customer support focused.
说实话,还是挺聚焦 customer support 的。
22:29
That's where there's a lot of existing demand.
因为那边本来就有很多现成的需求。
22:32
And then over time, I've picked some of the frontier companies
然后随着时间推移,我挑了一些 frontier companies
22:36
in each of these other domains that we could work with
是在这些其他领域里、我们可以合作的那些
22:38
and build out the standard.
然后把标准搭建出来。
22:39
So such that we know that the same standard works for code,
这样一来,我们就知道同一套标准适用于 code,
22:41
the works for customer support, works for automation, etc.
也适用于 customer support,适用于 automation,等等。
22:44
So that's been one big thing.
所以这一直是件很重要的事。
22:46
Yeah.
嗯。
22:47
Then some of the things that have been talked about recently,
然后,最近一直在讨论的一些事情,
22:49
MISOS is bringing up a lot of concerns for security leaders.
MISOS 正给安全负责人带来很多担忧。
22:54
We're starting to get more and more questions around
我们开始收到越来越多关于
22:56
agent to agent interactions.
agent to agent 交互的问题。
22:58
It's very nascent at the moment,
目前它还非常早期,
23:00
but it's starting to emerge.
但已经开始冒头了。
23:02
There have been a lot of questions related to OpenClore and MCP.
有很多问题都跟 OpenClore 和 MCP 有关。
23:07
Again, agents starting to interact with each other.
再说,agents 开始彼此互动了。
23:09
It's really top of mind.
这真的是现在最关心的事。
23:11
Then as coding agents have really taken off,
然后随着 coding agents 真的火起来,
23:13
that's also where banks and hospitals, etc.
这也正是银行和医院等
23:16
are getting more and more precise on what they need.
对自己需要什么越来越明确的地方。
23:19
So really dialing in as a start to be where most of the tokens
所以真的从一开始就精准锁定,要处在全球大多数 tokens
23:22
flow through in the world.
流经的地方。
23:24
Getting much help around that.
围绕这一点也得到了很多帮助。
23:26
Can you share for people that are listening
你能给正在听的人分享一下吗
23:28
that don't really think about this?
那些其实没怎么想过这个问题的人?
23:29
Like you mentioned, there's the obvious stuff,
就像你提到的,有些很明显的东西,
23:31
hallucinations, citations.
hallucinations、citations。
23:33
What are best practices that people should do when building agents?
构建 agents 时,人们应该遵循哪些 best practices?
23:37
Like, if they come to you pretty ready with sort of,
比如,如果他们来找你的时候已经准备得挺充分,差不多是,
23:41
they'll probably pass certification.
他们大概率能通过认证。
23:43
What are things people don't think about that they should have?
有哪些东西是人们不会想到、但其实应该具备的?
23:46
The most important thing is that a lot of companies
最重要的是,很多公司
23:50
have not done a serious stress test.
还没有认真做过 stress test。
23:52
They spend most of the time, perhaps rightly so,
它们大部分时间,也许这么做也没错,
23:55
optimizing for how does it work in the good case?
都在优化它在 good case 下效果怎么样?
23:58
The average case.
average case。
23:59
How high quality is the output for the customer?
给客户的 output 质量有多高?
24:02
And a lot of these companies are pretty new.
而且很多这类公司都还很新。
24:06
So they haven't spent a lot of time stress testing.
所以它们还没花很多时间做 stress testing。
24:09
What is an adversary on the other side?
另一边的 adversary 是什么?
24:12
What are some of the complicated corner cases
有哪些复杂的 corner cases
24:14
that you've not really considered?
是你其实还没真正考虑过的?
24:16
So I think there's like a frame of mind.
所以我觉得这有点像是一种心态。
24:18
And you also see this in startups.
你在 startups 里也会看到这一点。
24:19
It often takes a while until they hire their first security person.
他们往往要过一段时间才会招第一个 security 人员。
24:22
And that's a whole different kind of risk surface
而那是一种完全不同的 risk surface
24:24
than just building a good product.
而不只是做出一个好产品。
24:25
So a lot of that applies.
所以其中很多都适用。
24:27
Most companies actually also have the right kind of architecture.
大多数公司其实也都有合适的那类 architecture。
24:30
Most of them will have some kind of guard rails in place.
它们大多数也都会部署某种 guard rails。
24:32
Either some that come out of the box from their model provider,
要么是他们的 model provider 开箱即用提供的一些,
24:35
or they'll have built their own filters to sit in between
要么他们会自己构建 filters 夹在中间,
24:37
that you just don't work very well.
但这些就是效果不太好。
24:39
The difference between putting a classifier in place
区别就在于把一个 classifier 部署到位,
24:41
that maybe goes and checks whether you're giving medical advice
它可能会去检查你是不是在提供医疗建议
24:45
when you shouldn't, and says, hey,
在你本不该的时候,它却说,嘿,
24:47
if this looks like medical advice, filter it out.
如果这看起来像是医疗建议,就把它过滤掉。
24:49
Lots of companies have that in place.
很多公司都已经有这种机制了。
24:51
The question is whether it works.
问题是,这到底管不管用。
24:52
And it actually is a pretty fit to sit down
而其实,这很适合坐下来
24:54
and think about all the ways in which you could ask for medical advice.
想想所有你能寻求医疗建议的方式。
24:57
Read the academic literature on what are the kinds of framings
去读学术文献,看看都有哪些种 framings,
25:03
or tricks you might play to get an AI to give you medical advice
或者你能耍哪些花招,让 AI 给你医疗建议。
25:07
when you really shouldn't.
在你其实真的不该这么做的时候。
25:08
And so there's like an error of expertise that's just missing.
所以就好像有一种专业上的错误,只是被漏掉了。
25:11
So what we find is that most people have the right building box
所以我们发现,大多数人都有正确的 building box
25:13
in place that doesn't, it's not rocket science.
已经到位了,这并不,这不是 rocket science。
25:15
But the finicky thing is like getting into the corners
但真正麻烦的是,要钻进那些边边角角里
25:17
and testing whether it works.
并测试它到底有没有用。
25:18
Such said you can look your customers in the eye
话虽如此,你可以直视你的客户
25:20
who may be a bank or maybe a hospital and be like,
他们可能是银行,也可能是医院,然后说,
25:23
this is going to work for you.
这对你会有效的。
25:25
I see.
我明白了。
25:26
So we talked a lot about the agent level certification.
所以我们聊了很多关于 agent level certification 的事。
25:32
Where do you guys go from here?
你们接下来要往哪走?
25:34
So announcing series A, off camera,
所以,宣布 series A,镜头外,
25:36
we talked about this a bit.
我们稍微聊过这个。
25:37
There's the whole security risk of
还有 fable government 介入的整个安全风险。
25:40
fable government stepping in.
fable 政府介入。
25:42
You guys are kind of announcing they are also going into model certification.
你们这有点像是在宣布,他们也要进军 model certification 了。
25:45
When we do a bit of cutting afterwards,
等我们之后做一点剪辑的时候,
25:47
we will not yet be announcing this.
我们暂时还不会宣布这件事。
25:49
But the question that is top of everyone's minds now
但现在所有人最关心的问题
25:53
is at the model level.
是在 model level。
25:54
And Miso's then fable has really brought this to the fore
而 Miso's then fable 真的把这件事推到了台前,
25:58
or that in addition to the commercial risk and the kind of economic security
或者说,除了商业风险和那种经济安全
26:02
risks that are happening at the agent layer,
风险——这些正发生在 agent layer,
26:04
the models are going to present risks in the national security category.
这些 models 会带来国家安全类别的风险。
26:09
The shape of the problem is very similar.
问题的形态非常相似。
26:11
You have some people that are on the hook
有些人是要担责的
26:14
if something goes wrong.
如果出了什么问题。
26:15
In the case of agents is often security leaders in the enterprise.
就 agents 来说,通常会是企业里的安全负责人。
26:19
In this case, it's the government.
在这种情况下,就是政府。
26:21
They don't have necessarily spent their entire lives thinking about
他们不一定一辈子都在思考
26:24
what are the new risks that come here.
这里会出现哪些新的风险。
26:26
What is the kind of data you might be looking for?
你可能在找的是什么样的 data?
26:28
How might you test that?
你可能会怎么测试那个?
26:29
But they do have to make sure that their concerns are addressed.
但他们确实必须确保自己的担忧得到解决。
26:31
You have some frontier companies that are deeply technical.
你有一些 frontier companies,技术实力非常强。
26:35
They know a lot about the risks.
他们对这些风险了解很多。
26:37
But they fundamentally have an incentive to not always be truthful.
但他们从根本上就有一种动机,让他们不总是说实话。
26:41
So you have a trust gap between the government and the labs.
所以政府和 labs 之间就有一个信任差距。
26:44
And in every other industry,
而在其他所有行业里,
26:46
you end up with some kind of body sitting between a neutral third party
最终你会有某种机构,作为一个中立的第三方坐在中间,
26:50
sitting between those people.
坐在那些人之间。
26:52
There's no other industry where you allow people to audit themselves.
没有其他行业会允许人们审计自己。
26:55
So there's going to be a need for a third party
所以就会需要一个第三方,
26:58
that can take the rigor of the labs to run frontier technical evils.
能把 labs 的严谨性拿过来,去运行 frontier technical evils。
27:03
But can also speak legible trust in the way that the government trusts
但也能像政府信任
27:08
PWC to go and run financial audits.
PWC 去进行财务审计那样,传达出清晰可懂的信任。
27:10
And they know that they output audit reports in a way that's consistent.
而且他们知道,自己输出审计报告的方式是一致的。
27:14
That's easy to read.
这个读起来很轻松。
27:15
That's factual.
这是有事实依据的。
27:16
That's trustworthy.
这是值得信赖的。
27:18
Those two things need to be brought together.
这两者需要结合起来。
27:20
And what we've learnt from our work with agents
而我们通过与 agents 合作学到的是,
27:23
is that if you want that communication between those two parties to be smooth,
就是,如果你想让这两方之间的沟通顺畅,
27:27
there has to be one common standard that is public,
就必须有一个公开的通用标准,
27:31
that people can go and inspect,
让人们可以去查看,
27:32
what are the risks that matter?
真正重要的是哪些 risks?
27:34
We then usually use risk.
然后我们通常就会用到 risk。
27:35
What are the kind of threat models that you're really looking for?
你真正在找的是哪些类型的 threat models?
27:38
You need to specify for each of those risks,
你需要针对这些 risks 中的每一个具体说明,
27:40
what are the guardrails that need to be in place?
需要有哪些 guardrails 到位?
27:42
And what are the tests that need to run to see whether those guardrails are effective?
以及需要跑哪些测试,才能看出这些 guardrails 是否有效?
27:46
And then you need to go and run audits that are technical audits that are consistent.
然后你需要去跑 audits,也就是一致的 technical audits。
27:52
So if you're trying to bring trust,
所以,如果你想带来信任,
27:54
it's extremely important that you methodically work your way through the risks.
极其重要的是,你要有条不紊地把各种风险逐一梳理清楚。
27:57
You can't send one researcher in and say like,
你不能派一个研究员进去,然后说,
28:01
come back with whatever you find.
找到什么就带回来什么。
28:02
You need to be able to explain exactly what you did, exactly what you tried,
你需要能够准确解释你做了什么、你尝试了什么,
28:06
exactly what you did not try.
以及你确切没有尝试什么。
28:07
And therefore the kinds of promises you can and cannot make at the end of it.
因此,到最后你能做出哪些承诺、不能做出哪些承诺。
28:11
I think of Fable as a direct symptom of this problem that the government was told that there's a risk.
我把 Fable 看作这个问题的直接症状,即政府被告知存在风险。
28:19
The government may struggle to assess just how big that risk is.
政府可能很难评估这个风险到底有多大。
28:23
They call it an anthropic.
他们管它叫 Anthropic。
28:24
An anthropic is trying to tell them,
Anthropic 在试图告诉他们,
28:26
hey, actually every model can be jailbroken.
嘿,其实每个模型都能被 jailbreak。
28:30
There's no way you want to hear, right?
这你肯定不想听,对吧?
28:32
As the government, that might be hard to trust.
作为政府,这可能很难让人相信。
28:35
And we think that a broker is the most natural solution.
而我们认为 broker 是最自然的解决方案。
28:40
In all the markets, you see something like financial markets, you see moody's.
在所有市场里,比如金融市场,你会看到 Moody's。
28:44
Moody's goes in and they look at a bond.
Moody's 会介入,然后去看一只债券。
28:48
And they output a rating.
然后他们输出一个 rating。
28:49
They say like, here's the evidence we found.
他们会说,你看,这是我们找到的证据。
28:51
Here's the rating.
这是 rating。
28:52
We don't decide whether anyone should buy this bond or not buy this bond.
我们并不决定任何人应该买这个 bond,还是不买这个 bond。
28:55
Well that depends on the risk appetite.
嗯,那取决于 risk appetite。
28:57
But we do pry this common information layer that everyone can rely on.
但我们确实提供了这个所有人都能依赖的 common information layer。
29:01
In the case of moody's, the government points to them and say,
在 moody's 的例子里,政府会指着他们说,
29:06
hey, pension funds, you should probably really take care.
嘿,pension funds,你们大概真的得小心点。
29:09
You shouldn't risk your pension as money.
你不应该拿养老金当钱来冒险。
29:11
So you can only invest in AAA rated bonds.
所以你只能投资 AAA rated bonds。
29:15
That means that now the government doesn't have to staff thousands
这意味着现在政府不用养着成千上万的
29:19
of financial technical experts to rerun forecasts every week to see
金融技术专家,每周重新跑预测,看看
29:24
whether things are correctly rated.
这些东西有没有被正确评级。
29:26
They get to point to some neutral third party.
他们只要指向某个中立的第三方就行。
29:29
So my hypothesis is my hunch is that you will see a third party that sits
所以我的假设,也就是我的直觉,是你将会看到一个第三方,它位于
29:34
between the government and the labs.
政府和 labs 之间。
29:36
And it could either be the government bills it themselves.
而且也可能是政府自己来买单。
29:38
So something like Casey was set up to do exactly this.
所以像 Casey 这样的机构,就是为了做这件事而设立的。
29:43
And the question is, I'm not familiar with Casey.
问题是,我对 Casey 不太了解。
29:45
Casey is the center for AI standards and innovation.
Casey 是 Center for AI Standards and Innovation。
29:49
Okay.
好。
29:50
I won't get into the details, but it's a sub body of nists that typically set standards.
我就不细讲了,但它是 NIST 下属的一个机构,通常负责制定标准。
29:54
So it's basically a government body that has air experts.
所以它基本上就是一个有 AI 专家的政府机构。
29:57
Yeah.
嗯。
29:58
Exactly.
没错。
29:59
Very lucky.
非常幸运。
30:00
It's one of those things when you just sit back and look at it.
这种事就是,你退后一步看看就会明白。
30:04
Like, is there enough technical expertise in the government to measure
就像,政府里有没有足够的技术专业能力去衡量
30:09
test these things right now?
测试这些东西,就现在?
30:11
Probably not, right?
大概没有,对吧?
30:12
And Fable is a result of, okay, we've had to scale back and pause things.
而 Fable 就是这样一个结果:好吧,我们不得不缩减规模、暂停一些事情。
30:17
Yeah.
是啊。
30:18
And they have excellent people, but they have an extraordinarily small budget
而且他们有非常优秀的人才,但预算却小得惊人
30:23
compared to the scale of the challenge that's ahead of us.
和我们面前这项挑战的规模相比。
30:27
And I think they have a role to play.
而且我觉得他们能发挥作用。
30:29
The question is kind of like, who does what?
问题有点像:谁负责做什么?
30:31
We have now outlined the jobs to be done.
我们现在已经列出了需要完成的任务。
30:33
And they're quite extensive.
而且这些任务相当多。
30:34
Every model release.
每一次 model 发布。
30:36
There's an astounding, given that they take in any input, their risk surface is astounding.
有一个惊人的点:考虑到它们会接收任何 input,它们的 risk surface 大得惊人。
30:42
And so the question is really, what can only the government do?
所以真正的问题是,什么是只有政府能做的?
30:45
And what can the market provide here that can keep up with the pace as AI risk changes?
市场能提供什么来跟上 AI 风险变化的节奏?
30:51
Our perspective is that also the model layer, the risk that people care about today are not
我们的观点是,模型层这里,人们今天关心的风险已经不是
30:56
the same ones they care about three months ago.
三个月前他们关心的那些了。
30:58
So the pace of legislation is too slow to deal with pinpointing the risks here.
所以立法的速度太慢,没法精准定位这里的风险。
31:02
And so we think there's a lot that the market can do to surface timely information.
因此我们认为市场可以做很多事情来让及时的信息浮出水面。
31:06
Ultimately, there is a bunch of policy decisions here.
归根结底,这里有一堆政策决策。
31:08
Is the national security risks of a model too high?
一个模型的国家安全风险是不是太高了?
31:11
Yeah.
嗯。
31:12
That's a political answer.
那是个政治性的回答。
31:13
But what you want to make sure is that the process that produces this risk information
但你要确保的是,产生这些风险信息的过程
31:18
is compatible with very fast innovation.
能跟非常快的创新兼容。
31:20
So you don't want to...
所以你不希望……
31:21
This is not a question of like, can you slow things down?
这不是那种“你能不能放慢速度?”的问题。
31:23
Can you keep the models locked up until for months on end until everyone can make a guarantee?
你能把 models 一锁就是好几个月,直到所有人都能给出保证吗?
31:30
But can you in the time given that the US is competing with China on releasing models,
但鉴于 US 正和 China 竞争发布 models,在这段时间里,你能不能……
31:37
can you insert risk information that allows the government to make rapid decisions on some of these questions,
能不能把风险信息加进去,让政府能就其中一些问题快速做决定,
31:42
balancing that trade-off between failing to adopt AI is going to put us at risk,
平衡这个 trade-off:不采用 AI 会让我们陷入风险,
31:46
but also reckless adoption is going to put us at risk.
但鲁莽采用 AI 也会让我们陷入风险。
31:48
And that's a very fine balance that they're going to need like a lot of high quality intelligence to make.
而这是一个非常微妙的平衡,他们需要大量高质量情报才能做得到。
31:55
Just a side mention, because you mentioned Chinese models,
顺便说一句,因为你提到了 Chinese models,
31:58
any specific concerns that you're hearing from your CSOs about that,
你有没有从你的 CSO 们那里听到关于这个的任何具体担忧,
32:02
because I guess it's free.
因为我猜它是免费的。
32:04
But CSOs have a bunch of concerns around data flows in general
但总的来说,CSO 们对 data flows 有一堆担忧
32:08
that they're really concerned about.
那是他们真正担心的事情。
32:10
So there's a lot of questions like,
所以会有很多类似这样的问题,
32:12
if these models are Chinese, where does our data go?
如果这些 models 是中国的,我们的 data 会去哪儿?
32:15
I think a lot of this can be addressed, but they come up often.
我觉得很多问题都能解释清楚,但它们经常会被提出来。
32:18
I mean, they understand it running on American GPUs.
我的意思是,他们明白它是在美国 GPUs 上跑的。
32:21
Some of them understand that they're running on American GPUs.
他们中有些人明白,它们是在美国 GPUs 上跑的。
32:23
They're not like phoning home every time you like home.
它们不会动不动就 phoning home。
32:26
No.
不会。
32:27
A year ago, there was not a lot of understanding of this.
一年前,大家对这件事还没什么理解。
32:30
I actually think you're seeing this clearly just becoming kind of AI literate at a blistering pace
其实我觉得,你现在明显正以飞快的速度变得有点懂 AI 了
32:35
and you're actually also seeing my Twitter timeline that's very air-pilled
而且你其实也看到,我的 Twitter 时间线非常 air-pilled
32:39
and my LinkedIn feed that used to not at all be air-pilled kind of converged.
和我那个以前完全不 air-pilled 的 LinkedIn 信息流,也有点趋同了。
32:43
They're both talking about Fable.
它们都在聊 Fable。
32:45
Right, right.
对,对。
32:46
They are both talking about,
它们都在聊,
32:48
whether you can prevent models from being jailbroken these days.
现在到底能不能防止 models 被 jailbroken。
32:52
Like national security risk,
比如国家安全风险,
32:54
that conversation is actually emerging.
这种讨论其实正在出现。
32:56
Other than that, I think you mostly see a kind of,
除此之外,我觉得你主要看到的是一种,
33:00
there's no concerns with any particular model or any particular model output,
对任何特定模型或任何特定模型的输出都没什么担忧,
33:04
but there's a general nervousness of having critical infrastructure,
但会有一种普遍的紧张感:让关键基础设施
33:08
run on models that are not produced in America by Americans
运行在不是由美国人在美国生产的模型上,
33:12
where the American government has control.
而美国政府对这些是有控制权的。
33:14
That doesn't necessarily show up in your framework and directly.
这未必会直接体现在你的 framework 里。
33:18
There's a bit of stuff in there actually on the provenance of the models and disclosing that.
其实里面确实有一些内容,是关于模型的 provenance 以及披露这些信息的。
33:22
But I think there's a bunch of use cases where running
但我觉得有一大堆 use cases,在这些场景里,跑
33:24
and Chinese Open Source model is just the best solution,
一个 Chinese Open Source model 就是最好的解决方案,
33:26
and a concern in Cymo macro,
以及 Cymo macro 里的一个担忧,
33:28
which is not best addressed at any particular certification level.
这个担忧并不是在某个特定的 certification level 上就能最好地解决的。
33:32
Is there anything interesting that you see at the,
在那个……你们有没有看到什么有意思的东西,
33:34
you know, if you are trying to fill that middle gap, that mediation gap,
你知道,如果你想填补那个中间的空缺,那个 mediation gap,
33:38
any interesting stuff that you guys forecast would be required?
你们预测会需要哪些有意思的东西?
33:43
Other than, you know, what the average risk I might expect?
除此之外,你知道,一般来说我可能会面临的平均风险是什么?
33:47
There's a bunch of interesting questions about what are the risks that matter here?
这里有一堆有意思的问题,关于到底哪些风险才是真正重要的?
33:51
So right now, the risk of the day is cyber,
所以现在,眼下最受关注的风险是 cyber,
33:53
because it's very, very tangible.
因为它非常、非常具体、可感知。
33:55
And some of the risks that are also emerging as pretty real and pretty tangible
而且还有些风险也正在显现,变得相当真实、相当具体,
34:01
are things like child safety is becoming both extremely important,
比如儿童安全,它正变得极其重要,
34:05
but also politically important.
而且在政治上也很重要。
34:07
And then there's some of the risks that are coming down the pipeline that today
然后还有一些正在酝酿、即将到来的风险,它们今天
34:11
could have, but people spend a lot of time with the models,
本来可以,但人们花很多时间跟这些 models 相处,
34:13
see them coming down as things like risk that are ready to biology,
看到它们被归结为类似风险,而这些风险已经准备好用于 biology,
34:17
and specifically whether models will help adversaries produce biological weapons,
尤其是 models 会不会帮对手制造 biological weapons,
34:23
and making that extremely, extremely accessible,
并且让那件事变得极其、极其容易获取,
34:25
producing, making the chance of another COVID or worse that they make,
生产,增加另一场 COVID 或更糟的东西被他们造出来的几率,
34:31
COVID was not engineered to be bad,
COVID 并不是被设计成有害的,
34:33
as if you were trying to do that.
就好像你正试图那么做一样。
34:35
So I think those are some of the risks that are coming down the pipeline.
所以我觉得这些就是接下来会陆续出现的一些风险。
34:37
I think one other thing to just note is that agents are kind of deliberately narrow.
我觉得还有一点要提一下,就是 agents 其实是刻意做得很窄的。
34:43
So like when a Frontier agent company puts a chat button that interacts with customers,
所以比如当一家 Frontier agent 公司放一个跟客户互动的聊天按钮时,
34:49
they've really tried to narrow the topics it's interesting talking about,
他们真的会努力把它能聊的有意思的话题范围收窄,
34:53
such that if you ask it, like what do you think of the precedent,
以至于如果你问它,比如说你怎么看这个先例,
34:55
it will just decline.
它就会直接拒绝。
34:57
Which means there's a kind of risk area is somewhat smaller.
这意味着它的风险范围算是稍微小了一点。
34:59
For models, it is infinite.
但对 models 来说,它是无限的。
35:01
And so there's not a single expert out there
所以外面根本没有任何一个专家
35:05
who can competently evaluate the risks of cyber attacks,
谁有能力评估 cyber attacks 的风险,
35:09
and 15-year-olds having months-long conversations with a chatbot
以及 15 岁的孩子跟 chatbot 进行长达数月的对话,
35:13
and seeing whether it will, in fact, recommend suicide or something horrendous like that,
然后看看它是不是真的会建议自杀,或者类似那种可怕的事情,
35:17
and can evaluate the risks that terrorists can use AI to produce bioweapons.
以及能评估恐怖分子用 AI 制造 bioweapons 的风险。
35:23
The risk surface is just too big, and so the central challenge actually becomes,
这个 risk surface 实在太大了,所以核心挑战其实就变成了,
35:27
how do you get those subject matter experts to work within a one coherent framework,
你怎么让那些 subject matter experts 在一个统一的框架里工作,
35:33
that outputs one coherent report and rating that the world can gun inspect?
输出一份统一的报告和评级,让全世界都能去审查?
35:39
Because that global perspective is central, but there's not a single organization today that could produce that.
因为这种全球视角至关重要,但今天没有任何一个组织能做出这样的东西。
35:45
And you would be the presumptive one when you put out your model standards.
而当你发布自己的 model standards 时,大家就会默认是你。
35:49
We think there can be one company that can, with a consortium of experts,
我们认为,可以有一家公司,能够和一个专家联盟一起,
35:53
build one coherent standard.
制定出一套统一连贯的标准。
35:55
I think we've shown that across all of the enterprise risks today.
我认为,在如今所有企业风险方面,我们已经证明了这一点。
35:59
We think it could be one company that could, with a consortium, specify the audit rules,
我们认为,可以有一家公司,能够和一个联盟一起,明确规定 audit rules,
36:05
basically like their inputs and outputs that all these technical experts need.
基本上就是所有这些技术专家需要的 inputs 和 outputs。
36:09
What access do they need? How should they treat infrastructure security?
他们需要什么样的 access?他们应该如何处理 infrastructure security?
36:13
They can look at whether their evils are well-produced.
他们可以看看他们的 evals 是否做得好。
36:17
Without necessarily being able to say, hey, this is a threat or not a threat,
未必能直接说,嘿,这是威胁还是不是威胁,
36:21
but overall evaluating whether the evils are good or well-constructed,
但总体上是在评估这些 evals 好不好,或者构建得是否扎实,
36:25
that set of order rules basically becomes the interface for these experts,
那套 order rules 基本上就成了这些专家的 interface,
36:29
we think one clearing house could put together.
我们认为一个 clearing house 就能把它整合起来。
36:32
To be clear, when I say one company, I think of it as one company coordinating lots of this
说清楚一点,当我说一个公司时,我把它看作是一个公司在协调很多这类事情,
36:37
in the same way that when we saw our consortium, it's not like we say,
就像我们看到我们的 consortium 时,并不是说,
36:41
we have all the answers on agent security.
我们在 agent security 上拥有所有答案。
36:43
What we say is we are taking on the role of eliciting all of the concerns
我们说的是,我们是在承担一个角色,把所有担忧都引出来。
36:49
and being the secretary that puts it together and runs a tight house
而且还要当那个把它整合起来、把这一摊子管得很紧的秘书
36:53
such that the standard updates lockstep every quarter
让标准更新每季度都步调一致地推进
36:57
and that the order reports that come out in this case, 100 pages of order reports,
还有这种情况下出来的 order reports,100 页的 order reports,
37:01
uniform and crisp and clear to the level of detail that is required
统一、利落、清晰,细到所需的程度
37:05
for executives that need to make a clear go-no-go decision.
给需要做出明确 go-no-go decision 的高管们看的。
37:09
So that's kind of the role that we think we might play.
所以这大概就是我们觉得自己可能会扮演的角色。
37:11
I think in many ways, you're performing the role that O wasp used to do there
我觉得从很多方面看,你正在扮演 O wasp 以前在那里做的那个角色
37:15
and you sit like, you know, competition and partners,
而你就像,你知道的,处在竞争对手和合作伙伴的位置上
37:18
you can go more into like, healthy partner.
你可以再深入讲讲,比如健康的合作伙伴。
37:20
Yeah. So, first of all, O wasp is basically an open source community of security practitioners
对。所以,首先,O wasp 基本上是一个由安全从业者组成的 open source 社区
37:24
that are coming together to build frameworks for addressing the latest security concerns.
他们聚在一起,构建 frameworks,来应对最新的安全问题。
37:28
We think they are phenomenal at creating frameworks.
我们觉得他们在创建 frameworks 方面非常出色。
37:32
We've in fact, first of all, we're partners with them,
其实,首先,我们是他们的合作伙伴,
37:34
so we have a joint article to learn a lot from them.
所以我们有一篇联合文章,可以从他们那里学到很多。
37:37
We think a tremendous source of intelligence.
我们认为这是一个非常巨大的情报来源。
37:40
What O wasp does not do is building the machine that runs third-party audits
O wasp 不做的,是构建运行 third-party audits 的那套机器。
37:46
such that a company like Cursor or a company like Jake Morgan
这样一来,像 Cursor 这样的公司,或者像 Jake Morgan 这样的公司
37:52
could get a third party to go and review them against this and say,
就能让第三方对照这个来审查他们,然后说:
37:55
hey, you've passed the standard.
嘿,你们已经通过这个标准了。
37:56
And here is the report that you can use to build trust and preempt
这就是那份报告,你可以用它来建立信任,并提前化解
37:59
your partners or customers' questions.
你的合作伙伴或客户的问题。
38:01
So they fundamentally try to do something different.
所以,他们从根本上是在尝试做一些不一样的事。
38:04
They are part of the information gathering and intelligence gathering
他们是信息收集和情报收集的一部分,
38:07
and creating clarity, but the operational layer of turning this into promises
以及厘清情况,但把这一切变成承诺的 operational layer
38:11
is not the business they tried to be in.
并不是他们当初想进入的那个业务。
38:14
The standard is emerging and is doing very well.
这个标准正在成形,而且发展得非常好。
38:18
Was it necessary to then also do underwriting obviously it's in the name?
那还有必要也做 underwriting 吗?显然名字里就带着。
38:23
So presumably you thought about it first.
所以想必你们一开始就考虑过这件事。
38:25
I feel like if you just have enough consensus, you don't actually need the money angle,
我觉得,如果你有足够的 consensus,其实并不需要从钱的角度切入,
38:29
but it does happen.
但这种情况确实会发生。
38:30
I did want to also know you guys are a four-profit company too, right?
我还想知道,你们也是一家营利性公司,对吧?
38:35
It's not non-profit work.
这不是非营利工作。
38:37
There's a whole business side to it as well.
这里面还有一整套商业上的考量。
38:39
Yeah, yeah.
对,对。
38:40
Yeah, let's get into the money, but let's start from actually your question
对,我们来聊聊钱的事,不过还是从你那个问题说起吧
38:46
for a profit of a non-profit.
关于非营利组织的盈利。
38:48
In the security space today, cybersecurity, most of the standards are produced by non-profits.
如今在安全领域,也就是 cybersecurity,大多数标准都是由非营利组织制定的。
38:54
I think that's an issue.
我觉得这是个问题。
38:56
The question you have to ask yourself is,
你必须问自己的问题是,
39:01
how do you create good incentives for these standards to be good and keep up?
你怎么建立好的激励机制,让这些标准能做好并跟上发展?
39:08
Non-profits tend to not have these adverse profit incentives
非营利组织往往没有这些不利的利润激励,
39:12
where they hollow out their standard and create the race to the bottom,
也就是那种会掏空自己的标准、制造逐底竞争的激励,
39:16
but they're also not at all responsive by default to the communities that they serve.
但它们默认情况下也完全不会回应自己所服务的社区。
39:21
There's no problem.
这没什么问题。
39:22
They don't have customers that they serve,
它们没有自己服务的客户,
39:24
where they're going to ask what do you want, what do you want, what do you want?
不会去问你想要什么、你想要什么、你想要什么?
39:27
And when you look at the overall satisfaction with the security standards today,
而当你看看如今人们对 security standards 的整体满意度,
39:33
people tend to just not like them very much.
人们往往就是不太喜欢它们。
39:36
You do see in other domains that for-profit standards can serve the world quite well,
你在其他领域确实能看到,营利性标准也能很好地服务世界,
39:43
so there are examples we talked about movies before.
比如我们之前聊过的电影行业的例子。
39:47
It's not without flaws, but it is absolutely critical societal infrastructure
它并非没有缺陷,但它绝对是至关重要的社会基础设施
39:51
that gets run at a astounding scale today.
而且如今是以惊人的规模在运转。
39:54
Your credit score is FICO.
你的 credit score 就是 FICO。
39:56
It's also a for-profit business.
它也是一门营利性生意。
39:58
And when you go back even further in history,
再把时间往历史更早处推,
40:00
some of the crash-testing standards came out of insurance companies,
有些 crash-testing 标准就来自保险公司,
40:06
the insurance companies together founded the Institute of Highway Safety.
这些保险公司一起创立了 Institute of Highway Safety。
40:12
Because they were very interested in how can we use standards to drive down mortality and save money.
因为他们特别想知道,我们怎么用标准来降低死亡率、省钱。
40:18
Our name actually pays homage to the Underriders Laboratory, UL,
我们的名字其实是在致敬 Underriders Laboratory, UL,
40:23
which was started right around when the electricity came out,
它差不多是在电力刚出现的时候成立的,
40:26
houses started burning down,
那时候房子开始一栋栋烧毁,
40:28
insurers again were paying the bill.
保险公司又得掏钱买单。
40:32
And they were maybe also good people,
他们可能也是好人,
40:34
but their profit incentive was, let's prevent houses from burning down,
但他们的利润动机是:咱们得防止房子被烧毁,
40:37
let's test all the electrical products, the light bulbs,
我们来测试一下所有电器产品,灯泡,
40:40
all the light bulbs in here are probably UL tested, the toasters, etc.
这里所有灯泡可能都经过 UL 测试,烤面包机之类的。
40:43
And they set up an entity to create their standards.
然后他们成立了一个实体来制定自己的标准。
40:47
Today, UL has a for-profit entity and a non-profit entity,
如今,UL 有一个营利性实体和一个非营利性实体,
40:51
where they've recognized, they start a non-profit, they spun out a for-profit,
他们意识到,先成立一个非营利机构,再分拆出一个营利性机构,
40:55
because what they recognize was that they actually just serve customers well,
因为他们认识到,其实要想真正服务好客户,
40:58
you need a for-profit entity.
你就需要一个营利性实体。
41:00
The lesson here is one of the ways that the market can align incentives,
这里的教训是,这是市场能让 incentives 对齐的方式之一,
41:04
so you're both responsive to customers,
所以,你既要能对客户保持响应,
41:06
and not hollowing out your standard over time,
又不会随着时间推移把自己的标准掏空,
41:11
is to align it with insurers,
办法就是让它和保险公司对齐,
41:13
because they fundamentally have good incentives.
因为从根本上说,他们有良好的激励。
41:15
And so if you're a for-profit standard that works closely with insurers,
所以,如果你是一个与保险公司紧密合作的营利性标准,
41:20
you get the feedback loop in, such that you're really queuing into your customers,
你就能把 feedback loop 建立起来,这样你才是真正贴近你的客户,
41:24
but also have their interest at heart.
但同时也把他们的利益放在心上。
41:26
So that's the model that we're the kind of inspirational model
所以,这就是那种模式,我们算是那种有启发性的模式
41:29
that we've learned a lot from, and that's also where the name comes from.
我们从中学会了很多,这个名字也正源于此。
41:32
And in some ways, the term underwriting can both be associated with insurance,
而从某些方面来说,underwriting 这个词既可以和保险联系在一起,
41:36
but those are broad term for making decisions.
但这些都只是对做决策的宽泛说法。
41:39
If you're under a decision,
如果你要为一个决策担责,
41:41
you're fundamentally taking ownership for the consequences of it.
你本质上就是在为它的后果承担责任。
41:45
Yeah, I mean, does an insurance contract look like for you?
对,我是说,对你来说,保险合同长什么样?
41:48
Yeah, most of the demand comes today for insurance contracts,
对,如今大部分需求都来自保险合同,
41:52
sitting between people who built AI and people who are buying AI.
处在构建 AI 的人和购买 AI 的人之间。
41:56
Yes.
是的。
41:57
And what you want is the reason why people want insurers involved,
而你想要的是人们为什么希望保险公司参与进来的原因,
42:02
both for the traditional reasons, hey, if something goes wrong,
既有传统原因——嘿,如果出了什么问题,
42:05
we want to be compensated, but it's in particular because insurers can interest,
我们想得到赔偿,但尤其还因为保险公司会感兴趣,
42:10
can bring trust to the equation,
能给这个局面带来信任,
42:12
because insurers will take pay for the damages.
因为保险公司会为损失买单。
42:15
If they are willing to write an insurance policy,
如果他们愿意写一份保单,
42:17
that is them saying, hey, we think there's risk here, but that is manageable.
那就等于他们在说,嘿,我们认为这里有风险,但这是可控的。
42:21
And that is kind of a, they're incentive aligned with the enterprise adopting it,
而这在某种程度上是,他们的激励机制跟采用它的企业是一致的,
42:25
so that's a really a good signal to the market.
所以这对市场来说真是个很好的信号。
42:28
In the same way, actually,
同样地,实际上,
42:30
one of the things that Waymo tried to get the first permit to even operate in San Francisco
Waymo当初想要拿到在San Francisco运营的第一张许可,
42:34
was to get a lot of insurers to stack up a huge insurance policy
做的事情之一就是让很多保险公司一起叠了一份巨额的保险保单,
42:39
in the case of something went wrong, not because Google can't pay,
以防万一出了什么问题,不是因为Google赔不起,
42:42
but because it was very valuable to have a third party go and look at that data
而是因为让一个第三方去看那些数据是非常有价值的,
42:46
that are trusted by government, trusted by enterprises as conservative people
那些被政府信任、被企业信任的保守派
42:51
and say, hey, we've looked at it.
然后说,嘿,我们已经研究过了。
42:53
We're willing to take some of this on top of our cheap.
我们愿意在我们便宜的基础上再承担其中一部分。
42:55
So that's kind of the reason why people interested in it.
所以这大概就是人们会对它感兴趣的原因。
42:57
What it looks like is, in some ways,
它的样子是,在某些方面,
43:00
like everywhere in the insurance contract, you specify what are the perils you want to cover?
就像在保险合同的每一处,你都要明确你想覆盖哪些 perils?
43:04
How much do you want to cover them?
你想给它们提供多少保障?
43:06
Like, what limits?
比如,什么 limits?
43:08
What does it cost to cover that?
保这个要花多少钱?
43:10
And in the case of, if we take a really concrete example,
而且,如果举个非常具体的例子,
43:15
11 labs, but at first of its kind,
11 labs,但这是同类中的第一个,
43:19
AI agent insurance policy,
AI agent 保险单,
43:21
they work with some of the biggest enterprises that work with governments.
他们跟一些最大的、也跟政府合作的企业合作。
43:24
They're really interested in going above and beyond
他们真的很想做得超出预期,
43:26
and making promises to their customers.
并向他们的客户做出承诺。
43:28
So they wrote a policy that covers just some of the core concerns
所以他们写了一份保单,只覆盖了一些核心关切,
43:32
that their customers have been asking about.
而这些正是他们的客户一直在问的。
43:35
And the crucial thing was really to get lawyers of London,
而关键其实真的就是让 London 的律师们,也就是一直作为保险公司的那一方,我们的合作伙伴之一,来看这些数据,并作为第三方和我们一起说:嘿,我们觉得这里有些东西值得承保。
43:40
that was always insurer, one of our partners,
而实际情况其实就是这样。
43:42
to look at this data and be that third party alongside us to say,
然后他们会把那份合同展示给他们的客户,
43:47
hey, we think there's something here that's worth underwriting.
他们就能看到自己的保障额度是多少。
43:50
And that's actually what it looks like.
他们能看到它具体承保哪些内容。
43:53
And so they will show that contract to their customers
然后他们会把那份合同拿给客户看
43:55
and they can see how much they're covered for.
他们就能看到自己有多少保额。
43:57
They can see what exactly it covers.
他们能看到它具体保什么。
43:59
And that will also probably change next year.
而且明年这大概也会变。
44:01
They will want to write an insurance policy that might cover more.
他们会想出一份可能覆盖更多的保险单。
44:04
When you say, Lloyds, is it re-insurance?
你说 Lloyds,是指 re-insurance 吗?
44:06
Or are they sharing somehow at the same level?
还是说他们在同一层级上以某种方式分担?
44:10
Yeah, so typically the way new companies get into insurance
对,所以通常新公司进入保险行业的方式
44:15
is that they partner with insurers,
就是和保险公司合作,
44:18
such that the insurers take the majority
让保险公司承担大部分
44:21
or all of their financial risks.
或他们全部的财务风险。
44:23
Fundamentally, if insurance is useful because it brings trust,
从根本上说,如果保险有用是因为它能带来信任,
44:26
you have to be able to pay the bill.
那你必须得付得起账单。
44:28
Lloyds of London is 400 years old.
Lloyds of London 已经有 400 年历史了。
44:30
They've never not paid a claim.
他们从来没有哪一笔理赔没赔过。
44:31
They're extremely trusted.
他们极其值得信赖。
44:33
What Lloyds of London is struggled to do on their own
Lloyds of London 靠自己一直很难做到的是,
44:37
is to figure out which of the risks are real.
就是弄清楚哪些风险是真实存在的。
44:40
What should we be looking for?
我们应该寻找什么?
44:41
What are the technical controls and running the tests?
technical controls 和跑测试分别是什么?
44:44
So they use AAC1 as kind of the underlying framework.
所以他们会把 AAC1 当作一种底层 framework 来用。
44:48
And we produce a bunch of e-roll results
然后我们会产出一堆 e-roll 结果
44:51
that then directly feed in to inform the pricing.
这些结果会直接输入进去,为定价提供依据。
44:54
So this means that Leven Labs customers know
所以这意味着 Leven Labs 的客户知道
44:57
that that payment will be there.
那笔付款会到账。
44:59
They don't have to look to our series A.
他们不用指望着我们的 A 轮。
45:00
Do we think they have enough cash on the balance sheet?
我们觉得他们资产负债表上的现金够吗?
45:03
They will look at Lloyds.
他们会看看 Lloyds。
45:05
In Lloyds, famously very creative.
在 Lloyds,出了名地很有创意。
45:08
I think I remember some headline.
我好像记得有个标题。
45:10
They insured Jennifer Lopez's, but was something correct.
他们给 Jennifer Lopez 的某个东西投了保,但那是不是真的?
45:13
I think was it David Beckham's right foot?
我记得是不是 David Beckham 的右脚?
45:16
It's stuff like this.
就是这类东西。
45:18
Clearly not a large dataset.
显然不是一个很大的 dataset。
45:21
Exactly.
没错。
45:23
It's actually a remarkable institution
其实这是一个很了不起的机构
45:25
that's both kind of has some of the truly old school virtues
它既有某种真正老派的优点
45:29
of having been around for a long time.
也就是已经存在了很长时间。
45:31
They really operate like a trusted entity
他们运营起来真的就像是一个值得信赖的实体
45:34
and they have appetite to figure out the future.
而且他们有很强的意愿去弄清楚未来。
45:37
And I think there's a lot of recognition
而且我觉得,很多人都认识到
45:40
that both there is like tremendous amount of risk
一方面,存在着巨大的风险
45:42
in AI that is poorly understood today.
在 AI 里,而这一点今天还没有被充分理解。
45:44
So getting into this business carries real risks.
所以进入这个行业确实有真实风险。
45:48
But also this is where lots of the risk exposure
但同时,这也正是大量 risk exposure
45:51
will happen in the future.
会在未来出现的地方。
45:53
This is the one market where risk is truly growing.
这是唯一一个风险真正在增长的市场。
45:56
This is the one market that will also take out
这也是唯一一个会干掉
45:59
some of the existing markets.
一些现有市场的市场。
46:01
Take like auto insurance.
就拿汽车保险来说。
46:02
When there are no human drivers, how's that market going to look?
当没有人类司机时,那个市场会变成什么样?
46:05
Well, it's going to change.
嗯,这会改变的。
46:06
How are you going to assess?
你打算怎么评估?
46:07
You want to ensure we mo?
你想确保 we mo 吗?
46:09
All I say is the principles for how you ensure we mo
我要说的只是,确保 we mo 的原则
46:13
are very similar to how you ensure other kinds of AI.
和确保其他类型 AI 的方式非常相似。
46:15
So getting crash testing.
所以得做 crash testing。
46:17
That's what we do for customer share.
这就是我们为 customer share 做的事。
46:19
A lot more that will also need to happen for we mo,
对于 we mo 来说,还有很多事情也必须要发生,
46:21
which is not how you do it for human drivers.
而这不是你给人类司机做这件事的方式。
46:23
So there's this growing awareness
所以大家越来越意识到
46:25
that the world is changing very fast
这个世界正在飞速变化
46:27
and the only way to learn how to underwrite AI
而学会如何 underwrite AI 的唯一方式
46:30
is to write some policies.
就是写一些保单。
46:32
You may incur some losses and think of that as R&D expense, really.
你可能会承受一些损失,说真的,把它当成 R&D expense 就好。
46:36
But the question for them is like who are the trusted technical partners
但对他们来说,问题就像是:他们能和哪些值得信赖的技术合作伙伴
46:39
they can get into this business with that can help them navigate
一起进入这个业务,并且能帮他们摸清方向?
46:42
and make sure they don't make kind of foolish mistakes
还要确保他们不会犯那种愚蠢的错误
46:45
but also who is willing to hear the wisdom of they have.
但也要看谁愿意听取他们所拥有的智慧。
46:48
They've done this before.
他们以前就做过这种事。
46:49
They've seen who they were there when cyber came out.
他们见过,cyber 出现的时候他们就在场。
46:51
So there are lots of ways in which AI feels completely new
所以,AI 在很多方面让人感觉完全是新的
46:54
but there's also lots of ways in which risks look the same.
但风险也有很多方面看起来是一样的。
46:57
And so there's actually a tremendous amount of wisdom
所以其实有大量的智慧
46:59
sitting in some folks that may have gray hair,
就存在于一些可能已经头发花白的人身上,
47:01
but they really have like a keen sense of
但他们真的有一种很敏锐的感觉
47:05
how to quantify risk.
知道怎么量化风险。
47:07
Yeah.
是啊。
47:08
And the number is so basically like I want $50 million
然后这个数字基本上就像是,我想要 5000 万美元
47:11
with the coverage against these perils
要能覆盖这些 perils 的 coverage
47:14
and violence will give you a quote on it
然后 violence 会针对这个给你一个 quote
47:17
and then you have like a small markup or something
然后你会有个小小的 markup 之类的
47:19
and then you turn it around and do that.
然后你就转手这么做了。
47:21
So it's as simple as it is.
所以事情就这么简单。
47:22
You basically share some of that premium.
你基本上就是分一部分 premium 出去。
47:24
So X% goes to the people who do the pricing of it.
所以 X% 会给到做 pricing 的人。
47:27
It's kind of like a merchant bank for insurance type of thing.
这有点像保险行业的 merchant bank 那种东西。
47:32
Exactly.
没错。
47:33
You basically split the fee and you can think of the insurance supply chain
你基本上就是把 fee 拆开,然后你可以把 insurance supply chain 想成
47:36
as like there's bringing the capital.
就像有引入 capital 这一环。
47:38
There is doing the pricing and there's doing the distribution.
有做 pricing 的,也有做 distribution 的。
47:41
And typically you will pay out some X% of premium hair,
而且通常你会把 premium 的 X% 赔到这里,
47:44
Y% of premium hair and the rest of it will go here.
premium 的 Y% 赔到这里,剩下的会放到这里。
47:46
Does all the insurance will work like this
是不是所有保险都是这样运作的
47:48
or is there some point at which like so right now
还是说到了某个阶段,就像现在这样
47:50
you have equity capital at some point
你在某个时候有 equity capital
47:52
maybe you start raising debt or whatever.
也许你开始 raise debt 或者什么的。
47:55
And then you have enough of a bank account
然后你的银行账户里要有足够的钱
47:57
and enough history that say you've been operation for 10 years.
以及足够的历史记录,比如说你已经运营了 10 年。
48:00
That you don't utilize anymore.
你已经不再使用的那个。
48:01
That's totally an option.
这完全是一个选择。
48:02
And I could see some worlds where that makes sense.
而且我能想象在某些情境下这是说得通的。
48:05
Specifically if there are risks that we feel high confidence
具体来说,如果有一些风险,我们很有信心
48:07
they would want to ensure where the incumbent insurers
他们会想要承保,而现有保险公司
48:10
are too slow to find appetite
太慢找到承保意愿
48:12
or simply struggle to evaluate so that they don't want to do it.
或者干脆难以评估,以至于他们不想做。
48:15
But by and large,
但总的来说,
48:17
in general you do not want to compete with insurers
一般来说,你不会想跟保险公司竞争
48:19
on bringing risk capital to the game.
在把 risk capital 带进这个游戏这件事上。
48:22
For two reasons.
原因有两个。
48:23
One is that's fundamentally a cost of capital game.
第一,这本质上是一个 cost of capital 的游戏。
48:26
They have extremely low cost capital.
他们的 cost of capital 极低。
48:27
Starts with high cost of capital by and large.
而总体来说,起步时的 cost of capital 很高。
48:30
And two, you want to hedge your bets.
第二,你想 hedge your bets。
48:34
And it's very helpful then to also have a portfolio
而这个时候,同时拥有一个 portfolio 会非常有用。
48:37
of home insurance, of car insurance.
房屋保险、车险。
48:39
And we're not about to become a car insurer
而且我们并不打算变成一家车险公司,
48:42
or a home insurer.
或者一家房屋保险公司。
48:43
So they have some natural advantages
所以他们有一些天然优势,
48:44
which makes it much more likely that we'll partner.
这让我们合作的可能性大得多。
48:47
Yeah.
对。
48:48
And they bring the capital at scale
而且他们能带来大规模资本,
48:50
and we bring the technical.
而我们带来技术。
48:51
You're going to work with them for a long time.
你会跟他们长期合作的。
48:52
How are the discussions with the insurers as well?
还有,跟保险公司那边谈得怎么样了?
48:55
So basically they're going off of your certification, right?
所以基本上,他们是根据你的认证来的,对吧?
48:58
They're trusting the diligence on you
他们信任对你做的尽职调查,
49:00
that your certification is valid.
认为你的认证是有效的。
49:02
You tested the right things.
你测了该测的东西。
49:03
And they're backing the money that you have
而且他们愿意投钱,是因为你有
49:06
the right testing in place.
正确的测试到位。
49:08
So any interesting takeaways from working with insurers?
所以,跟保险公司合作下来,有什么有意思的收获吗?
49:11
I think the first thing is they feed into the standard as well.
我觉得第一点是,他们也会给 standard 提供输入。
49:14
So if there are things that they feel like they need
所以,如果有些事情他们觉得自己需要,
49:16
that they're not seeing,
但还没看到,
49:18
we are also taking that as input into the standard
我们也会把这些作为输入,纳入 standard。
49:21
because fundamentally we think a good standard
因为从根本上说,我们认为一个好的 standard,
49:24
is one that creates a really healthy promise ecosystem
就是能打造出一个非常健康的 promise ecosystem,
49:27
and we think insurers are a critical part of that.
而且我们认为保险公司是其中非常关键的一部分。
49:29
And again, there are the most well incentivized
再说一次,他们是最有动力这么做的
49:32
to they see all the lost data across any particular CSO
他们能看到任何特定 CSO 的所有丢失数据
49:35
knows their particular concerns.
每个特定的 CSO 都知道自己具体的顾虑。
49:37
Insurers see the concerns across the entire portfolio
保险公司能看到整个 portfolio 中的顾虑
49:40
and often have direct access
而且往往能直接获取
49:42
to what exactly happened,
到底发生了什么,
49:43
who is it fault, et cetera,
是谁的过错,等等,
49:44
as they do a part of their forensics.
因为这是他们 forensics 的一部分。
49:46
So they're actually like a great source of intelligence on this.
所以其实他们在这件事上是个很好的情报来源。
49:48
One of the big takeaways from cyber insurance,
cyber insurance 的一个重要启示是,
49:52
which is a market that didn't work that well
而这个市场其实做得并不太好,
49:54
was that the insurance and the technical expertise
就是保险和技术专业能力
49:57
was not married up.
并没有很好地结合起来。
49:59
Well, our conviction is that standards have to proceed insurance,
嗯,我们坚信,标准必须先于保险,
50:03
fundamentally,
从根本上说,
50:04
what everyone first and foremost want,
大家首先最想要的,
50:06
whether you are a CSO at Jig Morgan
不管你是 Jig Morgan 的 CSO
50:08
or a CSO at cursor
还是 cursor 的 CSO
50:11
or an underwriter
或者是一名 underwriter
50:14
at a Lloyds and London syndicate,
在 Lloyds and London syndicate,
50:16
is you want to not have an incident in the first place.
你想要的,都是一开始就别出事故。
50:20
You want to know that the risk is well managed
你想知道风险得到了很好的管理
50:22
and only then doesn't insurance that to make sense.
也只有这样,那样的保险才说得通。
50:24
So we'll see the standard ecosystem
所以我们会看到标准的 ecosystem
50:25
basically run ahead of the insurance
基本上就是跑在保险前面
50:27
and the reason why we,
而原因就是,我们,
50:29
you asked us a kind of why I also do insurance,
你问过我们,有点像为什么我也做保险,
50:31
this is kind of proving what we think
这有点像在证明我们认为
50:33
at a whole promise conference infrastructure
整个 promise conference infrastructure
50:35
ecosystem needs to look like
ecosystem 需要长什么样
50:37
and we think it's very compelling to bring that to life.
而且我们觉得,把它真正做出来非常有吸引力。
50:40
Even if we think the standard is kind of there,
即便我们觉得标准已经差不多在那儿了,
50:42
the core linchpin that unlocks the rest.
解锁其余部分的核心关键。
50:44
There's minimal claims here, right?
这里的 claims 很少,对吧?
50:46
This is one of those things
这就是那种情况
50:47
where people haven't really worked through
人们还没有真正想透
50:50
what it means to cover things.
覆盖这些东西到底意味着什么。
50:52
So for example,
所以举个例子,
50:53
I pay cursor $20 a month
我每个月给 cursor 付 20 美元
50:55
and I write,
然后我写,
50:56
I buy code something
我买 code 这种东西
50:57
that makes a plane crash
它能让飞机坠毁
51:00
causing $200 million worth of damage.
造成 2 亿美元的损失。
51:02
Do I claim $20?
那我是索赔 20 美元吗?
51:03
So do I claim $200 million?
那我是索赔 2 亿美元吗?
51:05
Yeah.
对。
51:07
So these are all great questions.
所以这些都是很好的问题。
51:10
And fortunately,
而且幸运的是,
51:12
all of insurance and legal history
整个保险和法律的历史
51:14
kind of helps answer some of those questions.
某种程度上能帮忙回答其中一些问题。
51:17
I think the first thing is people
我觉得第一点是,人们
51:18
have limits on that policy.
会受到那份保单的限额约束。
51:20
So if you want to claim $200 million,
所以如果你想索赔 2 亿美元,
51:22
you have to,
你就必须,
51:23
someone has to have paid a lot
得有人付过一大笔钱,
51:24
for that insurance policy upfront
而且是提前为那份保单付的。
51:25
to have $200 million a coverage.
每份 coverage 有 2 亿美元。
51:27
And ultimately,
而最终,
51:29
the way this works is that
它运作的方式是这样的:
51:32
you start from a lot of uncertainty.
你一开始是从一大堆不确定性出发的。
51:34
This is not just an insurance,
这不仅仅是一种保险,
51:35
but also like,
而且就像,
51:37
can you use,
你能不能用,
51:38
can anthropic use books
anthropic 能不能用书
51:39
from the internet to train up?
从互联网上拿来 train 吗?
51:41
Well, they can go and look at precedent,
嗯,他们可以去看看先例,
51:42
they can see.
他们能看出来。
51:43
But ultimately,
但最终,
51:44
these things get settled in court
这些事情会在法庭上解决,
51:45
and you hammer it out over time.
然后随着时间慢慢磨出来。
51:47
So you start from this,
所以你是从这一点开始,
51:48
like,
就像,
51:49
place of ambiguity,
模糊性所在的地方,
51:50
which is both why insurance
这既是为什么保险
51:52
can be hard to do early on,
在早期会很难做,
51:53
but it's also why people want insurance,
但这也是为什么人们想要保险,
51:55
because ambiguity slows down adoption.
因为模糊性会拖慢采用。
51:57
That also sets out the heads of the...
这也列出了……的要点
51:59
Yeah.
嗯。
52:00
In some ways,
在某种程度上,
52:01
actually, the first incident will help to establish a lot of this.
其实,第一起事件会有助于确立很多这方面的事情。
52:04
Exactly.
没错。
52:05
And there have been a number of incidents out there
而且现实中已经发生过不少事件了,
52:07
that have just not been insurance-covered,
只是这些都没有被保险覆盖,
52:08
so take the now old example from the Americana,
所以就拿 Americana 那个现在已算旧的例子来说,
52:13
where...
当时……
52:14
I was going to run it up.
我本来打算把这件事往上反映。
52:15
So I decided a refund policy.
所以我决定搞一个退款政策。
52:17
And the question was,
问题是,
52:19
at kind of the case where like,
有点像这种情况,就是,
52:21
hey, we have nothing to do with this.
嘿,这事跟我们没关系。
52:23
This chatbot messed up,
这个 chatbot 搞砸了,
52:25
but like, sorry.
但就是,抱歉。
52:27
And the courts were like,
然后法院就说,
52:28
no, if you put your chatbots
不,如果你让你的 chatbots
52:30
to interact with your customers,
去跟你的客户互动,
52:31
they make legally binding promises on your behalf.
他们代表你做出具有法律约束力的承诺。
52:33
That is now precedent for everything in the future,
这就成了未来一切事情的先例,
52:36
where you will...
到时候你会……
52:38
if someone wants to deploy a chatbot like that again,
如果有人想再次 deploy 那样的 chatbot,
52:40
they should not expect to be able to just pawn off
他们就不该指望能就这么甩锅,
52:43
and say, sorry, my chatbot lied.
然后说,抱歉,我的 chatbot 撒谎了。
52:45
It's nothing to do with me.
这跟我一点关系都没有。
52:46
I bought it from OpenAI.
我是从 OpenAI 买的。
52:47
No.
不。
52:48
If you're putting this in front of your customers,
如果你把这个摆到客户面前,
52:50
you are taking responsibility for it.
你就是在为它承担责任。
52:52
And so every court case
所以每一个法庭案件,
52:54
where the insurance is involved
只要保险牵涉其中
52:55
and not clarifies liability,
却没有厘清责任,
52:57
and liability is kind of the foundation for insurance.
而责任差不多就是保险的基础。
53:00
There's another reason why a standard insurance
还有一个原因,为什么标准保险
53:02
come together.
汇聚到一起。
53:03
Liability...
Liability……
53:04
I'll go on a little tangent here.
我先稍微跑个题。
53:06
Please. Please. Please.
请讲,请讲,请讲。
53:07
Liability often...
Liability 往往……
53:08
One of the core concepts is
其中一个核心概念是
53:09
whether someone was negligent.
某人是否 negligent。
53:11
Should they have seen this?
他们本应该看到这一点吗?
53:12
Should they have prevented this?
他们本该阻止这种情况吗?
53:14
And the question you...
而这个问题,你……
53:16
How do you judge that?
那你怎么判断呢?
53:17
Well, you basically judge whether they've met their duty of care.
嗯,基本上就是判断他们是否履行了注意义务。
53:20
What does that mean in practice?
这在实践中意味着什么?
53:22
Well, often they look to standards.
嗯,他们通常会参考标准。
53:24
So if there's a standard
所以,如果有一个标准
53:26
that is broadly adopted
是被广泛采用的
53:28
that says you must have a groundedness filter
说必须得有一个 groundedness filter
53:30
or you must have a jailbreak filter,
或者必须得有一个 jailbreak filter,
53:32
it becomes way harder to claim ignorance
就更难声称自己不知道
53:34
of these things existed.
这些东西存在过。
53:35
And so, setting standards
所以,制定标准
53:37
help clarify liability.
有助于厘清责任。
53:38
Courts will often point to standards
法院往往会援引标准
53:41
and being like,
然后说,
53:42
well, this seems like best practice to do
嗯,这看起来是该做的 best practice
53:44
is there for everyone to see.
就放在那里让所有人都能看到。
53:45
So, there's another way in which standards
所以,还有另一种方式,standards 在其中
53:47
are kind of civilization infrastructure
算是文明基础设施
53:49
that insurance can be built on,
保险可以建立在其上,
53:51
which promises can then be built on.
而承诺又可以建立在其上。
53:52
Yeah, until you get that.
是啊,直到你做到那一点。
53:54
We don't have to get certified
我们不是非得拿到认证
53:55
to write these to, you know,
把这些写到,你知道,
53:57
make these, like, bots and all these.
把这些做成,比如,bots 还有这些。
53:59
Right.
对。
54:00
But, like, basically,
但是,就,基本上,
54:02
whenever we go for the audit,
每次我们去做 audit 的时候,
54:04
I think people will start to shape up
我觉得大家就会开始规矩起来
54:06
and all this stuff.
还有所有这些事。
54:07
I wonder if, like, that means that
我在想,是不是,那意味着
54:09
you don't also then become, like,
那你也就不会同时变成,像是,
54:11
the approving authority for me to ship to production.
我 ship 到 production 的审批方。
54:14
You know, like, yes, you check once per quarter.
你知道,就像,对,你每个季度检查一次。
54:17
I want to ship once a day.
我想每天 ship 一次。
54:19
Yeah.
嗯。
54:20
And I don't know when one of my things breaks,
而且我也不知道我的某个东西什么时候会坏,
54:22
like, one of your certifications or not.
比如,你的某个 certification 有没有坏掉。
54:24
So, there's a couple of things.
所以,有几件事。
54:26
There's a couple of requirements in there
这里面有几条要求
54:28
that relate to how do you yourself,
跟你们自己怎么做有关,
54:31
where you have to tell...
也就是你得告诉……
54:32
It's like an ongoing monitoring.
这就像是一种 ongoing monitoring。
54:34
How are you yourself testing
你自己是怎么做 testing 的
54:35
before you make a major releases?
在你们做 major releases 之前?
54:38
We don't go and audit people every day,
我们不会每天去 audit 别人,
54:40
but at least there is now a trail where
但至少现在有一个 trail,能……
54:42
if you do a major mess up,
如果你捅了个大篓子,
54:44
then you're across there,
那你就到那边了,
54:45
and they come and ask you,
然后他们过来问你,
54:46
hey, you promised me that you were going to
嘿,你答应过我,你会
54:47
run these emails yourself.
亲自处理这些邮件。
54:49
And for lots of them, most...
而且对其中很多来说,大多数……
54:51
Most of the PRs that be emerged
大多数出现的 PRs
54:53
were not fundamentally all of the product experience
从根本上并不等同于整个 product experience
54:55
and some of them were,
而且其中一些是,
54:56
I think...
我觉得……
54:57
Sometimes you don't know.
有时候你并不知道。
54:58
And sometimes you don't know.
而且有时候你也不知道。
54:59
And this is also true.
这同样也是真的。
55:00
And this is also there are some inherent risks
而且这里面也有一些固有的风险
55:01
that everyone kind of...
每个人多多少少都……
55:02
Everyone knows that when they buy software,
每个人都知道,当他们买软件的时候,
55:03
there can be bugs,
可能会有 bug,
55:04
and this is just part of it.
而这只是其中一部分。
55:06
What they can,
他们能怎么样呢,
55:07
if you're selling to mom and pop shops,
如果你是在卖给夫妻店,
55:09
they may not care.
他们可能不在乎。
55:10
They're just like,
他们就会说,
55:11
well, I want to use your tool.
嗯,我就是想用你的工具。
55:13
So I'm just going to want to bring in
所以我就只是想带进来
55:14
to take that risk on.
去承担那个风险。
55:15
If you're selling to a big bank,
如果你是在向一家大银行推销,
55:17
they might be like,
他们可能会说,
55:18
sorry, we're making promises to our customers.
不好意思,我们在向客户做出承诺。
55:20
If you can't make a promise to us that we can pass on,
如果你不能向我们做出一个我们能转达出去的承诺,
55:22
we don't want to work with you.
我们就不想跟你合作。
55:24
And then it's up to you to say,
然后就得由你来说,
55:25
do I care for my agent
我到底在不在乎我的 agent?
55:27
to get used
被用起来
55:28
as critical infrastructure in this nation?
作为这个国家的 critical infrastructure?
55:30
If so,
如果是这样的话,
55:31
at least I can make promises about
至少我可以承诺,关于
55:33
what process I run.
我运行的是什么 process。
55:34
And then we can go and test that every quarter
然后我们可以每个季度去测试这一点
55:36
to be like,
然后说,
55:37
well, this seems like it's still...
嗯,这看起来好像还是……
55:39
kind of...
有点……
55:40
it still makes this standard.
它仍然让这成为标准。
55:42
So from my perspective,
所以从我的角度来看,
55:43
it's kind of a way to...
这有点算是……的一种方式
55:44
big companies by default kind of have
大公司默认情况下多多少少会有
55:46
some amount of trust
一定程度的信任
55:47
when they ship AI.
当他们发布 AI 的时候。
55:48
If you're a young company,
如果你是一家年轻公司,
55:50
if you're just signing out,
如果你只是 signing out,
55:51
but default, you have no trust.
但默认情况下,你是没有 trust 的。
55:52
And there are very few places
而且很少有地方
55:53
where you can go and get trust.
你可以去获得 trust。
55:55
So one of the things that most of our customers did
所以我们大多数客户做过的一件事
55:57
before they started working with us
在他们开始和我们合作之前
55:58
is that they would make their own security blog posts.
就是他们会自己写 security blog posts。
55:59
That's great.
这很棒。
56:00
But also,
但还有,
56:01
who's going to trust you,
谁会信任你呢,
56:02
saying,
说,
56:03
we're so secure.
我们特别 secure。
56:04
Like anyone can write that.
这话谁都能写。
56:05
But it's very hot.
但这非常火。
56:06
Where do you go and get that trust?
那你去哪儿才能得到那种信任呢?
56:08
And so I think making the
所以我觉得,做出那个
56:10
standards more legible
标准更清晰易懂
56:12
makes it easier for smaller companies
让小公司更容易
56:14
to prove that they're doing
证明自己在做
56:16
what they ought to be doing.
他们本该做的事。
56:17
Because the default assumption is
因为默认的假设是
56:19
that it's a wild west.
这里就是一片蛮荒西部。
56:20
Is there a roadmap you have
你有没有一个 roadmap
56:22
of like,
就是那种,
56:23
there's a lot of work to be done here, right?
这里还有很多工作要做,对吧?
56:24
This is the first one.
这是第一个。
56:26
Anything on the roadmap of
roadmap 上有没有
56:28
what you see is next,
你觉得接下来是什么,
56:29
what's coming,
接下来会有什么,
56:30
what's missing.
还缺什么。
56:31
I think when we
我觉得当我们
56:33
when we zoom out,
当我们把视角拉远,
56:34
as you want deals with agents,
既然你想处理 agents,
56:36
we will,
我们会的,
56:37
next up,
接下来,
56:38
we will deal with models.
我们会处理 models。
56:39
Next up from that,
再接下来,
56:40
we will deal with robotics,
我们会处理 robotics,
56:41
of which,
其中,
56:42
in some ways,
在某些方面,
56:43
Waymo is the first robot.
Waymo 是第一个 robot。
56:44
But the exact same problem is going to be
但完全相同的问题将会是
56:46
someone's going to develop a robot.
有人会开发一个 robot。
56:47
Someone's going to need some promises.
有人会需要一些承诺。
56:49
They're going to struggle
他们会很难
56:50
to make the promises.
去做出这些承诺。
56:51
And as you see this playing out
而当你看到这一切逐渐展开时
56:53
when like,
当,就像,
56:54
if you think
如果你觉得
56:55
fable concerns are bad,
对 fable 的担忧很糟,
56:56
like see when Waymo hits a dog,
就像看到 Waymo 撞到狗的时候,
56:57
and as if people lose their mind,
人们好像都疯了,
56:58
imagine when first robot
想象一下,当第一个 robot
57:00
knocks off a toddler
把一个小孩撞下
57:01
off a kitchen table.
厨房餐桌。
57:02
Yeah.
是啊。
57:03
You're going to see some
你会看到一些
57:04
real strict liability.
真正的 strict liability。
57:06
I mean, it's here, right?
我是说,它已经来了,对吧?
57:07
The crews got fully
那些团队被彻底
57:08
destroyed.
摧毁了。
57:09
No, it's all permits.
不,全是许可证。
57:10
They're gone.
都没了。
57:11
Yeah.
对。
57:12
Right.
对。
57:13
So physical AI,
所以 physical AI,
57:14
the level of stringency
严格程度
57:15
just goes up,
就是一路往上涨,
57:16
and up, and up,
越来越高,越来越高,
57:17
and up.
越来越高。
57:18
And so that's kind of like
所以这有点像是
57:19
the big picture,
大局,
57:20
agents,
agents,
57:21
models, robotics.
models、robotics。
57:22
I think within agents,
我觉得,在 agents 这个范围内,
57:24
the current set of agents
当前这批 agents
57:25
are well covered by this.
已经被这个很好地覆盖了。
57:26
But as the technology progresses,
但随着技术不断进步,
57:28
as agents get
随着 agents 变得
57:29
longer,
更长,
57:30
horizons,
时间范围,
57:31
new types of failure modes
新型的 failure modes
57:33
will emerge.
会出现。
57:34
And so it's mostly of,
所以它主要就是,
57:35
can you make sure
你能不能确保
57:36
that it's going to keep up
它能跟得上
57:37
when they appear?
当它们出现的时候?
57:38
And you also start to see
而且你也会开始看到
57:39
new modalities like today,
像如今这样的 new modalities,
57:41
world models,
world models,
57:42
it's mostly,
基本上,
57:43
kind of a research.
算是某种研究吧。
57:45
Question,
问题是,
57:46
there's no one who's
还没有谁
57:47
really using it.
真的在用它。
57:48
But they will also bring in
但他们也会引入
57:49
just new kinds of ways
只是新的方式
57:50
to create value,
去创造价值,
57:51
but also more risk surface
但也带来更多的 risk surface,
57:53
that no one knows how to grab
而没人知道该怎么抓住
57:54
with today.
用今天的方式。
57:55
You'll start to see true
你会开始看到真正的
57:56
agent to agent interactions
agent 与 agent 之间的互动
57:57
that are not mediated
是不经过中介的
57:58
by humans.
by humans.
58:00
There's going to be a bunch
会有很多
58:01
of interesting questions.
有趣的问题。
58:02
You're basically going to need
你基本上会需要
58:03
a new legal system.
一套新的法律体系。
58:04
How do they build trust
它们如何建立
58:05
amongst each other?
彼此之间的信任?
58:06
How one of the core things,
如何,其中一个核心的事情,
58:07
when humans trade
当人类交易
58:08
with each other is that you know
彼此之间的时候,你知道
58:09
that you have recourse,
你有追索权,
58:10
you can sue them.
你可以起诉他们。
58:11
How do you make sure
你怎么确保
58:12
that there is a persistent
有一个持续存在的
58:13
balance sheet
balance sheet
58:14
behind any agent such
在任何这样的 agent 背后
58:15
that if you trade with it
那就是,如果你跟它做交易
58:16
and it screws you,
结果它坑了你,
58:17
you can get your money back.
你能把钱要回来。
58:18
Those are some of the questions
这些就是其中一些问题
58:19
we're going to have to deal with.
我们将来得去处理。
58:20
And the technical testing
而技术测试
58:23
of multi-agent systems
针对 multi-agent systems 的
58:25
is also going to be
也将会是
58:26
interesting and complex.
有意思,也很复杂。
58:29
Very fun.
特别好玩。
58:30
Are there any perils
有没有哪些风险
58:31
that are uninsurable
是现在无法承保的
58:32
right now that people wish
但人们希望
58:34
that you would?
你们能承保?
58:35
Yeah, one of the places
对,其中一个地方
58:37
where there's a bunch of
那里有一大堆
58:39
appetite for insurance
对保险的兴趣
58:41
and not a lot of demand
而且需求不多
58:43
but not a lot of supply
但供给也不多
58:44
is when it comes
就是当谈到
58:45
to copyright.
copyright 的时候。
58:46
Ooh.
哦。
58:47
In some ways copyright
在某些方面,copyright
58:48
is kind of mundane.
其实有点平淡无奇。
58:49
It's always been an issue.
这一直是个问题。
58:50
There's a couple of reasons
有几个原因
58:51
for this.
导致这种情况。
58:52
The first is people
第一个原因是,人
58:54
who have trained
那些用受版权保护的材料
58:55
on copyrighted materials
训练过的人
58:56
almost always know
几乎总是知道
58:57
that they've done that.
他们做过那件事。
58:58
So if you want to buy
所以如果你想买
59:00
insurance for it,
它的保险,
59:01
it probably signals
这很可能说明
59:03
that you might be a high-risk customer.
你可能是个高风险客户。
59:05
The people
那些人
59:06
who are most interested
最感兴趣的
59:07
in getting insurance
是买保险
59:08
for copyright infringement
来应对 copyright infringement
59:09
are the people
就是那些人
59:10
who are most likely to have
最有可能已经
59:11
copyrighted.
取得版权。
59:12
It's like a lemon problem.
这就像个 lemon problem。
59:13
Exactly.
没错。
59:14
There's another side
还有另一面
59:15
too, right?
对吧?
59:16
Like, if you're building
就像,如果你在构建
59:17
on something so say I'm using
在某个东西上,比如说我在用
59:18
an open model,
一个 open model,
59:19
I don't know what it's
我不知道它
59:20
trained on, right?
是 trained on 什么的,对吧?
59:21
Yes.
是的。
59:22
And how far down that
那沿着这条
59:23
chain does copyright go?
链往下,版权能追到多深?
59:24
Yes.
是的。
59:25
Am I liable
我要担责吗
59:26
on my product
为我的产品
59:27
because
就因为
59:28
company X
company X
59:29
trained on copyright?
拿版权内容训练过?
59:30
But there's safety
但人多
59:31
in numbers.
就安全。
59:32
If everyone's doing it,
如果大家都这么做,
59:33
then you cook.
然后你就做饭。
59:34
I mean, and I would say
我是说,而且我会说
59:35
until, you know,
直到,你懂的,
59:36
Fable has rolled back
Fable 已经 rolled back
59:37
from everyone I used there.
从我在那儿用过的每个人那里。
59:38
Yeah.
嗯。
59:39
I think it's a hard question.
我觉得这是个很难的问题。
59:40
I don't know the answer to that.
我不知道这个问题的答案。
59:41
But I think your
但我觉得你的
59:42
your intuition is right
你的直觉是对的
59:43
and kind of like,
而且有点像,
59:44
what is the kind of
究竟是什么样的
59:45
duty of care?
duty of care?
59:47
And people
而人们
59:48
don't today think
今天并不
59:49
of it as
把它看作
59:50
customary that you go
按惯例你会去
59:51
and you, like,
然后你,就像,
59:52
dissect your open
剖析你的 open
59:53
models training data
models training data
59:54
and you check everything.
然后你检查所有东西。
59:55
In fact,
事实上,
59:56
lots of people use
很多人都在用
59:57
them.
它们。
59:58
It's seen as kind of
这被视为算是
59:59
generally acceptable
普遍可以接受的
60:00
to not check for this.
不去检查这一点。
60:01
And therefore, we're not
所以,我们不会
60:02
going to hold names
保留名字
60:03
because we also
因为我们也
60:04
really can't, right?
真的做不到,对吧?
60:05
We don't know the
我们不知道
60:06
training data.
training data。
60:07
That's the benefit.
这就是好处。
60:08
But I think no court is
但我觉得没有哪个法院会
60:09
going to get a copyright
接到版权
60:10
question because it's
问题,因为它
60:11
actually to get banned.
其实会被封禁。
60:12
Higher Nicholas
Higher Nicholas
60:13
Kalini and he can extract it.
Kalini,然后他就能把它提取出来。
60:14
Exactly.
没错。
60:15
Though he is in short
不过他确实
60:16
supply.
供不应求。
60:17
Yeah.
是啊。
60:18
He only has so many
他也就只有那么多
60:19
Kalini's.
Kalini's。
60:20
Exactly.
没错。
60:21
So I think this is
所以我觉得这是
60:22
also fair.
也有道理。
60:23
In the case of labs,
就 labs 来说,
60:24
there's a lot of interest for this.
对此有很多兴趣。
60:25
I wanted it is what makes
我想知道,这是不是只是让
60:26
just insurance suspicious of
保险对它
60:27
it.
起疑。
60:28
And so you have a
于是你就有了一个
60:29
lemons problem.
lemons problem。
60:30
Yeah.
嗯。
60:31
Is there like a theory
有没有那种理论
60:32
of insurance where adverse
关于保险,其中 adverse
60:33
selection dominates the risk
selection 主导了保险的 risk
60:35
share aspect of insurance?
share 那一方面?
60:37
Like, where does this
比如,这在
60:39
teach us insurance?
保险方面教会了我们什么?
60:40
A lot of
很多
60:41
insurance does come back to
保险确实会回到
60:42
like practical
比如那些实用的
60:43
versions of
版本,属于
60:44
microcomics 101.
microcomics 101。
60:45
Yeah.
是啊。
60:46
It's like, this is why
就像,这就是为什么
60:47
you do get worse.
你确实会变得更糟。
60:48
You need to pull health
你需要把健康拉起来
60:49
insurance because if you make
保险,因为如果你把
60:51
it too hyper-specific,
它弄得太过于具体,
60:52
then only people who are
那么只有那些
60:53
guaranteed to get the
一定会得这种
60:54
disease will sign
病的人,才会签下
60:55
off your insurance.
你的保险。
60:56
Exactly.
没错。
60:57
Same thing.
一回事。
60:58
The core problem is
核心问题在于
60:59
one of information is
信息
61:00
symmetry.
不对称。
61:01
If you were buying insurance,
如果你在买保险,
61:02
know something about the
你了解一些
61:03
risk that insurance did not
保险公司不知道的
61:04
know.
风险。
61:05
And so the question is actually
所以问题实际上是
61:06
and this comes back to the same
而这又回到了同一个
61:07
problem is if you rely, you
问题是,如果你依赖,你
61:08
can break out a lot of these
可以拆出很多这些
61:09
information is
信息就是
61:10
symmetries.
symmetries。
61:11
If there is some kind of
如果存在某种
61:12
testing that reveals the
能揭示出
61:13
underlying true risk.
底层真实风险的测试。
61:15
And so if you were able to,
所以,在你说的情况下,如果你能做出很好的诊断,判断出某人是否患病,或者某人患病的概率有多大,而且这个诊断是保险方信任的,那么他们可能就愿意承保。
61:18
in the case you mentioned, have
在你提到的那个案例里,能
61:20
good diagnosis of whether
很好地诊断出某人是否
61:22
someone has it or what the
得了它,或者
61:23
probability is that someone has it,
某人得了它的概率是多少,
61:24
that the insurance trust, then
那个 insurance trust,那么
61:26
they might be willing to
他们可能愿意
61:27
ensure it.
确保它。
61:28
But if they don't, if there's
但如果他们不这样做,如果
61:29
no kind of common information,
没有任何 common information,
61:30
then only the patient will
那就只有患者会
61:31
know if it sort of breaks it
知道它是不是有点把它
61:32
down.
拆解了。
61:33
So the question is again, how
所以问题又来了,你
61:34
do you create credible
怎么创造 credible
61:35
signaling between players?
signaling between players?
61:38
This is also the whole reason why
这也是 Moody's 存在的全部原因。
61:39
Moody's exists.
Moody's 就是在做 credible signaling。
61:40
Moody's just does credible
这也是为什么 Moody's 永远都不可能——Moody 必须保持独立。
61:41
signaling.
如果 Moody's 是由……拥有的
61:42
That's also why Moody's
这也是为什么 Moody's
61:43
could never, Moody has to be
永远不能,Moody 必须是
61:44
independent.
独立的。
61:45
If Moody's was owned by
如果 Moody's 是由
61:46
J.P. Morgan, then J.Morgan
J.P. Morgan,然后是 J.Morgan
61:48
cannot use it as a
不能把它当作一个
61:49
signaling mechanism.
signaling mechanism。
61:50
So a lot of
所以很多
61:53
the basics of standards and
标准和认证的
61:55
certification are just
基础其实只是
61:56
communication devices.
沟通工具。
61:57
It's just a trust gap.
这只是一个信任鸿沟。
61:58
And that's why you have to
这就是为什么你必须
62:01
think about what are the
想一想,什么是
62:02
incentives of the messenger.
信息传递者的动机。
62:03
And one another way you can
还有另一种办法,你可以
62:04
break a lot of this through
打破很多这种情况,通过
62:05
transparency.
透明度。
62:06
If you are transparent in how
如果你在如何
62:07
you operate, you just cannot
运作方面是透明的,你就不能
62:09
mess with others nearly as
去招惹别人几乎一样
62:10
easily.
容易。
62:11
You make it much more costly.
你让它变得代价高得多。
62:12
And that increases trust.
而这会增加信任。
62:13
And this is one of the
而这是其中一个
62:14
reasons why there's a change
原因:为什么这里会有一个 change
62:15
lock here.
lock。
62:16
Every little change.
每一个微小的改变。
62:17
Yeah, yeah.
对对。
62:18
You can go back and find.
你可以回去找找看。
62:19
And it means that if
而这意味着,如果
62:22
we were to make this
我们把这个
62:23
standard worse.
标准变得更差。
62:24
Oh, wow.
哇哦。
62:25
That's a lot of changes in
这会有很多改动,在
62:26
one update.
一次更新里。
62:27
Yeah.
嗯。
62:28
Okay.
好。
62:29
And a lot of this is just as
而且很多这些,只是随着
62:30
things get clearer.
事情变得更清楚。
62:31
You can see a lot of
你会看到很多
62:32
clarifications.
澄清说明。
62:33
You can see some revisions.
你会看到一些修订。
62:34
As things get hammered out,
随着事情逐渐敲定,
62:35
you want to change this.
你想改变这一点。
62:36
But if you make it all public,
但如果你把它全部公开,
62:38
you make it much harder to
你就更难
62:39
mess with people.
糊弄别人了。
62:40
Or at least you become found
或者至少你
62:41
out very easily.
很容易就会被发现。
62:42
And so this is a way of
所以这是一种
62:44
increasing, so reducing the
增加,从而减少……
62:46
information asymmetry is
information asymmetry 是
62:47
by just making more
就靠让更多
62:48
than information public.
比信息公开。
62:49
I like how you do know when
我喜欢你确实知道什么时候
62:51
future versions are coming.
未来版本要来了。
62:52
Yeah, I guess it's
嗯,我猜这
62:53
quite.
挺。
62:54
I mean, it's quite a
我是说,这挺
62:55
quite.
确实。
62:56
Yeah.
对。
62:57
Yeah.
对。
62:58
Yeah.
对。
62:59
That's amazing.
太厉害了。
63:00
Yeah.
对。
63:01
But this is also promise.
但这也是一种承诺。
63:02
Like we three now don't deliver
就像我们三个现在不兑现一样。
63:03
until I 15th.
直到15号。
63:04
I mean, you can just batch it
我是说,你直接把它 batch 起来就行
63:05
up and then whatever you got,
然后不管你拿到什么,
63:06
you just shut it.
就直接把它关掉。
63:07
Yeah.
嗯。
63:08
But it's kind of like we
但这有点像,我们
63:09
deposit some amount of trust
会存入一定量的信任
63:10
every time we meet this
每次我们遇到这个的时候
63:11
commitment.
承诺。
63:12
Yeah.
对。
63:13
And in the startup
而在 startup
63:14
line, it feels easy to ship
这条线上,感觉很容易
63:16
any version of a standard
把 standard 的任何一个版本
63:18
to quarter.
按季度交付。
63:19
And the enterprises who are
而那些已经
63:20
used to this like decade
这样习惯了差不多十年的企业
63:21
long cycle.
长周期。
63:22
We often get met with like
我们经常会遇到那种
63:24
incredulogy, like there's
难以置信的反应,就好像
63:25
just no way.
根本不可能。
63:26
And then you show them the
然后你给他们看
63:27
change log.
change log。
63:28
One thing I wanted to also
还有一件事,我也想
63:29
like try to really think
就试着真正去想一想
63:30
about is, you know, you say
它讲的是,你知道,你说
63:32
something about how if you
某种说法,就是如果你
63:33
have tests for the thing,
有对这个东西的测试,
63:34
then you can ensure it.
那你就能确保它。
63:35
Yes.
是的。
63:36
Right.
对。
63:37
And so really what your
所以,其实你的
63:38
standard is, what AI you see
标准是,你看到的 AI 是什么
63:39
is establishing their framework
正在建立他们的 framework
63:41
for the audits to happen
好让 audits 能进行
63:43
so that you can at least test
这样你至少能测试
63:44
like all these like
就像是所有这些,就像是
63:46
baseline standards of care
baseline standards of care
63:47
have been met.
都已经满足了。
63:48
And therefore people can
所以人们可以
63:49
ensure against standard
确保符合标准
63:50
risks that everyone has.
每个人都会有的风险。
63:51
I wonder if like there
我在想,是不是
63:52
needs to be, you need to
需要有,你得
63:53
develop other tests.
开发其他测试。
63:54
We've covered Mechin
我们过去讲过 Mechin
63:56
Terp in the past.
Terp。
63:58
Any interest in that?
你对那个有兴趣吗?
63:59
Or are there other
还是有其他
64:00
kinds of tests that we're
有哪些测试类型是我们
64:01
not thinking about?
没想到的?
64:02
Yeah.
对。
64:03
I think Mechin Terp
我觉得 Mechin Terp
64:04
is a big one.
是个很重要的方向。
64:07
A lot of interest in that.
大家对那方面很感兴趣。
64:09
I think everyone would
我觉得大家都会
64:10
agree that there's like
同意,有那种……
64:11
promising scientific
很有前景的科学
64:13
potential.
潜力。
64:14
We're still a while
我们还得再等一阵子,
64:17
a little bit away at least
至少还差一点点,
64:19
from this being like commercially
才能让它像商业上
64:21
available on demand, such
按需可用,到了
64:23
that there's like now a selection
就像现在有各种
64:25
of vendors you can go to.
供应商可以选的地步。
64:26
Good for you, I would say
要我说,这对你来说是好事。
64:27
it is commercially available.
它已经可以买到了。
64:28
Exactly.
没错。
64:29
We just were to agree with them.
我们本来就是要同意他们的。
64:31
We think that they worked
我们认为他们成功了。
64:32
and they're doing as
而且他们做得
64:33
tremendous.
非常棒。
64:34
We're not quite at a
我们还没到一个……
64:35
point where we could like
到了我们可能,像是
64:36
literally require it.
真的要求它的地步。
64:37
But it's the kind of thing
但这是那种
64:38
where you can imagine,
你可以想象,
64:39
relatively soon, you could
相对很快,你就能
64:40
put in an optional
加一个可选的
64:41
control for if people use
控制项,如果人们使用
64:43
that can turn up as a way
那可能会作为一种方式出现
64:44
to reduce risk.
来降低风险。
64:45
You at least get credit
至少你能得到认可
64:46
for it.
为此。
64:47
We can't require it because
我们不能要求这么做,因为
64:48
it would be hard to
那会很难
64:49
require everyone to become
要求每个人都成为
64:50
good-fighter customers.
good-fighter 客户。
64:51
Well, good that's
嗯,很好,那
64:52
credit to me.
功劳算我的。
64:53
This is the pass-fail, right?
这是 pass-fail,对吧?
64:54
Do I care about credit?
我在乎功劳吗?
64:55
It's the pass-fail,
就是 pass-fail,
64:56
but it's also a 100-page
但它也是一份 100 页的
64:57
order report that you'd be
order report,你会
64:58
surprised at how much
惊讶于 security
64:59
security is actually
实际上到底有多少
65:00
sit down and digest the
坐下来,消化一下
65:01
stuff.
这些东西。
65:02
Okay.
好的。
65:03
And I promise you that if
而且我向你保证,如果
65:05
someone is using Mechin Terp
今天有人在用 Mechin Terp,
65:06
today, they will have a
他们就会有一张
65:09
slide on it.
关于它的幻灯片。
65:10
It's easy if you have a
如果你有一个,就很简单。
65:11
third party saying,
第三方说,
65:12
yep, they have Mechin Terp
对,他们有 Mechin Terp
65:15
and actually someone
而且实际上有人
65:16
first.
先做了。
65:17
Just to, like, just to
只是为了,呃,只是为了
65:18
flesh spell it out for
给大家详细讲清楚
65:19
people, people have been
人们,人们一直在
65:20
following our Mechin Terp
关注我们的 Mechin Terp
65:21
podcast.
播客。
65:22
Yeah.
对。
65:23
It is literally like, oh,
这真的就像是,哦,
65:24
you're using, you know,
你在用,你知道,
65:25
GPT OSS.
GPT OSS。
65:26
It is activating these three
它会触发这三种
65:27
dangerous things.
危险的东西。
65:28
We monitor for it,
我们会监控这个,
65:29
and we log it out in
and we log it out in
65:30
whatever tool of choice.
whatever tool of choice.
65:31
Grace One has, like,
Grace One has, like,
65:32
signal and whatever.
signal and whatever.
65:33
And that's it.
And that's it.
65:34
That's the Mechin Terp
That's the Mechin Terp
65:35
based activation signal.
based activation signal.
65:37
Okay.
Okay.
65:38
Yeah.
嗯。
65:39
So I think Mechin Terp
所以我觉得 Mechin Terp
65:40
is interesting, and I think
挺有意思的,而且我觉得
65:41
if that promise
如果那个承诺
65:42
truly comes to fruition,
真的能兑现,
65:43
you can make stronger
你就能做出比
65:45
promises than you can
用 eVals 时
65:46
with eVals.
更强的承诺。
65:47
And so I think that's very
所以我觉得这非常
65:48
compelling.
有说服力。
65:49
Another thing that I think
另一件我觉得
65:50
will become increasingly
会变得越来越
65:52
important is just kind of
重要的事,其实就是
65:54
good old-school monitoring.
那种经典的老派 monitoring。
65:55
And slightly after the fact,
而且稍微事后一点,
65:57
one of the things you're
你要做的其中一件事
65:58
seeing with eVals,
从 eVals 来看,
65:59
some of the challenges
其中一些挑战
66:00
they're diverging is that
它们正在分化的是
66:01
the agents are starting
agents 开始
66:02
to become aware that they're
意识到自己
66:03
being eVals.
正在被 eVals。
66:04
Yeah.
嗯。
66:05
Evo in this.
Evo 在这其中。
66:06
Exactly.
没错。
66:07
Which is a problem.
可这就是个问题。
66:08
It means that they basically
这基本上意味着,
66:10
is they know they're being
就是他们知道自己正被
66:11
watched.
监视着。
66:12
And they don't do the thing
于是他们就不会去做那件
66:13
that they think they get
他们觉得自己会因此
66:14
punished for.
受到惩罚的事。
66:15
And by default,
而且默认情况下,
66:17
unless you know how to
除非你知道怎么
66:18
just kind of reduce the
稍微降低
66:19
of other awareness,
其他意识的,
66:20
you should trust eVals
你就应该少信任 eVals
66:21
less.
一点。
66:22
And one of the kind of
而且其中一种
66:23
truth things,
算得上真相的东西,
66:24
monitoring is the source of truth.
monitoring 是 source of truth。
66:25
Did you, in fact,
事实上,你有没有
66:26
give medical advice
提供医疗建议
66:27
and how could you do know
而且你怎么可能知道
66:28
how often have you
你有多经常
66:29
done that in the past?
在过去做过那件事?
66:30
How fast do you respond?
你响应得有多快?
66:31
How often do you detect
你多久检测一次
66:32
it?
它?
66:33
How fast do you detect this?
你检测这个有多快?
66:34
So I think that is also
所以我觉得那也是
66:35
a paradigm.
一种 paradigm。
66:36
It's slightly more
它稍微更
66:37
intrusive.
具侵入性。
66:38
You actually will look
你实际上会去看
66:39
at some customer data.
一些客户数据。
66:42
But I think we'll become
但我觉得我们会变得
66:43
more prevalent
越来越普遍
66:44
over time.
随着时间推移。
66:45
People talk about this.
人们在聊这个。
66:46
Like, we should not
就像,我们不应该
66:47
write about Evo awareness
写关于 Evo awareness 的内容
66:48
because it's going to leak
因为它会泄漏
66:49
into the data set and then
到 data set 里,然后
66:50
beat.
beat。
66:51
Like, we should just,
就是,我们干脆就,
66:52
like, we should never
就是,我们绝对不要
66:53
talk about it.
聊这件事。
66:54
Only meet in person
只当面见
66:55
and like talk offline
然后就是线下聊
66:56
unrequited.
单相思。
66:57
Like, did you guys see
就是,你们看到
66:58
the anthropic research
anthropic 的研究
67:00
where I think this was
我觉得这大概是在
67:02
a little anthropic
有点 anthropic
67:03
did that test.
做了那个测试。
67:04
What it took?
需要什么?
67:05
I can't remember the details
我不记得那边的细节了,
67:06
there, but they
但他们
67:07
ran some studies
做过一些研究
67:09
on misalignment.
关于 misalignment。
67:10
And then they took
然后他们
67:11
out the training data
拿出了 training data。
67:12
several days to less
好几天都在 LessWrong 上
67:13
wrong discussing
讨论
67:14
the alignment.
alignment。
67:15
And they ran the same test
然后他们跑了同一个 test。
67:16
again.
又一次。
67:17
And the failure rate went
而且 failure rate 降下来了。
67:18
down.
所以,这其实是一些证据,指向它已经学到了做这件事的 probability,或者说 propensity。
67:19
So it, in fact, was
所以,事实上,它确实就是
67:20
some evidence pointing
有些证据表明
67:21
towards it had learnt
它已经学到了
67:23
the probability
这个概率
67:24
or the propensity
或者说这种倾向
67:25
to do that.
去做那件事。
67:26
Yeah.
嗯。
67:27
I mean, there's the
我是说,有那个
67:28
high position effect
high position effect
67:29
and there's like the Luigi
还有像那个 Luigi
67:30
while Luigi effect.
while Luigi effect。
67:31
Correct.
对。
67:32
Which is like, you are
这就像是,你越是
67:33
the more you try to train
你越是想去 train
67:34
for it, you create the
为了它,你要创造出
67:35
opposite.
相反的东西。
67:36
Yes.
是的。
67:37
There you are.
这就对了。
67:38
That's exactly it.
正是如此。
67:39
The very success
正是这种成功
67:40
method is a result
方法,是
67:41
of high position, like the
高位置的结果,就像
67:42
the fact that you
你想要它
67:43
wanted this to exist
存在于这个世界上。
67:44
in the world.
不,它确实存在。
67:45
No, it does.
但与此同时它也会
67:46
But then it also
催生出相反的东西。
67:47
creates the opposite
我觉得那些
67:48
as well.
也一样。
67:49
I think people who are
我觉得那些人是
67:50
maybe newer to this space
可能刚接触这个领域
67:51
don't remember
不记得了
67:52
while Luigi, but I do
而 Luigi 呢,但我确实
67:53
think it's very, very
觉得这非常、非常
67:54
important for understanding
重要,对理解来说
67:56
that when you train
就是当你 train
67:57
for a thing, you also
一个东西时,你也会
67:58
train the opposite
train 出相反的东西
67:59
of the thing.
of the thing.
68:00
Because it's just a bit
因为这只是有点
68:01
flute.
长笛。
68:02
Yes.
是的。
68:03
Yes.
是的。
68:04
I think, you know,
我觉得,你知道,
68:05
just going back to where
就像回到我们
68:06
we were at, like, there's
之前聊到的地方,就像,有
68:07
a lot more than just
远不止
68:08
that.
这一点。
68:09
There's value in just
光是加进去
68:10
having added, right?
就有价值,对吧?
68:11
So your version of
所以你觉得
68:12
how fast can you measure
你能多快测量
68:13
stuff?
东西?
68:14
Do you have logging
你们有 logging 吗?
68:15
and give evals?
然后给 evals?
68:16
Do you see other parts
你还会看到 stack 的其他部分吗?
68:17
of the stack?
比如你用的 inference provider set,那些 services?
68:18
Like the inference
好。
68:19
provider set you use
我是在用中文吗?
68:20
the services?
服务本身?
68:21
Okay.
好。
68:22
Am I using Chinese
我是在用中文吗
68:23
model on their home API?
model 是在他们自家的 API 上吗?
68:24
Am I using
我是在用
68:25
through certified
通过认证的
68:26
vendor here?
vendor 这边吗?
68:27
Am I hosting
我是在自己 hosting
68:28
myself?
吗?
68:29
What am I doing on the
我到底在
68:30
inference and genocide?
inference 和种族灭绝上做什么?
68:31
There's just like so
就有,就像,这么
68:32
many levels of stuff
多层的东西
68:33
that gives information
能提供信息
68:34
that you can
你可以
68:35
standardize out, right?
standardize 掉的,对吧?
68:36
Yes.
是的。
68:37
With the,
有了这个,
68:38
you know,
你知道的,
68:39
we are actually
我们实际上
68:40
increasingly in addition
越来越,除了
68:41
to just the basic
基本的
68:42
chap us, you're increasingly
chap us 之外,你越来越
68:44
seeing big companies adopting
看到大公司采用
68:46
agent platforms, where they're
agent platforms,他们正在
68:48
building on top of Google's
基于 Google 的
68:50
agent studio.
agent studio 来搭建。
68:51
It comes with a bunch
它自带一堆
68:52
of like management.
像 management 之类的东西。
68:53
Management.
Management。
68:54
Everyone has
每个人都有
68:55
management.
management。
68:56
Exactly.
没错。
68:57
And there's even
而且甚至还有
68:58
levels you can host your own
层级,你可以自己 host 自己的。
68:59
management.
管理。
69:00
So, open your
所以,打开你的
69:01
agent SDK or hosted by
agent SDK,或者由
69:02
nthropec or Google does both.
nthropec 或 Google 托管,两家都做。
69:03
Correct.
没错。
69:04
These are just ways to strengthen the security guarantees you can make, and in some ways,
这些只是加强你能做出的 security guarantees 的方式,而且在某些方面,
69:10
it's a kind of bread and butter enterprise security.
这算是 enterprise security 的基本盘。
69:15
They love to host things on their own promises because they give them a really sensitive control.
他们喜欢把东西托管在自己的承诺上,因为这会给他们一种非常敏感的控制权。
69:21
I think you'll see just like you do in every other enterprise market.
我觉得你会看到,就像你在其他所有企业市场里看到的一样。
69:24
If you really sell to the enterprise, you start to compete in some of these security features,
如果你真的是卖给企业,你就会开始在一些这样的 security features 上竞争,
69:28
and this is also happening AI unsurprisingly.
而毫不意外,这在 AI 领域也在发生。
69:31
I think you're seeing some amount of enterprises wanting, and suppose I'm really grappling
我觉得你看到的是,有一些企业想要它,而我想,我真正在纠结的是,
69:36
with the thing that makes agents useful is their stochastic, and the thing that makes
让 agents 有用的,是它们的 stochastic 性;而让
69:41
them really hard to adopt is their stochastic, and these are just intentions.
它们真的很难被采用的,也是它们的 stochastic 性,而这些只是意图。
69:46
Leaders come out on different sides at that table.
在那张桌子上,领导者们会站到不同的立场上。
69:49
In part, depending on how much the CEO is trying to get the stock price to grow up by
这部分取决于 CEO 有多想让股价涨上去
69:53
saying they're AI native, and they're very much willing to take the risks.
说他们是 AI native,而且他们非常愿意承担这些风险。
69:56
Well, you see, we actually see phenomenal attention in the heads of the CSOS of the Fortune
嗯,你看,我们其实在 Fortune 1000 的 CSOS 脑子里看到极大的关注:一方面,CEO 会说,我们必须采用,否则我们就会变得无关紧要;而要是我们搞砸了,你就会被炒。
70:00
1000, where on the one hand, you have a CEO saying, we must adopt otherwise we're becoming
而这正是我们一次又一次看到的、核心的情绪关注;我们为他们解决的核心问题之一,就是在某种程度上把这种抽象的情绪担忧转变成一个 framework,给这种担忧带来一些自豪感和清晰度。
70:04
irrelevant, and if we fuck up, you're fired.
所以,这里面还有什么吗,我觉得我们好像漏掉了什么。
70:08
And that's the core emotional attention that we see showing up again and again and again,
而这就是我们一次又一次、一次又一次看到的核心情感关注点,
70:13
and one of the core problems that we solve for them is to take that abstract emotional
而我们要帮他们解决的一个核心问题,就是把这个抽象的
70:18
concern and turn it into a framework in some ways, just pride and clarity to that concern.
情感顾虑转化成一种 framework,某种程度上,只是给这个顾虑带来骄傲和清晰度。
70:23
So anything in here, something I think we skipped over.
所以这里面还有什么,我觉得我们漏掉了某个东西。
70:26
We talked a lot about agent language model, skipped over world models.
我们聊了很多关于 agent language model 的东西,但跳过了 world models。
70:32
You guys have voice, which is interesting with 11 labs, about generative media.
你们有 voice,这个跟 11 labs 一起看挺有意思的,是关于 generative media 的。
70:36
So generating images, videos, that's a category that actually has a lot of usage.
所以生成图像、视频,这其实是一个使用量很大的类别。
70:41
Is there anything in your current policy, is it separate policy, how do you see that
你们现在的 policy 里有没有相关的内容,是单独的 policy 吗,你怎么看那个领域?
70:46
space?
嗯。
70:47
Yeah.
就像,我们也稍微聊过版权或者音乐之类的。
70:48
It's like, we did talk a bit about copyrights or music as well.
嗯。
70:51
Yeah.
嗯。
70:52
A lot of the concerns that come up there either relate to copyright or there's a lot
那里冒出来的很多担忧,要么跟版权有关,要么很多都跟——我们就宽泛地叫它 safety 吧——有关。
70:58
related to, let's call it broadly safety.
所以比如这可能是 not safe work,或者就是非常露骨的内容,大概算是我们做过一些工作的核心事项之一。
71:01
So like this could be not safe work or just very graphic materials, kind of some of
standard 里也有一小部分明确涉及到那个视频。
71:06
the core things we have done some work on this.
我们在这方面还没做太多。
71:09
There's a little bit in the standard as well that deals explicitly with that video.
而且我觉得,要达到真正的 production——尤其是真正的 production,但要脱离 human in the loop——那还有一段路要走。
71:15
We've not done a lot in yet.
我们在这方面还没做太多。
71:17
And I think for proper production, that has still, especially proper production, but out
而且我觉得,对于真正的 production,这仍然,尤其是真正的 production,但去掉
71:23
of human in the loop, that's still got some ways to go.
human in the loop 的话,这仍然还有一段路要走。
71:25
It's obvious that it's coming, but it's very rare that it's like one shot deploy a video
很明显这是必然的,但像这样一次就把视频 deploy 到互联网上,还是很少见。
71:30
to the internet.
但最终这也发生了。
71:31
But eventually that also happened.
我们看到,你知道,Luma 有 Luma agent,或者说它还是挺 human in the loop 的。
71:33
We see like, you know, Luma has Luma agent or it's still pretty human in the loop.
嗯。
71:37
Yeah.
而随着技术逐渐成熟,这完全说得通:它会变得好到人们不想因为有 human in the loop 而拖慢速度。
71:38
And that just makes complete sense as the technology mature as an overtime, it will become
然后,需要做出承诺这件事就会崩掉。
71:42
so good that people will not want to slow things down by having a human in the loop.
好到人们不会想因为还要有 human in the loop 而放慢节奏。
71:46
And then the need to make promises will broke.
然后,需要做承诺这件事就会崩掉。
71:52
Why not just have prediction markets and everything, right, it's very EA adjacent.
为什么不干脆搞 prediction markets 之类的呢,对吧,这跟 EA 很沾边。
71:56
Yes.
对。
71:57
The core thing is that the people, prediction markets rely on public information.
核心问题是,prediction markets 依赖的是公开信息。
72:04
There is not a lot of public information, it's just inside it's trading on each side.
公开信息并不多,基本就是内幕,各方各自在交易。
72:09
That's illegal.
那是违法的。
72:10
There's leaked information, there's leaked information.
有泄露的信息,有泄露的信息。
72:13
The core challenge is that often you have private sensitive information and you need to
核心挑战在于,你常常有私密、敏感的信息,而你需要
72:19
convey confidence and trust around that.
围绕这些信息传递出信心和信任。
72:23
And you can of course, for some claims like, can any model be jail broken?
当然,对某些说法,比如,任何 model 都能被 jail broken 吗?
72:29
You could rely on public evidence because there will be lots of people being like, well,
你可以依赖公开证据,因为会有很多人会说,呃,
72:32
there's tons of studies and actually they all can, so that resolves fine.
有大量研究,而且实际上它们全都能,所以这个就好解决了。
72:35
I think that's good for, hey, there's new unreleased mythos model.
我觉得这对那种情况是好事,嘿,有个新的、还没发布的 mythos model。
72:41
How capable is it actually, prediction markets have not a lot to say because actually
它实际上有多强,prediction markets 也说不了太多,因为实际上
72:45
just no one knows.
就是没人知道。
72:46
And I think that's the core place where some of this breaks down is that actually lots
而我觉得,这正是其中一些东西会崩掉的核心地方,就是实际上
72:50
of the world's information that guys, some of these high level decision is private and
世界上很多信息,各位,其中一些高层决策是私密的,而且
72:54
often also just not known.
而且往往也根本不知道。
72:55
I think there's thing with prediction markets that people like is, it's not answering
我觉得 prediction markets 有个大家喜欢的地方,就是它不回答
72:59
the broad question.
那个宽泛的问题。
73:00
It's a specific rate, so will a model do this by this day or is a model capable to do this
它是一个具体的概率,所以某个模型会在某个日期之前做到这件事吗,或者某个模型有没有能力做到这件事
73:06
by then, right?
到那个时候,对吧?
73:07
So that's a little distinction.
所以这是个小小的区别。
73:10
Yeah.
对。
73:11
And often the most interesting question, if you're, say, the head of security at a bank,
而且通常最有趣的问题,如果你是,比如说,一家银行的安全负责人,
73:16
the question you're really trying to answer is, will this product, this agent do this bad
你真正想回答的问题是,这个产品、这个 agent 会不会做出这种不好的
73:21
thing that maybe primarily I care about, specifically in the setting that I care about?
事情——也许是我主要关心的,具体来说,是在我关心的那个 setting 里?
73:27
And the question is like, what's the closest that information may not exist anywhere?
而问题就像是,最接近的信息是什么,而这种信息可能哪里都不存在?
73:30
So prediction markets aggregate existing information, this information may not exist and you
所以 prediction markets 聚合的是已经存在的信息,而这种信息可能并不存在,而你
73:35
want some very specific and you're willing to pay for it.
想要一些非常具体的东西,并且愿意为它付费。
73:38
It's kind of where a third party audit comes in.
这大概就是 third party audit 发挥作用的地方。
73:40
We also don't really use prediction markets to figure out whether public companies have committed
我们其实也不会用 prediction markets 来判断上市公司有没有
73:45
fraud on their books, use audits.
在账目上搞 fraud,而是用 audits。
73:48
You probably could, but the information is not available and if so, it would be like just
你大概能,但信息拿不到,而且就算有,那也就像是只是在
73:53
trading on bias.
拿偏见来做交易。
73:54
I actually would have been really interesting to see whether prediction markets, 2001 were
其实那会非常有意思,看看 prediction markets 在 2001 年有没有
73:59
predicted and run going bankrupt and they kind of could you have told, could you have
预测到 Enron 会破产,而且它们有点像——你能看出来吗,你能不能
74:03
a sense from the like the the craziness of the CEO or some other trade that they're
从比如 CEO 的疯狂,或者某些其他交易里感觉到,他们
74:08
more likely to cut their books and others or enough insiders leak it than they're you
更可能做假账,或者别人也会,或者足够多 insiders 泄密,而不是他们让你
74:12
in.
参与进去。
74:13
That could also be right.
那也可能是对的。
74:14
Right.
对。
74:15
Right.
对。
74:16
Which is like, I mean, this is that that's the sort of the ideal dream of prediction
这就像,我是说,这就是 prediction markets 那种理想中的梦想。
74:17
markets.
你有 liquid markets,什么都有。
74:18
You have liquid markets and everything.
嗯。
74:19
Yeah.
然后你就可以组合出你确切要 offset 的那组风险。
74:20
And then you can compose your exact set of risks to offset.
是的。
74:24
Yes.
是的。
74:25
Right.
对。
74:26
Yes.
是的。
74:27
Yes.
是的。
74:28
Yes.
是的。
74:29
Yeah.
嗯。
74:30
And I think like prediction markets will bring lots of new information to it.
而且我觉得,prediction markets 会给它带来很多新的信息。
74:30
And so the thing is mostly not like which one is it and more like what are the types of
所以重点其实主要不是“它到底是哪一个”,而更像是“有哪些类型的
74:34
questions that prediction markets are really good at and what are the ones where the information
问题,是 prediction markets 真正擅长的,以及哪些问题是信息……
74:38
doesn't even exist inside us such that no one can in fact trade on it and used to get
甚至根本不存在于我们内部,以至于事实上没人能拿它来交易,而且过去常常会被
74:42
generated.
生成出来。
74:43
Okay.
好。
74:44
One self-serving question and then one open ended one on like the future of AI you see.
先问一个有点私心的问题,然后再问一个开放式的,比如说关于 AI 的未来,你懂的。
74:49
Self-serving question would be you so you have your standard.
有点私心的问题会是:你,所以你有你自己的标准。
74:51
Yes.
是的。
74:52
Right.
对。
74:53
I run, you know, a large AI engineer conference like there's been a lot of talk about us
我运营着一个,你知道,大型的 AI 工程师大会,就像一直有很多关于我们的讨论
74:56
certifying AI engineers, training programs, level one, level two, level three.
认证 AI engineers、培训项目、level one、level two、level three。
75:00
I was a CFA myself.
我自己就是 CFA。
75:02
So I know what that's what the finance industry does.
所以我知道,金融行业就是这么做的。
75:04
Yes.
是的。
75:05
Would it help if we I had AI engineer level one level two level three and then it would
如果我们有 AI engineer level one、level two、level three,会有帮助吗,然后它就会
75:10
they would like work with these guys.
他们就会愿意跟这些人合作。
75:11
I don't know.
我不知道。
75:12
If you think of the highest level objective as like accelerating secure deployment of agents,
如果你把最高层级的目标看成是加速 agents 的 secure deployment,
75:18
then that would totally help one of the things that happens often now is that folks build
那绝对会很有帮助;现在经常发生的一件事就是,大家会构建
75:22
agents, they bring it to the to the decision maker and the decision maker surface a bunch
agents,然后把它拿给决策者,而决策者会抛出一堆
75:28
of security considerations that they've not thought of and now it's not built to spec.
他们之前没想到的 security 方面的考虑,结果现在它就没按 spec 来做。
75:32
Now you have to go and re like add these filters etc.
现在你就得回去,重新加上这些 filters 之类的。
75:36
So if you shifted that left like if everyone knew what the spec they were building to everyone
所以如果你把这个左移,比如如果每个人都知道自己要照着什么 spec 来构建,每个人
75:40
knew the grading scheme.
都知道 grading scheme。
75:41
Yeah.
是啊。
75:42
That would be awesome if they're already trained so if I default you're the grading scheme
那会很棒,如果他们已经被训练过了,所以如果我默认你就是 grading scheme
75:45
right.
对。
75:46
I don't get to set the grid.
我没权设定 grid。
75:47
You set the grading scheme.
你来定评分方案。
75:48
We set the grading scheme and I think what's valuable is like if you can turn this
我们来定评分方案,我觉得有价值的是,如果你能把这个
75:51
into training programs.
变成培训项目。
75:53
Training programs.
培训项目。
75:54
Yeah.
是啊。
75:55
Which you're not doing.
但你现在没这么做。
75:56
We're not doing that.
我们不会那么做。
75:57
I think there's value in doing it.
我觉得这么做是有价值的。
75:58
There are others doing.
也有其他人在做。
75:59
I mean not to interrupt interrupt.
我是说,不是要打断,打断。
76:00
Yeah.
对。
76:01
Yeah.
对。
76:02
Yeah.
对。
76:03
Yeah.
对。
76:04
Yeah.
是啊。
76:05
And they want they want 100,000 deployed certified consultants.
而且他们想要 100,000 名已 deployed 的认证顾问。
76:09
Right.
对。
76:10
I think it's good for we will accelerate adoption if we have more people who know how to
我觉得这是好事,因为如果我们有更多人知道怎么
76:14
build secure agents and we are not working on this out of training people at the moment.
构建 secure agents,我们就能加速采用,而且我们目前既没在做这个,也没在培训人。
76:20
I think it's like very aligned with our mission.
我觉得这跟我们的使命非常一致。
76:22
We only have so much attention.
我们的注意力就这么多。
76:24
I tell you why I haven't done it.
我告诉你我为什么还没做这件事。
76:26
Yeah.
嗯。
76:27
It's not like I haven't thought about it before.
也不是说我以前没想过。
76:28
Yes.
对。
76:29
It's just being prescriptive about like, well, this is what you should know therefore like
它只是在很规定性地说,呃,这就是你应该知道的,所以呢,就像
76:33
the stuff that didn't include is what you don't need to know.
那些没包含进去的东西,就是你不需要知道的。
76:35
Yes.
对。
76:36
I'm like, that sucks.
我就觉得,这太烂了。
76:37
Yes.
对。
76:38
Yeah.
对。
76:39
I think it's like, you know, the very interesting, defensible thing you guys do is your opinionated
我觉得就是,你知道,你们做的一件非常有意思、也很有防御性的事,就是你们那份 opinionated
76:44
100 page report of here's what matters, right?
100 页报告,讲的是什么才是重要的,对吧?
76:48
Here's the like prescriptive definition of the requirements you need to be certified.
它就像是那种规范性的定义,说明你需要满足哪些要求才能获得认证。
76:54
So yeah.
所以,对。
76:55
And I think that's a choice.
而且我觉得那是一种选择。
76:56
I think recently that's a, that's a choice and I think that serves some audiences very
我觉得最近这是一种,这是一种选择,而且我觉得这能非常好地服务某些受众
77:02
well where if you're trying to deploy this into a bank or a hospital, et cetera, clarity
尤其是当你想把它 deploy 到银行或医院等等地方时,清晰度
77:06
of the boundaries is extremely valuable.
对边界的探索是极有价值的。
77:09
There's lots of other settings where being much more experimental, much more trying it
还有很多其他场景,更实验性一点、更多去尝试
77:13
out is just a benefit.
本身就是一种优势。
77:15
And so to me, this makes a ton of sense also it has to rewrite your curricula every
所以对我来说,这完全说得通,而且它还得每隔
77:18
freaking three months.
该死的三个月就重写一遍你的 curricula。
77:20
It's fine.
这没问题。
77:21
I do that.
我就是这么做的。
77:22
Makes it okay.
这样就可以了。
77:23
But yeah.
不过,对。
77:24
No, for me it's actually like genuinely like the consequences of getting it wrong and
不是,对我来说,这其实真的就像,搞错的后果,还有
77:31
like affecting somebody's career is is a big responsibility.
像是影响到某人的职业生涯,这真的是一个很大的责任。
77:34
Yeah.
对。
77:35
Yeah.
对。
77:36
I think that's exactly right.
我觉得这完全正确。
77:37
And I think a lot of our work actually goes like, we don't want to carry, we also don't
而且我觉得我们很多工作其实都是在说,我们不想承担,我们也不
77:42
think I was able to carry the kind of the true north of what's like secure and not secure.
觉得我能扛起那种,什么算安全、什么不算安全的真正方向。
77:48
But we can coordinate the forum where you list it all of that.
但我们可以协调一下这个论坛,让你把这些全都列在那里。
77:52
Yeah.
对。
77:53
You can sort of just can be crowdsourced to know it like for your example for what is
你可以说,就像你举的例子那样,这件事其实靠众包就能搞清楚,比如什么是
77:57
the engineer certification.
工程师认证。
77:58
I mean, this is a pretty big podcast.
我是说,这个播客还挺大的。
78:00
There's a lot of takes that people can have and you know, discussions that can, and
大家会有很多不同的观点,你知道,也会有很多讨论,而且
78:05
people reasonably disagree.
人们完全有理由不同意。
78:06
So who am I to say?
那我又有什么资格说呢?
78:07
Like that's a correct question.
就像,那才是一个正确的问题。
78:08
That's a wrong question.
那是个错误的问题。
78:09
Yeah.
嗯。
78:10
Right.
对。
78:12
Yeah.
嗯。
78:13
That you're first racing.
也就是说,你是最先起跑的。
78:15
Exactly.
没错。
78:16
And I think there's also you, it matters a lot about the promises.
而且我觉得,还有你,承诺这件事很重要。
78:22
So if the promises, hey, if you've taken my course, you will not fuck up.
所以如果说到承诺,嘿,如果你上过我的课,你就不会搞砸。
78:27
You can't make that promise clearly.
你没法明确地做出那种承诺。
78:30
You could make a promise of like here's the some important things that everyone should
你可以做出这样的承诺,比如:这些是一些重要的事情,每个人都应该
78:34
at least know.
至少知道。
78:35
And then you have to fill out the rest there, at least the promise changes.
然后剩下的部分你得补上,至少这个承诺就变了。
78:38
Of course, there's some subtlety in how do you communicate this stuff so that people
当然,这里面有一些微妙之处,就是你怎么把这些东西传达出去,好让人们
78:40
really get it.
真正明白。
78:41
But I think it's important to dial in and we have a section in our, stand on like, what
但我觉得把它调准很重要,而且我们有一个部分,在我们的,stand on,比如,什么
78:46
is the promise and what is the promise not?
那什么是承诺,什么又不是承诺?
78:48
Because it's impossible to guarantee that nothing will go wrong.
因为不可能保证什么事都不会出错。
78:51
If you need a guarantee that nothing will go wrong, you cannot work with Frontier AI,
如果你需要保证什么事都不会出错,就没法和 Frontier AI 合作,
78:54
but you can make some claims.
但你可以做出一些声明。
78:56
Yeah.
嗯。
78:57
For sure.
当然。
78:58
Cool.
好。
78:59
Wanted to end with OpenEnded.
想以 OpenEnded 收尾。
79:00
Where is AIUC going?
AIUC 接下来要往哪儿走?
79:01
I think you talked about model stuff, robotic stuff.
我记得你聊过 model 那些东西、机器人那些东西。
79:05
I just opened ended like, you know, what is, what is in the future for you guys?
我就是开放式地问,你懂的,你们未来会怎么样?
79:10
During your term, we've now started to work with some of the Frontier companies in
在你任期内,我们现在已经开始就那些正在起飞的类别,和一些 Frontier 公司展开合作。
79:14
the issue of the categories that are taking off.
我们会继续这项工作,确保覆盖所有真正在起飞的 use cases。
79:17
And we'll continue that work to make sure that we cover all of the use cases that are
一旦市场上第一个动起来,我们就会看到很多兴趣。
79:21
really taking off.
真的开始起飞了。
79:22
We see a lot of interest once the first one in the market moves.
一旦市场上第一个行动了,我们就会看到很多兴趣。
79:27
Lots of people want to follow them.
很多人想跟着他们。
79:28
And we think basically AC1 will get to a point where all of the Fortune 1000 will organize
而且我们觉得基本上,AC1 会发展到这样一个阶段,所有 Fortune 1000 公司都会围绕这个标准来组织
79:35
their risk processes around the standard.
他们的风险流程。
79:38
Then you have 50%.
然后你就有 50% 了。
79:39
No.
不。
79:40
I think there is some world where, probably by end of year, we might have representation
我觉得存在某种情况,可能到年底,我们可能会有代表席位
79:46
in our consortium for 50% of the Fortune 1000.
在我们的联盟里,代表 Fortune 1000 中的 50%。
79:48
So that's on the Asian layer, and then we think the model layer is going to be, it just
所以那是在 Asian layer 上,然后我们认为 model layer 将会是,它只是
79:53
brings, are now suffering the concerns that are most likely to slow down adoption of AI.
带来的,现在正被那些最有可能拖慢 AI 普及的担忧所困扰。
79:59
And then we think robotics comes after that.
然后我们认为,robotics 会紧随其后。
80:03
What are you hiring for?
你们在招什么岗位?
80:04
What's hard to hire for?
哪些岗位很难招?
80:05
We are hiring across the board, across go to market, and a number of technical staff.
我们各个方向都在招人,包括 go to market,还有不少技术岗。
80:10
The people who do really well on our technical team are folks who are really excited about
在我们技术团队里表现特别好的人,都是那些真的对
80:15
kind of being truly full stack.
有点像真正做 full stack 这件事感到兴奋的人。
80:16
So let's say when we started working with cursor, we had never done coding tools before.
比如说,我们刚开始和 cursor 合作的时候,之前从来没做过 coding tools。
80:24
So taking the standard and extending it, fleshing out, what does Frontier EVAs look
所以,把这个 standard 拿来扩展、充实,Frontier EVAs 对 long horizon coding agents 来说会是什么样子,并且把这个问题从跟 cursor 以及这个领域里的其他人合作,一直推进到去充实和塑造一个新版本的 standard。
80:29
like for long horizon coding agents, and taking that problem all the way from working
所以这就像是真正的 full stack、创业者型技术人做得极其出色的事情。
80:33
with cursor and other folks in the space, down to like fleshing out and shaping a new
而且如你所见,最难的部分是打造一个通用的 red teamer,能跨 Harvey 到 cursor 以及其间的所有人。
80:37
version of the standard.
一套一致的方法论,一套一致的 taxonomy,来定义 attacks 中都有哪些风险,
80:39
So that's like a truly a full stack entrepreneur-shaped technical people do extremely well.
所以这就像是真正的 full stack、创业者型的技术人员能做得极其出色的事情。
80:43
And as you see, the hard part is building one universal red teamer that works across
而如你所见,最难的部分是打造一个通用的 red teamer,能够覆盖
80:52
from Harvey to cursor and everyone between the bus.
从 Harvey 到 cursor,以及这中间的所有人。
80:54
One consistent methodology, one consistent taxonomy of what are the risks in the attacks,
一套一致的方法论,一套一致的 taxonomy,来界定攻击中的风险有哪些,
81:00
and making, we think that's fundamentally the best way to make consistent promises.
而且在做的时候,我们认为这从根本上就是做出一致承诺的最好办法。
81:04
Jim Morgan is buying both.
Jim Morgan 两个都要买。
81:06
They want to have one framework, one consistent way that this comes out, and the mechanics
他们想要有一个 framework,一种一致的产出方式,以及实现这件事的机制
81:09
of making that happen, you get to deal with a lot of the complexity of the real world.
而要做到这一点,你就得处理现实世界里的很多复杂性。
81:13
I think we have good answers in a bunch of that, but there's some pretty hard engineering
我觉得在很多这方面我们都有不错的答案,但也有一些相当难的工程
81:16
problems and the next and queuing on that.
问题,还有接下来和 queuing 的那些。
81:18
Can I push a little bit?
我能稍微追问一下吗?
81:19
Like, must you have one?
就是说,你们必须得有一个吗?
81:21
Why not just be like, okay, look, 40% of our use cases are coding agents.
为什么不干脆就说,好吧,你看,我们 40% 的 use cases 都是 coding agents。
81:26
So we would specialize in coding agents and that's the one with them.
那我们就专攻 coding agents,而那就是跟他们一起的那个。
81:28
Yes.
是的。
81:29
20% is like, rag.
20% 是像 rag 这样的。
81:30
Yes.
是的。
81:31
Just do rag.
就做 rag。
81:32
Yes.
是的。
81:33
I think there's some wisdom in that question.
我觉得这个问题里有点智慧。
81:38
It depends on what we found that there's a lot of value on is being able to, if the decision
这取决于我们发现的是,很多价值在于,如果买方的决策者——比如你说你是银行的风险主管,而你最大的风险是 non-encoding,或者在客服,或者别的什么,但前两大 use cases 却在别的地方,你还是想确保那个 framework 能对你那个最迫切的问题说点什么,否则你就赢不到那份信任。
81:44
maker on the buying side, saying you're the head of risk at a bank, and your biggest
现在,确实,很多最迫切的问题都出现在 adoption 很多的地方。
81:49
risk is non-encoding or in customer support, or whatever, the top two biggest use cases
所以,很好。
81:54
but is somewhere else, you want to still make sure that that framework has something
我们也一样。
81:57
to say about it to the burning question you have, otherwise you'll not earn that trust.
能就这件事,对你那个迫切的问题给出说法,否则你就赢不到那份信任。
82:00
Now, it's true that a lot of the burning questions follow where there's a lot of adoption.
现在,确实,很多迫切的问题都出现在采用度高的地方。
82:06
And so great.
那很好。
82:07
So do we.
我们也是。
82:08
So we do today do not cover every single edge, but we have a framework that we can add all
所以我们现在并没有覆盖每一个 edge,但我们有一个 framework,可以把所有这些
82:14
of these within.
都纳入其中。
82:16
We have one global taxonomy of risks and attacks.
我们有一套全球统一的 taxonomy,涵盖 risks 和 attacks。
82:19
That keeps adapting.
它一直在不断适应。
82:21
Like every time a new incident occurs that there's never been seen before, great.
就像每当出现一个以前从未见过的新 incident,太好了。
82:24
Let's go and update the taxonomy so we bake that in.
那我们就去更新 taxonomy,把它固化进去。
82:27
So I think we have one coherent universal approach.
所以我觉得我们有一套连贯且通用的方法。
82:31
It doesn't mean that we spend equal amounts of time on code and in certain niche use case,
但这并不意味着我们在 code 上、在某个 niche use case 上花同样多的时间,
82:37
we do spend time where people care.
我们确实会把时间花在人们在意的地方。
82:40
We think it's very valuable to have one language.
我们认为,能有一种共同语言是非常有价值的。
82:43
Yeah, yeah, it makes sense.
对,对,这有道理。
82:45
That's a, that's an important choice.
那是一个,那是一个重要的选择。
82:47
We were going to end actually by, I thought of one final ending closing question, which
我们其实本来打算以——我想到了最后一个收尾的结束问题——来结束,那就是
82:51
is take this however you want, right?
就是,你怎么理解都行,对吧?
82:53
Let's say one and a half years from now, opening a secret panel of five experts declares
假设一年半以后,开场时,一个由五位专家组成的秘密小组宣布
82:58
that we have reached HGI.
我们已经达到了 HGI。
82:59
Do you expect your business to change?
你预计你的业务会改变吗?
83:03
No.
不会。
83:04
I think there's some important way.
我觉得有某种重要的方式。
83:07
I think the, the last businesses to exist beyond the labs will be underwriting.
我觉得,那个,那个,在 labs 之外最后还会存在的业务会是 underwriting。
83:11
Well, there's one, there's one job that the labs can never do for themselves, which
嗯,有一个,有一个工作是 labs 永远没法为自己做的,那就是
83:18
is to be their own watchdog.
做自己的监督者。
83:20
There you go.
这就对了。
83:21
So I think it could kind of, to the extent that you believe this frame of like, you'll
所以我觉得它可能会有点,取决于你在多大程度上相信这种 frame,就是,你会
83:24
see hyper concentration, like the labs will kill the startups, which we can go into with
看到 hyper concentration,比如 labs 会干掉 startups,这个我们可以深入聊,包括
83:31
the pros and cons.
利弊。
83:32
I feel like the labs actually care a lot about this, right?
我觉得 labs 其实很在意这件事,对吧?
83:34
There's a whole superposition.
这里有一整套 superposition。
83:35
What do we do?
那我们怎么办?
83:36
And we have model smarter than us, and the tier above, right?
而且我们有比我们更聪明的 model,还有再上面那一层,对吧?
83:39
Model smarter than them training themselves.
比它们更聪明的 model 自己 training 自己。
83:41
The labs actually think about this all.
labs 其实都在考虑所有这些。
83:43
They think a lot about, I think there are some of the smartest people on these topics work
他们对这些事想得很多;我觉得,在这些议题上,有些最聪明的人就在这些实验室工作。
83:46
at the labs.
所以问题不在于他们是否在乎。
83:47
So the problem is not whether they care.
问题在于,他们都会被困在一场竞赛里,可能因此有动机偷工减料,也可能有动机向政府隐瞒信息,等等。
83:50
The problem is that they will all be stuck in a race, where they might have incentive
所以,有一种感觉像是永恒真理:你需要一个独立的第三方去检查那些数据,并分享信息。
83:54
to cut corners, and they might have incentive to withhold information from the government,
走捷径,而且他们可能有动机对政府隐瞒信息,
83:59
et cetera.
等等。
84:00
And so one kind of feels like eternal truth is that you need an independent third party
所以,某种程度上感觉像是一条永恒真理:你需要一个独立的第三方
84:04
to go and inspect that data and share information.
去检查那些数据并分享信息。
84:07
In this case, say with the government, it's more of an incentive problem than an interest
在这种情况下,比如说和政府那边,这更多是个激励问题,而不是利益问题。
84:11
problem.
我觉得从根本上说,他们都是想让这件事顺利推进。
84:12
I think they're fundamentally all trying to make this go well.
但我没听到的是,像 AGI,不管这个标签对你、对我、对他们意味着什么,
84:14
But what I'm not hearing is like AGI, whatever that label means to you, to me, to them,
从根本上并没有那种构成上的质变。
84:20
doesn't fundamentally have like a qualitative shift in make.
你还是得做。
84:24
You still have to.
而我觉得,唯一能让这成为质变的一件事,就是有一些定义
84:25
And I think the one thing that would make this a qualitative shift is there's some definitions
关于 AGI 的。
84:30
of AGI.
AGI 的。
84:31
It will just get nationalized.
它只会被收归国有。
84:32
It would be a threat to sovereignty.
那会对主权构成威胁。
84:34
Yes.
是的。
84:35
And at that point, it kind of maybe every company is the government, the government is
而到了那个时候,可能就有点像是,每家公司都是政府,政府就是每家公司。
84:39
every company.
我很想认真思考那个世界,但到了那个时候,你已经有点……我觉得我们不会行动得足够快。
84:40
I strongly was to think about that world, but at that point, you've kind of, I don't
对。
84:43
think we'll move fast enough.
觉得我们会推进得足够快。
84:44
Right.
对。
84:45
You know, we're not set to do that, but I have discussed this a lot.
你知道,我们本来没打算那么做,但这件事我聊过很多。
84:50
On the podcast.
在播客里。
84:51
Yeah.
对。
84:52
Yeah.
对。
84:53
I mean, you know, as far as the watchdog concern, I will also mention that because I have my
我是说,你知道,就监管方面的担忧来说,我还要提一下,因为我自己有
84:57
finance background, I often think about the scene in the big short where they talk to
金融背景,所以我经常会想到 the big short 里的那个场景,他们在跟
85:01
like moody's, but also standard and poorest.
比如 moody's,还有 standard and poorest 说话。
85:04
And then the lady and moody's is like, well, if I don't give you triple A rating, you're
然后 moody's 的那位女士就说,嗯,如果我不给你 triple A rating,你
85:07
just going to go down to standard and poor thing to it.
就只会一路降到 Standard and Poor 那种东西上去。
85:09
Yeah.
嗯。
85:10
So, so actually the watchdog is a natural monopoly because if you have race dynamics
所以,其实 watchdog 是一种 natural monopoly,因为如果你有 race dynamics
85:13
in watchdogs, then the watchdogs will compete each other to the lowest possible standard.
在 watchdogs 里,那 watchdogs 就会互相竞争,把标准压到尽可能低。
85:18
Correct.
对。
85:20
And so, I think what's one of the things, one of the reasons why we're very excited
所以,我觉得,有一点、我们非常兴奋的一个原因是
85:23
about having insurers be around this table is that insurers are the only ones that do
让保险公司坐在这张桌子旁,就在于保险公司是唯一不
85:29
not have this dynamic because they pay the bill as they keep lowering their prices.
会有这种动态的,因为他们一边不断降价,一边要买单。
85:32
Yeah.
对。
85:33
You will find them market clearing.
你会发现它们能实现 market clearing。
85:35
And this is not just moody's where they don't directly pay the bill as they make recommendations
而且这不仅仅是 Moody's 的问题,他们给出不准确的建议时,并不直接承担后果。
85:40
that are off.
所以,我们认为这个平衡因素相当重要。
85:41
So, we think that balancing factor is pretty important.
而且我觉得这也凸显出,没有什么系统是完美的。
85:45
And I think it also highlights that there's like no system that's perfect.
你需要对 Moody's 进行审查,你需要对监督机构进行审查。
85:48
You need scrutiny of moody's, you need scrutiny of the watchdogs.
确实。
85:52
For sure.
当然。
85:53
Beautiful.
太棒了。
85:54
Thank you so much for indulging.
非常感谢你愿意迁就我。
85:55
This is a beautiful conversation covering everything congrats on your success so far.
这是一场很棒的对话,涵盖了所有内容,恭喜你到目前为止取得的成功。
85:59
Thanks for having me.
谢谢你邀请我。
86:00
Yeah.
嗯。
86:01
Appreciate it.
很感谢。

Play Queue

☀️