Latent Space
Runway’s WorldPrompt and the Engineering of Real-Time Worlds
00:00
00:00
Okay, we're here with Anastasis from runway with me and Vibhu in his studio, welcome.
好,我们现在和来自 runway 的 Anastasis 在一起,我和 Vibhu 在他工作室里,欢迎。
00:07
Good to be here. Congrats on all your success and progress from runway, your opening offices
很高兴来到这里。恭喜你们在 runway 取得的所有成功和进展,你们在世界各地
00:12
all over the world. Did you envision this when you first started out?
开设办公室。你们刚起步的时候有想到会这样吗?
00:16
Not quite. I think even when we started with, I had this idea that, you know, it was more
不太算。我觉得就算我们刚开始做的时候,我就有这个想法,你知道,这更像是
00:21
a matter of when, not if, we were seeing the early-generative models of 2016-2017 and
时间早晚的问题,而不是会不会发生的问题,我们当时看到 2016-2017 的 early-generative models,然后
00:29
just extrapolating, assuming, you know, we, resolution, quality increases predictably
直接外推,假设,你知道,我们,resolution、quality 会随着时间
00:34
over time. There's going to be a point where most of content will be generated and that
可预测地提升。总会有那么一个时间点,大多数内容都会被生成,而那
00:40
was maybe the initial, the initial thesis of runway was we will need as a result of those
可能算是 runway 最初的、最初的论点:由于那些结果,我们将会需要
00:45
generative models to rethink how creative tools are made. And as we built out the research
用 generative models 来重新思考创作工具是怎么做出来的。
00:51
behind our generative models, it then became clear that they were used for far beyond
而当我们把 generative models 背后的研究搭建起来时,就很明显了,它们的用途远远不止于此。
00:55
that as well. And it is more obvious now with, like, the real-world stuff and the world
而且现在这一点更明显了,比如现实世界的东西,还有我们之后会聊到的 world models。
01:00
models that we'll talk about later. I'm just kind of curious how you go from a background
我就是有点好奇,你是怎么从像 ZokDok 这样的背景,还有你知道的 computer vision,转到 runway 的。
01:04
in, like, ZokDok and, you know, computer vision into runway. Like, take us back to that
比如,带我们回到早期的那次对话,作为 Chris,还有你知道的,你创始团队里的其他人。
01:10
early conversation as a Chris and, you know, whoever else is on your founding team.
我一直在这两个世界之间来回切换。
01:14
I was always splitting those two worlds. One was the, you know, I had my own art practice.
其中一个就是,你知道,我有自己的艺术实践。
01:19
I was making a lot of interactive art, I think, for a long time. And then on the other
我想,我有很长一段时间都在做很多互动艺术。
01:25
side, I was working in startups, and I was working as ML engineer, as a backup engineer
另一方面,我当时在创业公司工作,做的是 ML engineer,也做过 backup engineer
01:30
at different companies. I've always been interested in coding and commutation, and especially
在不同的公司。我一直对 coding 和 commutation 很感兴趣,尤其是
01:37
interesting simulation and brought it back into my early artwork as well. And at the
有趣的 simulation,而且我也把它带回到了我早期的艺术作品里。而在
01:41
same time, I was interested. The first time it has a few, right?
同一时间,我也很感兴趣。第一次它有几个,对吧?
01:45
Yes. Is there one that we should pull up? Just in case there's something that's, like,
是的。有没有哪个我们应该调出来?以防有什么东西,像是,
01:48
I just like to go down memory layer. Okay, what is this?
我就是喜欢在 memory layer 里往下走。好,这是什么?
01:51
So this was a project that I made, I think, about in 2015, where I built this software that
所以这是我在大概 2015 年做的一个项目,当时我做了这个 software,它
01:58
would give voice instructions to people in a gallery space. So it would basically coordinate
会给画廊空间里的人发出语音指令。所以它基本上就是协调
02:04
interactions between people. And so it will first give you an identity, like, you're an architect,
人与人之间的互动。所以它会先给你一个身份,比如,你是一个建筑师,
02:10
you're 30 years old, and you like sports. And then it would match you with another person,
你 30 岁,喜欢运动。然后它会把你和另一个人配对,
02:16
you have this completely generated interaction. Obviously, language models were not quite
你们之间会有这种完全生成出来的互动。显然,language models 当时还没那么
02:20
there at the time. And so it was a mix of some templates and some, like, some mark-of-change
到位。所以它是混合了一些 templates,还有一些,像是,一些 mark-of-change
02:26
generated text. And it would just completely simulate this small talk conversations between
生成的文本。然后它就会完全模拟这些闲聊对话,在
02:34
everyone in the in a gallery space. So it was always very fascinated on the one end with
一个画廊空间里的每个人之间。所以一方面,它一直非常着迷于
02:40
generative models and, like, the early machine learning work that was being out at the time,
generative models,还有,像是,当时正在进行的早期 machine learning 工作,
02:45
but at the same time, there was a separate thread of simulation and what it means.
但与此同时,还有另一条独立的线索,关于 simulation 以及它意味着什么。
02:49
Like, what can we learn about humans by creating those very simple models of their interactions
比如说,通过创建那些非常简单的、关于他们互动的 models,我们能了解到人类的什么呢?
02:55
and their behavior? Did you generate the prompts, or, you know, the 30-year-old, or whatever,
以及他们的行为?是你生成的 prompts 吗,还是,你知道,那个 30 岁的,或者随便什么,
03:00
was it you generating them, or how do you? Exactly. So the program would just generate those from,
是你生成它们的,还是你们怎么做的?没错。所以程序就会直接生成那些,从……
03:07
a lot of it would be kind of modlib style, just, you know, you have lists of different
很多都会是 modlib 风格的,就是,你知道,你有不同
03:12
professions, lists of different personalities, lists of different ages, things like that. And then
职业的列表,不同性格的列表,不同年龄的列表,诸如此类。然后
03:19
it would just combine those things together. And then maybe the next project we go is Ankanivale
它就会把这些东西组合在一起。然后也许我们接下来的项目是 Ankanivale
03:24
Ankanivode, which was one of the first projects that we built with one of my two co-founders,
Ankanivode,那是我们和我两个联合创始人之一一起做的第一批项目之一,
03:33
Chris. This was taking PixelPix HD, which was one of the early image-to-image models that
Chris。这是拿 PixelPix HD,它是早期的 image-to-image models 之一,那个
03:40
Nvidia released back in 2016 or 2017. And it was a model that would take a semantic map of a scene
Nvidia 早在 2016 还是 2017 年就发布过。它是一个模型,会拿一个场景的 semantic map,然后生成 photorealistic 的东西,我们就叫它 output 吧。显然,那还是非常非常早期的阶段。所以那些 output 的保真度不是很高,但我觉得,它是第一个能以 1K resolution 生成图像的 image generation model。而且它完全是在 self-driving datasets 上训练的。所以它支持的 semantic categories 只有,你知道,你在路上会遇到的那些东西。所以就是行人、交通标志、红绿灯、自行车、红绿灯。所以这其实是我们最早的迹象之一:我们做了这个,然后人们做出各种非常超现实的图像,比如一百万个行人、一百万个交通标志,或者像巨大的人类。这是一个迹象。
03:49
and then generate photorealistic, let's call it output. Obviously, very, very early days. So it was
然后生成 photorealistic 的,咱们姑且叫它 output。显然,那是非常非常早期的阶段。所以它
03:56
not very high fidelity outputs, but it was, I think, was the first image generation model that
并不是很高 high fidelity 的 output,但我觉得,它是第一个 image generation model,能
04:01
would generate at 1K resolution. And it was all trained on self-driving datasets. So the semantic
在 1K resolution 下生成。而且它完全是用 self-driving datasets 训练的。所以它的 semantic
04:09
categories it would support were only, you know, things you would encounter on the road. So it would
categories 能支持的,就只有,你知道,路上会遇到的那些东西。所以它会
04:14
be pedestrians, traffic signs, stoplights, bikes, stoplights. And so that was actually one of our
是行人、交通标志、红绿灯、自行车、红绿灯。所以那其实是我们最早的
04:21
first indications that we built this and people were making all this like very surreal imagery of
一个迹象,我们做出了这个,而人们在生成各种特别超现实的画面,比如
04:27
a million pedestrians or a million traffic signs or like gigantic humans. And it was an indication
一百万个行人,或者一百万个交通标志,或者像巨大的人类。而那是一个迹象,
04:34
that you could take a model that was trained on this very boring dataset, essentially, of like,
就是你可以拿一个在这个超级无聊的 dataset 上训练出来的 model,基本上,就是那种,
04:40
not that many interesting things happen when you're on the road. And then you can repurpose it and
在路上开的时候,其实不会发生太多有意思的事。然后你可以给它换个用途,然后
04:44
go very out of distribution and make something that was artistically compelling. And that was,
去做非常 out of distribution 的东西,做出一些在艺术上很有感染力的东西。而这其实,
04:50
it's a summary of the thesis of Runway in some ways that you can take the same generative models.
某种程度上,这算是 Runway 的核心论点的一个总结:你可以拿同样的 generative models,
04:54
And if you look at them from another direction, if you build interesting tools around them
而如果你从另一个方向看它们,如果你围绕它们做出有意思的工具,
04:59
and you give them to artists, they're going to do things that you don't expect.
然后把它们交给艺术家,他们就会做出你意想不到的东西。
05:02
Very cool. I like the, you know, UX of it basically. You're just giving an empty canvas,
非常酷。我喜欢它的,你知道,基本上就是它的 UX。你只是给了一块空白画布,
05:08
try whatever, do whatever. And then the other one, like, you see everyone with wired headphones,
想试什么就试什么,想做什么就做什么。然后另一个,就像,你看到每个人都戴着有线耳机,
05:12
you know, like that's the, that's a sign that it's, it's very Apple, your original ads. Yeah,
你知道,就像,这就是,这就是一个标志,说明它,它非常 Apple,你最初的那些广告。是的,
05:19
take us to today. You've been doing this for seven years at Runway. I've we got to this, you know,
带我们说到今天吧。你在 Runway 做这个已经七年了。我们走到了这一步,你知道,
05:24
like, how do we go from driving simulator data to all this? And you kind of cover the whole stack of
就像,我们是怎么从 driving simulator data 走到这一切的?而你差不多覆盖了整个 stack 的
05:32
generative media, you know, I mean, interestingly, we're almost back and, you know, we're full circle
generative media,你知道,我的意思是,有意思的是,我们几乎又回来了,而且,你知道,我们完整地绕了一圈
05:37
where we're now applying our models and going to be on creative tools into real world scenarios.
现在我们正把我们的 models 应用起来,并且要从创意工具进入真实世界场景。
05:43
But it was a, it was a long journey. It was very early on. We realized the, yeah, the first version
但那是一段,那是一段很长的旅程。那是在很早的时候。我们意识到,是的,第一版
05:49
of Runway was a way to easily use all the open source models of the day, things that pick,
Runway 的第一版是一种轻松使用当时所有 open source models 的方式,像是 pick,
05:56
to pick to and give them to artists. That was, that was initial idea is those models are too
to pick to,然后把它们交给艺术家。那是,那是最初的想法,就是那些 models 太
06:00
difficult to use if you're not a machine learning engineer, like what happens when you give them to
如果你不是 machine learning engineer,就很难用,比如你把这些东西交给
06:04
artists? Very quickly, we realized we needed to build a research org inside of Runway. And that
艺术家会怎样?很快我们就意识到,我们需要在 Runway 内部建立一个研究组织。而那
06:09
happened maybe on year one. And a lot of the mandate there was the image generation models of
大概是在第一年发生的。而那里的很多使命,就是当时的 image generation models、
06:16
the time, the video generation models of the time, or there were barely any video generations
当时的 video generation models,或者说几乎没有什么 video generations
06:20
all the time, but they were not quite there where they could be productionized and brought into
一直如此,但它们还远没到可以被 productionized 并带进
06:25
tools that would be part of creative workflows. So we need to push the frontier of the research.
那些会成为 creative workflows 一部分的工具里。所以我们需要推动研究的前沿。
06:31
And so maybe the first four years of Runway research was almost happening on the background
所以可能 Runway 的研究头四年几乎都是在幕后进行的,
06:37
until there was a moment in 2022 with light and effusion, with Dalai too, where, you know,
直到 2022 年出现了 light and effusion、Dalai too 的那个时刻,你知道,
06:45
there was that step function change, and you guys maybe remember. I started in space because
说明出现了那种 step function change,你们可能还记得。我一开始是在太空领域,因为
06:51
of basically the indefusion is the world of fusion. Because I was like, wow, this is not only
基本上,indefusion 就是 fusion 的世界。因为我当时想,哇,这不仅是
06:57
feasible, it is actually doable on consumer hardware. Exactly. I think the Delta is also huge.
可行的,而且实际上在 consumer hardware 上真的能做。没错。我觉得 Delta 也很大。
07:03
I learned picks to picks, like this was intro to ML, the TensorFlow, like Jupyter, Google
我学的是 pix2pix,就像这是 ML 入门,TensorFlow、Jupyter、Google
07:09
collab notebooks were like this, and then you have a sudden step function change, you know,
Colab notebooks 就像这样,然后突然就有了一个 step function 式的变化,你知道,
07:14
with diffusion and whatnot. Any other one sense that like there were clear examples of what early
还有 diffusion 什么的。还有没有别的让你觉得,有清晰的例子说明早期
07:20
diffusion were to get to here? Any other changes in key technology? Between 2018 and when we started
diffusion 是怎么走到这里的?还有没有其他关键技术的改变?在 2018 年到我们开始
07:28
in 2022. So one of the early work that we did in Runway was solving segmentation, image and video
2022 年之间。所以我们在 Runway 做的早期工作之一,就是解决 segmentation,图像和视频
07:36
segmentation. It was a very important problem because most VFX involves essentially separating
segmentation。这是一个非常重要的问题,因为大多数 VFX 本质上都涉及分离
07:41
rural subjects, yeah, rotoscoping, extremely manual process, nobody, nobody enjoys doing that.
乡村题材,对,rotoscoping,极其手工的过程,没人,没人喜欢干这个。
07:48
And so a lot of the early days of Runway was building this tool, it was called green screen,
所以 Runway 早期很多时候就是在做这个工具,它叫 green screen,
07:52
and it was for a long time the main thing that people were using Runway for. It ended up being used
而且在很长一段时间里,大家用 Runway 主要就是用它。它后来被用在了
07:57
in everything everywhere, all at once, a bunch of other kind of high visibility films and series,
Everything Everywhere All at Once、一堆其他那种高曝光度的电影和剧集里,
08:03
but that was essentially Runway for a long time was a post-production tool, until light
但本质上,Runway 在很长一段时间里就是一个 post-production 工具,直到 light
08:08
and diffusion generated, Gen 1, Gen 2 happened. Cool, I mean let's just go past that moment.
和 diffusion 生成,Gen 1、Gen 2 出现。酷,我是说,我们就先跳过那个时刻吧。
08:15
You've come a long way, then you started releasing your own models, maybe describe that journey as
你已经走了很远,然后你开始发布自己的 models,也许可以描述一下那段历程,随着
08:20
well. Yeah, so we go to the other point, yeah, in kind of mid 2022 when it became clear that we're
嗯。对,那我们说另一个点,对,大概在 2022 年中期,当时已经很清楚,我们
08:26
doing research at a fairly fairly small scale of of compute and it became clear that like scaling
做研究的 compute 规模其实相当相当小,而且已经很明显,就像 scaling
08:33
laws would apply to image and video gen in the same way that we're applying to language
laws 会以同样的方式,适用于 image 和 video gen,就像我们应用到 language
08:38
generation. So we made a big bet and I think, so at the time we signed this deal to build a cluster
generation 上那样。所以我们就下了一个大注,而且我觉得,当时我们签了这个协议,要建一个 cluster
08:45
of a thousand A100s which at the time we were a series B startup that was almost, you know,
由一千块 A100 组成的,当时我们还是一家 series B 阶段的初创公司,这几乎,你知道,
08:51
slightly irrational decision maybe but we really believed that if we trained a video model at the
可能有点不太理性的决定,但我们真的相信,如果我们用
08:56
large scale, we would get like a great model at the end and at the time the goal, you know,
large scale 来训练一个 video model,我们最后就能得到一个很棒的模型,而且当时那个目标,你知道,
09:03
the goal, we set the goal around fall of 2022 of what is, what does the latent diffusion, stable
那个目标,我们在 2022 年秋天左右设定了目标,就是 latent diffusion、stable 到底是什么,能做什么
09:10
diffusion moment look like for video and at the time the the best model of the time was called
video 的 diffusion moment 是什么样的,当时最好的 model 叫
09:16
COG video. I was one of the early kind of video models was very two six by two six resolution,
COG video。它算是最早期的那种 video model 之一,resolution 是 256x256,
09:23
very not not very high quality and so we decided we're going to build out this cluster and we're
quality 并不是很高,所以我们决定要搭建这个 cluster,而且我们
09:29
going to just invest in like in building out our around video model. It became clear as we're training
就打算投入进去,围绕 video model 来搭建。随着我们 training
09:36
Gen1 that it was it was difficult to get to fully, we wanted to build text to video but it became clear
Gen1,发现要完全做到很难;我们想 build text to video,但后来清楚了
09:45
to us that an easier starting point would be to to start from video to video because when you have
对我们来说,更简单的起点是从 video to video 开始,因为当你有
09:51
a stronger conditioning, it's it's basically an easier problem to re-style as an existing video
更强的 conditioning 时,基本上把现有的 video 做 re-style 是更容易的问题
09:56
versus generate a video from scratch and so we released Gen1 first back in, you know, it was
而不是从零开始 generate 一个 video。所以我们先发布了 Gen1,那是在,你知道,当时是
10:01
January of 2023. It's just a fun visual podcast honestly like you we can see February 2023
2023年1月。老实说,这就是一个挺好玩的视觉播客,就像你能看到的那样,2023年2月当时的情况就是这样。
10:08
was the state of stuff. I mean it's so interesting because at the time when you see those results,
我的意思是,这特别有意思,因为当时你看到那些结果时,你会觉得这太不可思议了,几乎就像 image generation 或 video generation 已经被解决了一样。
10:14
you think this is this is so incredible and like it's almost like image generation or video generation
然后几年后回头看,很明显,你很快就会对那些模型的结果习以为常。
10:20
is solved and then you look back a few years after and it's like it's obviously it's just like
但在我们刚开始看到那些结果的时候,你知道,感觉真的相当相当不可思议,而且你能得到的质量水平也很厉害。
10:26
you get used to results very quickly with those models but at the time when we started seeing
所以 Gen1 是一个 depth-conditioned video model,它会——它会接收一段输入视频,它会预测,它会先把它转换成 depth map,然后我们会……
10:33
those results it was you know it felt quite quite incredible and the level of like quality you could
那些结果,你知道,感觉真的真的非常不可思议,还有你能达到的那种质量水平
10:40
get and so the Gen1 was a depth-conditioned video model so it would turn it would it would take
get,所以 Gen1 是一个 depth-conditioned video model,所以它会转,它会,它会拿
10:49
an input video, it would predict it would first convert it into the depth map and then we would
一段 input video,它会 predict,它会先把它转成 depth map,然后我们会
10:56
generate pixels with a with a light and a fusion model. Yeah, very effective. I didn't realize how
用光照和一个 fusion model 来生成 pixels。对,非常有效。我没意识到那篇博客文章会这么
11:02
distracting the blog post would be sorry but one of my favorite examples actually on Gen1 was
让人分心,抱歉,但我在 Gen1 上最喜欢的例子之一其实是
11:11
both if you go up to mode 3 or mode 2 there was this storyboard use case where people would make
无论是升到 mode 3 还是 mode 2,都有这个 storyboard 用例,人们会做
11:18
with basically you can mess around with the main make a city out of pukes or out of boxes and then
基本上你可以随便摆弄主要的,用 pukes 或盒子搭出一座城市,然后
11:23
they would they would kind of shoot a video with their phone and then translate into a for the
他们会用自己的手机拍一段视频,然后转成……为了那个
11:28
realistic output. There was all these ways in which those models were starting to be used for
realistic output。那些模型开始被用于很多这样的方式,
11:32
storyboarding and also for really and then if you go to mode 4 like of taking kind of
storyboarding,还有真的……然后如果你到 mode 4,就像把那种
11:39
untextured 3d scenes and then turning them into for the realistic output. So we saw a lot of use
untextured 3d scenes 转成……为了 realistic output。所以我们看到了大量使用
11:45
cases early on where people that were familiar you know where power VFX editors would just
早期有些案例,当时那些人很熟悉,你知道,厉害的 VFX editors 就会直接
11:52
take a blender a render and then they would get translated in with Gen1 or create a scene in unity
拿 blender 渲染一下,然后他们会用 Gen1 转进去,或者在 unity 里创建一个场景
11:59
and then take a capture video of it and then and then translate into you know restart as it.
然后拍一段它的 capture video,然后、然后转换成,你知道,restart as it。
12:04
So I still think video to video is powerful. I think we we had a recent video to video model as well
所以我还是觉得 video to video 很强大。我觉得我们、我们最近也有一个 video to video model
12:10
and it's one of my favorite ways of using using those models is essentially using them to
而且它是我最喜欢的用、使用那些 models 的方式之一,基本上就是用它们来
12:16
use ground truth video as like the initial inspiration and then translate into into different styles
拿 ground truth video 当作最初的灵感,然后转换成、转换成不同的风格
12:22
or different outputs. I think we're going to go into like the rest of room we and catch people
或者不同的 outputs。我觉得我们会进入剩下的 room,我们,然后让大家
12:26
up to speed today. I did want to cover the let's call it the stable diffusion controversy or
今天跟上进度。我确实想聊一下,我们姑且称之为 stable diffusion 争议,或者
12:32
you know what happened with stability AI whatever you know I think there was a sort of two size
你知道 stability AI 发生了什么吧,不管怎样,你知道,我觉得这个故事有两种
12:37
of the story. I think there's part of that is a normal thing of like people you know joining
说法。我觉得其中一部分是那种很正常的事,就像人们,你知道,加入、
12:43
leave companies but what is the you know retrospective now that you know there's been some years
离开公司,但你知道,现在都过了好几年了,回顾起来是什么样的呢?
12:48
behind it. Yeah it's a very it's a very long story to to go into I think which I remember you
是啊,这是一个非常,是一个非常长的故事,我觉得要深入讲,我记得你
12:54
actually wrote a really long post about. We would probably cover the whole hour to to to to go
其实写过一篇很长的帖子讲这个。我们可能得花整整一个小时,才,才,才,才能
12:59
into it in a in more detail but essentially you know there was the latent diffusion paper that came
更详细地讲进去,但本质上,你知道,有那篇 latent diffusion paper 出现,
13:04
in I think that was at the at the end of 2021 and then Patrick Esser who was one of the
我觉得那是在 2021 年底,然后 Patrick Esser,他是
13:13
researchers behind latent diffusion and he worked at Runway at the time. He built latent diffusion
latent diffusion 背后的研究人员之一,当时他在 Runway 工作。他构建了 latent diffusion
13:18
collaboration with Robyn Rumbuck and a few other folks back in the Convis which was a lot like
和 Robyn Rumbuck 以及另外几个人当年在 Convis 的合作,那很像
13:26
the research group. Yeah and after releasing the early latent diffusion model they essentially
那个研究小组。对,而且在发布了早期的 latent diffusion model 之后,他们基本上
13:33
they were you know the goals keep working on versions of the model kind of scale it up incorporate
他们呢,你知道,目标就是继续做这个模型的各种版本,把它 scale up,纳入
13:39
new data incorporate new tasks and stable diffusion was basically the same model but trained on
新的数据,纳入新的任务;而 stable diffusion 基本上就是同一个模型,但训练时用了
13:44
more compute and then with a few more tricks like a classifier for guidance paper came at some point
更多的 compute,然后又加了一些技巧,比如 classifier for guidance 那篇 paper,是在某个时候出来的
13:50
I think at the early 2022. Which was a big prompting improvement. Yeah you know that improved results
我觉得是在 2022 年初。那是一个很大的 prompting 改进。对,你知道,那让结果变好了
13:57
I was trained on better data so like a aesthetic subset of fly on but it was effectively you know
我是在更好的数据上训练的,比如 fly on 的一个 aesthetic 子集,但实际上,你知道
14:04
the same the same underlying architecture and there was that big training run that happened on
同样的,同样的底层 architecture,而且还有那次大型的 training run,它发生在
14:10
stability cluster stability kind of finance finance that run and looking back at that story I
stability cluster stability 那种金融金融在运行,然后回顾那个故事,我
14:17
think it was the work to build and train that model was was done it was a it was a research project
觉得构建和训练那个模型的工作已经做完了,那是一个,那是一个研究项目
14:23
it was done as part of the continuation of the latent diffusion work it then I think the model
它是作为 latent diffusion 工作的延续的一部分完成的,然后我觉得那个模型
14:30
became very successful and it I think there were the and I think as a result of its success
变得非常成功,然后它,我觉得有那个,然后我觉得因为它的成功
14:39
their company tried to figure out the commercialization path for it but for us it was very important
他们公司试图找出它的商业化路径,但对我们来说,这非常重要
14:46
that we try to you know we make sure that we it was meant to be an open source research project
我们要努力,你知道,我们要确保它本来就是一个 open source 研究项目
14:52
and so the we decided that we should continue releasing versions of it since that was kind of
所以我们就决定应该继续发布它的各个版本,因为那算是
14:57
the original the original core of stability fusion and that led to releasing stable diffusion 1.5
最初的,最初的 stability fusion 的核心,然后这就导致了发布 stable diffusion 1.5
15:03
there was maybe a day of a bit of miscommunication there but ultimately that was resolved very quickly
可能有一天吧,那边有点沟通上的误会,但最终很快就解决了
15:09
within within hours so yeah there was not nice I just want to you know you have to you're actually
几个、几个小时内,所以对,那并不愉快。我只是想说,你知道,你得——你其实是
15:15
wanted to main players in in that sort of journey and so it's nice to hear from the source of
那段历程里的主要玩家之一,所以能从源头听到
15:20
like what happened yeah yeah I mean I think it's all it's all in the past now I would say and
到底发生了什么。是啊是啊,我的意思是,我觉得现在这都已经、都已经过去了,我会这么说,而且
15:27
like both companies you know stability took its own path runway took its own path yeah there's
就像两家公司,你知道,Stability 走了自己的路,Runway 也走了自己的路,是啊,还有
15:32
still I mean James Cameron is backing the new stability whatever they're doing with the Hollywood
不过我是说,James Cameron 在支持新的 Stability,不管他们在跟 Hollywood
15:37
studios right I don't know what they're doing I think one thing that impresses me and I'm happy to
studios 搞什么,对吧?我不知道他们在做什么。我觉得有一件事让我印象很深,而且我也乐意
15:42
move on is that back in the that time let's say like 2021-2022 there was this community of people
继续往下聊,就是回到那个时候,比如说 2021-2022 年,有那么一个社区,一群人
15:49
that you were involved in that was researching all this stuff right and like from everyone I
你参与的那个,当时在研究所有这些,对吧,而且就像从我
15:54
talked to was active then it seemed like it was fairly obvious that somebody would do the hero
聊过的每个人,当时都很活跃,然后看起来相当明显,有人会去做那个 hero
16:00
training run that would produce stable diffusion so I guess the question is you know like you had
training run,能产出 stable diffusion,所以我想问题就是,你知道,就像你当时有
16:06
the you had you had made investments you you you had the foresight is it accurate to say like
那个你当时有,你,你,你当时做了投资,你,你,你当时有先见之明,这么说准确吗,就像
16:11
that is reflective of like what people were thinking at the time or was it still very much like
这反映了当时人们像在想什么,还是说那会儿仍然非常像
16:16
well we'll use it as like a post-production tool or something I don't know you know like where
嗯,我们会把它当作后期制作工具之类的,我不知道,你知道,就像到底
16:21
where in the sentiments were we that maybe you can sort of think back to like what the community
在情绪上,我们当时处在哪儿?也许你可以稍微回想一下,比如当时社区
16:26
was like back then I remain this and I think of very fondly those early years from like
当时是什么样子。我至今记得这一点,而且我非常怀念那些早期岁月,从,嗯,
16:33
2018-2022 because it was a very small community that as you said were very convinced that this
2018到2022年,因为那是一个非常小的社区,就像你说的,他们都非常确信这个
16:38
was going to be a big thing and at the time you know anyone who because it was such a small circle
会变成一件大事。而当时,你知道,任何人,因为圈子这么小
16:44
and everyone who would like be part of that circle and like make projects with it would you know
而且每个想成为这个圈子一员、想用它做项目的人,你知道
16:51
immediately kind of get you know go viral so like they are right they're just there's some name
马上就会,你知道,火起来。所以他们说得对,只是,你知道,有个名字
16:56
on the you know GitHub door how can you face them where exactly yeah so so I remember one of the first
在你知道的 GitHub 上,老兄,你怎么面对他们,具体在哪儿,yeah,所以我记得最早
17:03
big viral moments of creative AI was there was the neural style transfer paper that there's
creative AI 的大爆红时刻之一,是有一篇 neural style transfer 的论文,还有
17:10
something dreaming I think it was called neural style transfer that was also deep-trained the
something dreaming,我记得它叫 neural style transfer,那也是 deep-trained 的
17:15
puppy slice which was also also really cool but there was this project that Jin Koga who was an
puppy slice,那也真的真的很酷,但有一个项目,Jin Koga,他当时是一个
17:23
early advisor of runway and one of those marketing guys big creative creative AI folks he literally
他是 Runway 的早期顾问,也是那种营销圈的人,大创意、创意 AI 圈的人,他真的就
17:31
just like shot a video of himself taking the New York subway kind of and going over the Williams
就拍了个视频,拍自己坐 New York 地铁,差不多吧,然后穿过 Williams
17:37
Oak Ridge and then stylized it with a I think in the in the style of Van Gogh or like one one
Oak Ridge,然后把它风格化成了——我觉得是用 Van Gogh 的风格,或者像某个
17:43
now a painter and and that was like at the at the time that was like so so cool and it went viral and
现在的画家,而且、而且当时那真的太太酷了,然后它就火了,而且
17:49
it was completely revelation the people that you could do those regenerative models and that was
这完全是个启示,让大家知道你可以做那些 regenerative models,而那还只是
17:54
only you know it was less than it was maybe 10 years ago so just like as an indication of like how
你知道,那还不到——大概也就是 10 年前,所以这就像是在说明,才
17:59
how how quick like things of Hong Kong it's really crazy like even since then you've kind of got
多快、多快、多快,Hong Kong 的那些事有多快,真的太疯狂了,就像从那时候起,你差不多已经有
18:05
people at every level of the stock you've got devs creatives artists obvious you've got everyone
stock 的每个层级都有人了,你有 devs、creatives、artists,显然,你有了所有人
18:09
using it and for people that tried stuff early they'll remember how hard it was to use regular
用它,而且对早期就试过这些东西的人来说,他们会记得用 regular diffusion 有多难
18:15
diffusion right like nowadays you can use your favorite chat GPT image or whatever give a sentence
对吧?就像现在,你可以用你最喜欢的 chat GPT image 或者随便什么,给一句话
18:21
get a beautiful output but diffusion was like you know the whole ultra HD 4K high resolution like
就能得到一个很漂亮的输出,但 diffusion 就像,你知道的,整个 ultra HD 4K high resolution 那种
18:27
prompting these things was very different anything you learned on the tooling side like from the offerings
prompting 这些东西是非常不一样的。任何你在 tooling 那边学到的东西,比如从这些 offerings
18:34
you guys have now so like creatives devs you really took the research and brought it to everyone
你们现在有的,所以就像创意人士、开发者,你们真的把研究带给了所有人
18:41
to use anything interesting there to share we had to build the entire model serving infrastructure
去用。有什么有趣的东西可以分享吗?我们得搭建整个 model serving infrastructure
18:47
for video diffusion models there was no nothing else already like as we had gen 2 was the first
给 video diffusion models 用,当时还没有任何其他现成的东西,就像我们有了 gen 2,它是第一个
18:53
text to video model I think out in the in the market so many things that we learned over time I
text to video model,我觉得是市场上最早推出的,所以随着时间推移我们学到了很多东西,我
18:59
think the biggest one was like we you know it was very clear early on that text to video was not
我觉得最大的一个就是,我们,你知道,很早就很清楚,text to video 不会是
19:05
going to be the answer like you like people wanted a lot more control than that and so we invested
答案,就像,人们想要比那多得多的控制,所以我们投资了
19:12
in like control billion type of those models very quickly you know how do you use the camera
在像 control billion 类型的那些模型上,非常快,你知道,你怎么使用 camera
19:17
trajectory as control how do you use an initial input frame as control so so that was a very
trajectory 作为控制,你怎么使用 initial input frame 作为控制,所以所以那是一个非常
19:23
early learning for us the text video was like gen 2 was not amazing you know step function
早期的学习对我们来说,text video 就像 gen 2 并不是很惊艳,你知道,step function
19:30
improvement in the quality of video models but it was you use much more an exploratory way because
在视频模型质量上的提升,但它更多是一种探索性的方式,因为
19:36
there was nothing to ground it to there was no reference that you could bring into it there was
没有什么可以 ground 它,没有可以参考的东西你可以带入,有
19:40
no you couldn't really control the camera motion you could control the object motion and so the first
不,你无法真正控制 camera motion,你可以控制 object motion,所以第一个
19:46
year in 2023 was really all about what are all the interesting ways in which we can condition those
2023 年真的全是关于我们能以哪些有趣的方式来 condition 这些模型,而且很多就是 post-training,算是基于 base model 去搞清楚,人们到底想怎么控制它们。所以很快就接连出现了一些功能,比如那个叫 motion brush 的,本来是说你可以直接画箭头,指定东西在场景里该往哪儿移动;还有 camera control,就是你可以直接描述你希望镜头在场景里怎么移动。因为我们跟来自加拿大的电影人合作,Runway 的大部分历史里,我们立刻就得到了这些反馈,然后决定这值得投入。所以 control ability 很早就成了一个很大的主题。
19:51
models and it was a lot of just post-training ground sort of of the base model to figure out like what
models 开始,而且很多就是围绕 base model 做 post-training,去搞清楚,像是
19:56
a how do people actually want to control them and so there was like this quick succession of
人们到底想怎么控制它们。然后就有一连串很快的
20:01
the we it was called motion brush was supposed you could like you could basically draw arrows and
那个……我们,它叫 motion brush,本来你可以,就是,你可以基本上画箭头,然后
20:07
dictate where things should move in the scene there was camera control that was you could just
指定场景里的东西该往哪儿移动。还有 camera control,就是你只要
20:11
describe like how you want the camera to move in the scene and because we work with filmmakers from
描述你想让镜头在场景里怎么移动。而且因为我们和来自
20:16
Canada most of the history of runway we immediately got this feedback and got this you know decided
Canada 的电影人合作,在 Runway 的大部分历史里,我们立刻就得到了这个反馈,然后得到这个,你知道,决定……
20:23
that this was worth investing in and so control ability became a big theme I think very early on
就是这值得投资,所以 control ability 很早就成了一个重要主题。
20:27
as we're building as we're building those models something fun that I haven't really talked about
在我们构建,构建那些模型的时候,有件挺好玩的事,我之前没怎么聊过太多
20:32
too much was just how gen 2 came to be out of gen 1 so so it was a bit strange because we announced
就是 gen 2 是怎么从 gen 1 演变出来的。所以,所以这有点奇怪,因为我们宣布
20:40
two months after gen 1 and very excited it was it was before gen 1 was even generally available
在 gen 1 发布两个月后,而且非常兴奋,那是在,那是在 gen 1 甚至还没有 generally available 之前
20:47
but gen gen 1 was a depth to video model so it would take a depth map and it would convert it into
但 gen,gen 1 是一个 depth to video model,所以它会接收一个 depth map,然后把它转换成
20:53
RGB and we couldn't get you know text or image to video to work directly and that's why we started
RGB,而且我们没法让,你知道,text to video 或 image to video 直接跑通,所以我们就开始
21:01
from depth to video but you know and we had discussions of like okay we need to spend the next six
从 depth to video 开始,但你知道,我们也讨论过,比如,好吧,接下来六个月我们得花
21:08
months actually investing in text video maybe increasing the compute scale or the model scale like
几个月真正投入到 text video 上,也许提升 compute scale 或 model scale,比如
21:14
training larger model and I had this weekend project idea which was what if I take a model that
training 更大的 model,然后我有个周末项目的点子,就是,要是我拿一个 model,那个
21:21
starts from text input and converts to depth maps and then and then use gen 1 to convert the depth
从 text input 开始,转换成 depth maps,然后,然后用 gen 1 把 depth
21:27
into RGB and so gen 2 was basically that and it worked pretty well I mean there were you know if
转成 RGB。所以 gen 2 基本上就是这样,效果还挺好。我是说,你知道,如果你
21:41
you with a knowledge that it has this like two stage pipeline you can tell in some cases that the
知道它有这么个 two-stage pipeline,有些情况下你能看出来,
21:47
structure of the video looks a bit off because you have to generate the depth first before you go
视频的结构有点不太对,因为你得先生成 depth,然后才能进入
21:52
into the output video but it worked and it allowed us to bring this to our users very quickly
输出的视频。但它确实行得通,也让我们能很快把这个带给用户。
22:01
but it's actually now it's interesting because like people are coming back to this almost two-stage
但现在其实挺有意思的,因为大家又回到这种几乎算是 two-stage 的
22:07
approach like if you look at the rive text to image model that came a few months ago it had this
方法了。比如你看几个月前出来的 rive text to image model,它就有这个
22:14
this planner model that would generate bounding boxes before it fed that into a diffusion
这个 planner model,会先生成 bounding boxes,然后再把它送进 diffusion
22:19
transformer yeah ideal ground also the same day I remember that was very strange they both of
transformer 啊,ideal ground 也是同一天,我记得那很奇怪,它们两个在同一天带着完全相同的创新出来了。
22:24
them came out the same day with the same exact innovation it's a small community I'm like this is
这个圈子很小,我当时就想,这完全是巧合吧。
22:28
like this is completely coincident all right people talk so yeah there's there's definitely
好吧,人们都在说,所以是的,这个方向绝对有点东西。
22:36
something into this approach and obviously now like every single video like video generation model
显然现在,就像每一个视频,像视频生成模型,在生产环境中底层都用了复杂的 prompt completion pipeline。
22:41
in production uses a complex prompt completion pipeline under the hood I think that's no secret that
我觉得这不是秘密,人类在 prompting 方面很糟糕,我觉得普遍如此。
22:47
there is human's a terrible prompting I think across the board yeah I think like the original
是的,我觉得就像最初的 Sora 那篇博客文章甚至告诉过你,你输入之后发生的事情就是重写你的 prompt,就是...
22:53
Sora one blog post even told you that what happens after your input is rewriting your prompt it's
Sora 的一篇 blog post 甚至告诉你,在你的 input 之后发生的事就是重写你的 prompt,它会把你到底想要什么描述得更详细。对,而且 DALL-E 3 paper 里也有类似的句子:Gen 2 勉强能生成连贯的视频,但如果你把它 scale up,你就会……按理说没有理由行不通。某种程度上,嗯,我觉得那一直就是 Runway 的那种思路,就是这种外推:你知道,就算我们从 2018 年开始,你看当时的结果,你得更多看趋势,比如我们在 2018 年处在哪,对比我们到了——你知道,当第一个 GAN 在 2020 年,或者 2014、2015 年出来的时候,你从 32 by 32 的人脸图像开始,然后到 2018 年的时候,你……
23:00
much more descriptive about what you would want exactly yeah and there was the Dalai three paper
更能详细描述你到底想要什么,对,还有那篇 DALL-E 3 paper
23:06
beforehand that was the kind of the the first public description of the fact that synthetic
之前,那算是第一次公开描述这个事实:synthetic
23:12
captions and really detailed captions work really well and then Sora kind of built built on that
captions 和非常详细的 captions 效果特别好,然后 Sora 差不多就是在此基础上发展出来的
23:19
yeah so it was 2023 we were kind of releasing all these updates to gen 2 like the camera control
对,所以那是 2023 年,我们当时在给 gen 2 发布各种更新,比如 camera control
23:24
motion brush and there was actually something very interesting about camera control because it was
motion brush,而且 camera control 其实有个很有意思的地方,因为那是
23:30
the first time that you felt that instead of like you were creating video you were creating a short
第一次让你感觉到,你不是在创作视频,而是在创作一个短
23:34
video you were actually navigating inside the world and I think camera control was maybe the seed of
视频,你实际上是在这个世界里穿行,而且我觉得 camera control 可能是……的种子
23:40
some of the ideas that we had around world models and really opening up that research direction
我们围绕 world models 的一些想法,并且真正打开了那个研究方向
23:46
we realized you know it was this era and this series of you know gen 1 and gen 2 models really
我们意识到,你知道,就是那个时代,以及你知道的 gen 1 和 gen 2 这一系列模型,真的
23:54
proved to ourselves yeah this is this is there so this is not the original camera control this was
我们向自己证明了,对,这个,这个就是,所以这不是最初的 camera control,这是
24:00
the update to camera control on top of gen 3 but yeah I think it made those models usable
是在 gen 3 基础上对 camera control 的更新,但没错,我觉得它让那些模型变得可用了
24:06
to filmmakers outside the so camera camera control was very was very popular and so we realized
对外的电影制作人来说,所以 camera,camera control 非常,非常受欢迎,于是我们意识到
24:14
you know there is one way of seeing those models which is you know you're just as a content creation
你知道,有一种看待这些模型的方式,就是,你知道,你只是把它们当成内容创作
24:19
machines and there is the other way which is your as your predicting video in order to predict video
机器,而另一种方式就是,你把它们当作在预测视频,为了预测视频
24:25
well you need to simulate the world in an increasing and increasing capacity and if scaling laws apply
嗯,你需要以越来越强、越来越强的能力去模拟世界,而如果 scaling laws 适用
24:32
on video just like they apply on language models then as we scale the computer we put in those
于视频,就像它们适用于 language models 一样,那么随着我们扩大投入这些模型中的 computer
24:38
models then they're going to be able to simulate physics they're going to be able to simulate human
模型,它们就能模拟物理,它们就能模拟人类
24:42
actions and dynamics increasingly well and predictably well that was the thesis about around
动作和动态越来越准确,而且可预测性越来越好,这就是当时的大致论点
24:49
our efforts on world models and we spin up this research group to just focus on on world models
关于我们在 world models 上的努力,我们成立这个研究小组,就专注于 world models
24:55
and how do we turn the video generation models that we're building into something broader and
以及我们如何把我们正在构建的 video generation models 变成更广泛的东西,
25:00
something that would be useful beyond also configuration as well and that was roughly when
一些除了 configuration 之外也有用的东西,那大概就是当时
25:06
yeah so that was in late 2023 interesting you know like I think a lot of people have been saying a lot
对,所以那是在 2023 年底,有意思的是,你知道,我觉得很多人一直在说很多
25:12
of video gen model companies have all pivoted to world models these days but like 2023 you're posting
很多 video gen model 公司现在都转向了 world models,但就像 2023 年你发帖
25:20
it it's debatable whether it's a pivot like arguably that's what you always had to do anywhere you're
这算不算 pivot 是有争议的,可以说你一直都得这么做,无论如何你
25:27
in a way an expansion of the applications of the models as they become more capable the early
在某种程度上是模型应用范围的扩展,随着模型能力越来越强,早期的
25:32
signs it seems like the original models you guys had people that say it's very not bitter lesson
有迹象表明,看起来你们最早的那些模型,有人说它非常不 bitter lesson
25:37
piled right you're adding rewriting prompts you're having all these one-off things but
堆起来,对吧,你在加东西、重写 prompts,你有一堆这种一次性的东西,但
25:42
that's just the state of the tech as it was versus the future of as you said you can scale it up is
那只是当时技术所处的状态,而未来就像你说的,你可以把它 scale up,是
25:48
you know we can scale up to world models yeah so it just became and and if you looked at the outputs
你知道,我们可以 scale up 到 world models,对,所以它就变成了,而且,而且如果你看那些输出
25:55
of gen 2 it was not I think it was not obvious to people that this would scale to become a general
gen 2 的,它并不,我觉得,人们并没有明显意识到这会 scale 成一种通用的
26:00
simulator of the world like you had very limited movement you had you know very low fidelity or
world simulator,比如动作非常有限,你知道,fidelity 非常低,或者
26:06
resolution like obvious mistakes in human anatomy like all kinds of limitations but it was just
resolution 很低,比如人体解剖上明显错误,各种局限,但它只是
26:12
you know the idea was that's just GPT2 and GPT2 you know you can barely generate like coherent
你知道,当时的想法是,那只是 GPT2,而 GPT2,你知道,你几乎只能勉强生成,像是连贯的
26:19
sentences similar gen 2 can barely create coherent video but if you scale it up you're gonna
Sora 也类似,Gen 2 几乎不能生成连贯的 video,但如果你把它 scale up,你就会
26:25
there is no reason why it shouldn't work in a way it's uh and I think that was that's that's always
没有理由它不该在某种程度上奏效,呃,我觉得那就是,那一直是
26:30
the mindset of kind of of runways like this extrapolation of like if you know like even when we start
Runway 的那种心态,就像这种 extrapolation,就像,如果你知道,即便我们从
26:37
in 2018 and you looked at the results of the day you need to look more at the trend of like
2018 年开始,你看着当时的结果,你更需要看的是趋势,就像
26:43
where we're in 2018 versus when we're at the you know when the first gun came out in 2020 for
我们在 2018 年处在哪儿,对比我们在,你知道,当第一个 GAN 在 2020 年出来的时候,对于
26:49
2014 or 2015 and you start from like 32 by 32 images of faces and then by the time in 2018 you
2014 或 2015,然后你从像 32 by 32 的人脸图像开始,然后到 2018 年的时候你
26:59
could generate you know three images at the one care resolution and it was kind of the same with
可以生成,你知道,三张 one care resolution 的图像,而且这跟
27:05
world models very early science of something much bigger yeah I was kind of thinking like it's
world models 还是很早期的科学,研究的是一个要大得多的东西。对,我有点觉得它就像
27:10
kind of diffusing into focus like if you look at our visible output from year to year to year it
有点像慢慢从弥散变得聚焦,就像如果你逐年看我们可见的输出,它
27:15
it looks like a diffusion process itself yeah especially watching the early like for all the
它本身看起来就像一个 diffusion process,对,尤其是看早期那些,就像所有那些
27:20
old blog posts you can really see the topiness the details like human civilization starting from
旧博客文章,你真的能看到那种 topiness、那些细节,就像人类文明从
27:26
random noise and then yes you know using it too yeah yeah we just run it as a realization yeah that's
random noise 开始,然后,对,你知道,也在用它,对,对,我们就把它当作一个 realization 来跑,对,这就是
27:31
how you know you're on track you know you're still noisy right yeah I like the way that you guys
你怎么知道自己在正轨上,你知道,你还是 noisy,对吧,对,我喜欢你们
27:35
phrased it when you uh announced it in June which is or that you had a sort of uh video I say
在,呃,六月宣布它时用的说法,也就是你们有个那种,呃,视频,我说
27:39
the human mind is no longer the center of AI or world this right which is uh you know let's let's
人类心智不再是 AI 或世界的中心,这句,对吧,也就是,呃,你知道,咱们,咱们
27:44
call it the past five years of of LLM based AI is very much like trying to emulate human preferences
就把它叫做过去五年基于 LLM 的 AI,很大程度上就像是在尝试模仿 human preferences
27:52
and human speech but now that's like mostly solved I think that's like some of the context of your essay
还有人类语音,但现在这基本上已经解决了,我觉得这算是你那篇文章的一些背景
27:58
which you also wrote around the time and now it's like the focus of some modeling the world accurately
你差不多也是那时候写的,而现在这有点像某些准确地对世界进行 modeling 的工作的重点
28:03
exactly yeah so the the way we see it is there is that uh that initial mission statement of deep
没错,是的,所以我们是这样看的,就是那个呃,那个 DeepMind 最初的使命宣言
28:10
mind which is uh solving intelligence and I used it to solve everything else but I think it's
也就是呃,solving intelligence,然后我用它来解决其他所有事情,但我觉得这
28:17
starting from everything else could be valuable of like starting from you know there is just so much
从其他所有事情开始可能是有价值的,就像从你知道的,有那么多
28:23
complexity and detail in the world that in order to that is it's hard to learn directly from
世界中的复杂性和细节,以至于为了……也就是说,很难直接学习
28:30
just human descriptions of the world like we're assuming that you know like language models
仅仅从人类对世界的描述,就像我们在假设,你知道,像 language models
28:36
learn from everything that humans have written about the world like our own understanding as of
从人类写下的关于世界的一切中学习,就像我们自己的理解,截至
28:41
you know the 2020s and there is just so much that we don't know and so much that's not captured
你知道,2020s 就是有太多我们不知道的东西,还有太多没有被捕捉到
28:46
by existing text uh about both the you know the low level dynamics of the world like we're not
被 existing text 呃,关于这两者,你知道的世界的 low level dynamics,就像我们并没有
28:53
describing in detail you know if you if I tell you describe like how do you tie your shoes
详细描述,你知道,如果我让你描述一下,比如你怎么系鞋带
28:58
that's a very difficult thing to to describe in words but it's very obvious thing to demonstrate
那是非常难用语言描述的事情,但演示起来却非常明显
29:05
and so I think there's been and there's you know more of x paradox like we're constantly underestimating
所以我觉得一直以来,而且你知道,更多是 x paradox 那种情况,就像我们一直在低估
29:11
all the complexity that goes into very like things that we do subconsciously as humans and we don't
所有那些我们作为人类潜意识里做的事情所涉及的 complexity,而我们并不
29:18
even necessarily always have the words to describe them and so in my mind the simulating the world and
甚至不一定总是有词去描述它们,所以在我看来,simulating 世界和
29:26
simulating uh physics simulating the dynamics of the world has always been kind of underestimated uh
simulating 呃 physics、simulating 世界的 dynamics,一直以来都有点被低估了 呃
29:33
compared to uh we place too much emphasis on the things that are easy to talk about but there's
相比之下,呃,我们太强调那些容易谈论的东西了,但世界上还有
29:41
just all this complexity and kind of richness of the world that if we just try and train directly
所有这些复杂性和丰富性,如果我们只是试着直接
29:47
on that observational data instead of training on how people describe the world we would learn
在那样的 observational data 上训练,而不是根据人们如何描述世界来训练,我们就会学到
29:52
something new that we would otherwise know you think that the present architectural paradigm is
一些新的东西,我们本来就会知道——你觉得现在的 architectural paradigm 是
29:57
fine you don't need like another layer like a japa you know like another famous uh New York AI
没问题的,你不需要再加一层,比如 JEPA,你知道,就像另一位著名的,呃,New York AI
30:04
leader would say you know we're a very pragmatic research lab if we have evidence that an approach
领袖会说的那样,你知道,我们是一个非常务实的研究实验室,如果我们有证据表明某个方法
30:10
works better than the approach that we're taking then we have no qualms to taking it we just have seen
比我们正在采用的方法效果更好,那我们就会毫不犹豫地采用它,我们只是看到
30:15
no indication that video prediction itself doesn't scale and even if you look now you know not just
没有任何迹象表明 video prediction 本身不能 scale,而且即便你现在去看,你知道,不只是
30:21
our work but the work of others you're seeing in robotics some of the most promising work
我们的工作,但还有你在 robotics 里看到的其他人的工作,一些最有前景的工作
30:26
starts from video prediction models and then you adapt them to also the action models for example
从 video prediction models 开始,然后你把它们适配到 action models 上,比如说
30:31
so there is very little evidence that you need something else and like your time is better spent on
所以几乎没有证据表明你需要别的东西,而且你的时间最好花在
30:39
a novel architectural change compared to improving data and improving the and scaling the current
一个全新的 architectural change 上,而不是改进 data、改进……以及 scaling 当前
30:46
the current approach and so we don't have any indication that you know the there is that counter
当前方法,所以我们没有任何迹象表明,你知道,有那个反
30:54
argument that I think there was a tweet by Jan Lecou in a few days ago that you know understanding
论点,我想几天前 Jan Lecou 发过一条推文,说你知道,理解
31:01
the dynamics of the world is very different than generating cute videos and your answer is no
世界的 dynamics 和生成可爱的视频非常不同,而你的回答是:不
31:06
they're the same thing they're the same videos are the same as understanding physics right because if
它们是一回事,它们是一样的,视频就和理解物理一样,对吧,因为如果
31:12
you want to generate you know obviously video models can cheat and like they could you could give
你想生成,你知道,很明显 video models 可以作弊,就像它们可以,你可以给出比如场景的一系列 diff shots,用一种不需要你真的去 simulate 困难 physics 的方式。总是有,就像,各种不同的方法,让你可以隐藏 model 的不足。重要的是,不要有点太被当前 video models 的表现骗到。很容易,
31:16
like successive diff shots of the scene a way that doesn't require you to actually simulate
像是场景的一连串 successive diff shots,一种不需要你真正去模拟
31:22
difficult physics there is always like all these different ways in which you can hide the
困难的 physics 的方式。总是有各种不同的办法,让你可以隐藏
31:27
deficiencies of the model and it's important not to be kind of too tricked by the performance of
model 的缺陷。而且很重要的一点是,不要被……骗得太厉害
31:33
the current video models it's easy to you know cherry pick examples and and think that video models are
当前 video models 的表现。你很容易,你知道,cherry pick 一些例子,然后觉得 video models 已经
31:39
further advanced than they actually are so there is a lot a lot more work that we need to do to improve
比实际上先进得多。所以,我们还有很多很多工作要做,来改进
31:43
those models but in my mind very similar to language and like we've you know do you go from
那些 models。但在我心里,这跟语言非常像。而且就像我们,你知道,从
31:50
barely coherent sentences to something that you know could call the conversation with a human
勉强算连贯的句子,走到某种你知道可以称之为跟人类对话的东西
31:55
to something that could can operate autonomously for for a day and like create an entire code basis
到某种能自主运行一整天、还能创建一整个 codebase 的东西
32:02
and the main difference there's obviously some architecture improvements on the way but the
而主要区别在于,显然一路上会有一些 architecture 的改进,但
32:06
main thing is scale and so it's the same bad for video and we have no indications that this is
主要就是 scale,所以对 video 来说也是同样糟糕,而且我们没有迹象表明这是
32:12
out trading like we have benchmarks that we use for measuring the physics of those models and we see
out trading,就像我们有 benchmarks,用来衡量那些 models 的 physics,而且我们看到
32:18
those predictably improve as we scale those models so there is if you want to pull up a physics IQ
那些会随着我们 scale 那些 models 而可预测地提升,所以,如果你想调出 physics IQ
32:25
is one of those benchmarks that measures how well does the model performance hold mechanics
它是那些 benchmarks 之一,衡量 model performance 在 mechanics
32:30
or fluid dynamics or optics if you've seen any emergence any scaling law around this yeah he's
或 fluid dynamics 或 optics 上保持得有多好,如果你看到过任何 emergence、任何围绕这个的 scaling law,对,他
32:37
saying there is the scaling law right exactly yeah so so the way those those those models those
在说存在 scaling law,对,没错,对,所以那些那些那些 models 那些
32:42
benchmarks work is you you know the researchers have gone and like captured a few videos that are
benchmarks 的工作方式是,你知道,研究人员会去,像是,捕捉一些视频,这些视频
32:49
representative of the physical phenomena and then you can take the first frame and then pass it
能代表这些物理现象,然后你可以取第一帧,然后把它
32:54
through any much to video model and then generate generate kind of a rollout that shows what should
输入到任何 much to video model 里,然后生成,生成一种 rollout,展示接下来应该
32:59
happen next so you have a ball hanging from the ceiling and then you use that as input and then
发生什么。比如你有一个球吊在天花板上,然后你把它作为输入,然后你
33:05
you the model predicts how the ball should fall on the ground and this measures you know we have an
模型就会预测这个球应该怎么落到地上,而这衡量的是,你知道,我们有
33:12
intuitive understanding of physics I know you know you can imagine what will happen next if I drop
intuitive understanding of physics,我知道,你知道,你可以想象接下来会发生什么,如果我把
33:17
this this bottle so it's measuring that same intuitive physics understanding of those models
这个,这个瓶子丢下去,所以它衡量的是那些模型身上同样的 intuitive physics understanding
33:22
and we've measured that at different model scales and we see and compute scales and we see that the
而且我们在不同的 model scales 上衡量了这一点,我们看 compute scales,我们看到
33:28
score and physics IQ predictably improves there's other you know tricks and techniques that you can
score 和 physics IQ 可预测地提升也差不多;还有别的,你知道,一些 tricks 和 techniques,你可以
33:33
make to improve the score even further but even scale alone helps in in the in the model learning
用来进一步提升 score,但光是 scale 本身就能帮到,让 model 学
33:38
better physics my main sympathy with Jan Lecune is the Plato's cave allegory right like you're
更好的 physics。我跟 Jan Lecune 最大的共鸣就是 Plato's cave allegory,对吧,就像你是在
33:46
like learning on the output of a thing not the internal process of a thing and it's very very noisy
学一个东西的输出,而不是这个东西的内部过程,而且这非常非常 noisy
33:51
and you know if only you could observe the internals of a thing it's hard to observe the internals
而且你知道,要是能观察一个东西的 internals 就好了;很难观察
33:56
of a human mind but you can very much observe or at least we have a whole branch of science
人的头脑的 internals,但你其实可以很好地观察,或者说至少我们有一整个科学分支
34:02
and physics that we're ignoring on how to model physics and in movement and you know gravity
还有 physics,我们一直忽略的,关于如何 model physics 和 movement,还有你知道的 gravity
34:08
and you know other interactions and we're just like throwing away all that and just saying just
而且你知道,其他的 interactions,我们就像把那些全扔了,然后就说,就
34:12
just scale data which is very much the lesson of unsupervised learning but it feels wrong
就只是 scale data,这很大程度上就是 unsupervised learning 的教训,但感觉不对
34:19
as the signal the main idea I think the history of machine learning is a large it feels wrong
作为 signal,我觉得 machine learning 的历史主要就是大规模,但这感觉不对
34:25
yeah it's a little less than right now it's a simple answer to that I guess how much can you scale so
对,现在它还有点不太对,对那个问题的简单答案,我猜就是你能 scale 到多大,所以
34:30
like even on let's say the video generation side like there's one side of video understanding
就像即使是在,比如说 video generation 这一侧,就好像有一侧是 video understanding
34:35
video generation are we still gonna have tools where it's like on generate two hours 20 hours
video generation,我们还会不会有那种工具,就像直接生成两个小时、20 个小时
34:41
there's a infer a way to do it it batches and such it together but like do we just keep scaling
有一种 infer 的方式来做,它会把东西 batches 到一起之类的,但就像,我们就只是继续 scaling 吗
34:48
do we just continue long generation consistency all that with scale and like tying it into where we're
我们就继续 long generation、consistency 这些,靠 scale,然后把它联系到我们
34:55
at now from we looked at runway two to 4.5 like technically what what advancements have we made
现在的位置,从我们看 Runway two 到 4.5,就像技术上,我们到底取得了什么什么进步
35:01
to today and then where do you see things still going so part of the answer is definitely scale
到现在,然后你觉得事情接下来还会往哪里走?所以答案的一部分肯定就是 scale
35:06
and that was we learned that lesson in a in a big way for with gen 3 so gen 3 was the
而那一次,我们在 gen 3 上算是狠狠地学到了这个教训,所以 gen 3 是那个
35:12
most released the year after like in 2024 that was a few months after Sora was released so
大多数是在第二年发布的,比如 2024 年,那是在 Sora 发布几个月之后,所以
35:20
yeah there's an interesting story that that that that came to be as well gen 3 for us was you know
对,这里面也有一个挺有意思的故事,事情也是这么来的,gen 3 对我们来说,你知道,
35:25
the the first time that we need really needed to build basically we had to learn all the lessons
那是我们第一次真正需要去搭建,基本上,我们得学会所有的经验教训
35:30
that the language model world learned in it to in in three years in the span of a few months
也就是 language model 领域在三年里学到的那些,要在短短几个月内
35:37
one of the biggest changes of Sora was using diffusion transformers instead of confidence
Sora 最大的变化之一,就是用 diffusion transformers 代替了 confidence
35:41
so a lot of the early you know the legendary fusion models were all um uh
所以很多早期的,你知道,那些传奇的 fusion models 全都是,呃,呃
35:47
components for the diffusion model part and the diffusion transformer paper came at some point
diffusion model 那部分的 components 和 diffusion transformer 那篇论文,是在 2023 年
35:53
in 2023 and it basically showed scaling laws for image uh diffusion transformers and we
某个时候出来的,它基本上展示了 image 呃 diffusion transformers 的 scaling laws,而我们
36:01
realized at that point that we need to invest in infrastructure for model parallelism for
那时意识到,我们需要投入 infrastructure 来做 model parallelism,为了
36:07
really scaling scaling training to larger than you know a few billion parameter models and we spent
真的 scaling、scaling training 到超过,你知道,几十亿 parameter 的 models,然后我们花了
36:14
maybe the you know most of the fall of 2023 building out our infrastructure for distributed training
大概,你知道,2023 年秋天的大部分时间,来搭建我们用于 distributed training 的 infrastructure
36:22
and we had a lot of full starts and a lot of failure in trying to scale um image and video
而且我们在试图 scale 呃 image 和 video
36:27
diffusion transformers and at that point uh you know February 2024 Sora comes out and the results are
diffusion transformers 时,有很多 full starts,也失败了很多,然后到那个时候,呃,你知道,2024 年 2 月 Sora 出来了,结果
36:35
very much superior to what gen 2 could produce there were a lot of you know a lot of chatter on twitter
比 gen 2 能生成的好太多了,Twitter 上有很多,你知道,很多讨论
36:42
about runway runways done uh like there is there is no way runway will catch up and if you remember
说到 runway,runway 已经完了,呃,就像,runway 根本不可能追上来,如果你还记得
36:49
also open AI in the early 2024 it felt very like it's a formidable opponent now but at that point
还有 open AI,在 2024 年初,它给人的感觉非常像是——它现在是个很强大的对手,但在那个时候
36:59
it you know there were on the top of their game you know nobody could even get close to them
它,你知道,他们当时正处于巅峰状态,你知道,甚至没人能接近他们
37:04
there was maybe Gemini was just the first version of Gemini had just released so when open AI came
当时可能 Gemini 才刚发布第一版 Gemini,所以当 open AI 带着
37:09
with Sora and there was such a big jump of equality it gave me there was like an existential crisis
Sora 出来,而且质量一下子提升了那么多,让我感觉就像经历了一场存在主义危机
37:16
for for a few hours but that I think the the amazing thing about runway and like I think the you
持续了几个小时,但我觉得,runway 最了不起的地方在于,而且我觉得,你
37:22
know we've been around eight years now which is almost were dinosaur in AI and we had to like
知道,我们已经做了八年了,现在在 AI 里几乎算是恐龙了,我们不得不,就像
37:28
there was a lot of those moments where to learn adapt very quickly and build out skillset in
有很多那样的时刻,需要快速学习、适应,并建立技能,在
37:33
the team that we didn't have and so you know if you ask anyone what is their favorite time at runway
我们当时没有的那支团队,所以你知道,如果你问任何人,他们在 Runway 最喜欢的时期是什么
37:38
that was there during that time it was that that push in like three months to get to a model better
在那段时间待过的人,都会说那就是那种在大约三个月里的冲刺,要做出一个比
37:43
than that Sora and it you know we scale 10x the model scale the the the the the model size and uh
那个 Sora 更好的 model,而且你知道我们把 model scale 扩大了 10 倍,那个那个那个那个 model size,还有呃
37:49
you know compute that we were training on uh we figure out model parallelism we had zero expertise
你知道我们训练用的 compute,呃,我们搞明白了 model parallelism,而我们完全零经验
37:55
in that and then we came out with Gen 3 during during that summer so uh so that was a big turning
在这方面,然后我们在那个夏天期间推出了 Gen 3,所以呃,所以那是一个很大的转折
38:01
point I think for the company where the the research or grew very quickly and we really started
点,我觉得对公司来说,那是 research org 成长得非常快的地方,然后我们真的开始
38:06
pursuing this vision of the general world model uh in in earnest I think after after Gen 3 was out
认真地追求 general world model 这个愿景,呃,我觉得是在 Gen 3 出来之后
38:12
yeah I mean that's the the amazing thing about building when you're building there's no stack
对,我的意思是,这就是构建最神奇的地方,当你在构建的时候,根本没有 stack
38:17
to you have to invent everything yourself you have to be completely full stack you know and now I
对你来说,所有东西都得你自己发明,你得完全 full stack,你知道,而现在我
38:22
now I think that they are inference specialists like foul or whatever that can help with like
现在我觉得他们是 inference 专家,像 foul 或者什么的,能帮忙做类似
38:27
model serving and I think you guys work with them as well but yeah like it's you know at the time
model serving,而且我觉得你们也跟他们合作,但是对,就像,你知道,当时
38:32
it was just it's very interesting to think about what you do when when Sora comes out and people
那只是,想想当 Sora 出来的时候你该做什么,这非常有意思,而人们
38:38
are questioning whether your company should still exist yeah and yeah there was no you know
都在质疑你的公司是否还应该存在,对,而且对,那时没有,你知道,那
38:42
there was no real I'm of diffusion models that we had to build the whole model serving infrastructure
没有真正的 diffusion models 的 inference,我们必须搭建整个 model serving infrastructure
38:47
and make things efficient and a few months after we released Gen 3 we released the turbo version
并让东西变得高效,而在我们发布 Gen 3 几个月后,我们发布了 turbo 版本
38:52
which I think was the first step distilled model and production there was a full trend that we covered
我觉得那是第一个投入 production 的 step distilled model。当时有一个完整的趋势,我们报道过
38:58
as well yeah so that allowed us actually to serve to serve those models at the larger scale because
而且,是啊,所以这其实让我们能够去 serve,去 serve 那些模型,在更大的 scale 上,因为
39:05
I think the first version of Gen 3 was was quite expensive to serve you know I think the the whole
我觉得 Gen 3 的第一个版本,是,是 serve 起来相当贵,你知道,我觉得整个
39:10
like trend in like consistency models lightning and turbo and all these things somehow didn't
像 consistency models、lightning 和 turbo 以及所有这些的 trend,不知怎么并没有
39:15
really stick around I don't know if you have any reflections on this because at the time I was like
真的留下来,我不知道你对这个有没有什么反思,因为当时我就觉得
39:20
well obviously everything should start with distilled model first and then you can upscale right
嗯,显然一切应该先从一个 distilled model 开始,然后你再 upscale,对吧
39:26
basically your your bigger models just turn into fancy upscalers but like you should always draft
基本上,你你更大的模型就变成了花哨的 upscalers,但就像你应该总是 draft
39:32
with a smaller model and faster model right because you can get it so quickly like near real time
用一个更小、更快的模型,对吧,因为你可以很快拿到它,像近乎 real time 一样
39:38
yeah I would not be so sure to say that didn't stick around I think that's it's it's likely to
是啊,我不会那么确定地说它没有留下来,我觉得那,它,它很可能会
39:46
I mean that there is a lot of step distilled models that are actively using production
我是说,现在有很多 step distilled models 正在 production 里被积极使用
39:50
there is still obviously a gap in quality compared to the you know the non distilled model
但跟你知道的那个 non distilled model 比,质量上显然还是有差距
39:58
but in my mind we're still you know there is a two to three year offset from language models
但在我看来,你知道,我们跟 language models 相比还是有两到三年的 offset
40:04
so the things that so it's just a matter of time before there is better distillation techniques
所以这些东西,所以只是时间问题,早晚会有更好的 distillation techniques
40:12
you know we use right now we have a real time model core character is that I think it's the
你知道,我们现在用的,我们有一个 real time model core character,就是,我觉得这是
40:16
largest deployment of real-time video models that's a step distilled model and it's actively
real-time video models 里最大的 deployment,那是一个 step distilled model,而且它正在被积极
40:22
being used it's a very specific use case compared to a general video model so this is a
使用,跟 general video model 相比,它是一个非常具体的 use case,所以这是一个
40:27
way this is a avatar it's just to see character yeah so this is a talking avatar model we were
方式,这是一个 avatar,就是来看 character 的,对,所以这是一个 talking avatar model,我们当时
40:34
able to you know we we optimized the the hell out of it and and it it generates a 24 FPS
能,你知道,我们我们把它往死里优化,它它现在能以 24 FPS 生成
40:42
and it's a it's a step distilled autoregressive video model so if we look at our world model direction
而且它是一个 step distilled autoregressive video model,所以如果我们看我们的 world model 方向
40:49
a big component of it is starting from the bi-directional diffusion that basically generates
其中很大一部分是从 bi-directional diffusion 开始,它基本上一次性生成整个视频,然后做成 autoregressive shows,让你一次生成一帧或几帧
40:54
entire video at once and making autoregressive shows so you generate one frame or a few frames at a
所以要让整个 pipeline 达到 real-time model,里面有很多东西要做
40:59
time so there's a lot that goes into that pipeline of getting to a real-time model it's first you
首先你得把它变成一个 causal autoregressive model
41:06
need to make it into a causal autoregressive model and then you turn it into you need to do some
然后你还得再做一些额外的 step distillation,才能真正做到 real-time
41:11
additional step distillation to get it to actually be real-time and I think that part is actually just
我觉得这部分其实才刚刚开始
41:17
just starting out we're very surprised if we're you know two years from now we don't primarily
如果我们,你知道,两年后不是主要……我们会非常惊讶
41:25
use real-time models to me real-time video generation it's just inevitable that you know it has
用 real-time models 来做 real-time video generation,这你知道就是不可避免的,它会有
41:30
much better user experience it's much cheaper to serve and you know the quality gap between
好得多的 user experience,服务起来也便宜得多,而且你知道,质量差距
41:36
the base model and the real-time model is only going to close as we as we figure out better
在 base model 和 real-time model 之间只会缩小,随着我们、随着我们想出更好的
41:41
distillation techniques and we made a lot of progress there internally on maintaining the quality
distillation techniques,而且我们在内部已经取得了很多进展,在保持
41:47
of the base model when we when we distill them how much of this is transferable so is it the same
base model 的质量,当我们、当我们 distill 它们的时候,这有多少是可迁移的?所以是同一个
41:52
base model like if you're doing diffusion across the whole sequence and you're converting it to step
base model 吗?比如如果你做的是整个 sequence 上的 diffusion,然后你把它转换成 step
41:57
autoregressive distillation is this like distillation where you still need to train both you can use
autoregressive distillation,这算不算那种你还是得同时 train 两者的 distillation?你可以用
42:03
the same base and converter what's that process like to go from regular model to something that's
同一个 base 和 converter?那个过程是怎样的,从 regular model 到某种……
42:09
real-time on a technical level so the nice thing about the fusion models is you have two access
从技术层面来说,real-time 的好处是,对于 fusion models,你有两种 access
42:14
of distillation so there is the you can distill to a smaller model which resembles what you do in
distillation 的,所以你可以 distill 到一个更小的模型,这类似于你在
42:19
LLIMS or you can distill in in terms of taking less steps less diffusion steps so you could take a
LLM 里做的,或者你可以 distill,在采取更少的 steps、更少的 diffusion steps 方面,所以你可以拿一个
42:26
model that generates in 50 steps and and and generating four steps and get to you have some
模型,它在 50 步内生成,然后生成四步,然后你会得到一些
42:32
performance degradation but very often you get comparable comparable up with so you can even take the
performance degradation,但很多时候你会得到相当的 performance,所以你甚至可以拿
42:40
the you know the large frontier model and distill it with step distillation and get to a real-time
那个,你知道,large frontier model,然后用 step distillation 对它进行 distill,并达到 real-time 的
42:47
performance and that's what we've seen so depending on the use case in some cases we might also
performance,这就是我们看到的,所以取决于 use case,在某些情况下我们可能也会
42:52
serve with a smaller model but in a lot of use cases we actually just use the
用更小的模型来 serve,但在很多 use case 中,我们实际上只是用
42:56
frontier model and we're able to make it work in real-time I think this might be a good time to
frontier model,而且我们能让它 real-time 地跑起来。我觉得现在可能是个好时机,来
43:01
cut over to his laptop to show off some of the real-time stuff that you're doing this is one of the
切到他的笔记本电脑上,展示一些你正在做的 real-time 的东西。这是其中一个
43:07
research updates that we we did recently so we've been working in in getting our general models
我们最近做的研究更新,所以我们一直在努力让我们的 general models
43:15
to different applications one of them that we think is very is very compelling is using
应用到不同的 applications 上。其中一个我们觉得非常非常引人入胜的,就是使用
43:23
general world models as essentially an interface a universal interface to software this is a
general world models 作为本质上的一种 interface,一种通往 software 的 universal interface。这是一个
43:30
a version of our world model that's called the interface world model and the idea is that it's
我们 world model 的一个版本,叫做 interface world model,思路是它
43:37
it essentially replaces the front end of a software application it renders the pixels directly
本质上替代了 software application 的 front end,它直接渲染 interface 的 pixels
43:45
of an interface and it's trained to predict what happens next as a result of a click or another
并且被训练来预测接下来会发生什么,作为点击或另一个动作的结果
43:52
interaction you have with an interface so this is all pixels it's there is no HTML CSS react that's
你和 interface 的交互,所以这全都是像素,根本没有 HTML、CSS、React 在
43:58
powering this interface this is directly at the output of our real-time video generation model
驱动这个 interface,这直接是我们 real-time video generation model 的输出
44:04
and it takes clicks directly as input and drags click and drag
它直接把点击作为输入,还有拖拽,点击和拖拽
44:11
right so it supports yeah clicks it supports drugs it's all supports scrolling
对,所以它支持,嗯,点击,它支持拖拽,还支持滚动
44:18
and the the amazing thing about this is that you can effectively describe in the prompt
而最厉害的地方在于,你可以在 prompt 里有效地描述
44:23
how you want different elements like what do you want the behavior of different elements to be
你想要不同元素怎么样,比如你想让不同元素的行为是什么样的
44:28
so it's almost your you can turn an interface from you know markup language description
所以这几乎就是,你可以把一个 interface 从,你知道,markup language 描述
44:34
of like an HTML interface and instead you can just describe the interface you know if I press
比如一个 HTML interface,转变成你直接描述这个 interface,你知道,如果我按下
44:40
this button I expect this to happen if I press this button this should happen and it's useful
这个 button,我期望的是,如果我按这个 button,这个应该发生,而且它很有用
44:45
we believe both for prototyping for like just testing like what different interactions would feel
我们相信,不管是做 prototyping,还是只是测试,比如不同的 interactions 会有什么感觉
44:49
like you can also add audio to it so it's a video audio generation model so you get essentially
就像你也可以给它加 audio,所以它是个 video audio generation model,所以你基本上能得到
44:57
can describe both what the visual outcome should be of your click and also what the if if there's a
可以描述你的 click 的 visual outcome 应该是什么,还有如果,如果有一个
45:03
sound effect that comes out of it so we believe that's going to be a much more flexible way of
sound effect 从中出来,所以我们相信那会是一种灵活得多的方式,来
45:08
building software just render it just yeah why why generate the code that generates the pixels
构建 software,直接 render 它,就直接,对,为什么,为什么要生成那些生成 pixels 的 code
45:16
just generate the pixels directly it's the end to end philosophy applying apply to to front ends
直接生成 pixels 就行了,这就是 end to end 哲学,应用到 front ends 上
45:25
so we think there is a few interesting use cases so you can build creative tools on top of it
所以我们觉得有一些有趣的 use cases,所以你可以在它上面构建创意工具
45:31
we think that you know for any kind of use case that involves a lot of exploration
我们觉得,你知道,对于任何需要大量探索的 use case
45:37
or like educational use case where you want to learn about any concept and you want some kind of
或者像是教育类的 use case,你想学习某个概念,然后你想要某种
45:43
visualization and and kind of open that exploration we think those these are very powerful
visualization,然后展开那种探索,我们觉得这些是非常强大的
45:51
approach you know you can imagine new new forms of design in kind of industrial design software
方法,你知道,你可以想象在 industrial design software 里出现新的设计形式
45:58
that could emerge as a result of of those models and this is all generated in kind of in in
这些可能会因为那些 models 而出现,而且这一切都是在某种程度上、在、在
46:06
in real time as well so you you can you can build a lot of interesting kind of camera transitions
实时生成,也是如此,所以你可以构建很多有趣的 camera transitions
46:13
and kind of forms of interaction that are very difficult to to build otherwise and one way in which
以及各种交互形式,这些用其他方式很难构建,而其中一种方式
46:19
we evaluate this is what if you try to generate the same interface with with cloud by just you
我们评估这个的方式是,如果你试着只用 cloud、只靠你自己来生成同样的 interface
46:27
have prompting cloud here's a new much reference of my interface that I made in Figma or that I
有 prompting cloud,这里有一个关于我 interface 的新参考,很多参考,是我在 Figma 里做的,或者是我
46:32
created somewhere else create this particular interaction which in this case it's you know
在别的地方创建的,来创建这个特定的 interaction,在这个情况下,它就是,你知道
46:38
drag that object upwards and beyond it being slower it's also very difficult to capture some
把那个 object 往上 drag,而且除了它比较慢之外,要捕捉某些
46:47
interactions by just fully with with just the elements so we think that this this is likely to be
interactions,仅仅完全靠这些 elements 来捕捉,也很难,所以我们觉得这,这很可能
46:54
the way that a lot of the future like software in the future will be created and one of the
会是未来很多像软件这样的东西会被创建的方式,而且其中一个
47:00
additional benefits is personalization might be a lot easier done with with those models like you
额外的好处是,personalization 用那些 models 来做可能会容易很多,就像你
47:06
can essentially try out different prompts based on whose whose visiting the interface you can
基本上可以根据谁,谁在访问这个 interface,来试不同的 prompts,你可以
47:12
more easily you know prompt engineer the the interface to have larger size tags for more
更容易,你知道,对 interface 做 prompt engineer,让它有更大尺寸的 tags,为了更多
47:19
accessibility reasons or you can make this or like if you have a particular study preferences so
出于 accessibility 的原因,或者你可以做这个,或者比如说如果你有特定的学习偏好,所以
47:25
so we're very excited about this approach is obviously early days and I think we'll need to
所以我们对这个方法非常兴奋,显然现在还处于早期阶段,我觉得我们需要
47:30
make it more cost effective as well to serve those models because you know running a real-time
让它也更具成本效益,来服务那些模型,因为你知道,运行一个 real-time
47:35
video model versus just purely rendering HTML there's obviously the composition needs are much higher
video model,而不是单纯渲染 HTML,显然 composition 需求要高得多
47:42
but we do see a lot of potential in this approach to building kind of front-end interfaces
但我们确实看到,这种方法在构建类似 front-end interfaces 方面有很大潜力
47:47
so we covered this similar thing with flipbook before with our ethernet episode with GROC video
所以之前我们和 flipbook 一起、在我们的 ethernet 那一期里、还有 GROC video,讲过类似的东西
47:53
and I think it's very engaging visually I think it's maybe very good for education but it's
而且我觉得它在视觉上非常吸引人,我觉得它可能对教育非常有用,但它
48:00
it does sound expensive I think there's an upper bound to how expensive it will be though right like
它听起来确实很贵,不过我觉得它贵到什么程度应该有个上限,对吧,就像
48:04
a you know the inference cost will go down over time you'll figure out ways to optimize it effectively
呃,你知道,inference 成本会随着时间降下来,你会找到有效优化它的方法。
48:08
when it pauses you don't you're not receiving human input you don't have to generate anything
当它暂停的时候,你不——你没有接收到人类输入,你不需要生成任何东西。
48:12
right so yeah I mean you could also like in this case you have ambient motion so there is
对,所以,嗯,我的意思是,你也可以,就像在这种情况下,你有 ambient motion,所以有……
48:19
parts of the of the screen that might you know if you're let's say you want to visit parties and
屏幕的某些部分可能会,你知道,如果你,比如说,你想去参加派对,然后……
48:26
then you you get this interface that you are walking or you have people walking or like things
然后你你会得到这个界面,你在走路,或者有人在走路,或者类似的事情……
48:31
happening but it's it's enough yeah it's obviously makes it more expensive because you need to
发生,但它已经足够了,是的,这显然让它更贵,因为你需要……
48:36
run the model all the time maybe you have some looping mechanisms so you don't need to do that
一直运行 model,也许你有一些循环机制,所以你不需要那样做。
48:41
but all those things I think it's definitely to figure out yeah I think our first consideration is
但所有那些事情,我觉得肯定得弄清楚,是的,我觉得我们的第一个考虑是……
48:46
let's make this clearly find some use cases where it's clearly much more compelling interaction
我们把这个搞清楚,找一些 use cases,让那里的 interaction 明显更有吸引力
48:52
compared to traditional interfaces and then it's a matter of time before it becomes more cost-effective
相比传统的 interfaces,然后它只是时间问题,迟早会变得更具成本效益
48:57
to serve yeah when it comes to the people walking you know I think the approach that makes the most
去服务,是啊,说到那些走路的人,你知道,我觉得最
49:01
sense to me is basically moon lake like oh I can't get my other name we're Chris Manning and
有道理的做法,对我来说基本上就是 moon lake,就像,哦,我想不起另一个名字了,我们是 Chris Manning 和
49:07
Fanny and Sun I don't know if you come across them where they basically map to some kind of
Fanny 和 Sun,我不知道你有没有遇到过他们,他们基本上会映射到某种
49:11
game engine I think it's unity or something or go go go and the you can obviously obviously
game engine,我觉得是 unity 或者什么的,或者 go go go,然后你显然显然可以
49:15
script some NPC behavior behind that and train on that whereas here you can really imagine whatever
在背后 script 一些 NPC 行为,然后在那上面 train,而在这里你可以真正想象任何东西
49:20
you want like that is a UI right like and and it feels like more tractable I guess to create a
你想要,就像那是一个 UI,对吧,而且感觉更容易处理,我猜,去创建一个
49:28
world model of software that is interactable because we have many of examples of that and you
software 的 world model,它是可交互的,因为我们有很多这样的例子,你可以,你知道,在那上面做你那些花哨的、我们自己的 environment 的东西,然后它就会 scaling up,去 embody 到真实世界的物理使用场景里,但这是一个很好的第一步;或者你知道,也有相反的情况,比如你有 1b models、350 million parameter language models,它就变得特别小,小到它们只是在预测,就像你知道的,鱼在游动。我是说 small models 不是 120b,很多都在 on device 上。但呃,不,我觉得它至少给我带来一个挺有意思的视角,至少车那个对我来说是这样,比如 applications,对吧,做那个的工作量,当然你一年只做一辆车的 model,但应用这个,也能省下手工做所有这些的成本,对吧,所以它打开了一个...
49:34
can you know do your fancy our own environment stuff on that then it is scaling up to embody
你能不能,你知道,在那上面做你自己那套花哨的、我们自己的环境之类的东西,然后它就会 scaling up 去 embody
49:40
in real world physical use cases but this is a nice first step or you know there's the opposite
到真实世界的物理用例里,但这是一个不错的第一步,或者你知道,还有相反的一面
49:44
of you have like 1b models 350 million parameter language models it just gets so small that they're
就是你会有像 1b models、350 million parameter language models,它变得太小了,以至于它们
49:51
just predicting like you know fishes moving I mean small models are not 120b so many on device
只是在预测,你知道,鱼在动之类的。我是说,small models 不是 120b,所以很多都在 on device 上
49:58
but uh no I think it like it puts an interest perspective at least the car one for me like
但是,呃,不,我觉得它,像,至少对我来说,汽车那个例子提供了一个有趣的视角,像
50:03
the applications right the amount of work to do that sure you only make one model your car per year
那些应用,对吧,做那个的工作量。当然,你每年只给你的车做一个 model
50:09
but applying this it's also a cost saving to have to manually make all this right so it opens up a
但应用这个,也能省下必须手动做所有这些的成本,对吧,所以它打开了一个
50:16
lot of possibilities too I'm curious if you extend this out to three years so where do you see things
也有很多可能性。我很好奇,如果你把这件事延伸到三年,那你会觉得事情会
50:23
going even further you know effectively the then game of something like interface world models is
走向哪里,甚至更进一步?你知道,实际上,像 interface world models 这种事情的终局就是
50:30
you have a fully neural operating system so I think uh under a car parathy has written about that
你有一个完全 neural 的 operating system,所以我觉得,呃,Andrej Karpathy 写过这个
50:37
quite quite a while back but it's uh you know I think to me it's it's a bit um it's a bit odd that
很久很久以前了,但,呃,你知道,我觉得对我来说,这有点,嗯,有点奇怪,就是
50:45
you know we have for example with an interaction with with with an LM of today you have this LM
你知道,比如我们跟今天的 LM 交互时,你有这个 LM
50:52
that can basically talk talk to you about anything it can you can take the conversation in direction
它基本上可以跟你聊任何东西,你可以把对话带向某个方向
50:57
you can it's very general so it can solve all those different tasks but you interact with it
你可以……它非常通用,所以它能解决所有那些不同的任务,但你跟它交互
51:02
through a very rigid interface and so to me it's just a matter of time before the interface itself
是通过一个非常僵化的 interface,所以对我来说,这只是时间问题,interface 本身……
51:07
becomes learnable and becomes you know part of the the whole loop of like you're not just delivering
变得 learnable,然后变成,你知道,整个 loop 的一部分,就像你不只是在交付
51:13
you're delivering an application end to end and that means you're delivering the language
你是在 end to end 地交付一个 application,这意味着你在交付 language
51:16
model but you also delivering the the renderer and the pixels and then and that's also a learnable
model,但你也在交付 renderer 和 pixels,然后,而且那也是 learnable 的
51:22
component and the concept of applications might not necessarily I think we'll need to figure out
component,而 applications 的概念可能不一定,我觉得我们需要搞清楚
51:28
new abstractions for software the the concept of application comes from this idea that you need
software 的新的 abstractions,application 的概念来自这个想法,就是你需要
51:35
you know separate kind of code basis to describe to for to power each individual
你知道,单独的某种 code basis,用来描述、用来驱动每一个单独的
51:41
tool and each individual application but you might you might think of something a lot more unified
tool 和每一个单独的 application,但你可能会,你可能会想到一个统一得多的东西
51:48
if you're if you have a video model that's actually generating the interface as you go
如果你,如果你有一个 video model,它实际上是在一边生成 interface
51:53
so it can take context from an LM and allow you to combine kind of different different
所以它可以从一个 LM 中获取 context,并允许你组合不同种类的
52:00
functional this is a traditional with live in different applications so so it's a way to solve
功能上,这是一种传统的,可以在不同的应用中运行,所以它是一种解决
52:05
you know software and to end effectively we also see this as a powerful way to train computer
你知道,软件端到端地有效地,我们也把这看作一种强大的方式来训练 computer
52:12
use agents as well so this is you know one way to see this as and in general with world models
use agents 也是,所以这是你知道的一种看待方式,并且一般来说,对于 world models
52:19
there is those two directions one is world models for humans and world models for for agents to
有那两个方向,一个是给人类的 world models,另一个是给 agents 的 world models,用来
52:24
train agents and so for every new work of world models that we do we have this with these both
训练 agents,所以对于我们做的每一个新的 world models 工作,我们都有这个,这两种
52:31
uses become possible so this is a powerful synthetic data generator for training computer use
用途都成为可能,所以这是一个强大的 synthetic data generator,用来训练 computer use
52:36
models it could become a live RL environment that you could use to do online RL with with a computer
models,它可以成为一个 live RL environment,你可以用它来和 computer 做 online RL
52:44
use agent and you can get wide diversity of different interactions kinds of interfaces just
使用 agent,你就能得到各种各样的不同 interactions,各种 interfaces,就
52:50
generate on the fly that to improve the how robust the your your your agent becomes so that's
generate on the fly,以提高你的你的你的 agent 变得多么 robust,所以那就是
52:57
the same also with the world models that we're working on for robotics use case as well
同样的,也适用于我们正在为 robotics use case 开发的 world models
53:02
is there research breakthrough that you're waiting for that would unlock the next set of use cases
有没有你正在等待的 research breakthrough,能解锁下一组 use cases
53:08
that you really want to pursue long context is very important one so being able to maintain
你真正想追求的那些。long context 是非常重要的一个,所以能够保持
53:14
consistency for a long periods of time and that depends on the use case so for our characters model
长时间的 consistency,这取决于 use case,所以对于我们的 characters model
53:20
for example or for the interface world model it's easier to maintain long sessions of interaction
例如,或者对于 interface world model,更容易保持长时间的 interaction sessions
53:25
if you go into more open and that world that you navigate and you take arbitrary actions in we
如果你进入更开放的世界,你 navigate,你采取 arbitrary actions,我们
53:32
like there is more the context of which you can and duration which you can generate
比如说,你能用的 context 更多,能生成的时长也更长
53:38
becomes limited much more quickly so we see more degradation and air accumulation happening
就会更快地变得受限,所以我们看到更多 degradation 和 air accumulation 发生
53:43
so the biggest challenge with auto regressive models is air accumulation is basically
所以 auto regressive models 最大的挑战就是 air accumulation,基本上就是
53:48
you're feeding generative frames back into the models to generate the next the next frames
你把生成的 frames 再喂回模型里,来生成下一帧、再下一帧
53:52
and if there is any small errors the accumulate over time that's not a new problem it's a problem that
而如果有任何小错误,它们会随着时间累积,这不是一个新问题,而是一个问题,
53:58
LMS also have and we've seen the you know the ability to generate now really really long outputs
LMS 也有,而且我们已经看到,你知道,现在生成非常非常长输出的能力
54:05
so it's a solid problem but it's definitely still a slow challenge yeah and what is the state of the art
所以这是个很实在的问题,但肯定仍然是个缓慢的挑战,对,那 state of the art 是什么
54:09
so for for GROC it would be like 10 to 20 seconds of context going in there for video
所以对 GROC 来说,大概是 10 到 20 秒的 context 输进去,用于视频
54:15
with the archives or smalls were able to generate up to 30 minutes of video author aggressively
用 archives 或 smalls 能生成最长 30 分钟的视频,或者就是那么激进
54:21
yeah but that's just for the advertisers yeah so if we look at the GWM worlds which is more
对,但那只是给广告主用的。对,所以如果我们看 GWM worlds,它更像是
54:28
our open-ended world exploration model it's on the order of a few minutes which is yeah so
我们的 open-ended world exploration model,它差不多就几分钟,对,所以
54:35
probably enough for people because you have to cut to the next scene anyway right yeah it's not
对人来说可能够了,因为反正你总得切到下一个场景,对吧?对,它不是
54:40
it's not the ideal game experience if you have to restart every few minutes so I think but I think
如果你每隔几分钟就得重启,那并不是理想的游戏体验,所以我觉得,但我觉得
54:46
it's yeah for certain kinds of game experiences you can you can work around it ideally
对,对于某些类型的游戏体验,你可以,你可以绕过它,理想情况下
54:51
you're able to just generate forever and it doesn't it doesn't degrade and I think that's a
你能一直生成下去,而且它不会,它不会退化,我觉得那只是
54:57
matter of time before we get there yeah genie has like one max one minute you know yeah this was your
时间问题,我们迟早会到那一步。对,genie 差不多最多也就一分钟,你知道吧,对,这是你的
55:03
you did a study on robotics I think I also have just your runway robotics page though
你做过一项关于 robotics 的研究吧,我记得。不过我这边也刚好有你们 Runway 的 robotics 页面。
55:10
is this better so last year we released gen 4.5 so that was our latest base model and
这样更好吗?所以去年我们发布了 gen 4.5,那是我们最新的 base model,然后
55:16
we've been as I mentioned we've been doing all this work in world models and
就像我提到的,我们一直在 world models 上做所有这些工作,而且
55:20
which essentially a lot of our approach to world models is how do you take a bi-directional
本质上,我们在 world models 上的很多方法就是,你怎么把一个 bi-directional
55:26
diffusion model and make it author-aggressive and make it accept actions so instead of being a video
diffusion model 做得 author-aggressive,并让它能接受 actions,所以它不再是一个视频
55:31
you watch it becomes a as a simulation that you step in and you can you know control it every step
你观看的视频,而变成一个你可以踏入的 simulation,你可以,你知道,在每一步都控制它
55:36
of the way you can explore counterfactuals like what happens if I take this action versus if I take
整个过程你都可以探索 counterfactuals,比如如果我采取这个 action 会怎样,对比如果我采取
55:41
this action and GWM1 was the it's the world model that we built on top of gen 4.5 so we did all
这个 action。而 GWM1 就是——它是我们建立在 gen 4.5 之上的 world model,所以我们做了所有
55:49
this author-aggressive and like distillation of the other-aggressive and then step distillation
这个 author-aggressive,有点像 other-aggressive 的 distillation,然后是 step distillation
55:55
on top of gen 4.5 and one of the biggest use case that we saw for GWM1 was in robotics one thing
在 gen 4.5 之上,而且我们看到 GWM1 最大的 use case 之一是在 robotics,有一件事
56:02
we like to say is we as we scaled video models we accidentally created one a state of the art model
我们喜欢说的是,我们在 scaling video models 的时候,不小心做出了一个 state of the art model
56:09
for robotics by just scaling video models so we realized at some point mid and last year that
给 robotics 用的,只是靠 scaling video models,所以我们去年年中某个时候意识到
56:17
robotics labs that are coming up to us and kind of asking to use video models for synthetic data
robotics labs 会来找我们,有点像在问能不能用 video models 来做 synthetic data
56:23
kind of asking us to post-training or video models to work really well for robotics so that they
有点像在让我们对我们的 video models 做 post-training,让它们在 robotics 上表现得特别好,这样他们
56:28
can they can use that to basically generate variations that was the first use case that we saw
就可以用那个来基本上生成各种 variations,那是我们看到的第一个 use case
56:33
and then increasingly became clear that the models will be useful beyond just creating synthetic
然后越来越清楚的是,这些 models 的用处会超出只是创建 synthetic
56:38
data to train robotic policies they would also be very useful as simulators so that means that you
用来训练 robotic policies 的数据,它们作为 simulator 也会非常有用,所以这意味着你
56:45
can use a video model online to test how your robotic action model performs so you can take an
可以在线用一个 video model 来测试你的 robotic action model 表现如何,所以你可以采取一个
56:52
action role and then get the outcome of the action inside the world model and then continue that loop
action role,然后在 world model 里面得到这个 action 的结果,然后再继续这个 loop
56:58
like this close loop simulation and you can use that to evaluate how well your robotics model works
就像这样的 close loop simulation,而且你可以用它来评估你的 robotics model 表现有多好
57:06
and the biggest thing that I think you need to solve if you want to build a simulator is establishing
而我觉得,如果你想构建一个 simulator,你需要解决的最大问题就是建立
57:11
real world correlation that if you take an action inside the world model if you take the same action
real world correlation,也就是说如果你在 world model 里面采取一个 action,如果你在
57:16
in the in the real world you get a similar outcome so that was that was the goal of some work that
真实世界里采取同样的 action,你会得到类似的结果,所以那就是我们今年早些时候
57:22
we did earlier this year so if you go to the first link so that was essentially wanted to establish
做的一些工作的目标,所以如果你点第一个链接,那本质上就是想要建立
57:28
that you know real real to seem correlation for a world model so that if you do a series of actions
就是,你要有一个 real-to-sim correlation 给一个 world model,这样如果你做一系列动作
57:35
inside the world model and if you do the same actions in the real world you get similar outcomes
在 world model 里面,如果你在真实世界里做同样的动作,你会得到相似的结果
57:39
and we took our GWM1 model and we we use some benchmark data that there is this rubberina
然后我们拿了我们的 GWM1 model,我们用了些 benchmark 数据,就是有这个 rubberina
57:47
benchmark that's very commonly used to evaluate how well do different action models perform
benchmark,它很常用来评估不同的 action model 表现有多好
57:52
and we use the same scenarios and settings and embodiments inside our world model and we measure
我们在我们的 world model 里用同样的场景、设置和 embodiments,然后我们测量
57:59
the correlation of how well did the action model perform inside the world model versus in the real world
action model 在 world model 里表现有多好,对比在真实世界里的表现,这其中的 correlation
58:04
and we saw that we could get very good correlation between between our world model and reality
然后我们发现,我们的 world model 和现实之间能得到非常好的 correlation
58:10
and that means that if you want to evaluate how well your robotic policies perform
这意味着,如果你想评估你的 robotic policies 表现得有多好
58:14
you can scale that much faster inside simulation instead of instead of having to do that
你可以在 simulation 里把这件事 scale 得快得多,而不是,而不是非得去做那件事
58:20
with actual physical hardware and so that was a first indication that our models could be quite
用实际的 physical hardware 来做,所以那是第一个迹象,说明我们的 models 可能相当
58:28
useful in robotics and we saw as we were working with robotics labs that that became like the first
在 robotics 里很有用,而且我们在和 robotics 实验室合作时看到,那差不多成了第一个
58:33
use case where they could use video models in a way that fit into their their kind of training pipeline
use case,让他们能以契合他们那种 training pipeline 的方式使用 video models
58:40
can I ask what the difference was from 4.5 to solving that so the same the real gap has always
我能问一下,从 4.5 到解决那个问题,区别是什么吗?所以,同样的,真正的 gap 一直都是
58:46
been the issue right you train a robotics model and video data it doesn't generalize the real world
问题所在,对吧?你用 video data 训练一个 robotics model,它没法 generalize 到 real world
58:51
and the simulation had an issue so seems like you solved it but how yeah so a big problem with
而 simulation 有个问题,所以看起来你把它解决了,但怎么做到的?对,所以一个大问题是
58:57
simulators is you know if if you're trying to simulate rigid objects like it works quite well
simulators 的话,你知道,如果,如果你是想 simulate rigid objects,那它还挺好用的
59:03
if you know if you if you can describe the physics of objects very accurately then you're able to
如果你知道,如果你能非常准确地描述物体的 physics,那你就能用 Isaac same 或 Mijok,或者某个传统 simulator,但对于更复杂的交互,比如说布料,或者像湿滑表面,你知道,带着所有那些你希望能用 manipulation、用一个能解决 manipulation 任务的 action model 来解决的复杂性,要为每一个这样的环境构建模拟版本、构建那个环境的 digital domain,是非常困难而且很耗时的。而有了 world model,你只需要提供 first frame,然后你就可以在 first frame 里 roll out 那些 poles。所以,你知道,我们把它和这些 methods 做了比较。
59:09
use Isaac same or Mijok or one of the traditional simulators but for more complex interactions with
使用 Isaac same 或 Mijok,或者传统 simulators 之一,但为了更复杂的交互,跟
59:18
with a cloth for example or like slippery surfaces you know with all the complexity that you you
比如说用一块布,或者像那种很滑的表面,你知道,带着所有那些复杂性,你你
59:26
want to be able to solve with a manipulation with an action model that solves manipulation tasks
想要能够用 manipulation 去解决,用一个能解决 manipulation tasks 的 action model
59:31
it's very difficult in soft time consuming to build you know if for each of those environments and
这非常困难,在 soft 里很耗时,去构建,你知道,如果针对每一个那些环境,而且
59:37
these are those tasks built the simulated version of that the digital domain of that of that
这些是那些任务,构建了那个的 simulated version,那个的 digital domain,那个那个
59:42
environment whereas with the world model you just need to provide the first frame and then you
环境,而用 world model 你只需要提供 first frame,然后你
59:47
just can roll out the poles inside the first frame so whereas you know we compared it to methods
就可以在 first frame 里 roll out 那些 poles,所以呢,你知道,我们把它和 methods 比较
59:52
that required like 3D scanning environment and then 3D scanning each individual object before you
那需要像先对环境做 3D scanning,然后对每个单独的物体做 3D scanning,之后你才能
60:00
can now you know you can bring a dissimulation whereas with a world model you just take a picture
现在你懂的,你可以搞一个 simulation,而有了 world model,你只要拍一张照片
60:07
of the of the environment and then and then you're able to test how how your policy performs
拍这个环境的,然后然后你就能测试你的 policy 表现如何
60:12
our general thesis on robotics is you know there is companies that are leveraging a lot of
我们在 robotics 上的总体论点是,你懂的,有些公司正在利用大量
60:18
teleoperation data to train robotics action models there is now companies aren't using
teleoperation data 来训练 robotics action models,现在有些公司并不使用
60:25
UMI data which is essentially human egocentric video where humans use kind of robotic
UMI data,它本质上就是人类 egocentric video,里面人类会用某种 robotic
60:34
reapers to perform the form manipulation tasks and then there is companies that are focusing on
reapers 来执行 form manipulation tasks,然后还有公司专注于
60:39
egocentric data which is you know you strap a GoPro on someone's head and then you kind of
egocentric data,也就是你懂的,你在某人头上绑一个 GoPro,然后你就有点像
60:45
capture them performing a task we think that you know and all those are great source of data for
捕捉他们执行任务,我们认为,你知道,而且所有那些都是很好的 data 来源,对于
60:50
training robotics models but the most plantable source of video data is their person video data it's
训练 robotics models,但最 plantable 的 video data 来源是他们的 person video data,它是
60:56
and if how how do we as humans learn how to perform different tasks a lot of it is by observing
那,我们人类是怎么学会执行不同任务的?很大程度上是通过观察
61:03
others perform those tasks we don't learn from first person we obviously do some trial and error
别人做这些任务。我们不是从第一人称学习的,我们显然会做一些试错
61:08
and like to learn different things but ultimately a lot of what we learn how to do in the
也喜欢学不同的东西,但归根结底,我们在
61:14
world we learn by watching other people do it and that's how when you pre-training a video model
世界上学会做的很多事情,都是通过看别人做来学会的,而这就是当你对一个 video model 做 pre-training 时
61:19
you're essentially doing that it's a lot of third-person video footage of people performing different
你本质上就是在做那件事——它包含大量第三人称视频素材,内容是人们在执行不同
61:25
tasks in the world people doing sports people doing household tasks and the our main thesis is that
任务:在世界上,人们做运动、做家务。而我们的主要论点是,
61:31
video pre-training once you do that you can then adapt a model to be useful in robotics use cases
video pre-training:一旦你做了这个,你就可以 adapt 一个 model,让它在 robotics use cases 里有用,
61:37
with lay away fewer hours of actual robotic data so you require way less calibration data which is
用少得多的实际 robotic data 小时数,所以你需要的 calibration data 也少得多,
61:44
very difficult to scale and even if you look at egocentric data which is a bit more easy to scale
而 calibration data 很难 scale;即使你看 egocentric data,它相比需要实际 hardware 的 television data 更容易 scale,
61:51
compared to television data which requires actual hardware is still three hours of magnitude
它在世界上存在的量仍然比 third-person video data 少三个数量级,所以我们的 thesis 是,
61:57
less of that that exists in the world compared to third-person video data out there and so our thesis
而且一般来说,最丰富的数据来源最终会赢,third-person video data pre-training 才是正确的起点,
62:04
and generally like the most plentiful source of data will ultimately ultimately win their
对于那种你知道你希望它们能 generalize、能处理新环境、新任务、没见过的东西的 models 来说。
62:09
person video data pre-training is the right starting point for models that you know you want
person video data pre-training 是合适的起点,对于那些你知道你希望
62:15
them to generalize and be able to deal with new environments new tasks things that you haven't seen
它们能够 generalize,并能处理新环境、新任务、你还没见过的东西的模型来说。
62:21
during training that's the motivation for why we think our models are especially useful in
在 training 期间,这就是我们为什么认为我们的模型特别适用于
62:27
robotics settings and we've seen that to be the case as well you said pre-training there so
robotics 场景的动机,而且我们也看到了确实如此。你刚说到了 pre-training,所以
62:33
maybe it's like third-person pre-training first-person SFT is it like a curriculum that you can
也许这就像是第三人称的 pre-training、第一人称的 SFT?这是不是像一种你可以
62:39
sort of introduce exactly so if we look at GWM worlds so GWM robotics so GWM robotics it starts
逐渐引入的 curriculum?没错。那如果我们看 GWM worlds,也就是 GWM robotics,GWM robotics 它是从
62:48
from gen 4.5 it's the same video diffusion backbone right exactly yeah so you start from the
gen 4.5 开始的,用的是同一个 video diffusion backbone,对吧?没错,是的,所以你从
62:54
base video model the one you're using to generate cats and dogs and other interesting stuff and then
base video model 开始,就是那个你用来生成猫、狗和其他有趣东西的模型,然后
63:00
you find you on a very small number of hours of robotic data so it's something on the order of
你在非常少的 robotic data 小时数上 fine-tune,所以大概是
63:09
kind of hundreds of hours compared to if you were to pre-training robotics model the current
几百个小时这个量级,相比你要 pre-training 一个 robotics model 的话,当前
63:14
pre-training go up to you know 100,000 or like millions of hours of data and you're able to get
pre-training 会用到,你知道,最多 10 万甚至几百万小时的 data,然后你就能很快得到
63:20
quite quite good performance quickly because the model leverages all the things that he has
相当相当好的 performance,因为模型会利用它学到的所有东西,
63:27
learned about the world the physics and human dynamics and the tasks that people care about
关于世界、物理、人类动态,以及人们关心的那些任务,
63:32
from pre-training and ultimately you you want those models to generalize you don't want to just
这些都是从 pre-training 来的。最终你、你就是希望这些模型能 generalize,你不希望它们只是
63:38
be able to perform the tasks that they have been during training and the diversity of actions
能执行它们在 training 期间做过的那些任务。而动作的多样性
63:44
and environments that you have with a pre-training video data set is much larger than what you can
以及你在一个 pre-training video data set 里拥有的环境多样性,远远超过你能
63:50
kind of realistically capture manually how's the scale looking like for the post-training like
现实中手动捕捉到的程度。post-training 的 scale 看起来怎么样,比如说
63:57
you still want to do is it like roughly 90% of the compute in regular video diffusion model and
你还是想做,那它大概是常规 video diffusion model 里 90% 的 compute 吗,然后
64:02
then scale up a lot or do it like we want different robotic models for different tasks or just
然后大幅 scale up,或者像我们想要的那样,为不同任务做不同的 robotic models,或者只是
64:09
the one base really good world model can also apply to robotics so currently we are post-training
那个单一的、非常好的 base world model 也能应用到 robotics,所以目前我们正在 post-training
64:16
our models for specific embodiments that we for particular partners so if they have a particular
我们的模型,针对特定 embodiments,是我们为特定合作伙伴做的,所以如果他们有一个特定的
64:24
kind of single arm robot or by manual robot or humanoid robot we would post-traine our GWM
那种 single arm robot,或者 by manual robot,或者 humanoid robot,我们就会 post-train 我们的 GWM
64:30
robotics model on their particular data set over time we see the different variants of GWM
robotics model 在他们特定的 data set 上,随着时间推移,我们看到 GWM 的不同变体
64:35
unifying like I would expect you know if a year from now or two years from now you have a
统一起来,就像我预期的那样,你知道,如果一年后或者两年后你有一个
64:41
single world model that can that can simulate manipulation tasks it can simulate navigation
单一的 world model,它可以,可以模拟 manipulation tasks,也能模拟 navigation
64:47
which is a lot of the gaming world models are navigational world models you're moving around the
其中很多 gaming world models 都是 navigational world models,你在四处移动
64:52
space and it will also simulate kind of human behavior so that's the character models so instead
空间,而且它还会模拟某种人类行为,所以那就是 character models,所以不是
64:57
of having three different models you have a single model that's able to you know ideally you're
有三个不同的模型,而是有一个单一模型,它能够,你知道,理想情况下你
65:02
able to simulate what it's like to be in the world you're moving out in the environment you're
能够模拟身处这个世界是什么感觉,你在环境里移动,你
65:08
maybe performing different tasks you're talking to other people and that happens with the same
也许在执行不同任务,你在和其他人交谈,而这一切都发生在同一个
65:13
a single kind of real-time video model that's generating that do you think you can solve self-driving
单一的那种 real-time video model 生成的过程中。你觉得你能解决 self-driving 吗
65:19
so if you are learning to drive a car in a simulator you have a world model your robot is basically
所以如果你是在 simulator 里学开车,你有一个 world model,你的 robot 基本上
65:26
you know car can manipulate so many axes how far off are you from something like that so
你知道,车可以操控这么多 axes,你离那样的东西还有多远,所以
65:31
really good data system world models are definitely being applied to self-driving research
非常好的 data system,world models 肯定正在被应用到 self-driving 研究里
65:38
right now mainly for evaluation use cases but our focus has been more on the robotic manipulation
现在主要是用于evaluation的场景,但我们的重点更多在robotic manipulation上
65:45
we've done some work on AV world models as well but yeah we do think that world models are
我们在AV world models上也做了一些工作,但确实我们认为world models
65:53
and video models are the best starting point for both simulators and also policy and the action
和video models是simulators以及policy和action models最好的起点
65:58
models so that's that's the other side to this is that once you have a great world model then you
所以这是这件事的另一面,就是一旦你有了一个很好的world model,你
66:04
can just add an action head and it can predict actions as well one way to think about it is if you
就可以直接加一个action head,它也能预测actions。一种理解方式是,如果你
66:10
take the starting frame of a scene with a robotic arm and you ask you know you prompt the model
取一个带有robotic arm的场景的起始帧,然后你prompt模型
66:16
generate the arm picking up an object it would and if it generates an accurate enough video then
生成手臂拿起一个物体,如果它生成的视频足够准确,
66:22
it should also be able to generate the exact poses in 3D that the arm should take to to perform
那它也应该能够生成手臂执行动作时应该采取的精确3D poses
66:29
the same action so this is the direction that's now the the popular term for it is world action
同一个 action,所以这就是现在的方向,现在流行的叫法就是 world action
66:35
models which is you're starting from a video model and then you're adding an action head to predict
models,也就是你从一个 video model 开始,然后加一个 action head 来预测
66:40
the actions and it becomes a policy essentially one thing I'm also impressed by is how much data
actions,本质上它就变成了一个 policy。还有一点让我印象很深的是,到底需要多少数据
66:47
you actually need to train these kinds of models you probably can't say exactly how much but like
你实际训练这类模型需要多少,你可能没法说确切是多少,但就像
66:52
you know the the original diffusion models and from what I know even of the open source Chinese
你知道,最初的 diffusion models,而且据我所知,甚至那些 open source Chinese
66:58
models is not that much data isn't it surprising what do you define as the much data yeah it just
models,其实并没有那么多数据,是不是挺让人惊讶的?你说的“那么多数据”是怎么定义的?对,它就是
67:05
counts goes into that is a token count still relevant so it's a bit more complicated and
counts 算进去,那 token count 还相关吗,所以这就有点复杂,而且
67:11
what is I mean you just gigabytes right yeah I feel like something that's interesting is it seems
什么,我的意思是你就是指 gigabytes 对吧?对,我觉得有意思的一点是,看起来
67:19
like the let's call it tokens to param counts in language models has really you know maybe
就像那个,我们姑且叫它 language models 里 tokens 到 param counts 的比值,真的,你知道,也许
67:25
there three years ahead or whatever seems to be a lot higher than video models still even though
它们领先三年或者什么的,看起来还是比 video models 高很多,尽管
67:30
technically video has more information you know per per a bit I don't know if it seems intuitive or
从技术上讲,video 有更多信息,你知道,每个每个 bit,我不知道这看起来直不直观,或者
67:37
maybe there's just a lot of like the variability between a pixel to the next pixel is not that high
也许只是有很多,像,一个 pixel 到下一个 pixel 之间的变化并没有那么高
67:44
so like maybe there's just a lot of information that is repeated my answer would be it's still very
所以,就像,也许只是有很多信息是重复的。我的回答会是,它现在还非常
67:49
early like like the training video models will scale way further than it is and and you'll have
早期,就像,就像 training video models 会 scale 得比现在远得多,而且而且你会有
67:57
capabilities that go much further than the car models can do so one thought experiment that I
capabilities,会走得比 car models 能做的远得多。所以有一个 thought experiment,我
68:03
like to use it's it's almost like the Turing test of video models are like the Turing test of world
喜欢用的,它,它几乎就像 video models 的 Turing test,就像 world 的 Turing test
68:08
models go I call it the lucid dream test is you have I mean like the actual person lucid lucid dream
模型会——我把它叫 lucid dream test,就是你有,我是说像真人那种 lucid、lucid dream
68:17
it comes from this is a lucid dreams right lucid dreams is telling your dreaming while you know
它来自这个:这就是 lucid dreams,对吧,lucid dreams 就是你在做梦的时候知道自己在做梦
68:22
yeah exactly so lucid dreaming is when you realize your dream inside a dream and then you basically
对,没错,所以 lucid dreaming 就是你在梦里意识到自己在做梦,然后你基本上
68:27
play around be able to control what happens no there's also an inference guy called lucid dreams yeah
可以随便玩,能控制发生什么。不,还有个搞 inference 的家伙叫 lucid dreams,对
68:32
oh quite a decision very prolific yeah so let's say you have a VR headset and and you're in a room
哦,挺有决断的,非常多产,对,那假设你有一个 VR headset,然后,然后你在一间房间里
68:40
with and you're wearing a VR headset and that VR headset you know most of today's VR headset have
里面,然后你戴着 VR headset,而那个 VR headset,你知道,现在大多数 VR headset 都有
68:46
a past room out so you can see directly what's in front of you in the world or you can obviously
一个 passthrough,所以你能直接看到现实世界里你面前的东西,或者你显然可以
68:52
render kind of something inside inside the VR headset and there's going to be a point where
在 VR headset 里面渲染出某种东西,然后会有那么一个时刻,那时候
68:57
those interactive real-time video models become good enough where you wear the headset
那些 interactive real-time video models 变得足够好,好到你戴上 headset
69:03
and you're in the same room and you're walking around and you're kind of and you're interacting
然后你在同一个房间里,到处走动,你有点,你在跟
69:08
with objects you're able to kind of move freely in that room and do and interact with any object
物体互动,你能在那个房间里自由移动,能做事情,能和任何物体互动
69:14
and at the end someone asked you did you were you using past remote or were you actually or was
最后有人问你,你当时是在用 past remote,还是你其实,还是说
69:20
this a rendered or generated footage and if you cannot tell for sure if that was what you were
这是 rendered 还是 generated footage,而如果你没法确定,那到底是不是你当时
69:28
seeing as you interacting with and moving around the world was generated or it was kind of
看到的,当你在这个世界里互动、移动时,那些是 generated 的,还是说它有点
69:34
past room out and was just what was happening in front of you that's an indication that the
past room out,就只是你眼前正在发生的事,那就说明那些
69:38
malls have become good enough and we're not we're not close to that yet and a lot of it is just this
malls 已经变得足够好,而我们还没,我们还没接近那一步,而且很多都只是这个
69:43
idea of really simulating dynamics and counterfactuals well like you know if you ask a video model to
真正把 dynamics 和 counterfactuals 模拟得很好这个想法,就像你知道,如果你让一个 video model 去
69:50
generate a person scoring a goal versus a person failing to score a goal it would do a better
生成一个人进球和一个人没进球的画面,它在生成进球这件事上会做得更好
69:56
job at scoring the goal because there is a bias from the training distribution there's a lot more
因为 training distribution 里有一种 bias,有更多
70:01
videos of the person succeeding at scoring the goal but if you have an interactive model you want
这个人成功进球的视频,但如果你有一个 interactive model,你会希望
70:06
it to be able to generate counterfactuals like if I take this action versus this action you want
它能生成 counterfactuals,就像如果我采取这个 action 还是那个 action,你会希望
70:12
to generate equally realistic outcomes so that's I think the big gap between video models and and
生成同样真实的结果,所以我觉得这就是 video models 和
70:18
what models is that idea of the counterfactual generation and if you want a great model for
world models 之间的大差距,就是 counterfactual generation 这个想法,而如果你想要一个很棒的模型用于
70:23
robotics you want to simulate failure very well because whether you're using it for evaluation or
robotics,你会想把失败模拟得非常好,因为不管你是用它来做 evaluation 还是
70:29
you're using it as an online RL loop in the future you want to be able to have the model kind of try
你是在把它当作一个 online RL loop 来用,未来你希望能让 model 去尝试、失败、再改进,所以要做到这一点,你需要能够 simulate 事情失败。
70:35
and fail to do things and and improve and so in order to do that you need to be able to simulate
这是唯一一个你有太多成功案例、却没有足够失败案例的领域,应该很容易生成失败。
70:41
things failing. This is the only domain where you have too many successful examples and not enough
我觉得像早期的 image video models 并不擅长做到像真人一样,对吧,你看到很多 high res 4k 的,像专业摄影,但不只是日常生活,像普通照片,对吧,所有东西看起来都像是专业生成的,像专业照片,但不只是普通的那种,你知道,桌子上乱糟糟的线缆。
70:46
bad examples should be easy to generate failure. I think like early image video models weren't good
好吧,所以有这些东西,你知道,我们还聊到过一件事,就是你们有视频。
70:55
at being human realistic right like you see a lot of the high res 4k like professional photography
在做得像人、逼真这点上,对吧,你会看到很多 high res 4K 的,就像专业摄影一样
71:02
but not just everyday life like normal picture right everything looks like it's professionally
但不只是日常生活,像普通照片,对吧,所有东西看起来都像是专业
71:07
generated like professional pictures but not just like normal like you know messy cables on a desk
生成的,就像专业照片,但不只是那种普通的,就像你知道,桌上乱糟糟的线缆
71:13
okay so there's there's this stuff you know one thing we also covered that you guys have video
好,所以有这些东西,你知道,我们还聊到过一件事,就是你们有 video
71:18
agents that you launch I guess how does the traditional let's call it frontier like you know
你启动的那些 agents,我猜,传统上,我们就叫它 frontier,怎么说呢,你知道
71:26
auto regressive LMS like feed in you know to to all this they're driving your robotics models or
auto regressive LLMs,就像喂进去,你知道,到所有这些里面,它们在驱动你的 robotics models,或者
71:32
they're driving other so your video agents production anything where you see the overlap
它们在驱动其他东西,所以你的 video agents、production,任何你看到重叠的地方
71:38
auto regressive and diffusion is called it yeah so harnesses are really important across all
auto regressive 和 diffusion,是这么叫的,对,所以 harnesses 在所有
71:43
those different use cases so we have this video agent which is essentially an LLAM that is
那些不同的 use cases 里都非常重要,所以我们有这个 video agent,它本质上是一个 LLM,
71:48
very effective at tool use of different image models video models and kind of helps you through
它非常擅长对不同 image models、video models 做 tool use,而且某种程度上帮你一路
71:53
creating a project end to end so you know very often in like a traditional kind of advertising flow
端到端地创建一个项目,所以你知道,很多时候在一种传统的 advertising flow 里
72:00
you have a brief you start from it and then you generate some a storyboard and then you generate
你有一个 brief,你从它开始,然后你生成一些 storyboard,然后你生成
72:04
a video a video agent or run away agent kind of helps you through that whole process and it helps
一个 video agent 或者 run away agent 差不多能帮你走完整个流程,而且它还能帮你
72:10
you also analyze performance data for example how well did this outperformers decide and then
你还能分析 performance data,比如说这些 outperformers 决定得怎么样,然后
72:17
generate me more of the based on those learnings figure out what what what to generate we think that
基于这些 learnings 给我多生成一些,搞清楚到底要生成什么、什么、什么。我们觉得
72:23
the harness is a very important piece of the pipeline as I mentioned all the video production
这个 harness 是 pipeline 里非常关键的一环。就像我提到的,所有 video production
72:30
all the production video models use some prompt completion that happens and we expect you know
所有 production video models 都会用到某种 prompt completion,然后我们预期,你知道
72:36
that to become more and more complex and more you know you generate longer and more detailed
它会变得越来越复杂,而且越来越,你知道,你会生成更长、更详细的
72:40
descriptions before you use the the diffusion transformer I do think eventually you know there's
描述,然后你才会去用这个 diffusion transformer。我确实觉得最终,你知道,会有
72:46
increasingly this unification into omnimodels where you have the you're training the models end to end
越来越朝着 omnimodels 统一,也就是你有一个……你在 end to end 训练模型
72:53
to both do authoritative text prediction and also diffusion as well so you're predicting the next
既能做权威性的 text prediction,也能做 diffusion,所以你在预测下一个
72:59
token of like you're doing maybe using some reasoning and planning of the scene and then
token,比如说你可能在做一些 reasoning,以及针对场景的 planning,然后
73:04
you're passing it into the diffusion head that's actually generating generating the pixels
你把它传进 diffusion head,它实际上在 generating、generating pixels
73:10
yeah I think it's currently maybe only Gemini and Quinn do it I'm not sure which of the Chinese
对,我觉得目前可能只有 Gemini 和 Quinn 这么做,我不确定哪些中国的
73:17
models are omnip but yeah it's not it's not a very well popularized modality I guess
models 是 omnip,不过对,它不是一个、不是一个很普及的 modality,我猜
73:25
it's an interesting use case when you think about it because not only do you have to end at
仔细想想,这其实是个挺有意思的 use case,因为你不仅得停在
73:29
like language model reason diffusion head generate you don't have to output there you can go back
比如 language model reason、diffusion head generate 这里,你不一定要在那里 output,你可以回去
73:35
into feed that output to the same model reason again on improvements and it can do a lot of loops
把那个 output feed 给同一个 model,再对改进做 reason,而且它可以跑很多 loops
73:43
just in its own I guess the question is like do we need that or can we just do agents scaffold
单就它本身来说,我猜问题是,我们到底需不需要那个,还是直接做 agents scaffold 就行?
73:50
like to it outside the model is there a big benefit to doing it in I think there's generally the trend
比如说放在 model 外面做,那放在 model 里面做有很大好处吗?我觉得总体上有一种趋势
73:56
of something is first done by a harness and then it comes part of the model right so you had
某个东西先是靠 harness 来做,然后它变成 model 的一部分,对吧,所以你有
74:02
chain of thought prompting where you had to do this super detailed system prompts to
chain of thought prompting,那时候你得写这种超级详细的 system prompts 来
74:07
take it up and now the the model basically generates the reasoning trace by itself before it gives
把它拉上去,而现在 model 基本上会自己生成 reasoning trace,然后才给出
74:12
you an answer and in the in video models similarly a lot of the video models of kind of the
你答案,而在 video models 里,类似地,很多 video models 有点像
74:17
the early days were single shot video models and you had to use some kind of orchestrators it turn
早期那种,都是 single shot video models,你得用某种 orchestrators,然后
74:22
you know generate multiple shots in parallel and then turn it into an actual video
你懂的,并行生成多个 shots,然后把它变成真正的 video
74:28
I mean comfy UI just all over the you know all this nodes by getting workflow
我是说,comfy UI 就是到处都是,你知道,全靠这些 nodes 来搭 workflow
74:34
and now you have multi-shot video generation where you have the you directly generate multiple
现在你有了 multi-shot video generation,你可以直接生成多个
74:40
shots and there is a benefit to that because then the video model learns some you know to generate
shots,这样做有个好处,因为这样 video model 就能学到一些,你知道,去生成
74:46
a single shot wall you need obviously to figure out a lot of stuff about the world to generate multi-shot
一个 single shot 的话,很明显你需要搞清楚世界上很多很多东西,才能生成 multi-shot
74:52
video wall you also need to basically get some like video editing instincts like you need to figure
video 的话,你还得基本上有一些,像是 video editing 的直觉,就像你得
74:57
out what is the right pacing of shots and also LMS are not that good at it like they're not that
搞清楚 shots 的正确 pacing 是什么,而且 LMS 在这方面也没那么好,就是它们没那么
75:03
great video editors if you ask LMS to take some videos and then kind of auto create a edited video
擅长做 video editors,如果你让 LMS 拿一些视频,然后有点像自动剪出一个 edited video
75:11
out of that it would it would feel uncanny so I don't think LMS are actually that good yet out
出来,那会,那会感觉很诡异,所以我觉得 LMS 其实还没那么好
75:18
being video editors and I think there's benefit to learning that end to end so I would expect
作为视频剪辑师,我觉得学习这种 end to end 是有好处的,所以我会预期
75:23
you know the training general is the things that you know you need the harness for eventually get
你知道,training 一般来说,就是你最终需要 harness 的那些东西,会
75:29
kind of injected into the into the model itself and you learn that end to end do you find that
逐渐被注入到 model 本身,然后你 end to end 地学习这些。你有没有发现
75:33
you need to hire engineers who can or researchers who are also artists to infuse that taste or do
你需要招那种能注入这种品味的工程师,或者本身也是艺术家的研究员,还是
75:40
you have kind of artists and residents to distill them we have we have a a a large creative team
你有那种艺术家和驻场人员来 distill 它们?我们有,我们有一个很大的创意团队
75:48
that's very actively involved in the in training those models like on the you know in every part of
非常积极地参与到那些 model 的 training 里,你知道,在每一个环节
75:53
the way and like how do you caption videos well so you capture the stuff that you need for like
而且比如说你们怎么做好 video caption,好捕捉到你需要的东西,比如
75:59
this cinematography the aesthetics the camera direction in as detailed ways as possible so that
这种电影摄影、美学、镜头调度,尽可能详细地,这样
76:05
you're able at inference time to actually at least see that through the model you know we have
在 inference time,你其实至少能通过 model 看到这一点,你知道,我们有
76:10
our creative team also does a lot of evaluation of like you know what constitutes a usable video out
我们的创意团队也会做很多评估,比如,你知道,什么算是一个可用的视频,从
76:16
of those models and so they're very involved kind of through every part of the process and I think
那些 model 里出来的,所以他们非常深入地参与整个流程的每个环节,而且我觉得
76:22
that's one of the special things of runways just that that makes between like kind of creatives
这是 Runway 特别的地方之一,就是它让,像是,创意人员
76:26
and researchers kind of sitting by by side by side and kind of working together to to build the
和研究人员,像是肩并肩坐在一起,一起合作来构建
76:31
next generation of our models I think that's been really really important piece to to you know how
我们下一代 models,我觉得这一直是非常非常重要的一环,对于,你知道,我们如何
76:36
we operate as a company yeah in some sense this all you can only do this in New York I mean you
作为一家公司运作,是的,从某种意义上说,这一切,你只能在 New York 做,我是说,你
76:41
have other offices but like you know I'm trying to find some poetic a significance in the fact that
在别处也有办公室,但就像,你知道,我试图在这个事实中找到某种诗意的意义,那就是
76:47
you are a big New York company as there's a few parts of being New York obviously there is
你们是一家大型 New York 公司,因为作为 New York 公司有几个方面,显然有
76:51
that intersection of all those different industries and like media kind of advertising like yeah this
那种所有不同行业的交汇,还有像媒体、广告那种,像,是的,这
76:58
is very advertising the art scene is New York not not too not to say anything about about the
非常广告,艺术圈是 New York,不是,不是太,不是说关于关于那个
77:04
summer fiscal but it's there is more more going on there is that component there's also I think
summer fiscal,但它是,那里有更多更多事情在发生,有那个成分,还有,我觉得
77:11
we benefit from being outsiders and thinking of things a bit differently like not being in the same
我们受益于作为局外人,用稍微不同的方式思考事情,像不在同一个
77:17
like high of mind of me recess ASI of of Bay Area and like taking and also taking your time to
像不在 Bay Area 其余地方那种 hive mind 里,以及像慢慢来,也花时间
77:25
you know to get where we are today like building the growing the team intentionally and bringing
你知道,达到我们今天的位置,像有意识地建设、壮大团队,并引进
77:30
people who are yeah both on the creative side and also an engineering research side there's
那些人,是的,既有创意方面,也有 engineering research 方面,有
77:35
obviously huge talent pool of amazing people in New York so that that hasn't really been been
显然 New York 有非常多很棒的人才,所以这其实一直都不算
77:40
a problem I mean congrats on everything what are you hiring for you know what should people look
问题。我是说,恭喜你这一切。你们在招什么岗位?你知道,大家应该期待
77:45
forward to for the future of runway we're hiring across the board I think this is probably the
Runway 未来的什么?我们各个方向都在招人,我觉得这可能是
77:51
most open roles we have had in the history of runway we're growing our research team quite
Runway 历史上开放职位最多的一次。我们的 research team 正在大幅
77:56
significantly so if you're if you're excited about video models if you're excited about world models
扩张,所以如果你对 video models 感兴趣,如果你对 world models 感兴趣,
78:01
if you're excited especially about robotics the robotics team we're hiring roles on the robotics
如果你尤其对 robotics 感兴趣,robotics team 我们正在招 robotics 相关的岗位,
78:06
across kind of software hardware and research so definitely definitely reach out and you know a lot
涵盖 software、hardware 和 research 这些方向,所以一定一定要联系我们。而且你知道,很多人
78:14
of people don't have direct robotics background but what should they have you know if they want to
没有直接的 robotics 背景,但他们应该具备什么呢?你知道,如果他们想
78:19
be useful in robotics so ideally some experience with learned policies with the just our
在 robotics 里能有用,所以理想情况下有一些 learned policies 的经验,就是我们的
78:26
else good for for robotics but we we tend to hire generalists as a philosophy and like people who
不然的话对 robotics 也有好处,但我们,我们倾向于把招 generalists 当成一种理念,喜欢那种
78:35
learn really quickly but some experience and the in kind of domain expertise in robotics is
学得特别快的人,但有一些经验,以及在 robotics 里那种 domain expertise,是
78:41
something that we're we're definitely looking for for the next months and then we're scaling the
我们接下来几个月肯定要找的东西,然后我们正在 scaling 这个
78:46
go-to-market team significantly there is a wide like very very active enterprise adoption happening
go-to-market 团队,规模大幅扩张;眼下有广泛的、非常非常活跃的 enterprise adoption 正在发生
78:53
around video models at the moment and we're really trying to respond to all the demand yeah great
围绕 video models,眼下我们真的在努力回应所有这些需求。对,很好
79:01
you want to talk about the open source robotics stuff sure we just random notes we had in video
你想聊聊 open source robotics 的东西吗?当然,我们只是在 video launch 里有一些随手记的笔记
79:08
launch cause well I guess it's interesting so you're a founding member AI labs to build open source
因为,嗯,我觉得这挺有意思的,所以你是 AI labs 的创始成员,来构建 open source
79:14
world models in physical AI anything much to talk on here is open research the biggest thing is
在 physical AI 里的 world models,这里其实没太多能细聊的,都是 open research,最重要的是
79:21
that as I mentioned while models are still early like there's still so much that we you can scale
就像我提到的,虽然模型还很早期,但还是有太多东西可以 scale
79:27
and those models further so much more advancements and things that we can figure out are how to
把这些模型进一步推进也是如此,还有多得多的进展,而我们要搞清楚的是怎么
79:33
improve those models further and I think this is it's important that some of this research happens
进一步改进这些模型;而且我觉得,很重要的一点是,有些研究得在
79:39
in the open and figuring out what is some incentives for different companies to come together to
公开环境里进行,还要搞清楚有什么激励能让不同公司走到一起
79:45
to actually bring some of that research into into the open and open source and so Cosmos
真正把其中一些研究带进公开环境和 open source,所以 Cosmos
79:49
Coalition was initially that we go find a within video to bring some of that research as open
Coalition 最初就是,我们去找一个 within video,把其中一些研究作为 open
79:56
source that could mean open weight model releases it could mean benchmarks that measure physics
source 推进;这可能意味着发布 open weight model,也可能意味着做衡量 physics 的 benchmarks
80:03
and things that people care about when building world models it could mean infrastructure so really
还有人们在构建 world models 时关心的事情,这可能指的是 infrastructure,所以其实
80:07
how do we grow the ecosystem of world models and make that something that also it's easier for
我们该怎么壮大 world models 的 ecosystem,并让它也能更容易让
80:13
a developer a researcher that's just starting out that is excited about world models to kind of
一个刚起步、对 world models 很兴奋的开发者、研究者,去
80:17
contribute to the field I think there's some amount of like is this also our response against
为
80:23
the Chinese world models that are being released you know or is there not part of the consideration
那些正在发布的中国 world models,你知道,还是说这不在考虑范围内?
80:29
I do think it's it's important for in video models you know if you look at the leaderboards
我确实觉得这很重要,对于 video models 来说,你知道,如果你看 leaderboards
80:36
of video models I would say right now the majority of models at the you know the top 10 top 20
在 video models 的 leaderboards 上,我会说现在大多数模型,在,你知道,top 10、top 20 里
80:43
are Chinese models there is only a handful companies that are made to the leaderboard from
都是中国模型,只有少数几家公司能登上 leaderboard,来自
80:48
like the U.S. or the West we're doing better with images but with video we're very behind right
像 U.S. 或者西方那样,我们在 images 上做得更好,但 video 上我们很落后,对吧
80:53
and so I think I think it's definitely important that we invest more probably as a community to make
所以我觉得,我觉得肯定很重要的一点是,我们得投入更多,可能作为一个社区,来确保
80:58
sure that we we can those models can you know we have competitive models out there well like what
我们,我们那些 models 能,你知道,我们有有竞争力的 models 在外面,那,像什么
81:03
was the stoppers from this distilling from them I don't know if that's the best long-term
从他们那里 distilling 的阻碍是什么?我不知道那是不是最好的长期
81:08
opportunity that you wanted by the performance that you can it's it's almost a bit of a pessimistic
机会,你想要的,靠你能做到的 performance,这,这几乎有点悲观
81:14
you know like you that you can get better you know you can this free data you know you might as well
你知道,就像你,你能变得更好,你知道,你可以用这个免费数据,你知道,你倒不如
81:19
like if they're doing it for life for the text language side and it might so do it for the
就像如果他们为了 life,为了 text language 这边做这件事,那也许也照样为
81:23
video side the other way yeah I mean I do think we're we're quite careful of training great models
video 这边,反过来,对,我的意思是,我确实觉得我们,我们对 training 很棒的 models 相当谨慎
81:29
without without the solution at the moment yeah yeah so anything you have to say on benchmarks and
without 目前解决方案 without 目前解决方案 对对 所以关于 benchmarks 和
81:34
evals like I feel like what I'm hearing is a lot of people really like arenas for video and image
evals 你有什么想说的 我感觉我听到的是很多人真的很喜欢 video 和 image
81:41
models customers and whatnot as well they only want the best on the leaderboard and they refer to
models 的 arenas 客户什么的 他们只想要 leaderboard 上最好的 而且他们提到
81:47
arenas a lot more than language models seem to do but any any notes on benchmarks what's lacking
arenas 的频率比 language models 似乎要高得多 但关于 benchmarks 有什么
81:53
how does that average person compare well these both look really hyper realistic more than that
notes 缺了什么 普通人怎么比较 这两个看起来都超级写实 除此之外 我们确实聊过比如
81:58
outside of we did talk about like robotic simulation the physics and all that with anything to say
robotic simulation 的 physics 之类的 有什么想说的吗
82:05
I actually think it's the opposite in some ways I think people generally creatives and artists
我其实觉得在某些方面是反过来的 我觉得人们 一般 creative 和 artists
82:09
and marketers like people that are using our platforms I think relied less on arena scores and
和 marketers 就是使用我们平台的人 我觉得他们更少依赖 arena scores 和
82:16
it's it's just so easy to you know generate with a bunch of different models and then compare
就是就是,真的很简单,你知道,用一堆不同的 models 生成,然后对比
82:22
the results visually like one nice thing about image and video models is you can immediately tell
结果,视觉上,就像 image 和 video models 有个好处,就是你马上就能看出来
82:27
video rights like what what feels good from a static standpoint like any artifacts and issues
video 嘛,对吧,比如从静态角度看,什么感觉好,像任何 artifacts 和问题
82:32
with the physics of those models you can immediately tell and so that's it's actually easier I would say
还有那些 models 的物理效果,你立刻就能看出来,所以我觉得,其实更容易
82:38
to evaluate as a human there is also those models and it is in language models where you have those
作为人类来评估那些 models 也是这样,而就是在 language models 里,你会有那些
82:45
those very complex kind of math and coding and tests where it becomes a lot more hard I think
那些非常复杂的 math、coding 和 tests,这时候就难多了,我觉得
82:52
for for humans to evaluate and can discriminate between the performance of for tier models at
对,对,人类来说,要评估并区分 for tier models 的表现,在
82:58
the time so I think in practice people just test out the same prompt with a bunch of different models
那个时候,所以我觉得实际上,人们就是用同一个 prompt 去试一堆不同的 models
83:04
and see what the results look like and right now in one way you can use our models and you can
然后看看结果是什么样。现在有一种方式是,你可以用我们的模型,也可以
83:08
use their part models as well so it's it's very easy to do that.
用他们的 part models,所以这事儿非常容易做。
83:12
Amazing. We're going to end with the AI runway AI summit the last sort of social societal issue
太棒了。我们要以 AI Runway AI Summit 收尾,最后一个算是社会层面的议题
83:17
I guess I don't know if this is a thing is the you are at the tension between sort of artists
我猜,我不知道这算不算一个现象,就是你处在一种张力之中,介于艺术家
83:24
creatives and AI a lot of people in that that community hate AI obviously the people that are
创意工作者和 AI 之间。那个圈子里很多人讨厌 AI,显然那些
83:28
in the runway community don't mind using tools it's just another brush but how if you see the sentiment
在 Runway 社区里的人不介意使用工具,这不过是另一支画笔。但你怎么看,如果你看这种情绪
83:37
I mean our perspective yes it's just another branch brush it's just another camera it's you know
我是说,我们的视角是,对,它只是另一支画笔,只是另一台相机,你知道
83:43
it's the latest of a long generation of technology and art and technology have kind of evolved
它是漫长一代技术里的最新一个,而艺术和技术在某种程度上一直在演进
83:49
together I think there's been a pretty significant shift over the past few months and it came
总的来说,我觉得过去几个月发生了一个相当显著的转变,而这个转变来自——
83:55
some of it you can see with a lot of public figures speaking out in favor of AI and being
其中一部分你能看到,很多公众人物站出来支持 AI,而且
84:02
you know like in Cannes you saw a few a few directors speaking in favor of AI we had run Howard
你知道,就像在 Cannes,你看到有几个、几个导演公开支持 AI,我们请来了 run Howard
84:08
in our film festival there is Marx such as a also adopting AI models so you have more of those
在我们的电影节上,还有 Marx such as a 也在采用 AI models,所以你会看到越来越多这样的
84:16
stories coming out every day of like a well-known figure kind of speaking in favor of AI
故事每天都在冒出来,比如某个知名人物多多少少在公开支持 AI
84:22
and it's just a matter of in my mind it's those models are becoming more and more demystified
而在我看来,这只是那些 models 正变得越来越去神秘化的问题
84:27
I would say I have also a bit of a hotake that one of the things that made the initial response
我想说,我也有点暴论,就是让最初反应
84:35
to those models maybe a bit more heated than it needed to be was this idea of text to video
对那些 models 来说可能比实际需要的更激烈的原因之一,是 text to video 这个概念
84:42
of you know you have a single text description and you get back a full video
就是,你知道,你有一个单独的 text description,然后就能得到一整段完整的视频
84:48
yeah there was the misconception obviously you can generate it to our feature-line phone
对,之前显然有个误解,就是你可以把它生成到我们的 feature-line phone 上
84:53
but the most of today now take a lot of references they take they are very controllable
但今天的大多数现在会接收很多 references,它们会接收,它们非常可控
84:58
and I think when people see a tool that allows efforts many degrees of freedom and control
我觉得当人们看到一个工具能带来很多 degrees of freedom 和控制时
85:05
they respond to it differently and you might as less that it's a generative model than
他们会对它有不同的反应,你可能不太会觉得它是一个 generative model,而更
85:11
than the fact that you can actually steer it to the direction that you want and so I think when
而是你能真正把它引导到你想要的方向这个事实。所以我觉得当
85:18
people look at you know complex workflow on top of those models when they look at you know
人们看到,你知道,这些模型之上的复杂 workflow 时,当他们看到,你知道
85:23
all the ways in which you can steer them and you can provide now with some of the latest models up
你能引导它们的所有方式,以及你现在用一些最新的模型可以提供的……
85:28
to 50 references like the conversation becomes a bit different because it feels much more like a
到 50 个 references 的时候,对话就变得有点不一样了,因为它感觉更像一个
85:34
storyboard a tool versus like something that a much call kind of entity that figures out like
storyboard 工具,而不是那种更像 much call 类型的 entity,会帮你搞清楚
85:41
your entire phone for you and he notes on like workflows changing for people in the field like
你的整个手机,为你搞定;他还提到,这个领域里人们的 workflows 正在变化,比如
85:47
think engineering at least has had a lot of people where they're like expectations have changed
想想 engineering,至少有很多人,他们的预期已经变了
85:53
I'm you know 10x 100x more productive and you can get a lot more done same thing is you know you're
你知道,我效率高了 10x、100x,你能完成多得多的事;同样,你知道,你在
85:58
making dev tools for creatives any notes there like there's some people that don't want to adopt
为创意人员做 dev tools,这方面有什么观察吗?比如有些人不想采用
86:05
some that do like anything yeah so I think in terms of like what people care about I see that
有些人则愿意,什么都行。对,所以我觉得,就人们关心什么而言,我看到
86:12
we've gone through a few stages so we started from the stage where the main thing that people
我们已经经历了几个阶段,一开始的那个阶段,人们最主要的事情是
86:16
were looking for was quality like as you know we scale those models the quality improved dramatically
我们当时追求的是质量,你知道,随着我们把那些模型 scale 上去,质量提升得非常显著
86:22
that's something that people still care about but it's it's now in addition to controllability
这仍然是大家关心的事情,但现在它变成了在可控性之外的另一个维度
86:27
like being able to steer those models with references with different kinds of inputs and storyboards
比如说能够用 references、不同类型的输入和 storyboards 来引导那些模型
86:33
and now my sense is increasingly people are going to care about latency more and more
而现在我的感觉是,人们会越来越在意 latency
86:38
as those models become better the ability to iterate very quickly becomes more important
随着那些模型变得更好,快速迭代的能力变得越来越重要
86:43
and like if you can you know with a single prompt generate 10 different outputs like almost instantly
如果你能用一个 prompt 几乎瞬间生成 10 个不同的输出
86:49
you can explore way faster than before and you get some of the magic that characterise the
你就能比以前快得多地去探索,你也能获得一些曾经定义
86:54
you know the creative tools for the past like Photoshop was instant and with we lost some of that
你知道,过去那些创意工具的那种魔力,比如 Photoshop 是即时的,而我们失去了一些那种感觉
87:00
with generative models you you're waiting for two minutes to get back a video and I think we're
用 generative models 的时候,你得等两分钟才能拿回一个视频,而且我觉得我们现在要
87:04
going to bring some of that back now with the real time models yeah exciting and exciting the last
用 real time models 把其中一些东西带回来,对,很兴奋,很兴奋,最后
87:12
thing we'll plug is this one right away you're finally doing this in SF yeah so we're very excited
我们要马上宣传的就是这个,你终于在 SF 做这件事了,对,所以我们非常兴奋
87:19
about this so this is in late September September 30th we're doing a summit on primarily focused
关于这个,时间是在九月下旬,9 月 30 号,我们要办一个峰会,主要聚焦
87:26
on physical AI and real time video generation we have panelists from Nvidia physical intelligence
在 physical AI 和 real time video generation 上,我们有来自 Nvidia、physical intelligence
87:31
the bot call deep mind yeah it's going to be I think a very interesting series of conversations
the bot call deep mind 的 panelists,对,我觉得这会是一系列非常有意思的对话
87:37
we're trying to make the panels really technical and elicit actual kind of substantive discussion
我们在努力让这些 panel 变得非常技术化,并引出真正有点实质性的讨论
87:45
and hopefully some some interesting kind of disagreements and interesting debates on things
而且希望能有一些有意思的分歧,以及在各种事情上有意思的辩论
87:50
and yeah there's tickets available hope people can join since you mentioned it what kind of
对了,还有票,希望大家能来参加。既然你提到这个,那人们应该思考、或者你预期会有哪些
87:58
disagreements and debates should people think about or do you expect so it seems like you know
分歧和争论呢?所以看起来,你知道,
88:05
there is a one debate right now in the robotics world is VLA's versus world action models
现在 robotics 世界里有一个争论,就是 VLA's 对 world action models。
88:12
so there is labs that are really really betting on one of those two directions there is
所以有些实验室真的真的在押注这两个方向之一。还有
88:18
what is the best source of data to train robotics models there's just the third party first party
训练 robotics models 最好的 data 来源是什么?就是 third party、first party,
88:23
that we talked about yeah there is you know the people who really believe in kind of further
我们聊过的。是啊,你知道,有些人真的相信要进一步
88:28
scaling teleop data versus leveraging more large scale video data so those are kind of some of the
scaling teleop data,而不是利用更大规模的 video data。所以这些算是其中一些,
88:34
and then there is you know the the world models debates of predict pixels directly versus
然后还有,你知道,world models 的争论,就是直接 predict pixels,还是
88:40
something like JEPA versus a more 3D based 3D based approach so I think we're at a nice time
有点像 JEPA 对比更 3D based 的 3D based 方法,所以我觉得我们现在正处在一个很好的时间点
88:47
in world models because there is still that kind of active debate happening or like what is the best
在 world models 里,因为那种活跃的争论还在发生,比如到底什么才是最好的
88:51
long term direction I feel very strongly that this video predict pixels directly and a scaling
长期方向。我非常强烈地觉得,这种 video 直接预测 pixels,以及 scaling
88:58
video generation models is the right approach but it's I think there is a lot of interesting
video generation models 才是正确的方向,但我觉得这里有很多有意思的
89:03
debate happening by by researchers on like what is the best best best kind of path to take
争论正在研究者之间发生,比如到底什么才是最好最好最好的路径
89:09
it's interesting that it's all on like sort of let's call it the policy layer and the data model layer
有意思的是,这些全都集中在,怎么说呢,我们姑且叫它 policy layer 和 data model layer 上
89:15
is the the physical side is completely solved like all the sensors all the actuators all these things
就是 physical side 已经完全解决了,像所有 sensors、所有 actuators,所有这些
89:20
there we we have everything that we need I don't think that's a self-light there
在这方面,我们已经有需要的一切了。我不觉得那是个 self-light
89:26
different problems you know it's like I want a dream of all these things and then I get you know
各种不同的问题,你知道,就像我梦想着所有这些东西,然后我得到了,你知道
89:32
you know I buy a robot or I try to assemble my own and I can't even get the you know the the
你知道,我买个 robot,或者我试着自己组装一个,结果我连,你知道,那个,那个
89:38
motors to like work right right again it's you're dealing with very sensitive equipment that has
motors 让它正常运转都搞不定,对吧,对吧,又来了,你面对的是非常敏感的设备,它有
89:44
voltage and power and like heat and all these things which you know abstracted away we're sitting
voltage、power,还有 heat 这类东西,这些你知道都被抽象掉了,我们坐
89:51
here we're talking about software and talking about models but like really you have to deal with
在这里聊 software,聊 models,但说真的,你还是得处理
89:55
those kinds of things too yeah and I think I'm I'm generally also not not opposed to incorporating
那类东西,是的,而且我觉得我总体上也不反对把
90:01
other modalities into our models like we've seen yes the simplest case is they can generate video
其他 modalities 纳入我们的 models,就像我们看到的,是的,最简单的情况是它们能生成 video
90:06
and audio at the same time so they can generate RGB and they can also generate they can generate sound
和 audio,同时生成,所以它们能生成 RGB,也能生成,它们能生成声音
90:12
and audio but my you know I everything about this as like what what does the maximalist version of
还有 audio,但我的,你知道,我关于这一切的理解就像,那个最大化的版本到底是
90:20
world model look like is you're incorporating more and more modalities from the universe and
world model 长什么样呢,就是你要把宇宙中越来越多的 modalities 都纳进来,
90:25
you're training a model on different scales of observation so yeah got a good good essay that
并且在不同尺度的 observation 上训练一个 model,所以对,有篇很好很好的文章,
90:32
people should be on yeah no meta released a model that was like six modalities in one yeah I
大家应该去看看。对,不是,Meta 发布了一个 model,像是六种 modalities 合一的,对,我
90:37
forget what the name of the the thing was but it was like yeah okay that is one of them but that
忘了那东西叫什么名字了,但它就像,对,好吧,那是其中一个,但那
90:41
is like a transformation of RGB in image bind image bind yeah yeah what are the modalities they
就像是 RGB 的一种 transformation,在 ImageBind 里,ImageBind,对,对,那有哪些 modalities?它们
90:47
have heat audio depth heat text whatever I am you is I do think like you might as well do ultraviolet
有 heat、audio、depth、heat、text,whatever,我是你,是,我确实觉得你不如也做 ultraviolet,
90:55
you might as well do like this whatever other modality you feel like because it's all data to the
你不如也把随便什么你觉得行的其他 modality 都做进去,因为对那个来说,全都是 data。
91:00
model yeah and a big bad is also that there is transfer between all those models yeah one of my
model 对,而且还有一个很大的坏处是,所有那些 model 之间都有 transfer,对。我最喜欢的一个例子,这个现在其实已经挺旧了,就是有一个 stable diffusion 的 fine tune 叫 refusion,对,是 music,对 对,就是把 stable diffusion 在 spectograms 上 fine tune,然后它变成了一个相当强的 music generator,对吧。所以可能有一些 spatial partner,还有像 spatial temporal 的部分,然后你看我们说的 video,它算是从不同的 scales 和不同的 modalities 里涌现出来的,所以某种程度上,你知道,model 做了一些 meta learning,让你能学得更快,如果你从一个只是用 images 训练出来的 model 开始,然后训练它去 predict audio,而如果你从零 train 的话,
91:05
favorite examples which is quite all that this at this point is there was this fine tune of stable
我最喜欢的例子之一,其实到这会儿也差不多就是这些了,就是有个 Stable Diffusion 的 fine-tune,叫 refusion,对,是做音乐的,对对,就是拿 spectrograms 去 fine-tune Stable Diffusion,然后它就成了一个相当能打的 music generator,对吧。所以可能有一些 spatial partner,然后像 spatial temporal 部分,你看我们说的那个 video,它有点是在不同 scales 和不同 modalities 下涌现出来的,所以有某种程度的,你知道,meta learning,是 model 已经做过的,这能让你学得更快,如果你从一个只是用 images 训练的 model 开始,然后训练它去 predict audio,如果你从 scratch 训练的话,还有……
91:11
diffusion that was called refusion yes was music yeah yeah just fine tuning stable diffusion on
diffusion,那个被叫做 refusion 的,对,是音乐,对对,就是 fine tuning stable diffusion 在
91:17
spectograms and it became a quite capable music generator right so there is probably some
spectograms 上,然后它变成了一个相当强的 music generator,对吧,所以大概有某种
91:24
spatial partner and so like spatial temporal part and see we're talking about the video that
spatial partner,然后像 spatial temporal part,而且你看,我们说的是 video,它
91:28
kind of emerged at different scales and different modalities and so there is some degree of
在 different scales 和 different modalities 中浮现出来,所以存在某种程度的
91:32
you know meta learning that the model has done that that allows you to learn faster if you if you start
你知道,meta learning,是 model 做过的,它让你能学得更快,如果你,如果你从
91:38
from a just a model train images and train it to predict audio and if you train from scratch and
从一个只是在 images 上 train 的 model 开始,然后 train 它去 predict audio,而如果你从 scratch train,而且
91:43
just just audio and there is some other interesting examples so there is this this project called
只是,只是音频,然后还有一些其他有趣的例子,所以有一个这个,这个项目叫
91:49
the well it's it's a data set of physics numerical simulations and physics and biology and a bunch
the well,它是一个 data set,包含 physics numerical simulations 和 physics 和 biology,以及一堆
91:58
of other domains so it's so it's essentially different physical systems across very different
其他领域,所以它,所以它本质上是不同的 physical systems,跨越非常不同的
92:04
scales of space and time from like astrophysics to low level kind of like atomistic interactions
空间和时间的尺度,从像 astrophysics 到低层次的、有点像 atomistic interactions
92:12
and we've seen we've done some some some some work on this and and we've seen that we can take
然后我们看到,我们做了一些,一些,一些,一些这方面的工作,而且,而且我们看到我们可以拿
92:18
our video model where you know real world video looks nothing like this and you can actually find
我们的 video model,你知道,真实世界的视频看起来完全不是这样,而你实际上可以找到
92:23
unit on on those numerical simulations and just treat them as RGB frames and you actually get
unit,在,在这些 numerical simulations 上,然后直接把它们当成 RGB frames,你实际上会得到
92:31
reasonable performance much quicker than if you just train from scratch. I think we've seen this
还不错的 performance,比起只是 train from scratch 要快得多。我想我们已经看到过这个
92:37
across languages yeah deep seco CR as well yeah deep seco CR like you don't have to organize text like
跨语言,对,deep seco CR 也是,对,deep seco CR,就像你不需要整理文本,就像
92:42
you can just do it in this images. There's a lot that happens in that base pre-training like
你可以直接在这些图像里做。在那个 base pre-training 里会发生很多事情,就像
92:47
there was an argument a long time ago people saying oh Jimin's have so many sensory representations
很久以前有个争论,人们说,哦,Jimin's 有那么多 sensory representations
92:52
right smell touch models have a hold two more modalities that will never like that we don't even
对吧,嗅觉、触觉,模型还多出两种 modalities,那永远不会,就像我们甚至都不
92:58
have data for it's like okay you take a qi sensor like you can you can try this stuff but actually
有数据。就像,好吧,你拿一个 qi sensor,就像你可以,你可以试试这些东西,但实际上
93:03
you know there's so much happening in just the base train run that you don't get as much from
你知道,光是在 base train run 里就发生了太多事情,以至于你从
93:08
these little things yeah exactly and I think that's what it solves is data scarcity so you don't
这些小东西里得不到那么多。对,没错,我觉得它解决的就是 data scarcity,所以你不
93:13
have as much you have so much video data available but you don't have like all factory data. The cool
会有那么多。你有那么多 video data 可用,但你没有像所有 factory data 那样的数据。酷的是
93:21
thing is it goes the other way too right so if you want to do physics like if you want to measure
问题是,它反过来也一样,对吧,所以如果你想做 physics,就像如果你想测量
93:26
this or you want to have a diffusion model do audio it transfers really well so like in your case
这个,或者你想让一个 diffusion model 来做 audio,它迁移得特别好,所以就像你那个情况
93:31
the little bit of post-training for robotics gets a video model to use its fundamentals in another
一点点针对 robotics 的 post-training 就能让一个 video model 把它的 fundamentals 用到另一个
93:37
domain so we can apply that to other stuff too yeah and if we look at you know like how do you
domain 里,所以我们也能把这个应用到其他东西上,对,然后如果我们看,你知道,比如你怎么
93:42
make those models more useful for scientific domains and if you know if you look at
让那些模型对 scientific domains 更有用,然后如果你知道,如果你看
93:47
awful they've had always very because of the data you know the limited amount of data that it needs
awful,它们一直都非常,因为 data,你知道,它需要的 data 量很有限
93:54
to be trained on it basically it's very fine tune architecture just to solve kind of
用来训练,基本上它是一个非常 fine tune 的架构,只是为了解决某种
94:01
protein structure prediction but if you take all those disparate sources of scientific data and
protein structure prediction,但如果你把那些各种不同的 scientific data 来源都拿来,然后
94:07
you bring them together under a single model like I think that's an approach that can help us
你把它们整合到一个单一的 model 下,我觉得这是一种能帮我们
94:12
kind of solve solve new kinds of problems across across science by leveraging all the all the
某种程度上解决、解决横跨科学领域的各种新问题,通过利用所有、所有
94:18
learnings from one model the or one set of the data to another so very early days for for that
从一个 model 或一组 data 到另一个的学习,所以对于、对于那个
94:25
direction but I do think that that's where ultimately what the endgame of kind of simulating the
方向来说还非常早期,但我确实觉得,最终模拟
94:30
world is you're not just using RGB you're using RGB as a starting point but you can incorporate more
世界的终局就是:你不只是用 RGB,你是把 RGB 当作起点,但你可以纳入更多
94:36
more modelities of the universe and leverage the transfer that happens from learning from one
更多宇宙的 modalities,并利用从一个
94:41
to the other I guess the follow up there is what's the drawback of Omni like why is everything
到另一个学习时发生的 transfer。我猜这里的后续问题是,Omni 的缺点是什么,就像为什么所有东西
94:47
not an Omni model so why not now and why would you start from language backbone or image video
都不是 Omni model,那为什么不是现在,以及你为什么会从 language backbone 或 image video 开始
94:54
backbone and then go Omni from there doesn't matter yeah we need to take it once so we need to
先做 backbone,然后再从那儿走到 Omni,这都无所谓,对,我们得先走这一步,所以我们需要
94:58
solve robotics first and then we can go into solve everything now yeah I mean I do think there is
先解决 robotics,然后我们才能去解决一切,现在,对,我是说,我确实觉得存在
95:06
a lot of open-ended research needs to happen for those Omni models there is you know there is
对于那些 Omni models,还需要进行大量 open-ended research,有,你知道,有
95:11
a lot of things that the required careful consideration when you're bringing multiple modalities into
很多事情需要仔细考虑,当你要把 multiple modalities 引入
95:16
a single model to predict but I think you know expect those to be and to be solvable wonderful you've
单一 model 来做 predict,但我认为,你知道,期待这些会是,并且是可以解决的,太好了,你
95:24
been very generous to your time congrats on your success and you know I'm excited for the AI
在时间上一直非常慷慨,恭喜你成功,而且你知道,我很期待 AI
95:29
summit or physical AI summit yeah thanks for having yeah and people should check out the film festival
summit,或者 physical AI summit,对,谢谢邀请,对,而且大家应该去看看那个电影节
95:35
if it's in town right yeah yeah you'll be going to be touring all over the place yeah next next
如果它来城里的话,对吧,对,对,你会到处巡演吧,对,下一个,下一个
95:39
year we're probably gonna do that so we do film festivals every May or June and we did the last one
今年我们大概会做那个,所以我们每年五月或六月都会办电影节,上一次我们办了
95:47
in New York LA Tokyo and that the AI engineer yeah fair yeah yeah yeah so yeah hopefully even more
在 New York、LA、Tokyo,还有那个 AI engineer,对,有道理,是啊是啊是啊,所以对,希望明年能有更多
95:56
place next year no I think like a Sunday you know you will be hosting the Oscars of AI video and you
地方。不,我觉得大概是个周日,你知道,你会主持 AI video 的 Oscars,而且你
96:01
know I think people should take this very seriously as like a potential career they can have
知道,我觉得大家应该非常认真地对待这件事,把它当成一个他们可以拥有的潜在职业
96:07
the Oscars of AI video will be called the Oscars all right thank you thank you
AI video 的 Oscars 就叫 Oscars,好吧,谢谢,谢谢