WEBVTT

00:00.000 --> 00:10.640
I'm super interested in that one because, like, I'm from the Python world and, like,

00:10.640 --> 00:15.520
wasn and running things in Python in the browser is super interesting and now we are

00:15.520 --> 00:19.760
hearing how to also run LLMs in the browser. That's going to be really super interesting

00:19.760 --> 00:23.560
talk. It's the stage years.

00:24.520 --> 00:29.520
It's the hours now. So, next to us, healthy. Remember,

00:29.520 --> 00:35.400
we're five. The first AI talk in four stems. Say, say, hello.

00:37.080 --> 00:42.280
Where's Peter? All right. So, whatever. All right. Thank you very much. And today,

00:42.280 --> 00:49.080
we're going to go first through some crazy things on this five. And if you guys

00:49.080 --> 00:54.120
free, go to our WISFI booth to take a look over the crazy part of, we're making an

00:54.120 --> 01:02.840
including a romance one version. Okay. So, today, I will go through a few brief on the WISFI stage

01:02.840 --> 01:07.880
on the hardware side. And then, Peter, we're going to the more deep part on the software side.

01:07.880 --> 01:16.840
So, then, I will dash off to another talk on the VFC side about AI. So, and, and myself is

01:16.920 --> 01:21.880
using, I'm the software guy. Make my money on the software guy. And then, I found it

01:22.440 --> 01:28.040
deep computing in doing the correct. And then, as I said, I'm very serious about software

01:28.040 --> 01:34.680
and the start losing money in the hardware. Probably same as a moment in 10 years. So,

01:34.680 --> 01:43.400
and then, we make a lot of hardware for the past three years. So, from 2020 to three, we made

01:43.480 --> 01:50.280
the first, the most expensive laptop, $5,000 laptops for 4 cores. So, that's much the history.

01:51.000 --> 01:58.440
So, it's stay in the British computer museum in Newton King. All right. So, and then, I do

01:58.440 --> 02:06.360
our 2020-4A core. And, and the second generation after, and the iPad, we've extended more

02:06.360 --> 02:13.560
of them, and they repair them in the core to his wife using the tablet, but as an anonymous

02:13.560 --> 02:19.560
core, and his wife rejected it. So, that's the story. So, we once are Ubuntu and Fedola.

02:19.560 --> 02:24.920
So, after that, I'm tired. I'm losing too much money and making our own laptop on the hardware.

02:24.920 --> 02:31.400
I decided I'd to focus on making the motherboard for framework. And, to come to frameworks

02:31.400 --> 02:37.160
up, move, and the risk-fight pool, you see how the motherboard works, and the risk-fight, and with the

02:39.160 --> 02:45.160
framework. So, actually, it's a repairable, replaceable, equitable. You know, multiple ISA architecture

02:45.160 --> 02:52.440
from ARM is a disaster, and risk-fight, and different risk-fight. And then, yeah, it's very good. And everything

02:52.440 --> 02:57.560
all the designs of Ubuntu. So, you guys can make a module for the risk-fighting wall and do whatever

02:57.720 --> 03:03.720
innovation you wanted to do. So, this is the first generation that we use, because it has

03:03.720 --> 03:09.800
it through the whole framework collaboration, and then use of the existing four core to make

03:09.800 --> 03:16.120
the motherboard, and then you launch it in a railroad price. And then what happened now, we see in

03:16.120 --> 03:22.360
that the challenge of the risk-fight associated is we don't know where to hit, what market is,

03:22.440 --> 03:28.040
and what compute power, and what resource-constraining, and we have time-constraining as well,

03:28.040 --> 03:36.520
so for most of the staff. And since the jet-GBT, and then I think everybody is clear,

03:36.520 --> 03:46.680
whoever missed chipset or ENI, nothing else better to do. So, that is the trend. So, and then we launch

03:47.000 --> 03:53.320
in 2025, we launch a chip-back version of a motherboard, and then we're good-spat. You can run

03:53.320 --> 04:01.320
our 14B model on the MPUSI, and we're powerful, and that's the look-lay of the chip,

04:02.520 --> 04:09.400
and that's the motherboard before the DADR, he's the supply chain. Now, if you want to get a

04:09.400 --> 04:21.320
60-40 DADR, probably costs $1,000. So, that is the height of AI. So, and 2020-26, you will see our

04:21.320 --> 04:27.960
16 core of risk-fight coming out with a customer AI instructions. So, and it's a two-caster,

04:27.960 --> 04:36.280
and support 32-GLP DADR is 60-top sparse, and FB-4, and it's pretty powerful, and fun to pay.

04:36.360 --> 04:43.000
And you will see a lot of customer feature, and I agree with you, and two different than.

04:44.440 --> 04:51.400
So, that's our target of selling in March. So, you guys can have a go on it, and see tell me

04:51.400 --> 04:58.760
what's wrong with it, okay? And all I say in that now, we'll, we'll, we'll, we'll find it out on

04:58.760 --> 05:07.400
this side. We don't have applications, low application at all, except GPT, right? So, it's on the web browser.

05:08.200 --> 05:15.320
And, and this I see, video player doesn't use it, and Chrome doesn't use it, and until last year

05:15.320 --> 05:20.840
and last year, I managed to convince reality, can you help me do a version of using the all-the-local

05:20.920 --> 05:27.960
computer of the AI? So, that's why I'm talking in the open media about video player. So,

05:27.960 --> 05:35.400
the model is crazy, all different kind of models, all right? And the AI framework is even more crazy,

05:35.400 --> 05:41.240
all right? And, and a lot of potting, and the framework, all the CPP stuff, and then I'm sure

05:41.240 --> 05:48.360
women tend to know about it, how painful is on the whole world, and all our hardware is potting.

05:48.440 --> 05:56.760
So, we, we do our comparisons on the Chrome side, it's very compressed, right? It's Chrome

05:56.760 --> 06:02.920
meme is on, on this fight, but there's no Chrome, my law. So, and Chrome meme is considered it's not

06:02.920 --> 06:08.200
a safe version of browser, because they don't have a security patch. So, we try to crack Chrome

06:08.200 --> 06:13.960
in and into talking to the local compile, and we'll see how it goes, right? That is,

06:13.960 --> 06:20.040
Chrome meme's money query. So, they have all the AI on the GPU side, and the CPU side,

06:20.040 --> 06:25.080
but doesn't have anything for this fight, at all, because this fight is so new that they have

06:25.080 --> 06:33.400
across the AI instructions, not just a vectorizer, so it's empty. And, Peter, we talk about all

06:33.400 --> 06:38.280
the framework, he knows more than I do. I just find it now, the problem, I'm no solution for it.

06:39.240 --> 06:45.720
And, that's the state of the AI side in the Chrome. So, basically, I will say that Chrome is the

06:45.720 --> 06:53.480
biggest app on any of the latest platform for an user, so we'll see the state of it, right? It's scary.

06:54.280 --> 06:59.480
So, and now, probably, I will pass it to Peter to explain more how he is going to

07:00.120 --> 07:08.680
fix all these for the local MPU, for the vectorizer, for the customer instruction coming

07:08.680 --> 07:14.680
last year, all the chips, and he's going to explain a lot more than I do. Peter, pass on to you.

07:16.520 --> 07:17.080
Thank you.

07:17.080 --> 07:32.040
All right, so a little bit about myself, so I'm a deliciously see the assembly sub-group chair,

07:32.040 --> 07:37.640
I actually cover out some of the proposals that were standardized, like the assembly and

07:37.640 --> 07:44.920
the election day. I'm currently championing the vector proposal as well. And, the former

07:45.000 --> 07:50.360
November time engineers, so I worked a little bit on V8, checkercore, maintainer, as well.

07:50.360 --> 07:59.560
If anybody remembers what that is. So, so the picture, sort of, of the previous slides,

07:59.560 --> 08:05.240
I mean, the reference to, so this is basically two ways to run inference in the browser. It's

08:05.240 --> 08:12.920
webinar and the assembly. So, I mean, for example, things like web-alarm would actually target

08:13.080 --> 08:18.280
one of the existing sort of ways, effectively. So, if you look at web-alarm, so there are usually

08:18.280 --> 08:22.920
ways to support work. There is an underneath, there is an MLLN time, and because we're talking about

08:22.920 --> 08:28.360
Linux, laptops, and maybe Android devices, that would be probably Linux or Android MLLN time.

08:30.200 --> 08:36.200
And the dispatcher part, typically, were sort of differentiated between what can run CPU and

08:36.360 --> 08:45.080
working run on NPO, some kind of accelerator. And the way it kind of, sort of expected to work,

08:45.080 --> 08:50.760
is that sort of on Linux and Android, you would have late RT, which is effectively a

08:50.760 --> 08:58.600
TensorFlow Lite. And the way then we'll go down lower, it's sort of fun. And I'm going to

08:58.600 --> 09:03.080
tip you side if you're going to take some back, which is what you would do to support your

09:03.080 --> 09:10.760
web-b64 or GCP, the inspectors enabled. And if you need to support something else, then you kind

09:10.760 --> 09:19.240
of start writing delegates. And this sort of, sort of, the accelerator side would be basically covered

09:19.240 --> 09:27.160
by existing matrix extension proposals. It's IME, IME, IME and VME, and depending on how

09:27.240 --> 09:31.080
they manage state, it might be actually full, sort of, like, back to the external unpack site,

09:31.080 --> 09:39.000
or it can actually be, it might still be have to impress the delegate. And so, in view,

09:39.000 --> 09:44.040
so the examples that you can give the SVN and space between PUs, it definitely have to

09:44.040 --> 09:48.440
be impressed with delegate. And bigger devices like the historian, the accelerator cards,

09:48.440 --> 09:56.200
but also fall into this category. So we kind of, he left us, I can say that. So we're going to

09:56.200 --> 10:02.760
disagree a little bit in this sense, because he thinks sort of sees us as, the moment that

10:02.760 --> 10:09.000
doesn't have to be lighter at T. So you basically, there's two options. You either have to choose

10:09.000 --> 10:13.400
a different runtime and teach programming to integrate with it. And currently, I'm actually

10:13.400 --> 10:20.520
not always going to integrate with this core ML and on-exantime. Or you basically follow the

10:20.520 --> 10:25.400
past of least resistance and you work with creating and delegate for your NPU and so on.

10:28.280 --> 10:31.800
And I think so, state of executive support and executive support and executive

10:31.800 --> 10:35.240
package is pretty good. At least, I mean, there might be still some operations missing,

10:35.240 --> 10:42.760
but generally it's made a big progress. And so the other side, so, also part of the problem

10:42.760 --> 10:49.800
with the webinar is that only programming supports seven and, but not only programming supports

10:49.880 --> 10:55.640
with the semantics. So, for example, Firefox would also be able to, you will build use with

10:55.640 --> 11:02.920
the semantics Firefox. So the state of the botanical works like this. So the botanical works

11:02.920 --> 11:12.760
has the CMD extension. And right now, sort of the, everything is relatively well until you

11:12.840 --> 11:19.880
can get to CMD. So the scalar code, lowest particular, the semantics works with floating

11:19.880 --> 11:24.760
point representation. That's actually not really well because despite has canonical lands,

11:25.320 --> 11:31.800
exactly the kind of the 71's. There's also half load proposal, which was somebody also

11:32.520 --> 11:38.120
would be able to lower to describe native half load, so this is a pretty good story besides the

11:39.080 --> 11:46.200
CMD. So the CMD is only 128 bit, and it's fixed length. So that presents some

11:47.320 --> 11:52.520
actually fun challenges because they're least likely for running something like

11:52.520 --> 11:56.280
TensorFlow, you would need to have vector operations without that, your performance,

11:56.280 --> 12:02.520
basically is completely unacceptable. And if you, if your registers are a lot bigger than

12:02.600 --> 12:10.680
128, there's also applications, implementations, discriminations, where taking less than full

12:10.680 --> 12:16.120
register of data doesn't actually increase your support. So basically, at that point, you might

12:16.120 --> 12:23.000
be decreasing the support of overall sort of execution. So the problem is actually not necessarily

12:23.000 --> 12:28.280
super unique to this file because there are other sections that there are larger than what

12:28.360 --> 12:36.920
128 bits from D, or they are flexible, like a CV example. So the easy proposal that

12:38.280 --> 12:44.200
vectors to the summary, so it will be basically bytecode level vectors that can be sort of

12:44.200 --> 12:53.720
relatively well mapped to existing RV instructions. And that's sort of the most part,

12:53.720 --> 12:59.480
it's kind of the plan. So unfortunately, this is the state kind of of the sort of stack and

12:59.480 --> 13:04.200
not necessarily a demo of how we fixed it yet. So hopefully next year we can come back and

13:04.200 --> 13:09.240
actually show that we fixed some of it, or most of it, probably not most of it, but at least some of it.

13:10.920 --> 13:17.480
So what sort of the next steps that can be sort of like taken in this, obviously the sort of

13:17.480 --> 13:23.160
ML runtime that will be integrated with the browser, and the support sort of further down,

13:23.160 --> 13:28.520
don't see it from that runtime. Well, and sort of in the vicinity, what's probably left to

13:28.520 --> 13:34.600
try is to actually implement the vector support. So for this, the proposal actually has been kind of

13:34.600 --> 13:40.600
lagging and sort of developed for a while, but the next step is to be able to change support and try

13:41.400 --> 13:50.600
to do runtime prototype. V8 already supports RV, but sort of like extending the support

13:50.600 --> 13:56.520
to use sort of the virtual vectors probably would be something to try. Yeah, anyway.

13:58.200 --> 14:01.800
Thank you. Questions?

14:01.800 --> 14:22.040
All right. So can you talk about sort of relationship between Web NN and Web GPU, because some

14:22.040 --> 14:28.120
communities actually prefer Web GPU to implement kind of the same functionality. So tiny grad for

14:28.120 --> 14:33.640
example has Web GPU backhand, and they kind of happy about it? Yeah, that's actually good question.

14:33.640 --> 14:45.720
So the Web Olam does Web GPU as well as with assembly, for example. So Web NN generally works

14:45.720 --> 14:52.760
a graph API, so you build a graph using Web NN and then basically execute it on a device accelerator.

14:52.840 --> 14:59.400
So Web GPU would definitely work for GPU like devices, but don't sell it for the NPUs and the

14:59.400 --> 15:06.360
like, because it's a generic GPU interface, and so that's kind of the problem there basically.

15:07.240 --> 15:10.360
So yeah. Any more questions?

