WEBVTT

00:00.000 --> 00:16.000
So, our penultimate talk got one more after this, but we have Yakov, tell us about Asik AI

00:16.000 --> 00:17.000
Sports on Linux.

00:17.000 --> 00:18.000
So, take a look.

00:18.000 --> 00:19.000
Thank you, Matthew.

00:19.000 --> 00:20.000
Okay.

00:20.000 --> 00:21.000
Hi, everyone.

00:21.000 --> 00:23.000
Thank you for joining us.

00:23.000 --> 00:29.000
So, today, I'm going to give you a brief overview of Colonel and Yutus-based

00:30.000 --> 00:33.000
AI Asik's sporting Linux.

00:33.000 --> 00:34.000
Okay.

00:34.000 --> 00:40.000
So, AI Asik, we can also say and use or whatever we are using these days.

00:40.000 --> 00:43.000
So, a little bit about myself.

00:43.000 --> 00:48.000
I basically work in permanent engineering.

00:48.000 --> 00:50.000
I do a lot of embedded Linux.

00:50.000 --> 00:56.000
I try to build like a Yakov Builder-Root OpenWT competitor,

00:57.000 --> 00:58.000
built system.

00:58.000 --> 01:00.000
And most of my work was network focused.

01:00.000 --> 01:04.000
But recently, I've been going more low level,

01:04.000 --> 01:10.000
and going towards more Linux hardware and enablement stuff.

01:10.000 --> 01:14.000
So, I work at BiteLab, BiteLab is a product design company.

01:14.000 --> 01:17.000
So, basically, BiteLab design hardware.

01:17.000 --> 01:21.000
And it also focuses on software development as well.

01:21.000 --> 01:29.000
And low amount of manufacturing, particularly focused for

01:29.000 --> 01:35.000
basically like development during design.

01:35.000 --> 01:36.000
Okay.

01:36.000 --> 01:42.000
So, to start off, we have to talk a little bit about hardware.

01:42.000 --> 01:47.000
So, there's enormous amount of hardware and Asik's for,

01:47.000 --> 01:52.000
well, basically, AI.

01:52.000 --> 01:56.000
And there's quite a lot of things we have to consider here.

01:56.000 --> 02:02.000
So, first of all, we need to consider what training needs for AI.

02:02.000 --> 02:08.000
So, for example, back propagation, the fundamental algorithm is still used.

02:08.000 --> 02:14.000
Here, I suppose the GPUs are good, but they are not really optimally in the sense that,

02:14.000 --> 02:21.000
for example, if you do inferencing, you don't want to have a leftover die unused.

02:21.000 --> 02:27.000
Also, there's a lot of work being done on, basically,

02:27.000 --> 02:32.000
enabling the low level mathematical functions to be optimised,

02:32.000 --> 02:34.000
sped up, and what not.

02:34.000 --> 02:36.000
That's on the training side.

02:36.000 --> 02:40.000
On the inference side, we have, you know, the standard stuff,

02:41.000 --> 02:42.000
as well as the training.

02:42.000 --> 02:45.000
So, we have a couple of matrix, multiplication for that.

02:45.000 --> 02:50.000
We have convolutions, and we have various different activation functions.

02:50.000 --> 02:53.000
So, like, sigmoid, like relu, whatever.

02:53.000 --> 03:00.000
Also, we have to consider that some Asik's are a purpose specific, right?

03:00.000 --> 03:03.000
So, for example, today we are not doing just text.

03:03.000 --> 03:06.000
We're also doing images between sound.

03:06.000 --> 03:09.000
Everything goes into Asik design, I would say.

03:09.000 --> 03:15.000
Here, we can also mention there's a big difference between approaches to Asik design,

03:15.000 --> 03:18.000
which I think most a lot of the talks were focused on.

03:18.000 --> 03:24.000
So, for example, we know that there's Stolic arrays being kind of like,

03:24.000 --> 03:29.000
in a position to some kind of meshed compute units, right?

03:29.000 --> 03:33.000
I suppose here there are some fundamental trade-offs, right?

03:33.000 --> 03:38.000
For example, in Stolic arrays, it's a bit harder to do conditional branching,

03:38.000 --> 03:40.000
blah, blah, blah, you know.

03:40.000 --> 03:47.000
And last but not least, we have the usual problem of how much power we use in our Asik,

03:47.000 --> 03:48.000
right?

03:48.000 --> 03:50.000
We want to cut down this.

03:50.000 --> 03:55.000
Actually, we want to improve the top's, the terror of the race in a second,

03:55.000 --> 03:57.000
provide, right?

03:57.000 --> 04:04.000
I'm just going to give you like a small overview of how things are progressing.

04:04.000 --> 04:08.000
This is kind of like a slide, which, you know, classic goes up.

04:08.000 --> 04:15.000
But I think what's most important here is that we went from, you know, general purpose stuff,

04:15.000 --> 04:19.000
like CPUs, we went over to GPUs because, you know,

04:19.000 --> 04:22.000
3D was very important and still is.

04:22.000 --> 04:26.000
But now, because of the advent of AI and machine learning,

04:26.000 --> 04:31.000
we're trying to optimize all of this and complexity rises.

04:31.000 --> 04:37.000
But that doesn't necessarily mean that NPUs are more complex or better than GPUs.

04:37.000 --> 04:42.000
They're just more focused on a particular task, right?

04:42.000 --> 04:46.000
So I'm just going to give you like a small couple of examples.

04:46.000 --> 04:49.000
So we can start off with Google's GPUs.

04:49.000 --> 04:51.000
I think everyone knows them.

04:51.000 --> 04:54.000
They started off around 2015.

04:54.000 --> 04:58.000
And basically, there's been quite a lot of development.

04:58.000 --> 05:02.000
There was the first one, which was quite revolutionary.

05:02.000 --> 05:04.000
And then we went over to recent trillion.

05:04.000 --> 05:08.000
And I think there's even one new one.

05:08.000 --> 05:12.000
Maybe this year or the next.

05:12.000 --> 05:15.000
There's also a new kid on the block here.

05:15.000 --> 05:19.000
There's 10 store-end, which the guys gave a talk about.

05:19.000 --> 05:25.000
So TPUs were, Google's ones were most focused on systolic design.

05:25.000 --> 05:27.000
So basically, data goes in.

05:27.000 --> 05:29.000
It kind of goes like a heartbeat.

05:29.000 --> 05:33.000
But here, we're doing like meshed compute units, right?

05:33.000 --> 05:35.000
Thanks, this course.

05:35.000 --> 05:38.000
But what's the most important here?

05:38.000 --> 05:42.000
And what I would like to draw your attention to is software stacks.

05:42.000 --> 05:44.000
And specifically on Linux.

05:44.000 --> 05:46.000
So what do we have here?

05:46.000 --> 05:47.000
This is important.

05:47.000 --> 05:51.000
We also have a couple of things we need to consider.

05:51.000 --> 05:53.000
We need to consider training and inference.

05:53.000 --> 05:56.000
So these are two different things.

05:56.000 --> 05:59.000
For example, suppose you want to beefy training, right?

05:59.000 --> 06:06.000
But maybe a returning is not as big as hardware or power for hardware.

06:06.000 --> 06:11.000
We can also do beefy inference, usually on the cloud, right?

06:11.000 --> 06:18.000
But I suppose a lot of focused recently is on the edge and local AI, right?

06:18.000 --> 06:21.000
So we want to do kind of lightweight inference.

06:22.000 --> 06:25.000
So maybe less than 10 tops.

06:25.000 --> 06:29.000
But still, it depends on what you actually want to do.

06:29.000 --> 06:36.000
If you want to analyze graphics, if you're doing like 30 frames per second or 60,

06:36.000 --> 06:38.000
it's a big difference.

06:38.000 --> 06:41.000
So on the software side, the stack is important.

06:41.000 --> 06:45.000
So what we want to consider on Linux is obviously the kernel support,

06:45.000 --> 06:47.000
and also the user space.

06:47.000 --> 06:50.000
What kind of APIs are exposed to the user space?

06:50.000 --> 06:55.000
So and here we'll find quite a lot of divergence, I would say.

06:55.000 --> 07:01.000
Last but not, well, there's enormous amount of frameworks, right?

07:01.000 --> 07:03.000
Basically, we all know TensorFlow.

07:03.000 --> 07:08.000
We know there's TensorFlow Lite, which is for inferencing,

07:08.000 --> 07:12.000
PyTorch, Keras, Jux, and so on, and the stack.

07:12.000 --> 07:15.000
Some are more low-level, like run times,

07:15.000 --> 07:19.000
and some are more focused towards model design and whatnot.

07:20.000 --> 07:24.000
And one, I would say, maybe overlooked aspect,

07:24.000 --> 07:27.000
is basically if you have some hardware,

07:27.000 --> 07:32.000
maybe it actually influences what you want to use on the model side.

07:32.000 --> 07:38.000
So for example, we have an example here, which is on the raw chip platform.

07:38.000 --> 07:42.000
And surprisingly, they're kind of like, well,

07:42.000 --> 07:45.000
solution, vendor solution,

07:45.000 --> 07:49.000
and they enforce a use of a particular format model for the weights,

07:49.000 --> 07:51.000
and whatnot, right?

07:51.000 --> 07:54.000
So this is also one aspect.

07:54.000 --> 07:58.000
So let's review on the kernel side, what do we have here?

07:58.000 --> 08:00.000
And this is interesting.

08:00.000 --> 08:04.000
So is there any kind of generic infrastructure for implementing AI

08:04.000 --> 08:06.000
accelerator devices on Linux?

08:06.000 --> 08:10.000
And there actually is something, and it's called

08:10.000 --> 08:11.000
right?

08:11.000 --> 08:14.000
And here you can see the manufacturing entry,

08:14.000 --> 08:19.000
basically the idea is for it to be kind of like a framework

08:19.000 --> 08:22.000
for all kinds of acceleration devices.

08:22.000 --> 08:28.000
We know that usually these were kind of connected to graphics,

08:28.000 --> 08:29.000
right?

08:29.000 --> 08:32.000
So they are usually rendered nodes,

08:32.000 --> 08:37.000
so you would go into, I don't know, like dev card or dev render,

08:37.000 --> 08:39.000
and then you would use that, right?

08:39.000 --> 08:45.000
But these are more, you know, very, very connected to GPUs.

08:45.000 --> 08:50.000
Sometimes you can even skip this kind of thoughts,

08:50.000 --> 08:54.000
and you can go direct hardware.

08:54.000 --> 08:57.000
And I'll give you like a brief overview of various different

08:57.000 --> 09:00.000
how various different vendors approach this.

09:00.000 --> 09:03.000
So Google, they have a couple of accelerators,

09:03.000 --> 09:07.000
development boards, and cards, PCI ones.

09:07.000 --> 09:10.000
And obviously there's a big difference between cloud

09:10.000 --> 09:12.000
GPUs and edge GPUs.

09:12.000 --> 09:16.000
I don't know if you probably have heard of CoreL AI, right?

09:16.000 --> 09:21.000
And together with this, Google basically had a particular

09:21.000 --> 09:24.000
kernel driver, as far as I can see,

09:24.000 --> 09:26.000
it was out of three.

09:26.000 --> 09:30.000
And this is the gasket and apex driver, right?

09:30.000 --> 09:34.000
It was, I suppose that it still has it used,

09:34.000 --> 09:37.000
especially for legacy hardware, right?

09:37.000 --> 09:41.000
But Google is now pushing us onto light RT,

09:41.000 --> 09:43.000
or TensorFlow Lite.

09:43.000 --> 09:48.000
And this framework uses the delegate mechanism,

09:48.000 --> 09:51.000
and I was surprised to see that it actually supports different

09:51.000 --> 09:52.000
vendors, not just Google.

09:52.000 --> 09:54.000
So for example, you can have Qualcomm hardware,

09:54.000 --> 09:58.000
MediaTek hardware, and even Apple hardware,

09:58.000 --> 10:03.000
but only if you have, for example, CoreML delegate, right?

10:03.000 --> 10:07.000
And what's, I suppose, important about this,

10:07.000 --> 10:11.000
is that vendors probably have to add support for this, right?

10:11.000 --> 10:14.000
Okay, so let's take a look at Rochtchip.

10:14.000 --> 10:15.000
What do we have at Rochtchip?

10:15.000 --> 10:17.000
But obviously focusing on the kernel.

10:17.000 --> 10:19.000
So we have two paths here.

10:19.000 --> 10:22.000
We have one, which is completely open source,

10:22.000 --> 10:25.000
and this is kind of a big thing for me.

10:25.000 --> 10:29.000
And I was quite surprised and excited to see this.

10:29.000 --> 10:33.000
So there's actually a driver, kernel driver,

10:33.000 --> 10:37.000
which was made by Tomel's, I think,

10:37.000 --> 10:41.000
and he did basically the whole reverse engineering

10:41.000 --> 10:46.000
of the Rochtchip NPU, which is around six stops, right?

10:46.000 --> 10:52.000
And you can basically do the whole exploration stack on this.

10:52.000 --> 10:57.000
I'll show you in a minute what I, what I, how it looks like, right?

10:57.000 --> 11:00.000
But what's important here is that we actually have,

11:00.000 --> 11:02.000
like, some kind of API description.

11:02.000 --> 11:04.000
You can find it in a kernel.

11:04.000 --> 11:07.000
It's Rocket Excel, right, header file.

11:07.000 --> 11:10.000
And that's pretty much it on the fourth side.

11:10.000 --> 11:15.000
Obviously, the OM Rochtchip has its own Open API, right,

11:15.000 --> 11:16.000
in the kernel side.

11:16.000 --> 11:20.000
And it basically has closed firmware blob,

11:20.000 --> 11:23.000
which the driver uploads to the NPU,

11:23.000 --> 11:28.000
and then it basically interacts between the NPU and the kernel, right?

11:28.000 --> 11:34.000
And this is very, something very similar to my work on network switching hardware,

11:34.000 --> 11:38.000
because switching hardware and the packet processing basics

11:38.000 --> 11:40.000
basically do the same thing, right?

11:40.000 --> 11:43.000
Nobody wants to expose their core.

11:44.000 --> 11:49.000
Okay, so, okay, on the MDN video and the Intel side,

11:49.000 --> 11:52.000
well, we all know what's up in there.

11:52.000 --> 11:54.000
I won't spend too much time here.

11:54.000 --> 11:57.000
There's also an interesting thing.

11:57.000 --> 12:02.000
There's the Intel NPU driver, which basically

12:02.000 --> 12:09.000
kind of respects the new driver Excel infrastructure in the kernel,

12:09.000 --> 12:13.000
so that was quite good to see.

12:13.000 --> 12:16.000
So, there's a couple of other stuff.

12:16.000 --> 12:21.000
There's obviously an enormous amount of new A6 and A6 designs.

12:21.000 --> 12:27.000
So, 10 stores, obviously, they have their own TTKMD stuff,

12:27.000 --> 12:30.000
which exposes a new node under 10 stores.

12:30.000 --> 12:32.000
There's Qualcomm.

12:32.000 --> 12:36.000
One interesting aspect about Qualcomm is that they inherited

12:36.000 --> 12:39.000
the whole NPU design is inherited from the DSP,

12:39.000 --> 12:43.000
and this is why, for example, the drivers from the NPU is under media, right?

12:43.000 --> 12:46.000
Instead of, I don't know, like DRM or something like that.

12:46.000 --> 12:53.000
And obviously, the node name is, you know, quite specific to Qualcomm, right?

12:53.000 --> 12:59.000
On the new hardware, so the NPU MSM was for the old hardware,

12:59.000 --> 13:02.000
the new hardware does some even more advanced stuff.

13:02.000 --> 13:05.000
It does, well, some kind of packet marshalling,

13:05.000 --> 13:08.000
and then you control the cores, and then you get data

13:08.000 --> 13:10.000
and shuffle data around it.

13:10.000 --> 13:11.000
What not?

13:11.000 --> 13:14.000
MediaTek obviously has their own.

13:14.000 --> 13:18.000
I was also, it was surprising to see that there's

13:18.000 --> 13:21.000
a mention of Iroha NPU, and I thought this was, oh, wow,

13:21.000 --> 13:24.000
it's an NPU, but no, it's a network pack,

13:24.000 --> 13:27.000
a processor unit for packet acceleration.

13:28.000 --> 13:32.000
So, on Apple side, New Orleans, well, I suppose,

13:32.000 --> 13:35.000
we might see it, might be reverse engineered.

13:35.000 --> 13:39.000
I thought maybe I saw Helinos guys did it, but we'll see.

13:39.000 --> 13:42.000
Okay, so we are done with the kernel.

13:42.000 --> 13:44.000
What do we have on the user space side?

13:44.000 --> 13:47.000
So, we'll take a look at various different vendors,

13:47.000 --> 13:51.000
but I would like to go back to Rockchip, because it's interesting.

13:51.000 --> 13:54.000
So, again, there's the two approaches.

13:54.000 --> 13:58.000
One is completely open source, and it's part of Mesa.

13:58.000 --> 14:02.000
So, Mesa 3D, there's actually a driver called Rockchip,

14:02.000 --> 14:05.000
which communicates with the Rockchip driver in the kernel,

14:05.000 --> 14:10.000
and basically exposes NPU functionality, right?

14:10.000 --> 14:14.000
On the close source side, well, on the vendor side, OEM side,

14:14.000 --> 14:18.000
we have, again, the same thing, we have some kind of open API,

14:18.000 --> 14:23.000
and we have the firmware blow, which is uploaded into the NPU.

14:23.000 --> 14:28.000
So, yeah, that's pretty much it for the Rockchip.

14:28.000 --> 14:34.000
On the Google TPU side, again, we have the lead edge-cat TPU.

14:34.000 --> 14:38.000
I suppose that this is kind of like the runtime.

14:38.000 --> 14:42.000
I couldn't really find out about the cloud part.

14:42.000 --> 14:48.000
I think this is mostly for local AI, local inference, anyway.

14:48.000 --> 14:53.000
So, we have the light RT, or as it was formerly named TensorFlow,

14:53.000 --> 14:57.000
and this is for the edge TPU, so edge local API.

14:57.000 --> 15:02.000
And we already told that they support multiple vendors.

15:02.000 --> 15:06.000
On the GPU side, we know everything about this.

15:06.000 --> 15:08.000
I won't go into much details.

15:08.000 --> 15:11.000
I will mention the Intel OpenVino stuff.

15:11.000 --> 15:14.000
So, if you recall on the kernel side,

15:14.000 --> 15:17.000
there was the Intel driver for MPUs,

15:17.000 --> 15:21.000
and we have kind of like an open firmware for that.

15:21.000 --> 15:22.000
And here it is.

15:22.000 --> 15:28.000
So, we actually do have a fully open source stack for AI

15:28.000 --> 15:32.000
inferencing edge-one, and it's actually on the Rockchip side.

15:32.000 --> 15:33.000
So, take a look.

15:33.000 --> 15:38.000
First off, we start with light RT, or TensorFlow Lite.

15:38.000 --> 15:40.000
We have the Tathlon delegate.

15:40.000 --> 15:45.000
And then we use the Mesa 3D rocket accelerator driver.

15:45.000 --> 15:49.000
Then we go to the kernel to the Excel rocket driver.

15:49.000 --> 15:54.000
And then we interface with the NPU on the rocket SOC.

15:54.000 --> 15:56.000
This is quite good, I would say.

15:56.000 --> 16:00.000
So, to kind of like wrapped things around.

16:00.000 --> 16:04.000
Basically, we have a couple of takeaways.

16:04.000 --> 16:07.000
There's extreme fragmentation throughout the whole stack.

16:07.000 --> 16:10.000
So, but can you even generalize ASIC?

16:10.000 --> 16:15.000
So, we know that some are focused towards more graphic stuff,

16:15.000 --> 16:19.000
some are inheriting stuff from the BSP, what not.

16:19.000 --> 16:23.000
Obviously, a vendor's specificness, and kind of like,

16:23.000 --> 16:26.000
nobody wants to do the hard work of mainlining stuff.

16:26.000 --> 16:32.000
So, every vendor has its own tooling, framework, and stuff there.

16:32.000 --> 16:36.000
I would say there's a big difference between inference and training.

16:36.000 --> 16:43.000
I would say that probably the more performant training stuff is,

16:43.000 --> 16:49.000
it's a higher probability that is completely closed.

16:49.000 --> 16:52.000
Yeah, so what's next?

16:52.000 --> 16:55.000
We have ASICs.

16:55.000 --> 16:59.000
So, I heard mention of language processing units.

16:59.000 --> 17:03.000
So, they're kind of like transformer optimized for particular models.

17:03.000 --> 17:08.000
There's always mention of power, latency, memory, everything is really important.

17:08.000 --> 17:11.000
But I would say that memory is a big deal, right?

17:11.000 --> 17:14.000
And especially memory latency, right?

17:14.000 --> 17:18.000
You want to shuffle data and largely to run quite fast.

17:18.000 --> 17:23.000
And here, I would like to highlight the community.

17:23.000 --> 17:25.000
Linux hardware enablement effort.

17:25.000 --> 17:27.000
So, to me, we saw so.

17:27.000 --> 17:31.000
He worked on rock chip, and he actually added support for this, right?

17:31.000 --> 17:35.000
So, I would say, that's very exciting.

17:35.000 --> 17:40.000
He's also doing work on the very Silicon Vmante MPUs at NaviV.

17:40.000 --> 17:47.000
So, I suppose we'll see more of this erosion engineering work in the kernel, right?

17:47.000 --> 17:51.000
And, obviously, our research hosts, the foundry,

17:51.000 --> 17:54.000
are doing their own ETF platform.

17:54.000 --> 17:59.000
So, excited to see new stuff there.

18:00.000 --> 18:03.000
And that's it. Thank you so much.

18:09.000 --> 18:12.000
We've got time for release a couple of questions.

18:12.000 --> 18:15.000
So, we are on right here.

18:16.000 --> 18:20.000
So, thank you so much for the overview of the software stack.

18:20.000 --> 18:27.000
I'm really curious about how powerful that hardware that you just went through is.

18:27.000 --> 18:34.000
I have a pretty good understanding of how CUDA and NVIDIA GPUs compare to AMD GPUs in Rockham.

18:34.000 --> 18:42.000
But, can I feasibly run any single digit billion parameter language model on a rock chip?

18:42.000 --> 18:45.000
I have absolutely no idea what rock chip is.

18:45.000 --> 18:49.000
And the same question for other kinds of hardware that you mentioned.

18:49.000 --> 18:55.000
So, just a brief overview of this space in terms of actual usability for running actual models.

18:58.000 --> 19:00.000
That's a good question. Thank you.

19:00.000 --> 19:03.000
So, the main focus I would say is that there's a big difference,

19:03.000 --> 19:06.000
like I said, between training and inference.

19:06.000 --> 19:09.000
But, inference, I can split into two parts.

19:09.000 --> 19:16.000
So, basically, like big inference with enormous models and small edge inference.

19:16.000 --> 19:20.000
So, the rock chip part is mostly focused on very small edge stuff.

19:20.000 --> 19:26.000
So, suppose the main reason why this talk came about is because the rock chip hardware supports cameras, right?

19:26.000 --> 19:34.000
And, basically, we have a client who wants to do like AI analysis of the camera feed, right?

19:34.000 --> 19:37.000
And, this stuff you can do, right?

19:37.000 --> 19:44.000
You can, I don't know, run maybe YOLO, right, for object detection and it works surprisingly good.

19:44.000 --> 19:49.000
I don't know about, you know, larger stuff like language models.

19:49.000 --> 19:53.000
I suppose someone else might know.

19:53.000 --> 19:56.000
But, yeah, I don't know.

19:56.000 --> 20:05.000
I don't think we will see kind of like fully open software stack for training.

