WEBVTT

00:00.000 --> 00:10.560
We have our next talk.

00:10.560 --> 00:20.960
So we have Duffing and Martin who are going to give a lovely talk about the various compute

00:20.960 --> 00:24.920
challenges based on using risk 5.

00:24.920 --> 00:27.160
And so take it away.

00:27.160 --> 00:28.160
Hello, everyone.

00:28.160 --> 00:35.160
I'm Duffing and today we are talking about all in a risk 5 and a risk 5, all in a AI.

00:35.160 --> 00:50.000
And my self is students, I'm faster on AI software of computing and so what do we do?

00:50.000 --> 00:55.480
We are making a consumer electronics based on risk 5 chips.

00:55.480 --> 01:07.840
And actually we built the world's first risk 5 laptop in 2021, which is based on our

01:07.840 --> 01:11.080
winner, D1, SOC.

01:11.080 --> 01:20.240
And after then we are making modular framework DIY laptop.

01:20.240 --> 01:27.560
So I think some of us may be know about framework.

01:27.560 --> 01:35.440
It's a software installable and a repairable and a risk replaceable parts, which means

01:35.440 --> 01:44.680
our components of this laptop can be replaced and the upgrade.

01:44.680 --> 01:56.200
And we made several members, which is powered by X86, I'm on the risk 5.

01:56.200 --> 02:11.960
For all risk 5 is our first class 6 and so in 2025 we built our first modular framework

02:11.960 --> 02:27.120
risk 5 laptop and these 4 cores, 1.5 gigahertz so after that we faced some problems.

02:27.120 --> 02:38.960
So our first is we do not know our target market and we do not know people wants how

02:38.960 --> 02:49.960
much computer power, people wants and we have limited money and we need launch our

02:49.960 --> 02:51.920
product on time.

02:51.920 --> 03:03.040
So we began to thinking and we decided to OIN AI because at the time almost most of the

03:03.040 --> 03:15.640
multiple risk 5 SOCs is coupled with AI co-processor on your processor so we think maybe

03:15.640 --> 03:26.680
it's time to OIN AI so there were two typical cases, the first one is the first chip

03:26.680 --> 03:39.800
late solution, the ISOC is introduced by ESWIN and it's has 8 cores and 2 gigahertz, 50

03:39.800 --> 03:51.320
tops AI computing power and it has dedicated NPU which is pretty similar to the NVIDIA

03:51.560 --> 04:08.360
architecture and it has a private tool to use computing power so we can say we cannot use TensorFlow

04:08.440 --> 04:28.200
or Pator directly on this laptop but we can run dedicated Lama not mainstream so this is the

04:28.280 --> 04:43.640
first case the product is launched in October 2025 and we have another case which is instruction

04:43.640 --> 05:03.640
draft which means we have a SOC made by space meet and it has two clusters, 2.5 gigahertz

05:03.640 --> 05:15.160
8 cores which is capable for our A23 profile and this is for the regular CPU and we have another

05:15.160 --> 05:29.720
two clusters is for is dedicated for the AI computing and we have 25 B's vector lens and

05:29.960 --> 05:42.200
much powerful TensorFlow processor which can deliver 60 tops AI computing power so

05:44.280 --> 05:57.480
that is much different compared to the NPU solution and they they deliver the execution

05:57.480 --> 06:09.160
parameter for ONIX runtime so it's capable to run ONIX runtime for all of the ONIX models and

06:10.600 --> 06:23.160
they also have a Lama.CPP but it's also not on the mainstream so we face too much

06:23.240 --> 06:33.560
problems so MPU is neither designed for all different kind of models no customized

06:33.560 --> 06:43.160
risk of AI extension instructions which means if we want to run AI models

06:43.320 --> 07:00.840
maybe the model is formatted in a copy or ONIX or GG UF I don't know but we want to run this

07:00.920 --> 07:13.320
model on this CPU but MPU and the customized AI extension is not fit in this case so

07:15.720 --> 07:26.200
and this company this semiconductor they do not have enough money and enough time so

07:26.840 --> 07:43.960
we decide to run a new AI framework so we get in touch with modular which is founded by

07:44.040 --> 07:55.960
Chrysler and this AI we are based on the AI based on the way they have a lot of money and

07:55.960 --> 08:11.320
they are extremely open and they have a very very open developer community and all of the

08:11.960 --> 08:25.880
compilers and AI software steps is open source so and this is plenty new and the

08:26.520 --> 08:38.680
technical depth so we decide to support module and max this modular open software

08:41.400 --> 08:44.520
adapted software on risk 5 chips

08:47.560 --> 08:51.720
okay I'll stop here and let Martin do the hard path

08:52.680 --> 08:59.960
all right thank you Delfang great job so deep computing is absolutely great at building hardware

08:59.960 --> 09:04.920
they get the security they build a motherboard they get your developer onboard it you can run the

09:04.920 --> 09:09.000
thing you can actually feel it it's not really hardware that's they are absolutely great at that

09:09.000 --> 09:15.080
but AI has a problem that you have a chip but that chip cannot run any model you want

09:15.080 --> 09:19.720
but most of the time like vendor will give you something and the stored of runs sort of

09:19.720 --> 09:25.800
doesn't run depends on what vendor gives you and okay fine like we all experience that right

09:26.360 --> 09:34.920
but I so I used to work on test oriented was until like a month ago I left but I'm right now in a

09:34.920 --> 09:42.440
community developer and and anyway I'm familiar with the with our connector and

09:44.200 --> 09:49.000
test run phase is a lot of the same same challenges as deep computing how to be facing like

09:49.080 --> 09:55.720
we have a chip their the chip is great I have a lot of faith and it's in a chip and how capable

09:55.720 --> 10:04.840
it is but running actual actual workload on there is hard so just as introduction to what's happening

10:05.960 --> 10:11.480
anti-installer and like we call the 106 processor it looks at this it's a great of course

10:12.040 --> 10:16.840
the brown part is compute and the green part is the DDR memory

10:18.840 --> 10:24.040
so it's a grid like you can do some systolic things like your classic systolic array of

10:24.040 --> 10:29.320
literature applies from like the 1980s and 1980s you can do matrix matrix multiplication

10:29.320 --> 10:35.400
very efficiently like this and like you load one of the matrix on one row or you load the other side

10:35.400 --> 10:39.800
of me other column you broadcast them you send them a row that that gives you matrix multiplication

10:41.640 --> 10:47.960
so but with your site course you can do something more fun you don't have to be brilliant and

10:47.960 --> 10:53.960
decide every data flow pattern into your chip you can kind of just hey you know I run software

10:53.960 --> 11:01.160
I am rick five I can control how that thing flows so what do you do you figure out there you can

11:02.200 --> 11:10.200
load matrix into one of the rows do I what's called ring we do's and that's how you get

11:10.200 --> 11:17.640
general general matrix vector multiplication that's what language models care but you know

11:17.640 --> 11:23.800
106 course on a really course they are five cores inside you got to the end up with the

11:23.800 --> 11:32.760
process that's moved data in and out and you get a unpacker matrix matrix and vector units

11:32.760 --> 11:38.920
and a packer the reason they did that was because hardware type of casting you don't need to

11:38.920 --> 11:45.320
expand Skype cycles on decontization reconization you just have hardware do that you know what

11:46.680 --> 11:53.560
and this is the and you write three kernels that's hard you're asking a developer to say hey

11:54.200 --> 11:59.160
no write one piece of code is ready hard enough you have to write three and one of those

11:59.160 --> 12:07.960
have to run three threads at the same time that's pretty hard right all right um how

12:08.040 --> 12:14.680
does that work um if you actually write right actual load of a code this is what you get so

12:16.360 --> 12:21.480
you have something something in it you enter loop that's your core compute do you wait for

12:21.480 --> 12:27.800
data to be available you you make sure that your compute side have the register ready you do the

12:27.800 --> 12:31.880
extra compute and you commit that because you have three cores you have this synchronization

12:32.440 --> 12:38.280
and then you have to prepare the output pipe and make it ready so so you don't raise the output

12:39.960 --> 12:45.560
the hardware is we call that we we have three cores that that doesn't do compute in

12:45.560 --> 12:52.200
self-intripe engine the math engine talks to your own register stuff and so that's why you

12:52.200 --> 12:56.120
have to do synchronization that's different deck I pack into register stuff like that

12:56.600 --> 13:03.400
so if you're familiar with hardware which was I assume most of the people here are

13:04.440 --> 13:10.680
it feels like a pipeline so there's three cores it controls three different hardware it's unpack

13:11.400 --> 13:17.080
math and pack right and you have different registers so effectively you're doing a pipelineing

13:17.640 --> 13:21.880
and at any same time your different cores are working on different different work load and just

13:21.880 --> 13:31.640
pipeline it's a pipeline you feed these are in compute out and yeah so this is the acquire

13:32.280 --> 13:41.080
synchronization pattern across your your three cores so but you know you can abuse this thing

13:43.640 --> 13:48.360
you know it's double buffered right like you're switching between register banks so you can just

13:48.520 --> 13:54.120
you know double buffer so I'll just abuse abuse the fact that register so we'll have the value be there

13:55.560 --> 14:01.480
for any loop that that's switching large enough the first two first two iterators you can do

14:01.480 --> 14:07.160
some expensive compute super components take six of them in the register and afterwards it will be

14:07.160 --> 14:13.960
layer as long as you don't override it so yeah that works so your dog I heard you like

14:13.960 --> 14:21.480
you have more machines so I put data for machines inside your data for machine but back to

14:21.480 --> 14:28.040
about a vector code that we just saw nothing was state four so if you look into the source code

14:28.040 --> 14:34.600
this is what it looks like there's constant register you can assign to and it makes your

14:34.600 --> 14:41.880
computation kind of state four and it's very efficient like trust me it's super efficient but it's

14:41.960 --> 14:47.800
super staple and therefore if you try to edit different stuff and you call the wrong edit function

14:47.800 --> 14:53.480
or forgot one edit function worst case your exponential function becomes logarithmic

14:56.120 --> 15:06.360
oops sorry where am I sorry and the chip is so super efficient it just vector

15:07.320 --> 15:14.600
so like NGPUs or school GPUs but the very very space efficient manner so

15:15.320 --> 15:23.720
any for any block like if an L spot always executed like GPUs so it only affects vector code if you

15:23.720 --> 15:29.400
try to stick actual print functions or scalar computations between the vector of a vector L

15:30.360 --> 15:36.760
both get executed and you have to be careful about divergence so go steeper

15:38.440 --> 15:43.800
actually inside the chip there's a macro expender unit so when you write vector code

15:43.800 --> 15:50.120
the compiler actually look for different different instructions sequences that that's repeating

15:50.120 --> 15:56.920
else it was a hey this is this is a single single macro like an execute and you know just

15:57.640 --> 16:02.520
tell the macro expender hey here's eight eight instructions please execute them while I'm doing

16:02.520 --> 16:10.600
looping logic this gives you double the theoretical maximum instruction throughput without much hardware

16:11.640 --> 16:17.880
that kind of like the healthy SP's does does is or over how looping and this is why you kind of see

16:17.880 --> 16:23.480
weird patterns where you you're incrementing the register value for for some reason that is why

16:23.480 --> 16:29.800
that's because the compiler will generate instruction sequences are identical for every loop and

16:29.800 --> 16:34.680
they will be able to do the macro thing so you'll talk I hear you like data for machines so I put

16:34.680 --> 16:41.480
your data for machines they have for machines they have for machines are like what's that architecture right

16:42.680 --> 16:47.160
but no these are hard problems that is the current standard surface of all stack

16:47.160 --> 16:52.760
the green part is GML because I work on that and I love that thing but we always face the same

16:52.840 --> 16:59.240
computation problem MP's are inherently hard these are efficient machines these are machines with

17:00.120 --> 17:06.680
many many many many tricks of their sleeves they it will squeeze every bit of efficiency out of the

17:06.680 --> 17:15.160
silicon the hard problem is how do you program the thing like you cannot ask your developers to be a

17:15.160 --> 17:22.200
PhD in order to program your chip that's just not viable the solution is we need a battery

17:22.680 --> 17:31.640
work whether whether it be deep computing they think it's modular and my thing is something else

17:31.640 --> 17:38.280
but we need a better better framework for computation and model of our computation

17:39.320 --> 17:43.000
with that thank you

17:43.320 --> 17:53.560
you've got time for a couple maybe one or two questions so some of my race they don't have any questions

17:58.680 --> 17:59.080
okay

18:01.240 --> 18:03.240
there

18:05.960 --> 18:09.640
so if I understood correctly you read the software and he does the hardware

18:09.720 --> 18:18.520
so kind of collaborating a decompany thing is really good to do hardware I'm a HPC program

18:18.520 --> 18:24.200
at heart I write code for different accelerators would decide to make this talk because you know

18:24.200 --> 18:30.040
hardware is hard software is also hard and someone kind of has to be the first one to fix these problems

18:30.120 --> 18:41.240
something I need to add that way deep computing make mudboard but we also build AI software

18:41.240 --> 18:55.400
stack to help our customers to have a much more mature AI software stacks yes so so way deep

18:55.480 --> 19:05.400
computing also builds software especially on Linux kernel and the AI software AI infrastructure

19:08.520 --> 19:11.720
and one else

19:11.720 --> 19:26.280
did I understand correctly that deep computing is building a machine or a piece of silicon

19:26.280 --> 19:32.920
where you have a mix of RVV implementations on the same board one with 256 bit vector like the

19:32.920 --> 19:40.360
execution and one with 1024 that's going to be wild that's going to be really cool

19:43.640 --> 19:58.200
so basically you have to implement to implement a software using vector extension

19:58.280 --> 20:13.800
which is vector links are not stick basically you have to do so but some engineers they do not have

20:13.880 --> 20:31.800
this this kind of field the build the software use SIMD method so how to say I think the

20:31.800 --> 20:46.360
communities may upgrade the legacy works to to to to build vector links are not stick

20:46.360 --> 20:53.720
software thank you very much that is all time we have

