WEBVTT

00:00.000 --> 00:15.600
Hello everyone, so my name is Samuel Deso, I am a founder of Aerithics and I work on monitoring

00:15.600 --> 00:23.200
on observability and I will present today is a summary of an experience I had last

00:23.200 --> 00:33.680
year. I had to implement LLN on a special environment. On a stage last year, it was a

00:33.680 --> 00:43.360
period of 25. I was not an expert on it, so I had to challenge to implement monitoring

00:43.360 --> 00:52.000
observability on a special environment and I have seen various tools. The black box for me

00:52.000 --> 00:59.600
can I do it or can I push it in or can I implement monitoring observability. So I had work

00:59.600 --> 01:07.840
many many times when I had success already. So what I will present today is a summary of

01:07.840 --> 01:15.120
my experience on how to implement observability for II or work run on an HPC. Because,

01:16.560 --> 01:24.160
the fact when you GPU utilization show 95 percent, it's a good fire, it's so wonderful,

01:24.880 --> 01:32.480
but at the same time when you are 75, I will training a job, it's a thing, you can add a question,

01:33.440 --> 01:44.080
what is the problem? The problem, as I said, when you are 95 percent of a GPU utilization,

01:44.080 --> 01:52.720
you can say, well, it looks great, but is it right? I'm not sure. Because, in fact, when we

01:52.720 --> 02:00.960
talk about a GPU utilization, it is a CPU, a lot of the average of AI workloads, because a GPU

02:00.960 --> 02:07.920
when it is at one of your person, it could be spining on a poorly optimized can,

02:07.920 --> 02:15.360
and it can wait in or memory or it can be blocked by the IO. So, actually, it's not the same thing

02:16.000 --> 02:24.160
as a polyclisting. So, if you are going to sum up the thing, you're not monitoring AI workloads,

02:24.240 --> 02:32.160
you're monitoring hardware, it's not the same thing. So, what are the three blind spots,

02:32.160 --> 02:42.000
or what you are missing in this case? So, the solution, it is a free layer, I have built

02:42.720 --> 02:48.240
in this experience, so it is a course layer collection, which can reveal this, I don't plan,

02:48.240 --> 02:55.440
so it is free layer. The first one is in the first sector. I have to check the health of my hardware,

02:55.440 --> 03:01.200
so it is a GPU utilization, which I run for the power of the temperature, and so on.

03:01.920 --> 03:09.440
The second layer, it is most workload. I have to evaluate the training efficiency by the

03:09.440 --> 03:16.800
front board, if I have a ratio of IO and put out the run ratio on the zone, on the sub layer,

03:17.920 --> 03:27.920
it is the model F. So, I can evaluate the progress of my running, in terms of first,

03:27.920 --> 03:31.920
in terms of gradients, in terms of covariance, on the what we are doing in your,

03:32.240 --> 03:39.680
so I will try to present my demo after of the lab as a criterion.

03:41.760 --> 03:47.280
On when I have worked on this project, I am understood,

03:47.840 --> 03:55.840
especially it is not for additionality. The specificity I have to think of it, the source one is the

03:55.840 --> 04:04.560
scale of a cardinality. The second one is the long run engine, because simulation run for

04:04.560 --> 04:12.000
day and four week on an intuition can prove a massive loss in time, on energy or on money.

04:13.280 --> 04:19.440
On the way over run challenge, it is known that it is introduced to the instance,

04:19.440 --> 04:25.520
especially, because it is a typical cycle for money to run, it is to run from compute.

04:28.880 --> 04:36.320
So, the question is, what do I actually monitor? We can't correlate, but we can't aggregate.

04:37.440 --> 04:43.440
I have put an example of a correlation issue, which we have to see after, or come with

04:43.440 --> 04:50.720
that in draw from it, so I have a draw by G or the node, a GPU matrix, let's try it.

04:51.600 --> 04:57.840
But you have to think on a three question to evaluate of the root cause.

04:58.640 --> 05:06.160
When the training, you have training, which is stagnant, is it the model? Is it the data pipeline?

05:07.040 --> 05:10.320
It is a network or let's try it. That's the first question.

05:11.280 --> 05:20.800
The second one, when you monitor, which consumer cluster resource is, when you put a drop,

05:20.800 --> 05:28.160
it's the second one. Under the third one, why did the business, to increase the training

05:28.320 --> 05:38.000
of the school year, when it was received? So, in the project, I already built.

05:39.840 --> 05:45.520
I have been in the study with the work path, the collection with a very

05:45.520 --> 05:50.640
specter, which collects the GPU matrix, and draw by counting the storage

05:51.600 --> 05:59.040
I have on the node export, for metric system. Under the storage, I have used the Victorian

05:59.040 --> 06:08.640
matrix, because in terms of ICAD90, which is a very performant, I have an extra,

06:08.640 --> 06:15.360
which my company will partner since last week with the Victorian matrix, only we will propose

06:15.360 --> 06:22.400
that training, just an extra. On the internal visualization, we have used the graphana,

06:22.400 --> 06:29.600
in terms of dashboard, in terms of alerting, and to distribute the tracing, we will get our base,

06:29.600 --> 06:32.480
and we have a, I have this, a project on the open-term experience.

06:36.160 --> 06:42.720
So, the question, why did the Victorian matrix for HPC? Because it is an open source, so it's a DB,

06:43.360 --> 06:50.000
which is very beautiful for ICAD90 team. The first advantage of the parameters,

06:50.800 --> 06:59.200
with a better complexion than parameters, we have an addition of a somewhat

06:59.200 --> 07:08.160
per second, which is very good on the advantage, it is a promture compatible, in fact, open source.

07:08.240 --> 07:20.480
It's really good. So, if you do this, I will be in a short, quickly demonstration.

07:23.280 --> 07:28.960
So, if you implement a V-stack, Monday, tomorrow morning, when you will go out of the job,

07:30.000 --> 07:36.160
there's a free thing you can do immediately to improve your monitoring. The first action

07:36.880 --> 07:45.200
is to achieve a time-sort core on your body matrix. You can also monitor storage

07:45.200 --> 07:53.680
higher latency per period, because the cellion is a bit clearer. On the LOSGPS, the latency is

07:53.680 --> 08:02.080
not that destroyed, you need to put a little problem. On the last one, you can try,

08:02.080 --> 08:12.160
just trigger in this routine, the job. So, I'm in a screenshot, I'm going to make a team dive

08:12.160 --> 08:20.240
on the one-to-demo. So, it is an example of how I have organized my monitoring on the

08:20.240 --> 08:26.000
side of the team, in terms of the year. So, you can see the first layer on the time of infrastructure,

08:26.640 --> 08:34.880
I'm checking the GPUTization, the tensile curve, etchorem usage. The second layer, it is a workload

08:34.880 --> 08:42.240
efficiency. So, I look at the two new footwear, the data-loading button, the one-to-one,

08:43.520 --> 08:52.880
it's unique, tiny breakdown on the second rhythm. First, the third layer is a model f. So, I look at the

08:52.880 --> 09:02.000
LOSGPS, the gradient norm. At the end, I have tried to cross, to make a correlation, a cross

09:02.000 --> 09:09.440
correlation between this field part. So, if you think you're going to go as well, I'm going to

09:09.520 --> 09:23.200
pray up. So, as a lab, I have built on how you can, after the tutorial, I have access to

09:23.200 --> 09:32.800
the repository to test. So, I present here the different materials I have used to implement

09:32.800 --> 09:39.840
my lab on the other, I have done the email project to LOSGPS. We're differentiating the different

09:39.840 --> 09:46.400
scenarios you can test. On the other hand, it's a texture. So, it's based on the training simulator

09:46.400 --> 09:54.960
that I have to pull, it's data on the VM Engine, it's in Valhya. And at the end, you can see the results

09:54.960 --> 10:07.600
in Dwarfana. So, I will do quickly, I think I have a minute time, just a minute. So, what I will present

10:07.600 --> 10:18.320
here quickly, so, you will find an explanation of this lab of AIOPSability for HBCM. So, the

10:18.320 --> 10:30.320
current Valhya I have described in this demo on the so on. So, it's more a summary on if I go

10:30.320 --> 10:45.680
on the next tutorial. So, yes. So, what I have shown present with my screenshots. So,

10:45.680 --> 10:56.560
actually, it works a sensor. So, you have this presentation, you have a new review of your training

10:56.560 --> 11:02.480
on the other, I have a spending in the different layers, so, in the first tutorial, you have

11:02.480 --> 11:09.600
seen the GPU utilization, the sensor co-authization. The second one is the workflow efficiency.

11:15.440 --> 11:22.240
As I found the help of my model, I thought, do I have a lot of graphics, we have a

11:22.240 --> 11:29.600
gang on the arm, on the so on, I have a very, also the credentials of the LCD, the long red,

11:29.600 --> 11:40.400
on a ZN, you have a course layer, a coalition, on the last poll. So, yes, it's just the last

11:40.480 --> 11:50.400
at the end of the sentence, and we are good. I will share the presentation after the

11:50.400 --> 11:59.200
for example, to you come all the presentation on the sources of this lab on these values.

12:00.000 --> 12:01.200
Thank you very much.

12:10.400 --> 12:18.880
One question? No. Any questions for Sam?

12:18.880 --> 12:20.160
Yeah, JP.

12:20.160 --> 12:26.160
Is this a spending in media or not?

12:26.160 --> 12:34.400
Is this a vendor at non-stakes, or a market in media?

12:34.400 --> 12:36.400
It is agnostic.

12:36.400 --> 12:37.360
Yeah, sure.

