WEBVTT

00:00.000 --> 00:11.640
All right, hi, my name is Ian Romanock, it's been a long time since I've been in front

00:11.640 --> 00:18.320
of a graphics dev room, it's good to be back up here, I'm actually awake now, so that's

00:18.320 --> 00:26.600
helped things much better than I was earlier today, so I'm going to talk about as part

00:26.600 --> 00:35.480
of the compiler work that I do on Mesa, I do a lot of running a thing called FossilDB,

00:35.480 --> 00:40.160
so Fossil, I'm going to start at the beginning in case for people who don't know, Fossil,

00:40.160 --> 00:48.280
a Fossil is a serialization format that was created by Valve Software, they use it to capture

00:48.280 --> 00:55.360
shaders and some auxiliary state from through the Vulcan API so that they can save those things

00:55.440 --> 01:00.440
offline and use them to do offline shader compilation, so that when people download a game,

01:00.440 --> 01:04.720
it comes with a shader cache, pre-populated, and the first time the user starts the game,

01:04.720 --> 01:12.880
it doesn't take an hour, now FossilDB is a collection of Fossils that the open source

01:12.880 --> 01:21.320
community has collected over the years, right now, the private, there's a public FossilDB

01:21.400 --> 01:26.200
and a private FossilDB, the private FossilDB unfortunately is the more interesting one,

01:26.200 --> 01:34.840
it has dumped from a little over 700 games and benchmarks and things like that and there's

01:34.840 --> 01:41.320
kind of in the ballpark of two and a half million shaders in it, so it's a big chunk of stuff,

01:41.320 --> 01:48.200
we don't use it for pre-filling shader caches, we use it for testing compiler changes, and

01:48.280 --> 01:55.880
there are, you'll see a lot of commit logs in Mesa where people will log, so some shader

01:55.880 --> 02:02.200
statistics about, you know, ran all the, try to compile all the Fossils and collect metrics

02:02.200 --> 02:07.560
from the shaders before your change, do the same thing after, compare changes in the metrics,

02:07.560 --> 02:13.880
and try to decide, did this change make things better, did it make things worse, we kind of

02:13.880 --> 02:20.360
use this as a proxy for real performance testing, because as long as compiling two and a half

02:20.360 --> 02:27.160
million shaders takes, actually running benchmarks on 700 applications would take orders of magnitude

02:27.160 --> 02:34.200
longer, so it's a pretty good approximation for performance, especially for things that maybe

02:35.080 --> 02:40.200
if you have a long series of changes that you're working on over a period of several years

02:40.200 --> 02:45.960
and you have, this change is going to make a half percent change, and that's going to make a half

02:45.960 --> 02:50.600
percent change, you may not actually be able to measure any of the difference across any of those,

02:50.600 --> 02:57.480
and you don't want to say, well over two years, we made 3% improvement like that, so this is

02:57.480 --> 03:03.160
something that is reproducible that you can actually collect the data in a reasonable way,

03:03.160 --> 03:17.240
but there are additional uses for fossil DB, using it, using it in conjunction with a bunch

03:17.240 --> 03:24.040
of internal validation passes that we have in the compiler that look for inconsistencies in the

03:24.040 --> 03:31.960
intermediate code, sort of as the shader is going through the whole process of compilation,

03:32.040 --> 03:42.040
can find a bunch of corner cases and invariants that get their violated, and we have a lot of

03:42.040 --> 03:48.120
regular functional test cases, but two and a half million shaders that were written by random people

03:48.120 --> 03:55.240
is going to catch a lot more stuff, and over the last year or so we found a whole bunch of extra

03:55.240 --> 04:01.160
bugs that the only thing that found them was running them through fossil DB, like it escaped through

04:01.240 --> 04:07.480
all of the CTA, hundreds of thousands of tests in the Vulcan CTS, all these things slipped right

04:07.480 --> 04:14.280
on by, but there was one shader in fossil DB. There's a couple of these that one shader in fossil DB

04:14.920 --> 04:24.840
hit, like, okay, cool. So it's, we can use it as a proxy for performance testing and also as a

04:24.920 --> 04:34.520
proxy for functional testing. Now, I use fossil DB a little bit differently than the way that most

04:34.520 --> 04:42.040
people do. I think that most people, when they have an MR, they'll collect data before the whole

04:42.040 --> 04:49.240
MR and after the whole MR on a single platform that they care about. I am a bit, especially in the

04:49.240 --> 04:56.040
presence of possible functional regressions, I run it on every platform on every commit.

04:57.880 --> 05:03.000
Every platform that we at least care about a little bit. And so what that means is

05:04.520 --> 05:08.040
I'm going to go through and I'm going to run all two and a half million of the shaders

05:08.840 --> 05:16.680
with the compiler configured for one of six different Intel platforms. So it takes a long time.

05:17.240 --> 05:25.480
It takes a long, long time to compile what? 15 million shaders now?

05:27.960 --> 05:34.680
So I wouldn't even attempt it on this laptop. It would take, I think I figured out it would probably

05:34.680 --> 05:41.800
take about 18 hours to do a single commit on this laptop. And this laptop is not that old.

05:41.800 --> 05:49.320
It gets low wattage, but still. So it turns out that compiling millions of shaders is the

05:49.320 --> 05:55.560
sort of problem that we refer to as embarrassingly parallel. They're all independent. You just

05:56.120 --> 06:01.800
you can completely, just have a bunch of workers, one worker for each CPU that is just going to

06:01.800 --> 06:08.360
pull another piece of work to do off the queue and keep doing that until everything is done.

06:08.360 --> 06:16.440
And this is the exact usage model that the Vulkan API was designed around. And the piece of

06:16.440 --> 06:26.200
software that we use to actually collect the shader metrics that's part of the fossilized

06:26.200 --> 06:33.800
package from Valve called fossilized replay, it already does this internally. So it would seem

06:33.800 --> 06:41.080
that the answer to make this go fast is, you know, we need cores, lots, lots and lots of cores.

06:43.240 --> 06:50.600
So over the years, I have one of my friends in the background says that I'm a, that was a

06:50.600 --> 06:55.960
pack rat, a pigeon, something. Anyway, he calls me something because I find things that are still

06:55.960 --> 07:02.600
useful that other people are getting rid of. And I go repurpose it. And over the years, I've collected

07:02.680 --> 07:10.680
a handful of machines that in their prime would have been ridiculously expensive and would have

07:10.680 --> 07:20.840
been considered like outstandingly powerful. So I have a 36 core 72 thread, Zion E5 V3.

07:21.880 --> 07:30.120
So CPU came out in 2014 and one of those CPUs would have cost $4,000 and the server that someone

07:30.120 --> 07:36.200
was sending off to eWaste has two of them. And like 128 gigs a RAM and it was just going to

07:36.200 --> 07:46.120
I'm like, I think I know what I can use this for. I, you know, I'll be honest, I try not to think

07:46.120 --> 07:51.880
about that. I try not to think about that. But thank you for bringing that up, Luke.

07:52.760 --> 08:00.680
Since then, I also picked up a 44 core 88 thread, Zion E5 V4. So the next generation,

08:01.400 --> 08:10.360
and then most recently a 64 core 128 thread, Zion platinum. Now, doing just like a quick

08:11.080 --> 08:16.360
sort of fat thumb estimate, it seems like even the slowest of these machines should be

08:17.000 --> 08:25.400
five to six times as powerful as like this laptop. Like they should be just like dividing

08:25.400 --> 08:31.000
numbers of cores. Like it should be at least five or six times as powerful. When I started

08:31.000 --> 08:36.840
running on these machines, that was not the speed up that I got. I was getting a little bit less

08:36.840 --> 08:44.920
than a four times speed up. And I kept hearing this voice in the back of my mind saying, why isn't

08:45.000 --> 08:52.680
it faster? It should be faster. But I had already gotten a really huge speed boost and I could

08:52.680 --> 08:56.840
had a big noisy machine somewhere in a room that I didn't have to think about. And those things,

08:56.840 --> 09:03.320
they each one of them sounds like a leaf blower. It's, it's ridiculous. For a while, we had it

09:03.320 --> 09:08.760
had one of them in the back of a big shared cube that we have. And every time it came out, people were like,

09:09.640 --> 09:19.720
come on. Yeah, yeah. So, but I just like, I had this voice in the back of my head wanting to know

09:19.720 --> 09:24.840
why isn't it faster, but I just like so many bad feelings. You just suppressed them.

09:26.120 --> 09:32.200
And until one day, you can't anymore. And what happened was on, I think it was on the 88th

09:32.200 --> 09:39.640
red system. I had H top running. And I noticed a couple of weird things while H top was running

09:39.640 --> 09:47.480
on that. One, there would be times where the system would be really fully loaded. And then there'd

09:47.480 --> 09:54.120
be like two threads active on the machine for a minute or two. And then it would burst back up

09:54.120 --> 09:58.680
to being really loaded. And then there would kind of had like the sinusoidal patterns of

09:59.480 --> 10:03.320
not being very utilized. And over a period of Iran, I don't know if this shows this yet. So you

10:03.320 --> 10:11.000
can see that on an 88th red machine, there's only about 51, like barely half the machine is used.

10:11.000 --> 10:18.600
What, what the heck? So it, I thought about for a second. And I pretty quickly figured out,

10:18.600 --> 10:24.520
you know, one, what was going on. And two, like, well, there's where all my performance is going.

10:25.480 --> 10:35.160
So the fossilized program that actually processes the fossils takes a single fossil at a time.

10:36.120 --> 10:40.840
And the fossil contains the dump of all the shaders from some application. And it'll take

10:40.840 --> 10:46.440
every shader that's in that fossil and it'll throw it at all the CPUs. And they'll just keep

10:46.440 --> 10:53.000
working away on it. But at some point, there's going to be fewer shaders left to work on

10:53.080 --> 10:59.160
than there are CPUs. And so during that time, you are wasting CPU time.

11:00.040 --> 11:06.520
And I don't think that this is something that the people who made fossilized replay ever even considered,

11:06.520 --> 11:13.880
because, you know, not that long ago, 8 or 12th red CPU machines, like that was the most common.

11:13.880 --> 11:18.360
And even still like, you know, raise your hand if you have a 128th red machine.

11:18.920 --> 11:26.360
All right, there's like five. So, you know, if you have an 8th red machine and half of it,

11:26.360 --> 11:31.320
and it's only half loaded for three minutes, like you've wasted 12 minutes of CPU time. Like,

11:31.320 --> 11:41.960
okay, you know what, maybe I should care about that, but I don't. But if you've got a 72th red machine,

11:41.960 --> 11:47.960
and it's only 50% loaded for three minutes, okay, you've wasted two and a half hours of CPU time.

11:48.520 --> 11:53.480
And that starts to feel like when you're talking about a runtime, you know, on this laptop,

11:53.480 --> 12:00.760
that's 18 hours, we're wasting two hours. Okay, all right, you got my intention now.

12:01.560 --> 12:09.080
So, how can I get that back? How can I use that time that's just wasted?

12:10.040 --> 12:14.760
So, my first thought was, I'll just modify fossilized replay.

12:14.920 --> 12:20.120
I'll modify it so that it can process multiple fossils at once. And as one is getting done,

12:20.120 --> 12:26.280
it'll start executing shaders from the other. But kind of the way that Vulcan is structured,

12:26.280 --> 12:32.840
and the way that you enable features and do things, it seemed like that would be very

12:32.840 --> 12:39.240
intrusive to fossilize replay. And it's really far out of the scope that that program was designed for.

12:39.240 --> 12:44.040
So, one, it seemed like it was going to be a big pile of work that I kind of didn't feel like doing.

12:44.760 --> 12:49.800
And two, I wasn't 100% confident that it would get accepted upstream.

12:50.600 --> 12:56.520
So, I just, I put that idea on the back burner while I tried to think of a better solution.

12:57.320 --> 13:04.120
And then kind of around that same time, someone else in the community made a piece of software called

13:06.680 --> 13:13.560
called DQPrunner that runs tests from the DQP test suite all in parallel. That will wait.

13:13.640 --> 13:19.240
I could maybe I could make a separate runner that just detects when, oh look,

13:19.240 --> 13:23.960
this instance of fossilized replay is getting about done. I'll go ahead and start another one

13:23.960 --> 13:31.160
to use up some of those available CPU resources, kind of in the same sense as what DQPrunner does.

13:32.600 --> 13:41.080
Now, the challenges are a little bit different. With DQPrunner, all the tests are inherently

13:41.160 --> 13:45.080
single threaded, so it's really easy to detect. Oh, one gets done, start a new one.

13:46.120 --> 13:53.400
But with sitting on top of fossilized replay, it itself is multi threaded, so it's kind of a

13:53.400 --> 13:58.200
different challenge of trying to figure out, well, when, how do I know when it's time

13:59.080 --> 14:06.360
before one is done to start another one overlapping it? And that, that was a head scratcher for a bit.

14:06.360 --> 14:14.040
And I did some internet searching of like, okay, what are good ways to detect when CPU load has

14:14.040 --> 14:21.160
dropped enough? And the answer I found was, don't, because GNU parallel already exists.

14:22.760 --> 14:28.200
It is literally the fully generalized version of that runner. It is the tool

14:29.320 --> 14:33.960
that people have spent a lot of time working on. And they have the other thing,

14:34.920 --> 14:38.920
one thing I find a little bit annoying about the program is the first time when you run it,

14:38.920 --> 14:44.920
it gives you this nag screen about if you use this in your academic work, please give us a shout out.

14:44.920 --> 14:53.800
So here is I am checking the box. So it is very configurable about how you can, how you can send

14:53.800 --> 15:00.760
work to it, how you can have it parcel out work and detect CPU load and do all kinds of things.

15:01.640 --> 15:07.800
But that's kind of, it's kind of the point where the bad news starts in that it's,

15:10.280 --> 15:17.880
I find it, I find parallel to be a bit like e-max. It is this really great tool that can do everything.

15:19.080 --> 15:25.640
And because it can do everything, like the simple tasks and the massively complicated tasks

15:25.640 --> 15:30.840
are about the same level of complexity. And so it kind of makes it difficult to know,

15:30.840 --> 15:36.600
if I want to accomplish a particular thing, what is actually the right way to do it.

15:38.600 --> 15:45.800
And so I started, I went off on kind of a false path. And that is kind of most of

15:46.680 --> 15:54.200
what I'm going to talk about now. So there was originally this script in the fossil DB

15:54.920 --> 16:00.680
repository that goes through and collects the big list of fossils that are going to be processed.

16:00.680 --> 16:08.440
And it's got its wrapped in this big for loop. And it starts a single instance of fossilized

16:08.440 --> 16:16.920
replay one at a time. So the starting point was I took this script and started making it work with

16:16.920 --> 16:26.280
with parallel. So my first change, my first thought was, well, what I'll do is I'll take

16:26.280 --> 16:33.000
this line of the script. The one that's actually doing the real work. And instead of actually

16:33.000 --> 16:38.760
executing that, I will dump that whole line of text out to a file. And that, each line of that

16:38.760 --> 16:44.840
file is what I'll send to parallel of like these are the pieces of work that I want you to do in

16:44.840 --> 16:52.920
parallel. And that is a super easy way to get started with it. But it doesn't work out very well

16:52.920 --> 16:57.400
in the long run. And it's definitely, I don't think the way it's intended to be used.

16:59.960 --> 17:07.960
The first thing that I tripped over is the old troll in any kind of unix shell scripting of

17:07.960 --> 17:13.480
as soon as you start having file names that have single quotes or double quotes or spaces in them,

17:14.440 --> 17:20.360
how do I get that in my script file so that it can get the quoting can get into the script file

17:20.360 --> 17:26.520
and be preserved. And out of the script file and be preserved. And the answer is, you can't,

17:26.520 --> 17:41.000
it's madness. You have better luck facing Kathulu. And yeah, so the better way to accomplish this

17:41.000 --> 17:47.960
is what you really want to do with parallel and I'll have another example of this in a bit.

17:49.080 --> 17:57.560
If you want to create a list of, a list of the work items that it's going to parcel out,

17:57.560 --> 18:02.600
and then a program that it's going to work on them with. So instead of creating the giant list of

18:02.600 --> 18:07.800
here's a bunch of command lines that I want handled in parallel, I create a list of,

18:09.800 --> 18:16.120
above here of, here's just the files I want to be operated on and then take each of these files

18:16.760 --> 18:21.160
and then this is the program with some extra command line invocation.

18:22.440 --> 18:28.760
This is how to operate on each line from this file. And that's sort of the more natural way for

18:28.760 --> 18:35.000
parallel to work. And I've found that a very easy way to adapt it into some other

18:36.600 --> 18:42.280
previously kind of inherently serial pipelines that that I have and some other things.

18:43.080 --> 18:49.160
Okay, so how much, all this work was to see if we could get some speed up. So how much speed up

18:49.160 --> 18:55.240
could I get? Well, I looked at at five different machines. I looked at the the two Zions. I looked at a

18:55.240 --> 19:04.520
big relatively modern desktop that I have. I looked at this laptop and then another fairly old

19:04.520 --> 19:12.360
desktop that I have. And the most interesting thing was so on the 12 core machine, there was no

19:12.360 --> 19:22.760
change on this laptop because it has performance cores, efficient cores, and low power efficient

19:22.760 --> 19:31.320
cores. There was a huge amount of variation in the run times just based on did a really tough

19:31.320 --> 19:37.320
shader get scheduled on one of the really slow cores. And so there was like, there was so much

19:37.320 --> 19:43.080
variation in the run times that I think there was that it's effectively the same either way, but it's

19:43.640 --> 19:50.760
I would have to do a bunch more 18 hour runs to find out I'm not gonna. But on the other machines,

19:51.480 --> 19:56.760
it was worth the effort. Even on the regular old desktop, it was a 4 percent sent,

19:57.480 --> 20:07.480
4 percent speed up, which is decent. On the massive machine, 42 percent, like that starts to feel like

20:07.480 --> 20:16.920
real time. It still doesn't quite get it up to the what the anticipated speed up, but

20:17.880 --> 20:26.920
it still felt pretty good. There are some areas of for future improvement to this.

20:27.720 --> 20:35.000
I haven't actually spent a lot of time optimizing this. Once I got it working and before I

20:35.000 --> 20:41.160
even measured what the speed up was, I could tell like, oh, it's going by a lot faster. I can get

20:41.160 --> 20:47.240
through a lot more of this done. It's fast enough, right? And so I didn't want to spend any more

20:47.240 --> 20:54.200
of my time to spend less of that time. But there are a few areas where I think it could be improved

20:54.200 --> 21:03.240
that I would be a pick at in the future. Parallel has a whole bunch of different modes where you can

21:03.320 --> 21:14.040
control when it is going to start new work. You can specify how many idle cores, how many

21:14.040 --> 21:23.080
what load percentage you can specify. Once don't start another worker, unless there's at at

21:23.080 --> 21:29.320
least 20 percent of memory available, like there's a whole, it is a massive chunk of the

21:29.320 --> 21:36.760
man page of here's all the different tools that you have to control how this does. So I think perhaps

21:37.480 --> 21:43.960
adjusting that load percentage, especially on the smaller core, on the CPUs with a smaller number

21:43.960 --> 21:53.480
of cores may help on those, but I haven't experimented with it yet. I also think that optimizing

21:53.560 --> 22:02.680
the order of the fossils might help because there are some, there is one from one application where

22:02.680 --> 22:10.440
that single fossil, it has so many shaders and some of them are massive, massive computers,

22:10.440 --> 22:17.160
it takes 20 minutes by itself. And so ordering it of putting some of like putting

22:17.160 --> 22:24.040
into my intuition says if you put a fossil that has some of these really, really long poll shaders

22:24.040 --> 22:31.000
in it followed by a bunch of them that go very fast. So we have that one fossil that takes 20 minutes,

22:31.560 --> 22:37.960
there's a few dozen that take under a second. So by kind of mixing them like that, that maybe that

22:37.960 --> 22:44.520
could help better utilize the CPUs. It seems worth experimenting with.

22:48.120 --> 22:55.160
Parallel also has a mode where kind of like this CC, you can say, I have all these computers.

22:56.040 --> 23:06.120
Go run it on all of them. And since I have all those computers. And so then that kind of opens up

23:06.120 --> 23:11.560
to make it interesting of, well, if I have just a handful of those 24 core desktops, you know,

23:12.280 --> 23:17.000
my current computer and my computer for a couple of years ago and my computer for a couple of

23:17.000 --> 23:22.200
years before that, well, I don't, I just spread it out across like use it as a cluster.

23:24.520 --> 23:28.600
It seems more likely that someone will have that in their house than will have one of those

23:29.320 --> 23:34.840
the ones. And then you can heat up multiple rooms in your home instead of just the one.

23:39.240 --> 23:49.320
Yes, yes, yes. Oh, and then playing around with some scheduling modes or something, I don't,

23:49.320 --> 23:54.520
I don't think Parallel has a way to do this, but if there's a way to try to schedule,

23:55.160 --> 24:02.600
I want this work item to run on the performance cores. Don't put it on deficient cores,

24:02.600 --> 24:08.680
because if it comes at the wrong place in the pipeline, I'm going to be waiting for it to finish

24:08.680 --> 24:16.360
because it's running on on the crummy cores. But once I, I kind of got to know this tool,

24:16.360 --> 24:22.200
it turned into one of those. When you got hammer and you got really good hammer, like everything

24:22.280 --> 24:28.120
starts to look like a nail. And I started finding a whole bunch of other places in my workflow,

24:28.120 --> 24:36.840
where oh, I could, I can use this tool to make this thing go faster. So I had a script that I wrote,

24:36.840 --> 24:46.440
oh my gosh, 10 years ago, maybe, maybe more to, I hate that every digital camera, all the

24:46.440 --> 24:53.720
file names are uppercase. It doesn't use optimized JPEG tables and it sets weird file mode bits.

24:53.720 --> 25:01.080
So I wrote a script that will use JPEG Tran to pull in all the JPEG images off the memory card,

25:01.080 --> 25:08.360
optimize the, the Huffman tables, fix the file names, fix the file file modes. And I've used the

25:08.360 --> 25:16.680
script for years and years and years. But what I, it has always kind of bugged me, but never enough

25:16.680 --> 25:22.600
to do anything about it that it's going to, it's going to do pull in like 80 pictures from my vacation,

25:23.320 --> 25:27.240
it's going to take a while. Like it's because it's going to do one at a time. It's going to take a while.

25:28.200 --> 25:36.200
But if I split this, if I split that one script up into two scripts where I have sort of a

25:36.200 --> 25:45.320
leaf script that does the actual work of running JPEG Tran on the file, figuring out the fixed

25:45.320 --> 25:52.200
file name and fixing the file mode, and then have the other part that finds all the files and just

25:52.200 --> 25:59.640
sends that to parallel. Like it's now basically bandwidth limited on how fast my SD card reader is

26:00.280 --> 26:07.720
and it goes a lot faster. And it was another problem of, you know, I'm sure I could like spawn

26:07.720 --> 26:11.800
a bunch of these in the background, but I have to figure out how many CPUs there about like, no,

26:11.800 --> 26:21.800
this, this change took me five minutes. So my actual scripts, they, I've been having trouble with

26:21.800 --> 26:30.120
internet here today. I think like everyone else, they will be in this GitLab repo for, for shared

26:30.120 --> 26:37.000
DB on the free desktop GitLab. There's a bunch of other scripts in there that I use for things too,

26:37.000 --> 26:45.320
but that's where the parallel fossil DB runner is. If you have some other machines with

26:46.280 --> 26:51.080
weird, you know, big numbers of course, I'd be very curious to hear if you get what kind of

26:51.080 --> 27:01.560
a speed up you get for the serialized version versus the parallel version. And that's it.

27:01.560 --> 27:05.480
Do you have other any questions? Yes, in the back.

27:05.720 --> 27:13.320
I sit in front of the steady. Instead of these. Someone took a look at the

27:13.320 --> 27:23.440
igen. OK. OK. I'm actually used the CC in, long time, OK.

27:24.440 --> 27:31.080
Yes. There's a evaluation if you're looking to spin up across lots of tasks across the

27:31.080 --> 27:41.080
So the suggestion was it's called hypercue.

27:41.080 --> 27:43.080
Hypercue did.

27:43.080 --> 27:45.080
Okay.

27:45.080 --> 27:47.080
Okay.

27:47.080 --> 27:49.080
Okay.

27:49.080 --> 27:53.080
I don't want them to stop like a slurring dog.

27:53.080 --> 27:55.080
Right.

27:55.080 --> 27:57.080
Right.

27:57.080 --> 27:59.080
Okay.

27:59.080 --> 28:01.080
Okay.

28:01.080 --> 28:03.080
Okay.

28:03.080 --> 28:08.080
So hypercue has a way to start a bunch of jobs across a cluster.

28:08.080 --> 28:09.080
Okay.

28:09.080 --> 28:10.080
Okay.

28:10.080 --> 28:11.080
Does it?

28:11.080 --> 28:22.080
So the one of the nice things about and some of the changes that I would have to make to the scripts that I have to make them work across a cluster using parallel is.

28:23.080 --> 28:29.080
There's a bunch of the sort of temporary file management that I'm still doing myself and parallel can do that.

28:29.080 --> 28:37.080
And when you let parallel do the temporary file management, it knows which files are temporary files and which things are result files.

28:37.080 --> 28:46.080
And so it knows which things to push out to the remote machines and then which things to pull back and which things to just delete.

28:46.080 --> 28:48.080
I have after it's done.

28:48.080 --> 28:50.080
And that seems convenient.

28:50.080 --> 28:52.080
But that does this.

28:52.080 --> 28:54.080
And so I mean it's important for this.

28:54.080 --> 28:55.080
You don't trust them either.

28:55.080 --> 28:57.080
Because you you parallel file systems.

28:57.080 --> 28:58.080
Okay.

28:58.080 --> 28:59.080
Okay.

28:59.080 --> 29:00.080
Yeah.

29:00.080 --> 29:01.080
I haven't domesticated this at all.

29:01.080 --> 29:02.080
Okay.

29:02.080 --> 29:04.080
It's just a very quite minimal.

29:04.080 --> 29:06.080
Single binary type situations.

29:06.080 --> 29:07.080
It's very quick to do.

29:07.080 --> 29:08.080
Okay.

29:08.080 --> 29:11.080
To try out the very similar to parallel.

29:11.080 --> 29:14.080
But I mean it's a it's a workload management.

29:14.080 --> 29:19.080
So you you would to you that one of these tests.

29:19.080 --> 29:21.080
This is a big big one.

29:21.080 --> 29:24.080
So it gets 20 course.

29:24.080 --> 29:25.080
Okay.

29:25.080 --> 29:26.080
Oh, okay.

29:26.080 --> 29:29.080
I wouldn't even schedule this optimally.

29:29.080 --> 29:30.080
Oh, interesting.

29:30.080 --> 29:33.080
Yeah, because that and that's the thing with.

29:33.080 --> 29:36.080
With using fossilized replay.

29:36.080 --> 29:38.080
Is that it.

29:38.080 --> 29:43.080
You can tell it how many course to use, but it you basically tell it all of them.

29:44.080 --> 29:45.080
Yeah.

29:45.080 --> 29:47.080
I don't know if you ever end them.

29:47.080 --> 29:50.080
This fossilized would have less than.

29:50.080 --> 29:51.080
The 64.

29:51.080 --> 29:54.080
I mean, especially on the on the really large machines.

29:54.080 --> 29:57.080
There are there are some some of the fossils that.

29:57.080 --> 30:00.080
Are quite small that maybe have 30 or 40.

30:00.080 --> 30:02.080
So maybe not.

30:02.080 --> 30:03.080
Such a big win.

30:03.080 --> 30:04.080
Yeah.

30:04.080 --> 30:06.080
Interesting.

30:06.080 --> 30:07.080
We'll go here in front.

30:07.080 --> 30:09.080
And then you'll be next.

30:09.080 --> 30:11.080
At the beginning when when you're looking at the problem.

30:11.080 --> 30:14.080
You just said, well, there is this one runner.

30:14.080 --> 30:18.080
But you don't be worked with jobs that are seeing the threat it.

30:18.080 --> 30:19.080
Right.

30:19.080 --> 30:21.080
Well, so that's so there is.

30:21.080 --> 30:25.080
Yeah, there's another runner thing called DQP runner.

30:25.080 --> 30:26.080
Yeah.

30:26.080 --> 30:30.080
And it's basically.

30:30.080 --> 30:32.080
Are you familiar with it or.

30:32.080 --> 30:33.080
No, no, no.

30:33.080 --> 30:35.080
And my question is.

30:35.080 --> 30:36.080
There's a.

30:36.080 --> 30:37.080
What didn't.

30:37.080 --> 30:40.080
Why didn't you try to basically make your problem fit?

30:40.080 --> 30:44.080
Make your replay to sing and spread it.

30:44.080 --> 30:48.080
And then you don't have to problem of guessing how much.

30:48.080 --> 30:49.080
First, basically.

30:49.080 --> 30:53.080
So the the reason for that is that.

30:53.080 --> 30:57.080
A big part there's a big part of work that happens.

30:57.080 --> 31:01.080
Once per fossil that if you did it.

31:01.080 --> 31:08.080
If you made it so that.

31:08.080 --> 31:10.080
So you would end up with one of two problems.

31:10.080 --> 31:15.080
You would either end up with invoking fossilized replay to compile.

31:15.080 --> 31:19.080
A single shader, which would have a lot of extra overhead.

31:19.080 --> 31:25.080
Because there's a there's a bunch of stuff that it has to do to set up Vulcan so that it can even compile it.

31:25.080 --> 31:30.080
Or you would run fossilized replay on a single thread.

31:30.080 --> 31:33.080
And so then if you have.

31:33.080 --> 31:36.080
If you have the.

31:36.080 --> 31:41.080
The the fossil from the the game I won't name that that takes 20 minutes.

31:41.080 --> 31:43.080
Now that that 20 minutes.

31:43.080 --> 31:46.080
I mean that's that's 20 minutes on the 88 thread machine.

31:46.080 --> 31:48.080
So it would take hours.

31:48.080 --> 31:52.080
And so even if you started that one first now you're going to end up with sort of.

31:52.080 --> 31:56.080
Like the transpose of the original problem of now you've got it.

31:56.080 --> 31:57.080
It's at the end and it's.

31:57.080 --> 31:59.080
So you do have a lot of fossils.

31:59.080 --> 32:01.080
But yeah, this one is too long.

32:01.080 --> 32:03.080
So the long tail at the end is still.

32:03.080 --> 32:04.080
Yeah, yeah, yeah.

32:04.080 --> 32:08.080
Yeah, and it's and it's there is a huge variation in in.

32:08.080 --> 32:11.080
I mean there are ones that take a quarter of a second.

32:11.080 --> 32:16.080
All the way up to the one that takes 20 minutes and every amount of time in between.

32:16.080 --> 32:18.080
And so it makes it.

32:18.080 --> 32:20.080
I see what you're saying.

32:20.080 --> 32:26.080
But there are other factors that that kind of make that more more of a challenging way to try to.

32:26.080 --> 32:30.080
To try to reshape the problem in in that way.

32:30.080 --> 32:33.080
I was going to go with him first and I'll come back to you.

32:33.080 --> 32:34.080
Which is.

32:34.080 --> 32:35.080
Would.

32:35.080 --> 32:37.080
We make a sense of trying.

32:37.080 --> 32:38.080
Focus on the.

32:38.080 --> 32:39.080
Master.

32:39.080 --> 32:40.080
Uh.

32:40.080 --> 32:41.080
A society.

32:41.080 --> 32:42.080
Presence.

32:42.080 --> 32:43.080
Yeah.

32:43.080 --> 32:44.080
Yeah.

32:44.080 --> 32:47.080
So that's that's kind of the idea of of optimizing the order of the.

32:47.080 --> 32:51.080
The fossils of like there's ones that we that I know are.

32:51.080 --> 32:52.080
Yeah, our hard.

32:52.080 --> 32:58.680
that I know are, yeah, are hard, start them first.

32:58.680 --> 33:08.880
Oh, well, so generally, I mean, I'm going to end up running all of them.

33:08.880 --> 33:18.200
So in terms of optimizing runtime, there's sort of two separate things there about picking

33:18.200 --> 33:25.680
a subset or reordering them is reordering them might allow you to run all of them more efficiently.

33:25.680 --> 33:32.160
And if there is a subset that you know is more likely to be problematic, then just running

33:32.160 --> 33:35.800
those, like you're just skipping a whole bunch of work.

33:35.800 --> 33:48.160
Yeah, so usually what, so the kind of the workflow that, let me go back.

33:49.160 --> 33:56.160
The workflow that I have gone through that ended up with almost all of these is when I'm going

33:56.160 --> 34:02.160
to test something of mine, I have to do the before run and when I do get at some point in the before run,

34:02.160 --> 34:11.160
there's some fossil that hits an assertion failure, hits a segment fault, has some kind of unexpected

34:11.160 --> 34:17.160
termination and that's like, okay, my whole, my whole work today is going to be different now than I thought

34:17.160 --> 34:20.160
it was going to be 10 minutes ago.

34:20.160 --> 34:21.160
And I have just one.

34:21.160 --> 34:27.160
I know, like, okay, this one is broken and so then I can do get bisect and instead of running all of

34:27.160 --> 34:33.160
them, I will just run that one until I figure out what, what committed is the, that broke

34:33.160 --> 34:40.160
that and then can usually figure out like, it's this app and this one and that one and that one.

34:40.160 --> 34:48.160
And can usually, I'll usually at least list all of them in the commit message of this change

34:48.160 --> 34:54.160
fixes, fixes this bad, this other bad commit on, on these various fossils.

34:54.160 --> 35:02.160
So in the fact, let's comment on the, on the background between, say, run the mold as parallel as it can versus a,

35:02.160 --> 35:06.160
a small, many of them have more shaders than you have course anyway.

35:07.160 --> 35:09.160
Right.

35:09.160 --> 35:13.160
And I mean, worst case, no, you have a easy one shaders.

35:13.160 --> 35:17.160
And so you run the first eight day and that's one for long period.

35:17.160 --> 35:22.160
I mean, but you could also, you don't have to go to single thread versus, right.

35:22.160 --> 35:24.160
Maximum, you go nine course.

35:24.160 --> 35:25.160
But that one.

35:25.160 --> 35:27.160
Yeah, and that would, that would be, I mean,

35:27.160 --> 35:34.160
so the suggestion is to, instead of going from the one extreme of all the course to the other

35:34.160 --> 35:38.160
extreme of just use one core, like, pick some middle ground.

35:38.160 --> 35:49.160
I think that kind of falls in with some of the other, like, trying to adjust, like the load setting on parallel and some of the other things to sort of adjust and tune.

35:49.160 --> 35:55.160
That could be combined with, if you had a more free and actual scheduler.

35:55.160 --> 35:56.160
Right.

35:56.160 --> 35:57.160
Like it's learned.

35:57.160 --> 35:58.160
Yeah.

35:58.160 --> 35:59.160
Yeah.

35:59.160 --> 36:04.160
But they actually sort of keep track of actually core number one, two three and four.

36:04.160 --> 36:06.160
I currently basically run in this god.

36:06.160 --> 36:11.160
So when I have run it, I had the display here.

36:11.160 --> 36:17.160
I mean, kind of why I didn't go much farther than I did is so on this, all right.

36:17.160 --> 36:27.160
So at this point, I was getting, like, 48 out of 88 threads are in use with just the work that I've done on this machine.

36:27.160 --> 36:31.160
It kind of hovers around 70, 75.

36:31.160 --> 36:36.160
And like, okay, that's, he need 90% utilization.

36:36.160 --> 36:39.160
Like, there's, there's not that much left to ring out of it.

36:39.160 --> 36:46.160
And how much, how much compute time am I going to save for how much time am I going to spend trying to get that extra mile?

36:46.160 --> 36:50.160
And, you know, in another six months, I'm going to find another e-waste machine.

36:50.160 --> 36:53.160
This one will be 256 threads.

36:54.160 --> 36:57.160
That's not as easy as you do.

36:57.160 --> 37:02.160
The big opening would be that this would be a paradise across the entire machine.

37:02.160 --> 37:03.160
Yeah.

37:03.160 --> 37:04.160
Yeah.

37:04.160 --> 37:08.160
We had a small, not too small work units, but right.

37:08.160 --> 37:12.160
Not that this low is young that you started that intervening is still.

37:12.160 --> 37:13.160
Yeah.

37:13.160 --> 37:19.160
And so that's sort of a general, like, when it comes time to schedule the things across multiple machines,

37:19.160 --> 37:28.160
the challenge there, and it's part of why I haven't done that yet, is there is such a variation in the power of the machines,

37:28.160 --> 37:37.160
is that, yeah, if you end up with the wrong work unit or the last work unit or whatever on the slowest machine,

37:37.160 --> 37:45.160
you're wasting less CPUs time in some sense, but it could actually make it take a little bit longer.

37:46.160 --> 37:47.160
Yeah.

37:47.160 --> 37:48.160
Yeah.

37:48.160 --> 37:59.160
And a lot of the way that I end up working is I'll have one of the machines doing some stuff and then some other experiments on one of the other machines.

37:59.160 --> 38:06.160
So I'll be using the machines, but for different, I'm parallelizing them in different ways.

38:06.160 --> 38:09.160
I was going to go over here quick.

38:09.160 --> 38:10.160
Yeah.

38:10.160 --> 38:22.160
So when you talked about phospholite replay, you said that if you want to run to phospholite together, you will have the problem of this matching features.

38:22.160 --> 38:23.160
Can you?

38:23.160 --> 38:28.160
I'm just wondering, can you actually create multiple contacts for the same device?

38:28.160 --> 38:29.160
Yeah.

38:29.160 --> 38:38.160
And so the question is, because of the different feature requirements of the fossils, could you create multiple contexts to account for that?

38:38.160 --> 38:41.160
That is basically what you would do.

38:41.160 --> 38:48.160
The problem, the program that we use to do this kind of thing for OpenGL shaders does exactly that.

38:48.160 --> 38:53.160
It creates a bunch of processes.

38:53.160 --> 38:59.160
It creates one process for each kind of OpenGL API there is.

38:59.160 --> 39:07.160
And then those create multiple threads, and it farms out the work to those in that way.

39:07.160 --> 39:09.160
To do basically that.

39:09.160 --> 39:12.160
And it's quite a bit of added complexity.

39:12.160 --> 39:15.160
And that's kind of the thing that I wasn't.

39:15.160 --> 39:21.160
I didn't feel 100% confident that that work would get accepted upstream.

39:21.160 --> 39:25.160
And I didn't want to have my own branch just kind of hanging around.

39:25.160 --> 39:30.160
But yeah, I think that that is the way that you would have to implement that.

39:30.160 --> 39:34.160
It was just because I think a pop-boss line has been played and was used by Steve playing.

39:34.160 --> 39:35.160
It does.

39:35.160 --> 39:36.160
Yeah.

39:36.160 --> 39:43.160
Up until when I think all type of line captures where it's like, then it just kind of never just turns it off.

39:43.160 --> 39:44.160
Right.

39:44.160 --> 39:49.160
So I was thinking if you have some things that maybe Steve can actually use.

39:49.160 --> 39:51.160
But yeah, it's been done that now.

39:51.160 --> 39:52.160
Yeah.

39:52.160 --> 39:56.160
Well, so they're, I mean, they had utilized it in a different way.

39:56.160 --> 39:59.160
And they were running it on their big servers elsewhere.

39:59.160 --> 40:03.160
So I suspect that they just had.

40:03.160 --> 40:09.160
I'm not sure how they deployed it, but I, I suspect they had different problems.

40:09.160 --> 40:10.160
Then.

40:10.160 --> 40:11.160
Um.

40:11.160 --> 40:13.160
What's your memory use?

40:13.160 --> 40:14.160
It's like I need it.

40:14.160 --> 40:15.160
Oh, so let's see.

40:15.160 --> 40:20.160
If there, if there's learning too many threads, then the schedule works usually just fine.

40:20.160 --> 40:21.160
Yeah.

40:21.160 --> 40:27.160
If you need one after the other, but at some point when you're just forming too many threads, you're running out of memory or crashing your passes.

40:27.160 --> 40:29.160
So more and then it gets slower, okay?

40:29.160 --> 40:33.160
So there is cash issues, but this like.

40:33.160 --> 40:36.160
So this is running for the, it's a serialized version.

40:36.160 --> 40:41.160
But even running in parallel, I have never seen this machine use more than eight gigs.

40:41.160 --> 40:42.160
Okay.

40:42.160 --> 40:46.160
Like, I don't, I don't understand how we're not using more memory than that.

40:46.160 --> 40:53.160
But yeah, that was, that was a thing that for sure surprised me.

40:53.160 --> 41:02.160
And if that were a problem, um, there are additional parameters that you can give to parallel to tell it like don't, don't over commit memory.

41:02.160 --> 41:09.160
In addition, so like, um, one of the things that parallel doesn't know how many new threads or how much memory usage.

41:09.160 --> 41:10.160
Right.

41:10.160 --> 41:12.160
We'll go to use when it starts one next come on.

41:12.160 --> 41:14.160
So there will be at least two.

41:14.160 --> 41:15.160
That's true.

41:15.160 --> 41:16.160
Yeah.

41:16.160 --> 41:22.160
You would, you would have to, you would have to, you may have to do some additional hand tuning and,

41:22.160 --> 41:27.160
and be careful about how you set the, um, so it's not in the slide.

41:27.160 --> 41:28.160
So it's always more grace.

41:28.160 --> 41:30.160
So like, you remember when your machine hit five?

41:30.160 --> 41:31.160
Yeah.

41:31.160 --> 41:32.160
Yeah.

41:32.160 --> 41:37.160
And, and on that machine, you know, when I got it, it had 128 gigs.

41:37.160 --> 41:42.160
Like, I've never even, never even come close to exhausting that.

41:42.160 --> 41:43.160
Yeah.

41:43.160 --> 41:44.160
Send it back.

41:44.160 --> 41:45.160
What's the level of package?

41:45.160 --> 41:47.160
Are you, are you within it?

41:47.160 --> 41:49.160
So I, I closed the level of it.

41:49.160 --> 41:50.160
Um.

41:51.160 --> 41:52.160
Well, on that.

41:52.160 --> 41:54.160
I mean, it's those.

41:54.160 --> 41:59.160
I, I, I, again, that can be a subject of the, I can recommend to two of us.

41:59.160 --> 42:01.160
I think it should work on the over the ones.

42:01.160 --> 42:02.160
Yeah.

42:02.160 --> 42:04.160
I have not looked at, at that.

42:04.160 --> 42:09.160
I mean, I know that they, that is, that is probably another thing.

42:09.160 --> 42:17.160
Another area to, uh, for future optimization is to look at the thermal usage and to see if we're actually.

42:17.160 --> 42:21.160
If overloading it is causing it to use slower clocks and that that's.

42:21.160 --> 42:32.160
I, I, I, I, I, I, I, I, I, I, I, I, I don't want to know how many watts have been spent.

42:32.160 --> 42:33.160
I, I, I, I, I don't want to know how many watts have been spent.

42:33.160 --> 42:35.160
I try not to think about it.

42:35.160 --> 42:40.160
Like, just, just, just those machines being on is a huge, is a huge draw.

42:40.160 --> 42:45.160
And especially when they're going full tilt, it's, it's kind of, it's kind of ridiculous.

42:46.160 --> 42:49.160
All right. Well, thanks, everyone.

42:49.160 --> 42:52.160
Uh, thanks for having me. I get, I guess you're all free to go home now.

42:52.160 --> 42:54.160
This is the, the end of the session.

42:54.160 --> 42:59.160
Thank you.

