WEBVTT

00:00.000 --> 00:11.240
So, the next talk. So, we're actually getting a lot of questions about Dr. Moder

00:11.240 --> 00:17.360
Ranner in all the Applamish channels. And the thought, like, they'd better speak for

00:17.360 --> 00:23.720
themselves. So, I'm actually happy to welcome Eric and Dorin, the talk about Dr. Moder

00:23.720 --> 00:29.960
Ranner. And also, there were some things promised about Lamissy P. I guess. And, yeah,

00:29.960 --> 00:30.960
so, let's hear it.

00:36.960 --> 00:47.720
Hi. Hi, Ron. I'm Dorin. I've joined Dr. 3 years and 10 months ago, initially in one of

00:47.720 --> 00:52.840
the court teams of Dr. Dextop, the one that dealt with the virtual machine that leaves under

00:52.840 --> 01:01.840
the hood of it, all my question, Rindos. And, since about one year, before about one year,

01:01.840 --> 01:05.840
so, I joined and started Dr. Moder Ranner.

01:07.840 --> 01:16.840
So, I'm Eric Curtin. I've worked on a bunch of different open-source things, I guess, we're

01:16.840 --> 01:24.840
getting to details. And now, I lead the Dr. Moder Ranner team. I'm working on all sorts of

01:24.840 --> 01:31.840
different initiatives at Dr. We can go ahead. So, yeah, I introduced Dr. Moder Ranner.

01:31.840 --> 01:42.840
Dr. Moder Ranner is not new, it's around maybe a year. And, yeah, the whole goal is kind

01:42.840 --> 01:51.840
of like, treat AI kind of, like, you would treat a container. And, we'll get into details

01:51.840 --> 01:57.840
both why that makes sense throughout the talk. I'm going to try and go through the slides.

01:57.840 --> 02:02.840
Very fast-paced, me and Dorin. We probably have too many slides, too many demos. And, if you

02:02.840 --> 02:09.840
want to ask us any questions for you, it's free to ask us in the hallway after. So, yeah, the goal

02:09.840 --> 02:19.840
is to package models in a standardized way, to achieve portability. So, what are the things

02:19.840 --> 02:24.840
that made Docker great is, you could have kind of container locally under local machine

02:24.840 --> 02:30.840
right? And then, you could scale out to like the same containers to Kubernetes nodes, which is

02:30.840 --> 02:37.840
quite useful. So, the same idea here, local clouder edge, simplified scaling. And, we also had

02:37.840 --> 02:46.840
Docker compose functionality if you wanted to define your applications in that way. And, yeah, all

02:46.840 --> 02:53.840
to benefits that come with containers. To be honest, run across Linux, macOS windows. I'm trying

02:53.840 --> 03:00.840
to think, because marketing always complain when I get this wrong. The top one is the

03:00.840 --> 03:08.840
logo, and the dragon is our new mascot. So, it's a dragon, and it's cool. So, there's

03:08.840 --> 03:12.840
one key feature I'd like to discuss, which I think is one of the nicer things about

03:12.840 --> 03:19.840
Docker model. So, the key difference between Docker model and all turn-dives is a lot of the

03:19.840 --> 03:24.840
alternatives, expect you to pull from one very specific type of repository over the public

03:25.840 --> 03:31.840
internet, something like the Lama repository or the hooking face repository. And, I

03:31.840 --> 03:36.840
believe there are some disadvantages. This approach, like, you kind of have a single point

03:36.840 --> 03:42.840
of failure. I don't know where these repositories behind cloud flare, for example, but let's

03:42.840 --> 03:48.840
say there are cloud flare goes down here. You can't pull the models anymore. You must keep

03:48.840 --> 03:55.840
their models on the public internet for a lot of these repositories. This is kind of

03:55.840 --> 04:00.840
sometimes you train a model with very sensitive data. It might be health that or, you

04:00.840 --> 04:06.840
know, internal company that you probably don't want that on to cross the public

04:06.840 --> 04:10.840
internet. So, there are good reasons you might want to keep certain models in

04:10.840 --> 04:17.840
house, because they're just like huge databases that you can query. Yeah, I'm pushing

04:18.840 --> 04:21.840
over public networks. It's more expensive, right? If you can just push and pull it

04:21.840 --> 04:27.840
within your internal network, that's generally cheaper. And, there's kind of an element

04:27.840 --> 04:33.840
of vendor lock in there. You're kind of tied to that repository. And, they're quite

04:33.840 --> 04:39.840
bespoke. They're both kind of a new type of storage repositories. And, the main goal is

04:40.840 --> 04:47.840
just to push and pull gigabytes of data. It's not really a problem we've

04:47.840 --> 04:52.840
haven't encountered before. So, all the great minds, Docker got together. We

04:52.840 --> 05:00.840
taught a lot of intents, thinking a lot of coffee. And, we're like, this kind of seems

05:00.840 --> 05:08.840
familiar. And, so I said, what if instead of asking people to use Olaima or

05:08.840 --> 05:14.840
hooking face repose, we asked people to use Docker hope. And, we all taught this

05:14.840 --> 05:19.840
was genius, because now we forced everyone into the Docker ecosystem, and we

05:19.840 --> 05:26.840
could all pair rent. So, we're like, yes, that's great idea. So, everyone

05:26.840 --> 05:34.840
agrees. But, no, I'm just joking. The goal was people can use any OCI

05:34.840 --> 05:42.840
registry of their liking. Or hooking this. Because hooking face has a lot of

05:42.840 --> 05:47.840
great models up there, and it's standardized. And, so, you can pull down a

05:47.840 --> 05:52.840
hooking face model in store, has an OCI artifact. I love to have Olaima

05:52.840 --> 05:57.840
on that slide. But, if anyone knows the history of Olaima, Olaima can

05:58.840 --> 06:04.840
start off as Olaima CPU wrapper. And, they kind of give it a Docker like

06:04.840 --> 06:10.840
experience. And, basically, within the last year, they

06:10.840 --> 06:13.840
threw away Olaima CPU and kind of wrote their own

06:13.840 --> 06:20.840
infront engine from scratch. So, they kind of forked the standard, which is

06:20.840 --> 06:24.840
GJWF, so that it only works with Olaima and nothing else, which I don't

06:24.840 --> 06:28.840
know. It's a very community friendly to be honest. So, that's why it's just

06:28.840 --> 06:34.840
or hooking face. Maybe in the future, that becomes Endolama. I don't know.

06:34.840 --> 06:39.840
So, why are OCI registries very suitable? There's many public OCI registries.

06:39.840 --> 06:44.840
Out there, there's Docker Help, GitHub registry, way. Most companies

06:44.840 --> 06:49.840
also have internal OCI registries, which are quite useful. Going back to my

06:49.840 --> 06:52.840
previous slide, there sometimes models are quite sensitive. You don't

06:52.840 --> 06:56.840
want them to leave because, you know, people's private data or whatnot.

06:56.840 --> 07:01.840
And, industry data, I'm sure you guys are modest, but industry data

07:01.840 --> 07:07.840
suggests 98% of the Fortune 500 software companies already have

07:07.840 --> 07:11.840
internal OCI registries. So, you don't have to spin up new infrastructure

07:11.840 --> 07:18.840
or buy new licenses or, you know, with unnecessarily, because at the end of

07:18.840 --> 07:22.840
the day, you're just pushing and pulling gigabytes of data, right?

07:22.840 --> 07:27.840
It's not that complex. And, yeah, it offers great

07:27.840 --> 07:31.840
interoperability. This is more or less the same point again. A lot of

07:31.840 --> 07:35.840
time when people deploy AI models at scale, they're using technologies like

07:35.840 --> 07:40.840
Docker, Docker composing Kubernetes. So, like, it's just naturally

07:40.840 --> 07:46.840
fits in, right? Yeah. So, who here is heard about cloud

07:46.840 --> 07:53.840
but they show hands? Okay, cool. I think cloud but it's really

07:53.840 --> 07:58.840
fruit. Don't use it. I agree. It's scary. And, actually, that's the

07:58.840 --> 08:04.840
wrong slide. It's called what's called it? Yeah, it's called

08:04.840 --> 08:10.840
loop and claw now. Yeah, I go back to don't use open

08:11.840 --> 08:20.840
claw. It is scary. I completely agree. It is almost too powerful.

08:20.840 --> 08:25.840
It can spawn many type of agents. It can talk to your messaging

08:25.840 --> 08:33.840
apps and whatever. But, I think, yeah, this is just a natural

08:33.840 --> 08:36.840
evolution of agents, right? And, it's pretty

08:36.840 --> 08:39.840
scary. So, we wrote a blog post recently about

08:39.840 --> 08:44.840
Docker model runner and open claw. Sorry,

08:44.840 --> 08:50.840
geez, there's too many names. I think, but anyway, the point was

08:50.840 --> 08:54.840
instead of sharing your sensitive data over the public

08:54.840 --> 08:59.840
internet, you just talk to Docker model runner and at least

08:59.840 --> 09:05.840
you're only speaking within your machine. That only

09:05.840 --> 09:10.840
scares one part of the problem. An important part of the problem.

09:10.840 --> 09:13.840
But, the point I'm trying to make is really scary

09:13.840 --> 09:16.840
technology and with agents and all this, it's getting even

09:16.840 --> 09:19.840
scary because we're kind of handing over to keys to do various

09:19.840 --> 09:23.840
tasks. If you're getting into the agents field, please

09:23.840 --> 09:26.840
containerize your workloads. Like, SC Linux, names,

09:26.840 --> 09:30.840
bases, C groups, set comp, all this stuff. It's perfectly

09:30.840 --> 09:34.840
suited to restrict what your agents can do. And, I got

09:34.840 --> 09:39.840
in LinkedIn. I thought I credited him. Oh, yeah.

09:39.840 --> 09:44.840
Muhammad Rabi, he wrote this diagram and he kind of got it.

09:44.840 --> 09:47.840
This is what we're trying to do. It's very important. Please

09:47.840 --> 09:51.840
containerize your AI workloads. I'll hand you back to

09:51.840 --> 09:58.840
our. All right. Yeah. This is a glimpse on how

09:58.840 --> 10:03.840
you can use Docker model runner. The first command shows how

10:03.840 --> 10:09.840
you can search for various models. And, as you can hopefully

10:09.840 --> 10:12.840
see it on the slide, both on Docker hub and on

10:12.840 --> 10:17.840
hacking phase. And then the next two commands show how

10:17.840 --> 10:24.840
you can actually pull from both these OSA registries.

10:24.840 --> 10:29.840
Then here are some of the main use cases. The one we

10:29.840 --> 10:34.840
like a lot is the privacy for sure. And then the

10:34.840 --> 10:40.840
flexibility. And the fact that you can play with a

10:40.840 --> 10:45.840
model on your machine on some not the most powerful

10:45.840 --> 10:50.840
server. And then just switch to some cloud provider or

10:50.840 --> 10:55.840
wind or even your beefy data center for more compute

10:55.840 --> 11:02.840
power. This is how we like to define Docker model runner.

11:02.840 --> 11:07.840
It's cost efficient, private, flexible, portable, and

11:07.840 --> 11:12.840
powerful. And this is it in action. The command my

11:12.840 --> 11:17.840
loop super familiar to you. But with that extra model

11:17.840 --> 11:22.840
added in it. Basically, from the standard Docker run

11:22.840 --> 11:27.840
alpine or some app. This time instead of running

11:27.840 --> 11:32.840
service or an application, you will know then one

11:32.840 --> 11:39.840
large language model. You can then talk to the model using

11:39.840 --> 11:46.840
this straightforward interactive CLI interface or

11:46.840 --> 11:51.840
by using some of the popular clients that are out there.

11:51.840 --> 11:58.840
Including open a web UI mainly for chatting with them. Open

11:58.840 --> 12:03.840
code works with it. Cloud code as well. And yeah,

12:03.840 --> 12:15.840
even the currently trending global. All right. What else?

12:15.840 --> 12:19.840
Oh, yeah. And on this slide, you can see the platform

12:19.840 --> 12:24.840
we have support for. Here, you can see the

12:24.840 --> 12:31.840
same video. Cuda, Wolken, M.D.s. Rock M, but also

12:31.840 --> 12:36.840
Ken and Musa. And I like to highlight that for some of

12:36.840 --> 12:41.840
these. The community members added support for them.

12:41.840 --> 12:45.840
But also different engines like VLLM, as

12:45.840 --> 12:49.840
Jelang or libraries like diffusers.

12:49.840 --> 12:53.840
But let's start with what you usually have in

12:53.840 --> 12:56.840
this. And that's the operating system.

12:56.840 --> 13:02.840
Docker model runner runs on Linux with Docker

13:02.840 --> 13:06.840
CE, the Docker engine. And then there you have

13:06.840 --> 13:10.840
again support for compute platforms like Cuda,

13:10.840 --> 13:15.840
Rock M and the one you see listed on here. On Windows,

13:15.840 --> 13:19.840
it comes bundled with Docker desktop. But you can also

13:19.840 --> 13:26.840
use it independently from it. And there you have both

13:26.840 --> 13:32.840
Cuda support and CPU of course. I think also

13:32.840 --> 13:37.840
AMD rock M nowadays. And Mac OS, again, bundled with

13:37.840 --> 13:41.840
Docker desktop. Regarding inference engines,

13:41.840 --> 13:46.840
we first started with LMSPP, which easily runs on

13:46.840 --> 13:51.840
all of these platforms mentioned before. Then we have

13:51.840 --> 13:57.840
VLLM. We initially added VLLM only for Linux.

13:57.840 --> 14:02.840
Then we added it for Windows as well. But it required

14:02.840 --> 14:06.840
the tiny trick because VLLM by design doesn't work on

14:06.840 --> 14:10.840
Windows. So what we did, we have launched the

14:10.840 --> 14:14.840
VLLM influence engine inside the container running

14:14.840 --> 14:20.840
in WLL2. That's working right because of the

14:20.840 --> 14:25.840
proper GPU pass through you can get within WLL2. Then

14:25.840 --> 14:31.840
we've already, I mean, not yet merged, but we also

14:31.840 --> 14:36.840
have VLLM on Mac OS. But not the standard VLLM,

14:36.840 --> 14:41.840
but instead plugin called VLLM Metal, which is an

14:41.840 --> 14:48.840
initiative started by Docker folks and folks from VLLM.

14:48.840 --> 14:54.840
And under the hood, you have an MLX engine to run the

14:54.840 --> 15:04.840
models. Next up, well, now we have models as new

15:04.840 --> 15:08.840
citizens in the Docker ecosystem. So we had to add them to

15:08.840 --> 15:13.840
the console as well. Like you have dependencies on

15:13.840 --> 15:18.840
some other services or applications or volumes or networks,

15:18.840 --> 15:23.840
today you can also depend on some model served by

15:23.840 --> 15:28.840
Docker model runner. For this, you can configure

15:28.840 --> 15:33.840
its context size, but also some allow the runtime

15:33.840 --> 15:36.840
in the same way. Like temperature,

15:36.840 --> 15:40.840
top case sampling, so on. I was mentioning before that

15:40.840 --> 15:44.840
you also get it, part of Docker desktop. And here is

15:44.840 --> 15:49.840
some of the simple UI you get with it. You can see your

15:49.840 --> 15:55.840
models, search for models, only on Docker hub here, but

15:55.840 --> 16:00.840
also inspect the requests. We provide some API

16:00.840 --> 16:04.840
recorder, which is super helpful for when you either

16:04.840 --> 16:09.840
trying to debug some influence requests, or simply to

16:09.840 --> 16:16.840
see what your AI agent, AI coding agent, does

16:16.840 --> 16:21.840
work better than how it interacts with the large language model.

16:21.840 --> 16:25.840
Now, about VLLM, Eric?

16:25.840 --> 16:29.840
Yes, so I saw the last talk, actually, the

16:29.840 --> 16:32.840
interview talk very interesting, but we spoke a lot about

16:32.840 --> 16:36.840
VLLM, so VLLM is kind of like the leading

16:36.840 --> 16:40.840
data center, great. In front of the engine right now.

16:40.840 --> 16:44.840
So Docker model runner has this concept of back

16:44.840 --> 16:47.840
in, so you can kind of pick whatever

16:47.840 --> 16:52.840
inference engine you like, so Lambda CPP is one, which

16:52.840 --> 16:56.840
probably supports the Y-Story hardware. The LLM is on

16:56.840 --> 16:59.840
S-G-Lang as one. There's a bunch more diffusers.

16:59.840 --> 17:02.840
They also different use cases, but this is kind of the

17:02.840 --> 17:06.840
leading data center, great one. It's very

17:06.840 --> 17:11.840
memory efficient. It does a lot of, memory efficient

17:11.840 --> 17:14.840
may be the wrong word, but it does a lot of caching to

17:14.840 --> 17:17.840
make sure you get maximum throughput from the

17:17.840 --> 17:21.840
front engine.

17:21.840 --> 17:24.840
Yes, so when we started Docker model runner, it only had

17:24.840 --> 17:27.840
the Lambda CPP back end, and this was, this

17:27.840 --> 17:32.840
Lambda CPP is great. I like that community a lot.

17:32.840 --> 17:36.840
But that only gave us functionality to run

17:36.840 --> 17:40.840
G-G-G-WF model type, which was created by Lambda CPP

17:40.840 --> 17:46.840
later, 4, 2, and that's great. But one of the things you see

17:46.840 --> 17:50.840
people do is under local machine, they might get a

17:50.840 --> 17:54.840
safe tensors, and they might do some sort of

17:54.840 --> 17:57.840
conversion transformation and to turn into G-G-WF,

17:57.840 --> 18:00.840
and then run it under local machine, and then in

18:00.840 --> 18:05.840
production, they would use the safe tensors type, which is

18:05.840 --> 18:09.840
okay, but that's not like, for example, when you use

18:09.840 --> 18:14.840
containers, use a container as is on your local machine

18:14.840 --> 18:17.840
or production was. So, what we like to be a

18:17.840 --> 18:21.840
Lambda is it's kind of closer to what's running in

18:21.840 --> 18:24.840
the data center, and it also enabled the safe

18:24.840 --> 18:28.840
tensors type of model, which is quite popular.

18:28.840 --> 18:31.840
G-G-WF and safe tensors are kind of the most popular

18:31.840 --> 18:34.840
formats, and obviously in the Docker model runner case,

18:34.840 --> 18:37.840
we encapsulate them in OCI artifacts.

18:37.840 --> 18:41.840
Yeah, and this gave us more standardization with the

18:41.840 --> 18:46.840
industry. Yeah, so I kind of described this recently.

18:46.840 --> 18:49.840
This is kind of the way things have been up to now.

18:49.840 --> 18:53.840
Like, people would hack on their local machine, the

18:53.840 --> 18:58.840
typical one, Lambda CPP. To me, Lambda CPP in Olam are

18:58.840 --> 19:00.840
kind of the same, because Lambda is kind of a

19:00.840 --> 19:03.840
re-implementation, and it's thought, whatever.

19:03.840 --> 19:07.840
So, people would use that locally. It's more tailored

19:07.840 --> 19:11.840
towards like a single user querying the model at a time.

19:11.840 --> 19:15.840
Single stream, low latency, which has its advantages.

19:15.840 --> 19:18.840
I don't know if you went to the first talk this morning,

19:18.840 --> 19:22.840
but, um, son, he spoke about multimodal.

19:22.840 --> 19:26.840
That's a very great use case where you want low latency.

19:26.840 --> 19:28.840
Start-up time super fast, it's a certain see

19:28.840 --> 19:31.840
process catered towards kind of like embedded on device

19:31.840 --> 19:34.840
inference. Then you have the other

19:34.840 --> 19:37.840
extreme, which is kind of like, even in video-hage,

19:37.840 --> 19:40.840
to undertale us like that, a central production grid,

19:40.840 --> 19:46.840
VLLM. Lambda CPP is like, do it like runs on calculators.

19:46.840 --> 19:50.840
I mean, it runs on the server. VLLM is kind of the

19:50.840 --> 19:53.840
opposite, and video is always kind of been the first

19:53.840 --> 19:57.840
class citizen in VDNX86. I think they brought in AMD

19:57.840 --> 20:01.840
recently, but that's only very recent, and it's more tailored

20:01.840 --> 20:05.840
towards multiple users and multiple streams, so many

20:05.840 --> 20:08.840
people sharing that hardware maximizing the throughput

20:08.840 --> 20:12.840
from that hardware. It's a bit slow, it's much

20:12.840 --> 20:15.840
slower to start-up, but that's okay, because

20:15.840 --> 20:19.840
time's up, okay. Sorry, sorry.

20:19.840 --> 20:26.840
But in here. So Docker created and donated

20:27.840 --> 20:32.840
VLLM plugin, which kind of changes the story. So now you

20:32.840 --> 20:36.840
can run VLLM on your MacBooks, which I think

20:36.840 --> 20:41.840
of Macmini costs $700, so that's great.

20:41.840 --> 20:44.840
I'll hand you back to Dwarntown.

20:44.840 --> 20:49.840
Yeah, this slide was to show that, yeah, supporting many

20:49.840 --> 20:53.840
APIs, makes us support a lot of clients, and some examples

20:53.840 --> 20:58.840
of these API, the fact of OpenAI, then the

20:58.840 --> 21:02.840
Olama, mainly for WebUY, and anthropic workload code.

21:02.840 --> 21:06.840
This is another glimpse of how you can see the

21:06.840 --> 21:10.840
requests from the interaction with the LLM, and

21:10.840 --> 21:17.840
fortunately, we don't have time for the hands-on, but

21:17.840 --> 21:21.840
yeah, you can come to us for any sort of questions. Thank you.

21:22.840 --> 21:25.840
There is actually one question in the chat, so I

21:25.840 --> 21:28.840
very quickly will ask the question, so would we

21:28.840 --> 21:31.840
allow LLM run well on IMD-3xHailer?

21:31.840 --> 21:34.840
It seems to hit rockmen in the comments.

21:34.840 --> 21:40.840
I don't get the question. Would VLLM run well on IMD-3xHailer?

21:40.840 --> 21:44.840
It seems to hit rock and curious about that. I'm not sure

21:44.840 --> 21:47.840
about VLLM.

21:47.840 --> 21:50.840
I'm just going to say, we did have a slide.

21:50.840 --> 21:55.840
This is where I don't know where it is, but we're a fully open

21:55.840 --> 21:59.840
source project, and we get this kind of query all the time.

21:59.840 --> 22:02.840
Please staff, please contribute. The question was basically,

22:02.840 --> 22:07.840
I have an IMD machine. VLLM doesn't work on my

22:07.840 --> 22:10.840
specific machine.

22:10.840 --> 22:13.840
Dorn doesn't own that machine, error curtain doesn't own that

22:13.840 --> 22:16.840
machine, so we really need the communities help.

22:16.840 --> 22:19.840
And it's a device specific problem.

22:19.840 --> 22:24.840
I've only tested VLLM on Kuda and there it works great.

22:24.840 --> 22:27.840
Again, not sure, but you're welcome to contribute.

22:27.840 --> 22:28.840
Thank you.

