WEBVTT

00:00.000 --> 00:08.000
OK, do you hear me, right?

00:08.000 --> 00:13.000
OK, so first of all, thank you very much for being here.

00:13.000 --> 00:14.000
I'm Danieli Mingol, I'm today.

00:14.000 --> 00:17.000
I'm going to talk about one AI, an open source framework

00:17.000 --> 00:20.000
for managing AI models at scale.

00:20.000 --> 00:22.000
A bit of introduction about myself.

00:22.000 --> 00:25.000
I'm Danieli. I work as a software developer for AI

00:25.000 --> 00:28.000
in a top enabled system. I work in the innovation team,

00:28.000 --> 00:33.000
so we try to implement good things into the product.

00:33.000 --> 00:35.000
Regarding my background, I have a machine in the science,

00:35.000 --> 00:38.000
measuring computer science, and I've been working over the last few years

00:38.000 --> 00:40.000
on different things in different industries,

00:40.000 --> 00:43.000
mobile gaming, logistics, delivery food,

00:43.000 --> 00:45.000
machine learning system, maybe test,

00:45.000 --> 00:47.000
a lot of experimentation that engineering things.

00:47.000 --> 00:51.000
Regarding my company, we are the first open source infrastructure

00:51.000 --> 00:54.000
as a service solution. We are more than 15 years of experience

00:54.000 --> 00:57.000
in building enterprise infrastructure software.

00:57.000 --> 00:59.000
We have a bunch of official and Europe,

00:59.000 --> 01:01.000
and thanks to open enabled,

01:01.000 --> 01:04.000
there are more than 5,000 clouds worldwide.

01:04.000 --> 01:06.000
When I talk about managing AI at scale,

01:06.000 --> 01:08.000
I talk about these scales.

01:08.000 --> 01:10.000
The Guardian, my speech,

01:10.000 --> 01:12.000
it will be mainly divided in six parts.

01:12.000 --> 01:14.000
A first part, where it's all about the product,

01:14.000 --> 01:16.000
open enabled, and the concept of AI factory,

01:16.000 --> 01:20.000
because this is related to the framework that we are currently working on.

01:20.000 --> 01:23.000
Then a second part about one AI,

01:23.000 --> 01:25.000
why we start building this framework?

01:25.000 --> 01:27.000
What problem we are trying to solve,

01:27.000 --> 01:29.000
and if you are able to solve these or not,

01:29.000 --> 01:31.000
then the architecture,

01:31.000 --> 01:35.000
I will give you an envelope review of the architecture of our framework,

01:35.000 --> 01:39.000
and the three pillars of the three pillars of our framework is based on

01:39.000 --> 01:41.000
that our marketplace, data,

01:41.000 --> 01:42.000
sorry, inference API.

01:42.000 --> 01:44.000
And then at the end, I will show you,

01:44.000 --> 01:48.000
from theory to fact, so I will show you the command

01:48.000 --> 01:50.000
and what are the expected, how to put.

01:50.000 --> 01:52.000
I will not be able to show you a demo.

01:52.000 --> 01:57.000
I also had a demo, but because my laptop is not working fine,

01:57.000 --> 02:00.000
but yes, it will be clear the same.

02:00.000 --> 02:04.000
So regarding OpenNable, OpenNable is an open cloud edge computing platform.

02:04.000 --> 02:06.000
We act as a virtual infrastructure manager,

02:06.000 --> 02:11.000
so this means that we enable you to run and orchestrate workflow

02:11.000 --> 02:15.000
that are related to your machine and so on.

02:15.000 --> 02:19.000
So it's simple, it's light, it does not require a Navy stock,

02:19.000 --> 02:23.000
it's sensible in the sense that you can import in your cloud

02:23.000 --> 02:26.000
different application and enhance its capabilities.

02:26.000 --> 02:29.000
It's multi-tenancy and multi-vehams out of the box,

02:29.000 --> 02:32.000
so you don't need to do anything regarding this.

02:32.000 --> 02:35.000
And if those of you know that are using Kubernetes,

02:35.000 --> 02:37.000
we also have an appliance in our marketplace.

02:37.000 --> 02:39.000
It's called one key, you can download it,

02:39.000 --> 02:44.000
and it will enable you to deploy openNable,

02:44.000 --> 02:47.000
sorry, Kubernetes cluster on top of openNable.

02:47.000 --> 02:52.000
Now, regarding AI factory and why this concept is important.

02:52.000 --> 02:57.000
Imagine that you have GPU, you have network,

02:57.000 --> 02:59.000
you have storage, you have CPU and so on.

02:59.000 --> 03:01.000
So you have the hardware, right?

03:01.000 --> 03:05.000
And you want to understand how to easily share these resources

03:05.000 --> 03:09.000
with all machine and so on such that you can create,

03:09.000 --> 03:12.000
isolated environment that in this slide are called tenant,

03:12.000 --> 03:16.000
where a person or a team not can do their own things.

03:16.000 --> 03:20.000
Developing AI models, running other service and so on,

03:20.000 --> 03:22.000
without interfering in each other.

03:22.000 --> 03:26.000
So we, in terms of openNable, we are in the middle there

03:26.000 --> 03:30.000
between your hardware and the visualization part on top.

03:30.000 --> 03:32.000
And we are helping you do in this.

03:32.000 --> 03:36.000
So one AI, the framework, it's related to enable you

03:36.000 --> 03:39.000
not to share your resources with different virtual machine

03:39.000 --> 03:41.000
in an efficient way.

03:41.000 --> 03:45.000
So regarding the problem, no, when why we start developing this framework?

03:45.000 --> 03:48.000
No, which problem we are trying to solve.

03:48.000 --> 03:51.000
First of all, we started noticing that there was the first problem

03:51.000 --> 03:53.000
related to model discovery.

03:53.000 --> 03:55.000
Okay, I want to train my model, I want to train my models,

03:55.000 --> 03:57.000
but where I put them, no.

03:57.000 --> 04:00.000
Is there like a sort of collection where I can find all my models,

04:00.000 --> 04:03.000
where I can easily know, put whatever I built there.

04:03.000 --> 04:06.000
Then the storage, okay, I find you in the model,

04:06.000 --> 04:09.000
I train the model, but this model are quite heavy, right?

04:09.000 --> 04:12.000
So we talk about more than one under GB sometimes.

04:12.000 --> 04:15.000
So I cannot download them every time, and then it's sensitive,

04:15.000 --> 04:17.000
and then exposing this model and so on.

04:17.000 --> 04:20.000
It's not efficient, right? A lot of bar.

04:20.000 --> 04:23.000
Regarding the deployment, the question is, okay, I have the GPU,

04:23.000 --> 04:29.000
but then I can leverage this GPU in order to reduce the inference

04:29.000 --> 04:33.000
time, the fine tuning, the training time, sorry, and so on.

04:33.000 --> 04:35.000
The first problem is this one.

04:35.000 --> 04:38.000
I have clients that are already developing.

04:38.000 --> 04:41.000
So I have a code base, for example, I deploy my model,

04:41.000 --> 04:44.000
and then what? Do I need to change my code base,

04:44.000 --> 04:49.000
such that it can communicate now with the new model deploy?

04:49.000 --> 04:53.000
That's why we start working on one AI.

04:53.000 --> 04:57.000
One AI is an open software to run AI models and inference on open nebula.

04:57.000 --> 05:00.000
And it's based mainly on three pillars.

05:00.000 --> 05:03.000
The argument face art marketplace that is a lightweight catalog,

05:03.000 --> 05:07.000
metadata first. What this means is that we already have

05:07.000 --> 05:10.000
this concept of collection of models on argument face rights.

05:10.000 --> 05:15.000
So what you can do, you can easily import this public or

05:15.000 --> 05:21.000
private collection of models from hugging face to your open nebula marketplace.

05:21.000 --> 05:24.000
But it's metadata first because you don't download the model.

05:24.000 --> 05:26.000
You only download the metadata.

05:26.000 --> 05:29.000
The model then will be materialized only when you want to deploy.

05:29.000 --> 05:33.000
It would just once and then you can share the images.

05:33.000 --> 05:37.000
Thanks to the shared file system data store with a lot of VMs.

05:37.000 --> 05:42.000
You download the model once and you reuse these model files a lot of times

05:42.000 --> 05:46.000
and you can share it with all the VMs that you want.

05:46.000 --> 05:50.000
Now that we have the model stored in our shared data store,

05:50.000 --> 05:53.000
it's time to expose it, right?

05:53.000 --> 05:58.000
And you can easily do that because in open nebula we have this apply on BLLM.

05:58.000 --> 06:02.000
That what it does in summary will mount the model into the VMs

06:02.000 --> 06:06.000
and then we'll expose it through an open AI API.

06:06.000 --> 06:10.000
So this means that you don't need to rewrite your clients and so on.

06:10.000 --> 06:14.000
Because if you already use open AI and tropic models on,

06:14.000 --> 06:18.000
you just need to change the URL, then the point and that's it.

06:18.000 --> 06:20.000
The level architecture is this.

06:20.000 --> 06:23.000
The user execute a command that is one AI run.

06:23.000 --> 06:26.000
Later I will show you some of the argument that you can pass to this command.

06:26.000 --> 06:31.000
What this command will do, it will generate a virtual machine template

06:31.000 --> 06:34.000
that is at the description of the virtual machine.

06:34.000 --> 06:38.000
Then the name of the virtual machine, the amount of memory,

06:38.000 --> 06:41.000
the number of CPU and so on, the network etc etc.

06:41.000 --> 06:45.000
Open nebula will read this virtual machine template

06:45.000 --> 06:48.000
and will instantiate the virtual machine.

06:48.000 --> 06:51.000
What I mean when I say instantiate the virtual machine.

06:51.000 --> 06:56.000
In this case, it will instantiate the virtual machine with the VLM image.

06:56.000 --> 07:01.000
Then it will attach all the disk that will contain the model files and so on.

07:01.000 --> 07:06.000
And it will also attach the PCI devices to the virtual machine.

07:06.000 --> 07:09.000
So the GPU and then of course, it will expose the service.

07:09.000 --> 07:12.000
With one command line at the end, you get this.

07:12.000 --> 07:14.000
Now in the picture below, you get an endpoint.

07:14.000 --> 07:17.000
And then you can communicate with your model.

07:17.000 --> 07:21.000
An important thing that is that no extra case run service is needed.

07:21.000 --> 07:23.000
You just run the CLI.

07:23.000 --> 07:28.000
It talks with open nebula and open nebula will schedule the VM to the right host.

07:28.000 --> 07:35.000
In this case, there are all the data, the GPU availability and also the GPU that you want.

07:35.000 --> 07:42.000
Now, regarding the first pillars before I was talking about this hugging phase marketplace.

07:42.000 --> 07:50.000
There is a picture there that show you what this marketplace in open nebula contains.

07:50.000 --> 07:57.000
This market will contain all the models fetched from your public or private collection in hugging phase.

07:57.000 --> 07:58.000
Just the metadata.

07:58.000 --> 08:01.000
So as you can see, these are only metadata.

08:01.000 --> 08:03.000
Very, very easy, just text.

08:03.000 --> 08:06.000
So it's very fetching is also quite rapid.

08:06.000 --> 08:11.000
What happens is that then in one eye, you pass the name of the model that you want to deploy.

08:11.000 --> 08:16.000
And what it will happen is that it will search for this model in the open nebula marketplace.

08:16.000 --> 08:21.000
It will import it into a shared file system data store.

08:21.000 --> 08:25.000
And if it's already present, of course, because if you are ready to load it,

08:25.000 --> 08:27.000
it will just reuse the image.

08:27.000 --> 08:30.000
Materials lies on deploy.

08:30.000 --> 08:33.000
Artifacts are copied only when you deploy.

08:33.000 --> 08:37.000
So you download once, the files, and you reuse multiple times.

08:37.000 --> 08:42.000
When you download the model, it is imported in a shared file system data store.

08:42.000 --> 08:51.000
That in summary, what is this data store that allows you to share what it contains with different virtual machines.

08:51.000 --> 08:53.000
So in the picture is very clear.

08:53.000 --> 08:59.000
You download the things one, download the images once, and then you can reuse multiple times.

08:59.000 --> 09:06.000
This is also cool because a lot of research lab, a lot of company already have their own high-performance storage,

09:06.000 --> 09:08.000
so you can also use them.

09:08.000 --> 09:13.000
The benefit, of course, is that you don't need to download the things every time and every time,

09:13.000 --> 09:17.000
because we are talking about very heavy model.

09:17.000 --> 09:25.000
And then, of course, it also have been reducing the storage cost, right?

09:25.000 --> 09:28.000
The third pillar is the inference, no?

09:28.000 --> 09:31.000
And the inference part, the lattice part, right?

09:31.000 --> 09:34.000
So one eye generates a template.

09:34.000 --> 09:37.000
The template is as a description of the virtual machine.

09:37.000 --> 09:39.000
How it should be insensitive.

09:39.000 --> 09:44.000
What will happen is that openable a create the virtual machine, assign the GPU, and start it.

09:44.000 --> 09:49.000
The inference engine, what we will do is just to add the model from the disk.

09:49.000 --> 09:55.000
Of course, you can also pass different configurations, so the maximum number of tokens in the output,

09:55.000 --> 10:02.000
the temperature of the model in order to manage its creativity and so on.

10:02.000 --> 10:07.000
Of course, these configs are passed as a context variable.

10:07.000 --> 10:11.000
And then, at the end, you get an end point that is reasonable.

10:11.000 --> 10:17.000
You know, if your codebase is already like talking with all the models that we know,

10:17.000 --> 10:21.000
because this is a standard way of communicating with the model.

10:21.000 --> 10:22.000
Okay, I talk a lot.

10:22.000 --> 10:26.000
This is just a practical example of a common that you can run.

10:26.000 --> 10:31.000
One eye run, you specify the model, the model will be imported if it's not present.

10:31.000 --> 10:36.000
Then you can specify the name of the ID or the inference engine.

10:36.000 --> 10:42.000
Right now, we only support the LLM, but of course, in the future, we are planning to support more inference engine.

10:42.000 --> 10:52.000
You also can pass the name of the network to associate to the VM, the number of CPU, the amount of memory, the number of GPU,

10:52.000 --> 10:55.000
the number of end data and the GPU type.

10:55.000 --> 11:00.000
What will happen here is that openable will do all the rest.

11:00.000 --> 11:08.000
So, in fact, you don't need any external customers, scheduler, whatever, because you pass the GPU count and the type.

11:08.000 --> 11:14.000
And the openable will automatically find the host that fulfill these requirements.

11:14.000 --> 11:22.000
Okay, regarding the GPU, that is also cool because before now, the presentation was about me,

11:22.000 --> 11:28.000
that we still don't support, but we are working on it, is that, of course, the VM can see the GPU right on the host.

11:28.000 --> 11:32.000
Because we pass it through the PCI pass through.

11:32.000 --> 11:36.000
So, the LLM can also, of course, leverage the GPU.

11:36.000 --> 11:44.000
Of course, if no GPU is present, it will use the CPU.

11:44.000 --> 11:49.000
Now, in theory here, I should show you like the demo.

11:49.000 --> 11:57.000
I don't know if maybe is there a way, but overall in the demo, what I was planning to show you was this.

11:57.000 --> 12:07.000
The demo was like a six minute demo divided in two parts, where I showed you how currently you can deploy a model together,

12:07.000 --> 12:14.000
not with the inference engine, expose a model using the common open API API in openable arena.

12:14.000 --> 12:16.000
You can already do it, right?

12:16.000 --> 12:18.000
But it's a manual process.

12:18.000 --> 12:28.000
You need to import the marketplace, you need to import the model from the marketplace into the shared data store,

12:28.000 --> 12:35.000
you need to go into the menu, check the virtual machine templates, all the list, find the inference engine,

12:35.000 --> 12:41.000
then you need to instantiate from the inference engine template, virtual machine templates,

12:41.000 --> 12:45.000
you need to instantiate the VM, so you need to assign a name, amount of memory, and so on, all these things,

12:45.000 --> 12:49.000
you need to attach the disk that contains the model files and so on.

12:49.000 --> 12:52.000
It can be done, but it's very manual, right?

12:52.000 --> 12:59.000
But thanks to the common line that we are currently working on, still in the element,

12:59.000 --> 13:05.000
I want to online this, you can do it almost like automatically.

13:05.000 --> 13:11.000
So in summary, one AI is a CLI that lets you discover a model from adding phase.

13:12.000 --> 13:16.000
It's important to say that in our interface, you can have also private collections, okay?

13:16.000 --> 13:20.000
So if you are a company that is developing and tuning training model, whatever,

13:20.000 --> 13:25.000
you can also have your private collection and import it into openable.

13:25.000 --> 13:32.000
You can store the models into a shared data store, so the load once and reduce multiple times,

13:32.000 --> 13:39.000
deploy an inference engine with that model and expose an open, an open-eye compatible API, okay?

13:40.000 --> 13:44.000
Just a list of the main benefits in common, same results.

13:44.000 --> 13:47.000
So this means that if you are into the MCP,

13:47.000 --> 13:52.000
wall, MCP2, MCP servers, and if you are into automation workflow and so,

13:52.000 --> 13:59.000
you can easily use this and put this like in your automation workflow, this common line.

13:59.000 --> 14:06.000
Your control, so this means your model, your infra openable is an open source cloud platform.

14:06.000 --> 14:11.000
So, not a keen, whatever, and all these sort of things.

14:11.000 --> 14:15.000
Reuse storage, one copy of an AI model serves many VMs.

14:15.000 --> 14:18.000
You reduce the storage cost.

14:18.000 --> 14:22.000
And then of course, as I said, in the first slide, openable will,

14:22.000 --> 14:25.000
out of the box will manage all the solution, the resource across the cluster, okay?

14:25.000 --> 14:28.000
So you don't need any other scale, Lorenzo.

14:29.000 --> 14:32.000
One AI, and the sheriff has been sent out, so our car rental in development,

14:32.000 --> 14:37.000
and we'll be available in a future openable release.

14:37.000 --> 14:46.000
So, well, can I, I, no, I know, but I have the debo, so maybe I can,

14:46.000 --> 14:51.000
we can just open for the, because the demo, I can not show it because it's not working.

14:51.000 --> 14:54.000
Yeah, yeah, maybe open for questions, such that.

14:54.000 --> 14:55.000
Thank you very much.

14:56.000 --> 14:57.000
Thank you.

14:59.000 --> 15:00.000
Oh, maybe.

15:01.000 --> 15:03.000
Hello, thank you for our talk.

15:03.000 --> 15:05.000
My question is about the shell effect.

15:05.000 --> 15:10.000
It's based on, um, it's based on some kind of, uh, me, you offset.

15:10.000 --> 15:11.000
Yes.

15:11.000 --> 15:13.000
It's based on virtual effects, yes.

15:13.000 --> 15:19.000
So what happens is that you have, uh, all the file, all the model of the files in your share

15:19.000 --> 15:25.000
file system data store, what happens is that when the virtual machine that contains the

15:25.000 --> 15:29.000
inference engine will mount the disk through virtual effects.

15:29.000 --> 15:35.000
So you're going to get, like, the VM can, no, can get visibility on what is story in the

15:35.000 --> 15:39.000
share, if I, if I, if I, if I, if I, if I, if I, if I, if I, if I, if I, if I answer your

15:39.000 --> 15:41.000
question, yeah, okay.

15:45.000 --> 15:47.000
So I actually also have a question.

15:47.000 --> 15:51.000
So I know that a lot of this type of clusters are kind of legacy and they don't really have

15:51.000 --> 15:54.000
GPU, so they have very underpowered GPUs.

15:54.000 --> 15:59.000
Uh, do you have any way of deploying inference engines other than VLLM?

15:59.000 --> 16:03.000
You know, that would actually be more CPU friendly, like LMSEPP or anything like that.

16:03.000 --> 16:04.000
Very good question.

16:04.000 --> 16:05.000
Thank you very much.

16:05.000 --> 16:08.000
Right now, no, in the sense that we only have the LLM.

16:08.000 --> 16:12.000
Of course, in the future, we are planning to, to implement some other inference

16:12.000 --> 16:18.000
engines that are more CPU prone, but right now, no, but of course, this is like an

16:18.000 --> 16:19.000
issue, right?

16:19.000 --> 16:22.000
Because as you say, not all the people have, not all the component, all the people,

16:22.000 --> 16:27.000
well, they don't have the GPU, so we should enable them also to run models,

16:27.000 --> 16:31.000
maybe not, you know, the top models, but something, you know, that can still

16:31.000 --> 16:32.000
bring value.

16:32.000 --> 16:37.000
Anybody else?

16:37.000 --> 16:41.000
Uh, I just wanted to ask where the GPU is located physically?

16:41.000 --> 16:42.000
Where?

16:42.000 --> 16:46.000
For OpenR, inabula, where's the GPU is located physically?

16:46.000 --> 16:52.000
Well, the GPU is located on the other, in the sense that OpenNembola is a cloud platform

16:52.000 --> 16:56.000
that you can install, no, in the meta, also, we support almost fully providers.

16:56.000 --> 17:01.000
And it will enable you to, it's sensitive, you have to machine orchestrate, we're from

17:01.000 --> 17:02.000
machine and so on.

17:02.000 --> 17:05.000
So, the GPU is like, it's not in OpenNembola servers.

17:05.000 --> 17:11.000
It will be on the else, customers are lost, they install OpenNembola on top, and what

17:11.000 --> 17:16.000
we do in terms of OpenNembola and be sure that they've used all the machines that are

17:16.000 --> 17:19.000
running on OpenNembola, okay?

17:19.000 --> 17:24.000
Can access to the GPU in order not to reduce the inference, inference, time,

17:24.000 --> 17:26.000
functioning, I mean, so.

17:26.000 --> 17:31.000
Okay, this is all I have got.

17:31.000 --> 17:34.000
Thank you.

17:34.000 --> 17:36.000
Thank you.

17:36.000 --> 17:39.000
Come next here, come next here.

