WEBVTT

00:00.000 --> 00:10.000
Yes, I built a prototype for it.

00:10.000 --> 00:14.000
Thank you to the question.

00:14.000 --> 00:20.000
As everybody knows Cloud is the rage in business

00:20.000 --> 00:25.000
and enterprises, everything is moving into the cloud

00:25.000 --> 00:29.000
and it gets even more with this whole AI stuff.

00:29.000 --> 00:37.000
And yeah, in the cloud, you typically have your two kinds of system software stacks.

00:37.000 --> 00:47.000
I can have VMs and you can have containers.

00:47.000 --> 00:54.000
And then if I'll take the closer look into what the trusted computing base would be for these two

00:54.000 --> 00:56.000
versions of solutions.

00:56.000 --> 01:00.000
Of course, for the virtualization, you have the hypervisor,

01:00.000 --> 01:08.000
as trusted computing base, and the individual gas OSs, as well as usually the kind of VM manager.

01:08.000 --> 01:14.000
Any one of these can be compromised in any way you can get information leaks,

01:14.000 --> 01:16.000
or in the case of the VM manager,

01:16.000 --> 01:22.000
and the attacker could start to control the VMs on this host system,

01:22.000 --> 01:24.000
which you typically don't want.

01:24.000 --> 01:29.000
And if the hypervisor's compromised, you could breach the isolation of the VMs.

01:29.000 --> 01:35.000
But even the gas OSs in the VM can be vulnerable to attacks, because if you leave,

01:35.000 --> 01:40.000
if this is compromised, it can't attack as of the VMs.

01:40.000 --> 01:48.000
Easily, it's still isolated, but it can leak the confidential information that we cite in the compromised VM.

01:48.000 --> 01:54.000
And for containers, we have a similar picture just that we have one isolation layer.

01:54.000 --> 01:59.000
Less, we have the host OS, which of course has to be trusted,

01:59.000 --> 02:04.000
and usually your container manager, because even if your host OS is secure,

02:04.000 --> 02:08.000
if it's for some reason, it's a back in the container manager that allows

02:08.000 --> 02:12.000
an attacker to remotely control the containers on your host system,

02:12.000 --> 02:15.000
you're also out of luck.

02:15.000 --> 02:20.000
So let's have a look at how the container solutions look,

02:20.000 --> 02:26.000
what are these solutions usually, and we can see about 80% is Docker.

02:26.000 --> 02:30.000
Probably not such a big surprise.

02:30.000 --> 02:40.000
And for the host OS, on Docker, we see 75% using Linux,

02:41.000 --> 02:58.000
making about 71% of all container-based cloud infrastructure being based on the Linux system.

02:58.000 --> 03:04.000
And on Linux, you usually initially didn't have any container support.

03:04.000 --> 03:12.000
So newing the evolution of Linux, there are several mechanisms added to your usual

03:12.000 --> 03:22.000
set of the usual traditional textbook unit system, which are C groups for restricting the usage of resources,

03:22.000 --> 03:29.000
for example, to restrict the usage of certain CPU costs to specific containers.

03:29.000 --> 03:35.000
Furthermore, you have namespaces, which reduce the visibility of the host fire system.

03:35.000 --> 03:38.000
You can easily completely hide the host fire system,

03:38.000 --> 03:44.000
but you can also, for example, usually if you're using Docker just for development,

03:44.000 --> 03:50.000
you may have a shared folder for your individual project that you are working on.

03:51.000 --> 03:58.000
And then, last but not least, we have second EBPF, which is used to restrict the source costs,

03:58.000 --> 04:05.000
a container might call, because if you allow everything a source call to be performed outside of the container,

04:05.000 --> 04:14.000
especially when it is allowed to do anything on the host system, you can easily break the isolation again.

04:14.000 --> 04:25.000
And then, yeah, the guys from Kern concept and backhausen Institute had a joint paper called Met Eagle,

04:25.000 --> 04:32.000
where they analyzed the security of these container solutions.

04:32.000 --> 04:38.000
And first, they stated that about 2.7 million lines of code in Linux,

04:38.000 --> 04:45.000
like 7.4 are for the container infrastructure alone.

04:45.000 --> 04:51.000
And there they found eight critical vulnerabilities for the second mechanism,

04:51.000 --> 04:55.000
three for three groups and 22 for namespaces.

04:55.000 --> 05:06.000
Critical meanings that these are really exploitable vulnerabilities that can lead to a breach of the isolation of the containers.

05:06.000 --> 05:16.000
And those, we have seen 33 critical TVs.

05:16.000 --> 05:24.000
And yeah, one month thing now here, okay, containers, we know they are lightweight, but they are not sectors that use the ends.

05:24.000 --> 05:32.000
But yeah, there we have first 50% being KVM, that's basically integrated part of the Linux kernel,

05:32.000 --> 05:37.000
so in some capacity it's still a Linux system involved.

05:37.000 --> 05:47.000
And we have 150,000 lines of code for KVM alone that's just the kernel part, not the virtual machine monitor,

05:47.000 --> 05:59.000
which is usually a chemo system, chemo in user space that provides the actual emulation for the machine interface that the guest OS expects.

05:59.000 --> 06:09.000
And there we see 24 critical vulnerabilities that have appeared since 2008 with a CBSS of more than 7.0,

06:09.000 --> 06:19.000
according to the NIST vulnerability database, and for QM we have 77 critical vulnerabilities.

06:19.000 --> 06:26.000
And if we move up the stack, we have, according to command linux.com,

06:26.000 --> 06:34.000
we are over 90% of guest OS's in those cloud setups are linux OS's,

06:34.000 --> 06:43.000
and then we have of course all the problems that the Linux kernel has in regards to security and vulnerabilities.

06:43.000 --> 06:50.000
Being that we have a huge trusted computing base of 14 million lines of code,

06:50.000 --> 06:57.000
of course not every kind of code is really actively used there, but it's a lot that will be used there,

06:57.000 --> 07:02.000
because then we also have diverse drivers and other stuff.

07:02.000 --> 07:06.000
And there we had more than 1,000 critical vulnerabilities,

07:06.000 --> 07:13.000
I looked it up again in the NIST and national vulnerability database,

07:13.000 --> 07:20.000
of course they are not all open yet. Now they've closed these issues,

07:20.000 --> 07:31.000
maybe not all, but most of them, but it shows that we have a very complex and huge trusted computing base,

07:31.000 --> 07:37.000
and due to the fact that it was originally was a unique system where you had a global namespace

07:37.000 --> 07:44.000
with everything being visible and just restricted by file system permissions in order

07:44.000 --> 07:50.000
to have containers, additional restriction techniques were introduced adding even more complexity

07:50.000 --> 07:56.000
of an already complex system.

07:56.000 --> 08:07.000
And these are prone to vulnerability that allows to breach the isolation.

08:07.000 --> 08:11.000
So I might ask, we are, isn't there any better solution yet?

08:11.000 --> 08:17.000
And I can recall Hermann Hattick, which I think most of you know,

08:17.000 --> 08:26.000
are always stated when we had at the German SIG group meeting a talk regarding cloud and unique kernels

08:26.000 --> 08:32.000
and hypervisor that yeah, we did this already 30 years ago,

08:32.000 --> 08:41.000
and our solution of course was, yes, micro kernels.

08:41.000 --> 08:46.000
So let's see how a micro kernel cloud architecture may look like.

08:46.000 --> 08:53.000
So I'd like to introduce Erlan OS, which is my take on building a cloud infrastructure based on the

08:53.000 --> 08:59.000
genet operating system framework.

08:59.000 --> 09:06.000
And first, we have a set of components which comprise our system.

09:06.000 --> 09:12.000
Each component is isolated, it's similar to an isolation to a container,

09:12.000 --> 09:17.000
maybe even close to a virtual machine, depending on the isolation level,

09:17.000 --> 09:19.000
the micro kernel may provide.

09:19.000 --> 09:25.000
In genet, you have at least four different micro kernels you can choose.

09:25.000 --> 09:33.000
And yeah, each of these components only interact with each other via RPCs by default.

09:33.000 --> 09:43.000
And the good part is each RPC can only be used if a capability was granted to the

09:43.000 --> 09:52.000
component that wants to invoke this RPC.

09:52.000 --> 10:02.000
And yeah, the access to these capabilities in a genet system usually controlled by the parent components.

10:02.000 --> 10:07.000
And again, for these parent components, it would be this component down here,

10:07.000 --> 10:11.000
which is usually the initial components that are started in a system.

10:11.000 --> 10:15.000
And those in no-minclature is not the one.

10:15.000 --> 10:22.000
The genet usually has because in my system, I've modified the unit component of genet with

10:22.000 --> 10:26.000
some additional features that I used for automatic resource partitioning.

10:26.000 --> 10:30.000
So don't get confused by the names here.

10:30.000 --> 10:38.000
And then, yeah, if we have, for example, this app 2 that wants now to get the,

10:38.000 --> 10:42.000
oh, there's some small mistake on the slide.

10:42.000 --> 10:47.000
Yeah, let's consider that here we would have a nick instead of tasking.

10:47.000 --> 10:52.000
And app 2 wanted to now get the nick capability.

10:52.000 --> 10:58.000
It would ask its parent, the screen component for this nick capability.

10:58.000 --> 11:04.000
And if in the configuration of this green component,

11:04.000 --> 11:09.000
the rule that allows app 2 to have that nick connection,

11:09.000 --> 11:20.000
then app 2 gets the capability and then can now use nick that is not working.

11:21.000 --> 11:29.000
And then you can have built trees in this kind with a parent component,

11:29.000 --> 11:33.000
child components that may also have child components.

11:33.000 --> 11:40.000
And in my system, I call them habitats because the components are originally called cells.

11:40.000 --> 11:45.000
And they have habitat of resources where they will start in.

11:45.000 --> 11:51.000
And these can also isolate different tenants or user sessions.

11:51.000 --> 11:56.000
Or maybe if you want to compose more complex container or VM scenario,

11:56.000 --> 12:00.000
we also have special libraries and not just one app.

12:00.000 --> 12:06.000
Maybe additional apps you could isolate and building such a tree.

12:06.000 --> 12:13.000
And this structure also guarantees that this components here, for example,

12:13.000 --> 12:19.000
can't see the yellow or blue component or communicate with them,

12:19.000 --> 12:25.000
then unless this component allows that,

12:25.000 --> 12:31.000
and in order to for this component to allow contact with this blue and yellow component,

12:31.000 --> 12:36.000
the gray component down here would have to allow it first.

12:36.000 --> 12:40.000
That's basically a capability delegation,

12:40.000 --> 12:48.000
or capabilities to decide here and then they are delegated down for subsets of resources.

12:48.000 --> 12:55.000
So the advantages of such an architecture is that we do not have an global namespace,

12:55.000 --> 13:01.000
so we don't need any restriction mechanisms that can possibly be breached.

13:01.000 --> 13:04.000
We are isolated by default.

13:04.000 --> 13:10.000
Then we have explicit access control to these capabilities.

13:10.000 --> 13:18.000
And in case for capabilities, it's a proven security mechanism for decades now.

13:18.000 --> 13:22.000
Now, and of course, since we're talking about microcolors,

13:22.000 --> 13:26.000
we're talking about the kernel size of about a few thousands,

13:26.000 --> 13:33.000
tens thousands of lines of code compared to the millions of lines of code for Linux.

13:33.000 --> 13:37.000
So we have a very small across the computing base.

13:37.000 --> 13:42.000
And in addition, our feature we get on microcolors,

13:42.000 --> 13:49.000
since we execute all drivers and system services in user space,

13:49.000 --> 13:53.000
based we have the advantage, if for example,

13:53.000 --> 13:58.000
this blue one would fail or the FS server here would fail.

13:58.000 --> 14:03.000
This could be recognized by the great component here,

14:03.000 --> 14:08.000
or the blue component itself, and if F server could be restarted.

14:08.000 --> 14:13.000
So you can have full content and a thought in a device

14:13.000 --> 14:18.000
to have our system server does not crash the kernel and the whole operating system,

14:18.000 --> 14:23.000
like for example, we saw a few years ago with some Windows update,

14:24.000 --> 14:30.000
ironically, a security feature that will update a security feature,

14:30.000 --> 14:36.000
crashed the Windows kernel on millions of machines and caused

14:36.000 --> 14:38.000
quite some economic damage.

14:38.000 --> 14:40.000
That's something we wouldn't have here.

14:40.000 --> 14:45.000
And of course, even the fact that we have this all in user space,

14:45.000 --> 14:50.000
we can have different variations of system services and drivers,

14:50.000 --> 14:53.000
so that we also gain more modularities,

14:53.000 --> 15:00.000
and you usually have with your monolithic kernel.

15:00.000 --> 15:04.000
And as the guys from cell 4, of course,

15:04.000 --> 15:08.000
we'll know you can certify or even formally prove

15:08.000 --> 15:13.000
the correctness of your microcolors, which you can't do with limits.

15:13.000 --> 15:18.000
It's a way to complex to do that.

15:19.000 --> 15:23.000
So one might ask, yeah, if these microcolors are so good,

15:23.000 --> 15:27.000
and we have them for 30 years by now.

15:27.000 --> 15:33.000
Why does everybody still use Linux?

15:33.000 --> 15:38.000
And I want to answer this question for my own experience and perspective,

15:38.000 --> 15:42.000
and I think it boils down to two major challenges,

15:42.000 --> 15:47.000
being compatibility and scalability performance.

15:47.000 --> 15:51.000
And that's how I put together because, usually,

15:51.000 --> 15:53.000
especially regarding microcolors,

15:53.000 --> 15:57.000
I think you can't really separate those two.

15:57.000 --> 16:02.000
And for compatibility, it stops with hardware support,

16:02.000 --> 16:04.000
who can't use a microchor on a cell out,

16:04.000 --> 16:07.000
wherever the hardware isn't supported.

16:07.000 --> 16:11.000
And there, for example, we need numerous support.

16:11.000 --> 16:15.000
This was one feature that Jeanette originally didn't have,

16:15.000 --> 16:20.000
I edited my variation of Jeanette in the Nova microcolors.

16:20.000 --> 16:23.000
This is especially needed for memory,

16:23.000 --> 16:28.000
bound and memory intensive applications like databases,

16:28.000 --> 16:32.000
because if you do not do this, you have to pay cash coherence,

16:32.000 --> 16:35.000
traffic, and the Numa costs that you have,

16:35.000 --> 16:39.000
then you communicate between different Numa regions,

16:39.000 --> 16:44.000
and knowing the Numa topology can help you avoid that.

16:45.000 --> 16:49.000
And of course, speaking of servers, we need drivers for nicks,

16:49.000 --> 16:51.000
and I'm talking to server nicks,

16:51.000 --> 16:57.000
and not your small Wi-Fi chip or onboard network cut in your laptop.

16:57.000 --> 17:02.000
And of course, also rate controllers, hard disk, and stuff,

17:02.000 --> 17:07.000
and increasingly more, especially since that I create,

17:07.000 --> 17:10.000
we also need drivers for accelerators,

17:10.000 --> 17:15.000
especially GPUs, NPUs, and so on.

17:15.000 --> 17:19.000
And regarding security,

17:19.000 --> 17:22.000
Windows already forces it,

17:22.000 --> 17:25.000
TPM support, or SQX,

17:25.000 --> 17:32.000
on-claves are a thing to further improve the security of your system,

17:32.000 --> 17:37.000
and you need to support for that too in your microcolors.

17:37.000 --> 17:42.000
One solution that a unit laps the tighter to go to the softest problem,

17:42.000 --> 17:47.000
because writing all drivers from scratch for hardware,

17:47.000 --> 17:51.000
may not even be documented at all by the hardware lender,

17:51.000 --> 17:54.000
is a very tedious task,

17:54.000 --> 18:01.000
and may even prove impossible for very big hardware like these GPUs.

18:01.000 --> 18:06.000
I know that I had a group of students that really dared to take on the task

18:06.000 --> 18:09.000
to write a driver for AMD GPUs,

18:09.000 --> 18:13.000
or by the test from scratch with what was documented.

18:13.000 --> 18:16.000
And after one year, we had to say,

18:16.000 --> 18:20.000
okay, that doesn't work out.

18:20.000 --> 18:23.000
So the idea here would be using Linux to have us

18:23.000 --> 18:27.000
inside our components in user space.

18:27.000 --> 18:31.000
And that's exactly how Gina does it with the device driver environment,

18:31.000 --> 18:35.000
which is currently based on Linux 6.12.

18:35.000 --> 18:38.000
And what it does is it provides a compatibility layer,

18:38.000 --> 18:42.000
which implements some parts of the Linux kernel APIs

18:42.000 --> 18:44.000
that is used by drivers, for example,

18:44.000 --> 18:48.000
for PCI, discovery, memory mapping,

18:48.000 --> 18:51.000
and memory management of your devices.

18:56.000 --> 19:00.000
But yeah, even if your microcolors runs on the server,

19:00.000 --> 19:03.000
and supports all your hardware,

19:03.000 --> 19:06.000
this doesn't help you if you don't have this software

19:06.000 --> 19:09.000
that you need for cloud operations.

19:09.000 --> 19:13.000
And yeah, it boils down that most software that is used in cloud,

19:13.000 --> 19:16.000
especially if it is something to do with web services,

19:16.000 --> 19:19.000
we have some kind of public status applications,

19:19.000 --> 19:28.000
or maybe even some runtime systems like JVM, Python, Ruby, and stuff.

19:28.000 --> 19:33.000
And this, you have to park to your microcolour operating system.

19:33.000 --> 19:37.000
But this is not all, you also need to support,

19:37.000 --> 19:42.000
maybe for parking, at least to ease the use of your tool channel

19:42.000 --> 19:44.000
that you have for your microcolour.

19:44.000 --> 19:48.000
And yeah, the servers are usually not controlled

19:48.000 --> 19:51.000
like a sculptor as directly on your device.

19:51.000 --> 19:56.000
No administrator really sits at the server, connect,

19:56.000 --> 20:01.000
keyboard to it in a monitor in these big server rooms,

20:01.000 --> 20:04.000
but they are controlled remotely,

20:04.000 --> 20:08.000
so you move some kind of tool support for remote management,

20:08.000 --> 20:14.000
something like a Docker console, SSH, or something.

20:14.000 --> 20:17.000
And this have to be part of two,

20:17.000 --> 20:23.000
but fortunately, Jeanette helps a bit with their Goa toolkit

20:24.000 --> 20:27.000
for porting to a party system,

20:27.000 --> 20:31.000
party software, and it's also supported common build systems

20:31.000 --> 20:36.000
and provides such nice features like automated debugging,

20:36.000 --> 20:43.000
support testing and publishing all this in these nice packages.

20:43.000 --> 20:50.000
And my system provides also a mocha where you can connect

20:51.000 --> 20:54.000
while there's a network to your machines

20:54.000 --> 20:58.000
that runs LNOS and manage the components at one time.

20:58.000 --> 21:01.000
You can start and new components,

21:01.000 --> 21:05.000
or destroy running components,

21:05.000 --> 21:09.000
and also monitor the resource utilization.

21:09.000 --> 21:13.000
But I have to, it's a prototype for academia.

21:13.000 --> 21:15.000
It's very bad once at the moment,

21:15.000 --> 21:20.000
I wouldn't say that it's already for production build set.

21:20.000 --> 21:24.000
But that's from the future.

21:24.000 --> 21:29.000
Yeah, then that was for the compatibility part.

21:29.000 --> 21:33.000
Now the other big challenge is scalability and performance.

21:33.000 --> 21:40.000
For example, here I measured the performance of the Jeanette heap,

21:40.000 --> 21:44.000
data structure, which is a common used by all user space applications,

21:45.000 --> 21:48.000
that do dynamic memory allocations.

21:48.000 --> 21:53.000
And I performed the last memory allocation benchmark from Lassen,

21:53.000 --> 21:59.000
which is a benchmark that performs memory allocations

21:59.000 --> 22:02.000
of eight to one kilobit chunks randomly,

22:02.000 --> 22:08.000
which should represent typical web server like memory allocations

22:08.000 --> 22:10.000
that you have.

22:10.000 --> 22:15.000
And yeah, in this case, I performed five million Mallock and three chords.

22:15.000 --> 22:19.000
Each in this benchmark, and one can see,

22:19.000 --> 22:24.000
if you run up the number of threats of your application,

22:24.000 --> 22:26.000
that uses a heap,

22:26.000 --> 22:31.000
you can see that the time it takes for allocation,

22:31.000 --> 22:34.000
an integral block goes up,

22:34.000 --> 22:36.000
which is not so good,

22:36.000 --> 22:41.000
because even if you have five million allocations,

22:41.000 --> 22:45.000
in total and would, yeah,

22:45.000 --> 22:49.000
these two built it over 31 CPU cores,

22:49.000 --> 22:54.000
you would expect it to be lower than with one CPU core.

22:54.000 --> 22:57.000
But unfortunately, that's not the case.

22:57.000 --> 23:05.000
And yeah, this boils down to some synchronization issues.

23:06.000 --> 23:08.000
Choose in the DNA system.

23:12.000 --> 23:16.000
But that's not the only bottleneck I discovered

23:16.000 --> 23:21.000
by creating my own version of Genet for Cloud.

23:21.000 --> 23:24.000
The other big problem is networking.

23:24.000 --> 23:28.000
There I discovered that the default socket implementation

23:28.000 --> 23:31.000
that Genet has doesn't scale well.

23:31.000 --> 23:35.000
It performed very, very simple benchmark

23:35.000 --> 23:38.000
that just measures the performance of the network stack

23:38.000 --> 23:40.000
without any, yeah,

23:40.000 --> 23:44.000
no noticeable impact from the application.

23:44.000 --> 23:48.000
It was just simple, very simple TCP echo benchmark,

23:48.000 --> 23:52.000
basically receiving a packet bar TCP

23:52.000 --> 23:57.000
and just sending it directly back to the plot.

23:58.000 --> 24:01.000
And compared against a state of the art,

24:01.000 --> 24:05.000
the PDK based network stack that was developed

24:05.000 --> 24:09.000
especially for these kinds of cloud operations,

24:09.000 --> 24:11.000
called Caladan.

24:11.000 --> 24:14.000
We can see, yeah, the tail latency,

24:14.000 --> 24:18.000
that is in this case in 1999.9s,

24:18.000 --> 24:23.000
percent tail tile of all latencies in the benchmark.

24:23.000 --> 24:28.000
We see, yeah, Caladan, you can't even see anything there

24:28.000 --> 24:33.000
because the Genet scalability is bottlenecked so much

24:33.000 --> 24:37.000
that we have orders of magnitude here.

24:37.000 --> 24:41.000
Here, if you have, yeah,

24:41.000 --> 24:44.000
named reasons and number of connections and 32 connections

24:44.000 --> 24:47.000
are not that much for solidifications.

24:47.000 --> 24:55.000
So what's the reason for that?

24:55.000 --> 25:00.000
I try to explain it as far as I've found out yet.

25:00.000 --> 25:07.000
One possible reason is that from the architectural perspective

25:07.000 --> 25:11.000
that in Genet and also other microcolonals

25:11.000 --> 25:14.000
that follow the original F4 like design,

25:14.000 --> 25:19.000
you have a server which has some entry point here,

25:19.000 --> 25:21.000
which is completely single threaded,

25:21.000 --> 25:25.000
and you have your client which performs an RPC

25:25.000 --> 25:27.000
for an RPC object,

25:27.000 --> 25:29.000
it has the capabilities for,

25:29.000 --> 25:33.000
but it does, it's called a system call for this RPC,

25:33.000 --> 25:36.000
which goes down into the kernel here,

25:36.000 --> 25:41.000
and then this kernel triggers this entry point

25:41.000 --> 25:43.000
and then performs this function.

25:43.000 --> 25:47.000
But this entry point has to be synchronized

25:47.000 --> 25:52.000
because if you allow your preferred access

25:52.000 --> 25:54.000
or an RPC object,

25:54.000 --> 25:57.000
you would have to think of it or you get

25:57.000 --> 26:00.000
inconsistent data, especially if you have

26:00.000 --> 26:03.000
what operations on this RPC object.

26:03.000 --> 26:07.000
But now since we only have one entry point,

26:07.000 --> 26:11.000
if you have 32 clients now that want to

26:11.000 --> 26:14.000
concurrently access this entry point,

26:14.000 --> 26:21.000
you see this will cause a traffic jam in this regard.

26:21.000 --> 26:23.000
And what makes it worse,

26:23.000 --> 26:27.000
at least I can say from the perspective of Genet,

26:27.000 --> 26:30.000
it uses test and set locks

26:30.000 --> 26:33.000
for locking internal data structures

26:33.000 --> 26:38.000
that are used to lock the calling thread here

26:38.000 --> 26:43.000
and for some resource management of the capabilities.

26:43.000 --> 26:49.000
And these are known in research for not scaling well

26:49.000 --> 26:53.000
when you have multi-core systems with many CPU cores,

26:53.000 --> 27:00.000
which way above it we talk more about hundreds of CPU cores here.

27:00.000 --> 27:04.000
And yeah, a quick solution would be to replace

27:04.000 --> 27:08.000
this non-scalable locks with more scalable locks like

27:08.000 --> 27:10.000
S locks or anything,

27:10.000 --> 27:15.000
and or use weight-free synchronization as possible.

27:15.000 --> 27:19.000
And an architectural attempt to first this problem

27:19.000 --> 27:24.000
was proposed by the factored operating system,

27:24.000 --> 27:29.000
where the idea is to not have a service with one

27:29.000 --> 27:30.000
or three point.

27:30.000 --> 27:35.000
But you have a fleet of service threads in your system.

27:35.000 --> 27:40.000
Some, for example, this name service has several fleets.

27:40.000 --> 27:44.000
And then all your requests to the service are distributed

27:44.000 --> 27:46.000
over the fleet members,

27:46.000 --> 27:50.000
which would be one architectural solution

27:50.000 --> 27:53.000
to solve this bottleneck problem.

27:53.000 --> 27:56.000
Yeah, and I have to fill up.

27:57.000 --> 27:59.000
Yeah, the future won't matter.

27:59.000 --> 28:02.000
I would envision would be, of course,

28:02.000 --> 28:04.000
plotting some server applications,

28:04.000 --> 28:08.000
advancing the tooling for my remote management system.

28:08.000 --> 28:13.000
And one big chunk, the most important challenge for me

28:13.000 --> 28:16.000
would be the redesign of genots network interface.

28:16.000 --> 28:19.000
For example, according to such the idea,

28:19.000 --> 28:21.000
like having fleets of genots,

28:21.000 --> 28:24.000
a big router component, which is usually

28:24.000 --> 28:26.000
a component handling or networking energy

28:26.000 --> 28:27.000
in that system.

28:27.000 --> 28:30.000
And on top of that, maybe the evaluation

28:30.000 --> 28:33.000
and reasoned down of genots page locator,

28:33.000 --> 28:36.000
I know from current concept, with their fiasco kernels,

28:36.000 --> 28:40.000
they had a scalability issue with their original memory locator

28:40.000 --> 28:44.000
and had a collaboration with the university of Hanover,

28:44.000 --> 28:47.000
where they integrated a more scalable page locator,

28:47.000 --> 28:49.000
which improves the performance of their system.

28:49.000 --> 28:54.000
So maybe we can do that with genotool.

28:54.000 --> 28:57.000
Yeah, and to finish my talk, yeah,

28:57.000 --> 29:00.000
we have seen current cloud infrastructure

29:00.000 --> 29:02.000
primarily based on Linux,

29:02.000 --> 29:05.000
which has a huge and complex trusted computing base,

29:05.000 --> 29:09.000
which is full with critical vulnerabilities,

29:09.000 --> 29:12.000
not all the time, but you can expect them.

29:12.000 --> 29:16.000
Yeah, and you know micro kernels of a greater

29:16.000 --> 29:19.000
security, safety, and modularity,

29:19.000 --> 29:23.000
but they see a little adoption,

29:23.000 --> 29:27.000
at least, until now, for cloud operations,

29:27.000 --> 29:31.000
this may be down to the effect of the lack of hardware

29:31.000 --> 29:33.000
that was for server-grade hardware,

29:33.000 --> 29:36.000
and a lack of tool and software support,

29:36.000 --> 29:39.000
especially when it comes to server applications

29:39.000 --> 29:43.000
that use legacy project style interfaces.

29:44.000 --> 29:46.000
And we also, on top,

29:46.000 --> 29:50.000
we have issues with the scalability of system services.

29:50.000 --> 29:55.000
But there's ongoing work to overcome those challenges.

29:55.000 --> 30:00.000
Genotlapse is more focused on providing better compatibility

30:00.000 --> 30:02.000
with improving their go-a tool chain

30:02.000 --> 30:06.000
and putting even more and more applications of libraries to genot,

30:06.000 --> 30:09.000
and we also have gaphood,

30:09.000 --> 30:12.000
which are more focused on improving the networking support,

30:13.000 --> 30:15.000
and the genotl systems they use,

30:15.000 --> 30:19.000
and can concept and back housing work with FRE,

30:19.000 --> 30:22.000
more from what I know,

30:22.000 --> 30:27.000
more focused on improving the architecture of micro kernels

30:27.000 --> 30:33.000
and developing new concepts to make micro kernels more scalable.

30:33.000 --> 30:35.000
So thank you for your attention,

30:35.000 --> 30:38.000
and I hope that you'll get in touch with me.

30:39.000 --> 30:40.000
Thank you.

30:51.000 --> 30:52.000
Thank you.

30:52.000 --> 30:55.000
As the IP statement, you know,

30:55.000 --> 30:58.000
I appreciate any help.

30:58.000 --> 31:00.000
Any questions?

31:00.000 --> 31:03.000
Would it make sense to,

31:03.000 --> 31:06.000
especially given the scalability of the network,

31:06.000 --> 31:10.000
and we want to take a little more close to abandon the cloud

31:10.000 --> 31:13.000
of our synchronous IPC,

31:13.000 --> 31:17.000
and finally move to asynchronous IPC there.

31:17.000 --> 31:20.000
I got the question,

31:20.000 --> 31:23.000
whether it makes sense to,

31:23.000 --> 31:27.000
we place synchronous IPC,

31:27.000 --> 31:32.000
especially if we have that scalability issue with asynchronous IPC,

31:32.000 --> 31:35.000
and, yeah, I have to say,

31:35.000 --> 31:37.000
I am also,

31:37.000 --> 31:39.000
of that opinion,

31:39.000 --> 31:44.000
I think asynchronous IPCs would be something we could,

31:44.000 --> 31:47.000
consider in the future, but of course,

31:47.000 --> 31:51.000
yeah, all the systems you all form micro kernels,

31:51.000 --> 31:54.000
are originally designed for synchronous IPCs,

31:54.000 --> 31:57.000
which means that we have to change

31:57.000 --> 32:01.000
a very fundamental part of the operating systems kernel,

32:01.000 --> 32:04.000
and do lots of redesigned stuff,

32:04.000 --> 32:07.000
because all your applications have to consider

32:07.000 --> 32:10.000
that course asynchronous now,

32:10.000 --> 32:14.000
and that's interesting to think about that,

32:14.000 --> 32:18.000
but I don't think that it will be in the near future,

32:18.000 --> 32:19.000
it's more important.

32:19.000 --> 32:21.000
There are not just alpha micro kernels.

32:21.000 --> 32:24.000
I know there are other micro kernels,

32:24.000 --> 32:28.000
I could also look for example into other micro kernels,

32:28.000 --> 32:31.000
and see whether there are micro-core candidates

32:31.000 --> 32:34.000
that could work with Jeanette for my system here.

32:34.000 --> 32:36.000
Now, as an example,

32:36.000 --> 32:38.000
sample and in the general perspective,

32:38.000 --> 32:41.000
yeah, if you already have an asynchronous

32:41.000 --> 32:43.000
IPC-based micro-channel,

32:43.000 --> 32:47.000
you could also try to build architecture like this on that.

32:47.000 --> 32:48.000
Thank you.

32:48.000 --> 32:49.000
Thank you.

32:49.000 --> 32:50.000
Thanks.

32:50.000 --> 32:52.000
Thanks a lot.

32:58.000 --> 33:00.000
Thank you.

