WEBVTT

00:00.000 --> 00:21.040
Okay, so we are ready to continue with the next session as we were talking in the previous

00:21.040 --> 00:22.040
session.

00:22.040 --> 00:29.960
We continue with the DNS topic, the last year we delivered in 3PA, what we call

00:29.960 --> 00:38.960
encrypted DNS, and this is a talk from Yuseppe and Ramon about the performance analysis.

00:38.960 --> 00:40.960
So, guys, you're up.

00:40.960 --> 00:42.960
Thank you very much.

00:42.960 --> 00:49.600
Thanks for coming, and welcome all of you.

00:49.600 --> 00:56.960
So, first, I'm interested in myself, I'm principal solution architect, in Red Hat,

00:56.960 --> 01:02.960
I'm passionate about technology, and I like to speak in events like that, number of

01:02.960 --> 01:10.840
communities, and I'm also a father of a young gamer, and passionate about videos,

01:10.840 --> 01:17.840
so it's our video, so I'm getting old.

01:17.840 --> 01:25.000
Hi, I'm Yuseppe Andreo, I'm senior format Red Hat, well, I remarked about my work that I

01:25.000 --> 01:29.280
spent a lot of time working with engineering teams on real-production problems with these

01:29.280 --> 01:35.280
guys, and well, I also teach classes in university about several topics.

01:35.280 --> 01:44.680
Okay, we start with the origins and motivations of these investigations, well, everybody

01:44.680 --> 01:49.920
have heard this sentence that it's always DNS, and now we can confirm that if it's not

01:49.920 --> 01:53.280
DNS, it's encrypted DNS.

01:53.280 --> 02:04.760
Okay, playing DNS assumes that the network is friendly, and such anyone on the network path

02:04.760 --> 02:10.360
can, can listen to your queries, can see your queries, can know to who you're listening, and

02:10.360 --> 02:12.800
even can change your money for your answers.

02:12.800 --> 02:17.960
And this is not theoretical, this is inherent to an encrypted DNS.

02:17.960 --> 02:23.880
And this is where a encrypted DNS comes to the rescue, and more than that, zero-trust architectures,

02:23.880 --> 02:30.600
because with that, we have ensure that all communications are authenticated, authorized, and

02:30.600 --> 02:35.600
encrypted, including DNS.

02:35.600 --> 02:42.600
Okay, encrypted DNS and zero-trust architectures, and coming from nowhere, there are two main

02:42.600 --> 02:45.000
tri-forces.

02:45.000 --> 02:52.640
One of that is the US-Memonandum, that mandates that every internal network that has

02:52.640 --> 02:59.080
super-for-hybrid warlobes, should adhere to the zero-trust architectures principles, and

02:59.080 --> 03:06.080
this says that internal network no longer means trusted by default, and in European Union,

03:06.080 --> 03:15.080
we have similar direction with the Cyber Resilience Act that translate the set of best practices

03:15.080 --> 03:20.360
into legal requirements, and also ensures that all the security requirements are in place

03:20.400 --> 03:24.360
for hardware and software products with digital elements.

03:24.360 --> 03:30.120
Well, to secure communication, to ensure communication in free-PA, we must provide certificates,

03:30.120 --> 03:36.720
certificates can be provided by the free-PA.ca, or can be your own certificates, this depends

03:36.720 --> 03:42.040
on your trust model, and most importantly, the encryption is not only provided at ruined

03:42.040 --> 03:49.440
time, it's provided from early stages of installation, and also at boot time, so we can

03:49.480 --> 03:56.440
assure that encryption is based on assumption from the zero.

03:56.440 --> 04:04.440
Okay, DNS is an abstract, it's visual, and it's visibility can be tool, and also a risk.

04:04.440 --> 04:13.000
This is a project that is called DNSMAT, and well, it's visualized geographically, all the

04:13.000 --> 04:20.200
DNS traffic around the world, and highlights the fundamental truth that is that anyone

04:20.200 --> 04:28.000
can see a geortrific, and this is where encrypted DNS can limit what attackers can infer

04:28.000 --> 04:31.000
or exploit by limiting all the visibility.

04:31.000 --> 04:36.000
Okay, come on.

04:37.000 --> 04:41.000
So, what is the alternative?

04:41.000 --> 04:49.000
So, we have to move from an encrypted to encrypted, and the implementation that free-PA has

04:49.000 --> 04:54.000
for the buying server is used DNS over TLS.

04:54.000 --> 04:55.000
What is that?

04:55.000 --> 05:05.000
We move from PUR, UDP, or TCP communication in plain text to TCP communication with encrypted messages.

05:05.000 --> 05:12.000
So, we are tunneling all the communication, and nobody can see, except the client on the

05:12.000 --> 05:13.000
server.

05:13.000 --> 05:16.000
What are the alternatives that we have for encrypted DNS?

05:16.000 --> 05:26.000
We have the DNS over HTTPS, that is very used in public clouds and sites that only

05:26.000 --> 05:29.000
allows HTTP traffic.

05:29.000 --> 05:37.000
And there is a new one that is not very mature jet, that is DNS over a quick, that will

05:37.000 --> 05:46.000
use HTTP three with UDP, but both of them have caveats versus the alternative of using

05:46.000 --> 05:49.000
DNS over TLS.

05:49.000 --> 05:51.000
Okay, so we want to encrypt that.

05:51.000 --> 05:55.000
What is the question that we will raise for the first time?

05:55.000 --> 05:58.000
What is the performance impact that we will have doing that?

05:58.000 --> 06:05.000
It is not just doing, but seeing what is the result of it.

06:05.000 --> 06:11.000
So, we find out, what is the procedure that we will have to test what is happening?

06:11.000 --> 06:18.000
When we move from an encrypted to an encrypted first, we want to know where we will do the test.

06:18.000 --> 06:28.000
So, we use some small instances in the cloud, because we didn't want to do something fancy.

06:28.000 --> 06:38.000
We will select a standard performance client from the DNS or work community that is DNS

06:38.000 --> 06:43.000
perf, almost everyone in the field knows this tool.

06:43.000 --> 06:47.000
And then the most important part is that we have to monitor what is happening.

06:47.000 --> 06:53.000
When we do the performance test, because if we don't monitor what is happening, we will only have some numbers,

06:53.000 --> 07:00.000
but we will not have the possibility to dig deep on what is the result.

07:00.000 --> 07:05.000
So, we create a big amount of records.

07:05.000 --> 07:12.000
We configure three for providing this kind of encrypted, an encrypted communication.

07:12.000 --> 07:19.000
We also use a Kubernetes cluster, because it is very interesting right now.

07:19.000 --> 07:25.000
And we modify some patterns, and we do some performance test to try different alternatives,

07:25.000 --> 07:34.000
and try to stress out the most of the bind server, and see what is happening across the chain.

07:35.000 --> 07:42.000
We measure two different things, firstly, the throughput, and another is the latency, the latency.

07:42.000 --> 07:55.000
We did a deep inspection of what is happening at some point, because we don't want like a median at the end.

07:55.000 --> 08:01.000
We want to have a detail of what is happening at every time that we will implement that.

08:01.000 --> 08:14.000
So, we use one second, and this is architecture that we built for this kind of test.

08:14.000 --> 08:20.000
Okay, I will expose the two architectors and the results.

08:20.000 --> 08:22.000
This is the first one.

08:22.000 --> 08:29.000
This is a client that is joined to a 90-end domain, that is listening on the DNS port,

08:29.000 --> 08:31.000
on the DNS standard.

08:31.000 --> 08:41.000
The rest is the collection of metrics, the parameters and how we visualize it via graphana.

08:41.000 --> 08:47.000
Okay, this is the stable compares, the three scenarios.

08:47.000 --> 08:52.000
This scenario is exactly the same, the only thing that changes the transport protocol.

08:52.000 --> 09:00.000
UDP has expected we obtain the highest throughput, and the lowest latency, in that case, even in the order of microseconds.

09:00.000 --> 09:06.000
In TCP, we see a reduction in throughput and increasing latency.

09:06.000 --> 09:13.000
This is due to reliability cost, congestion, return submission, ordering, etc.

09:13.000 --> 09:22.000
And the most penalized is the DNS over TLS about the reduction and also in lead on increase in latency.

09:22.000 --> 09:30.000
This is expected, but the question is not that this is lower, because everyone expects that,

09:30.000 --> 09:38.000
but how much is lower, and if the behavior is predictable under load and sustain.

09:39.000 --> 09:42.000
Okay, we don't slide the numbers into actual cost.

09:42.000 --> 09:47.000
We see a drop of 70% from transition to UDP to TCP.

09:47.000 --> 09:57.000
This is the cost of reliability set, UDP is extremely fast, microseconds, latency response, the latency, when we go to TCP is triple.

09:57.000 --> 10:03.000
This is due to the mechanism of TCP, three-way hand shape, and the establishing of connections, etc.

10:03.000 --> 10:09.000
That takes roughly three milliseconds to establish TCP to DOT.

10:09.000 --> 10:11.000
We see a drop of 23%.

10:11.000 --> 10:15.000
This is overhead for all the cryptographic operations.

10:15.000 --> 10:21.000
And if we see end-to-end, total capacity impact, we see a drop of 37%.

10:21.000 --> 10:31.000
This is less than this sum, because this one is calculated on a gross status fee one than the total.

10:31.000 --> 10:34.000
Okay, this is not 40%.

10:34.000 --> 10:37.000
And latency increases over 20%.

10:37.000 --> 10:44.000
But here context matters, and the topic is that the latency is incremented, but very few.

10:44.000 --> 10:49.000
So only 0.8 milliseconds if we compare to TCP.

10:49.000 --> 10:54.000
So in Sumari, system behaves really well.

10:54.000 --> 11:00.000
It has sustained the performance of 119k BS.

11:00.000 --> 11:12.000
That's an impressive result, and the system is able to absorb all the cryptotacks without any penalty or falling apart.

11:12.000 --> 11:14.000
Okay, this is next scenario.

11:14.000 --> 11:16.000
This is Kubernetes-based setup.

11:16.000 --> 11:22.000
This is really close to how this is really deployed on production environment.

11:23.000 --> 11:25.000
We use a forwarder.

11:25.000 --> 11:28.000
This is a pod, the product on the Kubernetes cluster.

11:28.000 --> 11:33.000
That talks via standard DNS to core DNS.

11:33.000 --> 11:37.000
And core DNS them forwards the queries to free by and back.

11:37.000 --> 11:41.000
Okay, and this communication is playing UDP.

11:41.000 --> 11:43.000
This is always this case.

11:43.000 --> 11:46.000
And the only thing that you can encrypt is the upstream leg.

11:46.000 --> 11:51.000
So you make the test with this leg, encrypted or unencrypt.

11:51.000 --> 11:59.000
And the rest is more or less the same scheme about recollection and gathering and visualization.

11:59.000 --> 12:04.000
Okay, and this is the thing that more or less can be surprised,

12:04.000 --> 12:09.000
because the performance is identically the same in both cases.

12:09.000 --> 12:18.000
This, what we'll explain later, the latency is increased, but marginally only 0.7 milliseconds.

12:18.000 --> 12:23.000
And the latency is the most affected by 4 milliseconds.

12:23.000 --> 12:30.000
If we translate this also, the surprises that both configurations have the same throughput.

12:30.000 --> 12:38.000
And this is due mainly to how this is done, core DNS versus free power.

12:38.000 --> 12:47.000
Core DNS is designed to be extensive, to be integrated with Kubernetes and free-based and auto-itative bind name server.

12:47.000 --> 12:55.000
And as such, is on to maximize throughput and to maximize the throughput of query handling.

12:55.000 --> 13:00.000
And having TLS encryption has practically any performance penalty.

13:00.000 --> 13:07.000
So this is totally marginally, in third or second, free-based care to extremely well.

13:07.000 --> 13:10.000
And the performance scaling factor is the core DNS.

13:10.000 --> 13:14.000
So is completely outside of free power binding domain.

13:14.000 --> 13:20.000
And, but I've seen it delivers the same throughput as the in-secule-based line.

13:20.000 --> 13:23.000
So the security gain is best-lead weight.

13:23.000 --> 13:26.000
The penalty that you can have in TLS.

13:26.000 --> 13:29.000
Okay.

13:29.000 --> 13:37.000
So we have recorded a demo, so you can see it later if you want.

13:37.000 --> 13:38.000
In this video.

13:38.000 --> 13:42.000
So we test on DNS resolution on clients.

13:42.000 --> 13:43.000
90.

13:43.000 --> 13:45.000
Okay.

13:45.000 --> 13:52.000
So you can see here, and we'll see that what we try to do is try to reproduce and make it reproducible.

13:52.000 --> 14:01.000
For even testing in the CACD pipelines, when there are some new features.

14:01.000 --> 14:08.000
So we first use direct communication from the pod to the DNS server.

14:08.000 --> 14:13.000
To double check that everything was working right with the IPA server.

14:13.000 --> 14:18.000
Then we will try to communicate through core DNS with the baseline.

14:18.000 --> 14:25.000
And the core DNS will not forward it to the IPA server because it is not configured for that.

14:25.000 --> 14:41.000
Then we will move from a standard configuration where the core DNS for that zone is forwarding using both an encrypted communication with the IPA server.

14:41.000 --> 14:45.000
And we did some small test and then the performance test.

14:45.000 --> 14:52.000
And we will see also that is how the core DNS is configured for the forwarding.

14:52.000 --> 15:00.000
And then you will see that we did some performance test running different scenarios.

15:00.000 --> 15:09.000
So we can double check where the results either from the client side and from the server side.

15:09.000 --> 15:17.000
So we can know if the server is having some issues with CPU memory or something like that or has some kind of networking issues.

15:18.000 --> 15:22.000
And we will also know what is the client scene from its side.

15:22.000 --> 15:29.000
So we will see with the DNS Perf we are using jobs, Kubernetes jobs and configuration.

15:29.000 --> 15:33.000
So we can check and modify that pretty easily.

15:33.000 --> 15:36.000
So we can do a lot of different scenarios.

15:36.000 --> 15:42.000
And then we upload the results, the login results.

15:42.000 --> 15:49.000
So we can see also in Prometheus and Grafana both both thus boards.

15:49.000 --> 16:01.000
And with that we find out that the result numbers were not just some gross information about what is happening.

16:01.000 --> 16:09.000
But the performance and the stability is quite impressive.

16:09.000 --> 16:20.000
And that's all. So if you have any questions.

16:20.000 --> 16:23.000
Any questions from anyone?

16:23.000 --> 16:25.000
Yeah, please.

16:25.000 --> 16:31.000
So from your question, you're testing to remember running into a bottleneck.

16:31.000 --> 16:34.000
And also the size of the project.

16:34.000 --> 16:41.000
Aspects, where was it mainly into the extra.

16:41.000 --> 16:45.000
Regarding UDP.

16:45.000 --> 16:49.000
When you were testing was done on the bottom of that.

16:49.000 --> 17:04.000
So the question is if what is the main issue when the performance is impacted?

17:04.000 --> 17:09.000
If it's the TCP or it is the encrypted communication.

17:09.000 --> 17:18.000
So we find out first that there are different performance parameters that should be taken into account when using TCP or UDP.

17:18.000 --> 17:25.000
Just by itself. So we have to tune both sides when we were doing the test.

17:25.000 --> 17:32.000
For sure, TCP has some performance impact.

17:32.000 --> 17:40.000
But on the other hand, and especially in cloud vendors, they have other limitations.

17:40.000 --> 17:47.000
In cloud regarding the number of packets per second that they can handle with the virtual networks that they are using.

17:48.000 --> 17:51.000
Depending on the situation and depending on where are the environments.

17:51.000 --> 17:59.000
You come face more impact by TCP or by by by encrypted configuration.

17:59.000 --> 18:05.000
Indeed, in some other tests that we did in other cloud providers.

18:05.000 --> 18:14.000
We face out that the TCP was more performance at UDP because of this because of the penalty impacted in packets per second.

18:14.000 --> 18:19.000
And in the number of contract connections.

18:19.000 --> 18:25.000
But of course, it has to be something related with encrypting.

18:25.000 --> 18:30.000
We didn't try to use hardware accelerator for encrypting.

18:30.000 --> 18:37.000
And we think it will enhance the communication performance on that.

18:38.000 --> 18:50.000
We find out at the end is that if the client is not so performant and we have seen here that code DNS is not so performant.

18:50.000 --> 18:55.000
Even TCP or the encrypted thing is not relevant for the final fears.

18:55.000 --> 18:59.000
So you have to be a man, what is the whole picture?

18:59.000 --> 19:06.000
To know what is the most impacted factor for the whole architecture.

19:06.000 --> 19:14.000
Well, we have done several performance tuning as someone said.

19:14.000 --> 19:20.000
And we will pull these two articles about with stretch really free time.

19:20.000 --> 19:26.000
One of the most interesting thing is that you can have a lot of threats.

19:26.000 --> 19:31.000
But if you are writing in a file log, all of them has to be synchronized.

19:31.000 --> 19:35.000
So we careful about that.

19:35.000 --> 19:40.000
Well, there is not yet.

19:40.000 --> 19:43.000
Can you explain to the difference?

19:43.000 --> 19:44.000
No.

19:44.000 --> 19:47.000
I think we use the usual one.

19:47.000 --> 19:49.000
I think it is 1.2.

19:49.000 --> 19:53.000
But it is with different DLS versions.

19:53.000 --> 19:56.000
So I think we use the default one from Fedora.

19:56.000 --> 19:58.000
That is 1.2.

19:58.000 --> 20:01.000
But it should go on point 3.

20:01.000 --> 20:03.000
Then it is 1.3.

20:03.000 --> 20:08.000
But we don't know exactly what core DNS is using for communicating with each other.

20:08.000 --> 20:10.000
It is also 1.3.

20:10.000 --> 20:13.000
We will double check.

20:13.000 --> 20:17.000
Well, for next investigations in future,

20:17.000 --> 20:22.000
we have been meant to test this setup with post quantum.

20:22.000 --> 20:26.000
So change all cryptographic protocols, etc.

20:26.000 --> 20:29.000
To see how much affects.

20:30.000 --> 20:33.000
Sir, we have signed for one question.

20:33.000 --> 20:38.000
Is there any other one?

20:38.000 --> 20:39.000
Yeah?

20:39.000 --> 20:40.000
Yeah.

20:40.000 --> 20:44.000
I was intentioned here to have a DLT3 services.

20:44.000 --> 20:47.000
Within the environment.

20:47.000 --> 20:49.000
Is this intended?

20:49.000 --> 20:50.000
I used to come back.

20:50.000 --> 20:52.000
So it is going to look up.

20:52.000 --> 20:54.000
DLT3.

20:56.000 --> 20:58.000
Yeah, it is.

20:58.000 --> 21:02.000
So the question if what is the intention of or what is the

21:02.000 --> 21:05.000
configuration exactly that we have tested.

21:05.000 --> 21:10.000
If it is for internal use or for just also for external communication.

21:10.000 --> 21:18.000
So another test that we didn't do is use the bind server as a general app

21:18.000 --> 21:19.000
stream server.

21:19.000 --> 21:21.000
We just used for that soon.

21:21.000 --> 21:27.000
But we think if this bind server has to forward one other.

21:28.000 --> 21:30.000
Up stream server.

21:30.000 --> 21:35.000
Either you can do use the same DLT communication or maybe.

21:35.000 --> 21:42.000
Then you have to also back forward to unencrypted DNS.

21:42.000 --> 21:49.000
So it should not impact a lot in the bind performance.

21:49.000 --> 21:54.000
Because what we have seen is that the main.

21:54.000 --> 21:59.000
The main factor is usually at the input of a DNS server.

21:59.000 --> 22:03.000
Not anything that they have to do because you can put a lot of different

22:03.000 --> 22:06.000
file threads for the output but on the input.

22:06.000 --> 22:11.000
You have the networking staff in the Linux kernel and so on.

22:11.000 --> 22:17.000
And it is harder to parallelize the input than the output.

22:17.000 --> 22:35.000
Now it is that situation.

22:35.000 --> 22:44.000
When courty Ns supports DOT at the input then you can have all of all of the communication

22:44.000 --> 22:46.000
and cryptone secure.

22:46.000 --> 22:52.000
But as far as I know courty Ns is not supporting encrypted communication yet.

22:52.000 --> 23:04.000
So we are all the different providers of DNS servers are working on some kind of solution for that kind of requirement.

23:04.000 --> 23:14.000
Thank you very much.

23:34.000 --> 23:40.000
Thank you very much.

