WEBVTT

00:00.000 --> 00:10.000
Thank you.

00:10.000 --> 00:11.000
So hello, everyone.

00:11.000 --> 00:14.000
Thanks for staying until the very end of the track.

00:14.000 --> 00:16.000
It's nice to see you here.

00:16.000 --> 00:19.000
So I'm going to talk about multi-stage retrieval.

00:19.000 --> 00:20.000
And so let's say search.

00:20.000 --> 00:22.000
But first of all, why do we knit this?

00:22.000 --> 00:23.000
Right?

00:23.000 --> 00:24.000
This is quite of a mouthful.

00:24.000 --> 00:26.000
Multi-stage kind of retrieval.

00:26.000 --> 00:29.000
So in order to understand what multi-stage retrieval is,

00:30.000 --> 00:32.000
I'm going to use an analogy, you know,

00:32.000 --> 00:34.000
and that is the interview process.

00:34.000 --> 00:37.000
We may be familiar with the interview process,

00:37.000 --> 00:40.000
you know, like having been through that quite a few times.

00:40.000 --> 00:43.000
Mostly probably on the candidate kind of side,

00:43.000 --> 00:46.000
but I'm going to look at it from the recruiter point of view.

00:46.000 --> 00:47.000
Okay?

00:47.000 --> 00:49.000
So the market, the market is huge.

00:49.000 --> 00:53.000
We have millions of people there and we want the best, right, for the job.

00:53.000 --> 00:55.000
So how do we do that?

00:55.000 --> 00:58.000
Are we going to interview the millions of people that we have out there?

00:58.000 --> 01:01.000
Now, we need like a process, you know, to do that,

01:01.000 --> 01:04.000
to get from a very huge kind of set of candidates

01:04.000 --> 01:06.000
to get the very best of them.

01:06.000 --> 01:09.000
So we can define something like this, right?

01:09.000 --> 01:11.000
Like a typical funnel.

01:11.000 --> 01:14.000
You are probably familiar with this from me commerce.

01:14.000 --> 01:17.000
So basically, the first thing that we can do is

01:17.000 --> 01:19.000
are resume kind of a scan.

01:19.000 --> 01:23.000
So I understand like the technologies that everyone has

01:23.000 --> 01:24.000
or some keywords.

01:24.000 --> 01:27.000
Then we're probably going to do like a phone screen call

01:27.000 --> 01:30.000
since you get more understanding of the candidates

01:30.000 --> 01:32.000
that would fit or not.

01:32.000 --> 01:35.000
And in the end, we're going to perform like once one interviews

01:35.000 --> 01:37.000
probably like several of them.

01:37.000 --> 01:39.000
While we're doing this, well, in the end,

01:39.000 --> 01:43.000
we have like a ton of resumes out there

01:43.000 --> 01:46.000
that we need to be able to filter quickly

01:46.000 --> 01:49.000
in order to get just some several ones

01:49.000 --> 01:52.000
that we will be using for a phone screen call

01:52.000 --> 01:54.000
and then in the end for the very best of them

01:54.000 --> 01:56.000
will be like spending in time to interview them.

01:56.000 --> 01:57.000
Why is that?

01:57.000 --> 01:59.000
Because basically we are going to reduce.

01:59.000 --> 02:02.000
Do you see that we are going to reduce substantially?

02:02.000 --> 02:04.000
Like the number of candidates that we get

02:04.000 --> 02:06.000
from one stage to the other.

02:06.000 --> 02:08.000
And that's some purpose because basically we have

02:08.000 --> 02:10.000
a limited amount of time, right?

02:10.000 --> 02:12.000
We can be interviewing every one of the candidates

02:12.000 --> 02:14.000
in order to get a song sign up.

02:14.000 --> 02:16.000
What we want to do is to perform like a resume scan

02:16.000 --> 02:18.000
for example, that's something that we can do in seconds

02:18.000 --> 02:21.000
probably just looking quickly through it.

02:21.000 --> 02:23.000
Then we can do a phone screen call that is going

02:23.000 --> 02:25.000
to take minutes and then we're going to really spend

02:25.000 --> 02:28.000
time from the most promising candidate to spend time

02:28.000 --> 02:31.000
into an interview that can, you know,

02:31.000 --> 02:34.000
like kind of spend several hours and several different interviews.

02:34.000 --> 02:38.000
So what you're seeing here is that we are using

02:38.000 --> 02:40.000
different kinds of tests.

02:40.000 --> 02:42.000
We're trying to get different kind of signals

02:42.000 --> 02:44.000
and the sending that there will be ones

02:44.000 --> 02:46.000
that will be like very fast to do.

02:46.000 --> 02:49.000
And then others that will be like,

02:49.000 --> 02:51.000
we'll provide like better signals,

02:51.000 --> 02:54.000
that's more powerful signals that this is

02:54.000 --> 02:56.000
like a promising candidate or not,

02:56.000 --> 03:00.000
but we'll take it like a long time to get this sign up for.

03:00.000 --> 03:02.000
So this is basically,

03:02.000 --> 03:04.000
the time is going to notice stage of a table.

03:04.000 --> 03:06.000
We need multi stage of a table because we have

03:06.000 --> 03:07.000
to match.

03:07.000 --> 03:10.000
We have basically like a huge number of candidates

03:10.000 --> 03:13.000
and we need to be able to solve through them.

03:13.000 --> 03:16.000
We have to little time because basically,

03:16.000 --> 03:18.000
you know, like the users or the customers

03:19.000 --> 03:21.000
like they put in some, they want it all,

03:21.000 --> 03:22.000
and they want it now.

03:22.000 --> 03:26.000
So we are always constrained by the amount of time that they use.

03:26.000 --> 03:29.000
It's going to be to find acceptable to wait.

03:29.000 --> 03:32.000
And then we have the good news is that we have like

03:32.000 --> 03:34.000
very good tools at our disposal.

03:34.000 --> 03:36.000
Like models are getting better.

03:36.000 --> 03:39.000
And we get a lot more options for us to choose from,

03:39.000 --> 03:41.000
but they come across.

03:41.000 --> 03:45.000
So we need to be very conscious on what we are using.

03:45.000 --> 03:48.000
And then it's because, you know, both in terms of latency,

03:48.000 --> 03:49.000
both in terms of money.

03:49.000 --> 03:54.000
So we better understand what the class for these models are.

03:54.000 --> 03:57.000
So this is basically what we, oh my gosh.

03:57.000 --> 04:00.000
I hope you can see that from my point of view,

04:00.000 --> 04:02.000
it's like a little bit washed out.

04:02.000 --> 04:05.000
But this is what multi stage we choose all this about, right?

04:05.000 --> 04:08.000
So we are starting with an initial phase,

04:08.000 --> 04:11.000
which we can have millions of billions of documents.

04:11.000 --> 04:14.000
And then we are applying techniques that are very fast.

04:14.000 --> 04:17.000
You know, like basic filtering is in tax.

04:17.000 --> 04:20.000
For example, being 25, then specters and specters and specters.

04:20.000 --> 04:26.000
You have seen that we have like techniques to apply like very fast filtering based on that.

04:26.000 --> 04:31.000
And very fast approximating kind of search using those.

04:31.000 --> 04:34.000
And semantic search that we can use, for example,

04:34.000 --> 04:37.000
with the semantic text skills like that as the search provides.

04:37.000 --> 04:43.000
Then from those, after casting like a wide net of what we are trying to catch,

04:43.000 --> 04:48.000
then we are going to get from those thousands that pass the initial filter

04:48.000 --> 04:51.000
that is needed, that is very cheap to perform.

04:51.000 --> 04:54.000
Then we are going to perform hybrid search, query rules, query scores,

04:54.000 --> 04:58.000
or learning to rank, which is a little bit more expensive to perform.

04:58.000 --> 05:03.000
But we have like a reduced amount of candidates for it to process.

05:03.000 --> 05:08.000
And then at the very end, a latest stage we are ranking face in which we can be talking about 10s

05:08.000 --> 05:12.000
or hundreds of documents in which we can put the pedal to the metal in the middle.

05:12.000 --> 05:17.000
So the metal in terms of the algorithms that we are going to use and the models that we are going to use.

05:17.000 --> 05:20.000
And the sending that they can be like much more expensive.

05:20.000 --> 05:27.000
But at least they will be used against like a very reduced number of candidates.

05:27.000 --> 05:29.000
Of course, they are trade-offs.

05:29.000 --> 05:35.000
The magical kind of words for search relevance, we call them precision.

05:35.000 --> 05:41.000
So we start with, we start with, we call basically casting a wide net trying to catch all the fishes like we can.

05:41.000 --> 05:44.000
So net the golden fish doesn't escape the net.

05:44.000 --> 05:49.000
And then as we are moving forward, we are going to increase precision.

05:49.000 --> 05:52.000
So we are trying to get the best of what works.

05:52.000 --> 05:57.000
And then in terms of course of the speed everything has, so it was as I was mentioning.

05:57.000 --> 06:05.000
Then we have like faster techniques on the very bottom that we are using to filter through like a huge amount of documents or candidates.

06:05.000 --> 06:12.000
And then we are going more slower as we go further up.

06:12.000 --> 06:15.000
So put it all together basically.

06:15.000 --> 06:22.000
You already know that we have these tools that are disposal and it's just the question of organizing them into a kind of pipeline.

06:22.000 --> 06:26.000
And which we can understand where it's of these tools fit.

06:26.000 --> 06:29.000
And that's also providing us the maximum benefit.

06:29.000 --> 06:33.000
I'm trying to minimize the cost of it.

06:33.000 --> 06:39.000
So this is fine as a kind of fundamentals way of looking into it.

06:39.000 --> 06:45.000
But basically we need to tell the search engine how we're going to do this.

06:45.000 --> 06:46.000
Right?

06:46.000 --> 06:51.000
So you are probably familiar already with the results in elastic search.

06:51.000 --> 06:53.000
By the way, this is not a busy on test.

06:53.000 --> 06:56.000
I'm not expecting you to read this.

06:56.000 --> 06:59.000
In case you can, that's perfect.

06:59.000 --> 07:03.000
This is just how a query looks like when using the query DSL.

07:03.000 --> 07:04.000
Right?

07:04.000 --> 07:05.000
It's a tree structure.

07:05.000 --> 07:08.000
You can see that it maps perfectly to JSON.

07:08.000 --> 07:15.000
And then we have like a single, like the, it all revolves against using like a single query.

07:15.000 --> 07:20.000
Which can be maths query, K&M, purpose, maintenance, neighbourhoods, attend query.

07:20.000 --> 07:22.000
But the query can also be composed of other queries.

07:22.000 --> 07:27.000
We can use like other structures like the Boolean query or these maps, etc.

07:27.000 --> 07:29.000
And we can change the results.

07:29.000 --> 07:33.000
Of queries using the script score, function score just to change all these.

07:33.000 --> 07:37.000
And then there's the retriever framework in which we can do the same thing.

07:37.000 --> 07:45.000
But you know, we have been working like a, like an ESQL in the elastic search query language.

07:45.000 --> 07:49.000
This is like a language that has been created.

07:49.000 --> 07:52.000
It started like being applied through the observability,

07:52.000 --> 08:00.000
I'm security kind of use cases precisely to try to make this kind of data analysis easier to do.

08:00.000 --> 08:02.000
So you're familiar with this plan.

08:02.000 --> 08:05.000
For example, they use like a like a singular kind of approach.

08:05.000 --> 08:08.000
And which basically it's, you can think of it's command.

08:08.000 --> 08:13.000
You know, we start like selecting all the data and then we start applying commands.

08:13.000 --> 08:16.000
Just like if they were a pipeline, right?

08:16.000 --> 08:19.000
Like a thing of it that's a, as a factory, right?

08:19.000 --> 08:26.000
And we say every single command that we apply is transforming the input into some other output.

08:26.000 --> 08:32.000
So the processing model for ESQL is basically, as I was mentioned, we start with all our data, right?

08:32.000 --> 08:38.000
We process it on steps and the output for a step is input for the next one.

08:38.000 --> 08:40.000
And this looks familiar, right?

08:40.000 --> 08:43.000
So what we were mentioned in a multi-stage retriever.

08:43.000 --> 08:48.000
So kind of the mental model that we have for ESQL maps perfectly in some multi-stage retriever.

08:48.000 --> 08:50.000
So we're going to start going to see all the results.

08:50.000 --> 08:52.000
Hills are on each of the steps.

08:52.000 --> 08:56.000
And yeah, each of the steps can be more expensive to apply, basically.

08:56.000 --> 09:01.000
Because it is being used for where I reduced amount of candidates.

09:01.000 --> 09:07.000
So we're going to talk about how this would look like using ESQL.

09:07.000 --> 09:11.000
So basically, we will start with our very simple query.

09:11.000 --> 09:13.000
It's like a star, right?

09:13.000 --> 09:14.000
It would be the equivalent.

09:14.000 --> 09:18.000
We're going to take everything from a catalog index that we have over there.

09:18.000 --> 09:22.000
We're starting with just a retrieval phase, right?

09:22.000 --> 09:24.000
We get everything there.

09:24.000 --> 09:27.000
And then we can apply some basic filtering.

09:27.000 --> 09:29.000
We're not talking about relevance yet.

09:29.000 --> 09:32.000
We are like filtering out all the things that we don't want.

09:32.000 --> 09:36.000
We're applying like very deep, very easy to do kind of filtering there.

09:36.000 --> 09:42.000
We're filtering on price and we're filtering on the category of items that we want to consider.

09:43.000 --> 09:45.000
Then the fun starts.

09:45.000 --> 09:48.000
We are going to start adding relevance to this, right?

09:48.000 --> 09:50.000
We are not just filtering out.

09:50.000 --> 09:52.000
We'll try to understand what the score is.

09:52.000 --> 09:56.000
How relevant is each of the candidates that we had.

09:56.000 --> 10:00.000
So what we are using there is, I think, the metadata underscore score.

10:00.000 --> 10:05.000
And that basically tells ESQL that I want to actually score things out.

10:05.000 --> 10:08.000
Right now I was only doing filtering for that.

10:08.000 --> 10:12.000
And then I'm going to introduce like a maths query there.

10:12.000 --> 10:17.000
You see our where, which is a filter and maths, which is a full text function.

10:17.000 --> 10:23.000
What is this does is basically it would be the equivalent to a maths in the query DSL.

10:23.000 --> 10:26.000
And that it comes with a trick.

10:26.000 --> 10:31.000
And that kind of full text function will modify the score.

10:31.000 --> 10:32.000
We'll update the score.

10:32.000 --> 10:36.000
So we can then sort on it and filter on it.

10:36.000 --> 10:39.000
This is basically like a top end kind of query.

10:39.000 --> 10:42.000
But make more or less simple.

10:42.000 --> 10:46.000
And of course, we can start like adding some combination.

10:46.000 --> 10:49.000
And if you're familiar with that Boolean query, we'll ask the search.

10:49.000 --> 10:51.000
Then you understand what this is.

10:51.000 --> 10:54.000
But in a quite simplified kind of way.

10:54.000 --> 10:56.000
And we can also apply boosting to it.

10:56.000 --> 11:01.000
We are basically considering that the product kind of title is much more important.

11:01.000 --> 11:06.000
And I'm giving it that boost of three over the description.

11:06.000 --> 11:10.000
And then of course, lexical search is really nice.

11:10.000 --> 11:15.000
But you know, semantic search is something that we need to take into account.

11:15.000 --> 11:19.000
Just to take meaning into account and not just lexical search.

11:19.000 --> 11:22.000
So we can of course include semantic search here.

11:22.000 --> 11:26.000
Here we are using a semantic text kind of field.

11:26.000 --> 11:30.000
So the kind of thing is this, but the same thing we could achieve with the kind of function.

11:30.000 --> 11:32.000
And calculating the text embedding as well.

11:32.000 --> 11:33.000
It would be like the same thing.

11:33.000 --> 11:36.000
It's just more simplified using this.

11:36.000 --> 11:41.000
And you can see that we are using fork and fuse kind of commands on top.

11:41.000 --> 11:49.000
And on the bottom of those two queries, you can think basically as the fork of unix.

11:49.000 --> 11:50.000
Right?

11:50.000 --> 11:54.000
We are doing basically here is like spanning multiple two queries.

11:54.000 --> 11:56.000
And then I'm going to join at the end.

11:56.000 --> 11:57.000
Right?

11:57.000 --> 11:58.000
Using fuse.

11:59.000 --> 12:02.000
So again, we have here in this like a basic example.

12:02.000 --> 12:04.000
We have a lexical query.

12:04.000 --> 12:05.000
We have semantic query.

12:05.000 --> 12:07.000
And then we are combining them.

12:07.000 --> 12:08.000
We can combine them of course.

12:08.000 --> 12:13.000
We have seen like previously like the inconsistencies of just having different like, like,

12:13.000 --> 12:15.000
scoring spaces, right?

12:15.000 --> 12:17.000
We can use for example, Lara Raff.

12:17.000 --> 12:23.000
Like, we're a reciprocal rank version in order not to consider the scores.

12:23.000 --> 12:27.000
But consider the relative position of the results that we get from the different queries.

12:27.000 --> 12:32.000
Or in case that we really understand what we are doing.

12:32.000 --> 12:36.000
And we really understand the scoring space for the different queries.

12:36.000 --> 12:38.000
We can apply a linear combination.

12:38.000 --> 12:39.000
Right?

12:39.000 --> 12:43.000
And which we are providing the actual weight that we want to use.

12:43.000 --> 12:46.000
And we can of course include like a like some normalizer.

12:46.000 --> 12:50.000
I'm like for post free trying to keep this simple.

12:50.000 --> 12:54.000
So I haven't put all the options that you can use here.

12:54.000 --> 13:00.000
But you can expect that for example the maths function or the linear combination.

13:00.000 --> 13:04.000
And I said all the different parameters that are expected from the,

13:04.000 --> 13:09.000
that you expect from the previous so for example.

13:09.000 --> 13:13.000
It would be nice if it would be just a semantic or lexical search.

13:13.000 --> 13:15.000
But that doesn't carry.

13:15.000 --> 13:16.000
Right?

13:16.000 --> 13:20.000
When we've got like, for example, if we are looking, we are searching over news.

13:20.000 --> 13:24.000
Then we need to take a listen for example into account.

13:24.000 --> 13:30.000
If we are talking about the commerce, maybe we are looking at the popularity of a different items.

13:30.000 --> 13:36.000
So we can just rely on the pure kind of lexical or semantic search.

13:36.000 --> 13:38.000
We need to incorporate more signals into it.

13:38.000 --> 13:43.000
It could be like a date as I was mentioning, popularity, you name it.

13:43.000 --> 13:49.000
So we need to find a way of mixing them into the actual lexical scoring that we have calculated.

13:49.000 --> 13:50.000
Already.

13:50.000 --> 13:57.000
And that's where it is real excels because basically it was built up from the ground up in some

13:57.000 --> 14:01.000
to take care of analytical and of sorority coming in the few cases.

14:01.000 --> 14:05.000
So we have like a very nice way of calculating expressions.

14:05.000 --> 14:11.000
And instead of doing like a script score, you know, like all the functional scoring things that you are,

14:11.000 --> 14:13.000
that you can do of course in the previous hall.

14:13.000 --> 14:16.000
But we can do them in a much more simple kind of way.

14:16.000 --> 14:18.000
It's much more intentional.

14:18.000 --> 14:24.000
And for my point of view, I think that's more understandable.

14:24.000 --> 14:28.000
As you can see that we have been like going through the retrieval phase.

14:28.000 --> 14:34.000
We've been to the fine tuning phase using like other signals that we have been combined here.

14:34.000 --> 14:38.000
And then we can go through there to find a really ranking kind of phase.

14:38.000 --> 14:42.000
The more expensive one in which we can like use a really ranking model.

14:42.000 --> 14:47.000
For example, here we are using a agenda model for we ranking.

14:47.000 --> 14:53.000
And we rank using our specific query on the different on different fields.

14:53.000 --> 14:57.000
For example, in this case, like a product and the description of the item.

14:57.000 --> 15:00.000
And just sorting again.

15:00.000 --> 15:06.000
As you can see, we have chosen like a 30 as the initial limit as the initial top end that we were considering.

15:06.000 --> 15:10.000
And then we are out putting out like much less results.

15:10.000 --> 15:14.000
And yeah, this is good.

15:14.000 --> 15:19.000
This is something that we can do as well with a query DSL and the retrieval kind of thing.

15:19.000 --> 15:23.000
But we wanted to go a little bit like a father and that.

15:23.000 --> 15:27.000
We want to push forward a little bit of search using ESKL.

15:27.000 --> 15:34.000
When you have something new, you really want to put it to the test and see where the limits are for that.

15:34.000 --> 15:39.000
So we started to throw well alums into the mix.

15:39.000 --> 15:42.000
And here you can see like an example that we are doing.

15:42.000 --> 15:49.000
We are doing a prompt to analyze asking a question of the different process we are getting.

15:49.000 --> 16:02.000
So in the end, what we will do is not only to just take like the best kind of candidates we got from all the different stages and the multi stage kind of pipeline that we have been defining.

16:02.000 --> 16:08.000
But what we are getting in the end is getting those results and asking a question to analyze them about it.

16:08.000 --> 16:12.000
Just to enrich it with some alums generated content.

16:12.000 --> 16:23.000
So this is like a quite an interesting query because in the end it achieves like a lot of different things.

16:23.000 --> 16:29.000
You know, in a single place, using a single language, using a single query.

16:29.000 --> 16:31.000
You can see that we have been filtering.

16:31.000 --> 16:34.000
We have been applying electrical search, semantic search.

16:34.000 --> 16:36.000
We have been combining the two of them.

16:36.000 --> 16:41.000
We have been applying like some rescoring kind of simple rules in this case.

16:41.000 --> 16:45.000
Then we have been rewrunking using a rewrunking model.

16:45.000 --> 16:49.000
And we are forming like some lalum augmentation on top of it.

16:49.000 --> 16:54.000
Like the kind of thing that this allows you is a rapid kind of iteration.

16:54.000 --> 16:56.000
It's basically it's like very fast to write.

16:56.000 --> 16:58.000
You can see like the results.

16:58.000 --> 17:00.000
You can go like one stage at a time.

17:00.000 --> 17:06.000
I find that like very kind of refreshing rather than fighting a JSON to be honest.

17:06.000 --> 17:10.000
But it could be just me.

17:10.000 --> 17:12.000
Come in soon.

17:12.000 --> 17:18.000
Of course, this is like the kind of things that I will be working tomorrow on.

17:18.000 --> 17:22.000
So there are like more things that we are preparing for this.

17:22.000 --> 17:25.000
So improving search support, okay?

17:25.000 --> 17:29.000
So we keep like including search capabilities into ESQL.

17:29.000 --> 17:34.000
You have seen like a that you can basically already use like multi-stage kind of retrieval.

17:34.000 --> 17:35.000
This in ESQL.

17:35.000 --> 17:39.000
But we want to keep adding like more support to it.

17:39.000 --> 17:41.000
Right now we don't have a sparse vector support.

17:41.000 --> 17:45.000
That's something that we have like a really that we are really looking forward to it.

17:45.000 --> 17:50.000
And we have seen like some of the advantages of using like a sparse vectors in some of the previous sessions.

17:50.000 --> 17:56.000
We can also have semantic text with a sparse vector basically like it the same thing.

17:56.000 --> 18:00.000
semantic text is already supported but for them specters only.

18:00.000 --> 18:02.000
We want to create high lighting.

18:02.000 --> 18:05.000
We'll probably like some high lighting support, okay?

18:05.000 --> 18:14.000
So we have like a the top tanks scale function there and also tanking that in order for you to perform the actual tanking.

18:14.000 --> 18:21.000
And then try to determine you could for example extract the top tanks or extract the tanking from a document.

18:21.000 --> 18:22.000
And score that yourself.

18:22.000 --> 18:25.000
That's something that you can totally do right now with ESQL.

18:25.000 --> 18:31.000
But we want to provide like a better kind of more use of friendly experience in terms of high lighting.

18:31.000 --> 18:34.000
And what is vector arithmetic?

18:34.000 --> 18:36.000
Yeah, we want to be able to multiply it.

18:36.000 --> 18:38.000
A vector etc.

18:38.000 --> 18:44.000
Right now you can of course use K&N but we want to provide you know and use like custom vectors in your active functions.

18:44.000 --> 18:48.000
Remember the script scores that we are doing for exact kind of vector search.

18:48.000 --> 18:54.000
That's something that you can do like much much more easy now using ESQL.

18:54.000 --> 18:58.000
But we want to provide you the ability as well to perform custom scoring.

18:58.000 --> 18:59.000
Right?

18:59.000 --> 19:07.000
And semantic search over close cluster search support that's coming almost now.

19:07.000 --> 19:10.000
Next week hopefully.

19:10.000 --> 19:12.000
Yeah?

19:12.000 --> 19:19.000
And we will want to just be supporting like search features that would be like the easy thing to do.

19:19.000 --> 19:20.000
Right?

19:20.000 --> 19:23.000
We want to push this a little bit further as I was mentioned.

19:23.000 --> 19:28.000
So we are going to include like the results diversifications using MMR.

19:28.000 --> 19:32.000
And we will have like a more interesting news in terms of multiple search.

19:32.000 --> 19:37.000
We've already seen some some great presentations about that.

19:37.000 --> 19:39.000
But we will start with the image support.

19:39.000 --> 19:43.000
That's something that you should be looking for because it's coming soon.

19:43.000 --> 19:46.000
So I've been starting for a while now.

19:46.000 --> 19:49.000
Hopefully everything's in team.

19:50.000 --> 19:53.000
Basically multi-stage retrieval is in ESQL.

19:53.000 --> 19:55.000
Why do we need multi-stage retrieval?

19:55.000 --> 19:59.000
Basically we want it because it allows us to use better ranking methods.

19:59.000 --> 20:03.000
Apply to the aids of the different stages that we are in.

20:03.000 --> 20:10.000
Either when we are casting a wide net or when we are trying to select like the better results out of it.

20:10.000 --> 20:18.000
Gradually filtering and the running expensive ranking on top of the canvas to keep like our costs in check.

20:18.000 --> 20:21.000
Both in terms of money, both in terms of latency.

20:21.000 --> 20:26.000
And ESQL will help like building like a similar like a very similar mental model.

20:26.000 --> 20:28.000
Not really just in for this.

20:28.000 --> 20:33.000
Basically we are using the piped query syntax which makes the processing steps explicit.

20:33.000 --> 20:40.000
No more nesting query sober and try to understand which query fits into each other.

20:40.000 --> 20:43.000
It's also already of course high results.

20:43.000 --> 20:45.000
We can do lexical search.

20:45.000 --> 20:48.000
We can do back to ourselves. We can fusion them together.

20:48.000 --> 20:54.000
And combine external retrieval aggregation, custom scoring without switching what we are doing.

20:54.000 --> 20:56.000
We don't have context switch.

20:56.000 --> 21:00.000
We don't need to go define something else.

21:00.000 --> 21:08.000
We just can't keep it right in on a single query using a single language instead of having things like all over the place.

21:09.000 --> 21:12.000
So this one might talk.

21:12.000 --> 21:19.000
We have time for questions here.

21:19.000 --> 21:20.000
And here you go.

21:20.000 --> 21:22.000
So you have three questions.

21:22.000 --> 21:25.000
One of the SLs we have to ask you out.

21:25.000 --> 21:26.000
The data is in the structure.

21:26.000 --> 21:30.000
All they have in performance and results.

21:30.000 --> 21:33.000
Well, results are exactly the same for performance.

21:33.000 --> 21:37.000
There are like, you know, there we have been like optimizing the query.

21:37.000 --> 21:40.000
The SL for like 15 years now.

21:40.000 --> 21:44.000
And for a query for ESQL is not exactly like the same thing.

21:44.000 --> 21:46.000
We keep it right in on performance.

21:46.000 --> 21:50.000
There are like some use cases and you need to take it into take into account of course.

21:50.000 --> 21:52.000
Like latency and output.

21:52.000 --> 21:56.000
We have like different results format as of now from the query itself.

21:56.000 --> 21:59.000
But we'll go is to have ESQL to perform better.

21:59.000 --> 22:01.000
And then we'll add ESL.

22:01.000 --> 22:04.000
That's the goal that we have for ESQL in general.

22:04.000 --> 22:07.000
Other questions?

22:07.000 --> 22:08.000
Yes.

22:09.000 --> 22:11.000
What do you need to perform to use?

22:11.000 --> 22:17.000
Can you just write for the risk of a cool rank of views around the next drafts?

22:17.000 --> 22:19.000
Can you repeat the question?

22:19.000 --> 22:20.000
Yes.

22:20.000 --> 22:22.000
You have this for for this blog, right?

22:22.000 --> 22:23.000
Yes.

22:23.000 --> 22:26.000
You have this for for this blog, right?

22:26.000 --> 22:27.000
Yes.

22:27.000 --> 22:30.000
You have this for for this blog, right?

22:30.000 --> 22:31.000
Yes.

22:31.000 --> 22:34.000
You have this for for this blog, right?

22:34.000 --> 22:35.000
Yes.

22:35.000 --> 22:38.000
Okay.

22:38.000 --> 22:43.000
So the question is why do we need to use for views, right?

22:43.000 --> 22:44.000
Yeah.

22:44.000 --> 22:45.000
Yeah.

22:45.000 --> 22:47.000
Basically.

22:47.000 --> 22:48.000
That's a very good question.

22:48.000 --> 22:51.000
So in the end, you have two options for doing that, right?

22:51.000 --> 22:55.000
Keeping to keep in mind that each of these queries,

22:55.000 --> 22:59.000
modify the underscore score, kind of omnisite a field,

22:59.000 --> 23:01.000
which we are using to calculate this score.

23:01.000 --> 23:04.000
In case we do like what you're saying, like saying, hey,

23:04.000 --> 23:07.000
let's not like we're at this in 124 views,

23:07.000 --> 23:11.000
and just keep them on a single kind of query.

23:11.000 --> 23:15.000
If you want to calculate like the score that is coming from

23:15.000 --> 23:18.000
each of the different queries, you will have to do that yourself.

23:18.000 --> 23:19.000
Which you can do?

23:19.000 --> 23:22.000
You can use like a score function in order to tell you.

23:22.000 --> 23:23.000
So so.

23:23.000 --> 23:24.000
Yeah.

23:24.000 --> 23:25.000
What would you discuss on that?

23:25.000 --> 23:27.000
It's a reasonable argument.

23:27.000 --> 23:28.000
Okay.

23:28.000 --> 23:30.000
I think you can get away without it.

23:31.000 --> 23:32.000
Happy to discuss this further.

23:32.000 --> 23:36.000
But in the end, again, what we're trying to do with focus is to make sure

23:36.000 --> 23:38.000
that you can like expand multiple queries.

23:38.000 --> 23:39.000
Yeah.

23:39.000 --> 23:40.000
Right?

23:40.000 --> 23:42.000
And then in the end, you will be, you will get those results.

23:42.000 --> 23:45.000
Think of it as a different kind of tables mixed together,

23:45.000 --> 23:47.000
without discriminator kind of count.

23:47.000 --> 23:48.000
That you will be using.

23:48.000 --> 23:50.000
And then you can treat them separately,

23:50.000 --> 23:53.000
understanding what score is coming from each of them,

23:53.000 --> 23:56.000
or you could implement, for example, a regular,

23:56.000 --> 23:59.000
or an efficient yourself, just getting the results from them,

23:59.000 --> 24:01.000
and having them differentiated.

24:01.000 --> 24:02.000
Hopefully.

24:02.000 --> 24:04.000
I think that I haven't convinced you.

24:04.000 --> 24:05.000
Oh, my.

24:05.000 --> 24:06.000
What?

24:06.000 --> 24:07.000
We can.

24:07.000 --> 24:09.000
We can maybe like talk about that later.

24:09.000 --> 24:12.000
I only care why we're going to list myself.

24:12.000 --> 24:13.000
Okay.

24:13.000 --> 24:14.000
Okay.

24:14.000 --> 24:15.000
Yeah.

24:15.000 --> 24:18.000
As I mentioned, like probably that detail is that each of these

24:18.000 --> 24:20.000
different, um, kind of queries.

24:20.000 --> 24:23.000
The maths, uh, kind of functions that you see there.

24:23.000 --> 24:25.000
They don't provide like an individual score out of them.

24:25.000 --> 24:28.000
They keep modifying, like, they only single,

24:28.000 --> 24:29.000
let's say it under score score.

24:29.000 --> 24:31.000
Um, I feel so.

24:31.000 --> 24:32.000
You don't read.

24:32.000 --> 24:34.000
If you read it, it's anonymous, right?

24:34.000 --> 24:36.000
Because it won't talk to anything,

24:36.000 --> 24:38.000
but then you decide the score of these questions.

24:38.000 --> 24:40.000
And then you care about the next question.

24:40.000 --> 24:43.000
You could do, like, exactly this,

24:43.000 --> 24:46.000
using, like, a, like, a semantic search function.

24:46.000 --> 24:50.000
I use it with score inside of each of the software

24:50.000 --> 24:51.000
and just, uh, with those.

24:51.000 --> 24:53.000
Can you ask the data sign?

24:54.000 --> 24:56.000
No, no, no, I don't like the discussion.

24:56.000 --> 24:57.000
But maybe there's other questions?

24:57.000 --> 24:58.000
Get that hotel.

24:58.000 --> 24:59.000
Okay.

24:59.000 --> 25:00.000
All right.

25:00.000 --> 25:01.000
We can discuss.

25:01.000 --> 25:02.000
Thanks.

25:02.000 --> 25:03.000
Thanks for the question.

25:03.000 --> 25:04.000
Okay.

25:04.000 --> 25:05.000
Next question.

25:05.000 --> 25:06.000
Okay.

25:06.000 --> 25:07.000
Thank you for coming to發 them.

25:07.000 --> 25:08.000
Hope you had a great weekend.

25:08.000 --> 25:09.000
Thank you.

25:09.000 --> 25:10.000
Thank you.

25:10.000 --> 25:11.000
Thank you.

25:11.000 --> 25:13.000
Thank you.

25:13.000 --> 25:14.000
Thank you.

