Connect AI to Billions of Legal Documents — Simon Eskildsen, turbopuffer & Jacob Lauritzen, Legora
스크립트
00:00:00Okay. Hi, everyone. Welcome to this 20-minute talk about connecting AI to loads of legal documents.
00:00:22My name is Jacob. I'm an engineer at Legora.
00:00:25Yeah, and I've got to step into frame here.
00:00:27I'm Simon. I'm the CEO and co-founder of TurboPuffer, a search engine that we work with Legora and others on.
00:00:36Super quickly, introduction to Legora.
00:00:41We're a collaborative AI platform for legal work, and so that means we have law firms that are clients,
00:00:47and we have in-house legal teams that are clients, and they use Legora to do reviews of contracts.
00:00:53They use it to go through an absurd amount of contracts and make sure that they all look good.
00:00:58They use them to create new contracts.
00:01:00They do legal research, which means looking over all potential law, and they collaborate inside Legora.
00:01:07So you can think of Legora sort of as a linear slash figma slash notion slash GitHub for legal work.
00:01:14It's a lot.
00:01:16We are one of the fastest-growing companies right now.
00:01:20We work extremely fast.
00:01:21Yeah, tons of numbers on the screen.
00:01:24I'll just skip through that.
00:01:26What we really want to talk about is search today.
00:01:29So at Legora, there's two types of search that we do.
00:01:32There is project search and legal research.
00:01:35Project search is basically projects in Legora.
00:01:39It's like the unit of work that you have.
00:01:41So if, let's say, you are SpaceX and you want to acquire Cursor, then that would be one project with your law firm.
00:01:50And so they would go into Legora, and they would upload all these documents, and the law firm that helped you would go through all of the employment agreements, all of the contracts with suppliers.
00:01:59I know Cursor is using TurboPuffer, so maybe there's a contract there they'd look at.
00:02:05But basically, you do the search confined to a project.
00:02:07And projects can be tens of documents to millions of documents.
00:02:10The other use case is legal research.
00:02:13And legal research is sort of a deep research-style workload where we'll search across tons of laws, previous cases, regulations, et cetera, et cetera, and people use this to answer questions such as, like, how do we handle this specific thing?
00:02:27And they'll also use it for litigation.
00:02:29Maybe they want to sue someone, or maybe they are getting sued, and they'll use legal research to support and help their case.
00:02:36So if we start at number one, project search.
00:02:39We've been through a little bit of a ride here on how we do search.
00:02:44Starting at, you know, hundreds of thousands of documents, all the way into two billions of documents.
00:02:49And we've tried a lot of different things.
00:02:52So first, we started with a very, very simple one, which is just a single elastic search cluster for all of our search workloads.
00:02:59That worked relatively well.
00:03:02It was sort of a simple setup.
00:03:03All of the tenants, you know, our clients, our users, would be on one big blob storage, where we store the raw documents,
00:03:10and on one big elastic search, where we would do all of their searching, the indexing and the searching.
00:03:15Super simple.
00:03:16Worked relatively well, initially.
00:03:19Then we wanted to enter the land of the free, and we got some new requirements.
00:03:24Americans only want processing to happen within the U.S., and Europeans only want it to happen within the EU,
00:03:30and Australians only want it to happen within Australia.
00:03:33So we had to basically, well, here's a scaling graph.
00:03:37We had to move to multiple elastic searches, and what we actually did was we took the entire setup,
00:03:41and we just basically iterated over the set that is EU, U.S., and Asia-Pacific,
00:03:47and so we just had this, like, multiplied by three or four.
00:03:50Kind of annoying.
00:03:51Lots of overhead, but it got us to where we needed to be.
00:03:57Then the next iteration of the story is enterprise.
00:04:00So really big banks, the biggest law firms in the world, they have really annoying requirements.
00:04:07And number one they have is they'll ask for full physical isolation of all of their data.
00:04:12There's probably a little bit of a war on what physical, you know, isolation actually means,
00:04:17but essentially it means they want their own database.
00:04:19They also want customer-managed encryption keys,
00:04:22and what that means is they basically have a Key Vault thing where they have an encryption key,
00:04:26and they give us access to read the key,
00:04:29and we then use that key to encrypt and decrypt all of their data addressed.
00:04:33And what that gives them is they can just revoke our access to their key,
00:04:37and then we can't decrypt their data anymore, and so it's safe.
00:04:39And so in a way, that gives enterprises a lot of control over all of their data
00:04:42because they control, you know, the keys to reading it.
00:04:47So we moved from Elasticsearch to Postgres,
00:04:50and I imagine a bunch of you guys are like,
00:04:52why would you ever put your vectors into Postgres?
00:04:55It actually worked surprisingly well,
00:04:57and the reason that we did this was we were already using Postgres for OLTP workloads,
00:05:02and so we sort of already had to do this split of multiple Postgresies and multiple blobs,
00:05:09and so it was really easy for us to try to shift all of our search into Postgres as well
00:05:13because then we only have one system.
00:05:15So the setup here was pg-vector, specifically disk-ANN,
00:05:19ts-vector for the search, so not BM25,
00:05:21which was, you know, we lost a little bit of retrieval performance there,
00:05:25and what we'd do is we would partition the table
00:05:27where we would store all the document chunks.
00:05:29We'd partition it super aggressively, like 4,000 partitions,
00:05:31and then each project, we'd basically hash the project key
00:05:34and we'd bin-pack them into the partitions.
00:05:37That actually worked relatively well,
00:05:39but it was expensive,
00:05:42and search performance wasn't super, super good,
00:05:44and what happened was when we scaled a lot,
00:05:49everything just broke and exploded.
00:05:51And so what happened was the,
00:05:53basically, you can imagine that you have a bunch of projects,
00:05:56and some of them, you spin up a project, you work on it,
00:05:58and then you close it and you basically never go back to it again.
00:06:01And we have a bunch of those where, like, they never get queried,
00:06:03and we have a bunch that get queried all the time
00:06:04because they're super active projects,
00:06:06and when we packed them into partitions,
00:06:08the cold ones and the hot ones would land on the same ones,
00:06:11and the partitions would get really, really big,
00:06:13and so when we queried them,
00:06:14Postgres would pull the partition, put it into memory,
00:06:17we'd do the stuff,
00:06:18and then we'd query another partition and another partition,
00:06:20and it would essentially thrash the cache all the time,
00:06:22and what that meant was our latencies would spike.
00:06:26So we went from, like,
00:06:27search and ingestion P99 of 100 milliseconds
00:06:31into 20 seconds,
00:06:33which you can imagine is a really bad user experience.
00:06:37So then we went to TurboPuffer
00:06:39in about, you know,
00:06:40we were about, I think, 400 million documents,
00:06:42something like that.
00:06:45And what we did with TurboPuffer was
00:06:48we did one namespace per project,
00:06:50and the advantages of moving to TurboPuffer
00:06:53is we got BM25,
00:06:55real BM25,
00:06:56much better relevancy,
00:06:58much better latencies,
00:06:59and it was extremely simple to operate
00:07:01because we could just have a single TurboPuffer cluster.
00:07:04You know, we didn't have to have a bunch of different ones
00:07:06like with Postgres,
00:07:07and it would, since it's blob-based,
00:07:10it could just query the blobs that we had anyway,
00:07:12and so much lower cost,
00:07:14and it was extremely simple to operate,
00:07:16and we didn't have this problem with the partitions
00:07:17because if a project's not used,
00:07:19it's just in blob,
00:07:20and so it's really easy.
00:07:22And Simon can talk a bit more about
00:07:23why that works so well.
00:07:28Yeah, so Legora has some of,
00:07:31and Lego in general has,
00:07:32by the way,
00:07:33if Jake and I have similar accents
00:07:36and maybe even look a bit similar,
00:07:37it's because we're both Danish.
00:07:41TurboPuffer has a particular architecture
00:07:43that supports these kinds of very regulated environments
00:07:45really, really well,
00:07:46but in order to understand that,
00:07:48we have to understand
00:07:48what kind of search engine is TurboPuffer.
00:07:50Why is it different
00:07:51than the ones that they used in the past?
00:07:53Since the very beginning of TurboPuffer,
00:07:56the design has more or less been the same.
00:07:58There may be changes in the future,
00:07:59but the design has stood the test of time.
00:08:02When you do a write to TurboPuffer,
00:08:04we write directly to object storage.
00:08:07There is no, like, disk replication,
00:08:09there's no PAXAs,
00:08:10there's none of that,
00:08:12direct to S3.
00:08:14That's the fundamental trade-off
00:08:16in TurboPuffer, right?
00:08:17Hundreds of milliseconds.
00:08:18If you're, like, Shopify
00:08:19and doing inventory reservations
00:08:21for a Kylie Jenner flash sale,
00:08:22not going to work.
00:08:24Very, very good for search
00:08:25because generally,
00:08:26when you're doing search,
00:08:27doing a slow write is fine
00:08:29as long as the read performance
00:08:30is adaptable and good.
00:08:32So that's what happens on write.
00:08:33It just goes in the write-ahead log.
00:08:35You can imagine you write
00:08:361.json, 2.json, 3.json.
00:08:38Obviously, it's a database,
00:08:40so it's not JSON,
00:08:40but for illustrative purposes,
00:08:42that's what happens.
00:08:43And in the background,
00:08:44we build the vector indexes,
00:08:47the text indexes,
00:08:48the columnar indexes,
00:08:49and so on to satisfy the queries
00:08:50that Jake and other customers have.
00:08:53So then at query time,
00:08:54we can go in
00:08:55and then the query reaches some namespace,
00:08:59and namespace is kind of our concept
00:09:00of a table.
00:09:02You can think of it as a directory
00:09:03on S3 that's isolated
00:09:04from everything else.
00:09:06We go to the node
00:09:07that is most likely to have it.
00:09:08It could go to any node, right?
00:09:09It could go to every single node,
00:09:11and they're all read replicas,
00:09:12but it would be go
00:09:12with some affinity to the node
00:09:14that has the highest probability
00:09:15of having it in cache.
00:09:16We check the memory cache
00:09:17for any objects,
00:09:18NVMe, SSD cache,
00:09:19and then finally to object storage.
00:09:21Everything in TurboPuffer
00:09:23is optimized around
00:09:24doing as much work
00:09:25in as few round trips
00:09:26as possible, right?
00:09:26S3 has a P99
00:09:27on a, like,
00:09:29one megabyte blob size
00:09:30of around 200 milliseconds,
00:09:32so you want to do
00:09:33as few round trips
00:09:33as possible, right?
00:09:34Ideally, you do around three,
00:09:36and everything in TurboPuffer
00:09:37in the database is...
00:09:39Oh, I'm going to need
00:09:40your fingerprint.
00:09:41You got it.
00:09:42Everything in TurboPuffer
00:09:44is designed around
00:09:46minimizing the number
00:09:46of round trips.
00:09:47This is also amazing
00:09:48for modern disks.
00:09:49If you do a lot of concurrency
00:09:50in few round trips,
00:09:51you utilize them optimally,
00:09:53and everything in TurboPuffer
00:09:54is designed around this.
00:09:56So why is this so good
00:09:57for a company like Legora?
00:09:59Well, Oblix Storage Native,
00:10:00if you design it around
00:10:02the atomic unit of separation
00:10:03being the namespace
00:10:04or the table,
00:10:05every single table
00:10:06could be encrypted
00:10:07with a different key.
00:10:09Every single namespace
00:10:10could be in a different bucket.
00:10:11We have customers
00:10:12that have thousands of buckets
00:10:16that they have namespaces in
00:10:17so that their customers
00:10:18get the warm IT fuzzies
00:10:20of having the bucket
00:10:21in their own cloud account.
00:10:23They can also be encrypted
00:10:24with their own keys.
00:10:25You can share buckets.
00:10:27You can do whatever configuration
00:10:28that you need
00:10:29at the namespace level.
00:10:31You can re-encrypt
00:10:31with different keys.
00:10:32You can move them around,
00:10:33and you can re-encrypt
00:10:34with other keys.
00:10:36For Legora in particular,
00:10:37this was really important
00:10:38for this full physical separation,
00:10:40right,
00:10:41an encryption separation.
00:10:42All of the namespaces
00:10:43needed to be physically at rest
00:10:45with different keys
00:10:46and as separated as possible.
00:10:48S3, GCS,
00:10:49Azure Bobstories,
00:10:50they passed that,
00:10:51and the other parts
00:10:53of the hierarchy also,
00:10:54except the NVMe SSD cache
00:10:56because in the SSD cache,
00:10:57we consider that
00:10:58to be volatile like memory,
00:11:00but your customers did not.
00:11:03So what we did
00:11:05was that we thought
00:11:06we were going to implement
00:11:08encryption into the disk cache,
00:11:10but instead,
00:11:10we just disabled the disk cache
00:11:11and saw how it fared,
00:11:12and the performance
00:11:13of Turbo Puffer
00:11:14even without the disk cache
00:11:15with just the memory cache
00:11:16was so good
00:11:17that we just kept it that way
00:11:18for some of the Legora workloads
00:11:19where we couldn't have
00:11:21the disk cache for multi-tenancy.
00:11:22Turbo Puffer will support
00:11:23that in the future,
00:11:24but it just goes to show
00:11:25the natural point
00:11:28where Turbo Puffer
00:11:29allows these encryption
00:11:30and storage and separation
00:11:31to become fully multi-tenancy native.
00:11:35I'll hand it back to you
00:11:36on what happened then.
00:11:39And then,
00:11:40drum roll please,
00:11:41latencies look like this.
00:11:44Is my mic working?
00:11:46No?
00:11:48Could I,
00:11:49or I'll stop screaming
00:11:49really loudly.
00:11:51It speaks for itself
00:11:52if you can't hear me.
00:11:53Okay.
00:11:55Latency has improved
00:11:55in order and magnitude basically,
00:11:57and these are median latencies,
00:11:58so P99 were even better.
00:12:01So obviously,
00:12:02this is a huge thing
00:12:02when you're doing,
00:12:03I mean,
00:12:03one thing is if you're doing
00:12:04a single sort of rack style thing,
00:12:06but if you have an agent
00:12:06that does 20 queries,
00:12:08100 queries,
00:12:09these really, really add up.
00:12:12So that was on the project side.
00:12:13And then,
00:12:13a more recent thing
00:12:15is legal research.
00:12:18So legal research
00:12:19is a kind of a difficult problem,
00:12:22and the reason it's difficult
00:12:23is that the corpus
00:12:25is extremely big,
00:12:26so we're racing
00:12:28towards 10 billion vectors,
00:12:29and we're growing extremely fast.
00:12:31We also have quite high read,
00:12:34so QPS can spike a lot,
00:12:36because we do a lot of fan out.
00:12:38Like, if you do a sort of
00:12:40legal research query,
00:12:41we will fan it out
00:12:42to a bunch of different queries,
00:12:43and we'll keep going.
00:12:44And the reason we do that
00:12:45is we need this heavy filtering,
00:12:47because essentially,
00:12:48it's kind of like a graph
00:12:49for a few different reasons.
00:12:51Firstly, it's hierarchical.
00:12:53You know,
00:12:54you have cities,
00:12:55and you have counties,
00:12:56and you have states,
00:12:56and you have federal law,
00:12:57and it's the same
00:12:58all around the world,
00:12:59and so you need to respect
00:13:01that authoritative sort of hierarchy.
00:13:03There's also some temporal validity,
00:13:05so one judge might overrule
00:13:07a decision that's been made
00:13:08somewhere else,
00:13:09and you need to also respect that
00:13:10and figure that out,
00:13:11and then sometimes,
00:13:12there's even, like,
00:13:12a new regulation
00:13:14that has exemptions
00:13:15or special cases
00:13:16of an old regulation,
00:13:18and so if you're finding this one,
00:13:19you need to find
00:13:19all the other ones as well.
00:13:20So you can imagine
00:13:21that it sort of
00:13:22explodes the search.
00:13:26And so we started
00:13:27on Elasticsearch for this,
00:13:28but also moving to TurboPuffer.
00:13:31Elasticsearch just got
00:13:32extremely expensive,
00:13:33because we had to have
00:13:34everything there,
00:13:35but with TurboPuffer,
00:13:37we can basically
00:13:39take different jurisdictions,
00:13:41and we can make them
00:13:42namespaces in TurboPuffer,
00:13:43and that means
00:13:43some of them,
00:13:45here's an example
00:13:45where, like,
00:13:45you have the EU
00:13:46that gets queried all the time
00:13:47that's super hot,
00:13:48and some of them,
00:13:49let's say Danish law,
00:13:50because we're Danish,
00:13:51no one cares, really.
00:13:52It's such a small country,
00:13:53so, like,
00:13:53it doesn't really get queried,
00:13:54and so that can just stay on Blob,
00:13:55and that's fine,
00:13:56and because it's sort of
00:13:58a deep research-style workload,
00:14:00if there's 500 milliseconds latency
00:14:02to fetch that cold blob,
00:14:04that's okay.
00:14:05That's fine.
00:14:05It's not really a big problem.
00:14:07So the way that TurboPuffer
00:14:08is designed
00:14:09lends itself super well
00:14:10to this super long scale
00:14:11of, like,
00:14:12cold, weird namespaces
00:14:14and a few that are
00:14:14really, really hot.
00:14:17Yeah,
00:14:17and Simon wants to talk
00:14:18more about that.
00:14:19Yeah,
00:14:20so I was talking about
00:14:24why the company
00:14:25is called TurboPuffer
00:14:26at another talk here
00:14:27earlier today,
00:14:28but one of the other explanations
00:14:30of the name of TurboPuffer
00:14:32is that it's about puffing
00:14:34into the different memory hierarchies
00:14:35and really mastering
00:14:36when data should be
00:14:37in particular memory hierarchies.
00:14:39So you can think about it here, right,
00:14:40of something like the EU law
00:14:42might be more or less part
00:14:44of almost every one
00:14:45of the legal research queries, right?
00:14:47So that probably sits
00:14:49closer to NVMe SSDs
00:14:51than memory, right?
00:14:52The economics kind of change
00:14:53as you move up
00:14:54and down this hierarchy.
00:14:56In memory,
00:14:57you want things
00:14:58that are queried a lot, right?
00:15:00Then the economics
00:15:01of memory are great.
00:15:02NVMe SSDs,
00:15:03you can do a lot of things
00:15:04directly on them,
00:15:05but the economics change
00:15:06as you move up and down
00:15:07this boundary,
00:15:07the latency changes,
00:15:08and the way that the database
00:15:09is architected
00:15:10to take advantage of it
00:15:11in terms of round trips
00:15:12versus random
00:15:13versus sequential,
00:15:14all changes
00:15:14as you navigate this hierarchy.
00:15:16TurboPuffer is a database
00:15:17that is really designed
00:15:19around the memory hierarchy
00:15:20and all of the smarts
00:15:21and TurboPuffer
00:15:22is that all of these namespaces
00:15:23are puffed in and out
00:15:24of the cache.
00:15:26You can think of this
00:15:26as we want to spend
00:15:27as much time,
00:15:28have as much data
00:15:29pushed as far down
00:15:30in this hierarchy
00:15:31as possible
00:15:31to get the best
00:15:33performance cost ratios.
00:15:35So how does that apply
00:15:36to search?
00:15:37Well,
00:15:38for something like
00:15:38vector search,
00:15:39for example,
00:15:40there's two fundamental ways
00:15:41to do vector search.
00:15:42One is to navigate it,
00:15:44basically design a graph.
00:15:45The problem with a graph
00:15:47on something like
00:15:47opic storage or disk,
00:15:48again,
00:15:49we want to have things
00:15:50as far down
00:15:51that memory hierarchy
00:15:51as possible.
00:15:53The problem with a graph,
00:15:54this is not a graph,
00:15:54this is a tree,
00:15:55but in a graph,
00:15:57you have to navigate
00:15:58from the center
00:15:59of the graph.
00:15:59And then every time
00:16:00you navigate
00:16:01through these nodes,
00:16:02you're doing 200 millisecond
00:16:04p99 to S3, right?
00:16:06And so you're trying
00:16:07to shrink the diameter
00:16:07of the graph,
00:16:08you're trying to do
00:16:08all these tricks
00:16:09to make the graph,
00:16:10but fundamentally,
00:16:11you're at odds
00:16:11with the fact
00:16:12that graph
00:16:13is about a random
00:16:14sequential trade-off
00:16:15that you have
00:16:16in memory
00:16:16and in registers,
00:16:17but not further down
00:16:19to memory hierarchy.
00:16:20The way TurboPuffer
00:16:21does it is organize
00:16:21it into clusters, right?
00:16:23Vectors,
00:16:23you can think of
00:16:24in two dimensions
00:16:24just as point
00:16:25is a massive
00:16:25coordinate system,
00:16:26and we can organize
00:16:27them into clusters.
00:16:29TurboPuffer then
00:16:30creates clusters
00:16:30of clusters
00:16:31and clusters
00:16:32of clusters
00:16:33of clusters
00:16:33to essentially organize
00:16:35all of the vector data
00:16:36in a tree.
00:16:37You can basically
00:16:38think of TurboPuffer
00:16:39as a very, very complicated
00:16:41B-tree, right?
00:16:42Because it's a tree
00:16:43on this geometry
00:16:44of this entire space
00:16:45and the clustering of it
00:16:46in an approximate way.
00:16:48Now, the root centroids
00:16:50further up the tree,
00:16:51you can imagine,
00:16:51are part of every single
00:16:52time you search, right?
00:16:54We're always trying
00:16:55to figure out
00:16:56which clusters
00:16:56that we're in,
00:16:57and we're always looking
00:16:58at the upper levels
00:16:58of the tree.
00:16:59So they're going to be
00:17:00further up the memory
00:17:01hierarchy, right?
00:17:02Closer to the registers,
00:17:03almost all in DRAM.
00:17:05Now, the leaves
00:17:05that have all of the
00:17:07actual legal cases
00:17:08of whatever long
00:17:09document it could be
00:17:09images,
00:17:10all of that
00:17:11is probably going
00:17:11to be on SSDs
00:17:12with that single
00:17:13one millisecond round
00:17:14trip at the end.
00:17:15It doesn't make sense
00:17:15to have all that
00:17:16puffed into DRAM.
00:17:17This is fundamentally
00:17:19the cheapest way
00:17:20that you can run
00:17:21a database, period.
00:17:23So for something
00:17:24like Legora
00:17:25or even Web Search,
00:17:27which is in the
00:17:27hundreds of billions
00:17:28or tens of billions
00:17:29depending on how much
00:17:29of the web you've scraped,
00:17:30this is fundamentally
00:17:31the cheapest way
00:17:32to do it,
00:17:32and we have customers
00:17:33that are indexing
00:17:35massive parts
00:17:35of the entire web
00:17:36into TurboPuffer,
00:17:37which is really also
00:17:38a part of what
00:17:39legal research is.
00:17:41Full text is also
00:17:42really respectful
00:17:44of the memory
00:17:45hierarchies.
00:17:46The way that text
00:17:46search works
00:17:47is essentially
00:17:48you can think of it
00:17:48as a hash map.
00:17:49You have a big document
00:17:50and then you take
00:17:51every single one
00:17:52of the tokens
00:17:53and you put them
00:17:53into the key
00:17:54in the hash map.
00:17:55The value in the hash map
00:17:56is some set
00:17:57with all of the document
00:17:58IDs that has that term.
00:18:00So then if you search
00:18:00for New York population,
00:18:02you're finding
00:18:02those three places
00:18:03in the hash map
00:18:04and then you're
00:18:04taking the three sets
00:18:05and doing an intersect
00:18:06on the sets.
00:18:08While you're intersecting,
00:18:09you're also trying
00:18:10to do some kind
00:18:11of scoring, right?
00:18:12A document
00:18:12that has York in it
00:18:14is probably more valuable
00:18:16than a document
00:18:17that has New in it
00:18:18because York
00:18:19is a more rare word.
00:18:20When people say BM25,
00:18:21this is the scoring
00:18:22that they're referring to.
00:18:24The art of full text search
00:18:26is one,
00:18:27we want to minimize
00:18:28the number of round trips.
00:18:29So first,
00:18:29you download the parts
00:18:30of the dictionary
00:18:30that are relevant,
00:18:31round trip one,
00:18:32maybe a round trip one
00:18:33before that
00:18:34to index into
00:18:35where the parts
00:18:35of the terms are.
00:18:37And then the second round trip
00:18:39is to get these massive lists.
00:18:41Try to make the list
00:18:42as small as possible
00:18:42by compressing them.
00:18:44But also,
00:18:44while you're doing
00:18:45the text search,
00:18:46you're trying to minimize
00:18:47the amount, again,
00:18:48of memory bandwidth
00:18:49that you want
00:18:49to intersect these lists.
00:18:51You can probably imagine
00:18:52that at some point,
00:18:53there's a point
00:18:54where you've seen
00:18:55so many documents
00:18:56with Population and York
00:18:57that have much higher scores
00:18:58that documents
00:18:59that just have new in it
00:19:00are irrelevant anymore.
00:19:01This is like a mega crash course
00:19:03in how text search works.
00:19:04And counterintuitively
00:19:05to most people,
00:19:06text search at web scale
00:19:08is more difficult
00:19:10and more competitionally expensive
00:19:11than doing vector search.
00:19:13I'll hand it over to you.
00:19:16Cool.
00:19:17So key learnings
00:19:18from what you heard today,
00:19:22retrieval is extremely important
00:19:23to Legora.
00:19:24It's key to legal reasoning.
00:19:27Turbo Puffer really excels
00:19:29when, for us,
00:19:31because it makes it
00:19:32extremely easy to operate.
00:19:34We have 70 plus tenants.
00:19:35We have 100.
00:19:36We have 200 tenants.
00:19:37You know,
00:19:37if we had to have
00:19:38separate elastic search databases
00:19:39for each of these,
00:19:40it would be just hell.
00:19:42But we can do this natively
00:19:43with Turbo Puffer
00:19:43with data residency
00:19:45and CMEG,
00:19:46et cetera, et cetera.
00:19:47And then it's
00:19:47extremely cost efficient
00:19:48generally when you have
00:19:49these types of workflows
00:19:51or workloads that we do
00:19:53where there's a long tail
00:19:54of cold indices,
00:19:56basically,
00:19:56that you don't need
00:19:57to query so much.
00:19:58And you're okay
00:19:59paying the small latency
00:20:01cost for it.
00:20:02So now with Turbo Puffer
00:20:05and four seconds to go,
00:20:06now we can focus
00:20:07on making Legora,
00:20:08we can focus on the product,
00:20:09making it really, really great,
00:20:10and not on scalability
00:20:10and infra.
00:20:11And also David,
00:20:12our CFO,
00:20:13is really happy
00:20:13about the cost.
00:20:14So it's great.
00:20:15Thanks, everyone.
커뮤니티 글
아직 글이 없습니다. 이 영상에 대한 첫 번째 글을 작성해 보세요!
이 영상에 대해 글쓰기