Large clusters for small models — Daniel Svonava, Superlinked
AAI Engineer
Computing/SoftwareSmall Business/StartupsInternet Technology
Transcript
00:00:00Reviewer: Denise RQ
00:00:12All right.
00:00:14I think you guys can hear me.
00:00:15I can certainly hear myself.
00:00:19Whoever came closer gets a T-shirt.
00:00:21I meant it.
00:00:22There's like a bag full of T-shirts over here.
00:00:25And also for questions.
00:00:26Maybe there will be some questions at the end.
00:00:28If you ask a question, you'll get the T-shirt as well.
00:00:31And if you can guess what is on the background of this slide,
00:00:35you get the T-shirt as well.
00:00:39Any guesses?
00:00:42What does that visualize?
00:00:44This picture in the background.
00:00:49No?
00:00:50Anybody has seen a transformer model?
00:00:55Yeah, positional encoding.
00:00:57Very good.
00:00:58You get the T-shirt, sir.
00:00:59All right.
00:01:00So today we'll discuss basically small open source models and how they're pretty good now and how they create unique challenges when you want to serve a bunch of them in your own cloud.
00:01:15Everything we'll discuss is kind of open source.
00:01:19Do it yourself.
00:01:20This is the kind of stuff you can just, you know, run a command and own the stack.
00:01:26So there is no proprietary, you know, pieces of the puzzle here.
00:01:31Let's get this underway.
00:01:33Okay.
00:01:34Okay.
00:01:35This works.
00:01:36Okay.
00:01:37So small models.
00:01:38What do we mean by small models?
00:01:40You know, depending who you ask, the way I think about it is basically models that you can run on two, three generations old NVIDIA hardware.
00:01:49The whole model fits into one GPU and therefore they are easy to serve.
00:01:56Those GPUs are available and they are affordable as well.
00:02:01And then most people think, okay, small models, there'll be some kind of trade-off in terms of, you know, quality of the results.
00:02:10And hopefully I'll be able to do a good job in this talk to convince you that actually for specific tasks, you can be at frontier or beyond frontier performance and get all the other obvious benefits, right?
00:02:22Orders of magnitudes of kind of cost savings and potentially quite big latency or throughput improvements, of course.
00:02:35So this is kind of one of the charts we like to show.
00:02:38This is the artificial analysis intelligence index over time.
00:02:42And what they typically don't show you is that there is like a breakdown of the open source models you should think about, right?
00:02:49There is the GLM 5.2 and so on, those kind of frontier open source models with let's say 750 billion parameters.
00:02:57But then there are the small open source models kind of trailing the big ones and trailing the frontier.
00:03:04You can see the frontier is kind of getting diminishing returns these days and the small models are catching up, right?
00:03:10So you see this kind of convergence, saturation on top and kind of growth of the small models.
00:03:16And, you know, let's say QN 3627B somewhere around the performance of GPT 5.1.
00:03:23So if you have a workflow, if you have a pipeline that, you know, can run with GPT 5.1,
00:03:29now you can move it to a small model and, you know, get all the benefits we discussed.
00:03:35So small models, not done anymore.
00:03:38Now, it is also about how you use the small models, right?
00:03:44So you can't just read that 27 billion parameter QN 36 as your kind of totally generalized, I can prompt you to do anything kind of model.
00:03:53No, you need to adopt the approach where you basically figure out slice of tasks from the generalized model workload.
00:04:00And then per task, you figure out which model in the open source fits the task the best.
00:04:05You run some evals, maybe some adaptation we'll discuss.
00:04:08And then, you know, that's how you kind of reach the right quality to actually push this into production.
00:04:15So here is some example of a contract review agent that uses, you know, nine different models.
00:04:20This is the kind of shape that you will see in your workloads and your agents as you move to using small models for your setup.
00:04:28You'll start to see that, okay, instead of kind of hammering one API with a bunch of different requests or one model,
00:04:36you'd rather use a fleet of models.
00:04:38And then your problem is, okay, how do I serve all of these different things in a way that my infra people don't go crazy, right?
00:04:45And this is just one of the agents that you might be running.
00:04:48And there might be, you know, 10 of these in your company.
00:04:51So how do we sort of, you know, that's the kind of expansion of infrastructure scope, let's say.
00:04:57Now, all of those different tasks that I mentioned, there is an open source model that's sitting there waiting to be used.
00:05:05From, you know, OCR to question answering on top of documents to labeling images, generating SQL, you know, reviewing code.
00:05:14There are open source models fine-tuned and trained for those tasks.
00:05:19You know, if you use an open source model that's trained to do OCR on receipts in Vietnamese, that project has seen the most receipts in Vietnamese, right?
00:05:28There's somebody who, like, took the time to gather as much data as possible.
00:05:32And on that task, that model will outperform pretty much anything else.
00:05:35And there is, you know, hundreds of thousands of models on hugging face that look like that, right?
00:05:41So it's just, it's all sitting there and it's all free, basically, mostly quite permissive licenses.
00:05:46So the models exist, you know, that's not the bottleneck.
00:05:50And, you know, we have been talking about, like, open source AI since 2024.
00:05:56And it's so far still not really happening.
00:05:59And to the extent it's happening in companies, it basically equals, like, open source AI equals AWS Bedrock.
00:06:05Except when you look at the model catalog in Bedrock, it's, like, very, you know, restrained in model types that are available.
00:06:14These models are old, often, you know, two, three years behind the state of the art.
00:06:19And when you do any kind of fine-tuning in Bedrock, you don't actually own the fine-tuned or trained artifacts.
00:06:25So you can't, you know, use it as an actual advantage in your business.
00:06:29It kind of stays serving from the Bedrock infra.
00:06:32So this on the proprietary now, if you do small models serving on open source infrastructure, VLLM, SGLang, different solutions,
00:06:41just know that these things are not tuned for any specific model or any specific hardware model combination.
00:06:48You'll have to do the tuning, right? This is the do-it-yourself.
00:06:51All of these tools ship with guides on how to actually do the tuning, the parameter sweep, tailoring to your traffic, and so on.
00:07:00This is a kind of open-ended research project every time you try to adopt one of these tools.
00:07:05So this is not really something that, sort of, you take it and it's like an engineering project,
00:07:10and a week later you have a high-performance surveying infrastructure.
00:07:14It doesn't work like that.
00:07:15And that's kind of the typical problem with open source tools, right?
00:07:19It's kind of, like, a little bit too much do-it-yourself.
00:07:21And then, on top of this not being kind of pre-tuned for small models, the small model workloads and traffic that uses a bunch of different models
00:07:32kind of flips the equation for inference kind of clusters, right?
00:07:37So normally, when you try to serve one big model, your problems are, how do I share that model across multiple GPUs?
00:07:44How do I have a router sitting on top that understands the state of all these workers, you know, the KVCache state and so on,
00:07:51and then makes a top-down routing decision of, okay, this request goes to this worker or this group of workers and so on, right?
00:07:59It's a very top-down setup.
00:08:01But if you have small and fast requests, and you have many of them, this sort of top-down routing becomes the bottleneck, right?
00:08:09Because the router has a little bit obsolete version of the worker state, and it's just really hard to saturate the workers
00:08:16if you have that kind of upfront decision on top that has to get it perfectly right in terms of, you know, balancing the local queues on each of these workers.
00:08:26Because there is many small requests, right?
00:08:29And, you know, like, we have experimented with the VLM and SGLang routers for small models and this sort of traffic,
00:08:37and it's very hard to get your GPU utilization beyond 20%, 30% under constant load.
00:08:44And the problem is that those batches are just not correctly sized, basically, because you have that routing bottleneck.
00:08:51And then the third problem is that with small models, you benefit a lot from LORAS and just model adaptation in general.
00:08:57And so the traffic that you have to serve, you know, contains, you know, people coming to you and saying,
00:09:03"Hey, I have 10 LORAS. How do I, you know, use this with our serving stack?"
00:09:08Or, "I have this custom fine-tune I made last night. You know, I want to serve this in production."
00:09:14And this conversation between the AI engineer and the infrastructure person in getting those, you know, LORAS up there,
00:09:21custom models up there, that's the thing that takes time.
00:09:24And basically, that's like the main killer in organizational velocity is talking, right?
00:09:30Like ideally, you would want the infrastructure engineers to do their job, and you would want those AI engineers to do their job.
00:09:37And they don't have to talk to operate on the day-to-day mode.
00:09:41So, you know, they're not blocking each other, basically.
00:09:45And this kind of model adaptation desire around small models kind of breaks that and creates a lot of back and forth.
00:09:51And that's a problem, right?
00:09:54So these are some challenges related to, okay, we have a bunch of small models.
00:09:58How do we have a cluster? How do we serve this efficiently?
00:10:02So we have been playing with this problem for a while.
00:10:05I'm Daniel, actually, from Superlinked. I kind of skipped the intro.
00:10:09So we are, you know, VC-backed company out of SF.
00:10:13And we have been building AI-powered search and document processing systems and agents for the last couple of years.
00:10:20And our main pain point has always been inference, specifically these problems that I have described.
00:10:26And so we have iterated and iterated and explored different topologies for clusters for running, you know, large, wide fleets of small models in different environments.
00:10:37Because sometimes you need to deploy together with some platform in some environment where who knows what is available there.
00:10:44You know, the small models make it easier because in whatever environment you can get some L4s or some kind of small GPU quota is much easier.
00:10:52So this is kind of -- I'll describe a little bit about the topology of the cluster that we have kind of converged to.
00:10:58And by the way, this whole thing is Apache 2.0, completely open source.
00:11:03You guys can just take it and wrap it and now you are an inference startup.
00:11:08This is open source from kind of the control plane all the way down to the thing that runs on the GPU.
00:11:14So we didn't pull any punches.
00:11:18And the topology is basically there is a gateway.
00:11:21And instead of having a router that kind of pre-decides what goes where, there is a gateway that parses some of the requests and attaches some metadata to the request.
00:11:30Inserts that request into a shared queue and into some side channels.
00:11:36I'll go a little bit into that.
00:11:37And then the workers pull from that centralized queue instead of kind of pushing the data down to the workers.
00:11:44And this way they can saturate themselves better.
00:11:47And then the worker setup -- I think I have a slide for that -- will describe how we basically absorb the complexity of different model architectures
00:11:57into kind of a coherent set of workers that, you know, don't have like competing Python requirements and stuff like that.
00:12:03So that's kind of the overall topology.
00:12:06And this is kind of life of a request.
00:12:11So maybe just I'll call out a couple of things from here.
00:12:15You know, one of the things we don't like about the OpenAI kind of API standard is the base64 encoded kind of JSON.
00:12:25Not good for small models, not good for high throughput.
00:12:28So we use a message pack throughout, like a binary format.
00:12:31This way we can also push all the multimodal data through the actual API gateway.
00:12:37So there is no, like, hey, you know, binary data over here and then request over here.
00:12:42And then the cluster needs access to your cloud storage to start loading some binary data, images or videos.
00:12:48We kind of encode it all and we push it through the gateway.
00:12:53And then the gateway kind of separates some of these heavier pieces to not clog the internal queue and defers it on cloud storage kind of in-flight while the request is in queue.
00:13:02So it kind of splits up some of these requests that are, let's say, over a megabyte and then uses cloud storage in the backend.
00:13:09But as a user, you push all your bits and bytes into the API layer and it's kind of clean interface because of that.
00:13:19So basically the whole stack is REST, so gateway REST, the worker is REST, and then over a socket, locally, it kind of attaches to different runtimes.
00:13:30And we have basically PyTorch, Candle, and SGLang as a runtime.
00:13:35And then when we do the optimization, I'll kind of go into that on how we make sure that whichever runtime we are using
00:13:43and whichever code is running in that runtime is the most efficient one.
00:13:47We have an auto-research loop for that, basically.
00:13:50But yeah, so life of a request kind of looks like that.
00:13:54And like one tidbit is that you really want to make sure that the gateway that's kind of the first thing that's hit by the request doesn't do too much work.
00:14:05Because then it becomes a bottleneck, right?
00:14:07So you don't even want to parse the whole request.
00:14:09You want to be able to kind of look at the packets and figure out the general shape of what's coming, do the annotation,
00:14:16and then you have the workers, however many workers you have, hundreds of GPUs that look at the queue state and then pull from there.
00:14:25And the queue use NATS Jetstream, and that thing can do, you know, million requests per second.
00:14:32Like, that's very hard for that to become a bottleneck.
00:14:35So, yeah, like ideally you don't want to serialize, deserialize as you go through all of these different components.
00:14:42That's basically the kind of obvious thing.
00:14:45This is a little animation that shows the idea behind the centralized queuing, right?
00:14:51So instead of the top-down router trying to, you know, fill in the local queues just right, which is basically impossible, you know,
00:15:01the whole idea is, hey, can we somehow centralize the queuing and can the workers rather pick up the task of forming their own batches
00:15:09batches with their own prediction of the cost of the batch and then, you know, become much more efficient?
00:15:15Now, one tidbit and kind of side note, once you kind of start working on these things, you realize that it's actually really hard to predict how many things to pick up from the shared queue for the batch to be really, like, really the optimal size.
00:15:30And so you would want some mechanism that sort of allows you to put some things back into the queue.
00:15:35If you figure out, oh, like I pulled a little bit too much.
00:15:37And that's a network hub, right?
00:15:39So that's a problem.
00:15:40And we have special optimization for that for machines that have multiple GPUs locally, right?
00:15:46So there is additional kind of machine local queuing element that takes advantage of the fact that the local processes that run on the multiple GPUs on one machine can kind of negotiate with the queue a little bit back and forth,
00:15:59which over the network, you know, there is like milliseconds extra that that would add.
00:16:03And so we don't do it over the network, only when we co-locate the workers on multi-GPU machines.
00:16:11And, you know, I mean, we are not talking about, like, 5% differences here, right?
00:16:16So, like, you centralize the queue and now you get double the throughput of the cluster.
00:16:19So this is significant.
00:16:22I mentioned three different runtimes.
00:16:25So, basically, it's either, you know, we write, let's say for models that are encoder only, we write the PyTorch code.
00:16:33And we kind of optimize it and we have an auto-research loop that optimizes it.
00:16:38Same for Candle.
00:16:39We started to play with Candle not too long ago.
00:16:42We still can't get it to perform anywhere near the PyTorch performance.
00:16:45So it's a little bit more of a research project.
00:16:47It's just the dependency, like, you know, the worker Docker image with PyTorch is like 12 gigabytes.
00:16:54And the worker basically binary, statically linked binary with Candle is maybe like 10% of that, right?
00:17:01And if you care about kind of waking up from the cold state and loading these images on a bunch of different machines,
00:17:08the, you know, going from 12 gigs to a gigabyte or something like this makes a huge difference.
00:17:13So that's kind of the motivation behind Candle.
00:17:15It's just the, getting the same performances from PyTorch is really hard.
00:17:20And then SG-Lang we have there as a kind of go-to baseline.
00:17:25Like we should perform as, at least as well as SG-Lang with the optimal tuning of all of those parameters that I mentioned that you have to do the tuning.
00:17:34Here are some numbers.
00:17:37So for example, when we wrap SG-Lang with the socket and with our kind of Rast sidecar,
00:17:43actually we can improve on the bare SG-Lang performance just because we kind of do something on the batching side that natively SG-Lang doesn't do.
00:17:54And probably you can make it to do that.
00:17:57If you do like, if you develop custom plugins into SG-Lang and stuff like that,
00:18:00like probably you can match our performance because, you know,
00:18:04you can just push the same logic into the SG-Lang core server.
00:18:07But now you are developing custom code that only works with SG-Lang.
00:18:11And the whole lesson here from small models is that the runtimes are super diverse, right?
00:18:16You don't want to necessarily get stuck with any one particular runtime because there is, you know,
00:18:22we have, I think on the order of 50 different adapters now that we parametrize for the different models.
00:18:28And so you need to somehow deal with this kind of underlying complexity.
00:18:32And it's probably not by building a bunch of plugins for one specific runtime.
00:18:36It's probably some kind of abstraction, which in our case is this Rast sidecar concept and then the socket.
00:18:47Now I'll talk about a couple different numbers, but in terms of like language around benchmarking, you know,
00:18:53the knee is this concept of like when you ramp up traffic on the server,
00:18:57when you sort of request more and more throughput from it,
00:19:01and it gives you more and more throughput, that's when you go kind of linearly up.
00:19:05And then at some point you hit this point where you kind of ask for more and more is not coming.
00:19:09So you kind of flatten out and the latency goes up.
00:19:12So we call that the knee and it's like a useful concept in benchmarking,
00:19:17because that's kind of the point of saturation, right?
00:19:20That's kind of the maximum performance without hurting latency.
00:19:23So just to give you some ideas of what is possible on relatively small hardware, right?
00:19:32And different types of small models.
00:19:34So this is measured on the RTX Pro 6000.
00:19:37We kind of work with NVIDIA L4, you know, A100s, RTX Pro 6000, H100, that sort of range.
00:19:46Again, those GPUs are much more readily available, kind of on demand in any cloud, basically.
00:19:52Most continents have quota, you know.
00:19:56And on this kind of stuff, you can basically get, for embedding models,
00:20:01even up to, let's say, hundreds of millions of parameters,
00:20:05you can get hundreds of thousands of tokens per second encoded into the embedding, right?
00:20:12So imagine you are sitting there now, like hitting your text embedding tree on OpenAI API.
00:20:17Instead you could be like having one GPU and push half a million tokens per second into the thing and get the vectors out.
00:20:26Right?
00:20:28Like, is this like connecting, right?
00:20:31You have half a million tokens that you are pushing into a single GPU that's not even that big per second,
00:20:38when you are getting out vector embeddings for your search system, as opposed to like pushing all of that into a managed embeddings endpoint somewhere and paying like orders of magnitude more money.
00:20:49Right?
00:20:50And you can get latencies like, you know, low tens of milliseconds for these calls, right?
00:20:55Like if you use, you know, cohere, open AI, APIs, and so on, these are hundreds of milliseconds, right?
00:21:02And this is not rocket science.
00:21:03You know, you can have just like massive cost saving, massive latency improvements, and relatively easy operation with like handful of GPUs and some infra around them, right?
00:21:17So this like really low hanging fruit, if you start anywhere with open source models, small models, embeddings are like no brainer, right?
00:21:26But it doesn't end there.
00:21:27So let's say you want to look at named entity recognition.
00:21:33You want to look at, let's say multi-vector search, even generation, right, of text or structured outputs and so on.
00:21:42You can be getting, you know, thousands of tokens per second output from, you know, task specific generative models as well per, like, let's say half a thousand per second for one GPU there at the bottom.
00:21:58And so let's say you are generating synthetic data, you are generating annotations for your fine tuning, for your evals, you know, don't do that on a managed endpoint.
00:22:11That's a perfect task because you have it kind of under control.
00:22:14You can survey the quality.
00:22:16That's a perfect task for open source model on your own infra.
00:22:20And then, you know, like, if the infra you have around those GPUs is like reasonable, you'll get linear scaling with the number of those GPUs.
00:22:31Now, another sort of idea, if you are into small model serving, is that you don't, you know, normally, you have kind of worker pool per model, right?
00:22:44You have a set of workers, set of nodes, they have GPUs, you kind of bring those up, you preload the models.
00:22:51The models load for tens of minutes because they are hundreds of billions of parameters.
00:22:55And so you are happy, okay, they finally loaded, now I have a worker pool.
00:22:59This mentality doesn't really work with small models.
00:23:02Yeah, yeah, quickly.
00:23:04How, how, what's the time left?
00:23:06Six minutes over.
00:23:07Oh, six minutes over.
00:23:08Okay.
00:23:09All right.
00:23:10So pack models on the same GPU is faster.
00:23:12This is a story of how you still want to pin some models, but you want to also do basically lazy loading and eviction as a kind of function of memory pressure.
00:23:26You want to figure out how to combine the two.
00:23:28There is a little bit about kind of auto research.
00:23:32We have auto research loops for adding support for new models and for their performance.
00:23:39We build a lot of internal tooling to do the measurement to feed into those auto research loops to basically push the numbers forward.
00:23:47And maybe perhaps most importantly, when we ship support for a model, it has all the tuning done, right?
00:23:53So there is no, okay, let's do a parameter sweep.
00:23:55We bundle basically a config for end-to-end the whole cluster.
00:24:00This is a setup for the auto research loop.
00:24:03There is like a meta loop that builds the harness that then runs the loop.
00:24:08And there is a dashboard on top that helps you understand how it works.
00:24:12We have custom UIs for that.
00:24:15And one of the outputs of that was a Lora that took 80 cents to train and it improved quality of retrieval on German legalistic as a proof of concept by 18%.
00:24:27And that's it.
00:24:29So small models are good.
00:24:30They are relatively easy to serve.
00:24:32They are actually much cheaper, faster.
00:24:35It's S smart.
00:24:36And that QR code goes to the GitHub repo of our cluster that I just described.
00:24:43Give us a star.
00:24:44And happy self-hosting.
00:24:45Thank you.
00:24:46Thank you.
00:24:47Thank you.
Community Posts
No posts yet. Be the first to write about this video!
Write about this video