Training Taste — Thais Castello Branco, Taste Labs
AAI Engineer
Computing/SoftwareSmall Business/StartupsInternet Technology
Transcript
00:00:00Okay, amazing. It's great to meet everyone. I'm Thais. I'm the founder of Taste Labs. For those
00:00:19of you who don't know us, we came out of Stout a few weeks ago, and our whole mission is basically
00:00:23how do we end AI Slop? It's my personal enemy. And so we really believe that to solve this problem of
00:00:31Slop, we have to like decode subjective domains. There's been so much effort being put into getting
00:00:36models and agents amazing at things like coding and math, and it's time that we put all that same
00:00:41effort into making them great at things like design and writing. And so design is this first pillar that
00:00:46we're starting with, and it's been incredibly exciting. We work primarily in two ways. So we
00:00:52work a lot with the Frontier Labs on how do we evaluate their models, understand where they're
00:00:57breaking, understand what could be better about them, and then construct the right either post-training
00:01:01data or our environments to basically fix that problem. And part of this is like how do you take
00:01:05something as fuzzy and large as design and break it down to a level that you can identify what is
00:01:11best solved through each method? What are elements of design that are almost like, once you kind of
00:01:15boil down the problem, become so specific that they almost become deterministic. So for example, if you're
00:01:20trying to train a model to be good at selecting color palettes or have contrast or alignment. Those
00:01:25are things that if you define the problem in the context in a specific enough way, you can get to
00:01:30an answer that's like pretty objective or that at least most experts would agree to. But maybe other
00:01:34things like aesthetics, you naturally will see this expert disagreement. And so then you want to lean
00:01:39on to things that are closer to data. So anyway, we spend a lot of time thinking about all those problems.
00:01:43But on the other side is also, without even touching the model layer, right? How do we actually help
00:01:48agents and app layer companies produce better things? And there's a lot that goes into that,
00:01:53right? You have these different sets of problems at the application layer, because you're using an
00:01:56off-the-shelf model that tends to collapse in terms of style, tends to collapse to the mean. So how do we
00:02:02force that creativity back to the system? How do we avoid these patterns of slop, which we'll talk about a lot
00:02:07today? How do you understand like user preferences or brand preferences preference so that you can
00:02:13maintain adherence to that style? So there's lots of things that actually need to be solved as context
00:02:19or judgment or verification at the app layer, which is why we kind of work across both.
00:02:25But maybe I'll start with more of a philosophical question of like, how do you define something that
00:02:29is great? Like, how do you define greatness? And for something like math, it's easier, right? Because
00:02:35there's kind of one objective answer. And great is the same as correct. But then for something like
00:02:40writing or design, it's much harder, right? Like, how do you define what's like a great tweet or what's
00:02:45a great art piece or what's a great website? I don't know what's the last time that you interacted with
00:02:52a poem or walked into a coffee shop and for some reason it kind of like hit different and it felt
00:02:56very special. But probably it's a combination of things that it felt very unique. It felt almost a little
00:03:02different. It kind of called your attention. It felt like it was made with a lot of care and attention
00:03:06to detail and craft. And it almost had this sense of like authenticity. And I think that's a lot
00:03:12of what AI is missing today. It's like, how do we take things that are not necessarily average, right?
00:03:16How do we produce things that are purposely like out of distribution? And slop is kind of the opposite
00:03:21of that, right? I think it is hard to define what is great sometimes, but I think it's pretty easy to
00:03:26define what is slop in the sense that most people would agree. I think the sense of like repetition,
00:03:30of kind of soullessness, is something that all of us feel right now when using AI. And I think it's
00:03:34quite magical, by the way, that AI has gotten to a point that any human on the planet that is not
00:03:39even a designer, that is not an engineer, can click a button and suddenly make an entire PowerPoint or
00:03:44make a website or make a web app. That's pretty cool. But it comes with consequences, right? It comes with
00:03:49consequences of suddenly now the cost of generation is basically going to zero. But the average person
00:03:56hasn't necessarily honed their taste. Like think about the amount of effort and work that a designer
00:04:01puts in throughout their life to like build up their taste, right? Like there's all this process of like
00:04:06getting exposed to many things and learning to like spot patterns and learning to develop a point of
00:04:11view and like kind of do things in a courageous way that maybe are a little bit against the norm.
00:04:15Learning what not to do and how to like have restraint and that's very hard. Like the average person
00:04:21doesn't necessarily have the the time nor the skills to go and develop taste in everything,
00:04:25let's say in design. And so I think it would be a bad case scenario for us to just like be like,
00:04:30okay, the way to fix slop is for everyone to have taste because I don't think that's necessarily
00:04:34realistic. I think how do we how can we understand this better so that we can make even for the average
00:04:39person the ability to create something great and to understand maybe their own taste
00:04:44easy, more easy. So that's that's a lot of what we're we're focusing on. So yeah, I think this
00:04:50phenomenon of slop, by the way, is not new. If you were in the internet as social media emerged,
00:04:55you probably saw a lot of slop before that. But I do think that AI has been this kind of like
00:05:00accelerating force right of like being able to create things very easily with a click of a button
00:05:04and the like thoughtlessness around it. And there's kind of these three characteristics that I would say
00:05:08repeat and slop. So a repetition. So you start seeing the same thing many, many, many times.
00:05:15The second is lack of fit, which I actually think is very related. So
00:05:19fit is kind of this ability for something to feel correct for a specific context, right,
00:05:23for a specific moment in time, for a specific person. But suddenly, if you have repetition,
00:05:27and let's say one person asked for a website for their pet shop, and the other one asked for a website
00:05:32for their finance firm, and somehow those designs converge and look the same, that's quite odd,
00:05:38right? Like if that wasn't, if you were actually crafting that with care, that wouldn't, you
00:05:42wouldn't converge necessarily on those things. And so this lack of fit and lack of understanding of
00:05:46context is actually a huge problem that like leads to slop. And the third is maybe low intent,
00:05:51which is probably a mix of, yeah, you're gonna have a bunch of people prompting really quickly,
00:05:54and maybe just wanting to one shot something. But I think there's actually this like intent
00:05:58interpretation piece that's missing in the systems that we're building. Like,
00:06:01how can you help your user, right? Like, how can you help them better understand the intent that they
00:06:05have, so that you can add more color and add more context onto what you're trying to create?
00:06:12Okay, and I'm a big believer, by the way, that you, in order to fix something, you first have to
00:06:17measure it, and you first have to understand it. I think that's exactly why we're so focused on,
00:06:21like, how do we turn these domains into something a bit more verifiable, so that we can attach a
00:06:27measure to it. So you'll go on a little bit of a research journey with me here now, but
00:06:31we basically wanted to figure out, can we measure slop? Like, can we actually measure this quantitatively
00:06:36and spot this? And what does that like look like? So we analyzed over two million websites from the
00:06:43past like 10 years, kind of like way back machine style to try to understand all the trends across
00:06:48like design? How is the internet changing? How is like design changing over time? And two things were
00:06:55interesting. And we also, by the way, then kind of synthetically generated a set of design websites,
00:07:00so we could kind of like compare. Like, how does human-made sites compare to AI generated ones? And
00:07:07there were a few things that were interesting. So one was that you already kind of saw a bit of a collapse
00:07:13on the internet before even AI. So you saw kind of the internet becoming more homogenous, using more
00:07:18similar color palettes, using more similar layouts, which is probably a function of more,
00:07:24I would say this kind of trend spreading more quickly, let's say. But with AI, I think you saw
00:07:28this repetition happening a lot more and being almost more like identified kind of regardless of
00:07:33context. So even in completely different buckets, you saw patterns that were very similar. So we built
00:07:38this, I call this probes, but basically we did two things. So we did this like pattern mining on all
00:07:44this data to understand like, what are features that we can extract from all these sites? What are
00:07:47all these characteristics that we can make more objective, right? Colors, typography, layout, audience,
00:07:53like how can we like distill this down into things that become almost like structured? And then how do we
00:07:58train up these like probes? So think of these as like baby classifiers. Like how do we train the ability to
00:08:03spot this one characteristic? And for all these slop sites, we identified, we started identifying like,
00:08:09what are the probes that basically mean this site is very likely to be AI slop? And especially when
00:08:16you start combining them and you see the frequency of multiple of these happening at once, it became
00:08:20very likely that you could actually like measure and predict slop. And we saw a super high, basically,
00:08:26ability to do that prediction, which was really cool to see. This performed better, by the way,
00:08:30than like most LLM as a judge methods of like asking an LLM to like judge if that is a great human
00:08:36quality versus like AI generated slop. So that was pretty cool to see. And I think it kind of shows
00:08:40this pattern that we see in AI really being an actual quantitative thing that we can see in slop,
00:08:47which I find really cool. But obviously, we don't want to stop there, right? We don't want to just
00:08:51measure slop. We want to also solve it. And so there's a few, I think I mentioned this before,
00:08:56but like the, as the cost of production basically goes to zero, I think the thing that becomes
00:09:01expensive and matters more than ever is judgment. I don't even want to use the word taste here,
00:09:07is judgment. I think it's this ability to discern what's right, is this ability to break down a
00:09:11problem so that you can actually understand it and create solutions for it. And so, yes,
00:09:16there's a side of judgment that is human judgment that I actually think is more valuable than ever.
00:09:19But there's also this side of like, how do we build the right tools and systems to like
00:09:23fix pieces of this problem, right?
00:09:27So yeah, how do we fight slop, my enemy? And by the way, I think there's a lot of conversation going
00:09:34around how do you fight slop at the model layer? Like, how do we make models better? How do we make
00:09:39models have a higher bar, which don't get me wrong, it has to be solved and we're working very hard to
00:09:43solve that too. But I actually think this problem of inference time is equally, if not even more,
00:09:48important. Because that's actually when you interact with the end user. And this kind of back and forth
00:09:53of how do you understand this context and intent happens at the moment of inference time. So I don't
00:09:57think that we can ignore and just make models better and not solve this so the rise stop will keep
00:10:01existing. So maybe breaking down a few of those pieces and kind of a few of the ways that we've
00:10:07thought about solving this or a few solutions that we built to solve this. But I think, for example,
00:10:11for something like repetition, one of the things that we're working on is, I've nicknamed it,
00:10:15I don't know if that's going to be the official name, but like the creativity API. How can we
00:10:18create a system that almost becomes an inspiration machine for your agent so that it can produce
00:10:23something that's actually out of distribution instead of something that is in that same average
00:10:27and kind of mean that we're seeing happen with like the swap sites. So this is one of the ways that
00:10:32practically if we can intentionally produce something that's out of distribution, you can improve this
00:10:36like overall quality. And by the way, I don't think that this can be something just like randomness.
00:10:43It's not just about like turning up a temperature of a model and kind of fingers crossed hoping for
00:10:47the best. I think it's much more like how do we understand even like what are rules or expectations
00:10:52in specific domains? Like let's say that you ask for a slide deck for for the picture of your startup.
00:10:57Like what does a good pitch deck look like? And then how do you almost like intentionally break rules
00:11:03to create things that are more creative, right? Because usually creativity isn't like randomness,
00:11:07isn't doing something that completely feels off for that situation. It's like you intentionally maybe
00:11:12diverge on a couple of things while maintaining kind of um adherence to to expectations of that
00:11:19category, let's say for others. So that's one of the things we're working on. The second one on this
00:11:23problem of fit, I think um it's interesting, but brands as probably a lot of you who are designers
00:11:28know, take so much effort to create great brands. Like great brands are the work of dozens of designers,
00:11:35uh putting in a lot of like craft and thought and care um and so we've almost like already pre-done the
00:11:40work of defining what is great for that specific company and then we're not using it well. So this
00:11:45like brand adherence actually I think is a huge problem and one of the things that can very more
00:11:49easily let's say like raise that bar of quality. So I'll touch on an example on this one specifically
00:11:54and then same with like intent and judgment. I think the baby classifiers was a good example.
00:11:58Um like how it how we can actually like use this to even become a gate for slop and not let your agent
00:12:04uh ship slop. But so the brand API is the first product that we're releasing to to the public. This
00:12:09is already in in beta testing with a bunch of uh our design partners and essentially what it does is
00:12:14it can take let's say a brand URL and extract this into like very specific components that are good for an
00:12:21agent to follow. So basically how do we turn something as fuzzy as a brand into something so
00:12:25structured that it becomes easy to uh for your agent to follow that but also for you to judge against
00:12:30it right because I think the piece that we can't forget here is this judgment and verification.
00:12:35So yes this goes and helps your agent to produce something better uh but how can we also add a
00:12:40way for you to judge okay is the agent actually staying on track? Is it actually performing well to
00:12:44adhere to this brand or how is it failing or where is it failing? So this is the first flow I would say
00:12:49that we we are seeing that is really helping to improve quality. Um and what's cool is of course
00:12:54we're talking here about an example of a brand that already exists but let's say you have an agent
00:12:59uh or you have an app and uh the person that is using your app actually doesn't have a brand. Let's say
00:13:03they're an average consumer. Can we actually one of the things that we're creating is basically like a
00:13:07repository like an index of brands uh of pre almost like pre-created brand systems so that if they want
00:13:13something that feels dreamy why not retrieve a dreamy brand system that already has been thought
00:13:17out to be cohesive instead of doing like a generative approach the moment of that might end up not so
00:13:23great or might end up again in those pillars of slop. And I want to show you a real example of this in
00:13:28action. So um there's this company that I think is awesome called the General Intelligence Company of
00:13:31New York. They have a sick website you guys should check it out. Um but basically if you ask Claude Design to
00:13:36create a slide deck uh in their branding the the middle one is basically what it comes up with.
00:13:41So the one on the left is is the original brand uh this is kind of the the default. And if you kind
00:13:47of use this extraction actually in the process it creates something that's way more high fidelity with
00:13:51the original um and that even like in the details I would say like feels right. So this is just to show
00:13:57an example of it in in action. Um but yeah I think we I think all of us would agree that like human human
00:14:06taste and kind of the peak of human craft is always going to be like deeply valuable and that right now
00:14:13I think the challenge is we are almost even not earning the right to debate this like how can we
00:14:18have uh models like reach this like pinnacle of taste. I don't think it's about that at all. It's like
00:14:23how do we first just like raise the bar? Like the bar is kind of really I would say on the ground and so
00:14:28I think all of this work that we're putting into like how do we decompose a problem and how do we measure it is
00:14:33exactly so that we can at least like improve this bar of quality. And I think we have to start with that.
00:14:41That's it. Thank you very much for for the time. This is this is awesome.
00:15:03We'll be right back.
Community Posts
No posts yet. Be the first to write about this video!
Write about this video