Training Taste — Thais Castello Branco, Taste Labs

AAI Engineer
Computing/SoftwareSmall Business/StartupsInternet Technology

Transcript

00:00:00Okay, amazing. It's great to meet everyone. I'm Thais. I'm the founder of Taste Labs. For those
00:00:19of you who don't know us, we came out of Stout a few weeks ago, and our whole mission is basically
00:00:23how do we end AI Slop? It's my personal enemy. And so we really believe that to solve this problem of
00:00:31Slop, we have to like decode subjective domains. There's been so much effort being put into getting
00:00:36models and agents amazing at things like coding and math, and it's time that we put all that same
00:00:41effort into making them great at things like design and writing. And so design is this first pillar that
00:00:46we're starting with, and it's been incredibly exciting. We work primarily in two ways. So we
00:00:52work a lot with the Frontier Labs on how do we evaluate their models, understand where they're
00:00:57breaking, understand what could be better about them, and then construct the right either post-training
00:01:01data or our environments to basically fix that problem. And part of this is like how do you take
00:01:05something as fuzzy and large as design and break it down to a level that you can identify what is
00:01:11best solved through each method? What are elements of design that are almost like, once you kind of
00:01:15boil down the problem, become so specific that they almost become deterministic. So for example, if you're
00:01:20trying to train a model to be good at selecting color palettes or have contrast or alignment. Those
00:01:25are things that if you define the problem in the context in a specific enough way, you can get to
00:01:30an answer that's like pretty objective or that at least most experts would agree to. But maybe other
00:01:34things like aesthetics, you naturally will see this expert disagreement. And so then you want to lean
00:01:39on to things that are closer to data. So anyway, we spend a lot of time thinking about all those problems.
00:01:43But on the other side is also, without even touching the model layer, right? How do we actually help
00:01:48agents and app layer companies produce better things? And there's a lot that goes into that,
00:01:53right? You have these different sets of problems at the application layer, because you're using an
00:01:56off-the-shelf model that tends to collapse in terms of style, tends to collapse to the mean. So how do we
00:02:02force that creativity back to the system? How do we avoid these patterns of slop, which we'll talk about a lot
00:02:07today? How do you understand like user preferences or brand preferences preference so that you can
00:02:13maintain adherence to that style? So there's lots of things that actually need to be solved as context
00:02:19or judgment or verification at the app layer, which is why we kind of work across both.
00:02:25But maybe I'll start with more of a philosophical question of like, how do you define something that
00:02:29is great? Like, how do you define greatness? And for something like math, it's easier, right? Because
00:02:35there's kind of one objective answer. And great is the same as correct. But then for something like
00:02:40writing or design, it's much harder, right? Like, how do you define what's like a great tweet or what's
00:02:45a great art piece or what's a great website? I don't know what's the last time that you interacted with
00:02:52a poem or walked into a coffee shop and for some reason it kind of like hit different and it felt
00:02:56very special. But probably it's a combination of things that it felt very unique. It felt almost a little
00:03:02different. It kind of called your attention. It felt like it was made with a lot of care and attention
00:03:06to detail and craft. And it almost had this sense of like authenticity. And I think that's a lot
00:03:12of what AI is missing today. It's like, how do we take things that are not necessarily average, right?
00:03:16How do we produce things that are purposely like out of distribution? And slop is kind of the opposite
00:03:21of that, right? I think it is hard to define what is great sometimes, but I think it's pretty easy to
00:03:26define what is slop in the sense that most people would agree. I think the sense of like repetition,
00:03:30of kind of soullessness, is something that all of us feel right now when using AI. And I think it's
00:03:34quite magical, by the way, that AI has gotten to a point that any human on the planet that is not
00:03:39even a designer, that is not an engineer, can click a button and suddenly make an entire PowerPoint or
00:03:44make a website or make a web app. That's pretty cool. But it comes with consequences, right? It comes with
00:03:49consequences of suddenly now the cost of generation is basically going to zero. But the average person
00:03:56hasn't necessarily honed their taste. Like think about the amount of effort and work that a designer
00:04:01puts in throughout their life to like build up their taste, right? Like there's all this process of like
00:04:06getting exposed to many things and learning to like spot patterns and learning to develop a point of
00:04:11view and like kind of do things in a courageous way that maybe are a little bit against the norm.
00:04:15Learning what not to do and how to like have restraint and that's very hard. Like the average person
00:04:21doesn't necessarily have the the time nor the skills to go and develop taste in everything,
00:04:25let's say in design. And so I think it would be a bad case scenario for us to just like be like,
00:04:30okay, the way to fix slop is for everyone to have taste because I don't think that's necessarily
00:04:34realistic. I think how do we how can we understand this better so that we can make even for the average
00:04:39person the ability to create something great and to understand maybe their own taste
00:04:44easy, more easy. So that's that's a lot of what we're we're focusing on. So yeah, I think this
00:04:50phenomenon of slop, by the way, is not new. If you were in the internet as social media emerged,
00:04:55you probably saw a lot of slop before that. But I do think that AI has been this kind of like
00:05:00accelerating force right of like being able to create things very easily with a click of a button
00:05:04and the like thoughtlessness around it. And there's kind of these three characteristics that I would say
00:05:08repeat and slop. So a repetition. So you start seeing the same thing many, many, many times.
00:05:15The second is lack of fit, which I actually think is very related. So
00:05:19fit is kind of this ability for something to feel correct for a specific context, right,
00:05:23for a specific moment in time, for a specific person. But suddenly, if you have repetition,
00:05:27and let's say one person asked for a website for their pet shop, and the other one asked for a website
00:05:32for their finance firm, and somehow those designs converge and look the same, that's quite odd,
00:05:38right? Like if that wasn't, if you were actually crafting that with care, that wouldn't, you
00:05:42wouldn't converge necessarily on those things. And so this lack of fit and lack of understanding of
00:05:46context is actually a huge problem that like leads to slop. And the third is maybe low intent,
00:05:51which is probably a mix of, yeah, you're gonna have a bunch of people prompting really quickly,
00:05:54and maybe just wanting to one shot something. But I think there's actually this like intent
00:05:58interpretation piece that's missing in the systems that we're building. Like,
00:06:01how can you help your user, right? Like, how can you help them better understand the intent that they
00:06:05have, so that you can add more color and add more context onto what you're trying to create?
00:06:12Okay, and I'm a big believer, by the way, that you, in order to fix something, you first have to
00:06:17measure it, and you first have to understand it. I think that's exactly why we're so focused on,
00:06:21like, how do we turn these domains into something a bit more verifiable, so that we can attach a
00:06:27measure to it. So you'll go on a little bit of a research journey with me here now, but
00:06:31we basically wanted to figure out, can we measure slop? Like, can we actually measure this quantitatively
00:06:36and spot this? And what does that like look like? So we analyzed over two million websites from the
00:06:43past like 10 years, kind of like way back machine style to try to understand all the trends across
00:06:48like design? How is the internet changing? How is like design changing over time? And two things were
00:06:55interesting. And we also, by the way, then kind of synthetically generated a set of design websites,
00:07:00so we could kind of like compare. Like, how does human-made sites compare to AI generated ones? And
00:07:07there were a few things that were interesting. So one was that you already kind of saw a bit of a collapse
00:07:13on the internet before even AI. So you saw kind of the internet becoming more homogenous, using more
00:07:18similar color palettes, using more similar layouts, which is probably a function of more,
00:07:24I would say this kind of trend spreading more quickly, let's say. But with AI, I think you saw
00:07:28this repetition happening a lot more and being almost more like identified kind of regardless of
00:07:33context. So even in completely different buckets, you saw patterns that were very similar. So we built
00:07:38this, I call this probes, but basically we did two things. So we did this like pattern mining on all
00:07:44this data to understand like, what are features that we can extract from all these sites? What are
00:07:47all these characteristics that we can make more objective, right? Colors, typography, layout, audience,
00:07:53like how can we like distill this down into things that become almost like structured? And then how do we
00:07:58train up these like probes? So think of these as like baby classifiers. Like how do we train the ability to
00:08:03spot this one characteristic? And for all these slop sites, we identified, we started identifying like,
00:08:09what are the probes that basically mean this site is very likely to be AI slop? And especially when
00:08:16you start combining them and you see the frequency of multiple of these happening at once, it became
00:08:20very likely that you could actually like measure and predict slop. And we saw a super high, basically,
00:08:26ability to do that prediction, which was really cool to see. This performed better, by the way,
00:08:30than like most LLM as a judge methods of like asking an LLM to like judge if that is a great human
00:08:36quality versus like AI generated slop. So that was pretty cool to see. And I think it kind of shows
00:08:40this pattern that we see in AI really being an actual quantitative thing that we can see in slop,
00:08:47which I find really cool. But obviously, we don't want to stop there, right? We don't want to just
00:08:51measure slop. We want to also solve it. And so there's a few, I think I mentioned this before,
00:08:56but like the, as the cost of production basically goes to zero, I think the thing that becomes
00:09:01expensive and matters more than ever is judgment. I don't even want to use the word taste here,
00:09:07is judgment. I think it's this ability to discern what's right, is this ability to break down a
00:09:11problem so that you can actually understand it and create solutions for it. And so, yes,
00:09:16there's a side of judgment that is human judgment that I actually think is more valuable than ever.
00:09:19But there's also this side of like, how do we build the right tools and systems to like
00:09:23fix pieces of this problem, right?
00:09:27So yeah, how do we fight slop, my enemy? And by the way, I think there's a lot of conversation going
00:09:34around how do you fight slop at the model layer? Like, how do we make models better? How do we make
00:09:39models have a higher bar, which don't get me wrong, it has to be solved and we're working very hard to
00:09:43solve that too. But I actually think this problem of inference time is equally, if not even more,
00:09:48important. Because that's actually when you interact with the end user. And this kind of back and forth
00:09:53of how do you understand this context and intent happens at the moment of inference time. So I don't
00:09:57think that we can ignore and just make models better and not solve this so the rise stop will keep
00:10:01existing. So maybe breaking down a few of those pieces and kind of a few of the ways that we've
00:10:07thought about solving this or a few solutions that we built to solve this. But I think, for example,
00:10:11for something like repetition, one of the things that we're working on is, I've nicknamed it,
00:10:15I don't know if that's going to be the official name, but like the creativity API. How can we
00:10:18create a system that almost becomes an inspiration machine for your agent so that it can produce
00:10:23something that's actually out of distribution instead of something that is in that same average
00:10:27and kind of mean that we're seeing happen with like the swap sites. So this is one of the ways that
00:10:32practically if we can intentionally produce something that's out of distribution, you can improve this
00:10:36like overall quality. And by the way, I don't think that this can be something just like randomness.
00:10:43It's not just about like turning up a temperature of a model and kind of fingers crossed hoping for
00:10:47the best. I think it's much more like how do we understand even like what are rules or expectations
00:10:52in specific domains? Like let's say that you ask for a slide deck for for the picture of your startup.
00:10:57Like what does a good pitch deck look like? And then how do you almost like intentionally break rules
00:11:03to create things that are more creative, right? Because usually creativity isn't like randomness,
00:11:07isn't doing something that completely feels off for that situation. It's like you intentionally maybe
00:11:12diverge on a couple of things while maintaining kind of um adherence to to expectations of that
00:11:19category, let's say for others. So that's one of the things we're working on. The second one on this
00:11:23problem of fit, I think um it's interesting, but brands as probably a lot of you who are designers
00:11:28know, take so much effort to create great brands. Like great brands are the work of dozens of designers,
00:11:35uh putting in a lot of like craft and thought and care um and so we've almost like already pre-done the
00:11:40work of defining what is great for that specific company and then we're not using it well. So this
00:11:45like brand adherence actually I think is a huge problem and one of the things that can very more
00:11:49easily let's say like raise that bar of quality. So I'll touch on an example on this one specifically
00:11:54and then same with like intent and judgment. I think the baby classifiers was a good example.
00:11:58Um like how it how we can actually like use this to even become a gate for slop and not let your agent
00:12:04uh ship slop. But so the brand API is the first product that we're releasing to to the public. This
00:12:09is already in in beta testing with a bunch of uh our design partners and essentially what it does is
00:12:14it can take let's say a brand URL and extract this into like very specific components that are good for an
00:12:21agent to follow. So basically how do we turn something as fuzzy as a brand into something so
00:12:25structured that it becomes easy to uh for your agent to follow that but also for you to judge against
00:12:30it right because I think the piece that we can't forget here is this judgment and verification.
00:12:35So yes this goes and helps your agent to produce something better uh but how can we also add a
00:12:40way for you to judge okay is the agent actually staying on track? Is it actually performing well to
00:12:44adhere to this brand or how is it failing or where is it failing? So this is the first flow I would say
00:12:49that we we are seeing that is really helping to improve quality. Um and what's cool is of course
00:12:54we're talking here about an example of a brand that already exists but let's say you have an agent
00:12:59uh or you have an app and uh the person that is using your app actually doesn't have a brand. Let's say
00:13:03they're an average consumer. Can we actually one of the things that we're creating is basically like a
00:13:07repository like an index of brands uh of pre almost like pre-created brand systems so that if they want
00:13:13something that feels dreamy why not retrieve a dreamy brand system that already has been thought
00:13:17out to be cohesive instead of doing like a generative approach the moment of that might end up not so
00:13:23great or might end up again in those pillars of slop. And I want to show you a real example of this in
00:13:28action. So um there's this company that I think is awesome called the General Intelligence Company of
00:13:31New York. They have a sick website you guys should check it out. Um but basically if you ask Claude Design to
00:13:36create a slide deck uh in their branding the the middle one is basically what it comes up with.
00:13:41So the one on the left is is the original brand uh this is kind of the the default. And if you kind
00:13:47of use this extraction actually in the process it creates something that's way more high fidelity with
00:13:51the original um and that even like in the details I would say like feels right. So this is just to show
00:13:57an example of it in in action. Um but yeah I think we I think all of us would agree that like human human
00:14:06taste and kind of the peak of human craft is always going to be like deeply valuable and that right now
00:14:13I think the challenge is we are almost even not earning the right to debate this like how can we
00:14:18have uh models like reach this like pinnacle of taste. I don't think it's about that at all. It's like
00:14:23how do we first just like raise the bar? Like the bar is kind of really I would say on the ground and so
00:14:28I think all of this work that we're putting into like how do we decompose a problem and how do we measure it is
00:14:33exactly so that we can at least like improve this bar of quality. And I think we have to start with that.
00:14:41That's it. Thank you very much for for the time. This is this is awesome.
00:15:03We'll be right back.

Key Takeaway

Eliminating AI slop requires decomposing subjective domains like design into measurable, deterministic elements and applying structured brand systems during inference time.

Highlights

  • Taste Labs emerged from Stout to eliminate AI slop by decoding subjective domains such as design and writing.

  • Analyzing over two million websites from the past decade revealed a pre-existing internet homogeneity that AI generation accelerates significantly.

  • Probes and small classifiers trained on pattern-mined features identify AI-generated slop more effectively than typical LLM-as-a-judge approaches.

  • AI slop manifests primarily through repetition, lack of contextual fit, and low user intent interpretation.

  • The brand API extracts existing web URLs into structured components for agent adherence and automated quality gating.

Timeline

Decoding Subjective Domains and the Problem of AI Slop

  • Taste Labs focuses on solving subjective domains like design and writing rather than just coding and math.
  • AI slop results from zero-generation costs combined with average users lacking the time to hone professional taste.
  • Repetition, lack of fit for specific contexts, and low intent characterize AI slop.

Models collapse to the mean when generating content off the shelf, stripping away unique craft, care, and authenticity. While the democratization of creation tools allows anyone to generate a website or presentation instantly, it creates widespread uniformity and soulless patterns.

Quantifying and Measuring Slop Across Millions of Websites

  • Analyzing over two million websites across a ten-year span tracked how the internet became increasingly homogenous over time.
  • Feature extraction isolated specific visual characteristics like colors, typography, and layouts into structured data probes.
  • Small classifiers trained on specific patterns predict AI slop with higher accuracy than general LLM-as-a-judge methods.

Measuring slop requires turning fuzzy design concepts into verifiable, quantitative metrics. Pattern mining historical web data uncovers how layout and style trends converge, allowing systems to flag likely AI-generated slop effectively.

Inference-Time Solutions and the Brand API

  • Inference-time tools matter equally to model-layer improvements for managing real-time user context and intent.
  • The brand API extracts existing website branding into structured components that guide agent outputs toward higher fidelity.
  • Pre-created brand repositories provide cohesive styling options for consumers who lack predefined brand guidelines.

Fighting slop requires an inspiration machine or creativity API that forces agents to produce out-of-distribution content instead of falling back on statistical averages. Extracting precise brand elements ensures that generated slide decks and web assets maintain high fidelity to professional standards.

Community Posts

No posts yet. Be the first to write about this video!

Write about this video