Transcript
00:00:00Agents can build most of what you ask it for, but there's still a lot it can't do on its own.
00:00:04People have built tools that fix those gaps, and we went through them to find the ones that
00:00:08actually hold up. Eight of them changed how we work, and two of those have even been trending
00:00:13on GitHub. One of them lets you build those scrollable product landing pages that you've
00:00:17been seeing a lot. Another one solves the biggest problem with these agents, which is that they say
00:00:22a feature is finished when it actually isn't. If this is your first time, we're a software company,
00:00:27and this is our channel, AI Labs. And in this video, we're going over all eight GitHub repos,
00:00:32what each one changes in your setup, and how to use them. Now you can take an object from an image
00:00:37and turn it into a 3D model made entirely from code, then add that model to your landing page,
00:00:42animate it in different ways, and let people interact with it directly on your site. The
00:00:46skill that does this is called Image23JS, and it's also one of the top trending GitHub repos. And
00:00:52because the entire model is made from code, your agent can adjust it the way you want. You can
00:00:57ask it to change the colors, adjust the lighting, or add an animation without having to use extensive
00:01:02software to animate 3D models yourself. Now to use it, you just need to give your agent the image you
00:01:07want it to recreate, and ask it to use Image23JS to turn it into a 3D model. The agent will then build
00:01:14the object in stages. As it builds, it compares the model against your image. If it finds differences
00:01:19between the model and your image, it will fix them before adding more detail. This whole workflow burns a lot
00:01:24of tokens and takes a long time. The model might not be as detailed as you want on the first try, but you can
00:01:30keep asking the agent to improve specific parts and get closer to the level of detail shown in the product
00:01:35demos. Another problem with agents is that they mark a task as complete even though it isn't. They
00:01:40usually check whether the app loads correctly and whether the design was implemented properly, but there are still
00:01:45things that can fail behind that. The only way to catch those is to give the agent more access to what's actually
00:01:51happening in the browser. So to solve that problem, there's a tool called reticle. Once the agent says
00:01:56it's finished the app, this tool checks whether the app actually works the way it should. It lets the agent
00:02:01follow what happens inside the app as it runs, so it can find out why something isn't working. Once it's
00:02:06installed, whenever you ask your agent to build an app, it will use this tool to check what it built, even if you
00:02:11don't mention it in your prompt. The agent automatically opens the app in the browser and uses it the way you
00:02:17would, so it can check whether the app actually works. After the agent has tested the app through
00:02:22reticle, it basically gives three conclusions on each part it tested. These three conclusions are either
00:02:27it worked, it didn't work, or there wasn't enough information to tell. When something fails, reticle
00:02:32explains what went wrong, so the agent can fix the problem and try again. And when you open the app
00:02:37yourself, you'll see a pop-up in the corner for reticle that you can interact with. You can also
00:02:41look through the previous checks to see what went wrong and how the agent fixed those problems. But
00:02:46before we move to the next repo, it would be great if you subscribe to the channel and hit the hype
00:02:51button. This small gesture of support goes a long way for us. You've probably noticed how long the
00:02:56replies from models have been getting lately, especially the summaries they give at the end of their work.
00:03:01This is more prominent in Claude models. Some of that explanation isn't even helpful because of how it's
00:03:06structured. The replies get so long and the words are laid out so oddly that it becomes hard to understand
00:03:11what the agent is trying to say. So for that, there's a tool called Chisel that tells the agent
00:03:16to use fewer words and avoid writing code the task doesn't need. The aim is to reduce the tokens the
00:03:21agent spends on unnecessary explanations and code. Chisel's author has published test results as well.
00:03:27In his example of changing how a search box works, the agent used about 1,500 tokens for its answer
00:03:33without Chisel. With Chisel, the same task used about 600. You might have already heard of tools like
00:03:38Caveman and Ponytail that do similar work. Chisel's author compared those tools as well and the agent
00:03:43used fewer tokens overall with Chisel. It installs as a plugin for Claude code. Once you install it,
00:03:49you'll get four skills along with hooks. For those who don't know about hooks, they're commands that run
00:03:54automatically at particular points in a session. Chisel installs three of them. One runs when the session
00:03:59starts, another when you send a prompt, and the last one after the agent uses a tool. But if you're using
00:04:04Claude code, you get an extra benefit, which is that Chisel also reduces the amount of information
00:04:09the agent has to read. When tools send information back, Chisel removes repeated or unnecessary text
00:04:14before the agent reads that information. This helps keep the context from filling up with text the agent
00:04:20doesn't need. And with less text to read, the agent uses fewer tokens, which also helps reduce the cost of
00:04:25the session. With these skills installed, whenever you give any task to the agent, it will return a much
00:04:30smaller summary than it used to, using more direct wording so that you don't have to struggle to
00:04:35understand what it means. The outputs are clearly structured, first telling you what it generally did,
00:04:40then the key things it changed, and finally what it skipped, all laid out properly. There are a lot of
00:04:46skills for improving your designs, and each one comes with its own design principles. But that's exactly what
00:04:51makes it hard to choose, because there are so many of them, and figuring out which one fits the style
00:04:56your design needs takes real work. For solving this problem, there's a tool called UISkills, which is
00:05:01a library of design skills from different authors. It lets the agent search through those skills and
00:05:06gets whichever skill fits the design you want the best. It's available as both an MCP and a tool you run
00:05:13from the terminal, and you can choose between the two. We went with the MCP option because its tools
00:05:18stay available in the agent's session. With the terminal option, you need to tell the agent when
00:05:22to use UISkills, either in your prompt or in your project instructions. So once you've connected the MCP,
00:05:28it'll be listed among your agent's MCP tools. After that, you can ask your agent to use UISkills for the
00:05:35design you want to make. The agent first searches the library for skills that fit what you asked for in the
00:05:41prompt. From there, it narrows down to the relevant ones and gets their instructions to use while
00:05:45building the site, so you don't have to choose and install every design skill yourself before you
00:05:50start. But before we move on to the next one, let's first have a word by our sponsor, Manufact.
00:05:56Getting your app to actually work inside Claude means shipping an MCP server, and that's a whole
00:06:00project on its own, between a protocol that keeps changing plus hosting and deployment that's easy to
00:06:05get wrong. Manufact made the entire process faster and easier. We built the MCP app with
00:06:11their open source MCP use SDK, connected our GitHub, and now every commit push deploys the server
00:06:17automatically. Then we set up a custom domain, pasted the URL straight into Claude, and our app's real UI
00:06:22showed up live in the chat. We ran it against ChatGPT too, and so we could test both clients. And
00:06:28honestly, seeing our own app running inside Claude was so gratifying. Manufact also runs publishing checks,
00:06:33so your app is ready for the Claude connectors marketplace and the ChatGPT app store, not stuck in review. It's like
00:06:39Versal, but for MCP. And MCP use has over 10,500 GitHub stars and over 9 million downloads. Deploy your
00:06:47first MCP server for free at Manufact.com. If you haven't explained exactly how you want your app
00:06:53to work, the agent will decide those details for you. Ouroboros is a tool that helps the agent work
00:06:58through those details before building the app. It then checks the app and sends any problems back to the
00:07:03agent to fix. To install it, you need to follow the setup commands for the agent you're using. It
00:07:08includes a command that you run from the terminal, and the setup connects that command to your agent,
00:07:13so it can run the workflow for you. Once you've set it up, you need to describe the app you want
00:07:17to build and ask the agent to use Ouroboros. It starts by interviewing itself, it generates
00:07:22questions about your request, then answers them from what you've already described and what it can see
00:07:27in your project. It only fills in small missing details on its own when those choices are easy to
00:07:32change later and stay within what you asked for. If a decision could change what the app is supposed
00:07:37to do, it pauses and asks you. It also records any assumptions so you can see what was decided.
00:07:42Those answers become a plan that describes what the finished app needs to do. Before building starts,
00:07:47Ouroboros reviews that plan and works on any problems it finds, and the agent starts building
00:07:52only after the plan passes. Once the app has been built, Ouroboros checks whether it does what the
00:07:57plan says. The agent's instructions for building the app never include how those checks are run,
00:08:02or what results they expect. If something doesn't work as planned, Ouroboros tells the agent to fix
00:08:07that part. The agent keeps what already worked, and those checks run again to make sure the changes
00:08:12haven't broken anything. You can also open the Ouroboros dashboard to follow the agent's work as it
00:08:17happens. If the agent stops making progress or reaches its limit on attempts, Ouroboros will tell
00:08:22you why the work stopped. There are a lot of design skills available, but fewer are focused on how
00:08:27iPhone apps should look and behave. When you ask an agent for those, it tries to recreate the look
00:08:33of liquid glass, which is the glass-like style Apple uses for buttons and menus by adding blurs and
00:08:38backgrounds. But the problem is that you get something closer to a transparent button than the real thing,
00:08:44even though the agent tells you it's the same, and the apps it builds don't keep up with the
00:08:48functional changes Apple has been making either. So for iOS apps, there's a repo called SwiftUI
00:08:54Skills. It contains the instructions for building your screens properly with SwiftUI, which is Apple's
00:08:59own system for building them, and it's kept updated with the latest changes Apple has been making.
00:09:04There are two skills in the repo. The liquid glass skill tells the agent how to use Apple's built-in
00:09:09buttons and menus in iOS 26, and which common mistakes to avoid. The other one covers how to
00:09:14arrange your app's content for the iPhone Duo, which is the new phone that Apple just launched recently,
00:09:19so your design works at that wider size as well. Once you've installed the skill you need,
00:09:24you can ask your agent to follow its instructions while building your app. It will then use Apple's
00:09:28actual patterns, which gets you much closer to a native app than the agent manages on its own.
00:09:34When you've created a lot of skills, it gets hard to know which skills are still helping the agent do
00:09:39better work. This becomes harder after a model update, because the newer model may already do what
00:09:44a skill tells the agent to do without needing those instructions. So a skill that helped before might
00:09:49now be adding to the token cost without improving the work. You might think Claude Code's new skill
00:09:54doctor command already checks this. That command shows how many tokens each skill takes up in the
00:09:59context and which skills the agent hasn't used. But that doesn't tell you whether a skill actually helps
00:10:04the agent do better work. Caliper is a tool that basically tests if the skill that you've created
00:10:10actually does what it was made for and is actually having any impact on your workflow or not. It can
00:10:15also compare the agent's work with and without the skill to see whether those instructions actually
00:10:20help. Caliper runs from the terminal and the repo includes two skills that help your agent use this
00:10:25tool. The first one is Grill Skill, which interviews you about what your skill should do. So the tests
00:10:30actually reflect what you intended. The second is the Evaluate Skill, which helps your agent create and
00:10:35run those tests, which are called evals. Each test includes a prompt and a description of the result
00:10:41expected from that prompt. The agent saves those tests alongside your skill and Caliper keeps records
00:10:46of the runs in a folder called .caliper, so you can look through the results later. So to use Caliper,
00:10:52you just need to ask your agent to use it on a skill you've created. The agent will then invoke
00:10:57the Evaluate Skill first and prepare the tests for the skill. If it gets stuck in writing tests,
00:11:02it will then invoke the Grill Skill and interview regarding it. Once the tests are written, it will
00:11:06launch multiple runs with different prompts, few using skill and few without the skill, and see how
00:11:12the impact compares. At the end, your agent gives you a report showing what happened across multiple runs
00:11:17and what you can do next. The report helps you see which tasks the skill handles reliably and where the skill
00:11:22still needs work. Comparing runs with and without the skill also shows whether the improvement is
00:11:27worth the extra tokens. Caliper can also make the same comparison for MCP servers and other
00:11:33instructions in your setup, so you can check more than just the skills you've installed. Antislop is
00:11:38basically a tool that's built to remove all the bad coding patterns agents tend to use when writing code.
00:11:43It looks for specific coding mistakes and reports them to the agent so they get corrected without you
00:11:48having to read the code yourself. It checks the code using a set of rules written by the author. Some of
00:11:53them catch unnecessary work that can slow the app down, while others reflect how he prefers to write
00:11:58code. You can follow the instructions in the repo to set up Antislop in your project, so when you have
00:12:03the skill installed, whenever you give your agent any task, it will run these checks and fix all the issues
00:12:09that it comes across. You can also ask it to run a final review once you have completed what you're working on.
00:12:14Now everything we build and show you in all our videos can be found in AI Labs Pro, which is our
00:12:19community. So if you found value in what we do and want to support the channel, this is the best way
00:12:24to do it. The link's in the description. That brings us to the end of this video. If you'd like to support
00:12:29the channel and help us keep making videos like this, you can do so by using the super thanks button below.
00:12:35As always, thank you for watching and I'll see you in the next one.
Community Posts
No posts yet. Be the first to write about this video!
Write about this video