Loophole: Adversarial Agents To Stress Test Your Morality — Brendan Rappazzo, Morgan Stanley

AAI Engineer
Computing/SoftwareSmall Business/StartupsInternet Technology

Transcript

00:00:00I'll be talking about my project loophole. And I'm actually a machine learning researcher at Morgan
00:00:17Stanley, but this has nothing to do with Morgan Stanley. This is just an open source project I've
00:00:23been building for fun. And to give sort of the high level flavor to start, it's really this game you
00:00:30can play that's built on top of this adversarial agent framework. So you specify your morals, one
00:00:37agent codifies that into a legal system, and then these two adversarial agents try to find
00:00:42contradictions in your morals. And lately I've been building different extensions on top. But I
00:00:49wanted to, you know, start with sort of the origin story and how I came up with this idea. And so this
00:00:56really started, you know, a long time ago I had sent my DNA into 23andMe for ancestry testing. And I
00:01:05kept hearing about, you know, more recently how DNA samples can be used, of course, to help solve
00:01:11crimes and all these forensics and cold cases. And I was thinking about how I had sort of opted out of
00:01:17everything because, you know, I was scared of the kind of slippery slope and how my DNA would be used.
00:01:26But, you know, there are certain cases that I would be okay with. And it's sort of interesting. I was
00:01:31thinking, like, you know, if someone could present to me case by case, you know, we'll use your DNA to
00:01:37sell, you know, help solve this cold case or this murder, I could sort of say yes or no. And I know where the
00:01:43definition of, like, and the nuance of my morals are. And, but, you know, of course, like, enumerating all of
00:01:51this case by case is really sort of cognitively prohibitive. Like, there's not a good way to do this
00:01:58currently. And then I was thinking sort of more zoomed out that there's a lot of analogies to sort of the
00:02:04legal system as a whole. So, you know, one way to think of what a legal system is in our society is
00:02:10really just a way that we are trying to codify our own moral beliefs. And I think sort of in a similar
00:02:17way, like, finding the true nuance of our morals and what the law should be is this really hard
00:02:23translation task. And I think often we kind of err on the side of being too general because, you know,
00:02:30finding that nuance is really difficult. And if you try to have a perfect translation, it can lead to
00:02:36these kind of weird corner cases or weird failure modes. And I even think kind of like peer-to-peer when
00:02:42we're relating to each other politically, a lot of times the disagreements are more kind of fighting over
00:02:48core values which we don't really disagree on instead of exploring the really nuanced points of our morals.
00:02:57And I think, you know, following that kind of broader legal example, I think in the, you know,
00:03:02the English common law system, that's why we sort of lean on case law so heavily because we know that
00:03:09finding this nuance and these nuanced boundaries is difficult. And so we kind of rely on smart judging to
00:03:16interpret and apply the law correctly. And, you know, of course, even with like the Supreme Court,
00:03:22things can get elevated and we can decide whether a law is valid at all.
00:03:28And so that was sort of, you know, the idea is like, can you take your morals and can you do this kind of
00:03:34synthetic case law generation? So, you know, this is overwhelming to do by hand, but it seems like the
00:03:42new generation of LLMs are finally sort of smart enough to do this kind of high level moral reasoning.
00:03:48And so that was sort of the starting point for this game. And I just want to take you through sort of
00:03:55the initial release of the game, the setup, how it works, and then also talk about some of the different
00:04:00branches I've been building on top of this open source project because I think it could go in some
00:04:05interesting directions. And so at a high level, how the game works is you and natural language,
00:04:12and this all happens, it's sort of like a terminal based game, you specify your morals. And it could be
00:04:19your general morals or maybe about a specific subject. And then there's one agent that takes those morals
00:04:24and drafts sort of a really rich legalese, codified legal system. And then it just operates in this loop
00:04:31where one agent is instructed to try and find loopholes in your system, so something that is
00:04:38immoral but legal. And another is prompted to find overreach, so things that are actually moral but
00:04:45illegal given your system. And a judging agent looks at the, your morals, the produced legal code,
00:04:52and these sort of synthetic case law examples. At first sees can it auto patch, so like maybe the
00:04:58original draft of your, your legal code sort of was an imperfect translation and there's not really
00:05:05a contradiction and it can just sort of auto do this update. Or maybe it's really kind of an
00:05:10under specification of your morals or some kind of contradiction in your morals. And in that case,
00:05:16it raises it to you as the user to sort of be the judge and make a determination.
00:05:22Is it still on for you? It disappeared for me.
00:05:26So I know that, you know, if you, I hope if you're curious about the game, you'll play it. It's all on
00:05:31GitHub, but I just wanted to show some examples. And this is a lot of text, so it's more about just
00:05:37showing the kind of shape of the input and output. So this is sort of how you would provide your input and going back to the DNA example you might, you know,
00:05:44Okay. So I know that, you know, if you, I hope if you're curious about the game you'll play it, it's all on GitHub, but I just wanted to show some examples and this is a lot of text, so it's more about just showing the kind of shape of the input and output. So this is sort of how you would provide your input and going back to the DNA example, you might, you know, specify some number of moral principles.
00:06:10And then the sort of codified legal system, again, just kind of looking at the shape has this really like legalese, you know, preamble, articles, sections, really trying to be, you know, write it in precise legal language.
00:06:24And then these, the different kind of synthetic case laws get suggested. So in this case, it's talking about, this is a loophole it found where an insurance company
00:06:34trained a predictive machine learning model, not on your DNA, but on artifacts of the DNA. And so it's saying, you know, this is actually immoral, but currently legal given your system.
00:06:46And in this case, it's, it found that it could do sort of this auto patching. And then you get this sort of like get style difference of, of your original legal system. And then the difference it had to make to ensure,
00:06:58you know, this was consistent with your morals. And then this is a example of overreach. And in this case, it found that it couldn't do the auto patch. So it's talking about,
00:07:08you know, someone submits their DNA for genetic research, but the researcher finds they have a rare but treatable genetic disorder.
00:07:17But currently, your morals kind of say this shouldn't be allowed that they could disclose this disease to the, the person submitting. And so this was raised to the user to me to kind of make a judgment. And then similarly, when you make the judgment, you get this, this patch legal system.
00:07:33And so, you know, it's just sort of a fun game. And I posted on Twitter and shared it open source on GitHub. And for me, at least, it was by far the most viral post I've had. And it sort of made me think like, I think a lot of people just said it was sort of fun.
00:07:50You could stress test your morals, see if you have any interesting contradictions. But it also made me think, you know, is there maybe something more here? Like, could this be,
00:07:59be, you know, have more like practical or bigger scope implications? And so I'll just talk about three different branches on kind of exploring the first and sort of leaving the legal area, and really more practical is thinking about sort of an auto way to make constitutions for chatbots or really, you know, for agents in general, where, you know, say you're a company, and you want to have a agent or chatbot that's, that's customer facing,
00:08:29and you want it to sort of adhere to a moral code, but also have things that will and will not talk about. I've kind of in one branch formulated it. So you in a similar way, write your morals, you write what the chatbot should and not talk about. And then it tries to write this codified system prompt, and then you have these kind of adversarial agents trying to get it to either talk about something it shouldn't, or refuse to talk about something it should.
00:08:56And I see it as this sort of analogy or analogous method to GEPA, but really aimed at kind of building these codified system prompts.
00:09:03The second use case that I'm particularly interested in is thinking of it as a way to sort of do more ad hoc or decentralized contracts. So I think in a simple case, say like you can specify how you want your data or privacy to be handled online, and you can go through this sort of adversarial game to get this codified legal system of how you want your data handled online.
00:09:30And if you go to, you know, say Apple releases a new terms of service or something, you can run the contradictions between your legal system and between Apple's terms of service and like surface any interesting contradictions or like synthetic cases where this would lead to a difference between how, you know, your morals, what you want and what the company is doing.
00:09:57And you know, in the case that it's a big company, maybe you can't really change anything, it's not a negotiation, but you can at least be sort of have better information about the contract you're signing.
00:10:08But I also think in the case of, you know, thinking more decentralized, like if you're trying to have contracts without, you know, some central authority kind of enforcing them, and you're trying to maybe do contracts across different countries.
00:10:23Thinking about like if you can specify your morals and how you want to like interface, you know, maybe it's just like contracted work, how you want your work to be paid for and the different morals surrounding that, and the other party can do the same.
00:10:40And then you both get this kind of stress tested codified contract and then you can kind of find the disagreements if there are any and surface them before you agree and then you can kind of be more confident in the contract as a whole.
00:10:55And the last thing and maybe the kind of more aspirational angle is thinking about smarter government or more efficient government.
00:11:07I think there would be a lot of different privacy issues and logistical issues but sort of ignoring those for now and just thinking big picture.
00:11:15I think for voters or constituents, you know, this could be a really interesting way if you define your morals, you have this stress tested legal code, sort of any new bill or politician that comes out,
00:11:28you could kind of run your contract against theirs and surface, you know, what are the cases you would disagree or are interesting points that are kind of immoral to you or a contradiction.
00:11:41I also think, you know, relating to one another, it's like a more, I think we all have a lot of nuance in the way we feel about things and this is a way to kind of get to that nuance instead of arguing over just values,
00:11:57which is, you know, often the values are not in contradiction.
00:12:01And then I think maybe a little more practically for legislators, you could imagine if you want to propose a bill and you can have like a simulation of all the other legislators in a legislative body,
00:12:15you could sort of stress test it before submission.
00:12:18And so the third branch I've been building on this project is I tried to do this for the U.S. Senate.
00:12:24And so what I did is I first had Claude go through all current U.S. Senators and look at, you know, kind of all their voting history and anything else that was public and build their kind of moral system and then ran it through the loophole process to get a codified sort of legal code.
00:12:43And then on this system you can, you know, take any current bill that's being proposed or even propose your own and submit it and you can have Claude sort of simulate how each senator would vote.
00:12:56And so here you can see like a breakdown of some senators, which way they're leaning and sort of the reasoning behind the vote.
00:13:07And I think, you know, it's sort of interesting just to think about like seeing what, you know, a proposed piece of legislation, how people would vote.
00:13:16But also this sort of becomes, and I think on theme of the conference, its own verifiable domain or loop.
00:13:22And you could think about even kind of hill climbing the bill towards getting like a super majority or whatever you need it to pass.
00:13:30And so in this case, like this Medicare bill I was testing, you know, it found that I think it originally started at like a 50/50 vote and it found ways to hill climb the language of the bill such that it passed with 52 votes.
00:13:45And I think, you know, this is an example of it can find like the sort of the core tenants of the bill and it can try to find like run the bill against each senator's contract and find is there any way I can change the language such that I don't violate sort of the core tenants or morals of the bill and kind of do those auto patching that way.
00:14:09And then it can also find, you know, kind of rank order the changes that would need to be in place to maximize votes and you as a user can kind of choose the trade offs that way.
00:14:23And then the last thing I've been trying out more recently with this branch is actually looking at, you know, kind of even bigger picture like can this lead to an even more efficient government where you have every sort of constituent
00:14:38in a state or whatever the district is sort of have their legal code and then you could just submit any bill and actually measure sort of the agreement between like the actual voters.
00:14:50And so for this I took the, NVIDIA has this really great data set of USA personas and so I took 500 personas per state and it's supposed to be sort of well representative of the state's population.
00:15:03Did the same process of having them given the persona, draft their morals, draft their sort of legal contract and then take any bill you're interested in and kind of run it against each state.
00:15:15And you can also, you know, measure how much people like this bill or how much it's in agreement with their morals and then also do this hill climbing where you kind of optimize the bill for the people.
00:15:25And so just to conclude, you know, at minimum, I think it's a pretty fun game, I'm biased, but it's a lot of fun to just try out different, you know, things you care about, put in your morals, see if there's any contradictions.
00:15:40You know, often it will raise some really interesting questions and then once you kind of provide that nuance, the game, you know, the agents won't be able to find any more contradictions and you can kind of feel good that you have like a consistent, nuanced moral system.
00:15:56But I am interested in, you know, exploring could this be, are there kind of real applications here for some kind of like decentralized or better contracts and maybe even for legislators as a way to sort of stress test your bills and even think about how to write better laws that are, you know, better for the people in your district or more representative of what the people in your district want.
00:16:19And so this QR code is to the Senate simulator, so I encourage you if you're interested to play. And the other one is to my website, which has the full GitHub loophole. And please, you know, play with it, fork it. I'd love to have other contributors. Thank you.
00:16:49Thank you.

Key Takeaway

Loophole uses adversarial multi-agent frameworks to translate abstract moral principles into formal legal code, enabling automated policy stress-testing, contract alignment, and legislative optimization across 500 state personas.

Highlights

  • Loophole is an open-source adversarial agent framework that stress-tests human moral principles by translating them into formal legal codes and finding contradictions.

  • The core game loop involves one agent drafting legalese, two adversarial agents searching for loopholes and overreaches, and a judging agent either auto-patching or escalating to the user.

  • An initial release of the project on GitHub and Twitter became the creator's most viral post.

  • A second branch of the project applies codified system prompts and adversarial agents to enforce custom rules and boundaries for AI chatbots.

  • Another extension simulates the voting behavior of all current U.S. Senators by converting their public records and voting histories into codified moral contracts.

  • Testing a Medicare bill through the Senate simulator allowed the system to optimize the legislative language and successfully increase passing votes from 50 to 52.

  • A population-scale simulation using 500 NVIDIA personas per state measures public alignment on proposed bills and automatically hill-climbs legislative text to maximize constituent agreement.

Timeline

Origin of Moral Codification and Adversarial Game Design

  • The project originates from a desire to granularly specify consent rules for genetic testing data instead of relying on broad opt-out frameworks.
  • Legal systems function as society-wide translations of collective moral beliefs.
  • The game architecture employs multiple LLM agents to translate natural language morals into formal legalese and aggressively probe for logical inconsistencies.

Abstract moral principles are notoriously difficult to translate into rigid rules without creating bizarre corner cases or overreaches. Traditional common law relies heavily on case law and smart judges to navigate this nuance. This project automates that exact translation task by having one agent draft legal codes while adversarial agents uncover loopholes and overreaches.

Core Mechanics and Synthetic Case Law Generation

  • Players input general or specific moral principles through a terminal interface.
  • The system generates structured legal code featuring preambles, articles, and sections.
  • Synthetic case laws identify hidden contradictions, triggering automatic patches or user-driven moral judgments.

Inputting moral principles generates a rich legalese contract. The adversarial loop exposes edge cases, such as an insurance company using DNA artifacts rather than direct samples to bypass restrictions. When simple auto-patching fails to resolve the discrepancy, the system elevates the dilemma to the user to make a definitive moral ruling.

Chatbot Guardrails and Decentralized Smart Contracts

  • Codified system prompts manage customer-facing chatbot boundaries through adversarial stress-testing.
  • Adversarial agents compare personal data privacy contracts against corporate terms of service to expose hidden clauses.
  • Decentralized cross-border contracts use adversarial matching to surface disagreements before agreements are finalized.

Beyond personal morals, the framework adapts to enterprise use cases like generating rigorous system prompts that withstand prompt injection and boundary testing. It also allows individuals to compare their personal privacy rules directly against new corporate terms of service or negotiate cross-border work agreements without central authorities.

Legislative Simulation and Constituent Optimization

  • Claude reconstructs the moral and voting frameworks of all current U.S. Senators based on public records.
  • A simulated Medicare bill successfully hill-climbs its language to increase supporting votes from 50 to 52.
  • State-level populations modeled using 500 NVIDIA personas per state measure bill agreement and optimize legislation for constituents.

Scaling the framework to government operations involves building simulated profiles of every U.S. Senator and running proposed legislation against their codified contracts. The system automatically modifies bill language to maximize support without violating core tenets. Applying this to 500 NVIDIA personas per state enables large-scale measurement of constituent alignment and legislative optimization.

Community Posts

No posts yet. Be the first to write about this video!

Write about this video