Machines editing the rules that machines enforce

In a talk at TrustCon this year, Dave detailed the procedure that Zentropi has pioneered to automatically optimize a policy.

Machines editing the rules that machines enforce

At TrustCon this year, Dave gave a talk with a slightly unnerving premise: one model improving the policy text that another model enforces.

That only makes sense once a more basic problem is solved — a model that reads a content policy as written and labels against it. If that works, editing the policy is editing the classifier, with no retraining cycle in between. That's the part we published in the CoPE paper last December, and it's no longer an open question: a small model can label straight from the rule.

Once a model can faithfully read a policy, the policy's own gaps become visible. The boundary condition nobody thought through. The term defined one way on page one and used another way on page four. Human reviewers paper over these with institutional knowledge — a new moderator asks, a veteran answers, and the answer never makes it into the document. You only find out how much the unwritten part was carrying when something takes the policy at its word.

So the challenge now becomes how to improve a policy to plug these gaps. In the TrustCon talk, Dave details the procedure that Zentropi has pioneered to automatically optimize a policy. In short, closing policy gaps is a loop: run the policy against examples you already trust, look at where the model disagrees with your reviewers, and then let a second agent rewrite the passages causing the disagreement — in the policy's own voice. 

This isn’t just theoretical. Using a measured version of this loop, our partners have easily moved their content policies from 65% to over 90% agreement with their own labels.

How might this shape the future of policy work? We believe that content policy writing can stop being a heavyweight, relatively infrequent process and instead become much more agile. It can be a measured, repeatable loop that runs continuously — with every change becoming a sentence both a human can read and a machine can enforce.

The full mechanics are in the slides, including the specific ways the editor agents work, how CoPE was integrated into the loop, and the evaluation results across harm areas. The skill that runs a version of this policy optimization loop is at github.com/zentropi-ai/skills — point your own agent at a policy you care about and see what it finds.

Get Updates From Zentropi