PatchTrial
A private system for verifying AI-generated code repairs, rather than trusting a patch because it looks plausible and compiles.
- My role
- Designed and built the system. Private, with no public access.
- Stack
- LLM toolingAutomated verificationTest executionPython
The problem
A model-generated patch arrives looking correct. It reads well, it applies cleanly, and it compiles. None of that is evidence that it fixed the defect or that it did not break something else. At any real volume a human cannot read every patch carefully enough, so the decision to trust a repair has to be made by something that produces evidence instead of an opinion.
What I built
A system that puts AI-generated repairs through automated verification before they are accepted, so the output of the process is evidence about a patch rather than a verdict about it.
The hard part
Verification has to be adversarial toward the patch. A check that a repair can satisfy by looking plausible is not a check. It has to be able to fail, and failing has to mean something specific about the change.
Evidence
- SourceRestrictedPrivate repository. There is no public demo and no read access, and I would rather say that than stage one.
- Walkthrough on requestRestrictedI will talk through the architecture, the verification model and the trade-offs directly in a conversation.
Happy to walk through the architecture and the trade-offs directly.