Introduction
Loopjacking reveals a critical flaw in human-in-the-loop AI systems where a user's approval for one operation can be silently reused to authorize something entirely different. The vulnerability arises from how software preserves, represents, and executes decisions across workflow boundaries.
What Happened
Researchers demonstrated loopjacking through controlled experiments. In one scenario, an administrator approved a small mock transfer, but the workflow later executed a much larger payment to a different recipient. The approval record still reflected the original request, yet the system accepted the substituted operation. Another case involved an incomplete approval view that omitted runtime arguments, allowing execution beyond what was shown to the reviewer. A working control in OpenAI Agents SDK preserved approval boundaries across state saves and restores, rejecting any operation change.
Why This Matters
When approvals are not tightly bound to the exact operation that executes, attackers can influence outcomes while appearing to have legitimate authorization. The implications extend to automated finance, file transfers, and command execution in agentic systems. Without strict alignment between approved intent and executed action, human oversight becomes ineffective.
Key Takeaways
Developers must ensure the approval interface and execution engine agree on the exact operation being authorized. Four practical design choices support this: build the approval view from the operation that will execute, bind decisions to protected operation records, compare the resolved operation against the approval immediately before execution, and make approval consumption and audit records precise with expiry and reuse limits. A cryptographic digest can help, but only if it covers the full, tamper-protected operation description.
Conclusion
Loopjacking does not require prompt injection to succeed, but it does require a mismatch between what a human approved and what the system executes. For reviewers, approval should carry concrete meaning. Software's responsibility is to preserve that meaning through every state transition. An audit trail showing who authorized which operation, and that the released operation stayed within that authorization, is the minimum standard for trustworthy AI workflows.










Discussion
Join the conversation
Thoughtful reactions, questions, and follow-up ideas help shape the next story.