When the input fights back.
A document carries instructions the user never gave. Does the agent treat them as data—or follow them?
THE PROVING GROUND FOR AI AGENTS
Hostile inputs. Broken tools. Unexpected actions. Put your agent through the hard parts before the real world does.
Interactive demo. No sign-up. No live agent required.
Illustrative scenarios · scripted outcomes · no live model calls
Simulate the unexpected
Compare what holds
Keep the evidence
01 / THE TESTING GROUND
Once an agent can take action, every tool, document, and decision becomes part of the test. Start with the moments that matter.
A document carries instructions the user never gave. Does the agent treat them as data—or follow them?
An API times out halfway through a task. Does the agent recover, stop, or confidently invent a result?
A routine request reaches for restricted data. Does the agent stay inside the user’s permissions?
02 / THE APPROACH
One failure is a story.
A repeatable test is something you can improve.
Choose a task, its tools, and the actions that should be off-limits. Write down what a successful outcome looks like.
Introduce a hostile document, a failed tool, or an unexpected request. Change one condition and compare the outcome.
Inspect the sequence of actions, identify the decision that mattered, and keep the test for the next iteration.
03 / RESEARCH DIRECTION
We’re exploring how to evaluate agents across realistic, multi-step tasks—and how to test safeguards without losing the ability to do useful work.
Read a foundational benchmark: AgentDojoQUESTIONS WORTH TESTING
Can a test discover failures across multiple tools?
Does a safeguard hold up when the scenario changes?
Can the agent stay useful while respecting its boundaries?
Faultyard is in development. This page demonstrates the proposed testing experience.
04 / A FEW DETAILS
Here’s where things stand.
Faultyard is an early product concept. The interactive demo on this page works entirely in your browser and plays through predefined scenarios. It does not evaluate a real model or connect to your systems.
Teams building agents that read external content, call tools, or act on business data. A useful first pilot would focus on one workflow, clear permissions, and a small set of measurable failure cases.
No. This demo uses synthetic inputs and scripted outcomes. It requests no API keys and sends no agent or customer data to a model.
No. A test provides evidence about a particular scenario and configuration. The research direction is to broaden that evidence and make its limits clear; no finite test suite can guarantee safe behaviour in every situation.
LET’S FIND THE FIRST BREAKING POINT
Bring one workflow. Define the boundaries.
Help shape the first Faultyard pilot.