Latest research / September 25, 2026
Six boundaries to check when sandboxing coding agents
A fresh checkout can still share history, memory, and host services. What building coding-agent benchmarks taught us to check.

See what agents choose, where they get stuck, and what to improve in your docs, examples, and tooling.
Record what coding agents do. See what to improve.
From the researchers behind the original What Claude Code Chooses study.
Latest research / September 25, 2026
A fresh checkout can still share history, memory, and host services. What building coding-agent benchmarks taught us to check.
Two products

01 / Choices, findings, playbooks
Measure when coding agents choose your product, what they choose instead, and why. Get findings and prioritized playbooks for your team.

02 / Changes, tests, results
Make a change to your docs, skills, or tooling. Rerun the same task to see what gets better or worse, then iterate.
More from our research
Our original study / February 2026
Our original What Claude Code Chooses study put numbers behind agents’ developer-tool decisions. Today, we help developer tool teams measure and improve how coding agents find, understand, and use their products.
Read the original studySafety / April 2026
Six development tasks leave the security defaults unspecified. The report checks the resulting code against 33 exploit tests.
Read studyNew Amplifying research on coding agents, delivered by email.
Tell us what you’re working on. We’ll show you where Amplifying can help.