Explore/observability/AgentSpy: Making AI Agent Behavior Observable
A

Christoph Bühler, Matteo Biagiola, Luca Di Grazia, Guido Salvaneschi/AgentSpy: Making AI Agent Behavior ObservableUnknown

AI agents built on large language models (LLMs) run shell commands, read and write files, and reach the network, typically with their user's privileges. However, what an agent does during an execution is difficult to understand: tests assert on the result, and the agent's trajectory records only what the agent reports about itself, which may omit behavior executed by its subprocesses. We present AgentSpy, an approach that observes an agent from outside the agent. AgentSpy runs the agent in an isolated environment, configured by a declarative specification, and records the system calls and network traffic of the agent and of every process it executes. Based on this monitoring, AgentSpy supports two families of analyses: conformance analyses, which measure obligations, i.e., what an agent execution should do, and safety analyses, which check prohibitions, i.e., what an agent execution must never do. We instantiate one analysis of each family. The reliability analysis uses rules to summarize each run by the environment resources the agent uses: the commands it executed, the files it accessed, and the hosts it contacted. The security analysis applies deterministic rules to the system calls of an execution. For reliability, we evaluated AgentSpy on 77 tasks with the codex harness and three recent LLMs, executing each task three times. Sets of repeated runs of the same task are more similar than sets that include runs of another task in 92.2% of the comparisons. Among tasks for which all three runs pass outcome-based tests, the agent performs task-unrelated activities in 18% of the cases, reads the grading files in 7%, and does not use the developers' guidance in 17%. For security, the generic rules of AgentSpy detect four of five attack categories we considered, with no false positives across 50 runs.

observability
GitHubCompare
Refreshed 10h ago
OverviewActivity52wAlternativesDocs
Stars0
Forks0
HF Downloads—30d
Last commit—
Refreshed10h ago
Project healthUnknownNo activity data.
Production readinessResearch / EarlyBest for exploration and prototyping.
Risk notesUnknown licenseVerify license before production use.
AgentHub Score
55 / 100
Composite score from 6 signals. How we score →
Active project
55Score
Growth
40C
Activity
30C
Documentation
70C+
Maturity
45C
Community
42C
Production
58C
GitHub stars · 0 days observed0 not enough history
snapshots
not enough history
Repository activity · 0 days observednot enough history from pushed_at
inactivepushed
not enough history
not enough history
Practical assessment
Should you use it?

✓ Best for

  • Agent tracing in production
  • LLM cost tracking
  • Debugging agent failures

◎ Strengths

  • Integrates with major frameworks
  • OpenTelemetry compatible

✕ Not ideal for

  • Small prototype projects
  • Teams without existing observability culture

⚠ Watch-outs

  • Data retention costs at scale
  • Privacy implications of storing LLM traces
Technical details
What's inside
Language—
License—
Sourcearxiv
Open source✗ No
Commercial use—
Docs—
Demo—

AgentHub Score

55
Score 55/100
Below average

Alternatives

B
BOHM: Zero-Cost Hierarchical Attribution for Compound AI Systems
0 · observability
55
T
Traccia: An OpenTelemetry-Based Governance Platform for AI Systems
0 · observability
55
F
Formal Methods Meet LLMs: Auditing, Monitoring, and Intervention for Compliance of Advanced AI Systems
0 · observability
55
I
In-IDE Toolkit for Developers of AI-Based Features
0 · observability
55
Compare all →

Recent activity

Latest commit ——
Indexed by AgentHub crawler10h ago
Monitor for new releasesongoing