agent-otel-bridge v0.5.2

agent-otel-bridge · OpenTelemetry for AI agents

See what your AI agents actually do. Without slowing them down.

agent-otel-bridge turns the lifecycle hooks of AI coding tools into standard OpenTelemetry traces and metrics. A 145 KB native hook returns in about 150 µs and fails open, so telemetry never blocks an agent turn.

cargo install agent-otel-bridge agent-otel-client
  • Apache-2.0
  • Linux · macOS · Windows
  • OTLP traces & metrics
  • GenAI semantic conventions
A multi-agent trace in your observability tool An illustrated flamegraph: a Claude Code agent span contains tool spans classified as inspection, build and test, and state mutation, and a child Codex agent span linked by W3C trace context. trace · feature/email-validation 0 ms1.2 s2.4 s3.6 s invoke_agent claude-code git diff edit · mutate cargo test · verify invoke_agent codex (child) rg · search apply_patch traceparent 00-4bf9…a3ce-00f0…67aa-01 vcs.branch.name feature/email-validation hook overhead ~150 µs per call · fail-open at 3 ms SigNoz Grafana Datadog any OTLP backend Illustrative trace, not recorded production data.

For technology leadership

You can't govern agent work you can't see.

AI coding tools now take part in delivery, but their activity is scattered across incompatible local logs. Leaders can't yet answer basic questions: what agents are doing, where the time goes, when they change state, and how close teams are to provider limits.

Use the stack you already own

Standard OTLP flows into SigNoz, Grafana, Datadog, Honeycomb, or any OpenTelemetry collector. There is no new vendor and no new console.

Governance · one view of agent activity

Behaviour, not just volume

Commands are classified into ten behavioural archetypes, from inspection and search to build-and-test, state mutation, and network transfer, so dashboards separate reading from changing.

Risk · see when agents mutate state

Plan capacity, not surprises

Quota metrics for the tools actually installed report remaining headroom and time to reset, so teams can plan agent capacity rather than discover limits mid-task.

Cost · headroom before limits hit
In one sentence

agent-otel-bridge gives leadership an evidence base for AI-assisted engineering — adoption, behaviour, and capacity — on the observability standard the organization already invests in, at a cost to developers measured in microseconds.

For adopters

From install to your first trace.

You need Rust's cargo (or a release archive) and an OpenTelemetry collector endpoint. Ready-made SigNoz dashboards are included.

  1. Install the CLI and hook

    Both crates are published on crates.io. Pre-built archives are also available from GitHub Releases.

    cargo install agent-otel-bridge
    cargo install agent-otel-client
  2. Stage the local runtime

    Installs versioned binaries to an isolated per-user path, records SHA-256 integrity hashes, and registers canonical hook paths idempotently.

    agent-otel-bridge local install
  3. Register hooks for your agents

    Detects installed Antigravity, Claude Code, Codex, Grok, and Pi, and configures their lifecycle hooks. You can also target one client.

    agent-otel-bridge install-hooks
    # or: --client claude | codex | antigravity
    agent-otel-bridge hooks status
  4. Point at your collector

    Standard OpenTelemetry environment variables choose the destination and resource attributes.

    OTEL_EXPORTER_OTLP_ENDPOINT=http://127.0.0.1:4318
    OTEL_RESOURCE_ATTRIBUTES=deployment.environment=dev
  5. Verify end to end

    A five-point diagnostic checks environment, IPC, the collector, the observability UI, and every hook registration.

    agent-otel-bridge doctor

Import the SigNoz dashboards for agent observability, developer velocity, fleet governance, fleet operations, and agent SRE loops from contrib/dashboards/signoz ↗.

For technologists

Two tiers, so the agent never waits for telemetry.

Hooks run synchronously inside every agent turn, and script-based hooks pay process start-up costs on each call. agent-otel-bridge splits the work: a tiny native client on the hot path, an asynchronous batcher off it.

agent-otel-bridge architecture An agent lifecycle hook calls the 145 kilobyte agent-hook binary, which writes a three-byte frame over local IPC and returns immediately. A background daemon classifies, enriches, micro-batches, and exports OTLP to a collector and on to observability backends. Agent hook PreToolUse · PostToolUse Invocation · Stop TIER 1 · HOT PATH agent-hook 145 KB native binary 3 ms fail-open watchdog returns {} · exit 0 TIER 2 · ASYNC DAEMON Archetype engine · ~2.8 µs Context harvester · no subprocesses Micro-batcher · 50 spans / 200 ms Adaptive quota metrics OTLP protobuf · keep-alive pool OTel Collector :4318 HTTP · :4317 gRPC Backends SigNoz · Grafana · Datadog 3-byte IPC
Fail-open is a guarantee, not a hope
A dedicated OS thread enforces a 3 ms deadline. If IPC stalls or the daemon is down, the hook prints {} and exits 0; the agent turn continues.
Native IPC on every platform
Non-blocking Unix domain sockets on Linux and macOS, overlapped named pipes on Windows, carrying a 3-byte frame of event and client identifiers.
GenAI semantic conventions
invoke_agent and execute_tool spans with gen_ai.* attributes, plus workspace, VCS, archetype, and quota attributes.
Cross-agent distributed tracing
W3C traceparent propagates through environment variables, so Antigravity → Claude Code → Codex chains appear as one trace.
Context without spawning processes
Branch, worktree, and sanitized remote are read directly from .git; credentials are stripped from URLs, and git is never spawned on the telemetry path.
Limitations are documented
Fail-open can drop spans under stall, visibility depends on tool calls, and provider quota data differs by tool. Each has a stated mitigation in LIMITATIONS.md ↗.
145 KBhook binary (budget < 300 KB)
~150 µshook cold-start impact per call (budget < 1 ms)
101 µsp99 named-pipe IPC round trip on Windows
62,943spans per second, ProtoJSON parser throughput

Figures are the project's own measurements; methodology and raw data are in docs/BENCHMARKS.md ↗. Reproduce them with agent-otel-bench.

Status and boundaries

What you can rely on today.

An honest picture for platform teams evaluating agent telemetry.

Release0.5.2, published on crates.io as agent-otel-bridge and agent-otel-client, with release archives on GitHub. Pre-1.0: expect attribute and configuration changes between minor versions.
PlatformsLinux, macOS, and Windows.
Agent toolsGoogle Antigravity, Claude Code, OpenAI Codex, xAI Grok, and Pi.
Data handlingRuns locally and exports only to the OTLP endpoint you configure. Remote URLs are sanitized of credentials before export.
LicenceApache-2.0.

Questions

Before you instrument.

Short answers for platform engineers, security reviewers, and engineering managers.

Will it slow down my agent?

The hook is designed for microsecond overhead and has a hard 3 ms fail-open deadline. If anything goes wrong, the agent continues and only telemetry is lost.

Do I need a specific observability vendor?

No. It emits standard OTLP traces and metrics to any OpenTelemetry collector. SigNoz dashboards are provided as a starting point.

Does it capture prompts or source code?

It is designed around lifecycle events, tool names, command archetypes, token counts, workspace and VCS context, and quota metrics rather than conversation content, and it strips credentials from remote URLs. Review the telemetry dictionary ↗ for the complete attribute list, and filter in your collector where policy requires.

Why not use each tool's native telemetry?

Native telemetry differs by tool and rarely links parent and child agents. The bridge gives one semantic model and one trace across tools. The trade-offs are discussed in WHY_AGENT_OTEL_BRIDGE.md ↗.

Does the daemon run forever?

No. It shuts down after 30 minutes without agent activity.

Put AI-assisted engineering on the same footing as production software.

Instrument one workstation, open a trace, and decide from evidence how agents are really being used.

Built by Samuel Mota · Move the Needle

Why agent-otel-bridge exists

Open source · Apache-2.0 · 0.5.2 on crates.io

Value

See what AI coding agents actually do, including adoption, behaviour, and quota headroom, in the observability stack you already run.

Strategy

Agent activity is now part of software delivery, so it deserves the same evidence base as production systems. Standard OTLP and GenAI semantic conventions avoid yet another vendor console.

Technical frontier

A 145 KB native hook with a 3 ms fail-open watchdog on the hot path, a three-byte IPC frame, and an asynchronous micro-batcher that classifies commands into behavioural archetypes and links multi-agent chains with W3C trace context.