arthuursantos opened a new issue, #13774: URL: https://github.com/apache/apisix/issues/13774
### Description > Drafted by Claude through back-and-forth with @arthuursantos. The direction and prioritization are his; the prose is Claude's. ## Description As an operator running APISIX as an AI gateway, I want APISIX to ingest the OpenTelemetry **usage metrics** emitted by agents that do **not** flow through the proxy (e.g. Claude Pro/Max/Team subscription, which authenticates directly against Anthropic), so that subscription-backed token/cost usage lands in the **same** metrics and logging pipeline as the traffic `ai-proxy` already meters — giving me one gateway-owned view of total LLM consumption instead of two disconnected ones. > Scope note: this is about **unifying accounting**, not about gaining visibility. Visibility already exists today (see Alternatives considered). The gap this closes is having subscription usage and proxied usage in one pipeline, under one set of labels. ## Problem Today `ai-proxy` meters token usage **only for requests it proxies**: usage is extracted from the upstream response on the wire (`extract_usage` / `merge_usage` in `apisix/plugins/ai-providers/base.lua`) and surfaced via the normalized set of `ctx.var` fields — `llm_prompt_tokens`, `llm_completion_tokens`, `llm_total_tokens`, `llm_cache_read_input_tokens`, `llm_cache_creation_input_tokens`, `llm_reasoning_tokens` — which then feed the Prometheus LLM metrics and the logger plugins. Subscription-backed agents (Claude Code with a Pro/Max/Team plan, IDE assistants, etc.) call the provider directly and never traverse the gateway, so their consumption is invisible to APISIX. Operators running a mixed fleet (some traffic via `ai-proxy` API keys, some via subscriptions) have no single gateway-owned place that reflects total LLM usage or lets them normalize subscription usage against API pricing under the same labels. ## Source of the data Per the Claude observability docs, the Claude Code CLI emits **three independent OTLP signals**, each with its own exporter: - **Metrics** (`OTEL_METRICS_EXPORTER`) — "Counters for **tokens, cost**, sessions, lines of code, and tool decisions." - **Log events** (`OTEL_LOGS_EXPORTER`) — structured records per prompt / API request / API error / tool result. - **Traces** (`OTEL_TRACES_EXPORTER`, beta) — spans; `claude_code.llm_request` carries token counts as attributes. It exports over OTLP HTTP (`OTEL_EXPORTER_OTLP_PROTOCOL=http/protobuf`, default port `4318`), enabled by `CLAUDE_CODE_ENABLE_TELEMETRY=1`. By default only structural/numeric data is exported; prompt/response content is opt-in (`OTEL_LOG_USER_PROMPTS`, `OTEL_LOG_TOOL_*`). The relevant signal for usage accounting is therefore **metrics**, not traces. ## Prior art Mozilla AI's Otari implemented this in [mozilla-ai/otari#378](https://github.com/mozilla-ai/otari/issues/378): a master-key-protected `POST /v1/usage/external-events` endpoint that ingests normalized usage metadata from local agents/collectors (mapping Claude Code's OTel export — `request_id` → `source_event_id`, cache-token accounting, etc.). Key constraints there, which align with Claude Code's default export posture: - accepts **only** metadata and numeric usage fields — never prompts, completions, tool input, or tool output; - events are idempotent via a unique `(source, source_event_id)`; - imported events do **not** trigger reservation/reconciliation/refund/spend-mutation and do not affect budget limits. ## Proposed solution (APISIX-idiomatic) Otari persists events into a `usage_logs` table because it is a DB-backed application. APISIX is a stateless streaming gateway with no usage store, so the natural adaptation is: **receive the usage event, normalize it, and inject it into the telemetry path `ai-proxy` already owns** — no new storage layer. 1. **Ingestion endpoint.** Expose an authenticated receiver (likely via the `public-api` plugin mechanism) that accepts external usage events. Auth on par with other sensitive endpoints (dedicated key / consumer). 2. **Wire format = OTLP metrics (`http/protobuf`).** This is exactly what Claude Code emits, so an operator can point `OTEL_EXPORTER_OTLP_ENDPOINT` at APISIX with no intermediate transform. (An optional compact JSON schema could be a secondary input — see open questions.) 3. **Map to the existing normalized model.** Translate the token/cost counters onto the same `llm_*` `ctx.var` fields listed above. 4. **Emit into existing sinks.** Re-emit through the current Prometheus LLM metrics and logger plugins (`http-logger`, `loki-logger`, etc.), tagged with a `source` label (e.g. `claude-code`) so proxied vs. imported usage is distinguishable in the same dashboards. 5. **Safety constraints** (mirroring Otari): imported events are accounting-only and must be inert with respect to `ai-rate-limiting` reservation/budget paths; content must never be accepted or persisted. ## Scope / non-goals - **Not** routing subscription traffic through the gateway or handling subscription OAuth tokens — separate concern, and constrained by provider ToS. - **Not** a persistent usage database inside APISIX — events flow to the existing observability sinks. - **Not** ingesting request/response content — usage metadata/counters only. - **Not** consuming the `traces` or generic `logs` signals — the `metrics` signal is the target. ## Alternatives considered 1. **Existing `opentelemetry` plugin** (`apisix/plugins/opentelemetry.lua`). Does **not** address this. It is an OTLP **exporter** of the **traces** signal (`opentelemetry.trace.exporter.otlp`, `batch_span_processor`, `tracer_provider`), producing spans for requests that traverse an APISIX route and shipping them outbound to a collector. Wrong direction (export, not ingest), wrong signal (traces, not the metrics counters that carry usage/cost), and no data source for agents that bypass the gateway. 2. **Standalone OpenTelemetry Collector.** Because Claude Code speaks standard OTLP, an operator can point its metrics exporter at a standalone collector that forwards to Prometheus/Datadog/Grafana today — **no APISIX change needed**. This fully solves *visibility*. What it does **not** do is unify that usage with `ai-proxy`'s own `llm_*` metrics under one gateway-owned pipeline and label set. This feature exists specifically to close that unification gap; operators who only need visibility should use a standalone collector instead. ## References - Prior art: mozilla-ai/otari#378 - Source telemetry: [Claude — Observability with OpenTelemetry](https://code.claude.com/docs/en/agent-sdk/observability) -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
