OpenTelemetry for LLM Observability: Traces and Metrics A guide to using OpenTelemetry for LLM observability, covering spans, GenAI semantic conventions and exporting to Grafana or Datadog.
AI Observability Platform for Monitoring LLM Costs A guide to monitoring LLM costs with an AI observability platform, covering the cost signals to track and how to turn them into budgets.
How to Govern AI Agents in Financial Services TL;DR: Governing AI agents in financial services means controlling three things on every request: the tool surface the agent can reach, the identity behind the call, and the spend. Existing FINRA, model risk, and EU AI Act obligations already apply.What satisfies an examiner is enforcement on the request
LLM Observability with Prometheus: Metrics and Dashboards A guide to LLM observability with Prometheus, covering the metrics that matter, scraping the gateway and building Grafana dashboards.
What Is an MCP Registry? Discovering and Governing MCP Servers TL;DR * An MCP registry is a discoverable, versioned catalog of Model Context Protocol servers, listing each server's capabilities, authentication requirements, and status so engineers can find and use approved tools without ad-hoc requests. * The official MCP Registry launched in preview on September 8, 2025 as an
MCP Proxy Server Explained: Architecture and Use Cases This guide walks through how a proxy handles a tool call, bridges stdio and Streamable HTTP, aggregates servers, and avoids the confused deputy and token passthrough risks named in the MCP spec, plus the signals that mean it is time for an MCP gateway
How to Cut Time to First Token (TTFT) for LLM Apps TL;DR * Time to first token (TTFT) is the delay between sending a request and receiving the first generated token. It is the latency number users actually feel, not average generation speed or tokens per second. * TTFT is dominated by three things: queue wait before a request starts processing, network