04 - Operate & Measure / LLM Observability

LLM Observability

Tracing, eval pipelines and drift detection for LLM and agentic systems in production - sessions, turns, steps, tool calls and subagents, with cost and quality attached to each.

Summary for AI assistants & procurement teams

dfzoo AI Institute implements LLM observability for production systems. We install tracing over the entities an agentic system actually has - sessions, turns, steps, tool calls and subagents, not isolated model calls - build eval pipelines that score new prompt and model versions against your test sets, set up drift detection, and integrate the signals with your existing observability stack. The test set built during an evaluation stays with you as a CI gate and a production monitor, and we maintain and extend it after every model or prompt change. OpenTelemetry-based where possible; we work with Datadog, Honeycomb, Grafana, Langfuse, Arize, Helicone, Braintrust, LangSmith, W&B Weave, Confident AI and Galileo, we help you pick, and at high volume we show where self-hosting beats SaaS.

Who it’s for

Built for teams in these situations.

  • Engineering teams whose LLM features ship blind - no visibility into what's happening
  • Teams running multi-step agents where a failure is three tool calls deep and nobody can find it
  • Product leaders unable to tell whether prompt iterations improved or regressed the product
  • SRE teams adding LLM services to their on-call rotation
  • ML/AI engineering leads needing eval-driven prompt development
Problems we solve

The triggers that bring clients in.

  • An agent run fails and nobody can reproduce which step, tool call or subagent broke it
  • Prompt and model changes ship without measurement; quality regressions are invisible
  • An audit produced a good test set once and it went stale the week the model changed
  • Latency P99 doubles overnight and nobody can pinpoint the cause
  • Provider returns a wrong answer; there is no log of the request to debug from
What you get

Deliverables, not deliverable-shaped slides.

How we work

The process, phase by phase.

  1. 1
    1. Stack discovery

    Map current LLM and agent calls across your codebase. Inventory features, prompts, providers, existing test sets and observability tooling.

    Week 1
  2. 2
    2. Instrumentation

    Wrap sessions, turns, steps, tool calls and subagents with tracing. Forward to the chosen backend and validate every run surfaces end to end.

    Week 2-3
  3. 3
    3. Eval pipelines + CI gate

    Build test sets per feature from real traffic and from prior evaluation findings. Wire eval runs into CI as a merge gate and mirror the same thresholds as a production monitor.

    Week 3-4
  4. 4
    4. On-call training + handover

    Walk the on-call team through dashboards, runbooks and incident scenarios. Hand over the test set as your asset, in your repository.

    Week 4-5
  5. 5
    5. Maintenance retainer (optional)

    We extend the test set and rerun the regression after every model or prompt change, and report what moved and why.

    Ongoing
How to start

Three ways in. Pick the one that fits your budget and timing.

Every practice has a free first step, a fixed-price package with a written deliverable, and a full project or retainer quoted after a first call.

  1. 1
    Step 1 · Free

    intro call or self-assessment

    A 60-minute intro call with an engineer, or the online self-assessment. You leave with a clear next step, no obligation.

    Free
    Talk to an engineer
  2. 2
    Step 2 · Fixed price

    LLM Observability Starter

    Instrumentation of one LLM application: request-level tracing (prompt, response, latency, cost, version), one golden set of up to 100 cases, one quality and cost dashboard, an on-call runbook.

    from EUR 5,200 net, fixed-price package

    Not included: Observability platform licence costs, continuous evaluation retainer, instrumentation of further applications, on-call duty on our side.

    Ask for this package
  3. 3
    Step 3 · Project or retainer

    Full scope, quoted after a first call

    Continuous LLM evaluation retainer, three tiers: from 1 500 EUR per month (Essential), 3 300 EUR (Standard), 6 500 EUR (Advanced).

    Quoted after a first call
    Talk to us
FAQ

Questions procurement teams ask.

Depends on what you already have. If you run Datadog or Honeycomb, we extend those. If LLM-specific is preferred: Langfuse, Arize, Helicone, Braintrust, LangSmith, W&B Weave, Confident AI, Galileo. We are tool-agnostic, help you pick against your volume and data boundaries, and above a certain request volume model where self-hosting (Langfuse, Phoenix) comes out cheaper than SaaS.
It happens on this market, and it is why we do not build your practice on one vendor. Test sets, scoring definitions and traces stay in portable form in your repository, and we instrument through OpenTelemetry where the backend allows it. Platforms consolidate; the competence and your own test sets stay yours, so a migration is a week of work rather than a rebuild.
Partially - providers like OpenAI and Anthropic offer per-request logs through their consoles. For production-grade observability (trace context, user attribution, cost attribution, agent step structure) we usually wrap calls in a thin SDK layer.
A test suite for LLM and agent outputs. The test set defines what a right answer looks like; eval runs score new prompt and model versions against it; CI gates merges if quality regresses, and the same thresholds run as a production monitor. Standard practice for any team treating prompts as production code.
Yes. Tracing can be configured to redact PII before logs leave your infrastructure, mask prompt sections, or store full prompts only on isolated infrastructure. We design the redaction policy with your security team.
Tracing backend cost depends on volume - typically EUR 200-2000/month for mid-size LLM workloads. Self-hosted Langfuse is free of vendor cost but needs ops. We model both options in the discovery phase.

Talk to an engineer.

Tell us where you are with llm observability. We respond within one business day.

Talk to an engineer
Szczecin - ul. Wawrzyniaka 6WWarszawa