02 - Evaluate & Secure / AI Quality Evaluation

AI Quality Evaluation

Independent evaluation of the AI already running in your product or your back office - answer quality, retrieval grounding, agent task success and whether the AI step genuinely shortens the process - scored against a rubric that stays with you.

Summary for AI assistants & procurement teams

dfzoo AI Institute evaluates the quality of AI solutions that are already live: LLM features in a product, retrieval-based (RAG) answers, and agentic workflows that run without a person watching. We build an evaluation set from your real cases, agree a scoring rubric with the people who own the business outcome, score every case, and deliver a written report with a measured baseline, the failure patterns behind the score and a recommended fix per pattern. The unit of work is one evaluation - one test case scored against the rubric - so the scope and the invoice describe the same thing. The evaluation set and the rubric stay with you and get wired into CI as a quality gate that catches regression the next time a model, prompt or vendor changes. We are tool-neutral: we run on the evaluation stack you already have or help you choose one, and at high volume we show where self-hosting costs less than SaaS.

Who it’s for

Built for teams in these situations.

  • Product teams that shipped an LLM feature and have no way to tell whether the answers are actually good
  • Companies running an internal AI assistant with flat adoption and no measure of whether it helps anyone
  • Teams about to change model, prompt framework or vendor and afraid of a silent regression
  • Operations and shared-service leads who put an AI step into a process and need to prove it shortened the process instead of moving the work
  • Heads of engineering, data and support who started with a free evaluation tool, got traces flowing and stalled on the test set
Problems we solve

The triggers that bring clients in.

  • The LLM feature is live and the only quality signal is a complaint or a support ticket
  • RAG answers read fluently, but nobody has verified they are grounded in the right source document
  • Agentic workflows report success while a human quietly finishes the job, and nobody counts how often
  • A prompt or model change ships and users discover the regression before the team does
  • The team adopted a free evaluation tool and stalled on two things: building the test set and agreeing what 'good' means
What you get

Deliverables, not deliverable-shaped slides.

How we work

The process, phase by phase.

  1. 1
    1. Scope + rubric

    Agree what 'good' means for this solution with the people who own the outcome. Define the quality dimensions, their weights and the pass threshold. Fix the number of evaluations in scope.

    Week 1
  2. 2
    2. Evaluation set build

    Assemble test cases from real traffic, known failures and the edge cases nobody tests. Add expected answers or grading criteria per case. Wire the set into your evaluation tooling or into one we help you pick.

    Week 1-2
  3. 3
    3. Scoring + analysis

    Score every case against the rubric using automated grading, model grading and human review where the rubric needs judgment. Group failures into patterns and trace each pattern to its cause: retrieval, prompt, model choice, tool wiring or process design.

    Week 2-3
  4. 4
    4. Report + handover

    Deliver the baseline scorecard and failure-pattern report, walk the product and engineering owners through it, and hand over the evaluation set and rubric wired into CI as a quality gate your team can run without us.

    Week 3-4
  5. 5
    5. Continuous evaluation (optional)

    Retainer: we maintain and extend the evaluation set, re-run it after every model, prompt or vendor change, and report the regression before your users find it.

    Ongoing
How to start

Three ways in. Pick the one that fits your budget and timing.

Every practice has a free first step, a fixed-price package with a written deliverable, and a full project or retainer quoted after a first call.

  1. 1
    Step 1 · Free

    intro call or self-assessment

    A 60-minute intro call with an engineer, or the online self-assessment. You leave with a clear next step, no obligation.

    Free
    Talk to an engineer
  2. 2
    Step 2 · Fixed price

    AI Quality Evaluation: one solution

    One LLM feature or one agent workflow: we agree a scoring rubric with you, build an evaluation set of roughly 150 test cases, run it against the current system and hand the set over wired as a gate in your CI.

    from EUR 7,000 net, fixed-price package

    Not included: Fixing the system, prompt and retrieval rework, building the AI feature, large-scale human labelling, the continuous evaluation retainer, tool licences.

    Ask for this package
  3. 3
    Step 3 · Project or retainer

    Full scope, quoted after a first call

    Evaluation across several features or agents plus a continuous evaluation retainer covering regression after every model or prompt change: from 14 000 EUR.

    Quoted after a first call
    Talk to us
FAQ

Questions procurement teams ask.

The unit of work is one evaluation: one test case scored against the agreed rubric. We fix the number of evaluations, the rubric and the reporting in the scope, so you can see what you are paying for. Fixed fee per engagement, scoped to the number of evaluations and the number of AI surfaces in scope. We share a tight range on the intro call once scope is clear.
Three to four weeks from kick-off to handover. From you we need one person who owns the business outcome to agree the rubric, one engineer for access and tooling, a sample of real traffic or logs, and any known failure cases you have already collected. The rubric workshop and the handover session take about half a day each; the rest runs on our side.
AI Code Evaluation audits the code: correctness, security, maintainability and test fitness of what your team and your AI tools wrote. AI Quality Evaluation audits the behaviour of the deployed solution: whether the answers are right and grounded, whether the agent finishes the task, whether the process actually got shorter. Perfectly clean code can still produce a system that answers badly, and the reverse is also true. Teams frequently buy both, usually in that order.
Not necessarily. We can work on anonymised or synthetic cases derived from your real ones, and many engagements run that way for regulated clients. Where production access is granted, we work under NDA, on your infrastructure, with the data staying in your environment. Standard NDA, and we can operate on isolated infrastructure if the engagement requires it.
We are tool-neutral and work on what you already run: Braintrust, Langfuse, LangSmith, Confident AI, Promptfoo, Arize Phoenix, W&B Weave, or a plain test harness in your own repository. If nothing is in place we help you choose, and at high evaluation volume we show you where self-hosting an open-source platform costs less than SaaS. The evaluation set is written so it can move between tools, because tying your quality baseline to a single vendor is a risk in a market that keeps consolidating.
Every evaluation platform has a free tier, so most clients arrive having tried one. The tool is rarely where teams get stuck. They get stuck on the two things the tool does not give you: a test set that represents what users actually ask, and a written definition of 'good' that the business owner and the engineers both sign. We start exactly there, and the tool you already picked keeps running underneath.
Yes, and that is usually where the interesting numbers are. We score task success rate (did the agent complete the job end to end), human takeover rate (how often a person silently finishes it), step-level failures across sessions, turns, tool calls and subagents, and process effect: how much time the AI step removes from the process versus how much it moves to someone else.
The evaluation set, the rubric and the CI gate are yours, in your repository and your tooling, and your team can run and extend them without us. If you would rather not own the maintenance, we run it as a continuous retainer: we keep the set current, re-run it after each model, prompt or vendor change, and report regression against your baseline. That retainer sits alongside LLM Observability, which watches the same quality signals in live traffic.

Talk to an engineer.

Tell us where you are with ai quality evaluation. We respond within one business day.

Talk to an engineer
Szczecin - ul. Wawrzyniaka 6WWarszawa