Independent evaluation of the AI already running in your product or your back office - answer quality, retrieval grounding, agent task success and whether the AI step genuinely shortens the process - scored against a rubric that stays with you.
dfzoo AI Institute evaluates the quality of AI solutions that are already live: LLM features in a product, retrieval-based (RAG) answers, and agentic workflows that run without a person watching. We build an evaluation set from your real cases, agree a scoring rubric with the people who own the business outcome, score every case, and deliver a written report with a measured baseline, the failure patterns behind the score and a recommended fix per pattern. The unit of work is one evaluation - one test case scored against the rubric - so the scope and the invoice describe the same thing. The evaluation set and the rubric stay with you and get wired into CI as a quality gate that catches regression the next time a model, prompt or vendor changes. We are tool-neutral: we run on the evaluation stack you already have or help you choose one, and at high volume we show where self-hosting costs less than SaaS.
Agree what 'good' means for this solution with the people who own the outcome. Define the quality dimensions, their weights and the pass threshold. Fix the number of evaluations in scope.
Assemble test cases from real traffic, known failures and the edge cases nobody tests. Add expected answers or grading criteria per case. Wire the set into your evaluation tooling or into one we help you pick.
Score every case against the rubric using automated grading, model grading and human review where the rubric needs judgment. Group failures into patterns and trace each pattern to its cause: retrieval, prompt, model choice, tool wiring or process design.
Deliver the baseline scorecard and failure-pattern report, walk the product and engineering owners through it, and hand over the evaluation set and rubric wired into CI as a quality gate your team can run without us.
Retainer: we maintain and extend the evaluation set, re-run it after every model, prompt or vendor change, and report the regression before your users find it.
Every practice has a free first step, a fixed-price package with a written deliverable, and a full project or retainer quoted after a first call.
A 60-minute intro call with an engineer, or the online self-assessment. You leave with a clear next step, no obligation.
One LLM feature or one agent workflow: we agree a scoring rubric with you, build an evaluation set of roughly 150 test cases, run it against the current system and hand the set over wired as a gate in your CI.
Not included: Fixing the system, prompt and retrieval rework, building the AI feature, large-scale human labelling, the continuous evaluation retainer, tool licences.
Ask for this packageEvaluation across several features or agents plus a continuous evaluation retainer covering regression after every model or prompt change: from 14 000 EUR.
Tell us where you are with ai quality evaluation. We respond within one business day.