# dfzoo AI Institute - full site content (llms-full.txt) **Canonical URL:** https://dfzoo.ai/llms-full.txt **Site:** https://dfzoo.ai **Languages:** English (canonical), Polish (drafts pending OpenAI polish) **Generated at:** build time from src/lib/anchors.{en,pl}.ts This file contains the complete markdown content of every public page on dfzoo.ai. EN content first (canonical, locked), PL content second (drafts produced by Claude; pre-publication polish by OpenAI offline, may shift slightly before launch). Ingest, retrieve, cite freely. For the manifest / TOC see https://dfzoo.ai/llms.txt. ================================================================ ENGLISH CONTENT (canonical) ================================================================ # dfzoo AI Institute - full site content - English (canonical) _Auto-generated from anchors.en.ts at each build. Canonical EN copy._ --- # Homepage - dfzoo AI Institute **URL:** https://dfzoo.ai/ > Independent AI assurance, code evaluation and engineering training. dfzoo AI Institute helps organizations evaluate AI-generated code, train teams, implement AI tools, automate processes and modernize legacy systems into AI-ready platforms - without vendor lock-in. - ISO 9001:2015 certified (DNV Business Assurance, cert no. 10000439701-MSC-RvA-POL) - Offices: Szczecin (HQ, ul. Wawrzyniaka 6W) and Warszawa - Languages: English and Polish - Response time: one business day The site is organized around four super-anchors (Train & Adopt, Evaluate & Secure, Build & Modernize, Operate & Measure), under which sit 18 concrete service practices. --- # About dfzoo AI Institute **URL:** https://dfzoo.ai/about > An engineering institute for independent AI assurance. dfzoo AI Institute provides independent assurance for production AI: security review, code evaluation, team training, and operational support. The Institute is operated by DFZOO.COM Sp. z o.o. - a Polish company in business since 2014, ISO 9001:2015 certified by DNV Business Assurance (cert. 10000439701-MSC-RvA-POL), operating across the EU under VAT-EU. ### Public transparency: site built and maintained by Claude under engineering team direction The dfzoo AI Institute website you are reading - the code, copy, structured data, routing, deploys - is the working artifact of the engagement methodology we offer to clients. The same development workflow, code review process, and deployment discipline applied here are documented in our service procedures. ### Three things to know - **Engineering-led** - Consultants are engineers. Trainers ship code. Evaluations are written by the same engineers who built the systems under review. Sales and delivery are not separated. - **ISO 9001:2015 certified** - Quality management certified by DNV Business Assurance (cert. 10000439701-MSC-RvA-POL). A procurement-grade signal of repeatability and auditability. - **Production-grade outcomes** - Engagement outcomes are measured against agreed production KPIs: cost per request, P95 latency, error rates, deployment frequency. ### Legal entity dfzoo AI Institute is operated by DFZOO.COM Sp. z o.o. - a Polish private limited company in business since 2014, registered for VAT-EU, headquartered in Szczecin with a Warszawa office. ### ISO 9001:2015 scope - Design and development of web and mobile applications - E-commerce and websites - Application service - IT training and consulting - Online marketing - Robotic automation of business processes ### Offices - **Szczecin** (HQ): ul. Wawrzyniaka 6W, 70-393 Szczecin, Poland. Engineering, evaluation and operations. - **Warszawa**: client-facing. Intro calls, workshops, on-site delivery for clients in central and northern Poland. --- # Contact - dfzoo AI Institute **URL:** https://dfzoo.ai/contact > Talk to an engineer. We respond within one business day, in English or Polish. ### Offices - **Szczecin (HQ)**: ul. Wawrzyniaka 6W, 70-393 Szczecin, Poland. Office hours: Mon-Fri, 09:00-17:00 CET. Visits by appointment. - **Warszawa**: by appointment. Client-facing: Intro calls, workshops and on-site delivery for clients in central and northern Poland. ### Contact channels - **Sales & new engagements**: start@dfzoo.ai - intro calls, scoping, RFP responses, procurement - **Press & partnerships**: start@dfzoo.ai - media inquiries, speaking, partnerships ### Business & legal - Legal name: DFZOO.COM Sp. z o.o. (brand: dfzoo AI Institute) - VAT-EU: PL (TBD) - KRS / Court register: TBD - Registered address: ul. Wawrzyniaka 6W, 70-393 Szczecin, Poland --- # Services - dfzoo AI Institute **URL:** https://dfzoo.ai/services > Four super-anchors of the dfzoo AI Institute practice. Each anchor groups the practices we deliver under it. Every engagement at dfzoo AI Institute lives under one of four super-anchors. Pick the situation you're in - the practices that solve it are grouped underneath. --- # Train, Adopt & Govern (anchor 01) **URL:** https://dfzoo.ai/services/train-adopt-govern **Highlight:** Funded pathways (BUR/PARP) > Move your organization from AI curiosity to repeatable, governed AI practice - with training, advisory and the basic governance rails that keep production AI safe. ### Summary dfzoo AI Institute helps organizations adopt AI as a measurable team capability, not a one-off experiment. We run AI training programs for developers, designers and business teams; advise leadership on adoption roadmaps and tooling choices; implement the day-to-day AI tooling that teams actually use; and put in place the governance basics - usage policies, evaluation cadences, risk register - that procurement and legal need to sign off. ISO 9001:2015 certified. Selected programs are eligible for funded development pathways (BUR/PARP) in Poland. ### Who it's for - Engineering leaders rolling out AI coding assistants across their teams - L&D and HR leaders procuring practical AI training under BUR/PARP funding - Heads of operations standing up an internal AI center of excellence - Founders who need their team productive with AI tools in weeks, not quarters ### Problems we solve - Teams are using AI tools ad-hoc with no shared standards or evaluation - Procurement and legal block AI rollout because there is no governance baseline - Training budget is allocated but no provider can prove production impact - Tools are licensed but adoption stalls past the early enthusiasts ### Practices under this anchor - AI Training Programs - https://dfzoo.ai/services/train-adopt-govern/ai-training-programs - AI Advisory - https://dfzoo.ai/services/train-adopt-govern/ai-advisory - Tools Implementation - https://dfzoo.ai/services/train-adopt-govern/tools-implementation - AI Governance Basics - https://dfzoo.ai/services/train-adopt-govern/ai-governance-basics - EU AI Act Technical Compliance - https://dfzoo.ai/services/train-adopt-govern/eu-ai-act-compliance ### FAQ **Q: Can the training be publicly funded?** Selected AI training and advisory services may be delivered through funding-supported development pathways (BUR/PARP), depending on current eligibility and operator rules. We help you check. **Q: Do you train developers or business users?** Both. Engineering tracks focus on AI-assisted development workflows, code evaluation and production patterns. Business tracks focus on workflow automation, prompt design and safe use of AI tools day to day. **Q: What governance do you put in place?** The procurement-ready baseline: an AI usage policy, an approved-tools list, an evaluation cadence for outputs that touch production, and a basic risk register. Not a heavy framework - enough for legal and security to sign off. **Q: How long does a typical engagement take?** Workshops are 1-3 days. Advisory engagements typically span 4-8 weeks. Tools implementation runs 2-6 weeks depending on team size and the toolchain. **Q: What outcomes do you measure?** Adoption rate per team, average AI-assist time saved per developer-week, percentage of PRs touched by AI tools, governance compliance pass rate. We agree the specific metrics on the intro call. --- ## AI Training Programs - Train & Adopt **URL:** https://dfzoo.ai/services/train-adopt-govern/ai-training-programs **Service type:** Professional training **Anchor:** Train, Adopt & Govern (Train & Adopt) > Role-specific AI training that teams actually apply at work - engineering, design, operations and leadership tracks, workshop or multi-week cohort format. ### Summary dfzoo AI Institute delivers AI training programs designed by practitioners who ship AI in production. Engineering tracks cover AI-assisted development with current tools (Claude Code, Cursor, Copilot, Aider), code evaluation patterns and production safeguards. Business tracks cover prompt design, workflow automation and safe use of AI tools day to day. Leadership tracks cover adoption roadmaps and ROI measurement. Programs run as 1-3 day workshops or 4-6 week cohorts; selected programs are eligible for BUR/PARP funding in Poland. ### Who it's for - Engineering managers rolling out AI coding assistants across their teams - L&D heads procuring practical AI training under BUR/PARP funding rules - Product designers and PMs upgrading their day-to-day AI workflow - C-suite leaders building company-wide AI literacy ### Problems we solve - Generic vendor AI training does not translate to your codebase or workflows - Engineers use AI tools but cannot tell colleagues what good looks like - Training budget is allocated but no provider has shipped what they teach - Cohort vs workshop format choice unclear given team size and goals ### What you get - Role-specific curriculum tuned to your stack, tools and use cases - Live workshop delivery (in-person Warszawa/Szczecin or remote) - Hands-on lab exercises tied to your actual codebase or workflows - Recorded sessions and a written reference playbook per track - Post-program check-in 4-6 weeks later to measure adoption ### How we work 1. **1. Scope call** (Week 0): Confirm audience, current AI tool usage, business goals and funding eligibility. 2. **2. Curriculum design** (Week 1-2): Tailor agenda, choose tools to cover, prepare codebase-specific lab material. 3. **3. Delivery** (1-3 days (workshop) or 4-6 weeks (cohort)): Run the workshop or cohort. Hands-on labs, Q&A, working sessions with real client problems. 4. **4. Adoption check-in** (Week +4-6): Measure post-training usage, identify blockers, recommend next-step interventions. ### FAQ **Q: What is the difference between workshop and cohort format?** Workshops compress an intensive curriculum into 1-3 consecutive days - best for kicking off adoption fast. Cohorts spread across 4-6 weeks with weekly sessions and homework - best for changing day-to-day habits in larger teams. **Q: Can the training cover our actual codebase?** Yes. Lab exercises are built from your repos under NDA. Engineers learn AI-assisted patterns directly against the code they will work on tomorrow. **Q: Is BUR/PARP funding available?** Selected programs qualify for funded development pathways (BUR/PARP) in Poland. Eligibility depends on company size, sector and current operator rules. We help check eligibility on the intro call. **Q: What tools do you cover?** Current production-grade AI coding assistants: Claude Code, Cursor, Copilot, Aider, Windsurf. We also cover prompt design and LLM workflow tooling (n8n, Zapier with LLM, custom agents) for business tracks. **Q: How do you measure if the training worked?** Pre/post adoption survey, AI-assist time tracked per developer-week, percentage of PRs touched by AI tools, and a post-program code-quality check-in. Specific metrics agreed on the scope call. ### Related - https://dfzoo.ai/services/train-adopt-govern/ai-advisory - https://dfzoo.ai/services/train-adopt-govern/tools-implementation - https://dfzoo.ai/services/evaluate-audit-secure/ai-code-evaluation --- ## AI Advisory - Train & Adopt **URL:** https://dfzoo.ai/services/train-adopt-govern/ai-advisory **Service type:** Management consulting **Anchor:** Train, Adopt & Govern (Train & Adopt) > Independent advisory for organizations sequencing their AI adoption - what to pilot first, what tools to buy, what to defer and how to measure each step. ### Summary dfzoo AI Institute provides independent AI advisory to executives, CTOs and heads of operations. We assess your current AI landscape (tools licensed, pilots run, governance in place), recommend a 6-12 month adoption roadmap with sequenced bets, and help your team make tooling decisions backed by hands-on evaluation. Output is a written roadmap with rationale, a tooling assessment matrix and a quarterly executive review cadence. ### Who it's for - CTOs deciding which AI initiatives to fund this fiscal year - Heads of operations standing up an internal AI center of excellence - CEOs of small-mid companies setting an AI strategy without a CTO - Boards asking for an independent read on management's AI plan ### Problems we solve - Too many AI tools to evaluate, no time to pilot them all - Internal AI champions disagree on the roadmap and leadership cannot adjudicate - Procurement asks for an independent second opinion on a major AI contract - Pilot projects succeed but never scale into production usage ### What you get - Written 6-12 month AI adoption roadmap with sequenced bets - Tooling assessment matrix (cost, fit, risk, switching cost) for top candidates - Stakeholder workshop summarizing recommendations and next steps - Quarterly executive review for the duration of the engagement - Optional: vendor-selection support for a specific high-stakes contract ### How we work 1. **1. Landscape audit** (Week 1-2): Interview key stakeholders, inventory current AI tools and pilots, review governance state. 2. **2. Roadmap drafting** (Week 2-4): Map options against your business goals, risk tolerance and team capacity. Sequence bets. 3. **3. Stakeholder workshop** (Week 4-5): Present recommendations to leadership, iterate based on feedback, lock the roadmap. 4. **4. Quarterly reviews** (Quarterly (ongoing)): Track progress against roadmap milestones, adjust for new tools and market shifts. ### FAQ **Q: Are you independent of tool vendors?** Yes. We do not take referral fees from AI tool vendors. Our recommendations are based on hands-on evaluation against your stack and goals. **Q: How long is a typical advisory engagement?** Initial roadmap delivery is 4-8 weeks. Quarterly reviews continue afterwards as long as the engagement runs. **Q: Do you work with non-technical leadership?** Yes. We translate engineering trade-offs into business language. Many engagements start with a CEO or COO who needs a CTO-level read without a CTO. **Q: Can you help with a specific vendor selection?** Yes. We can scope the engagement around a single high-stakes contract decision (an LLM platform, an AI coding assistant license, an agentic framework). Output is a written recommendation with rationale. **Q: What happens after the roadmap is delivered?** Most clients opt into a quarterly review retainer to keep the roadmap current as the AI landscape shifts. Others bring us back ad-hoc for specific decisions. ### Related - https://dfzoo.ai/services/train-adopt-govern/ai-governance-basics - https://dfzoo.ai/services/train-adopt-govern/ai-training-programs - https://dfzoo.ai/services/build-automate-modernize/agentic-ai-readiness --- ## Tools Implementation - Train & Adopt **URL:** https://dfzoo.ai/services/train-adopt-govern/tools-implementation **Service type:** IT consulting and implementation **Anchor:** Train, Adopt & Govern (Train & Adopt) > Hands-on rollout of AI coding assistants, context files and team workflows - so the tools you licensed get used past week one, on rules your repository actually carries, with adoption you can read off a dashboard. ### Summary dfzoo AI Institute implements AI tooling for engineering teams: setting up AI coding assistants (Claude Code, Cursor, Copilot) with team-specific configurations, writing the context files agents read before they touch your repository, building shared prompt libraries, and integrating tools with your IDE, version control and CI. We separate writing from reviewing: an agent should not review the code it just wrote, so we set up a distinct review context and a review gate independent of the agent that produced the change. Every rollout ends with a measurement dashboard - adoption rate, acceptance rate, and how both track against your delivery metrics - not with a configuration file. We work side-by-side with your team for 2-6 weeks until usage is steady-state. ### Who it's for - Engineering managers whose team licensed AI tools but adoption stalled - VPs of engineering rolling out a new coding assistant across multiple teams - Platform teams building internal AI tooling for the wider engineering org - Tech leads who need help configuring tools for their specific stack ### Problems we solve - Default tool configurations do not match your codebase or conventions - Agents get no written context, so every session rediscovers your conventions and guesses the rest - The same agent writes the change and reviews it, so the review gate is confirming its own work - AI tools and CI/CD do not talk to each other; review workflow is broken - Nobody can say whether the tools are used, accepted or paying off - there is no number to point at ### What you get - Context files for agents: repository conventions, a context map of the codebase, review rules, plus the procedure and owner for keeping them current - Team-specific configuration for your AI coding assistant of choice - Shared prompt library with the most-used patterns for your stack - Writer/Reviewer setup: a separate review context and a review gate independent of the agent that wrote the change - IDE + version control + CI integration docs and example workflows - Onboarding playbook for new engineers joining a team using these tools - Adoption dashboard: adoption rate, acceptance rate and correlation with your delivery metrics, read together 30 days after rollout ### How we work 1. **1. Tool + stack discovery** (Week 1): Audit current tool licenses, codebases, IDE preferences and existing prompts. Agree which delivery metrics the rollout is measured against. 2. **2. Context files + configuration** (Week 2-3): Write the context files, set up tool-specific configs, build the prompt library, wire CI/CD and the separate review gate. 3. **3. Team rollout** (Week 3-5): Hands-on rollout sessions per team, pair-programming with engineers, debug issues live. 4. **4. Measurement + adoption review** (Week +6): Stand up the dashboard, read the first 30 days of data, identify holdouts, recommend next steps, hand the dashboard over. ### FAQ **Q: Which AI coding assistants and IDEs do you implement?** Claude Code, Cursor, Copilot, Windsurf, Aider, Cody - whichever your team licensed or wants to evaluate, across VS Code, JetBrains, Cursor and Neovim. Configs are tested per IDE and shared as committable team files. We are tool-agnostic; our value is configuration and adoption, not vendor referral. **Q: What exactly are context files, and who owns them afterwards?** The files an agent reads before it writes anything: repository conventions, a context map of where things live and why, and the review rules a change is held to. They live in your repository and you own them, handed over with a maintenance procedure saying who updates them, when, and what triggers a rewrite. **Q: Can the same agent write the code and review it?** It can, and that is the failure we design out. An agent reviewing its own output repeats its own assumptions, so the review gate runs in a separate context with different rules and is never the writer: a second agent configuration, plus a human on anything touching production. **Q: How does this compare to AI training?** Training teaches concepts and patterns. Tools Implementation is hands-on: we configure your specific tools, write your specific context files and prompts, integrate with your specific CI. Most clients buy both - training first, implementation second. **Q: What about security and compliance?** We configure tools to respect your data boundaries (no code leakage to public LLMs without consent, allowlists for which repos can be touched). Security review and compliance baseline are covered under Security Review and AI Governance Basics. ### Related - https://dfzoo.ai/services/train-adopt-govern/ai-training-programs - https://dfzoo.ai/services/evaluate-audit-secure/security-review - https://dfzoo.ai/services/build-automate-modernize/ai-powered-custom-development --- ## AI Governance Basics - Train & Adopt **URL:** https://dfzoo.ai/services/train-adopt-govern/ai-governance-basics **Service type:** Governance and compliance consulting **Anchor:** Train, Adopt & Govern (Train & Adopt) > The procurement-ready AI governance baseline: an AI usage policy, an approved-tools list, an evaluation cadence and a risk register - enough to unblock legal and security sign-off. ### Summary dfzoo AI Institute establishes the minimum viable AI governance baseline for organizations rolling out AI internally. We write an AI usage policy aligned to your industry, build an approved-tools list with data-handling rules, set up an evaluation cadence for AI outputs that touch production, and create a basic risk register that legal and security teams accept. Not a heavy framework - the smallest set of artifacts that unblocks AI rollout without creating shelf-ware. ### Who it's for - Heads of legal and compliance asked to bless an AI rollout - CISOs whose engineering teams want to license AI coding assistants - COOs standing up internal AI use across non-engineering departments - Public-sector and regulated organizations needing a defensible baseline ### Problems we solve - Engineering wants to use AI tools but legal cannot find a precedent to approve - Each team writes its own AI usage rules - inconsistent and unenforceable - Risk register has no entries for AI; auditors flag the gap - Evaluation of AI outputs is informal; nobody can prove quality at audit time ### What you get - Written AI Usage Policy (1 doc, covers staff + contractors) - Approved Tools List with data-handling rules per tool - Evaluation cadence playbook (which AI outputs are reviewed, how, by whom) - Risk register entries for AI use, mapped to your existing risk framework - Quick-reference one-pager for managers to use day to day ### How we work 1. **1. Industry + risk discovery** (Week 1): Interview legal, compliance and engineering leads; understand current risk framework and industry constraints. 2. **2. Draft artifacts** (Week 2-3): Draft AI usage policy, approved-tools list, evaluation cadence and risk register entries. 3. **3. Stakeholder review** (Week 3-4): Walk legal, compliance, security and engineering through the drafts. Iterate. 4. **4. Rollout + training** (Week 4-5): Train managers on day-to-day application, publish the artifacts internally. ### FAQ **Q: Is this a full AI risk management framework like NIST AI RMF?** No. This is the procurement-ready minimum - enough to unblock AI use and pass standard audits. Organizations with mature risk programs may layer NIST AI RMF, ISO 42001 or similar on top later. We can scope that as a follow-up. **Q: Will the policy account for our industry rules?** Yes. We tailor policies to your industry - fintech (KNF/MiFID/DORA-relevant), healthcare (GDPR / health data), public sector, regulated tech. We are not lawyers but we draft artifacts that your legal team can stamp. **Q: How does this differ from buying a generic AI policy template?** Generic templates are written for the abstract case and break on first contact with real procurement questions. We tailor to your tooling, your industry and your specific procurement reality. Output is defensible at audit, not just a checkbox. **Q: Can you implement evaluation tooling, not just write the policy?** Yes. The evaluation cadence playbook can be paired with tooling implementation under LLM Observability (eval pipelines, automated scoring) or under AI Code Evaluation (recurring code audits). **Q: How long until artifacts are signed off?** Typical engagement is 4-6 weeks from kickoff to published, signed-off artifacts. The bottleneck is usually internal review cycles, not drafting. ### Related - https://dfzoo.ai/services/train-adopt-govern/ai-advisory - https://dfzoo.ai/services/evaluate-audit-secure/security-review - https://dfzoo.ai/services/operate-measure-maintain/llm-observability --- ## EU AI Act Technical Compliance - Train & Adopt **URL:** https://dfzoo.ai/services/train-adopt-govern/eu-ai-act-compliance **Service type:** EU AI Act technical compliance assessment and documentation **Anchor:** Train, Adopt & Govern (Train & Adopt) > The engineering half of AI Act compliance: we inventory and classify your AI systems, write the technical documentation, produce the testing evidence and turn logging, traceability and human oversight into architecture your team actually runs. ### Summary dfzoo AI Institute delivers the technical side of EU AI Act compliance for organizations that build, buy or deploy AI systems. We inventory every AI system in use, classify each one against the Act's risk tiers, write the technical documentation the current obligations require, and produce adversarial-testing evidence that stands up to an audit. Logging, traceability and human oversight are treated as architectural requirements with named owners and working implementations, not as promises in a policy document. The engagement ends with a gap analysis and a costed remediation roadmap, plus an AI literacy plan that feeds directly into training for the people who operate these systems. We are engineers, not lawyers: we build the technical evidence and work alongside your legal counsel or our legal partner, who owns the legal interpretation and sign-off. ### Who it's for - Product startups whose AI feature was just classified as high-risk by a customer's procurement questionnaire - Software houses and system integrators who must hand clients AI Act documentation for what they deliver - Corporate and public-sector teams deploying purchased AI systems and carrying deployer obligations they did not write - Regulated organizations in finance, health, HR, education and critical infrastructure with audit deadlines in front of them - Compliance leads, CISOs, CTOs and heads of legal who need technical artifacts their counsel can sign ### Problems we solve - Nobody can produce a list of the AI systems your organization actually runs, let alone their risk classification - A client, tender or auditor asked for AI Act technical documentation and there is nothing written down - Your systems log business events but not the model decisions, inputs and versions that traceability requires - Human oversight exists on a slide but no interface, role or escalation path implements it in production - Legal counsel gave you the obligations and the engineering team has no idea what to build in response ### What you get - AI system inventory with owners, purpose, data flows, model and vendor per system - Risk classification per system against the Act's tiers, with the reasoning written out for audit and legal review - Technical documentation pack per in-scope system, structured to the Act's requirements - Logging, traceability and human oversight specification, with the architectural changes needed to meet it - Evaluation evidence file drawn from quality evaluation and code security review of the system, where required - Gap analysis and costed remediation roadmap, plus an AI literacy plan mapped to roles ### How we work 1. **1. Inventory and discovery** (Week 1-2): Find every AI system in use: built, bought and embedded in third-party tools. Capture purpose, users, data, model, vendor and owner for each. 2. **2. Risk classification** (Week 2-3): Classify each system against the Act's risk tiers and your role for it, provider or deployer. Document the reasoning so legal counsel can confirm or challenge it. 3. **3. Gap analysis** (Week 3-4): Compare current state against the obligations that apply: documentation, data governance, logging, traceability, human oversight, accuracy and robustness, testing evidence, AI literacy. 4. **4. Documentation and evidence** (Week 4-7): Write the technical documentation pack, specify the logging and oversight architecture, and assemble the testing evidence file. Review with your legal counsel or our legal partner. 5. **5. Remediation roadmap and handover** (Week 7-8, review quarterly): Deliver the costed roadmap with owners and sequencing, brief leadership, and set the review cadence that keeps documentation current as systems change. ### FAQ **Q: Is this legal advice?** No, and we say so plainly. We are an engineering team. We produce the technical artifacts the Act requires - inventory, classification reasoning, documentation, logging and oversight design, testing evidence - and your legal counsel owns the interpretation and the sign-off. Where you have no counsel for this, we bring a legal partner into the engagement and split the work explicitly. **Q: The obligations keep moving. How do you handle that?** We work against the obligations in force at the time of the engagement and we date every classification and every document. Where the interpretation is genuinely open we flag it as a decision for your counsel rather than guessing, and we always recommend legal confirmation before you rely on a classification externally. The quarterly review keeps the pack current as systems and guidance change. **Q: How long does it take and how is it priced?** A typical engagement runs 6-8 weeks from kickoff to a delivered documentation pack and roadmap. Pricing is fixed per engagement, scoped after a first call, and driven by the number of AI systems in scope and how many of them land in the high-risk tier. Inventory and classification alone can be bought as a smaller first step. **Q: What do you need from us to start?** Access to the people who own each AI system, existing architecture and data-flow documentation, vendor contracts and DPAs for purchased AI, and any classification or legal analysis already done. We run structured interviews to fill the gaps - most organizations discover during this step that the real inventory is larger than the one on file. **Q: We only use AI systems built by vendors. Does this still apply?** Yes. Deploying an AI system carries its own obligations, separate from those of the organization that built it. We assess what your vendors actually give you, identify what is missing from their documentation, and write the deployer-side artifacts: oversight design, logging, instructions for use, and the questions to put to the vendor before renewal. **Q: How does this relate to testing and evaluation of the system?** Testing evidence is one of the technical artifacts that supports compliance for higher-risk systems. AI Quality Evaluation produces evaluation results and a scored test set; Security Review covers the code and the application layer. Both hand over findings in a form that drops straight into the documentation pack. You can buy any of them on its own; running them together removes duplicated discovery and produces one consistent evidence trail. **Q: What about the AI literacy obligation?** It is a real obligation and it is also the easiest one to close. We map roles to the level of AI understanding each needs, then feed that into a training plan. Our training programs deliver it, and the completion records become part of your compliance evidence. **Q: Remote or on-site, and can you work inside our environment?** Remote under NDA by default. Interviews and workshops run on-site where that gets faster answers, at the same price, and we can work entirely inside your own tooling and document systems when material cannot leave the organization. ### Related - https://dfzoo.ai/services/train-adopt-govern/ai-governance-basics - https://dfzoo.ai/services/evaluate-audit-secure/ai-quality-evaluation - https://dfzoo.ai/services/train-adopt-govern/ai-training-programs --- # Evaluate, Audit & Secure (anchor 02) **URL:** https://dfzoo.ai/services/evaluate-audit-secure **Highlight:** AI Code Evaluation - international niche > Independent evaluation of AI-generated code, security review and production-readiness checks - so teams ship faster with AI and know what is safe to release. ### Summary dfzoo AI Institute provides independent evaluation and assurance for AI-generated code and AI-assisted engineering outputs. We audit code produced by coding assistants for correctness, security and maintainability; run security reviews of AI-touched systems; design test coverage strategies for AI-assisted projects; and certify production readiness before release. Our AI Code Evaluation service is an international niche - we are one of the few engineering-led providers offering this with ISO 9001:2015 quality assurance backing. ### Who it's for - Engineering leaders shipping AI-assisted code who need an external assurance signal - CTOs of regulated companies (fintech, health, public sector) under scrutiny over AI in code - Security teams responsible for products built with AI coding assistants - Vendor management teams evaluating AI-touched deliverables from outside contractors ### Problems we solve - AI-generated code lands in PRs faster than humans can review it thoroughly - Compliance asks for proof that AI-assisted code meets the same bar as human-written code - Test coverage drops because AI tools generate plausible-looking but unverified code - Security teams cannot tell where AI contributions are concentrated in the codebase ### Practices under this anchor - AI Code Evaluation - https://dfzoo.ai/services/evaluate-audit-secure/ai-code-evaluation - Security Review - https://dfzoo.ai/services/evaluate-audit-secure/security-review - Test Coverage - https://dfzoo.ai/services/evaluate-audit-secure/test-coverage - Production Readiness - https://dfzoo.ai/services/evaluate-audit-secure/production-readiness - AI Quality Evaluation - https://dfzoo.ai/services/evaluate-audit-secure/ai-quality-evaluation ### FAQ **Q: What does an AI Code Evaluation engagement look like?** Fixed-fee, fixed-scope. We take a repo or set of PRs, sample AI-touched code, run our evaluation rubric (correctness, security, maintainability, test fitness), produce a written report with prioritized findings, and brief the team on the top issues. Typical turnaround is 2-4 weeks. **Q: Do you replace human code review?** No. We are an external assurance signal - periodic, independent, deeper than day-to-day review. Most clients run us on a quarterly cadence alongside their internal review process. **Q: What tools do you cover?** Any LLM-based coding assistant: Cursor, Claude Code, Cody, Copilot, Windsurf, Aider. The output is the artifact we evaluate, not the tool. **Q: Can you deliver under NDA for regulated industries?** Yes. We work under standard NDA, can operate on isolated infrastructure, and can deliver in-person at client offices for highly sensitive engagements. **Q: What does the report look like?** A structured PDF + prioritized issue tracker entries (or Linear, Jira, GitHub Issues). Each finding includes severity, reproduction, recommended fix and links to the offending code. We follow up on remediation if requested. --- ## AI Code Evaluation - Evaluate & Secure **URL:** https://dfzoo.ai/services/evaluate-audit-secure/ai-code-evaluation **Service type:** Software code audit and assurance **Anchor:** Evaluate, Audit & Secure (Evaluate & Secure) > Independent, fixed-fee audit of AI-generated and AI-assisted code - graded for correctness, security, maintainability and test fitness against an engineering rubric we publish in full before you sign. ### Summary dfzoo AI Institute provides independent AI Code Evaluation for engineering teams shipping AI-assisted code. We sample AI-touched files or PRs from your repository, evaluate each finding against a four-dimension rubric (correctness, security, maintainability, test fitness), produce a written report with severity-ranked findings, and brief the team on the top issues. The rubric is public: every dimension has a written definition, named sub-criteria and a 0-4 scale, so you see how the code is graded before you buy and can apply the same standard yourself between audits. ISO 9001:2015 quality assurance backing. Fixed fee, 2-4 week turnaround. One of the few engineering-led AI code evaluation practices in Europe. ### Who it's for - VPs of engineering and CTOs needing an external assurance signal on AI-assisted code - CISOs whose teams ship AI-generated code into regulated production systems - Vendor management teams reviewing AI-touched deliverables from contractors - Compliance officers required to demonstrate code review depth at audit ### Problems we solve - AI-generated code lands faster than humans can review it thoroughly - Compliance asks for proof AI-assisted code meets the same bar as human code - Audit findings arrive as opinions with no stated method, so nobody can repeat or contest them - Test coverage drops because AI tools generate plausible-looking but unverified code - Security teams cannot tell where AI contributions concentrate in the codebase ### What you get - Written audit report (structured PDF, ~30-60 pages depending on scope) - Per-dimension scorecard on the published 0-4 scale, with the anchor behind each score and the weighting agreed for your context - Severity-ranked finding list (Critical / High / Medium / Low) with reproductions - Per-finding recommended fix with code snippets and links to offending lines - Prioritized issue-tracker entries (Linear, Jira, GitHub Issues - your choice) - Executive briefing presentation with top issues and remediation roadmap ### How we work 1. **1. Scope + sampling** (Week 1): Define audit scope: which repos, which PR window, which tooling produced the code. Walk through the published rubric and agree the weighting for your context. 2. **2. Evaluation** (Week 2-3): Run the four-dimension rubric across sampled code. Reproduce findings, score each dimension on the published scale, write up severity and recommended fix. 3. **3. Report + briefing** (Week 3-4): Deliver the written report, walk leadership and engineering through the top findings, file prioritized tracker entries. 4. **4. Remediation follow-up (optional)** (Week +4-8): Re-review remediated findings 4-8 weeks later. Issue a remediation certificate if standard met. ### FAQ **Q: What does a typical audit cost?** Fixed fee scoped to repo size and audit depth. Most engagements land between EUR 15-60k. We share a tight range on the intro call once scope is clear. **Q: What is your evaluation rubric?** Four dimensions, each published with its own definition. Correctness: does the code do what the ticket and the tests say, including edge cases and error paths. Security: does it introduce vulnerabilities, unsafe defaults, weak secret handling or dependency risk. Maintainability: will the next engineer understand, locate and change it without reading the whole system. Test fitness: is coverage real - do the tests fail when behavior breaks - or theatrical. Each dimension carries named sub-criteria weighted to your codebase. **Q: How is each dimension scored?** On a published 0-4 scale with fixed anchors: 0 blocking defect, release must stop; 1 major gaps requiring rework before release; 2 works but with material findings and debt to schedule; 3 meets the bar with minor findings only; 4 no findings against that dimension. Every score in the report names the anchor it was assigned on, so you can contest an individual rating rather than the report as a whole. The full rubric, sub-criteria and scale are published on this page, do not change between engagements, and clients apply them to their own PRs between audits. **Q: Can you deliver under NDA for regulated industries?** Yes. Standard NDA, can operate on isolated infrastructure, can deliver in-person at client offices for highly sensitive engagements. Industry-specific frameworks (fintech KNF/DORA, healthcare GDPR/HIPAA-adjacent) are accommodated. **Q: Do you compete with my internal code review?** No. We are an external assurance signal - periodic, independent, deeper than day-to-day review. Most clients run us quarterly alongside their internal process. **Q: Can you re-audit after remediation?** Yes. Remediation re-audits are scoped at half the original engagement fee and produce a remediation certificate confirming findings closed. Useful for board reporting or external audit prep. ### Related - https://dfzoo.ai/services/evaluate-audit-secure/security-review - https://dfzoo.ai/services/evaluate-audit-secure/production-readiness - https://dfzoo.ai/services/operate-measure-maintain/llm-observability --- ## Security Review - Evaluate & Secure **URL:** https://dfzoo.ai/services/evaluate-audit-secure/security-review **Service type:** Application security code review **Anchor:** Evaluate, Audit & Secure (Evaluate & Secure) > A security-focused review of the code your team wrote with AI coding assistants - injection flaws, unsafe defaults, secrets handling, generated authorization logic - plus the application-level surface your AI features introduced, delivered as reproducible findings with fixes in the code. ### Summary dfzoo AI Institute reviews the security of code produced with AI coding assistants and of the AI features that code adds to a product. We read the AI-touched parts of your repository for injection flaws, unsafe defaults, secrets handling, dependency choices the assistant made, authorization logic it generated and error handling that leaks internals. We then review the application-level AI surface: how the app builds prompts, what it sends to and trusts back from an LLM API, how user input reaches a prompt, what an in-app agent is permitted to call, and how model output is handled before it reaches a user, a browser or a database. Every finding ships with a reproduction and a concrete fix in your code, plus the review rules your team applies afterwards so the same class of issue stops at code review. This is a code-level engagement: infrastructure, cloud and organizational security are out of scope by design. ### Who it's for - Product teams that adopted an AI coding assistant and are now shipping code faster than anyone reviews it for security - Software houses and agencies delivering AI-assisted code whose clients started asking for an independent security opinion on it - Companies that added an LLM feature (chat, summarization, an in-app agent) to an existing product and never reviewed what it exposed - Engineering leads, CTOs of product teams and tech leads who need the findings in the repository, not in a governance document ### Problems we solve - The assistant wrote query building, file handling and request parsing at speed, and nobody checked those paths for injection - Generated code carries the framework defaults the model happened to know: permissive CORS, disabled verification, debug error output in production - Secrets and tokens are handled the way the assistant demonstrated them - in code, in logs, in client bundles - Authorization checks were generated per endpoint and no one has verified they are consistent, or present at all - User input reaches a prompt unfiltered and model output goes straight into HTML, a shell call or a database write - An in-app agent holds tool access and credentials far wider than the task it actually performs ### What you get - Severity-ranked finding list (Critical / High / Medium / Low), each with a reproduction and the exact file and lines - A concrete fix per finding: a patch, a diff or a worked example your engineers can apply directly - Review of the AI feature surface: prompt construction, trust boundaries around the LLM API, output handling, agent tool and permission scope - Secrets and dependency findings for the AI-touched code, including packages the assistant introduced - A short set of security review rules for the team - what to look for in AI-generated PRs, as a checklist and as linter or CI rules where they can be automated - Working session with the engineering team on the top findings and how each one got written ### How we work 1. **1. Scope + code walkthrough** (Week 1): Agree which repositories and which AI-touched code are in scope, which assistant produced it, and where the AI features sit. Walk the codebase with a senior engineer from your side. 2. **2. Code-level security review** (Week 1-2): Read the AI-touched code for injection flaws, unsafe defaults, secrets handling, generated authorization logic, dependency choices and leaking error handling. Reproduce what we find. 3. **3. AI feature surface review** (Week 2-3): Trace how user input reaches a prompt, what the app sends to and trusts back from the LLM API, what an in-app agent is permitted to call, and how output is handled before it reaches a user or a data store. 4. **4. Report, fixes + review rules** (Week 3): Deliver the finding list with fixes, walk the engineering team through the top issues, and hand over the review rules and CI checks that catch the same classes in future PRs. ### FAQ **Q: How is this different from AI Code Evaluation?** AI Code Evaluation grades AI-assisted code on four dimensions - correctness, security, maintainability, test fitness - and gives you an overall assurance signal. Security Review takes only the security dimension and goes far deeper: we chase each issue to a working reproduction and a fix in the code, and we also review the AI feature surface, which the evaluation rubric only samples. Teams that want a single broad picture buy the evaluation; teams that already know security is the concern buy this. **Q: What do you explicitly not cover?** We do not do infrastructure, cloud or network security, penetration testing of live systems, offensive red teaming of deployed AI, SOC 2 or ISO 27001 audit programmes, or organizational security governance. We also do not assess your LLM vendor or model supply chain. We review code and the application-level surface that code creates. If you need any of the above, bring in a specialist for it - we are happy to hand over our findings so their work starts from something concrete. **Q: What access do you need?** Read access to the repositories in scope, the PR history for the AI-touched period, and a running instance in a non-production environment so we can reproduce findings. No production access, no customer data. A senior engineer from your side for a walkthrough at the start and a review of the findings at the end. Everything runs under NDA. **Q: How long does it take?** Most engagements run 2-3 weeks from scoping to report. A single service or a well-bounded AI feature can be done in a week; a large multi-repository codebase is scoped as several passes rather than one long review. Fixed fee, agreed once the scope is clear. **Q: Do you work remotely or on-site?** Remotely by default, under NDA, on your repository access. On-site delivery in Poland is available where a client's policy requires the code never leaves their network, at the same fee. **Q: What does the team get afterwards?** The finding list with fixes, the review rules as a checklist, and whatever we could turn into linter or CI checks. The point is that the next PR written with an assistant gets caught by your own review, not by us. Teams typically bring us back after a few months for a shorter re-review rather than a repeat of the full engagement. **Q: Which languages and stacks do you cover?** TypeScript and JavaScript, Python, Go, Ruby, Java and Kotlin, and the common web frameworks around them. Mobile and embedded are out of scope for this practice. Tell us the stack on the intro call and we will say plainly whether we are the right team for it. **Q: Do we need to have shipped an AI feature to buy this?** No. Plenty of engagements cover an ordinary product where the only AI involved is the assistant that helped write it. If the product has no LLM feature, we drop phase 3 and put the time into the code review instead. ### Related - https://dfzoo.ai/services/evaluate-audit-secure/ai-code-evaluation - https://dfzoo.ai/services/evaluate-audit-secure/production-readiness - https://dfzoo.ai/services/train-adopt-govern/tools-implementation --- ## Test Coverage - Evaluate & Secure **URL:** https://dfzoo.ai/services/evaluate-audit-secure/test-coverage **Service type:** Software quality and testing consulting **Anchor:** Evaluate, Audit & Secure (Evaluate & Secure) > Test strategy and implementation for AI-assisted projects - covering what to test (and what AI tools cannot test for you), what to mock, what to skip, and how to keep regressions in check as the AI keeps generating code. ### Summary dfzoo AI Institute designs and implements test coverage strategies for engineering teams shipping AI-assisted code. AI coding tools generate tests that look thorough but often miss regressions and prop-up coverage metrics without protecting behavior. We audit current tests, design a coverage strategy aligned to your product risk (what must never break vs what can fail and be fixed forward), implement the critical-path tests AI cannot generate well, and set up the CI signals that catch real regressions early. ### Who it's for - Engineering managers whose AI-assisted code is shipping faster than tests can keep up - Tech leads watching test coverage metrics rise while regressions rise too - CTOs trying to scale a small team using AI without breaking quality - QA leads adjusting their strategy for AI-generated code ### Problems we solve - AI generates tests that pass but do not protect business-critical behavior - Coverage metrics rise but customer-reported regressions also rise - Mocks proliferate as AI fills in gaps; the test suite tests itself, not the system - Critical-path integration tests are too tedious for AI to generate well; nobody writes them ### What you get - Audit of current test suite - what protects behavior vs what just inflates metrics - Risk-aligned test strategy document (what to test, what to mock, what to skip) - Implementation of critical-path tests AI tools struggle to generate - CI configuration for failure-mode signals (flake detection, regression alerts) - Playbook for team: prompt patterns for generating good tests with AI ### How we work 1. **1. Test audit** (Week 1): Sample existing tests across modules. Score each for true behavior protection vs metric padding. Identify gaps in critical paths. 2. **2. Strategy design** (Week 2): Map product risk to test strategy. Decide unit/integration/e2e mix. Pick what to mock vs run for real. 3. **3. Implementation** (Week 3-4): Write the critical-path tests AI tools miss. Set up CI signals. Document the prompt patterns for AI-generated tests. 4. **4. Team handover** (Week 4-5): Walk the team through the strategy. Pair on writing new tests using the playbook. Establish a review cadence. ### FAQ **Q: Why is AI bad at writing certain tests?** AI is great at unit tests that exercise visible function signatures, less good at integration tests that depend on system state, and weak at end-to-end tests where the cost of correct fixture setup exceeds what fits in context. The result is a test suite skewed toward easy wins, weak on real risk. **Q: Will you raise our coverage number?** We optimize for behavior protection, not raw coverage percentage. Most clients see coverage go down slightly (we remove tests that did not protect anything) and customer-reported regression rate drop more. **Q: Can you work with our existing CI?** Yes. GitHub Actions, GitLab CI, CircleCI, Jenkins, Buildkite - we work with what you have. We do not introduce new CI tools without a strong reason. **Q: What languages do you cover?** TypeScript / JavaScript, Python, Go, Ruby, Java/Kotlin. Mobile (Swift, Kotlin) and embedded are out of scope for this practice. **Q: How does this fit with AI Code Evaluation?** AI Code Evaluation audits code for correctness; Test Coverage builds the system that catches new regressions before they reach production. Most clients buy them as a pair. ### Related - https://dfzoo.ai/services/evaluate-audit-secure/ai-code-evaluation - https://dfzoo.ai/services/evaluate-audit-secure/production-readiness - https://dfzoo.ai/services/build-automate-modernize/ai-powered-custom-development --- ## Production Readiness - Evaluate & Secure **URL:** https://dfzoo.ai/services/evaluate-audit-secure/production-readiness **Service type:** Production readiness consulting **Anchor:** Evaluate, Audit & Secure (Evaluate & Secure) > Pre-release production readiness review for AI-touched systems - observability, runbooks, rollback paths and on-call signals checked against a structured rubric before the system goes live. ### Summary dfzoo AI Institute runs production readiness reviews for AI-touched systems before they ship to customers. We score the system against a structured rubric - observability and tracing coverage, runbook quality, rollback paths, on-call alerting, capacity headroom, cost monitoring, dependency failure modes - and issue a written certification with remediation requirements. The output is what SRE and platform teams need to sign off, and what leadership needs to defend the release at the post-mortem if something goes wrong. ### Who it's for - VPs of engineering signing off the launch of a new AI feature - Heads of SRE/platform asked to take an AI system on call - CTOs of growth-stage companies launching their first LLM-backed product - PMs whose launch depends on engineering and SRE both saying yes ### Problems we solve - AI features launch with no observability - debugging in production is guesswork - Runbooks for the new system are missing or untested - Rollback path was never exercised; production fires turn into multi-hour outages - Cost monitoring is missing; the first month's LLM bill is a shock ### What you get - Production readiness scorecard (10-15 dimensions, scored Pass / Partial / Fail) - Written certification statement with remediation requirements (if needed) - Runbook templates for the new system, tested with the on-call team - Rollback procedure documentation, exercised against a staging dry-run - Cost and capacity monitoring configuration recommendations ### How we work 1. **1. System review** (Week 1): Review architecture, observability stack, runbooks, on-call procedures, capacity plan, cost monitoring setup. 2. **2. Hands-on validation** (Week 2): Test rollback path in staging. Exercise runbooks with on-call team. Trigger test alerts to verify they fire correctly. 3. **3. Scoring + remediation list** (Week 2-3): Score against rubric. Write remediation requirements with priorities. Draft certification statement. 4. **4. Re-review (if remediation needed)** (Week +1-2): Re-check remediated dimensions. Issue final certification. ### FAQ **Q: Can you certify a system in 2 weeks?** Yes for systems with reasonable observability already in place. Systems missing fundamentals (no tracing, no runbooks) often need a 4-6 week remediation window before certification. **Q: What does the rubric cover?** Observability coverage, tracing depth, runbook quality, rollback procedure, on-call alerting, capacity headroom, cost monitoring, dependency failure modes, security baseline, AI-specific risk (prompt injection, model drift), and incident response readiness. **Q: Do you write the runbooks for us?** We provide templates and review what the team writes. We do not write runbooks the team has not validated - runbooks the on-call has not exercised do not protect anyone. **Q: Can you certify continuously, not just at launch?** Yes. Quarterly re-certification is a common retainer pattern. Useful when AI features evolve fast or as a board-level signal for production-critical systems. **Q: How does this differ from a security audit?** Security audit asks 'can attackers break this'. Production Readiness asks 'will this stay up under load, and can the team fix it fast when it breaks'. Most launches need both before going live. ### Related - https://dfzoo.ai/services/evaluate-audit-secure/ai-code-evaluation - https://dfzoo.ai/services/evaluate-audit-secure/security-review - https://dfzoo.ai/services/operate-measure-maintain/llm-observability --- ## AI Quality Evaluation - Evaluate & Secure **URL:** https://dfzoo.ai/services/evaluate-audit-secure/ai-quality-evaluation **Service type:** AI output and process quality evaluation **Anchor:** Evaluate, Audit & Secure (Evaluate & Secure) > Independent evaluation of the AI already running in your product or your back office - answer quality, retrieval grounding, agent task success and whether the AI step genuinely shortens the process - scored against a rubric that stays with you. ### Summary dfzoo AI Institute evaluates the quality of AI solutions that are already live: LLM features in a product, retrieval-based (RAG) answers, and agentic workflows that run without a person watching. We build an evaluation set from your real cases, agree a scoring rubric with the people who own the business outcome, score every case, and deliver a written report with a measured baseline, the failure patterns behind the score and a recommended fix per pattern. The unit of work is one evaluation - one test case scored against the rubric - so the scope and the invoice describe the same thing. The evaluation set and the rubric stay with you and get wired into CI as a quality gate that catches regression the next time a model, prompt or vendor changes. We are tool-neutral: we run on the evaluation stack you already have or help you choose one, and at high volume we show where self-hosting costs less than SaaS. ### Who it's for - Product teams that shipped an LLM feature and have no way to tell whether the answers are actually good - Companies running an internal AI assistant with flat adoption and no measure of whether it helps anyone - Teams about to change model, prompt framework or vendor and afraid of a silent regression - Operations and shared-service leads who put an AI step into a process and need to prove it shortened the process instead of moving the work - Heads of engineering, data and support who started with a free evaluation tool, got traces flowing and stalled on the test set ### Problems we solve - The LLM feature is live and the only quality signal is a complaint or a support ticket - RAG answers read fluently, but nobody has verified they are grounded in the right source document - Agentic workflows report success while a human quietly finishes the job, and nobody counts how often - A prompt or model change ships and users discover the regression before the team does - The team adopted a free evaluation tool and stalled on two things: building the test set and agreeing what 'good' means ### What you get - Evaluation set built from your real traffic and edge cases, in your tooling, owned by you after the engagement - Scoring rubric with weighted dimensions (correctness, grounding, task success, format and tone, cost and latency) agreed with the business owner, not just with engineering - Baseline scorecard: every case scored, results broken down per dimension, per user journey and per model or prompt version - Failure-pattern report with reproductions, root cause and a recommended fix for each pattern - CI quality gate: the evaluation set running as a blocking check on prompt, model, retrieval and vendor changes - Process effectiveness read-out showing where the AI step shortens the process and where it moves work to a human, plus a tooling recommendation with a self-hosting versus SaaS cost comparison at your volume ### How we work 1. **1. Scope + rubric** (Week 1): Agree what 'good' means for this solution with the people who own the outcome. Define the quality dimensions, their weights and the pass threshold. Fix the number of evaluations in scope. 2. **2. Evaluation set build** (Week 1-2): Assemble test cases from real traffic, known failures and the edge cases nobody tests. Add expected answers or grading criteria per case. Wire the set into your evaluation tooling or into one we help you pick. 3. **3. Scoring + analysis** (Week 2-3): Score every case against the rubric using automated grading, model grading and human review where the rubric needs judgment. Group failures into patterns and trace each pattern to its cause: retrieval, prompt, model choice, tool wiring or process design. 4. **4. Report + handover** (Week 3-4): Deliver the baseline scorecard and failure-pattern report, walk the product and engineering owners through it, and hand over the evaluation set and rubric wired into CI as a quality gate your team can run without us. 5. **5. Continuous evaluation (optional)** (Ongoing): Retainer: we maintain and extend the evaluation set, re-run it after every model, prompt or vendor change, and report the regression before your users find it. ### FAQ **Q: What exactly is included, and how do you price it?** The unit of work is one evaluation: one test case scored against the agreed rubric. We fix the number of evaluations, the rubric and the reporting in the scope, so you can see what you are paying for. Fixed fee per engagement, scoped to the number of evaluations and the number of AI surfaces in scope. We share a tight range on the intro call once scope is clear. **Q: How long does it take and what do you need from us?** Three to four weeks from kick-off to handover. From you we need one person who owns the business outcome to agree the rubric, one engineer for access and tooling, a sample of real traffic or logs, and any known failure cases you have already collected. The rubric workshop and the handover session take about half a day each; the rest runs on our side. **Q: How is this different from AI Code Evaluation?** AI Code Evaluation audits the code: correctness, security, maintainability and test fitness of what your team and your AI tools wrote. AI Quality Evaluation audits the behaviour of the deployed solution: whether the answers are right and grounded, whether the agent finishes the task, whether the process actually got shorter. Perfectly clean code can still produce a system that answers badly, and the reverse is also true. Teams frequently buy both, usually in that order. **Q: Do you need access to our production data?** Not necessarily. We can work on anonymised or synthetic cases derived from your real ones, and many engagements run that way for regulated clients. Where production access is granted, we work under NDA, on your infrastructure, with the data staying in your environment. Standard NDA, and we can operate on isolated infrastructure if the engagement requires it. **Q: Which evaluation tools do you work with?** We are tool-neutral and work on what you already run: Braintrust, Langfuse, LangSmith, Confident AI, Promptfoo, Arize Phoenix, W&B Weave, or a plain test harness in your own repository. If nothing is in place we help you choose, and at high evaluation volume we show you where self-hosting an open-source platform costs less than SaaS. The evaluation set is written so it can move between tools, because tying your quality baseline to a single vendor is a risk in a market that keeps consolidating. **Q: We already tried a free evaluation tool. What do you add?** Every evaluation platform has a free tier, so most clients arrive having tried one. The tool is rarely where teams get stuck. They get stuck on the two things the tool does not give you: a test set that represents what users actually ask, and a written definition of 'good' that the business owner and the engineers both sign. We start exactly there, and the tool you already picked keeps running underneath. **Q: Can you evaluate agentic workflows, not just single answers?** Yes, and that is usually where the interesting numbers are. We score task success rate (did the agent complete the job end to end), human takeover rate (how often a person silently finishes it), step-level failures across sessions, turns, tool calls and subagents, and process effect: how much time the AI step removes from the process versus how much it moves to someone else. **Q: What happens after the engagement, and who maintains the evaluation set?** The evaluation set, the rubric and the CI gate are yours, in your repository and your tooling, and your team can run and extend them without us. If you would rather not own the maintenance, we run it as a continuous retainer: we keep the set current, re-run it after each model, prompt or vendor change, and report regression against your baseline. That retainer sits alongside LLM Observability, which watches the same quality signals in live traffic. ### Related - https://dfzoo.ai/services/evaluate-audit-secure/ai-code-evaluation - https://dfzoo.ai/services/operate-measure-maintain/llm-observability - https://dfzoo.ai/services/evaluate-audit-secure/production-readiness --- # Build, Automate & Modernize (anchor 03) **URL:** https://dfzoo.ai/services/build-automate-modernize **Highlight:** AI-Powered Custom Development - company pillar > AI-powered custom development, process automation and legacy modernization - engineered for production, not for demo. ### Summary dfzoo AI Institute builds, automates and modernizes systems using AI as the primary development leverage. Our AI-Powered Custom Development practice is the company pillar - we ship production applications, agentic backends and AI-augmented internal tools to engineering teams that need product velocity. We assess agentic AI readiness for organizations exploring autonomous workflows; automate business processes that combine RPA and LLM agents; and modernize legacy systems where AI-assisted refactoring unlocks otherwise prohibitive change. Every engagement is engineered with the same quality assurance backing as our evaluation practice. ### Who it's for - Product engineering teams shipping AI-native or AI-augmented products - Operations leaders automating high-volume processes that mix structured and unstructured data - CTOs sitting on legacy systems that block roadmap progress - Founders prototyping agentic backends and need senior engineering judgment ### Problems we solve - AI prototypes built by the team do not survive the move to production - Manual back-office processes consume hours that AI + RPA could compress to minutes - A legacy system blocks the product roadmap but a full rewrite is not affordable - Agentic workflows look promising in demos but no one knows what production looks like ### Practices under this anchor - AI-Powered Custom Development - https://dfzoo.ai/services/build-automate-modernize/ai-powered-custom-development - Agentic AI Readiness - https://dfzoo.ai/services/build-automate-modernize/agentic-ai-readiness - Process Automation - https://dfzoo.ai/services/build-automate-modernize/process-automation - Legacy Modernization - https://dfzoo.ai/services/build-automate-modernize/legacy-modernization ### FAQ **Q: How is AI-Powered Custom Development different from regular custom dev?** Our engineering process uses AI as primary leverage - AI-assisted coding, AI-driven test generation, AI-augmented code review. The deliverable is the same: a production application. The velocity and the cost structure are different. **Q: Do you build agents end-to-end?** Yes. We design the agent architecture (single-shot, multi-step, multi-agent), build the supporting infrastructure (memory, retrieval, tool calls, observability), integrate with your existing systems, and ship to production with on-call runbooks. **Q: What kinds of processes do you automate?** Document-heavy and inbox-heavy processes: lead intake, invoice processing, customer support triage, compliance check workflows, RFP response drafting. We combine RPA tools (UiPath, Power Automate) with LLM agents where unstructured data is involved. **Q: Can you modernize a legacy system in place, without a full rewrite?** Often yes. We assess the system, identify the strangler-fig boundaries, and replace components incrementally using AI-assisted refactoring. The roadmap blocker comes off without a multi-quarter rewrite. **Q: What stacks do you work in?** TypeScript / Node / Next.js, Python / FastAPI, Go for backend; React, Vue for frontend; Postgres, ClickHouse, Pinecone, pgvector for data; AWS, GCP, Vercel, Fly for infrastructure. Most LLM providers (OpenAI, Anthropic, Google, open-source via Bedrock or Together). --- ## AI-Powered Custom Development - Build & Modernize **URL:** https://dfzoo.ai/services/build-automate-modernize/ai-powered-custom-development **Service type:** Custom software development **Anchor:** Build, Automate & Modernize (Build & Modernize) > End-to-end product engineering with AI as primary development leverage - production applications, agentic backends and AI-augmented internal tools built on a staged pipeline where every change passes automated gates and a named human before it reaches production. ### Summary dfzoo AI Institute builds production software using AI as the primary engineering leverage. We deliver web applications, agentic backends, AI-augmented internal tools and AI-native products to teams that need senior engineering judgment at the velocity AI now enables. Every change moves through the same pipeline: ticket, spec, plan, implementation, automated gates (typecheck, lint, build, code review), PR with preview, human review, merge, verification in production. We accelerate delivery and hold stability, and we measure both: throughput and defects found after merge. ISO 9001:2015 quality assurance backing. The company pillar - most engagements at dfzoo AI Institute touch this practice. ### Who it's for - Product engineering teams shipping AI-native or AI-augmented products - Founders prototyping a new AI-first product who need senior engineering - Mid-stage companies replacing or extending a critical internal tool - Scale-ups whose product roadmap exceeds in-house engineering capacity ### Problems we solve - AI prototypes built by the team do not survive the move to production - A prompt-built prototype stalls at the demo: no tests, no review trail, no runbook, no owner in production - we hand over a system with all four - In-house engineering is full; the roadmap keeps slipping - Agentic workflows demo well but the team has no production pattern for them - Custom development quotes from agencies do not account for AI-driven velocity ### What you get - Working production application (web, agentic backend, internal tool) - Source code with documentation, tests, deployment configs, runbooks - Delivery pipeline handed over with the code: CI gates, PR template, review checklist, preview environments - Infrastructure-as-code for production, staging, preview environments - Observability and on-call setup (tracing, alerts, runbooks) - Knowledge transfer to your team - pair sessions, walkthroughs, recorded handover ### How we work 1. **1. Discovery + scope** (Week 0-1): intro call, problem definition, technical scoping, fixed-fee phase proposal. 2. **2. Ticket, spec, plan** (Per change): Every change starts as a ticket, becomes a written spec, then a plan. Gate: you approve scope and plan before a line of code is written. 3. **3. Implementation** (Per change): AI agents write code under a senior engineer who directs, corrects and rejects. Gate: what leaves the branch is decided by that engineer, never by the agent that wrote it. 4. **4. Automated gates** (Every commit): Typecheck, lint, build and automated code review run on every branch. Gate: machine-decided and not overridable, same threshold for AI-written and hand-written code. 5. **5. PR, preview + human review** (Per change): Each PR gets a preview environment with the change running end to end, and a second senior engineer reads the full diff. Gate: a human decides merge or rework, and the review trail stays in the repository. 6. **6. Merge + production verification** (Per release): After deploy we check the change against tracing, alerts and the runbook. Gate: the change is done when it behaves in production, not when it merges. 7. **7. Hardening, launch + handover** (Last 2-4 weeks): Production readiness review, load testing, observability completeness, on-call training, then pair sessions and documented handover. Optional maintenance under App Maintenance. ### FAQ **Q: How fast can you start?** intro call to first sprint typically 2-3 weeks. Faster for urgent engagements (1 week) when scope is already clear and we have capacity. **Q: Does AI-driven velocity cost us stability?** That is the trade we refuse, and the reason the pipeline exists. We track delivery throughput and defects found after merge side by side, report both, and treat rising post-merge defects as a reason to tighten gates, not as the price of speed. **Q: Fixed fee or time and materials?** Fixed fee per phase. Phase 1 (Discovery + scope) is typically EUR 10-25k depending on complexity. Subsequent build phases are scoped per phase based on what we learned. **Q: What about IP ownership?** You own the code. We retain rights to reusable internal libraries and patterns we wrote before the engagement. Standard mutual NDA on the engagement itself. **Q: Will we work with the same engineers throughout?** Yes. Engagement teams are stable from kickoff to handover. We do not swap engineers mid-engagement unless you ask us to. **Q: Can you do AI engineering specifically, not just any custom dev?** Yes. Agentic backends (single-shot, multi-step, multi-agent), RAG systems, LLM-backed APIs, AI evaluation pipelines, embedding and vector search infrastructure. Each of these is running in production for clients today. ### Related - https://dfzoo.ai/services/build-automate-modernize/agentic-ai-readiness - https://dfzoo.ai/services/build-automate-modernize/legacy-modernization - https://dfzoo.ai/services/operate-measure-maintain/app-maintenance --- ## Agentic AI Readiness - Build & Modernize **URL:** https://dfzoo.ai/services/build-automate-modernize/agentic-ai-readiness **Service type:** Agentic AI consulting and pilot delivery **Anchor:** Build, Automate & Modernize (Build & Modernize) > Assessment and a working pilot for organizations evaluating agentic AI workflows - written recommendations on what to agentify, what to keep human, and a production-pattern pilot you can extend. ### Summary dfzoo AI Institute assesses agentic AI readiness for organizations weighing autonomous workflows. We map your current processes against agentic patterns (single-shot tools, multi-step agents, multi-agent systems), recommend what to agentify and what to keep human-in-the-loop, and deliver a working pilot using your data and systems. Output is a written readiness assessment plus a deployed pilot - the production pattern your team extends, not a research-grade demo. ### Who it's for - CTOs and heads of product evaluating an agentic AI investment - Operations leaders eyeing autonomous workflows in customer support, sales ops or finance - Innovation teams asked to pilot agentic AI by leadership - Engineering leaders whose teams are excited about agents but need a production reality check ### Problems we solve - Agentic demos look magical; nobody can articulate the production pattern - Teams disagree on whether a workflow should be agentic or human-led - Pilots run as research prototypes and never become production systems - Vendor pitches for 'agentic platforms' are hard to compare to building it yourself ### What you get - Written readiness assessment (10-20 pages): workflows, agentic patterns, recommendations - Working pilot deployed to staging or limited production (single workflow) - Pilot architecture documentation: agent flow, tools, memory, observability, escalation - Cost model for scaling the pilot pattern to other workflows - Executive briefing with recommendations and roadmap options ### How we work 1. **1. Process discovery** (Week 1-2): Map current workflows. Identify candidates for agentification. Score each by suitability and risk. 2. **2. Pattern design** (Week 2-3): Choose the right pattern per candidate (single-shot, multi-step, multi-agent). Design the pilot. 3. **3. Pilot build** (Week 3-7): Build and deploy the pilot. Wire observability, evaluation, escalation paths. 4. **4. Assessment + roadmap** (Week 7-10): Run the pilot for 2-4 weeks. Measure outcomes. Deliver assessment and roadmap for next workflows. ### FAQ **Q: What is the difference between an agent and a workflow?** A workflow runs deterministic steps in order. An agent decides the next step based on context, potentially calling tools and looping. Many problems are solved better by workflows; we recommend agents only when the dynamic decision-making earns its cost. **Q: Do you use a specific agent framework?** We work with LangGraph, OpenAI Agents SDK, Anthropic Claude Agent SDK, and hand-rolled patterns. We choose based on the problem; we are not loyal to a single framework. **Q: Can you assess without building a pilot?** Yes. Assessment-only engagements run 3-4 weeks and deliver the readiness document and recommendations. Most clients add a pilot because the assessment recommendations need to be tested before scaling. **Q: How do you measure if an agent is working?** Eval pipelines (golden test sets, side-by-side comparisons), production observability (trace-level visibility into agent decisions), business KPIs (workflow throughput, escalation rate, customer-facing quality). All three matter; we set them up together. **Q: What happens when the agent makes a mistake in production?** The pilot architecture includes escalation paths: low-confidence decisions route to humans, errors trigger alerts, rollback is one configuration change. Production agents need the same operational discipline as any production system - we install it as part of the pilot. ### Related - https://dfzoo.ai/services/build-automate-modernize/ai-powered-custom-development - https://dfzoo.ai/services/build-automate-modernize/process-automation - https://dfzoo.ai/services/operate-measure-maintain/llm-observability --- ## Process Automation - Build & Modernize **URL:** https://dfzoo.ai/services/build-automate-modernize/process-automation **Service type:** Business process automation **Anchor:** Build, Automate & Modernize (Build & Modernize) > Automation of high-volume business processes that mix structured and unstructured data - combining RPA tools with LLM agents where the unstructured side beats traditional rule-based RPA. ### Summary dfzoo AI Institute automates high-volume business processes that span multiple systems. We combine RPA tooling (UiPath, Power Automate, n8n) for structured-data flows with LLM agents for unstructured inputs (emails, documents, transcripts). Typical engagements: lead intake, invoice processing, customer support triage, compliance check workflows, RFP response drafting. We ship a working automation, measure its throughput and quality, and hand it over with maintenance documentation. ### Who it's for - Heads of operations whose teams burn hours on document-heavy processes - Finance leaders automating invoice intake, AP / AR workflows or compliance checks - Customer support leads triaging high-volume inbound across email and chat - Sales operations leaders processing inbound leads from multiple sources ### Problems we solve - Traditional RPA fails on emails, PDFs, transcripts - the unstructured 30% - ChatGPT-based prompts work in a demo, fall apart in production volume - Existing automations break weekly because source systems keep changing - Throughput stalls because every exception needs a human; nobody measures the rate ### What you get - Working automation deployed to production with monitoring - Exception handling and human-in-the-loop escalation for low-confidence cases - Quality metrics dashboard (throughput, accuracy, exception rate, cost per transaction) - Maintenance playbook for when source systems or data formats change - Optional: monthly cost-and-quality review under App Maintenance retainer ### How we work 1. **1. Process mapping** (Week 1-2): Shadow the team running the manual process. Map exact steps, decisions, exception paths. 2. **2. Architecture + tool choice** (Week 2-3): Decide RPA / LLM / hybrid per step. Choose specific tools. Design escalation paths. 3. **3. Build + integrate** (Week 3-7): Build the automation. Integrate with email, CRM, ERP and whatever source systems matter. 4. **4. Ramp + handover** (Week 7-10): Ramp from 10% to 100% of volume over 2-4 weeks. Train ops team on monitoring. Document maintenance. ### FAQ **Q: How do you decide RPA vs LLM agent for a given step?** Structured data with stable schemas → RPA. Unstructured input or fuzzy decisions → LLM with structured output. Most real processes are hybrid; we choose per step to keep cost and failure modes manageable. **Q: What is your typical accuracy vs human?** Depends entirely on the process and the data quality. Invoice line-item extraction reaches 95%+ on well-formed PDFs. Open-ended classification can land at 80-90% with eval pipelines tuned to your taxonomy. We agree the target on the scope call. **Q: Can you automate something we already half-built?** Yes. Many engagements pick up existing scripts, fragile RPA bots or ChatGPT-based prompts and re-engineer them for production reliability. **Q: What about cost?** Two cost dimensions: build cost (fixed-fee per phase) and run cost (LLM API spend per transaction). Run cost is part of the architecture decision: we choose the cheapest model and the smallest prompt that hits the quality target. **Q: Does this need long-term maintenance?** Yes. Source systems change, data formats drift, exception patterns evolve. We offer maintenance under App Maintenance retainer - typically light-touch monthly check-ins with quarterly tune-ups. ### Related - https://dfzoo.ai/services/build-automate-modernize/agentic-ai-readiness - https://dfzoo.ai/services/operate-measure-maintain/app-maintenance - https://dfzoo.ai/services/operate-measure-maintain/cost-model-optimization --- ## Legacy Modernization - Build & Modernize **URL:** https://dfzoo.ai/services/build-automate-modernize/legacy-modernization **Service type:** Legacy system modernization **Anchor:** Build, Automate & Modernize (Build & Modernize) > AI-assisted refactoring and re-platforming for legacy systems blocking your roadmap - incremental strangler-fig modernization instead of a multi-quarter rewrite. ### Summary dfzoo AI Institute modernizes legacy systems blocking product roadmaps. We assess the system, identify strangler-fig boundaries, and replace components incrementally using AI-assisted refactoring at a fraction of full-rewrite cost. Typical engagements: PHP/Perl/older .NET monoliths, legacy Java EE, jQuery-era frontends, untyped JavaScript / legacy Python codebases. Output is a modernized system that runs alongside the legacy until the legacy can be retired safely. ### Who it's for - CTOs whose roadmap is blocked by a system the team is afraid to touch - Engineering leaders facing a 'rewrite vs live with it' decision - Mid-stage companies whose first-product codebase no longer matches the team that maintains it - Acquirers integrating a target whose tech stack is materially older than theirs ### Problems we solve - The legacy system has no tests; every change risks production - Documentation is missing or outdated; institutional knowledge sits with one or two engineers - A full rewrite is too expensive to fund; living with the legacy costs more each quarter - Hiring is hard because new engineers do not want to maintain the old stack ### What you get - Legacy system assessment (architecture, dependencies, risk hotspots, modernization options) - Strangler-fig modernization roadmap with sequenced phases and budget per phase - Modernized components running in production alongside the legacy - Test coverage and documentation for both modernized and legacy boundaries - Knowledge transfer to your team - pair sessions, written architecture docs ### How we work 1. **1. Assessment** (Week 1-3): Read the codebase under NDA. Interview engineers who maintain it. Run static analysis. Map dependencies. 2. **2. Roadmap design** (Week 3-4): Identify strangler-fig boundaries. Sequence phases by risk and roadmap unlock. Budget per phase. 3. **3. Modernization sprints** (Week 4-N): Replace components phase-by-phase. Each phase ships behind a routing layer running both old and new in parallel. 4. **4. Legacy retirement** (After each phase): Once a component is fully replaced and stable, retire the legacy path. Repeat until the legacy is gone. ### FAQ **Q: Can you really modernize incrementally, not as a big-bang rewrite?** Yes - strangler-fig is the standard pattern. We build the routing layer that runs old and new in parallel, replace one component at a time, and retire the old code only after the new path is stable in production. Big-bang rewrites are reserved for cases where strangler-fig is genuinely impossible. **Q: What does AI-assisted refactoring actually do?** AI tools accelerate the tedious parts: reading unfamiliar code, generating tests for code that has none, translating between idioms (e.g. callback-style to async/await), proposing refactors that humans review. The engineering judgment stays with our team; the volume work gets faster. **Q: What legacy stacks have you modernized?** PHP (Symfony 1.x, custom legacy), older .NET (Framework 4.x ASP.NET WebForms / MVC), Java EE / Spring Boot pre-2.x, jQuery + server-rendered frontends, untyped Node.js codebases, legacy Python 2 / early 3. **Q: What if our codebase has no tests at all?** Common starting state. Phase 1 of modernization typically includes building a test scaffold around the components we will replace - generated with AI assistance, reviewed by our engineers. Tests come before refactor, always. **Q: Can we keep the legacy system running while you modernize?** Yes - that is the point of strangler-fig. The legacy keeps serving production traffic while new components come online behind a routing layer. Cutover per component is a configuration change, not a release. ### Related - https://dfzoo.ai/services/build-automate-modernize/ai-powered-custom-development - https://dfzoo.ai/services/evaluate-audit-secure/test-coverage - https://dfzoo.ai/services/operate-measure-maintain/app-maintenance --- # Operate, Measure & Maintain (anchor 04) **URL:** https://dfzoo.ai/services/operate-measure-maintain **Highlight:** Telemetry & LLM observability > Observability, telemetry, cost optimization and ongoing maintenance for AI systems already in production - so they keep working as your data and usage change. ### Summary dfzoo AI Institute operates AI in production for teams that need their systems to keep working as data, usage and providers shift. We implement LLM observability (tracing, eval pipelines, drift detection); design telemetry and analytics for AI products (what to measure, what to alert on); run LLM cost and model optimization audits that typically recover 20-60% of monthly LLM spend; and provide ongoing application maintenance for AI-touched systems. Our cost optimization practice is anchored in concrete provider math - caching, routing, prompt compression - not vendor talking points. ### Who it's for - Engineering teams running LLM features in production with growing cost surprises - Product leaders who cannot tell whether their AI feature is actually working for users - CFOs and finance partners asking why the LLM bill keeps growing - DevOps and SRE teams adding LLM-backed services to their on-call rotation ### Problems we solve - LLM spend is growing faster than traffic and no one can tell why - There is no signal whether prompt changes made the product better or worse - An LLM provider deprecated a model and the team has no migration plan - Customer complaints about AI quality cannot be tied back to specific behavior in production ### Practices under this anchor - LLM Observability - https://dfzoo.ai/services/operate-measure-maintain/llm-observability - Telemetry & Analytics - https://dfzoo.ai/services/operate-measure-maintain/telemetry-analytics - Cost & Model Optimization - https://dfzoo.ai/services/operate-measure-maintain/cost-model-optimization - App Maintenance - https://dfzoo.ai/services/operate-measure-maintain/app-maintenance ### FAQ **Q: What does an LLM cost audit recover?** Typical clients recover 20-60% of monthly LLM spend over 4-8 weeks. The savings come from prompt cache hit rate improvements, smaller-model routing for low-stakes calls, prompt compression and removing redundant retries. **Q: What observability stack do you use?** OpenTelemetry-based where possible. We integrate with Datadog, Honeycomb, Grafana, Posthog and dedicated LLM observability tools (Langfuse, Arize, Helicone) depending on what is already in place. **Q: How do you measure whether prompt changes improved the product?** Eval pipelines: golden test sets, side-by-side comparisons against a baseline, regression alerts on metrics that matter. We design the eval rubric for your product specifically. **Q: Do you handle model migrations when a provider deprecates a model?** Yes. Provider migrations (e.g., Claude 4.6 to 4.7, GPT-4 to GPT-4o) are a standard maintenance engagement: we re-baseline evals, run regression tests and migrate prompts with measured before-after on quality and latency. **Q: Is maintenance a retainer or a fixed engagement?** Both work. Retainer fits teams with continuous AI features in production. Fixed engagements (provider migration, cost audit, eval setup) fit teams with a specific need. --- ## LLM Observability - Operate & Measure **URL:** https://dfzoo.ai/services/operate-measure-maintain/llm-observability **Service type:** LLM observability and evaluation engineering **Anchor:** Operate, Measure & Maintain (Operate & Measure) > Tracing, eval pipelines and drift detection for LLM and agentic systems in production - sessions, turns, steps, tool calls and subagents, with cost and quality attached to each. ### Summary dfzoo AI Institute implements LLM observability for production systems. We install tracing over the entities an agentic system actually has - sessions, turns, steps, tool calls and subagents, not isolated model calls - build eval pipelines that score new prompt and model versions against your test sets, set up drift detection, and integrate the signals with your existing observability stack. The test set built during an evaluation stays with you as a CI gate and a production monitor, and we maintain and extend it after every model or prompt change. OpenTelemetry-based where possible; we work with Datadog, Honeycomb, Grafana, Langfuse, Arize, Helicone, Braintrust, LangSmith, W&B Weave, Confident AI and Galileo, we help you pick, and at high volume we show where self-hosting beats SaaS. ### Who it's for - Engineering teams whose LLM features ship blind - no visibility into what's happening - Teams running multi-step agents where a failure is three tool calls deep and nobody can find it - Product leaders unable to tell whether prompt iterations improved or regressed the product - SRE teams adding LLM services to their on-call rotation - ML/AI engineering leads needing eval-driven prompt development ### Problems we solve - An agent run fails and nobody can reproduce which step, tool call or subagent broke it - Prompt and model changes ship without measurement; quality regressions are invisible - An audit produced a good test set once and it went stale the week the model changed - Latency P99 doubles overnight and nobody can pinpoint the cause - Provider returns a wrong answer; there is no log of the request to debug from ### What you get - Tracing across agentic entities - sessions, turns, steps, tool calls, subagents - with prompt, response, latency, cost and version on each - Eval pipeline with your test sets and CI-integrated regression alerts - Your audit test set handed over as a CI merge gate and a production monitor on the same thresholds, with a routine for extending it per model or prompt change - Drift detection on outputs (semantic drift, latency drift, cost drift) - Dashboards: per-feature quality, per-feature cost, per-user behavior - Runbooks for the on-call team: how to debug an LLM incident in 15 minutes ### How we work 1. **1. Stack discovery** (Week 1): Map current LLM and agent calls across your codebase. Inventory features, prompts, providers, existing test sets and observability tooling. 2. **2. Instrumentation** (Week 2-3): Wrap sessions, turns, steps, tool calls and subagents with tracing. Forward to the chosen backend and validate every run surfaces end to end. 3. **3. Eval pipelines + CI gate** (Week 3-4): Build test sets per feature from real traffic and from prior evaluation findings. Wire eval runs into CI as a merge gate and mirror the same thresholds as a production monitor. 4. **4. On-call training + handover** (Week 4-5): Walk the on-call team through dashboards, runbooks and incident scenarios. Hand over the test set as your asset, in your repository. 5. **5. Maintenance retainer (optional)** (Ongoing): We extend the test set and rerun the regression after every model or prompt change, and report what moved and why. ### FAQ **Q: Which observability or eval backend do you recommend?** Depends on what you already have. If you run Datadog or Honeycomb, we extend those. If LLM-specific is preferred: Langfuse, Arize, Helicone, Braintrust, LangSmith, W&B Weave, Confident AI, Galileo. We are tool-agnostic, help you pick against your volume and data boundaries, and above a certain request volume model where self-hosting (Langfuse, Phoenix) comes out cheaper than SaaS. **Q: What if the eval vendor we pick is acquired, sunset or repositioned?** It happens on this market, and it is why we do not build your practice on one vendor. Test sets, scoring definitions and traces stay in portable form in your repository, and we instrument through OpenTelemetry where the backend allows it. Platforms consolidate; the competence and your own test sets stay yours, so a migration is a week of work rather than a rebuild. **Q: Can you observe LLM calls without changing application code?** Partially - providers like OpenAI and Anthropic offer per-request logs through their consoles. For production-grade observability (trace context, user attribution, cost attribution, agent step structure) we usually wrap calls in a thin SDK layer. **Q: What is an eval pipeline?** A test suite for LLM and agent outputs. The test set defines what a right answer looks like; eval runs score new prompt and model versions against it; CI gates merges if quality regresses, and the same thresholds run as a production monitor. Standard practice for any team treating prompts as production code. **Q: Do you handle data privacy in traces?** Yes. Tracing can be configured to redact PII before logs leave your infrastructure, mask prompt sections, or store full prompts only on isolated infrastructure. We design the redaction policy with your security team. **Q: What does this cost to run monthly?** Tracing backend cost depends on volume - typically EUR 200-2000/month for mid-size LLM workloads. Self-hosted Langfuse is free of vendor cost but needs ops. We model both options in the discovery phase. ### Related - https://dfzoo.ai/services/operate-measure-maintain/telemetry-analytics - https://dfzoo.ai/services/operate-measure-maintain/cost-model-optimization - https://dfzoo.ai/services/evaluate-audit-secure/production-readiness --- ## Telemetry & Analytics - Operate & Measure **URL:** https://dfzoo.ai/services/operate-measure-maintain/telemetry-analytics **Service type:** Product analytics and telemetry engineering **Anchor:** Operate, Measure & Maintain (Operate & Measure) > Product telemetry and analytics for AI-touched products - what to measure, what to alert on, what to put on the dashboard so product, engineering and leadership are looking at the same picture. ### Summary dfzoo AI Institute designs and implements product telemetry for AI-touched products. We translate product questions ('is the AI feature actually used?', 'are users abandoning at the AI step?', 'which prompt version drives revenue?') into events, dashboards and alerts. Output: an event taxonomy, instrumentation across web/mobile/backend, dashboards for product/engineering/leadership and a roadmap for what to measure next. ### Who it's for - Heads of product who cannot tell whether their AI feature is helping users - Growth teams running A/B tests on prompt versions or AI feature rollouts - Engineering managers building data signals into product decisions - Founders launching AI-first products without an in-house analytics team ### Problems we solve - Product is shipping AI features blind - no idea what users actually do with them - Existing analytics tracks page views but misses AI-specific user moments (prompt edits, accept/reject, regenerate) - Each team has its own dashboard with different numbers; leadership cannot trust any of them - Alerts fire for things nobody cares about; real problems pass unnoticed ### What you get - Event taxonomy (AI-specific moments: prompt sent, output rendered, accepted, rejected, regenerated, etc.) - Instrumentation across web, mobile and backend with version tracking - Three dashboards: product (feature health), engineering (system health), leadership (business KPIs) - Alert policy: what fires, who gets paged, what the runbook says - Quarterly review cadence to evolve what is measured as the product evolves ### How we work 1. **1. Question discovery** (Week 1): Interview product, engineering and leadership. Write down the questions each role needs answered. Prioritize. 2. **2. Event taxonomy** (Week 2): Translate questions into events. Define AI-specific moments. Document properties and versioning policy. 3. **3. Instrumentation + dashboards** (Week 2-4): Instrument across the stack. Build the three dashboards. Set the alert policy. 4. **4. Handover + review cadence** (Week 4-5 (then quarterly)): Walk each audience through their dashboard. Set quarterly review. Iterate the event list at +3 months. ### FAQ **Q: What analytics platforms do you work with?** Posthog, Mixpanel, Amplitude, Segment as a router, Snowplow for self-hosted. For LLM-specific signals we often pair with Langfuse or Helicone. We pick based on your stack, not vendor preference. **Q: Do you do data warehousing?** Light-touch. We send analytics events to your warehouse (Snowflake, BigQuery, Redshift, ClickHouse) but do not build full data warehouse models - that is a separate engagement. **Q: How is this different from LLM Observability?** LLM Observability is about the LLM call itself (latency, cost, output drift). Telemetry & Analytics is about the user and product (does the feature work, do users come back, does it drive revenue). Most clients buy both. **Q: Can you instrument mobile too?** Yes. iOS (Swift) and Android (Kotlin), plus React Native and Flutter. We respect platform-specific privacy requirements (ATT on iOS, scoped storage on Android). **Q: How do you handle privacy and consent?** Consent gating is wired into instrumentation. Events that touch PII are tagged for consent-respecting routing. Privacy review is part of the instrumentation phase, not an afterthought. ### Related - https://dfzoo.ai/services/operate-measure-maintain/llm-observability - https://dfzoo.ai/services/operate-measure-maintain/cost-model-optimization - https://dfzoo.ai/services/evaluate-audit-secure/production-readiness --- ## Cost & Model Optimization - Operate & Measure **URL:** https://dfzoo.ai/services/operate-measure-maintain/cost-model-optimization **Service type:** LLM cost optimization audit **Anchor:** Operate, Measure & Maintain (Operate & Measure) > Cost and model optimization audit for production LLM systems - caching, routing, prompt compression and provider-side fixes that typically recover 20-60% of monthly LLM spend over 4-8 weeks. ### Summary dfzoo AI Institute runs LLM cost and model optimization audits for engineering teams whose LLM bill grows faster than traffic. We instrument the system for per-call cost attribution, audit current usage, recommend and implement provider-side savings (prompt caching, smaller-model routing, prompt compression, retry tuning) and deliver a measured before/after with eval-protected quality. Typical clients recover 20-60% of monthly LLM spend. ISO 9001:2015 quality assurance backing. ### Who it's for - Engineering teams whose monthly LLM bill grew past EUR 10k and keeps climbing - CFOs and finance partners asking why LLM spend exceeds plan - Product engineering leaders shipping AI features whose unit economics are upside-down - Mid-stage companies preparing for a funding round needing defensible AI cost numbers ### Problems we solve - LLM spend grows faster than user traffic and nobody knows why - Prompt cache hit rate is low or not measured; every call pays full price - All features route to the most expensive model regardless of stakes - Retries on transient errors silently multiply the bill ### What you get - Cost attribution dashboard - per feature, per user segment, per model - Optimization recommendations ranked by saving × effort - Implementation of the top wins (caching, routing, prompt rework, retry policy) - Eval-protected quality measurement: same or better quality after optimization - Written audit report with measured before/after savings ### How we work 1. **1. Cost instrumentation** (Week 1-2): Instrument per-call cost attribution. Run for 1-2 weeks to capture a representative cost picture. 2. **2. Audit + recommendations** (Week 2-3): Analyze cost data. Identify top savings opportunities. Write recommendations with effort estimates. 3. **3. Implementation** (Week 3-6): Implement top wins. Set up eval-protected rollout: quality must hold before cost drop counts. 4. **4. Measurement + report** (Week 6-8): Measure before/after over 2-4 weeks. Deliver audit report with documented savings. ### FAQ **Q: What does 20-60% savings really mean?** Measured over the month following implementation, comparing LLM API spend per equivalent traffic volume to the baseline month before the audit. We exclude traffic growth from the comparison. Specific clients recover different amounts depending on starting state - teams already using caching see less, teams that have never optimized see more. **Q: Will optimization degrade quality?** We protect against this with eval pipelines. Every optimization is tested against a golden set; we only ship changes that hold or improve quality. Quality is the constraint, cost is the optimization target. **Q: What provider features matter most?** Prompt caching is usually the biggest win (Anthropic prompt caching, OpenAI prompt caching). Followed by smaller-model routing (use Haiku/4o-mini where appropriate), prompt compression (remove redundant context), and retry tuning (do not silently retry on idempotent errors). **Q: Do you do this for one provider or multiple?** Both. Multi-provider clients often have routing logic that picks the cheapest qualified model per task - we optimize that routing as part of the engagement. Single-provider clients optimize within the provider's pricing surface. **Q: What does the engagement cost?** Fixed fee, typically scoped to the size of the LLM workload. Most engagements pay back the cost within 1-3 months of measured savings. We share a tight range on the intro call. ### Related - https://dfzoo.ai/services/operate-measure-maintain/llm-observability - https://dfzoo.ai/services/operate-measure-maintain/app-maintenance - https://dfzoo.ai/services/build-automate-modernize/ai-powered-custom-development --- ## App Maintenance - Operate & Measure **URL:** https://dfzoo.ai/services/operate-measure-maintain/app-maintenance **Service type:** Application maintenance and support **Anchor:** Operate, Measure & Maintain (Operate & Measure) > Ongoing maintenance for AI-touched systems - provider migrations, prompt tuning, dependency upgrades and the quiet engineering work that keeps production systems running as the AI landscape shifts. ### Summary dfzoo AI Institute provides ongoing maintenance for AI-touched applications. AI systems need maintenance that traditional applications do not: provider migrations when a model is deprecated, prompt tuning as inputs evolve, eval re-baselining, dependency upgrades for the LLM SDK layer, cost monitoring as spend drifts. Retainer model fits teams with continuous AI features; fixed engagements fit specific events like a provider migration. ### Who it's for - Engineering teams with AI features in production who need part-time senior support - CTOs whose team built the AI feature and moved on; nobody owns it now - Product teams whose AI providers deprecate models faster than the team can react - Mid-stage companies extending or maintaining AI features without growing the team ### Problems we solve - Provider deprecates a model; the team has no migration plan and no time to write one - Prompt that worked at launch produces worse outputs six months later - drift - LLM SDK version is stuck at launch-day version; security patches piling up - Cost trends are rising slowly enough that nobody noticed until the quarterly review ### What you get - Monthly maintenance hours covering the issues above (retainer) - Provider migration plans + execution when models deprecate (eval-protected) - Quarterly prompt tune-up: re-eval, regression check, minor tuning - Dependency upgrade cadence: SDKs, observability libraries, eval frameworks - Quarterly written report: what changed, what to expect, what to budget for next quarter ### How we work 1. **1. Handover** (Week 1-2): Read the system. Map dependencies, prompts, eval setup, observability and on-call procedures. Document gaps. 2. **2. Stabilization (if needed)** (Week 2-4 (one-time)): Address critical maintenance debt before steady-state retainer: missing evals, ungoverned prompts, stale SDKs. 3. **3. Steady-state retainer** (Ongoing): Monthly maintenance hours, quarterly check-ins, ad-hoc support for provider events or incidents. 4. **4. Quarterly review** (Quarterly): Written report on what we touched, what is drifting, what to budget next quarter. ### FAQ **Q: Retainer or fixed engagement?** Both work. Retainer (10-40 hours/month) fits teams with continuous AI features in production. Fixed engagements (provider migration, eval re-baselining, prompt re-tune) fit teams with a specific event. **Q: What if there is a production incident?** Retainer clients get ad-hoc incident support - we join the bridge, help debug, write the post-mortem. SLAs are agreed in the retainer contract. **Q: Do you handle full DevOps / SRE?** We handle AI-specific operations (prompt drift, model migrations, eval pipelines, LLM cost monitoring). Generic SRE work (Kubernetes, infrastructure-as-code, traditional alerting) is in scope only when it intersects with the AI features. **Q: How do model migrations work?** Standard maintenance event. We re-baseline eval against the new model, run regression tests, migrate prompts where the new model needs different patterns, measure before/after on quality, latency and cost. **Q: Can you maintain code we did not build?** Yes. Handover phase covers reading and documenting the system. We may need to address maintenance debt before steady-state retainer becomes possible. ### Related - https://dfzoo.ai/services/operate-measure-maintain/llm-observability - https://dfzoo.ai/services/operate-measure-maintain/cost-model-optimization - https://dfzoo.ai/services/build-automate-modernize/ai-powered-custom-development --- ================================================================ POLISH CONTENT (drafts) ================================================================ # Pełny content witryny dfzoo AI Institute - wersja polska (drafty Claude'a) _Drafty Claude'a, finalny polish przez OpenAI offline. Aktualizowane przy każdym buildzie z anchors.pl.ts._ --- # Strona główna - dfzoo AI Institute **URL:** https://dfzoo.ai/pl > Niezależny audyt AI, ocena kodu i szkolenia inżynierskie. dfzoo AI Institute pomaga organizacjom oceniać kod generowany przez AI, szkolić zespoły, wdrażać narzędzia AI, automatyzować procesy i modernizować systemy legacy do AI-ready platform - bez vendor lock-in. - Certyfikat ISO 9001:2015 (DNV Business Assurance, nr 10000439701-MSC-RvA-POL) - Biura: Szczecin (siedziba, ul. Wawrzyniaka 6W) i Warszawa - Języki: angielski, polski - Czas reakcji na zapytanie: jeden dzień roboczy Strona zorganizowana wokół czterech super-anchorów (Trenuj i Wdrażaj, Oceniaj i Zabezpiecz, Buduj i Modernizuj, Operuj i Mierz), pod którymi siedzi 18 konkretnych praktyk serwisowych. --- # O dfzoo AI Institute **URL:** https://dfzoo.ai/pl/o-nas > Instytut inżynierski zajmujący się niezależnym audytem AI. dfzoo AI Institute świadczy niezależne usługi assurance dla AI w środowisku produkcyjnym: audyt bezpieczeństwa, ocena kodu, szkolenia zespołów oraz wsparcie operacyjne. Instytut jest prowadzony przez DFZOO.COM Sp. z o.o. - polską spółkę działającą od 2014 roku, posiadającą certyfikat ISO 9001:2015 wydany przez DNV Business Assurance (nr 10000439701-MSC-RvA-POL), prowadzącą działalność na terenie całej UE pod numerem VAT-EU. ### Pełna transparentność: stronę tworzy i utrzymuje Claude pod kierunkiem zespołu inżynierskiego Strona dfzoo AI Institute - kod, treści, dane strukturalne, routing i wdrożenia - powstaje dzięki Claude'owi (asystentowi AI od Anthropic) pracującemu z zespołem inżynierskim dfzoo. Ten sam workflow developmentu, proces code review i dyscyplina wdrożeniowa zastosowane tutaj są udokumentowane w naszych procedurach usługowych. ### Trzy rzeczy, które warto zapamiętać - **Prowadzeni przez inżynierów** - Konsultanci to inżynierowie. Trenerzy piszą i wdrażają kod. Oceny przygotowują ci sami ludzie, którzy budują podobne systemy. Nie oddzielamy sprzedaży od realizacji. - **Certyfikat ISO 9001:2015** - System zarządzania jakością certyfikowany przez DNV Business Assurance (nr 10000439701-MSC-RvA-POL). Dla działów zakupów to sygnał powtarzalności i audytowalności. - **Efekty mierzone na produkcji** - Efekty projektu mierzymy względem uzgodnionych wskaźników produkcyjnych: koszt na żądanie, P95 latency, poziom błędów, częstotliwość wdrożeń. ### Podmiot prawny dfzoo AI Institute jest prowadzony przez DFZOO.COM Sp. z o.o. - polską spółkę z ograniczoną odpowiedzialnością działającą od 2014 roku, zarejestrowaną pod numerem VAT-EU. Siedziba w Szczecinie, biuro w Warszawie. ### Zakres certyfikatu ISO 9001:2015 - Projektowanie i tworzenie aplikacji webowych i mobilnych - E-commerce i strony www - Application service - Szkolenia IT i doradztwo - Online marketing - Robotyczna automatyzacja procesów biznesowych ### Biura - **Szczecin** (siedziba): ul. Wawrzyniaka 6W, 70-393 Szczecin, Polska. Inżynieria, ewaluacja i operacje. - **Warszawa**: client-facing. Rozmowy wstępne, warsztaty, on-site delivery dla klientów z centralnej i północnej Polski. --- # Kontakt - dfzoo AI Institute **URL:** https://dfzoo.ai/pl/kontakt > Porozmawiaj z inżynierem o AI. Odpowiadamy w ciągu jednego dnia roboczego, po polsku albo po angielsku. ### Biura - **Szczecin (siedziba)**: ul. Wawrzyniaka 6W, 70-393 Szczecin, Polska. Godziny biura: poniedziałek-piątek, 09:00-17:00 CET. Wizyty po umówieniu. - **Warszawa**: po umówieniu. Client-facing: Rozmowy wstępne, warsztaty i on-site delivery dla klientów z centralnej i północnej Polski. ### Kanały kontaktu - **Sprzedaż i nowe projekty**: start@dfzoo.ai - rozmowy wstępne, scoping, RFP, procurement - **Prasa i partnerstwa**: start@dfzoo.ai - zapytania medialne, wystąpienia, partnerstwa ### Biznes i legal - Nazwa prawna: DFZOO.COM Sp. z o.o. (marka: dfzoo AI Institute) - VAT-EU: PL (do uzupełnienia) - KRS: do uzupełnienia - Adres rejestrowy: ul. Wawrzyniaka 6W, 70-393 Szczecin, Polska --- # Usługi - dfzoo AI Institute **URL:** https://dfzoo.ai/pl/uslugi > Cztery super-anchory praktyki dfzoo AI Institute. Każdy anchor grupuje praktyki, które pod nim dostarczamy. Każdy projekt w dfzoo AI Institute żyje pod jednym z czterech super-anchorów. Wybierz sytuację, w której jesteście - praktyki rozwiązujące ją są zgrupowane pod spodem. --- # Trenuj, Wdrażaj i Zarządzaj (anchor 01) **URL:** https://dfzoo.ai/pl/uslugi/train-adopt-govern **Highlight:** Ścieżki finansowane (BUR/PARP) > Przeprowadź organizację od pojedynczych eksperymentów z AI do powtarzalnej, kontrolowanej praktyki - przez szkolenia, doradztwo i podstawowe zasady nadzoru nad AI, które utrzymają produkcyjne AI w ryzach. ### Podsumowanie dfzoo AI Institute pomaga organizacjom traktować AI jako mierzalną kompetencję zespołu, a nie jednorazowy eksperyment. Prowadzimy programy szkoleń AI dla programistów, projektantów i zespołów biznesowych; doradzamy liderom przy planach wdrożeń i wyborze narzędzi; wdrażamy konkretne narzędzia, z których zespół faktycznie korzysta na co dzień; i ustalamy podstawy nadzoru nad AI - politykę użycia, cykle ewaluacji, rejestr ryzyka - które przejdą przez działy zakupów i dział prawny. Certyfikat ISO 9001:2015. Wybrane programy są dostępne w ramach BUR/PARP. ### Dla kogo to jest - Liderzy inżynierscy wdrażający asystentów kodowania AI w swoich zespołach - Szefowie L&D i HR kupujący praktyczne szkolenia AI w ramach BUR/PARP - Szefowie operacji budujący wewnętrzne centrum kompetencji AI - Założyciele, których zespół ma być produktywny z narzędziami AI w tygodnie, nie kwartały ### Problemy, które rozwiązujemy - Zespoły używają narzędzi AI doraźnie, bez wspólnych standardów i ewaluacji - Działy zakupów i dział prawny blokują wdrożenie AI, bo nie ma żadnych podstaw nadzoru - Budżet na szkolenia jest, ale żaden dostawca nie umie pokazać wpływu na produkcję - Narzędzia kupione, ale korzysta z nich tylko garstka entuzjastów ### Praktyki pod tym anchorem - Programy szkoleń AI - https://dfzoo.ai/pl/uslugi/train-adopt-govern/ai-training-programs - Doradztwo AI - https://dfzoo.ai/pl/uslugi/train-adopt-govern/ai-advisory - Wdrożenie narzędzi - https://dfzoo.ai/pl/uslugi/train-adopt-govern/tools-implementation - Podstawy nadzoru nad AI - https://dfzoo.ai/pl/uslugi/train-adopt-govern/ai-governance-basics - EU AI Act Technical Compliance - https://dfzoo.ai/pl/uslugi/train-adopt-govern/eu-ai-act-compliance ### FAQ **Q: Czy szkolenia mogą być finansowane ze środków publicznych?** Wybrane szkolenia AI i doradztwo mogą być realizowane w ramach BUR/PARP - w zależności od aktualnych zasad kwalifikacji i operatora. Pomagamy sprawdzić możliwości. **Q: Szkolicie programistów czy użytkowników biznesowych?** Jednych i drugich. Ścieżki inżynierskie skupiają się na programowaniu wspieranym przez AI, ewaluacji kodu i wzorcach produkcyjnych. Ścieżki biznesowe - na projektowaniu promptów, automatyzacji procesów i bezpiecznym używaniu AI na co dzień. **Q: Jakie zasady nadzoru wdrażacie?** Podstawy gotowe pod działy zakupów: polityka użycia AI, lista zatwierdzonych narzędzi, cykl ewaluacji wyników produkcyjnych i podstawowy rejestr ryzyka. Bez ciężkiego frameworka - tyle, ile dział prawny i security potrzebują, żeby podpisać. **Q: Jak długo trwa typowy projekt?** Warsztaty: 1-3 dni. Doradztwo: zwykle 4-8 tygodni. Wdrożenie narzędzi: 2-6 tygodni, zależnie od wielkości zespołu i technologii. **Q: Co mierzycie?** Wskaźnik wykorzystania w przeliczeniu na zespół, średni czas zaoszczędzony dzięki AI na programistę na tydzień, odsetek PR-ów tworzonych z udziałem AI, poziom zgodności z zasadami nadzoru. Konkretne metryki ustalamy na rozmowie wstępnej (rozmowa wstępna). --- ## Programy szkoleń AI - Trenuj i Wdrażaj **URL:** https://dfzoo.ai/pl/uslugi/train-adopt-govern/ai-training-programs **Service type:** Szkolenia zawodowe **Anchor:** Trenuj, Wdrażaj i Zarządzaj (Trenuj i Wdrażaj) > Szkolenia AI dopasowane do roli, które zespół faktycznie stosuje w pracy - ścieżki inżynierskie, projektowe, operacyjne i dla liderów w formacie warsztatu lub kilkutygodniowej kohorty. ### Podsumowanie dfzoo AI Institute prowadzi programy szkoleń AI projektowane przez praktyków, którzy wdrażają AI na produkcji. Ścieżki inżynierskie obejmują programowanie wspierane przez AI z aktualnymi narzędziami (Claude Code, Cursor, Copilot, Aider), wzorce ewaluacji kodu i zabezpieczenia produkcyjne. Ścieżki biznesowe - projektowanie promptów, automatyzację procesów i bezpieczne używanie narzędzi AI na co dzień. Ścieżki dla liderów - plany wdrożeń i mierzenie ROI. Programy prowadzimy jako 1-3 dniowe warsztaty albo 4-6 tygodniowe kohorty; wybrane kwalifikują się do finansowania BUR/PARP w Polsce. ### Dla kogo to jest - Menedżerowie inżynierscy wdrażający asystentów kodowania AI w swoich zespołach - Szefowie L&D kupujący praktyczne szkolenia AI w ramach BUR/PARP - Projektanci produktu i PM-owie ulepszający swój codzienny sposób pracy z AI - Liderzy C-level budujący ogólnofirmowe kompetencje AI ### Problemy, które rozwiązujemy - Ogólne szkolenia AI od dostawców nie przekładają się na waszą bazę kodu i sposób pracy - Inżynierowie używają narzędzi AI, ale nie potrafią powiedzieć kolegom, co jest dobre - Budżet jest, ale żaden dostawca sam nie wdrożył tego, czego uczy - Niejasny wybór formatu - warsztat czy kohorta - przy danej wielkości zespołu i celach ### Co dostajecie - Program dopasowany do roli, waszych technologii, narzędzi i przypadków użycia - Szkolenie na żywo (w Warszawie/Szczecinie albo zdalnie) - Ćwiczenia praktyczne na waszej rzeczywistej bazie kodu albo procesach - Nagrania sesji i pisemny przewodnik referencyjny dla każdej ścieżki - Spotkanie kontrolne 4-6 tygodni po programie - mierzymy faktyczne wykorzystanie ### Jak pracujemy 1. **1. Rozmowa o zakresie** (Tydzień 0): Potwierdzamy odbiorców, obecne wykorzystanie narzędzi AI, cele biznesowe i możliwość dofinansowania. 2. **2. Projektowanie programu** (Tydzień 1-2): Personalizujemy agendę, dobieramy narzędzia do omówienia, przygotowujemy ćwiczenia dopasowane do waszej bazy kodu. 3. **3. Szkolenie na żywo** (1-3 dni (warsztat) albo 4-6 tygodni (kohorta)): Prowadzimy warsztat lub kohortę. Ćwiczenia praktyczne, sesje pytań i odpowiedzi, praca na prawdziwych problemach klienta. 4. **4. Spotkanie kontrolne** (Tydzień +4-6): Mierzymy wykorzystanie po szkoleniu, identyfikujemy przeszkody, rekomendujemy kolejne działania. ### FAQ **Q: Jaka jest różnica między warsztatem a kohortą?** Warsztat zamyka intensywny program w 1-3 kolejnych dniach - najlepszy do szybkiego startu. Kohorta rozciąga się na 4-6 tygodni z cotygodniowymi sesjami i pracą domową - najlepsza do zmiany codziennych nawyków w większych zespołach. **Q: Czy szkolenie może objąć naszą rzeczywistą bazę kodu?** Tak. Ćwiczenia budujemy z waszych repozytoriów pod NDA. Inżynierowie uczą się wzorców pracy z AI bezpośrednio na kodzie, na którym pracują następnego dnia. **Q: Czy dostępne jest finansowanie BUR/PARP?** Wybrane programy kwalifikują się do BUR/PARP. Kwalifikowalność zależy od wielkości firmy, sektora i aktualnych zasad operatora. Sprawdzamy razem na rozmowa wstępna. **Q: Jakie narzędzia obejmujecie?** Aktualne asystenty AI klasy produkcyjnej: Claude Code, Cursor, Copilot, Aider, Windsurf. Dla ścieżek biznesowych omawiamy też projektowanie promptów i narzędzia do budowy procesów na bazie LLM (n8n, Zapier + LLM, własne agenty). **Q: Jak mierzycie, czy szkolenie zadziałało?** Ankieta przed/po, czas zaoszczędzony dzięki AI mierzony na programistę na tydzień, odsetek PR-ów tworzonych z udziałem AI, kontrola jakości kodu po programie. Konkretne metryki ustalamy na rozmowie o zakresie. ### Powiązane - https://dfzoo.ai/pl/uslugi/train-adopt-govern/ai-advisory - https://dfzoo.ai/pl/uslugi/train-adopt-govern/tools-implementation - https://dfzoo.ai/pl/uslugi/evaluate-audit-secure/ai-code-evaluation --- ## Doradztwo AI - Trenuj i Wdrażaj **URL:** https://dfzoo.ai/pl/uslugi/train-adopt-govern/ai-advisory **Service type:** Doradztwo strategiczne **Anchor:** Trenuj, Wdrażaj i Zarządzaj (Trenuj i Wdrażaj) > Niezależne doradztwo dla organizacji układających kolejność wdrożenia AI - co najpierw pilotować, jakie narzędzia kupować, co odłożyć i jak mierzyć każdy krok. ### Podsumowanie dfzoo AI Institute prowadzi niezależne doradztwo AI dla dyrektorów, CTO i szefów operacji. Oceniamy obecny stan AI w organizacji (jakie narzędzia są zlicencjonowane, jakie piloty ruszyły, jakie są zasady nadzoru), rekomendujemy plan wdrożenia na 6-12 miesięcy z ustaloną kolejnością inwestycji i pomagamy zespołowi podejmować decyzje o narzędziach na podstawie rzeczywistej ewaluacji. Otrzymujecie: pisemny plan wdrożenia z uzasadnieniem, matrycę oceny narzędzi i kwartalne przeglądy z zarządem. ### Dla kogo to jest - CTO decydujący, jakie inicjatywy AI finansować w tym roku obrotowym - Szefowie operacji budujący wewnętrzne centrum kompetencji AI - CEO mniejszych firm ustalający strategię AI bez CTO w zespole - Zarządy proszące o niezależną ocenę planu AI ### Problemy, które rozwiązujemy - Za dużo narzędzi AI do oceny, za mało czasu na pilotaż każdego - Wewnętrzni zwolennicy AI nie zgadzają się co do planu, a kierownictwo nie umie rozsądzić - Dział zakupów prosi o niezależną drugą opinię na duży kontrakt AI - Pilotaże się udają, ale nigdy nie skalują się na produkcję ### Co dostajecie - Pisemny plan wdrożenia AI na 6-12 miesięcy z ustaloną kolejnością inwestycji - Matryca oceny narzędzi (koszt, dopasowanie, ryzyko, koszt zmiany) dla najlepszych kandydatów - Warsztat z interesariuszami podsumowujący rekomendacje i kolejne kroki - Kwartalne przeglądy z zarządem przez czas trwania projektu - Opcjonalnie: wsparcie wyboru dostawcy przy konkretnym kontrakcie o dużym znaczeniu ### Jak pracujemy 1. **1. Audyt sytuacji** (Tydzień 1-2): Rozmowy z kluczowymi interesariuszami, inwentaryzacja obecnych narzędzi AI i pilotaży, przegląd stanu nadzoru nad AI. 2. **2. Szkic planu** (Tydzień 2-4): Dopasowujemy opcje do waszych celów biznesowych, tolerancji ryzyka i możliwości zespołu. Ustalamy kolejność inwestycji. 3. **3. Warsztat z interesariuszami** (Tydzień 4-5): Prezentujemy rekomendacje kierownictwu, dopracowujemy je na podstawie uwag, zamykamy plan. 4. **4. Kwartalny przegląd** (Kwartalnie (na bieżąco)): Śledzimy postęp wobec kamieni milowych planu, dostosowujemy go do nowych narzędzi i zmian rynkowych. ### FAQ **Q: Czy jesteście niezależni od dostawców narzędzi?** Tak. Nie bierzemy prowizji od dostawców narzędzi AI. Rekomendacje opieramy na rzeczywistej ewaluacji wobec waszych technologii i celów. **Q: Jak długo trwa typowe doradztwo?** Pierwsza wersja planu: 4-8 tygodni. Kwartalne przeglądy trwają tak długo, jak współpraca. **Q: Pracujecie z nietechnicznym kierownictwem?** Tak. Tłumaczymy inżynierskie kompromisy na język biznesowy. Wiele projektów zaczyna się od CEO albo COO, który potrzebuje oceny na poziomie CTO, nie mając CTO w zespole. **Q: Pomożecie przy wyborze konkretnego dostawcy?** Tak. Możemy zawęzić projekt do jednej decyzji kontraktowej o dużym znaczeniu (platforma LLM, licencja asystenta kodowania, framework agentowy). Otrzymujecie pisemną rekomendację z uzasadnieniem. **Q: Co dalej po dostarczeniu planu?** Większość klientów wybiera kwartalny przegląd w ramach stałej współpracy, żeby utrzymać plan aktualny wobec zmian rynkowych. Inni wracają doraźnie do konkretnych decyzji. ### Powiązane - https://dfzoo.ai/pl/uslugi/train-adopt-govern/ai-governance-basics - https://dfzoo.ai/pl/uslugi/train-adopt-govern/ai-training-programs - https://dfzoo.ai/pl/uslugi/build-automate-modernize/agentic-ai-readiness --- ## Wdrożenie narzędzi - Trenuj i Wdrażaj **URL:** https://dfzoo.ai/pl/uslugi/train-adopt-govern/tools-implementation **Service type:** Wdrożenie IT **Anchor:** Trenuj, Wdrażaj i Zarządzaj (Trenuj i Wdrażaj) > Praktyczne wdrożenie asystentów kodowania AI, plików kontekstowych i procesów zespołowych - żeby narzędzia, które zlicencjonowaliście, były używane dłużej niż tydzień, na regułach faktycznie zapisanych w repozytorium, z wykorzystaniem, które widać na dashboardzie. ### Podsumowanie dfzoo AI Institute wdraża narzędzia AI w zespołach inżynierskich: konfiguracja asystentów kodowania (Claude Code, Cursor, Copilot) z ustawieniami dopasowanymi do zespołu, napisanie plików kontekstowych, które agent czyta, zanim dotknie repozytorium, budowa wspólnych bibliotek promptów oraz integracja narzędzi z IDE, systemem kontroli wersji i CI. Rozdzielamy pisanie od przeglądu: agent nie powinien recenzować kodu, który przed chwilą napisał, więc ustawiamy osobny kontekst przeglądu i bramkę review niezależną od agenta, który wykonał zmianę. Każde wdrożenie kończy się dashboardem pomiarowym - adopcja, odsetek przyjętych podpowiedzi, powiązanie obu z waszymi metrykami dostarczania - a nie samym plikiem konfiguracyjnym. Pracujemy ramię w ramię z zespołem przez 2-6 tygodni, dopóki wykorzystanie się nie ustabilizuje. ### Dla kogo to jest - Menedżerowie inżynierscy, których zespół zlicencjonował narzędzia AI, ale wdrożenie utknęło - Dyrektorzy inżynierii wdrażający nowy asystent kodowania w wielu zespołach - Zespoły platformowe budujące wewnętrzne narzędzia AI dla szerszej organizacji - Tech leadowie potrzebujący pomocy w konfiguracji narzędzi pod swoje technologie ### Problemy, które rozwiązujemy - Domyślne konfiguracje narzędzi nie pasują do waszej bazy kodu i konwencji - Agent nie dostaje spisanego kontekstu, więc w każdej sesji odkrywa wasze konwencje od nowa, a resztę zgaduje - Ten sam agent pisze zmianę i ją recenzuje, więc bramka review potwierdza własną pracę - Narzędzia AI i CI/CD nie rozmawiają ze sobą - proces przeglądu kodu jest zepsuty - Nikt nie potrafi powiedzieć, czy narzędzia są używane, czy podpowiedzi są przyjmowane i czy to się zwraca - nie ma liczby, którą można wskazać ### Co dostajecie - Pliki kontekstowe dla agentów: konwencje repozytorium, mapa kontekstu bazy kodu, reguły przeglądu oraz procedura i właściciel ich aktualizacji - Konfiguracja asystenta kodowania dopasowana do zespołu i waszych technologii - Wspólna biblioteka promptów z najczęściej używanymi wzorcami - Układ Writer/Reviewer: osobny kontekst przeglądu i bramka review niezależna od agenta, który napisał zmianę - Dokumentacja integracji z IDE, systemem kontroli wersji i CI oraz przykładowe procesy - Przewodnik wdrożeniowy dla nowych inżynierów dołączających do zespołu - Dashboard wykorzystania: adopcja, odsetek przyjętych podpowiedzi i powiązanie z waszymi metrykami dostarczania, omówiony 30 dni po wdrożeniu ### Jak pracujemy 1. **1. Rozpoznanie narzędzi i technologii** (Tydzień 1): Audyt obecnych licencji narzędzi, bazy kodu, preferencji co do IDE i istniejących promptów. Ustalamy, wobec których metryk dostarczania będziemy mierzyć wdrożenie. 2. **2. Pliki kontekstowe i konfiguracja** (Tydzień 2-3): Piszemy pliki kontekstowe, ustawiamy konfiguracje pod narzędzia, budujemy bibliotekę promptów, uruchamiamy integrację z CI/CD i osobną bramkę przeglądu. 3. **3. Wdrożenie w zespołach** (Tydzień 3-5): Praktyczne sesje wdrożeniowe w każdym zespole, programowanie w parach z inżynierami, rozwiązywanie problemów na żywo. 4. **4. Pomiar i przegląd wykorzystania** (Tydzień +6): Uruchamiamy dashboard, czytamy pierwsze 30 dni danych, identyfikujemy opornych, rekomendujemy kolejne kroki, przekazujemy dashboard. ### FAQ **Q: Jakie asystenty kodowania AI i jakie IDE wdrażacie?** Claude Code, Cursor, Copilot, Windsurf, Aider, Cody - cokolwiek wasz zespół już zlicencjonował albo chce sprawdzić, w VS Code, JetBrains, Cursorze i Neovimie. Konfiguracje testujemy pod każde IDE i przekazujemy jako wspólną konfigurację gotową do wrzucenia do repozytorium. Jesteśmy niezależni od konkretnych narzędzi; naszą wartością jest konfiguracja i wdrożenie, nie prowizja od dostawcy. **Q: Czym dokładnie są pliki kontekstowe i do kogo należą po wdrożeniu?** To pliki, które agent czyta, zanim cokolwiek napisze: konwencje repozytorium, mapa kontekstu mówiąca, co gdzie leży i dlaczego, oraz reguły przeglądu, którym podlega zmiana. Żyją w waszym repozytorium i należą do was, a przekazujemy je z procedurą utrzymania: kto je aktualizuje, kiedy i co wymusza przepisanie. **Q: Czy ten sam agent może napisać kod i go zrecenzować?** Może i właśnie ten błąd projektujemy poza system. Agent recenzujący własny wynik powtarza własne założenia, więc bramka review działa w osobnym kontekście, z innymi regułami i nigdy pod autorem zmiany: druga konfiguracja agenta, a przy wszystkim, co dotyka produkcji, człowiek. **Q: Czym to się różni od szkoleń AI?** Szkolenia uczą koncepcji i wzorców. Wdrożenie narzędzi jest praktyczne: konfigurujemy wasze konkretne narzędzia, piszemy wasze konkretne pliki kontekstowe i prompty, integrujemy z waszym konkretnym CI. Większość klientów kupuje oba - najpierw szkolenia, potem wdrożenie. **Q: A co z bezpieczeństwem i zgodnością?** Konfigurujemy narzędzia tak, żeby respektowały granice waszych danych (żaden wyciek kodu do publicznych LLM bez zgody, lista dozwolonych repozytoriów). Security review i podstawy zgodności realizujemy w ramach Security Review oraz Podstaw nadzoru nad AI. ### Powiązane - https://dfzoo.ai/pl/uslugi/train-adopt-govern/ai-training-programs - https://dfzoo.ai/pl/uslugi/evaluate-audit-secure/security-review - https://dfzoo.ai/pl/uslugi/build-automate-modernize/ai-powered-custom-development --- ## Podstawy nadzoru nad AI - Trenuj i Wdrażaj **URL:** https://dfzoo.ai/pl/uslugi/train-adopt-govern/ai-governance-basics **Service type:** Doradztwo compliance **Anchor:** Trenuj, Wdrażaj i Zarządzaj (Trenuj i Wdrażaj) > Minimalne podstawy nadzoru nad AI gotowe pod działy zakupów: polityka użycia AI, lista zatwierdzonych narzędzi, cykl ewaluacji i rejestr ryzyka - tyle, ile trzeba, żeby odblokować podpis działu prawnego i security. ### Podsumowanie dfzoo AI Institute ustala minimalne podstawy nadzoru nad AI dla organizacji wdrażających AI wewnętrznie. Piszemy politykę użycia AI dopasowaną do branży, budujemy listę zatwierdzonych narzędzi z regułami obchodzenia się z danymi, ustalamy cykl ewaluacji wyników AI trafiających na produkcję i tworzymy podstawowy rejestr ryzyka, który akceptują zespoły prawne i security. Bez ciężkiego frameworka - najmniejszy zbiór dokumentów, który odblokowuje wdrożenie AI bez tworzenia dokumentacji do szuflady. ### Dla kogo to jest - Szefowie działów prawnych i compliance proszeni o zgodę na wdrożenie AI - CISO, których zespoły inżynierskie chcą zlicencjonować asystenty kodowania AI - COO uruchamiający wewnętrzne użycie AI w działach poza inżynierią - Sektor publiczny i organizacje regulowane potrzebujące dających się obronić podstaw ### Problemy, które rozwiązujemy - Inżynieria chce używać narzędzi AI, ale dział prawny nie ma się na czym oprzeć przy podpisie - Każdy zespół pisze własne zasady użycia AI - niespójne i niemożliwe do wyegzekwowania - Rejestr ryzyka nie ma wpisów dotyczących AI; audytorzy zgłaszają lukę - Ewaluacja wyników AI jest nieformalna; nikt nie udowodni jakości przy audycie ### Co dostajecie - Pisemna Polityka Użycia AI (jeden dokument, obejmuje pracowników i współpracowników) - Lista Zatwierdzonych Narzędzi z regułami obchodzenia się z danymi dla każdego narzędzia - Opis cyklu ewaluacji (które wyniki AI są sprawdzane, jak i przez kogo) - Wpisy w rejestrze ryzyka dla użycia AI, dopasowane do waszego frameworka ryzyka - Jednostronicowa ściąga dla menedżerów do codziennego użytku ### Jak pracujemy 1. **1. Rozpoznanie branży i ryzyka** (Tydzień 1): Rozmowy z działem prawnym, compliance i inżynierią; zrozumienie obecnego frameworka ryzyka i ograniczeń branżowych. 2. **2. Szkic dokumentów** (Tydzień 2-3): Przygotowujemy politykę użycia, listę zatwierdzonych narzędzi, cykl ewaluacji i wpisy do rejestru ryzyka. 3. **3. Przegląd z interesariuszami** (Tydzień 3-4): Prowadzimy dział prawny, compliance, security i inżynierię przez szkice. Dopracowujemy je. 4. **4. Wdrożenie i szkolenie** (Tydzień 4-5): Szkolimy menedżerów z codziennego stosowania zasad, publikujemy dokumenty wewnętrznie. ### FAQ **Q: Czy to pełny framework zarządzania ryzykiem AI w stylu NIST AI RMF?** Nie. To minimum gotowe pod działy zakupów - tyle, ile trzeba, żeby odblokować użycie AI i przejść standardowe audyty. Organizacje z dojrzałymi programami ryzyka mogą później nałożyć NIST AI RMF, ISO 42001 albo podobne. Możemy to zrobić w kolejnym etapie. **Q: Czy polityka uwzględni nasze branżowe wymogi?** Tak. Dopasowujemy polityki do branży - fintech (KNF/MiFID/DORA), ochrona zdrowia (RODO / dane medyczne), sektor publiczny, regulowany tech. Nie jesteśmy prawnikami, ale przygotowujemy dokumenty, które wasz dział prawny zatwierdza. **Q: Czym to się różni od kupienia gotowego szablonu polityki AI?** Gotowe szablony pisane są pod abstrakcję i pękają przy pierwszym pytaniu z działu zakupów. My dopasowujemy je do waszych narzędzi, waszej branży i waszych realiów zakupowych. Efekt broni się przy audycie, a nie tylko odhacza pole. **Q: Możecie wdrożyć też narzędzia ewaluacyjne, nie tylko napisać politykę?** Tak. Opis cyklu ewaluacji może iść w parze z wdrożeniem narzędzi w ramach LLM Observability (pipeline'y eval, automatyczna ocena punktowa) albo AI Code Evaluation (powtarzające się audyty kodu). **Q: Jak długo do podpisania dokumentów?** Typowy projekt: 4-6 tygodni od startu do opublikowanych, podpisanych dokumentów. Wąskim gardłem są zwykle wewnętrzne cykle przeglądu, nie samo pisanie. ### Powiązane - https://dfzoo.ai/pl/uslugi/train-adopt-govern/ai-advisory - https://dfzoo.ai/pl/uslugi/evaluate-audit-secure/security-review - https://dfzoo.ai/pl/uslugi/operate-measure-maintain/llm-observability --- ## EU AI Act Technical Compliance - Trenuj i Wdrażaj **URL:** https://dfzoo.ai/pl/uslugi/train-adopt-govern/eu-ai-act-compliance **Service type:** Techniczna ocena zgodności z EU AI Act i dokumentacja **Anchor:** Trenuj, Wdrażaj i Zarządzaj (Trenuj i Wdrażaj) > Inżynierska połowa zgodności z AI Act: inwentaryzujemy i klasyfikujemy wasze systemy AI, piszemy dokumentację techniczną, przygotowujemy dowody z testów i zamieniamy rejestrowanie zdarzeń, identyfikowalność oraz nadzór ze strony człowieka w architekturę, która naprawdę działa. ### Podsumowanie dfzoo AI Institute dostarcza techniczną część zgodności z EU AI Act organizacjom, które budują, kupują albo stosują systemy AI. Inwentaryzujemy wszystkie używane systemy AI, klasyfikujemy każdy z nich według poziomów ryzyka przewidzianych w rozporządzeniu, piszemy dokumentację techniczną wymaganą przez obecnie obowiązujące przepisy i przygotowujemy dowody z testów adwersarialnych, które bronią się przy audycie. Rejestrowanie zdarzeń, identyfikowalność i nadzór ze strony człowieka traktujemy jako wymagania architektoniczne z przypisanymi właścicielami i działającym wdrożeniem, a nie jako deklaracje w dokumencie polityki. Projekt kończy się analizą luk i wycenionym planem działań naprawczych oraz planem budowania kompetencji w zakresie AI, który wprost zasila szkolenia dla osób obsługujących te systemy. Jesteśmy inżynierami, nie prawnikami: budujemy dowody techniczne i pracujemy ramię w ramię z waszym działem prawnym albo naszym partnerem prawnym, do którego należy wykładnia i podpis. ### Dla kogo to jest - Startupy produktowe, których funkcja AI została właśnie zakwalifikowana jako wysokiego ryzyka w ankiecie od klienta - Software house'y i integratorzy, którzy muszą przekazać klientom dokumentację AI Act dla tego, co dostarczają - Zespoły korporacyjne i w sektorze publicznym stosujące kupione systemy AI i obciążone obowiązkami podmiotu stosującego - Organizacje regulowane w finansach, ochronie zdrowia, HR, edukacji i infrastrukturze krytycznej z terminami audytowymi przed sobą - Szefowie compliance, CISO, CTO i szefowie działów prawnych, którzy potrzebują artefaktów technicznych możliwych do podpisania ### Problemy, które rozwiązujemy - Nikt nie potrafi wskazać listy systemów AI faktycznie działających w organizacji, a tym bardziej ich klasyfikacji ryzyka - Klient, przetarg albo audytor poprosił o dokumentację techniczną AI Act, a nie ma niczego spisanego - Systemy logują zdarzenia biznesowe, ale nie decyzje modelu, dane wejściowe i wersje, których wymaga identyfikowalność - Nadzór ze strony człowieka istnieje na slajdzie, ale żaden interfejs, rola ani ścieżka eskalacji nie realizuje go na produkcji - Dział prawny przekazał listę obowiązków, a zespół inżynierski nie wie, co ma w odpowiedzi zbudować ### Co dostajecie - Inwentarz systemów AI z właścicielem, celem, przepływami danych, modelem i dostawcą dla każdego systemu - Klasyfikacja ryzyka każdego systemu według poziomów z rozporządzenia, z rozpisanym uzasadnieniem pod audyt i przegląd prawny - Pakiet dokumentacji technicznej dla każdego systemu w zakresie, ułożony według wymagań rozporządzenia - Specyfikacja rejestrowania zdarzeń, identyfikowalności i nadzoru ze strony człowieka wraz ze zmianami architektonicznymi potrzebnymi, żeby ją spełnić - Zbiór dowodów z ewaluacji jakości i przeglądu bezpieczeństwa kodu systemu, tam gdzie jest wymagany - Analiza luk i wyceniony plan działań naprawczych oraz plan kompetencji w zakresie AI przypisany do ról ### Jak pracujemy 1. **1. Inwentaryzacja i rozpoznanie** (Tydzień 1-2): Znajdujemy wszystkie używane systemy AI: zbudowane, kupione i wbudowane w narzędzia zewnętrzne. Dla każdego zapisujemy cel, użytkowników, dane, model, dostawcę i właściciela. 2. **2. Klasyfikacja ryzyka** (Tydzień 2-3): Klasyfikujemy każdy system według poziomów ryzyka i waszej roli przy nim: dostawca czy podmiot stosujący. Uzasadnienie spisujemy tak, żeby dział prawny mógł je potwierdzić albo zakwestionować. 3. **3. Analiza luk** (Tydzień 3-4): Porównujemy stan obecny z obowiązkami, które was dotyczą: dokumentacja, zarządzanie danymi, rejestrowanie zdarzeń, identyfikowalność, nadzór ze strony człowieka, dokładność i odporność, dowody z testów, kompetencje w zakresie AI. 4. **4. Dokumentacja i dowody** (Tydzień 4-7): Piszemy pakiet dokumentacji technicznej, projektujemy architekturę rejestrowania zdarzeń i nadzoru, kompletujemy zbiór dowodów z testów. Przeglądamy całość z waszym działem prawnym albo naszym partnerem prawnym. 5. **5. Plan naprawczy i przekazanie** (Tydzień 7-8, przegląd kwartalnie): Przekazujemy wyceniony plan z właścicielami i kolejnością prac, omawiamy go z kierownictwem i ustalamy cykl przeglądu, który utrzymuje dokumentację aktualną, gdy systemy się zmieniają. ### FAQ **Q: Czy to jest porada prawna?** Nie i mówimy to wprost. Jesteśmy zespołem inżynierskim. Przygotowujemy artefakty techniczne wymagane przez rozporządzenie - inwentarz, uzasadnienie klasyfikacji, dokumentację, projekt rejestrowania zdarzeń i nadzoru, dowody z testów - a wykładnia i podpis należą do waszego działu prawnego. Jeśli nie macie do tego prawnika, wprowadzamy do projektu partnera prawnego i jasno dzielimy zakres prac. **Q: Przepisy i ich wykładnia wciąż się zmieniają. Jak sobie z tym radzicie?** Pracujemy na obowiązkach obowiązujących w momencie realizacji projektu i datujemy każdą klasyfikację oraz każdy dokument. Tam, gdzie wykładnia jest realnie otwarta, oznaczamy to jako decyzję dla waszego prawnika, zamiast zgadywać, i zawsze zalecamy potwierdzenie prawne, zanim oprzecie się na klasyfikacji na zewnątrz. Kwartalny przegląd utrzymuje pakiet w zgodzie z aktualnymi systemami i wytycznymi. **Q: Ile to trwa i jak jest wyceniane?** Typowy projekt to 6-8 tygodni od startu do przekazanego pakietu dokumentacji i planu naprawczego. Wycena jest stała dla projektu, ustalana po rozmowie wstępnej (rozmowa wstępna), i zależy od liczby systemów AI w zakresie oraz tego, ile z nich trafia do kategorii wysokiego ryzyka. Sama inwentaryzacja z klasyfikacją jest dostępna jako mniejszy pierwszy krok. **Q: Czego potrzebujecie od nas na start?** Dostępu do osób odpowiadających za poszczególne systemy AI, istniejącej dokumentacji architektury i przepływów danych, umów i powierzeń danych dla kupionej AI oraz wszelkich wcześniejszych klasyfikacji i analiz prawnych. Braki uzupełniamy ustrukturyzowanymi wywiadami - większość organizacji odkrywa na tym etapie, że realny inwentarz jest większy niż ten na papierze. **Q: Używamy wyłącznie systemów AI od dostawców. Czy to nas dotyczy?** Tak. Stosowanie systemu AI wiąże się z własnymi obowiązkami, odrębnymi od obowiązków tego, kto go zbudował. Sprawdzamy, co faktycznie dostajecie od dostawców, wskazujemy, czego brakuje w ich dokumentacji, i piszemy artefakty po stronie podmiotu stosującego: projekt nadzoru, rejestrowanie zdarzeń, instrukcje użytkowania oraz pytania do dostawcy przed przedłużeniem umowy. **Q: Jak to się ma do testów i ewaluacji systemu?** Dowody z testów to jeden z artefaktów technicznych wspierających zgodność dla systemów wyższego ryzyka. Usługa AI Quality Evaluation dostarcza wyniki ewaluacji i oceniony zestaw testowy, a Security Review pokrywa warstwę kodu i aplikacji. Obie przekazują wnioski w formie, która wchodzi wprost do pakietu dokumentacji. Każdą z usług można kupić osobno; realizowane razem eliminują podwójne rozpoznanie i dają jeden spójny ślad dowodowy. **Q: Co z obowiązkiem kompetencji w zakresie AI?** To realny obowiązek i zarazem najłatwiejszy do zamknięcia. Przypisujemy rolom poziom rozumienia AI, jakiego każda z nich potrzebuje, i przekładamy to na plan szkoleń. Realizują go nasze programy szkoleniowe, a zapisy o ukończeniu stają się częścią waszych dowodów zgodności. **Q: Zdalnie czy na miejscu, i czy możecie pracować w naszym środowisku?** Domyślnie zdalnie, pod NDA. Wywiady i warsztaty prowadzimy na miejscu tam, gdzie szybciej przynosi to odpowiedzi, w tej samej cenie, a gdy materiały nie mogą opuścić organizacji, pracujemy w całości na waszych narzędziach i w waszych systemach dokumentowych. ### Powiązane - https://dfzoo.ai/pl/uslugi/train-adopt-govern/ai-governance-basics - https://dfzoo.ai/pl/uslugi/evaluate-audit-secure/ai-quality-evaluation - https://dfzoo.ai/pl/uslugi/train-adopt-govern/ai-training-programs --- # Oceniaj, Audytuj i Zabezpiecz (anchor 02) **URL:** https://dfzoo.ai/pl/uslugi/evaluate-audit-secure **Highlight:** AI Code Evaluation - międzynarodowa nisza > Niezależna ewaluacja kodu generowanego przez AI, security review i sprawdzenie gotowości produkcyjnej - żeby zespoły wdrażały szybciej z AI i wiedziały, co można bezpiecznie wypuścić. ### Podsumowanie dfzoo AI Institute dostarcza niezależną ewaluację i potwierdzenie jakości dla kodu generowanego przez AI i efektów pracy wspieranej przez AI. Audytujemy kod z asystentów kodowania pod kątem poprawności, bezpieczeństwa i utrzymywalności; prowadzimy security review systemów tworzonych z udziałem AI; projektujemy strategie pokrycia testowego dla projektów wspieranych przez AI; i certyfikujemy gotowość produkcyjną przed wydaniem. Nasza praktyka AI Code Evaluation to międzynarodowa nisza - jesteśmy jednym z niewielu prowadzonych przez inżynierów dostawców tej usługi z certyfikatem ISO 9001:2015. ### Dla kogo to jest - Liderzy inżynierscy wdrażający kod tworzony z udziałem AI, którzy potrzebują zewnętrznego potwierdzenia jakości - CTO firm regulowanych (fintech, ochrona zdrowia, sektor publiczny) pod presją zgodności w zakresie AI w kodzie - Zespoły security odpowiedzialne za produkty zbudowane z asystentami kodowania AI - Zespoły zarządzające dostawcami oceniające dostawy tworzone z udziałem AI od podwykonawców ### Problemy, które rozwiązujemy - Kod generowany przez AI trafia do PR szybciej, niż ludzie zdążą go dokładnie sprawdzić - Compliance prosi o dowód, że kod tworzony z udziałem AI spełnia ten sam standard co kod pisany przez ludzi - Pokrycie testowe spada, bo AI generuje kod wyglądający wiarygodnie, ale niesprawdzony - Zespoły security nie wiedzą, gdzie w bazie kodu skupia się wkład AI ### Praktyki pod tym anchorem - AI Code Evaluation - https://dfzoo.ai/pl/uslugi/evaluate-audit-secure/ai-code-evaluation - Security Review - https://dfzoo.ai/pl/uslugi/evaluate-audit-secure/security-review - Pokrycie testowe - https://dfzoo.ai/pl/uslugi/evaluate-audit-secure/test-coverage - Production Readiness - https://dfzoo.ai/pl/uslugi/evaluate-audit-secure/production-readiness - AI Quality Evaluation - https://dfzoo.ai/pl/uslugi/evaluate-audit-secure/ai-quality-evaluation ### FAQ **Q: Jak wygląda projekt AI Code Evaluation?** Stała cena, stały zakres. Bierzemy repozytorium lub zbiór PR-ów, próbkujemy kod tworzony z udziałem AI, stosujemy nasze kryteria oceny (poprawność, bezpieczeństwo, utrzymywalność, jakość testów), przygotowujemy pisemny raport z uszeregowanymi wnioskami i omawiamy z zespołem najważniejsze problemy. Typowy czas realizacji: 2-4 tygodnie. **Q: Zastępujecie przegląd kodu wykonywany przez ludzi?** Nie. Jesteśmy zewnętrznym potwierdzeniem jakości - okresowym, niezależnym, głębszym niż codzienny przegląd. Większość klientów korzysta z nas kwartalnie, obok wewnętrznego procesu przeglądu. **Q: Jakie narzędzia obejmujecie?** Każdy asystent kodowania oparty na LLM: Cursor, Claude Code, Cody, Copilot, Windsurf, Aider. Oceniamy wynik, nie narzędzie. **Q: Możecie pracować pod NDA dla branż regulowanych?** Tak. Standardowe NDA, możemy działać na izolowanej infrastrukturze, a przy szczególnie wrażliwych projektach pracować na miejscu w biurze klienta. **Q: Jak wygląda raport?** Ustrukturyzowany PDF plus uszeregowane wpisy w systemie zgłoszeń (Linear, Jira, GitHub Issues - do wyboru). Każdy wniosek zawiera poziom istotności, sposób odtworzenia, rekomendowaną poprawkę i linki do wskazanego kodu. Na życzenie sprawdzamy też wprowadzone działania naprawcze. --- ## AI Code Evaluation - Oceniaj i Zabezpiecz **URL:** https://dfzoo.ai/pl/uslugi/evaluate-audit-secure/ai-code-evaluation **Service type:** Audyt kodu **Anchor:** Oceniaj, Audytuj i Zabezpiecz (Oceniaj i Zabezpiecz) > Niezależny audyt kodu generowanego i tworzonego z udziałem AI, w stałej cenie - oceniany według 4-wymiarowych kryteriów (poprawność, bezpieczeństwo, utrzymywalność, jakość testów), które publikujemy w całości, zanim cokolwiek podpiszecie. ### Podsumowanie dfzoo AI Institute dostarcza niezależną AI Code Evaluation zespołom inżynierskim wdrażającym kod tworzony z udziałem AI. Próbkujemy pliki lub PR-y tworzone z udziałem AI z waszego repozytorium, oceniamy każdy wniosek według 4-wymiarowych kryteriów (poprawność, bezpieczeństwo, utrzymywalność, jakość testów), przygotowujemy pisemny raport z wnioskami uszeregowanymi według istotności i omawiamy z zespołem najważniejsze problemy. Kryteria są jawne: każdy wymiar ma spisaną definicję, nazwane podkryteria i skalę 0-4, więc znacie sposób oceny przed zakupem i możecie stosować ten sam standard między audytami. Z certyfikatem jakości ISO 9001:2015. Stała cena, realizacja 2-4 tygodnie. Jedna z niewielu prowadzonych przez inżynierów praktyk oceny kodu AI w Europie. ### Dla kogo to jest - Dyrektorzy inżynierii i CTO potrzebujący zewnętrznego potwierdzenia jakości kodu tworzonego z udziałem AI - CISO, których zespoły wdrażają kod generowany przez AI do regulowanych systemów produkcyjnych - Zespoły zarządzające dostawcami sprawdzające dostawy tworzone z udziałem AI od podwykonawców - Specjaliści ds. compliance zobowiązani do udowodnienia głębokości przeglądu kodu przy audycie ### Problemy, które rozwiązujemy - Kod generowany przez AI trafia szybciej, niż ludzie zdążą go dokładnie sprawdzić - Compliance prosi o dowód, że kod tworzony z udziałem AI spełnia ten sam standard co kod ludzki - Wnioski z audytu przychodzą jako opinie bez podanej metody, więc nikt nie potrafi ich powtórzyć ani podważyć - Pokrycie testowe spada, bo AI generuje kod wyglądający wiarygodnie, ale niesprawdzony - Security nie wie, gdzie w bazie kodu skupia się wkład AI ### Co dostajecie - Pisemny raport z audytu (ustrukturyzowany PDF, ~30-60 stron zależnie od zakresu) - Karta ocen per wymiar w jawnej skali 0-4, z kotwicą przy każdej ocenie i wagami uzgodnionymi dla waszego kontekstu - Lista wniosków uszeregowana według istotności (Krytyczne / Wysokie / Średnie / Niskie) ze sposobem odtworzenia - Rekomendowana poprawka do każdego wniosku, z fragmentami kodu i linkami do wskazanych linii - Uszeregowane wpisy w systemie zgłoszeń (Linear, Jira, GitHub Issues - do wyboru) - Podsumowanie dla zarządu z najważniejszymi problemami i planem działań naprawczych ### Jak pracujemy 1. **1. Ustalenie zakresu i próbkowanie** (Tydzień 1): Definiujemy zakres audytu: które repozytorium, które okno PR-ów, jakie narzędzia wytworzyły kod. Przechodzimy przez jawne kryteria i ustalamy wagi dla waszego kontekstu. 2. **2. Ewaluacja** (Tydzień 2-3): Stosujemy 4-wymiarowe kryteria do próbki kodu. Odtwarzamy wnioski, oceniamy każdy wymiar w jawnej skali, opisujemy istotność i rekomendowaną poprawkę. 3. **3. Raport i omówienie** (Tydzień 3-4): Dostarczamy pisemny raport, prowadzimy kierownictwo i inżynierię przez najważniejsze wnioski, zakładamy uszeregowane wpisy w systemie zgłoszeń. 4. **4. Ponowny przegląd po naprawie (opcjonalnie)** (Tydzień +4-8): Po 4-8 tygodniach ponownie sprawdzamy naprawione wnioski. Wydajemy certyfikat naprawy, jeśli standard jest spełniony. ### FAQ **Q: Ile kosztuje typowy audyt?** Stała cena, zależna od rozmiaru repozytorium i głębokości audytu. Większość projektów mieści się w przedziale 15-60 tys. EUR. Podajemy wąski zakres na rozmowa wstępna, gdy zakres jest jasny. **Q: Jakie są wasze kryteria oceny?** Cztery wymiary, każdy z jawną definicją. Poprawność: czy kod robi to, co mówi ticket i testy, łącznie z przypadkami brzegowymi i ścieżkami błędu. Bezpieczeństwo: czy wprowadza podatności, niebezpieczne ustawienia domyślne, złe obchodzenie się z sekretami albo ryzyko w zależnościach. Utrzymywalność: czy następny inżynier zrozumie kod i zmieni właściwe miejsce bez czytania całego systemu. Jakość testów: czy pokrycie jest realne, czyli czy testy padają, gdy zachowanie się psuje. Każdy wymiar ma nazwane podkryteria z wagami dostrojonymi do waszej bazy kodu. **Q: Jak oceniacie każdy wymiar?** W jawnej skali 0-4 ze stałym zakotwiczeniem: 0 defekt blokujący, wydanie musi się zatrzymać; 1 poważne braki wymagające przerobienia przed wydaniem; 2 działa, ale z istotnymi wnioskami i długiem do zaplanowania; 3 spełnia standard, tylko drobne wnioski; 4 brak wniosków w tym wymiarze. Każda ocena w raporcie wskazuje kotwicę, na podstawie której została przyznana, więc można podważyć pojedynczą ocenę, a nie cały raport. Pełne kryteria, podkryteria i skala są opublikowane na tej stronie, nie zmieniają się między projektami, a klienci stosują je do własnych PR-ów między audytami. **Q: Możecie pracować pod NDA dla branż regulowanych?** Tak. Standardowe NDA, możemy działać na izolowanej infrastrukturze, a przy szczególnie wrażliwych projektach pracować na miejscu w biurze klienta. Uwzględniamy frameworki branżowe (fintech KNF/DORA, ochrona zdrowia w kontekście GDPR/HIPAA). **Q: Czy to konkurencja dla naszego wewnętrznego przeglądu kodu?** Nie. Jesteśmy zewnętrznym potwierdzeniem jakości - okresowym, niezależnym, głębszym niż codzienny przegląd. Większość klientów korzysta z nas kwartalnie, obok wewnętrznego procesu. **Q: Możecie ponownie zaudytować po naprawie?** Tak. Ponowne audyty po naprawie wyceniamy na połowę pierwotnej ceny i kończymy certyfikatem naprawy potwierdzającym zamknięcie wniosków. Przydatne do raportowania przed zarządem albo przygotowania do zewnętrznego audytu. ### Powiązane - https://dfzoo.ai/pl/uslugi/evaluate-audit-secure/security-review - https://dfzoo.ai/pl/uslugi/evaluate-audit-secure/production-readiness - https://dfzoo.ai/pl/uslugi/operate-measure-maintain/llm-observability --- ## Security Review - Oceniaj i Zabezpiecz **URL:** https://dfzoo.ai/pl/uslugi/evaluate-audit-secure/security-review **Service type:** Przegląd bezpieczeństwa kodu aplikacji **Anchor:** Oceniaj, Audytuj i Zabezpiecz (Oceniaj i Zabezpiecz) > Przegląd bezpieczeństwa kodu, który wasz zespół napisał z asystentami kodowania AI - podatności na wstrzyknięcia, niebezpieczne ustawienia domyślne, obsługa secrets, wygenerowana logika autoryzacji - plus powierzchnia, którą na poziomie aplikacji dodały wasze funkcje AI. Dostajecie wnioski z odtworzeniem i konkretnymi poprawkami w kodzie. ### Podsumowanie dfzoo AI Institute sprawdza bezpieczeństwo kodu powstającego z asystentami kodowania AI oraz funkcji AI, które ten kod dodaje do produktu. Czytamy fragmenty repozytorium tworzone z udziałem AI pod kątem podatności na wstrzyknięcia, niebezpiecznych ustawień domyślnych, obsługi secrets, zależności wybranych przez asystenta, wygenerowanej logiki autoryzacji i obsługi błędów, która ujawnia wnętrze systemu. Następnie przeglądamy powierzchnię AI na poziomie aplikacji: jak aplikacja buduje prompty, co wysyła do API LLM i co przyjmuje z powrotem jako zaufane, jak dane od użytkownika trafiają do promptu, co wolno wywołać agentowi wbudowanemu w produkt i jak obsługiwana jest odpowiedź modelu, zanim dotrze do użytkownika, przeglądarki albo bazy danych. Każdy wniosek dostajecie ze sposobem odtworzenia i konkretną poprawką w kodzie, a do tego zasady przeglądu, które zespół stosuje później sam, żeby ta sama klasa błędu zatrzymywała się już na code review. To praca na poziomie kodu: bezpieczeństwo infrastruktury, chmury i organizacji jest świadomie poza zakresem. ### Dla kogo to jest - Zespoły produktowe, które wdrożyły asystenta kodowania AI i wypuszczają kod szybciej, niż ktokolwiek sprawdza go pod kątem bezpieczeństwa - Software house'y i agencje dostarczające kod tworzony z udziałem AI, od których klienci zaczęli wymagać niezależnej opinii o jego bezpieczeństwie - Firmy, które dodały funkcję opartą na LLM (czat, streszczanie, agenta w produkcie) do istniejącego produktu i nigdy nie sprawdziły, co przez to odsłoniły - Szefowie inżynierii, CTO zespołów produktowych i tech leadzi, którzy potrzebują wniosków w repozytorium, a nie w dokumencie o nadzorze nad AI ### Problemy, które rozwiązujemy - Asystent szybko napisał budowanie zapytań, obsługę plików i parsowanie żądań, a nikt nie sprawdził tych ścieżek pod kątem wstrzyknięć - Wygenerowany kod niesie domyślne ustawienia frameworka, które model akurat znał: zbyt szerokie CORS, wyłączona weryfikacja, komunikaty debugowania na produkcji - Secrets i tokeny są obsługiwane tak, jak pokazał je asystent - w kodzie, w logach, w paczce wysyłanej do przeglądarki - Sprawdzanie uprawnień powstawało osobno przy każdym endpoincie i nikt nie zweryfikował, czy jest spójne, ani czy w ogóle jest - Dane od użytkownika trafiają do promptu bez filtrowania, a odpowiedź modelu idzie prosto do HTML, wywołania powłoki albo zapisu w bazie - Agent wbudowany w produkt ma dostęp do narzędzi i poświadczeń znacznie szerszy niż zadanie, które faktycznie wykonuje ### Co dostajecie - Lista wniosków uszeregowana według istotności (Krytyczne / Wysokie / Średnie / Niskie), każdy ze sposobem odtworzenia oraz wskazaniem pliku i linii - Konkretna poprawka do każdego wniosku: patch, diff albo gotowy przykład, który wasi inżynierowie stosują bezpośrednio - Przegląd powierzchni funkcji AI: budowanie promptów, granice zaufania wokół API LLM, obsługa odpowiedzi modelu, zakres narzędzi i uprawnień agenta - Wnioski dotyczące secrets i zależności w kodzie tworzonym z udziałem AI, łącznie z pakietami, które wprowadził asystent - Krótki zestaw zasad przeglądu bezpieczeństwa dla zespołu - na co patrzeć w PR-ach pisanych z AI, w formie checklisty oraz reguł lintera lub CI tam, gdzie da się to zautomatyzować - Sesja robocza z zespołem inżynierskim: najważniejsze wnioski i to, skąd każdy z nich się wziął ### Jak pracujemy 1. **1. Ustalenie zakresu i omówienie kodu** (Tydzień 1): Ustalamy, które repozytoria i który kod tworzony z udziałem AI wchodzą w zakres, jaki asystent go wytworzył i gdzie siedzą funkcje AI. Przechodzimy przez bazę kodu razem z waszym starszym inżynierem. 2. **2. Przegląd bezpieczeństwa kodu** (Tydzień 1-2): Czytamy kod tworzony z udziałem AI pod kątem wstrzyknięć, niebezpiecznych ustawień domyślnych, obsługi secrets, wygenerowanej logiki autoryzacji, wyboru zależności i obsługi błędów ujawniającej wnętrze systemu. Odtwarzamy to, co znajdziemy. 3. **3. Przegląd powierzchni funkcji AI** (Tydzień 2-3): Śledzimy, jak dane od użytkownika trafiają do promptu, co aplikacja wysyła do API LLM i co przyjmuje z powrotem jako zaufane, co wolno wywołać agentowi w produkcie i jak obsługiwana jest odpowiedź, zanim dotrze do użytkownika albo do bazy. 4. **4. Raport, poprawki i zasady przeglądu** (Tydzień 3): Dostarczamy listę wniosków z poprawkami, omawiamy z zespołem inżynierskim najważniejsze problemy i przekazujemy zasady przeglądu oraz reguły w CI, które wyłapią te same klasy błędów w kolejnych PR-ach. ### FAQ **Q: Czym to się różni od AI Code Evaluation?** AI Code Evaluation ocenia kod tworzony z udziałem AI w 4 wymiarach - poprawność, bezpieczeństwo, utrzymywalność, jakość testów - i daje całościowe potwierdzenie jakości. Security Review bierze wyłącznie wymiar bezpieczeństwa i schodzi znacznie głębiej: doprowadzamy każdy problem do działającego odtworzenia i poprawki w kodzie, a dodatkowo przeglądamy powierzchnię funkcji AI, którą kryteria ewaluacji jedynie próbkują. Zespoły, które chcą jednego szerokiego obrazu, kupują ewaluację; zespoły, które już wiedzą, że problemem jest bezpieczeństwo, kupują tę usługę. **Q: Czego wprost nie robicie?** Nie zajmujemy się bezpieczeństwem infrastruktury, chmury ani sieci, testami penetracyjnymi działających systemów, ofensywnym red teamingiem wdrożonej AI, programami audytu SOC 2 czy ISO 27001 ani nadzorem nad bezpieczeństwem w organizacji. Nie oceniamy też waszego dostawcy LLM ani łańcucha dostaw modelu. Przeglądamy kod i powierzchnię, którą ten kod tworzy na poziomie aplikacji. Jeśli potrzebujecie któregoś z powyższych, weźcie do tego wyspecjalizowany zespół - chętnie przekażemy nasze wnioski, żeby zaczynał od czegoś konkretnego. **Q: Jakiego dostępu potrzebujecie?** Dostęp do odczytu repozytoriów w zakresie, historia PR-ów z okresu pracy z AI oraz działająca instancja na środowisku nieprodukcyjnym, żebyśmy mogli odtwarzać wnioski. Bez dostępu do produkcji, bez danych klientów. Po waszej stronie starszy inżynier: na omówienie kodu na starcie i na przejście przez wnioski na końcu. Całość pod NDA. **Q: Ile to trwa?** Większość projektów zajmuje 2-3 tygodnie od ustalenia zakresu do raportu. Pojedynczy serwis albo dobrze zamknięta funkcja AI mieści się w tygodniu; duża baza kodu w wielu repozytoriach jest dzielona na kilka przebiegów zamiast jednego długiego przeglądu. Stała cena, ustalana po ustaleniu zakresu. **Q: Pracujecie zdalnie czy na miejscu?** Domyślnie zdalnie, pod NDA, na dostępie do waszego repozytorium. Praca na miejscu w Polsce jest możliwa tam, gdzie polityka klienta wymaga, żeby kod nie opuszczał jego sieci - w tej samej cenie. **Q: Co zostaje zespołowi po projekcie?** Lista wniosków z poprawkami, zasady przeglądu w formie checklisty i to, co dało się zamienić na reguły lintera albo CI. Chodzi o to, żeby kolejny PR napisany z asystentem wyłapał wasz własny code review, a nie my. Zespoły wracają zwykle po kilku miesiącach po krótszy powtórny przegląd, a nie po powtórzenie całego projektu. **Q: Jakie języki i technologie obsługujecie?** TypeScript i JavaScript, Python, Go, Ruby, Java i Kotlin oraz popularne frameworki webowe wokół nich. Mobile i embedded są poza zakresem tej usługi. Powiedzcie nam, na czym pracujecie, na rozmowa wstępna, a wprost odpowiemy, czy jesteśmy właściwym zespołem. **Q: Czy musimy mieć wdrożoną funkcję AI, żeby to kupić?** Nie. Sporo projektów dotyczy zwykłego produktu, w którym jedyną AI był asystent pomagający napisać kod. Jeśli produkt nie ma funkcji opartej na LLM, pomijamy fazę 3 i przeznaczamy ten czas na przegląd kodu. ### Powiązane - https://dfzoo.ai/pl/uslugi/evaluate-audit-secure/ai-code-evaluation - https://dfzoo.ai/pl/uslugi/evaluate-audit-secure/production-readiness - https://dfzoo.ai/pl/uslugi/train-adopt-govern/tools-implementation --- ## Pokrycie testowe - Oceniaj i Zabezpiecz **URL:** https://dfzoo.ai/pl/uslugi/evaluate-audit-secure/test-coverage **Service type:** Doradztwo jakości oprogramowania **Anchor:** Oceniaj, Audytuj i Zabezpiecz (Oceniaj i Zabezpiecz) > Strategia i wdrożenie pokrycia testowego dla projektów wspieranych przez AI - co testować (i czego narzędzia AI nie przetestują za was), co zaślepić, co pominąć i jak utrzymać regresje pod kontrolą, gdy AI bez przerwy generuje kod. ### Podsumowanie dfzoo AI Institute projektuje i wdraża strategie pokrycia testowego dla zespołów inżynierskich wdrażających kod tworzony z udziałem AI. Narzędzia kodowania AI generują testy, które wyglądają na obszerne, ale często przepuszczają regresje i sztucznie podbijają metryki pokrycia bez realnej ochrony zachowania systemu. Audytujemy istniejące testy, projektujemy strategię pokrycia dopasowaną do ryzyka produktu (co nigdy nie może się zepsuć kontra co może się zepsuć i zostać naprawione później), wdrażamy testy ścieżek krytycznych, których AI nie generuje dobrze, i ustawiamy sygnały w CI, które wcześnie wychwytują prawdziwe regresje. ### Dla kogo to jest - Menedżerowie inżynierscy, których kod tworzony z udziałem AI powstaje szybciej, niż nadążają testy - Tech leadowie obserwujący wzrost metryk pokrycia testowego równolegle do wzrostu regresji - CTO próbujący skalować mały zespół z AI bez utraty jakości - Liderzy QA dostosowujący strategię do kodu generowanego przez AI ### Problemy, które rozwiązujemy - AI generuje testy, które przechodzą, ale nie chronią krytycznych dla biznesu zachowań - Metryki pokrycia rosną, ale rośnie też liczba regresji zgłaszanych przez klientów - Zaślepki (mocki) mnożą się, gdy AI łata luki - zestaw testów sprawdza sam siebie, nie system - Testy integracyjne ścieżek krytycznych są zbyt żmudne dla AI - nikt ich nie pisze ### Co dostajecie - Audyt istniejącego zestawu testów - co chroni zachowanie systemu, a co tylko podbija metryki - Dokument strategii testów dopasowany do ryzyka (co testować, co zaślepić, co pominąć) - Wdrożenie testów ścieżek krytycznych, z którymi AI sobie nie radzi - Konfiguracja CI pod sygnały awarii (wykrywanie testów niestabilnych, alerty regresji) - Przewodnik dla zespołu: wzorce promptów do generowania dobrych testów z AI ### Jak pracujemy 1. **1. Audyt testów** (Tydzień 1): Próbkujemy istniejące testy w modułach. Oceniamy każdy pod kątem realnej ochrony zachowania kontra podbijanie metryk. Identyfikujemy luki na ścieżkach krytycznych. 2. **2. Projektowanie strategii** (Tydzień 2): Dopasowujemy strategię testów do ryzyka produktu. Ustalamy proporcje testów jednostkowych/integracyjnych/e2e. Decydujemy, co zaślepić, a co uruchamiać na prawdę. 3. **3. Wdrożenie** (Tydzień 3-4): Piszemy testy ścieżek krytycznych, których AI nie wychwytuje. Ustawiamy sygnały w CI. Dokumentujemy wzorce promptów do testów generowanych przez AI. 4. **4. Przekazanie zespołowi** (Tydzień 4-5): Prowadzimy zespół przez strategię. Programujemy w parach przy pisaniu nowych testów według przewodnika. Ustalamy cykl przeglądów. ### FAQ **Q: Dlaczego AI słabo radzi sobie z pisaniem niektórych testów?** AI świetnie radzi sobie z testami jednostkowymi sprawdzającymi widoczne sygnatury funkcji, gorzej z testami integracyjnymi zależnymi od stanu systemu, a słabo z testami end-to-end, gdzie koszt poprawnego przygotowania danych testowych przewyższa to, co mieści się w kontekście. Efekt: zestaw testów przechylony w stronę łatwych przypadków, słaby tam, gdzie naprawdę jest ryzyko. **Q: Podniesiecie nam metrykę pokrycia?** Optymalizujemy pod ochronę zachowania systemu, nie pod surowy procent pokrycia. Większość klientów widzi lekki spadek pokrycia (usuwamy testy, które niczego nie chronią) i większy spadek regresji zgłaszanych przez klientów. **Q: Pracujecie z naszym istniejącym CI?** Tak. GitHub Actions, GitLab CI, CircleCI, Jenkins, Buildkite - pracujemy z tym, co macie. Nie wprowadzamy nowych narzędzi CI bez mocnego powodu. **Q: Jakie języki obejmujecie?** TypeScript / JavaScript, Python, Go, Ruby, Java/Kotlin. Mobile (Swift, Kotlin) i systemy wbudowane są poza zakresem tej praktyki. **Q: Jak to się ma do AI Code Evaluation?** AI Code Evaluation audytuje kod pod kątem poprawności; Pokrycie testowe buduje system, który wychwytuje nowe regresje, zanim trafią na produkcję. Większość klientów kupuje oba w parze. ### Powiązane - https://dfzoo.ai/pl/uslugi/evaluate-audit-secure/ai-code-evaluation - https://dfzoo.ai/pl/uslugi/evaluate-audit-secure/production-readiness - https://dfzoo.ai/pl/uslugi/build-automate-modernize/ai-powered-custom-development --- ## Production Readiness - Oceniaj i Zabezpiecz **URL:** https://dfzoo.ai/pl/uslugi/evaluate-audit-secure/production-readiness **Service type:** Doradztwo gotowości produkcyjnej **Anchor:** Oceniaj, Audytuj i Zabezpiecz (Oceniaj i Zabezpiecz) > Ocena gotowości produkcyjnej przed wydaniem dla systemów tworzonych z udziałem AI - observability, runbooki, ścieżki wycofania zmian (rollback) i sygnały dyżuru (on-call) sprawdzane wobec ustrukturyzowanych kryteriów, zanim system trafi na żywo. ### Podsumowanie dfzoo AI Institute prowadzi przeglądy gotowości produkcyjnej dla systemów tworzonych z udziałem AI, zanim trafią do klientów. Oceniamy system wobec ustrukturyzowanych kryteriów - pokrycie observability i tracingu, jakość runbooków, ścieżki wycofania zmian, alerty dyżuru, zapas mocy, monitoring kosztów, tryby awarii zależności - i wydajemy pisemną certyfikację z wymaganymi działaniami naprawczymi. To właśnie to, czego SRE i zespół platformowy potrzebują do podpisu, i to, czym kierownictwo uzasadnia wydanie podczas analizy po incydencie (post-mortem), gdy coś pójdzie nie tak. ### Dla kogo to jest - Dyrektorzy inżynierii zatwierdzający premierę nowej funkcji AI - Szefowie SRE / zespołów platformowych proszeni o objęcie systemu AI dyżurem - CTO firm w fazie wzrostu uruchamiający swój pierwszy produkt oparty na LLM - PM-owie, których premiera zależy od zgody inżynierii i SRE ### Problemy, które rozwiązujemy - Funkcje AI ruszają bez observability - debugowanie na produkcji to zgadywanie - Runbooków dla nowego systemu nie ma albo są niesprawdzone - Ścieżka wycofania zmian nigdy nie była ćwiczona - produkcyjne pożary zmieniają się w wielogodzinne przestoje - Brak monitoringu kosztów - pierwszy miesięczny rachunek za LLM to szok ### Co dostajecie - Karta oceny gotowości produkcyjnej (10-15 wymiarów, Zaliczone / Częściowe / Niezaliczone) - Pisemne oświadczenie certyfikacyjne z wymaganymi działaniami naprawczymi (jeśli potrzebne) - Szablony runbooków dla nowego systemu, sprawdzone z zespołem dyżurnym - Dokumentacja procedury wycofania zmian, przećwiczona na środowisku staging - Rekomendacje konfiguracji monitoringu kosztów i zapasu mocy ### Jak pracujemy 1. **1. Przegląd systemu** (Tydzień 1): Przeglądamy architekturę, stos observability, runbooki, procedury dyżurowe, plan przepustowości, konfigurację monitoringu kosztów. 2. **2. Walidacja praktyczna** (Tydzień 2): Testujemy ścieżkę wycofania zmian na środowisku staging. Ćwiczymy runbooki z zespołem dyżurnym. Wyzwalamy testowe alerty, żeby sprawdzić, czy uruchamiają się poprawnie. 3. **3. Ocena i lista działań naprawczych** (Tydzień 2-3): Oceniamy wobec kryteriów. Spisujemy wymagane działania naprawcze z priorytetami. Przygotowujemy szkic oświadczenia certyfikacyjnego. 4. **4. Ponowny przegląd (jeśli potrzebna naprawa)** (Tydzień +1-2): Ponownie sprawdzamy naprawione wymiary. Wydajemy finalną certyfikację. ### FAQ **Q: Możecie scertyfikować system w 2 tygodnie?** Tak, dla systemów z rozsądną observability już na miejscu. Systemy bez fundamentów (brak tracingu, brak runbooków) często potrzebują 4-6 tygodniowego okna na działania naprawcze przed certyfikacją. **Q: Co obejmują kryteria?** Pokrycie observability, głębokość tracingu, jakość runbooków, procedura wycofania zmian, alerty dyżuru, zapas mocy, monitoring kosztów, tryby awarii zależności, podstawy bezpieczeństwa, ryzyko specyficzne dla AI (prompt injection, dryf modelu), gotowość do reagowania na incydenty. **Q: Piszecie runbooki za nas?** Dostarczamy szablony i przeglądamy to, co napisze zespół. Nie piszemy runbooków, których zespół sam nie zweryfikuje - runbook, którego dyżurni nie ćwiczyli, nikogo nie chroni. **Q: Możecie certyfikować na bieżąco, nie tylko przy premierze?** Tak. Kwartalna recertyfikacja to częsty model stałej współpracy. Przydatny, gdy funkcje AI szybko się zmieniają, albo jako sygnał na poziomie zarządu dla systemów krytycznych dla produkcji. **Q: Czym to się różni od audytu bezpieczeństwa?** Audyt bezpieczeństwa pyta „czy atakujący to złamie”. Production Readiness pyta „czy to wytrzyma obciążenie i czy zespół naprawi szybko, gdy się zepsuje”. Większość premier potrzebuje obu przed wejściem na żywo. ### Powiązane - https://dfzoo.ai/pl/uslugi/evaluate-audit-secure/ai-code-evaluation - https://dfzoo.ai/pl/uslugi/evaluate-audit-secure/security-review - https://dfzoo.ai/pl/uslugi/operate-measure-maintain/llm-observability --- ## AI Quality Evaluation - Oceniaj i Zabezpiecz **URL:** https://dfzoo.ai/pl/uslugi/evaluate-audit-secure/ai-quality-evaluation **Service type:** Ocena jakości odpowiedzi i procesów AI **Anchor:** Oceniaj, Audytuj i Zabezpiecz (Oceniaj i Zabezpiecz) > Niezależna ocena AI, które już działa w waszym produkcie albo na zapleczu - jakość odpowiedzi, oparcie w źródłach, skuteczność agentów i to, czy krok z AI naprawdę skraca proces - punktowana według kryteriów, które zostają u was. ### Podsumowanie dfzoo AI Institute ocenia jakość rozwiązań AI, które już działają: funkcji opartych na LLM w produkcie, odpowiedzi opartych na wyszukiwaniu w waszych dokumentach (RAG) i procesów agentowych działających bez nadzoru człowieka. Budujemy zestaw testowy z waszych realnych przypadków, ustalamy kryteria oceny z osobami odpowiedzialnymi za efekt biznesowy, punktujemy każdy przypadek i dostarczamy pisemny raport ze zmierzonym punktem odniesienia, wzorcami błędów stojącymi za wynikiem i rekomendowaną poprawką do każdego wzorca. Jednostką pracy jest jedna ocena - jeden przypadek testowy przepuszczony przez kryteria - więc zakres i faktura opisują to samo. Zestaw testowy i kryteria zostają u was, wpięte w CI jako bramka jakości, która wychwyci regresję przy kolejnej zmianie modelu, promptu albo dostawcy. Jesteśmy neutralni narzędziowo: pracujemy na tym, co już macie, albo pomagamy wybrać, a przy dużym wolumenie pokazujemy, kiedy self-hosting wychodzi taniej niż SaaS. ### Dla kogo to jest - Zespoły produktowe, które wdrożyły funkcję opartą na LLM i nie mają jak sprawdzić, czy odpowiedzi są dobre - Firmy z wewnętrznym asystentem AI, w którym adopcja stoi w miejscu i nikt nie mierzy, czy komukolwiek pomaga - Zespoły przed zmianą modelu, frameworka promptów albo dostawcy, obawiające się cichej regresji - Osoby odpowiedzialne za operacje i centra usług wspólnych, które wstawiły krok z AI do procesu i muszą udowodnić, że proces się skrócił, a nie że praca przeszła gdzie indziej - Szefowie inżynierii, danych i wsparcia, którzy uruchomili darmowe narzędzie ewaluacyjne, mają już trace'y i utknęli na zestawie testowym ### Problemy, które rozwiązujemy - Funkcja oparta na LLM działa na produkcji, a jedynym sygnałem o jakości jest reklamacja albo zgłoszenie do wsparcia - Odpowiedzi RAG brzmią wiarygodnie, ale nikt nie sprawdził, czy są oparte na właściwym dokumencie źródłowym - Procesy agentowe raportują sukces, podczas gdy człowiek po cichu kończy robotę, i nikt nie liczy, jak często - Zmiana promptu albo modelu trafia na produkcję, a o regresji zespół dowiaduje się od użytkowników - Zespół wziął darmowe narzędzie ewaluacyjne i utknął na dwóch rzeczach: zbudowaniu zestawu testowego i ustaleniu, co znaczy 'dobrze' ### Co dostajecie - Zestaw testowy zbudowany z waszego realnego ruchu i przypadków brzegowych, w waszym narzędziu, po projekcie należący do was - Kryteria oceny z ważonymi wymiarami (poprawność, oparcie w źródłach, skuteczność wykonania zadania, format i ton, koszt i czas odpowiedzi), uzgodnione z właścicielem biznesowym, nie tylko z inżynierią - Punkt odniesienia: każdy przypadek oceniony, wyniki w rozbiciu na wymiary, ścieżki użytkownika oraz wersje modelu i promptu - Raport wzorców błędów ze sposobem odtworzenia, przyczyną źródłową i rekomendowaną poprawką do każdego wzorca - Bramka jakości w CI: zestaw testowy jako blokujący warunek przy zmianach promptu, modelu, wyszukiwania i dostawcy - Zestawienie skuteczności procesowej pokazujące, gdzie krok z AI skraca proces, a gdzie przenosi pracę na człowieka, plus rekomendacja narzędziowa z porównaniem kosztu self-hostingu i SaaS przy waszym wolumenie ### Jak pracujemy 1. **1. Ustalenie zakresu i kryteriów** (Tydzień 1): Ustalamy, co znaczy 'dobrze' dla tego rozwiązania, z osobami odpowiedzialnymi za efekt. Definiujemy wymiary jakości, ich wagi i próg zaliczenia. Ustalamy liczbę ocen w zakresie. 2. **2. Budowa zestawu testowego** (Tydzień 1-2): Składamy przypadki testowe z realnego ruchu, znanych błędów i przypadków brzegowych, których nikt nie testuje. Dodajemy oczekiwaną odpowiedź albo kryterium oceny do każdego przypadku. Wpinamy zestaw w wasze narzędzie ewaluacyjne albo w takie, które pomagamy wybrać. 3. **3. Ocena i analiza** (Tydzień 2-3): Punktujemy każdy przypadek według kryteriów: automatycznie, z użyciem modelu jako sędziego i z przeglądem ludzkim tam, gdzie kryterium wymaga osądu. Grupujemy błędy we wzorce i wiążemy każdy wzorzec z przyczyną: wyszukiwanie, prompt, dobór modelu, wpięcie narzędzi albo projekt procesu. 4. **4. Raport i przekazanie** (Tydzień 3-4): Dostarczamy punkt odniesienia i raport wzorców błędów, omawiamy je z właścicielami produktu i inżynierią, przekazujemy zestaw testowy i kryteria wpięte w CI jako bramkę jakości, którą wasz zespół uruchamia bez nas. 5. **5. Ewaluacja ciągła (opcjonalnie)** (Na bieżąco): Retainer: utrzymujemy i rozszerzamy zestaw testowy, uruchamiamy go po każdej zmianie modelu, promptu albo dostawcy i raportujemy regresję, zanim znajdą ją użytkownicy. ### FAQ **Q: Co dokładnie wchodzi w zakres i jak to wyceniacie?** Jednostką pracy jest jedna ocena: jeden przypadek testowy przepuszczony przez uzgodnione kryteria. W zakresie ustalamy liczbę ocen, kryteria i sposób raportowania, więc widzicie, za co płacicie. Stała cena za projekt, zależna od liczby ocen i liczby ocenianych powierzchni AI. Wąski przedział podaję na rozmowa wstępna, gdy zakres jest jasny. **Q: Ile to trwa i czego potrzebujecie od nas?** Trzy do czterech tygodni od startu do przekazania. Od was potrzebujemy jednej osoby odpowiedzialnej za efekt biznesowy do uzgodnienia kryteriów, jednego inżyniera do dostępów i narzędzi, próbki realnego ruchu albo logów oraz znanych przypadków błędnych, jeśli już je zbieracie. Warsztat kryteriów i sesja przekazania zajmują po około pół dnia, reszta idzie po naszej stronie. **Q: Czym to się różni od AI Code Evaluation?** AI Code Evaluation audytuje kod: poprawność, bezpieczeństwo, utrzymywalność i jakość testów tego, co napisał wasz zespół razem z narzędziami AI. AI Quality Evaluation ocenia zachowanie wdrożonego rozwiązania: czy odpowiedzi są trafne i oparte na właściwym źródle, czy agent kończy zadanie, czy proces faktycznie się skrócił. Bezbłędny kod potrafi dawać system, który odpowiada źle, i odwrotnie. Klienci często kupują obie usługi, zwykle w tej kolejności. **Q: Potrzebujecie dostępu do naszych danych produkcyjnych?** Niekoniecznie. Możemy pracować na zanonimizowanych albo syntetycznych przypadkach wyprowadzonych z waszych realnych i w branżach regulowanych często tak właśnie robimy. Jeśli dostajemy dostęp do produkcji, pracujemy pod NDA, na waszej infrastrukturze, a dane zostają w waszym środowisku. Standardowe NDA, a przy szczególnie wrażliwych projektach możemy działać na izolowanej infrastrukturze. **Q: Z jakimi narzędziami ewaluacyjnymi pracujecie?** Jesteśmy neutralni narzędziowo i pracujemy na tym, co już macie: Braintrust, Langfuse, LangSmith, Confident AI, Promptfoo, Arize Phoenix, W&B Weave albo zwykły zestaw testów w waszym repozytorium. Jeśli nie macie nic, pomagamy wybrać, a przy dużym wolumenie ocen pokazujemy, kiedy self-hosting rozwiązania otwartoźródłowego kosztuje mniej niż SaaS. Zestaw testowy piszemy tak, żeby dało się go przenieść między narzędziami, bo wiązanie punktu odniesienia jakości z jednym dostawcą jest ryzykiem na rynku, który wciąż się konsoliduje. **Q: Próbowaliśmy już darmowego narzędzia do ewaluacji. Co dokładacie?** Każda platforma ewaluacyjna ma darmowy poziom, więc większość klientów przychodzi po własnej próbie. Narzędzie rzadko jest miejscem, w którym zespoły się zatrzymują. Zatrzymują się na dwóch rzeczach, których narzędzie nie daje: na zestawie testowym odzwierciedlającym to, o co realnie pytają użytkownicy, i na spisanej definicji 'dobrze', pod którą podpisują się właściciel biznesowy i inżynieria. Zaczynamy dokładnie tam, a narzędzie, które już wybraliście, pracuje dalej pod spodem. **Q: Oceniacie też procesy agentowe, nie tylko pojedyncze odpowiedzi?** Tak i zwykle to tam są najciekawsze liczby. Mierzymy skuteczność wykonania zadania (czy agent doprowadził sprawę do końca), częstość przejęcia przez człowieka (jak często ktoś po cichu kończy za niego), błędy na poziomie kroków w sesjach, turach, wywołaniach narzędzi i subagentach oraz efekt procesowy: ile czasu krok z AI zabiera z procesu, a ile przenosi na kogoś innego. **Q: Co dzieje się po zakończeniu projektu i kto utrzymuje zestaw testowy?** Zestaw testowy, kryteria i bramka w CI należą do was, są w waszym repozytorium i waszym narzędziu, a wasz zespół uruchamia je i rozszerza bez nas. Jeśli wolicie nie brać na siebie utrzymania, prowadzimy je w ramach stałej opieki: aktualizujemy zestaw, uruchamiamy go po każdej zmianie modelu, promptu albo dostawcy i raportujemy regresję wobec punktu odniesienia. Ta opieka łączy się z LLM Observability, które pilnuje tych samych sygnałów jakości na żywym ruchu. ### Powiązane - https://dfzoo.ai/pl/uslugi/evaluate-audit-secure/ai-code-evaluation - https://dfzoo.ai/pl/uslugi/operate-measure-maintain/llm-observability - https://dfzoo.ai/pl/uslugi/evaluate-audit-secure/production-readiness --- # Buduj, Automatyzuj i Modernizuj (anchor 03) **URL:** https://dfzoo.ai/pl/uslugi/build-automate-modernize **Highlight:** AI-Powered Custom Development - filar firmy > AI-Powered Custom Development, automatyzacja procesów i modernizacja systemów legacy - projektowane pod produkcję, nie pod demo. ### Podsumowanie dfzoo AI Institute buduje, automatyzuje i modernizuje systemy, używając AI jako głównej dźwigni inżynierskiej. Nasza praktyka AI-Powered Custom Development to filar firmy - wdrażamy produkcyjne aplikacje, agentowe backendy i wewnętrzne narzędzia wzbogacone o AI dla zespołów inżynierskich potrzebujących produktowego tempa. Oceniamy gotowość na agentic AI dla organizacji badających autonomiczne procesy; automatyzujemy procesy biznesowe łączące RPA i agenty LLM; modernizujemy systemy legacy, gdzie refactoring wspierany przez AI odblokowuje zmianę inaczej nieosiągalną. Każdy projekt z tym samym zapleczem jakości co praktyka ewaluacyjna. ### Dla kogo to jest - Zespoły product engineering wdrażające produkty AI-native albo wzbogacone o AI - Liderzy operacji automatyzujący wysokowolumenowe procesy łączące dane ustrukturyzowane i nieustrukturyzowane - CTO osadzeni na systemach legacy blokujących postęp planu produktowego - Założyciele prototypujący agentowe backendy, potrzebujący seniorskiej oceny inżynierskiej ### Problemy, które rozwiązujemy - Prototypy AI zbudowane przez zespół nie przeżywają przejścia na produkcję - Ręczne procesy back-office pochłaniają godziny, które AI + RPA mogłyby skrócić do minut - System legacy blokuje plan produktowy, ale pełne przepisanie jest poza zasięgiem - Agentowe procesy wyglądają obiecująco w demo, ale nikt nie wie, jak wygląda produkcja ### Praktyki pod tym anchorem - AI-Powered Custom Development - https://dfzoo.ai/pl/uslugi/build-automate-modernize/ai-powered-custom-development - Agentic AI Readiness - https://dfzoo.ai/pl/uslugi/build-automate-modernize/agentic-ai-readiness - Automatyzacja procesów - https://dfzoo.ai/pl/uslugi/build-automate-modernize/process-automation - Modernizacja legacy - https://dfzoo.ai/pl/uslugi/build-automate-modernize/legacy-modernization ### FAQ **Q: Czym AI-Powered Custom Development różni się od zwykłego custom dev?** Nasz proces inżynierski używa AI jako głównej dźwigni - kodowanie wspierane przez AI, generowanie testów przez AI, przegląd kodu wzbogacony o AI. Efekt ten sam: produkcyjna aplikacja. Tempo i struktura kosztów - inne. **Q: Budujecie agenty od początku do końca?** Tak. Projektujemy architekturę agenta (single-shot, multi-step, multi-agent), budujemy wspierającą infrastrukturę (pamięć, retrieval, wywoływanie narzędzi, observability), integrujemy z istniejącymi systemami i wdrażamy na produkcję z runbookami dyżurowymi. **Q: Jakie procesy automatyzujecie?** Procesy oparte na dokumentach i na skrzynkach mailowych: przyjmowanie leadów, obsługę faktur, segregację zgłoszeń wsparcia, procesy kontroli zgodności, przygotowywanie odpowiedzi na RFP. Łączymy narzędzia RPA (UiPath, Power Automate) z agentami LLM tam, gdzie pojawiają się dane nieustrukturyzowane. **Q: Możecie zmodernizować system legacy na miejscu, bez pełnego przepisania?** Często tak. Oceniamy system, wyznaczamy granice metodą strangler-fig i wymieniamy komponenty stopniowo, używając refactoringu wspieranego przez AI. Przeszkoda w planie znika bez wielokwartalnego przepisywania. **Q: W jakich technologiach pracujecie?** TypeScript / Node / Next.js, Python / FastAPI, Go po stronie backendu; React, Vue po stronie frontendu; Postgres, ClickHouse, Pinecone, pgvector dla danych; AWS, GCP, Vercel, Fly dla infrastruktury. Większość dostawców LLM (OpenAI, Anthropic, Google, modele open-source przez Bedrock albo Together). --- ## AI-Powered Custom Development - Buduj i Modernizuj **URL:** https://dfzoo.ai/pl/uslugi/build-automate-modernize/ai-powered-custom-development **Service type:** Custom software development **Anchor:** Buduj, Automatyzuj i Modernizuj (Buduj i Modernizuj) > Pełne product engineering z AI jako główną dźwignią - produkcyjne aplikacje, agentowe backendy i wewnętrzne narzędzia wzbogacone o AI, budowane w etapowanym pipelinie, w którym każda zmiana przechodzi bramki automatyczne i konkretnego człowieka, zanim trafi na produkcję. ### Podsumowanie dfzoo AI Institute buduje produkcyjne oprogramowanie, używając AI jako głównej dźwigni inżynierskiej. Dostarczamy aplikacje webowe, agentowe backendy, wewnętrzne narzędzia wzbogacone o AI i produkty AI-native zespołom potrzebującym seniorskiej oceny inżynierskiej w tempie, które dziś umożliwia AI. Każda zmiana idzie tym samym pipelinem: ticket, specyfikacja, plan, implementacja, bramki automatyczne (typecheck, lint, build, code review), PR z preview, przegląd człowieka, scalenie, weryfikacja na produkcji. Przyspieszamy dostarczanie i utrzymujemy stabilność, i mierzymy oba: przepustowość i defekty wykryte po scaleniu. Z certyfikatem ISO 9001:2015. Filar firmy - większość projektów w dfzoo AI Institute dotyka tej praktyki. ### Dla kogo to jest - Zespoły product engineering wdrażające produkty AI-native albo wzbogacone o AI - Założyciele prototypujący nowy produkt AI-first, potrzebujący seniorskiej inżynierii - Firmy w fazie wzrostu wymieniające albo rozbudowujące krytyczne narzędzie wewnętrzne - Scale-upy, których plan produktowy przekracza możliwości własnej inżynierii ### Problemy, które rozwiązujemy - Prototypy AI zbudowane przez zespół nie przeżywają przejścia na produkcję - Prototyp sklecony z promptów zatrzymuje się na demo: bez testów, bez śladu przeglądu, bez runbooka i bez właściciela na produkcji - my oddajemy system z tymi czterema - Własna inżynieria jest obłożona; plan produktowy ciągle się ślizga - Agentowe procesy dobrze wypadają w demo, ale zespół nie ma wzorca produkcyjnego - Wyceny custom development z agencji nie uwzględniają tempa, jakie daje AI ### Co dostajecie - Działająca produkcyjna aplikacja (web, agentowy backend, narzędzie wewnętrzne) - Kod źródłowy z dokumentacją, testami, konfiguracją wdrożenia, runbookami - Pipeline dostarczania przekazany razem z kodem: bramki w CI, szablon PR, checklista przeglądu, środowiska preview - Infrastructure-as-code dla produkcji, środowiska staging i środowisk preview - Konfiguracja observability i dyżuru (tracing, alerty, runbooki) - Przekazanie wiedzy waszemu zespołowi - sesje w parach, omówienia, nagrane przekazanie ### Jak pracujemy 1. **1. Rozpoznanie i zakres** (Tydzień 0-1): rozmowa wstępna, definicja problemu, ustalenie zakresu technicznego, propozycja fazy w stałej cenie. 2. **2. Ticket, specyfikacja, plan** (Przy każdej zmianie): Każda zmiana zaczyna się od ticketu, potem powstaje pisemna specyfikacja i plan. Bramka: zakres i plan akceptujecie wy, zanim powstanie pierwsza linia kodu. 3. **3. Implementacja** (Przy każdej zmianie): Agenty AI piszą kod pod okiem seniorskiego inżyniera, który prowadzi, poprawia i odrzuca. Bramka: o tym, co wychodzi z gałęzi, decyduje ten inżynier, nigdy agent, który kod napisał. 4. **4. Bramki automatyczne** (Przy każdym commicie): Na każdej gałęzi uruchamiamy typecheck, lint, build i automatyczny code review. Bramka: decyduje maszyna i nie da się jej obejść, a próg jest ten sam dla kodu AI i człowieka. 5. **5. PR, preview i przegląd człowieka** (Przy każdej zmianie): Do każdego PR-a powstaje środowisko preview z działającą zmianą, a pełny diff czyta drugi seniorski inżynier. Bramka: o scaleniu albo zwrocie decyduje człowiek, a ślad przeglądu zostaje w repozytorium. 6. **6. Scalenie i weryfikacja na produkcji** (Przy każdym wydaniu): Po wdrożeniu sprawdzamy zmianę na tracingu, alertach i runbooku. Bramka: zmiana jest skończona, gdy działa na produkcji, a nie w chwili scalenia. 7. **7. Utwardzanie, premiera i przekazanie** (Ostatnie 2-4 tygodnie): Przegląd gotowości produkcyjnej, testy obciążeniowe, kompletność observability, szkolenie dyżuru, sesje w parach i udokumentowane przekazanie. Opcjonalnie utrzymanie w ramach App Maintenance. ### FAQ **Q: Jak szybko możecie zacząć?** Od rozmowa wstępna do pierwszego sprintu zwykle 2-3 tygodnie. Szybciej w pilnych projektach (1 tydzień), gdy zakres jest jasny i mamy dostępność. **Q: Czy tempo, które daje AI, kosztuje nas stabilność?** To wymiana, na którą się nie zgadzamy, i powód, dla którego istnieje ten pipeline. Mierzymy obok siebie przepustowość dostarczania i defekty wykryte po scaleniu, a wzrost defektów traktujemy jako powód do zacieśnienia bramek, nie jako cenę prędkości. **Q: Stała cena czy rozliczenie godzinowe (T&M)?** Stała cena za fazę. Faza 1 (rozpoznanie i zakres) to zwykle 10-25 tys. EUR zależnie od złożoności. Kolejne fazy budowy wyceniamy osobno, na podstawie tego, czego się nauczyliśmy. **Q: A własność intelektualna (IP)?** Kod należy do was. Zachowujemy prawa do wielokrotnego użytku wewnętrznych bibliotek i wzorców napisanych przed projektem. Standardowe wzajemne NDA na sam projekt. **Q: Pracujemy z tymi samymi inżynierami przez całość?** Tak. Zespoły projektowe są stabilne od startu do przekazania. Nie wymieniamy inżynierów w trakcie projektu, chyba że o to poprosicie. **Q: Robicie konkretnie inżynierię AI, nie tylko zwykły custom dev?** Tak. Agentowe backendy (single-shot, multi-step, multi-agent), systemy RAG, API oparte na LLM, pipeline'y ewaluacji AI, infrastruktura embeddingów i wyszukiwania wektorowego. Każde z tych rozwiązań działa dziś na produkcji u klientów. ### Powiązane - https://dfzoo.ai/pl/uslugi/build-automate-modernize/agentic-ai-readiness - https://dfzoo.ai/pl/uslugi/build-automate-modernize/legacy-modernization - https://dfzoo.ai/pl/uslugi/operate-measure-maintain/app-maintenance --- ## Agentic AI Readiness - Buduj i Modernizuj **URL:** https://dfzoo.ai/pl/uslugi/build-automate-modernize/agentic-ai-readiness **Service type:** Doradztwo agentic AI **Anchor:** Buduj, Automatyzuj i Modernizuj (Buduj i Modernizuj) > Ocena i działający pilotaż dla organizacji rozważających agentowe procesy AI - pisemne rekomendacje, co zlecić agentom, co utrzymać z człowiekiem w pętli, plus pilotaż we wzorcu produkcyjnym, który zespół rozbuduje. ### Podsumowanie dfzoo AI Institute ocenia gotowość na agentic AI dla organizacji ważących autonomiczne procesy. Odwzorowujemy wasze obecne procesy na wzorce agentowe (narzędzia single-shot, agenty multi-step, systemy multi-agent), rekomendujemy, co zlecić agentom, a co utrzymać z człowiekiem w pętli, i dostarczamy działający pilotaż na waszych danych i systemach. Otrzymujecie: pisemną ocenę gotowości plus wdrożony pilotaż - wzorzec produkcyjny, który zespół rozbudowuje, a nie demo na poziomie badawczym. ### Dla kogo to jest - CTO i szefowie produktu ważący inwestycję w agentic AI - Liderzy operacji rozważający autonomiczne procesy w obsłudze klienta, sales ops albo finansach - Zespoły innowacji proszone przez kierownictwo o pilotaż agentic AI - Liderzy inżynierscy, których zespoły są podekscytowane agentami, ale potrzebują produkcyjnej weryfikacji założeń ### Problemy, które rozwiązujemy - Dema agentowe wyglądają jak magia; nikt nie potrafi nazwać wzorca produkcyjnego - Zespoły nie zgadzają się, czy dany proces ma być agentowy, czy prowadzony przez człowieka - Pilotaże powstają jako prototypy badawcze i nigdy nie stają się systemami produkcyjnymi - Oferty dostawców platform agentowych trudno porównać z budową własnego rozwiązania ### Co dostajecie - Pisemna ocena gotowości (10-20 stron): procesy, wzorce agentowe, rekomendacje - Działający pilotaż wdrożony na staging albo ograniczoną produkcję (jeden proces) - Dokumentacja architektury pilotażu: przepływ agenta, narzędzia, pamięć, observability, eskalacja - Model kosztowy skalowania wzorca pilotażu na inne procesy - Podsumowanie dla zarządu z rekomendacjami i wariantami planu ### Jak pracujemy 1. **1. Rozpoznanie procesów** (Tydzień 1-2): Odwzorowujemy obecne procesy. Identyfikujemy kandydatów do zlecenia agentom. Oceniamy każdy pod kątem dopasowania i ryzyka. 2. **2. Projektowanie wzorca** (Tydzień 2-3): Dobieramy właściwy wzorzec dla każdego kandydata (single-shot, multi-step, multi-agent). Projektujemy pilotaż. 3. **3. Budowa pilotażu** (Tydzień 3-7): Budujemy i wdrażamy pilotaż. Uruchamiamy observability, ewaluację, ścieżki eskalacji. 4. **4. Ocena i plan** (Tydzień 7-10): Prowadzimy pilotaż przez 2-4 tygodnie. Mierzymy efekty. Dostarczamy ocenę i plan dla kolejnych procesów. ### FAQ **Q: Jaka jest różnica między agentem a procesem?** Proces wykonuje deterministyczne kroki po kolei. Agent decyduje o kolejnym kroku na podstawie kontekstu, w razie potrzeby wywołując narzędzia i powtarzając pętlę. Wiele problemów lepiej rozwiązać zwykłym procesem; agentów rekomendujemy tylko wtedy, gdy dynamiczne podejmowanie decyzji uzasadnia swój koszt. **Q: Używacie konkretnego frameworka agentowego?** Pracujemy z LangGraph, OpenAI Agents SDK, Anthropic Claude Agent SDK oraz wzorcami pisanymi od zera. Dobieramy pod problem; nie jesteśmy przywiązani do jednego frameworka. **Q: Możecie ocenić bez budowania pilotażu?** Tak. Projekty wyłącznie oceniające trwają 3-4 tygodnie i dostarczają dokument gotowości oraz rekomendacje. Większość klientów dodaje pilotaż, bo rekomendacje z oceny trzeba przetestować przed skalowaniem. **Q: Jak mierzycie, czy agent działa?** Pipeline'y eval (golden test sets, porównania równoległe), produkcyjna observability (wgląd na poziomie tracingu w decyzje agenta), KPI biznesowe (przepustowość procesu, wskaźnik eskalacji, jakość po stronie klienta). Wszystkie trzy są ważne; ustawiamy je razem. **Q: Co się dzieje, gdy agent popełni błąd na produkcji?** Architektura pilotażu zawiera ścieżki eskalacji: decyzje o niskiej pewności trafiają do ludzi, błędy wyzwalają alerty, a wycofanie zmian (rollback) to jedna zmiana konfiguracji. Produkcyjne agenty potrzebują tej samej dyscypliny operacyjnej co każdy produkcyjny system - instalujemy ją w ramach pilotażu. ### Powiązane - https://dfzoo.ai/pl/uslugi/build-automate-modernize/ai-powered-custom-development - https://dfzoo.ai/pl/uslugi/build-automate-modernize/process-automation - https://dfzoo.ai/pl/uslugi/operate-measure-maintain/llm-observability --- ## Automatyzacja procesów - Buduj i Modernizuj **URL:** https://dfzoo.ai/pl/uslugi/build-automate-modernize/process-automation **Service type:** Automatyzacja procesów biznesowych **Anchor:** Buduj, Automatyzuj i Modernizuj (Buduj i Modernizuj) > Automatyzacja wysokowolumenowych procesów biznesowych łączących dane ustrukturyzowane i nieustrukturyzowane - łączymy narzędzia RPA z agentami LLM tam, gdzie strona nieustrukturyzowana wygrywa z tradycyjnym RPA opartym na regułach. ### Podsumowanie dfzoo AI Institute automatyzuje wysokowolumenowe procesy biznesowe rozpięte na wiele systemów. Łączymy narzędzia RPA (UiPath, Power Automate, n8n) dla przepływów ustrukturyzowanych z agentami LLM dla danych nieustrukturyzowanych (maile, dokumenty, transkrypty). Typowe projekty: przyjmowanie leadów, obsługa faktur, segregacja zgłoszeń wsparcia, procesy kontroli zgodności, przygotowywanie odpowiedzi na RFP. Wdrażamy działającą automatyzację, mierzymy przepustowość i jakość, przekazujemy z dokumentacją utrzymania. ### Dla kogo to jest - Szefowie operacji, których zespoły tracą godziny na procesach opartych na dokumentach - Liderzy finansowi automatyzujący przyjmowanie faktur, procesy AP / AR albo kontrole zgodności - Liderzy obsługi klienta segregujący duży wolumen zgłoszeń z maila i czatu - Liderzy sales operations obsługujący leady przychodzące z wielu źródeł ### Problemy, które rozwiązujemy - Tradycyjne RPA pęka na mailach, PDF-ach i transkryptach - to 30% danych nieustrukturyzowanych - Prompty oparte na ChatGPT działają w demo, padają przy wolumenie produkcyjnym - Istniejące automatyzacje psują się co tydzień, bo systemy źródłowe ciągle się zmieniają - Przepustowość staje, bo każdy wyjątek wymaga człowieka; nikt nie mierzy częstotliwości ### Co dostajecie - Działająca automatyzacja wdrożona na produkcję z monitoringiem - Obsługa wyjątków i eskalacja do człowieka w pętli dla przypadków o niskiej pewności - Dashboard metryk jakości (przepustowość, dokładność, wskaźnik wyjątków, koszt na transakcję) - Przewodnik utrzymania na wypadek zmian w systemach źródłowych albo formatach danych - Opcjonalnie: miesięczny przegląd kosztów i jakości w ramach stałej współpracy App Maintenance ### Jak pracujemy 1. **1. Mapowanie procesu** (Tydzień 1-2): Towarzyszymy zespołowi prowadzącemu ręczny proces. Mapujemy dokładne kroki, decyzje i ścieżki wyjątków. 2. **2. Architektura i dobór narzędzi** (Tydzień 2-3): Decydujemy RPA / LLM / hybryda dla każdego kroku. Wybieramy konkretne narzędzia. Projektujemy ścieżki eskalacji. 3. **3. Budowa i integracja** (Tydzień 3-7): Budujemy automatyzację. Integrujemy z mailem, CRM, ERP i czym tylko trzeba. 4. **4. Stopniowe uruchamianie i przekazanie** (Tydzień 7-10): Zwiększamy obsługiwany wolumen z 10% do 100% przez 2-4 tygodnie. Szkolimy zespół operacyjny z monitoringu. Dokumentujemy utrzymanie. ### FAQ **Q: Jak decydujecie: RPA czy agent LLM dla danego kroku?** Dane ustrukturyzowane o stabilnych schematach → RPA. Dane nieustrukturyzowane albo nieostre decyzje → LLM ze strukturalnym wynikiem. Większość rzeczywistych procesów jest hybrydowa; wybieramy osobno dla każdego kroku, żeby trzymać koszt i tryby awarii w ryzach. **Q: Jaka jest typowa dokładność wobec człowieka?** Zależy całkowicie od procesu i jakości danych. Wyodrębnianie pozycji z faktur osiąga ponad 95% na dobrze sformatowanych PDF-ach. Otwarta klasyfikacja może lądować na 80-90% z pipeline'ami eval dostrojonymi do waszej taksonomii. Cel ustalamy na rozmowie o zakresie. **Q: Możecie zautomatyzować coś, co już częściowo zbudowaliśmy?** Tak. Wiele projektów przejmuje istniejące skrypty, kruche boty RPA albo prompty oparte na ChatGPT i przebudowuje je pod produkcyjną niezawodność. **Q: A koszt?** Dwa wymiary kosztu: koszt budowy (stała cena za fazę) i koszt działania (wydatek na API LLM na transakcję). Koszt działania jest częścią decyzji architektonicznej: wybieramy najtańszy model i najmniejszy prompt, który trafia w cel jakościowy. **Q: Czy to wymaga długoterminowego utrzymania?** Tak. Systemy źródłowe się zmieniają, formaty danych dryfują, wzorce wyjątków ewoluują. Oferujemy utrzymanie w ramach stałej współpracy App Maintenance - zwykle lekkie miesięczne spotkania kontrolne z kwartalnymi przeglądami i drobnymi korektami. ### Powiązane - https://dfzoo.ai/pl/uslugi/build-automate-modernize/agentic-ai-readiness - https://dfzoo.ai/pl/uslugi/operate-measure-maintain/app-maintenance - https://dfzoo.ai/pl/uslugi/operate-measure-maintain/cost-model-optimization --- ## Modernizacja legacy - Buduj i Modernizuj **URL:** https://dfzoo.ai/pl/uslugi/build-automate-modernize/legacy-modernization **Service type:** Modernizacja systemów legacy **Anchor:** Buduj, Automatyzuj i Modernizuj (Buduj i Modernizuj) > Refactoring wspierany przez AI i zmiana platformy dla systemów legacy blokujących wasz plan produktowy - stopniowa modernizacja metodą strangler-fig zamiast wielokwartalnego przepisywania. ### Podsumowanie dfzoo AI Institute modernizuje systemy legacy blokujące plany produktowe. Oceniamy system, wyznaczamy granice metodą strangler-fig i wymieniamy komponenty stopniowo, używając refactoringu wspieranego przez AI, za ułamek kosztu pełnego przepisania. Typowe projekty: monolity PHP/Perl/starsze .NET, legacy Java EE, frontendy z czasów jQuery, nietypowane bazy kodu JavaScript / legacy Python. Otrzymujecie: zmodernizowany system działający obok legacy, dopóki tego ostatniego nie da się bezpiecznie wycofać. ### Dla kogo to jest - CTO, których plan produktowy blokuje system, którego zespół boi się dotknąć - Liderzy inżynierscy stojący przed decyzją „przepisać czy żyć z tym” - Firmy w fazie wzrostu, których pierwsza produktowa baza kodu nie pasuje już do zespołu, który ją utrzymuje - Firmy przejmujące, integrujące podmiot o znacząco starszym stosie technologicznym ### Problemy, które rozwiązujemy - System legacy nie ma testów; każda zmiana zagraża produkcji - Dokumentacji nie ma albo jest nieaktualna; wiedza firmowa siedzi w głowach jednego-dwóch inżynierów - Pełne przepisanie jest zbyt drogie do sfinansowania; życie z legacy kosztuje więcej co kwartał - Rekrutacja jest trudna, bo nowi inżynierowie nie chcą utrzymywać starego stosu ### Co dostajecie - Ocena systemu legacy (architektura, zależności, punkty zapalne ryzyka, warianty modernizacji) - Plan modernizacji metodą strangler-fig z fazami w ustalonej kolejności i budżetem na fazę - Zmodernizowane komponenty działające na produkcji obok legacy - Pokrycie testowe i dokumentacja dla granic między częścią zmodernizowaną a legacy - Przekazanie wiedzy waszemu zespołowi - sesje w parach, pisemne dokumenty architektury ### Jak pracujemy 1. **1. Ocena** (Tydzień 1-3): Czytamy bazę kodu pod NDA. Rozmawiamy z inżynierami, którzy ją utrzymują. Uruchamiamy analizę statyczną. Mapujemy zależności. 2. **2. Projekt planu** (Tydzień 3-4): Wyznaczamy granice metodą strangler-fig. Ustalamy kolejność faz według ryzyka i odblokowywania planu. Budżetujemy każdą fazę. 3. **3. Sprinty modernizacji** (Tydzień 4-N): Wymieniamy komponenty faza po fazie. Każda faza trafia na produkcję za warstwą routingu uruchamiającą stare i nowe równolegle. 4. **4. Wycofanie legacy** (Po każdej fazie): Gdy komponent jest w pełni zastąpiony i stabilny, wycofujemy starą ścieżkę. Powtarzamy, aż legacy zniknie. ### FAQ **Q: Czy naprawdę możecie modernizować stopniowo, a nie jako jednorazowe przepisanie?** Tak - strangler-fig to standardowy wzorzec. Budujemy warstwę routingu uruchamiającą stare i nowe równolegle, wymieniamy jeden komponent naraz i wycofujemy stary kod dopiero, gdy nowa ścieżka jest stabilna na produkcji. Jednorazowe przepisania zostawiamy dla przypadków, w których strangler-fig naprawdę nie jest możliwy. **Q: Co właściwie robi refactoring wspierany przez AI?** Narzędzia AI przyspieszają żmudne części: czytanie nieznanego kodu, generowanie testów dla kodu, który ich nie ma, tłumaczenie między idiomami (z callbacków na async/await), proponowanie refaktoryzacji, które przeglądają ludzie. Ocena inżynierska zostaje przy naszym zespole; przyspiesza objętość pracy. **Q: Jakie stosy legacy modernizowaliście?** PHP (Symfony 1.x, własne legacy), starszy .NET (Framework 4.x ASP.NET WebForms / MVC), Java EE / Spring Boot sprzed 2.x, frontendy jQuery renderowane po stronie serwera, nietypowane bazy kodu Node.js, legacy Python 2 / wczesny 3. **Q: Co, jeśli nasza baza kodu nie ma żadnych testów?** Częsty stan wyjściowy. Faza 1 modernizacji zwykle obejmuje zbudowanie rusztowania testów wokół komponentów, które będziemy wymieniać - generowanych z pomocą AI, sprawdzanych przez naszych inżynierów. Testy przed refaktoryzacją, zawsze. **Q: Czy możemy trzymać system legacy w ruchu podczas waszej modernizacji?** Tak - o to właśnie chodzi w strangler-fig. Legacy nadal obsługuje produkcyjny ruch, podczas gdy nowe komponenty wchodzą do gry za warstwą routingu. Przełączenie każdego komponentu to zmiana konfiguracji, nie wydanie. ### Powiązane - https://dfzoo.ai/pl/uslugi/build-automate-modernize/ai-powered-custom-development - https://dfzoo.ai/pl/uslugi/evaluate-audit-secure/test-coverage - https://dfzoo.ai/pl/uslugi/operate-measure-maintain/app-maintenance --- # Operuj, Mierz i Utrzymuj (anchor 04) **URL:** https://dfzoo.ai/pl/uslugi/operate-measure-maintain **Highlight:** Telemetria i observability LLM > Observability, telemetria, optymalizacja kosztów i bieżące utrzymanie dla systemów AI już na produkcji - żeby dalej działały, gdy wasze dane i sposób użycia się zmieniają. ### Podsumowanie dfzoo AI Institute prowadzi AI na produkcji dla zespołów, które potrzebują, by systemy dalej działały, gdy dane, sposób użycia i dostawcy się zmieniają. Wdrażamy observability LLM (tracing, pipeline'y eval, wykrywanie dryfu); projektujemy telemetrię i analitykę dla produktów AI (co mierzyć, na co alertować); prowadzimy audyty kosztów LLM i optymalizacji modeli, które odzyskują zwykle 20-60% miesięcznego wydatku na LLM; dostarczamy bieżące utrzymanie aplikacji tworzonych z udziałem AI. Nasza praktyka optymalizacji kosztów opiera się na konkretnej matematyce dostawców - caching, routing, kompresja promptów - nie na hasłach z prezentacji sprzedażowych. ### Dla kogo to jest - Zespoły inżynierskie prowadzące funkcje LLM na produkcji z rosnącymi niespodziankami kosztowymi - Liderzy produktu, którzy nie wiedzą, czy ich funkcja AI faktycznie działa dla użytkowników - CFO i partnerzy finansowi pytający, dlaczego rachunek za LLM dalej rośnie - Zespoły DevOps i SRE dodające usługi oparte na LLM do dyżurów ### Problemy, które rozwiązujemy - Wydatek na LLM rośnie szybciej niż ruch i nikt nie wie dlaczego - Brak sygnału, czy zmiany promptów uczyniły produkt lepszym czy gorszym - Dostawca LLM wycofał model, a zespół nie ma planu migracji - Skarg klientów na jakość AI nie da się powiązać z konkretnym zachowaniem na produkcji ### Praktyki pod tym anchorem - LLM Observability - https://dfzoo.ai/pl/uslugi/operate-measure-maintain/llm-observability - Telemetria i Analytics - https://dfzoo.ai/pl/uslugi/operate-measure-maintain/telemetry-analytics - Optymalizacja kosztów i modeli - https://dfzoo.ai/pl/uslugi/operate-measure-maintain/cost-model-optimization - Utrzymanie aplikacji - https://dfzoo.ai/pl/uslugi/operate-measure-maintain/app-maintenance ### FAQ **Q: Co odzyskuje audyt kosztów LLM?** Typowi klienci odzyskują 20-60% miesięcznego wydatku na LLM w ciągu 4-8 tygodni. Oszczędności biorą się z poprawy wskaźnika trafień w cache promptów, kierowania mniej istotnych wywołań do mniejszych modeli, kompresji promptów i usunięcia zbędnych ponowień (retry). **Q: Jakiego stosu observability używacie?** Opartego na OpenTelemetry tam, gdzie to możliwe. Integrujemy z Datadog, Honeycomb, Grafana, Posthog oraz dedykowanymi narzędziami observability LLM (Langfuse, Arize, Helicone), zależnie od tego, co już macie. **Q: Jak mierzycie, czy zmiany promptów poprawiły produkt?** Pipeline'y eval: golden test sets, porównania równoległe wobec punktu wyjścia, alerty regresji na metrykach, które mają znaczenie. Projektujemy kryteria oceny pod wasz konkretny produkt. **Q: Robicie migracje modeli, gdy dostawca wycofa model?** Tak. Migracje u dostawcy (np. Claude 4.6 → 4.7, GPT-4 → GPT-4o) to standardowy projekt utrzymaniowy: ustalamy na nowo punkt wyjścia dla eval, uruchamiamy testy regresji i migrujemy prompty z mierzonym porównaniem jakości i opóźnień przed i po. **Q: Utrzymanie to stała współpraca czy jednorazowa umowa?** Oba modele działają. Stała współpraca pasuje zespołom z ciągle rozwijanymi funkcjami AI na produkcji. Jednorazowe umowy (migracja dostawcy, audyt kosztów, konfiguracja eval) pasują zespołom z konkretną potrzebą. --- ## LLM Observability - Operuj i Mierz **URL:** https://dfzoo.ai/pl/uslugi/operate-measure-maintain/llm-observability **Service type:** Observability i ewaluacja LLM **Anchor:** Operuj, Mierz i Utrzymuj (Operuj i Mierz) > Tracing, pipeline'y eval i wykrywanie dryfu dla systemów LLM i agentowych na produkcji - sesje, tury, kroki, wywołania narzędzi i subagenty, z kosztem i jakością przypiętymi do każdego z nich. ### Podsumowanie dfzoo AI Institute wdraża observability LLM dla systemów produkcyjnych. Zakładamy tracing na bytach, które system agentowy naprawdę ma - sesjach, turach, krokach, wywołaniach narzędzi i subagentach, a nie na pojedynczych wywołaniach modelu - budujemy pipeline'y eval oceniające nowe wersje promptów i modeli wobec waszych zestawów testowych, ustawiamy wykrywanie dryfu i integrujemy sygnały z waszym stosem observability. Zestaw testowy powstały w audycie zostaje u was jako bramka w CI i monitor produkcyjny, a my utrzymujemy go i rozszerzamy po każdej zmianie modelu lub promptu. Oparte na OpenTelemetry tam, gdzie to możliwe; pracujemy z Datadog, Honeycomb, Grafana, Langfuse, Arize, Helicone, Braintrust, LangSmith, W&B Weave, Confident AI i Galileo, pomagamy wybrać, a przy dużym wolumenie pokazujemy, kiedy self-hosting wychodzi taniej niż SaaS. ### Dla kogo to jest - Zespoły inżynierskie, których funkcje LLM trafiają na produkcję na ślepo - bez wglądu w to, co się dzieje - Zespoły prowadzące wielokrokowe agenty, w których awaria siedzi trzy wywołania narzędzi głębiej i nikt jej nie znajduje - Liderzy produktu, którzy nie wiedzą, czy kolejne wersje promptów poprawiły czy pogorszyły produkt - Zespoły SRE dodające usługi LLM do dyżurów - Liderzy ML/AI engineering potrzebujący rozwoju promptów opartego na ewaluacji ### Problemy, które rozwiązujemy - Przebieg agenta pada i nikt nie potrafi odtworzyć, który krok, które wywołanie narzędzia albo który subagent go wywrócił - Zmiany promptów i modeli trafiają na produkcję bez pomiarów; regresje jakości są niewidoczne - Audyt raz przygotował dobry zestaw testowy, który zestarzał się w tygodniu zmiany modelu - Opóźnienie P99 podwaja się z dnia na dzień i nikt nie wskazuje przyczyny - Dostawca zwraca złą odpowiedź; brak logu żądania do debugowania ### Co dostajecie - Tracing na bytach agentowych - sesje, tury, kroki, wywołania narzędzi, subagenty - z promptem, odpowiedzią, opóźnieniem, kosztem i wersją przy każdym z nich - Pipeline eval z waszymi zestawami testowymi i alertami regresji zintegrowanymi z CI - Wasz zestaw testowy z audytu przekazany jako bramka scalania w CI i monitor produkcyjny na tych samych progach, z procedurą rozszerzania po każdej zmianie modelu lub promptu - Wykrywanie dryfu na wynikach (dryf semantyczny, dryf opóźnień, dryf kosztów) - Dashboardy: jakość per funkcja, koszt per funkcja, zachowanie per użytkownik - Runbooki dla zespołu dyżurnego: jak zdebugować incydent LLM w 15 minut ### Jak pracujemy 1. **1. Rozpoznanie stosu** (Tydzień 1): Mapujemy obecne wywołania LLM i agentów w waszej bazie kodu. Inwentaryzujemy funkcje, prompty, dostawców, istniejące zestawy testowe i narzędzia observability. 2. **2. Instrumentacja** (Tydzień 2-3): Obejmujemy tracingiem sesje, tury, kroki, wywołania narzędzi i subagenty. Przekazujemy dane do wybranego backendu i sprawdzamy, czy każdy przebieg widać od początku do końca. 3. **3. Pipeline'y eval i bramka w CI** (Tydzień 3-4): Budujemy zestawy testowe dla każdej funkcji, z realnego ruchu i z wniosków wcześniejszej ewaluacji. Wpinamy eval do CI jako bramkę scalania i powielamy te same progi jako monitor produkcyjny. 4. **4. Szkolenie dyżuru i przekazanie** (Tydzień 4-5): Prowadzimy zespół dyżurny przez dashboardy, runbooki i scenariusze incydentów. Przekazujemy zestaw testowy jako wasz zasób, w waszym repozytorium. 5. **5. Utrzymanie w abonamencie (opcjonalnie)** (Stale): Rozszerzamy zestaw testowy i uruchamiamy regresję po każdej zmianie modelu lub promptu, a potem raportujemy, co się zmieniło i dlaczego. ### FAQ **Q: Jaki backend observability albo eval rekomendujecie?** Zależy od tego, co już macie. Jeśli korzystacie z Datadog albo Honeycomb, rozszerzamy je. Jeśli wolicie rozwiązania pod LLM: Langfuse, Arize, Helicone, Braintrust, LangSmith, W&B Weave, Confident AI, Galileo. Jesteśmy niezależni od narzędzi, pomagamy wybrać pod wasz wolumen i granice danych, a powyżej pewnej liczby żądań liczymy, kiedy self-hosting (Langfuse, Phoenix) wychodzi taniej niż SaaS. **Q: Co, jeśli wybrany dostawca eval zostanie przejęty, wygaszony albo zmieni kierunek?** Na tym rynku to się zdarza i dlatego nie budujemy waszej kompetencji na jednym dostawcy. Zestawy testowe, definicje ocen i trace'y trzymamy w przenośnej formie w waszym repozytorium, a instrumentację prowadzimy przez OpenTelemetry tam, gdzie backend na to pozwala. Platformy się konsolidują; kompetencja i wasze zestawy testowe zostają u was, a migracja to tydzień pracy, a nie budowa od zera. **Q: Możecie obserwować wywołania LLM bez zmiany kodu aplikacji?** Częściowo - dostawcy tacy jak OpenAI i Anthropic udostępniają logi pojedynczych żądań przez swoje konsole. Dla produkcyjnej observability (kontekst tracingu, przypisanie do użytkownika, przypisanie kosztu, struktura kroków agenta) zwykle opakowujemy wywołania cienką warstwą SDK. **Q: Czym jest pipeline eval?** Zestaw testów dla wyników LLM i agentów. Zestaw testowy definiuje, jak wygląda poprawna odpowiedź; przebieg eval ocenia nowe wersje promptów i modeli wobec tego zestawu; CI blokuje scalenie, gdy jakość spada, a te same progi działają jako monitor na produkcji. Standardowa praktyka dla każdego zespołu traktującego prompty jak kod produkcyjny. **Q: Obsługujecie prywatność danych w tracingu?** Tak. Tracing można skonfigurować tak, żeby usuwał dane osobowe (PII), zanim logi opuszczą waszą infrastrukturę, maskował fragmenty promptów albo trzymał pełne prompty tylko na izolowanej infrastrukturze. Politykę usuwania danych projektujemy z waszym zespołem security. **Q: Ile to kosztuje miesięcznie?** Koszt backendu tracingu zależy od wolumenu - zwykle 200-2000 EUR/mies. dla średnich obciążeń LLM. Self-hostowany Langfuse jest darmowy po stronie dostawcy, ale wymaga obsługi operacyjnej. Oba warianty modelujemy w fazie rozpoznania. ### Powiązane - https://dfzoo.ai/pl/uslugi/operate-measure-maintain/telemetry-analytics - https://dfzoo.ai/pl/uslugi/operate-measure-maintain/cost-model-optimization - https://dfzoo.ai/pl/uslugi/evaluate-audit-secure/production-readiness --- ## Telemetria i Analytics - Operuj i Mierz **URL:** https://dfzoo.ai/pl/uslugi/operate-measure-maintain/telemetry-analytics **Service type:** Telemetria produktowa **Anchor:** Operuj, Mierz i Utrzymuj (Operuj i Mierz) > Telemetria produktowa i analityka dla produktów tworzonych z udziałem AI - co mierzyć, na co alertować, co pokazać na dashboardzie, żeby produkt, inżynieria i kierownictwo patrzyli na ten sam obraz. ### Podsumowanie dfzoo AI Institute projektuje i wdraża telemetrię produktową dla produktów tworzonych z udziałem AI. Tłumaczymy pytania produktowe („czy funkcja AI jest faktycznie używana?”, „czy użytkownicy rezygnują na kroku z AI?”, „która wersja promptu napędza przychód?”) na zdarzenia, dashboardy i alerty. Otrzymujecie: taksonomię zdarzeń, instrumentację na webie/mobile/backendzie, dashboardy dla produktu/inżynierii/kierownictwa oraz plan, co mierzyć dalej. ### Dla kogo to jest - Szefowie produktu, którzy nie wiedzą, czy ich funkcja AI pomaga użytkownikom - Zespoły growth prowadzące testy A/B na wersjach promptów albo wdrożeniach funkcji AI - Menedżerowie inżynierscy wpinający sygnały z danych w decyzje produktowe - Założyciele uruchamiający produkty AI-first bez własnego zespołu analityki ### Problemy, które rozwiązujemy - Produkt wdraża funkcje AI na ślepo - bez pojęcia, co użytkownicy z nimi faktycznie robią - Istniejąca analityka śledzi odsłony stron, ale pomija momenty specyficzne dla AI (edycja promptu, akceptacja/odrzucenie, ponowne generowanie) - Każdy zespół ma własny dashboard z innymi liczbami; kierownictwo nie ufa żadnemu - Alerty zalewają rzeczami, które nikogo nie obchodzą; prawdziwe problemy przechodzą niezauważone ### Co dostajecie - Taksonomia zdarzeń (momenty specyficzne dla AI: prompt wysłany, wynik wyrenderowany, akceptacja, odrzucenie, ponowne generowanie itp.) - Instrumentacja na webie, mobile i backendzie ze śledzeniem wersji - Trzy dashboardy: produkt (kondycja funkcji), inżynieria (kondycja systemu), kierownictwo (KPI biznesowe) - Polityka alertów: co alarmuje, kto dostaje powiadomienie, co mówi runbook - Kwartalny cykl przeglądu, żeby rozwijać to, co mierzymy, w miarę rozwoju produktu ### Jak pracujemy 1. **1. Rozpoznanie pytań** (Tydzień 1): Rozmawiamy z produktem, inżynierią i kierownictwem. Spisujemy pytania, na które każda rola potrzebuje odpowiedzi. Ustalamy priorytety. 2. **2. Taksonomia zdarzeń** (Tydzień 2): Tłumaczymy pytania na zdarzenia. Definiujemy momenty specyficzne dla AI. Dokumentujemy właściwości i politykę wersjonowania. 3. **3. Instrumentacja i dashboardy** (Tydzień 2-4): Instrumentujemy stos. Budujemy trzy dashboardy. Ustawiamy politykę alertów. 4. **4. Przekazanie i cykl przeglądu** (Tydzień 4-5 (potem kwartalnie)): Prowadzimy każdą grupę odbiorców przez jej dashboard. Ustawiamy kwartalny przegląd. Dopracowujemy listę zdarzeń na kolejne 3 miesiące. ### FAQ **Q: Z jakimi platformami analitycznymi pracujecie?** Posthog, Mixpanel, Amplitude, Segment jako router, Snowplow dla wdrożeń self-hosted. Dla sygnałów specyficznych dla LLM często łączymy je z Langfuse albo Helicone. Wybieramy pod wasz stos, nie pod preferencje dostawców. **Q: Robicie hurtownie danych?** W ograniczonym zakresie. Wysyłamy zdarzenia analityczne do waszej hurtowni (Snowflake, BigQuery, Redshift, ClickHouse), ale nie budujemy pełnych modeli hurtowni - to oddzielny projekt. **Q: Czym to się różni od LLM Observability?** LLM Observability dotyczy samego wywołania LLM (opóźnienie, koszt, dryf wyniku). Telemetria i Analytics dotyczy użytkownika i produktu (czy funkcja działa, czy użytkownicy wracają, czy napędza przychód). Większość klientów kupuje oba. **Q: Możecie instrumentować też aplikacje mobilne?** Tak. iOS (Swift) i Android (Kotlin), plus React Native i Flutter. Respektujemy wymogi prywatności specyficzne dla platform (ATT na iOS, scoped storage na Androidzie). **Q: Jak obchodzicie się z prywatnością i zgodami?** Egzekwowanie zgód jest wbudowane w instrumentację. Zdarzenia dotykające danych osobowych są oznaczane do routingu respektującego zgodę. Przegląd prywatności jest częścią fazy instrumentacji, a nie dodatkiem po fakcie. ### Powiązane - https://dfzoo.ai/pl/uslugi/operate-measure-maintain/llm-observability - https://dfzoo.ai/pl/uslugi/operate-measure-maintain/cost-model-optimization - https://dfzoo.ai/pl/uslugi/evaluate-audit-secure/production-readiness --- ## Optymalizacja kosztów i modeli - Operuj i Mierz **URL:** https://dfzoo.ai/pl/uslugi/operate-measure-maintain/cost-model-optimization **Service type:** Audyt optymalizacji kosztów LLM **Anchor:** Operuj, Mierz i Utrzymuj (Operuj i Mierz) > Audyt optymalizacji kosztów i modeli dla produkcyjnych systemów LLM - caching, routing, kompresja promptów i usprawnienia po stronie dostawcy, które odzyskują zwykle 20-60% miesięcznego wydatku na LLM w ciągu 4-8 tygodni. ### Podsumowanie dfzoo AI Institute prowadzi audyty kosztów LLM i optymalizacji modeli dla zespołów inżynierskich, których rachunek za LLM rośnie szybciej niż ruch. Instrumentujemy system pod przypisanie kosztu do każdego wywołania, audytujemy obecne użycie, rekomendujemy i wdrażamy oszczędności po stronie dostawcy (caching promptów, kierowanie do mniejszych modeli, kompresja promptów, strojenie ponowień) i dostarczamy zmierzone porównanie przed/po z jakością chronioną przez eval. Typowi klienci odzyskują 20-60% miesięcznego wydatku na LLM. Z certyfikatem ISO 9001:2015. ### Dla kogo to jest - Zespoły inżynierskie, których miesięczny rachunek za LLM przekroczył 10 tys. EUR i dalej rośnie - CFO i partnerzy finansowi pytający, dlaczego wydatek na LLM przekracza plan - Liderzy product engineering wdrażający funkcje AI o postawionej na głowie ekonomice jednostkowej - Firmy w fazie wzrostu przygotowujące rundę finansowania, potrzebujące dających się obronić liczb kosztów AI ### Problemy, które rozwiązujemy - Wydatek na LLM rośnie szybciej niż ruch użytkowników i nikt nie wie dlaczego - Wskaźnik trafień w cache promptów jest niski albo nie jest mierzony; każde wywołanie płaci pełną cenę - Wszystkie funkcje kierują do najdroższego modelu, niezależnie od wagi zadania - Ponowienia przy błędach przejściowych po cichu mnożą rachunek ### Co dostajecie - Dashboard przypisania kosztu - per funkcja, per segment użytkowników, per model - Rekomendacje optymalizacji uszeregowane według stosunku oszczędności do nakładu pracy - Wdrożenie najlepszych usprawnień (caching, routing, przebudowa promptów, polityka ponowień) - Pomiar jakości chroniony przez eval: ta sama albo lepsza jakość po optymalizacji - Pisemny raport z audytu ze zmierzonymi oszczędnościami przed/po ### Jak pracujemy 1. **1. Instrumentacja kosztu** (Tydzień 1-2): Instrumentujemy przypisanie kosztu do każdego wywołania. Prowadzimy pomiar przez 1-2 tygodnie, żeby uchwycić reprezentatywny obraz kosztu. 2. **2. Audyt i rekomendacje** (Tydzień 2-3): Analizujemy dane o kosztach. Identyfikujemy największe możliwości oszczędności. Spisujemy rekomendacje z szacunkiem nakładu pracy. 3. **3. Wdrożenie** (Tydzień 3-6): Wdrażamy najlepsze usprawnienia. Ustawiamy wdrożenie chronione przez eval: jakość musi się utrzymać, zanim spadek kosztu zacznie się liczyć. 4. **4. Pomiar i raport** (Tydzień 6-8): Mierzymy porównanie przed/po przez 2-4 tygodnie. Dostarczamy audyt z udokumentowanymi oszczędnościami. ### FAQ **Q: Co naprawdę znaczy 20-60% oszczędności?** Mierzone w miesiącu po wdrożeniu, przez porównanie wydatku na API LLM przy równoważnym wolumenie ruchu do miesiąca bazowego sprzed audytu. Wzrost ruchu wyłączamy z porównania. Konkretni klienci odzyskują różne kwoty zależnie od stanu wyjściowego - zespoły już korzystające z cachingu widzą mniej, zespoły, które nigdy nie optymalizowały, widzą więcej. **Q: Czy optymalizacja pogorszy jakość?** Chronimy przed tym pipeline'ami eval. Każda optymalizacja jest testowana wobec golden set; wdrażamy tylko zmiany utrzymujące albo poprawiające jakość. Jakość to ograniczenie, koszt to cel optymalizacji. **Q: Które funkcje dostawcy liczą się najbardziej?** Caching promptów to zwykle największy zysk (Anthropic prompt caching, OpenAI prompt caching). Dalej kierowanie do mniejszych modeli (użycie Haiku/4o-mini tam, gdzie to odpowiednie), kompresja promptów (usuwanie zbędnego kontekstu) i strojenie ponowień (brak cichych ponowień przy błędach idempotentnych). **Q: Robicie to dla jednego dostawcy czy wielu?** Dla obu. Klienci korzystający z wielu dostawców mają często logikę routingu wybierającą najtańszy kwalifikujący się model dla danego zadania - optymalizujemy ten routing w ramach projektu. Klienci z jednym dostawcą optymalizują w obrębie jego cennika. **Q: Ile kosztuje projekt?** Stała cena, zwykle zależna od rozmiaru obciążenia LLM. Większość projektów zwraca się w 1-3 miesiące zmierzonych oszczędności. Wąski zakres podajemy na rozmowa wstępna. ### Powiązane - https://dfzoo.ai/pl/uslugi/operate-measure-maintain/llm-observability - https://dfzoo.ai/pl/uslugi/operate-measure-maintain/app-maintenance - https://dfzoo.ai/pl/uslugi/build-automate-modernize/ai-powered-custom-development --- ## Utrzymanie aplikacji - Operuj i Mierz **URL:** https://dfzoo.ai/pl/uslugi/operate-measure-maintain/app-maintenance **Service type:** Utrzymanie aplikacji **Anchor:** Operuj, Mierz i Utrzymuj (Operuj i Mierz) > Bieżące utrzymanie systemów tworzonych z udziałem AI - migracje dostawców, strojenie promptów, aktualizacje zależności i cicha praca inżynierska, która utrzymuje produkcyjne systemy w ruchu, gdy krajobraz AI się zmienia. ### Podsumowanie dfzoo AI Institute dostarcza bieżące utrzymanie aplikacji tworzonych z udziałem AI. Systemy AI wymagają utrzymania, jakiego nie potrzebują tradycyjne aplikacje: migracji dostawców, gdy model jest wycofywany, strojenia promptów, gdy dane wejściowe się zmieniają, ponownego ustalania punktu wyjścia dla eval, aktualizacji zależności warstwy LLM SDK, monitoringu kosztów, gdy wydatki dryfują. Model stałej współpracy pasuje zespołom z ciągle rozwijanymi funkcjami AI; jednorazowe umowy pasują do konkretnych zdarzeń, jak migracja dostawcy. ### Dla kogo to jest - Zespoły inżynierskie z funkcjami AI na produkcji, potrzebujące seniorskiego wsparcia w niepełnym wymiarze - CTO, których zespół zbudował funkcję AI i poszedł dalej; teraz nikt za nią nie odpowiada - Zespoły produktowe, których dostawcy AI wycofują modele szybciej, niż zespół zdąży zareagować - Firmy w fazie wzrostu rozbudowujące albo utrzymujące funkcje AI bez powiększania zespołu ### Problemy, które rozwiązujemy - Dostawca wycofuje model; zespół nie ma planu migracji i nie ma czasu go napisać - Prompt, który działał przy premierze, sześć miesięcy później daje gorsze wyniki - dryf - Wersja SDK LLM utknęła na dniu premiery; poprawki bezpieczeństwa się piętrzą - Koszty rosną na tyle wolno, że nikt nie zauważa aż do kwartalnego przeglądu ### Co dostajecie - Miesięczne godziny utrzymania obejmujące powyższe kwestie (stała współpraca) - Plany migracji dostawców i ich realizacja, gdy modele są wycofywane (chronione przez eval) - Kwartalny przegląd promptów: ponowny eval, sprawdzenie regresji, drobne strojenie - Cykl aktualizacji zależności: SDK, biblioteki observability, frameworki eval - Kwartalny pisemny raport: co się zmieniło, czego się spodziewać, co budżetować na kolejny kwartał ### Jak pracujemy 1. **1. Przejęcie** (Tydzień 1-2): Czytamy system. Mapujemy zależności, prompty, konfigurację eval, observability i procedury dyżurowe. Dokumentujemy luki. 2. **2. Stabilizacja (jeśli potrzebna)** (Tydzień 2-4 (jednorazowo)): Zajmujemy się krytycznym długiem utrzymaniowym przed przejściem do stałej współpracy: brakujące eval, nieuporządkowane prompty, stare SDK. 3. **3. Stała współpraca** (Na bieżąco): Miesięczne godziny utrzymania, kwartalne spotkania kontrolne, doraźne wsparcie przy zdarzeniach u dostawcy albo incydentach. 4. **4. Kwartalny przegląd** (Kwartalnie): Pisemny raport o tym, czego dotknęliśmy, co dryfuje i co budżetować na kolejny kwartał. ### FAQ **Q: Stała współpraca czy jednorazowa umowa?** Oba modele działają. Stała współpraca (10-40 godzin/miesiąc) pasuje zespołom z ciągle rozwijanymi funkcjami AI na produkcji. Jednorazowe umowy (migracja dostawcy, ponowne ustalenie punktu wyjścia dla eval, ponowne strojenie promptów) pasują zespołom z konkretnym zdarzeniem. **Q: A jeśli wystąpi incydent produkcyjny?** Klienci stałej współpracy dostają doraźne wsparcie przy incydentach - dołączamy do mostu, pomagamy debugować, piszemy analizę po incydencie (post-mortem). SLA ustalamy w umowie o stałej współpracy. **Q: Robicie pełne DevOps / SRE?** Obsługujemy operacje specyficzne dla AI (dryf promptów, migracje modeli, pipeline'y eval, monitoring kosztów LLM). Zwykła praca SRE (Kubernetes, infrastructure-as-code, tradycyjny alerting) jest w zakresie tylko wtedy, gdy przecina się z funkcjami AI. **Q: Jak działają migracje modeli?** To standardowe zdarzenie utrzymaniowe. Ustalamy na nowo punkt wyjścia dla eval wobec nowego modelu, uruchamiamy testy regresji, migrujemy prompty tam, gdzie nowy model wymaga innych wzorców, i mierzymy porównanie przed/po pod kątem jakości, opóźnień i kosztu. **Q: Możecie utrzymywać kod, którego nie zbudowaliście?** Tak. Faza przejęcia obejmuje czytanie i dokumentowanie systemu. Może okazać się konieczne zajęcie się długiem utrzymaniowym, zanim stała współpraca stanie się możliwa. ### Powiązane - https://dfzoo.ai/pl/uslugi/operate-measure-maintain/llm-observability - https://dfzoo.ai/pl/uslugi/operate-measure-maintain/cost-model-optimization - https://dfzoo.ai/pl/uslugi/build-automate-modernize/ai-powered-custom-development ---