← Back to stories
AI

Presentation: Building Reusable Evaluation Frameworks for Agentic AI Products

Susan Chang explains how Elastic transitioned from siloed, ad-hoc AI agent evaluations to a unified, production-grade framework. She discusses balancing LLM-as-a-judge with deterministic rules, bridging Python data science evals with TypeScript production code, and implementing deep tracing to catch regressions across complex RAG and cybersecurity workloads while preserving domain context.

You're reading a preview. The full article is published by InfoQ on their website.

Read the full story on InfoQ