Evals for workflows, not just nodes
Per-node evals stay green while the workflow does the wrong thing. Evaluating workflows end to end: representative cases, workflow-level scoring over recorded runs, and a regression gate.
3 posts tagged with this.
Per-node evals stay green while the workflow does the wrong thing. Evaluating workflows end to end: representative cases, workflow-level scoring over recorded runs, and a regression gate.
A runnable workflow-runtime repo with one workflow.yaml, POC and production variants, LangGraph/MSAF implementations, and failure scenarios that make the architecture testable.
How we represent human work as a graph, run it with LLMs, and what changes when the workflow has to survive hours, days, and thousands of concurrent runs.