Technical report · 2026
Evaluating large language models for regulated enterprise workflows
A practical evaluation framework for LLM systems in banking, insurance, and government contexts — metrics, test harnesses, and a taxonomy of failure modes.
In preparation