Reasoning Reliability Audit of an Enterprise AI Model

The goal of this project was to test an enterprise AI model to see how it performs on advanced mathematical problems, nontrivial quantitative tasks, and multi-step long-horizon tasks. 

Each of the tasks was checked with the Lean theorem prover for correctness. 

The verifiers had to be deterministic, meaning they had to be rule-based with no use of AI. 

The deliverable was:

  1. a suite of tasks in each category with their verifiers,
  2. an infrastructure in Python that implements the verification with a dashboard that summarizes the results, and 
  3. a report with recommendations.

Skills and deliverables: Artificial Intelligence, Mathematics, Statistics, Data Analysis, Python, Technical Writing

ai verification long horizon task
reasoning ai verification audit with the Lean theorem prover
Scroll to Top