Aviation Compliance SaaS Cuts AI Model QA from Manual Spot Checks to Automated Field-Level Validation with Claude on AWS Bedrock
At a glance
Aviation compliance software runs on trust. This aviation compliance SaaS company digitizes FAA-required aircraft records for business and general aviation, but adopting generative AI to extract maintenance data carried real risk: a misread compliance field is not a rounding error, it is a wrong airworthiness fact relied on by a mechanic or owner.
NMD built the company a field-level model evaluation harness on AWS Bedrock that scores every candidate model against ground-truth aircraft records before it touches a customer document. The outcome: the company swaps in newer, cheaper Claude models with evidence, not guesswork.
The company now approves AI models for production the way a regulator approves a part: with field-by-field proof. That repeatable gate lets the company ship faster, cheaper extraction across every report type without ever gambling with an aircraft’s compliance record.
Industry
Use Case
Intelligent Document Processing, Compliance & Governance, Responsible AI / AI Governance
Solution implemented
- Ground-truth evaluation harness scores extraction accuracy field by field per report type.
- Benchmarked Claude Sonnet 4.6 and Opus 4.6 against prior model generations on AWS Bedrock.
- Compared LLM-only versus Textract-hybrid extraction pipelines on accuracy and per-page cost side by side.
- Captured token usage and processing latency per run to inform model and pipeline choice.
- Automated scoring pipeline replaces a manual, one-person side-by-side report comparison process for QA.
- Established acceptance criteria that now enable repeatable, evidence-gated production model release decisions.
The value equation
- Company approves new AI models for production with proof, not guesswork or spot checks.
- Field-level accuracy scoring lets the company adopt cheaper models the moment they qualify.
- Evaluation harness protects FAA compliance trust as AI-driven extraction scales.
- Repeatable model QA replaces a manual process that once relied on one person's judgment.
Company Snapshot
A HealthTech-adjacent-regulated aviation SaaS company (Aviation/Transportation vertical) providing FAA-compliant aircraft records and logbook software, now backed by an evidence-gated AI model evaluation process for compliance data extraction.
Location
United States
Customer Situation
This aviation compliance SaaS platform is the system of record aviation operators trust to prove FAA compliance across hundreds of customers and thousands of tracked maintenance items per aircraft, spanning airworthiness directives, service bulletins, and inspections. To extend that trust to generative AI extraction, the company needed proof, not intuition, that a given model read a maintenance report correctly before that data ever reached a customer’s compliance file. A wrong field here is not a UX complaint, it is a compliance record a mechanic, owner, or regulator may rely on when an aircraft is bought, sold, financed, or certified airworthy. This same problem sits under every regulated industry now weighing generative AI for document-heavy compliance work.
NMD Solution
NMD reviewed the company’s manual, ad hoc model comparison process and found no repeatable way to prove one model’s extraction was more trustworthy than another before production use, especially as newer, cheaper models released to the market at an increasing pace. The solution required a ground-truth evaluation harness built on AWS Bedrock that scores every candidate model, including Claude Sonnet 4.6 and Opus 4.6, field by field across real aircraft report formats, providers, and page-layout variations, capturing accuracy, cost per page, and processing latency side by side before any model is promoted into the company’s production compliance extraction pipeline.
What We Delivered
The company runs every new model through an automated, field-level evaluation harness before production promotion. Each candidate is scored against ground-truth aircraft records across report types, providers, and document formats, with accuracy, cost per page, and processing latency compared side by side across the full model set under consideration for release, including newly available Claude model generations as they ship. The company is achieving near-100% field accuracy on clean reports and high-90s accuracy on its most variable formats, with every model swap now backed by evidence instead of a manual spot check.The platform runs a full model evaluation engine on Bedrock: developers define a scenario, select models and embedding configurations, and run it live or virtually, receiving cost, latency, and accuracy scores per model in a single session. The evaluation logic (rate-limited memory writes with completion polling) also improved the reliability of the company’s core memory benchmark scores.
Ready to Transform Your Transportation Business?
See how New Math Data can transform your transportation business with AWS-powered innovation.