Back to Systems
Internal R&DEvaluation & GovernanceActive

M-Class Harness

A governed evaluation methodology, scoring framework, and review standard for coding, reasoning, and multi-step AI workflows.

As AI agents move beyond single responses, evaluation must measure sustained work: state tracking, tool use, failure recovery, handoff quality, and evidence trails — under a consistent, governed scoring standard.

Status
Internal R&D
Evidence
Prototype implemented
Activity
Active
Type
Governed Evaluation Methodology
Category
Evaluation & Governance
Owner
Deep Bound Research Lab
Class
Evaluation Harness
Related
long-horizon-harnessboundaryex1

Problem Space

Most AI evaluations are too short to expose failures in sustained reasoning, context retention, tool discipline, and recovery from bad intermediate states — and lack a consistent review standard for scoring what they do capture.

System Direction

M-Class defines the evaluation methodology and scoring framework: scored artifacts, evidence-led review patterns, and governed review standards. Execution environments for long-horizon tasks are provided separately by the Long-Horizon Harness; runtime supervision belongs to TripSitter.

Public Capabilities

  • 01Evaluation methodology and review standards
  • 02Scoring frameworks for sustained work
  • 03Coding and reasoning task review
  • 04Evidence-led artifact inspection
  • 05Public-safe benchmark packaging
Disclosure Boundary

M-Class is presented publicly as an evaluation research program. Internal scoring rubrics, prompts, traces, and harness mechanics are not disclosed.

What Is Not Disclosed

Private implementation details, security-sensitive internals, and unreleased runtime architecture are intentionally not disclosed.