Terminal X published its first Terminal X Bench scorecard for new frontier models, grading Opus 5.5, Sol 6 and Luna 6 separately on retrieval, processing and answer generation using medium thinking, the same test sets and the same grading prompts. Opus 5.5 led retrieval (80% average versus 56% for Sol 6 and 51% for Luna 6) and processing but was kept out of production because it refused routine finance prompts, and Sol 5.6 remains the answer generation model after its answers were preferred over Sol 6 in 68% of head to head comparisons.
Our readInvestment firms get a stage level record of which foundation models the platform tests, which one runs answer generation, and why a new release was held back, which supports third party model change control in a model risk program. It indicates new model releases are gated on per stage testing before routing changes.