The AI Post
Agents & CodingOpen ModelsEnterpriseFundraisingGenerative MediaGovernanceInferenceInfrastructureLegal & SafetySector Impact
← Front Page Model Releases · Ant Group · OpenAI · Anthropic

Vals AI scores Ant Group's Ling 3.0 Flash Fin at 54.9% on finance test

The evaluator put the model fourth among open-weight entries and said it costs $0.045 per task, the cheapest in the top 20.

Vals AI said Ant Group's Ling 3.0 Flash Fin scores 54.9% on its Finance Agent v2 benchmark, at $0.045 per task. The evaluator ranked the model fourth among open-weight entries and 16th overall, and called it the cheapest in the top 20 by a wide margin.

Vals AI put the score within one point of GPT-5.6 Terra and Claude Sonnet 5, at roughly 25 to 75 times less per task. Those comparisons come from the evaluator's own runs, according to its posts, and the paper has not reproduced them.

The model is specialised, not stronger overall, Vals AI said. LegalBench falls from 79.7% to 68.7%, MedScribe from 80.9% to 75.6% and MedCode from 32.3% to 29.3%. The evaluator did not name the model those declines compare with, and gave no dates for the runs.

Long-horizon coding is where the model falls away, on the evaluator's numbers. It scores 3.4% on Vibe Code Bench, 2.9% on Code Migration and zero on HLAB. Vals AI gave no figure for how rivals score on those three.

Open-weight finance models that price this low could change what a bank pays for routine analysis, if the scores hold outside one evaluator's harness. Neither Ant Group nor Vals AI has published the task mix behind Finance Agent v2.

Sources 3 sources

  1. Commentary @ValsAIpost on X
  2. Commentary @ValsAIpost on X
  3. Commentary @ValsAIpost on X