ES
← Back to Portfolio
Scientific Machine Learning September 2026

Enunciado, How Faithfully Language Models Turn a Problem Statement into a Solvable Model

The field reports whether a generated model ran and calls that correct. Enunciado measures faithfulness instead: a corpus of 20 authored optimization statements across five complexity tiers, every trap covered and four controls without one, is handed to language models, and the formal model each returns is judged by an oracle that is not a language model: executable, structural and property layers over a solver, with a duality certificate. Sixteen models from four providers, 320 calls, every row complete. The best local model, phi4, is faithful on 5 of 20. The output cap decides the reasoning models: DeepSeek-V4-Pro is faithful on 2 of 20 at 8192 tokens and 11 of 20 at 32768. The first version of this measurement read +0.000 and was wrong.

Corpus
20 authored optimization statements, 4 per tier over 5 complexity tiers, every trap covered, 4 controls without a trap; the bake verifies every reference solves, every claimed optimum matches the solver and every property relation holds
Measurement
16 models from 4 providers (Anthropic, Z.AI, DeepSeek, 11 open-weight models through Ollama on one 8 GB laptop GPU), 320 calls, every row complete, a 640-record ledger
The three repositories
planteo (the representation) and copela (the harness: provider seam, append-only ledger, four verdict layers, budget guard), both published on PyPI from their own repositories; Enunciado declares no package
Defects caught by gates, not by review
10, including 3 wrong claimed optima, a solver wrapper raising on the infeasible case, two sweeps interleaved in one ledger, a +0.000 that meant no measurement; and 4 times a gate caught itself
Workbench
14 methods in 4 groups, a Duality tab as a four-part optimality certificate, model matrices, a cap-sensitivity table, ledger-computed caveats; 207 browser checks pass against the live origin on 0.08.000 in dark, light and Spanish, the Spanish pass walking every tab and failing on unaccented words; the wiki carries 17 figures exported from the same components
Not built, not claimed
The other three target families, the four manuscripts (one registered as pending), and any frontier-model measurement beyond the sixteen
Architecture diagram of Enunciado, How Faithfully Language Models Turn a Problem Statement into a Solvable Model
#formalization #optimization #autoformalization #llm-evaluation #benchmark #pyomo #solvers #ollama #reproducible-research

Business Context

Anyone who hands an optimization, a simulation or an experiment design to a language model needs to know how often the returned model is the problem they stated, not merely a model that solves. Enunciado gives that number per model, per tier and per trap, with the caveats attached to the same ledger, so a team can pick a model for a formalization task on measured faithfulness rather than on a leaderboard that scored compilation. It also shows what moves the number: the output cap decides the reasoning models, the local lane silently shifts its default context, and a metamorphic refutation can be an unbounded candidate misread by three layers, each of them a recorded finding rather than a footnote.

Strategic Value

The product is the first adopter of a narrative-to-formal measurement archetype, and its record of what the gates caught is its strongest argument. Ten defects were caught by gates rather than by review: three wrong claimed optima; a solver wrapper that raised on the deliberately infeasible case instead of reporting infeasibility; two sweeps that shared one ledger and interleaved records from different code versions; a report that printed a gap of +0.000 when it meant no measurement; vendor names found in the CLI by the seam test; a deep link that answered 404 while rendering correctly. Four times a gate caught itself: the design-document guard read a prose cross-reference as a duplicate requirement, the UI gate passed its artifacts-loaded check while the page showed a router error, its theme switch wrote a key nothing reads so both passes ran in one theme, and its test server modelled the host wrongly in two different ways. Both Claude gaps of +0.050 rest on one refutation each that lands exactly on the reference whole-number optimum in a statement that never fixes integrality; read as allowed they are 0.000, the page states it, and the definition is a recorded decision. The other target families (mathematical and simulation modelling, experiment design, machine-learning framing), the manuscripts, and any measurement of a frontier model beyond the sixteen are not built and not claimed. Release 0.08.000 put the six pages in the standard order with App first, split the Benchmark into six tabs, and wrote the Spanish surface with its accents, 1,813 words decided in context.

The Challenge

Enunciado is the Spanish word for the statement of a problem: the text as it is handed to you, before anyone has decided what the variables are. Turning it into a formal, solvable model is the step the published evaluations skip. Four separate literatures report the artifact ran and call that correct, and each, read closely, admits the measured faithfulness is lower: in optimization modelling, reaching the reference objective does not imply a correct model (arXiv:2508.10047); in statement formalization there is a 3.0 to 29.0 point gap between compiling and being faithful, the strongest agent compiling 89.5 per cent of the time and being faithful 60.5 (arXiv:2606.31002); in experiment design every model is weak at datasets, baselines and metrics (arXiv:2608.03501); in simulation modelling the model runs but the causal reasoning is weak and no single model dominates (arXiv:2605.28994). A 2025 position paper argues these are one problem and supplies none of the machinery (arXiv:2509.09810). The community benchmarks cannot serve: they carry 8.13 to 54.0 per cent error rates, two of the most cited cannot be redistributed by licence, and the adjacent machine-learning family is contaminated.

Our Approach

Three repositories, one product. planteo (its own repository, on PyPI) is the representation: dimensions on every quantity, provenance on every element, and a record of what the statement left open. copela (its own repository, on PyPI) is the harness: a provider seam, an append-only ledger, four verdict layers and a budget guard that refuses to call a model without a declared budget. Enunciado is the product and declares no package: the corpus, the bake, the measurement and the web surface. The corpus is 20 authored cases, four per tier over five complexity tiers, every trap covered and four controls with no trap, written rather than imported for the reasons above. The bake verifies that every reference solves, every claimed optimum matches the solver and every property relation holds on the reference; it caught three wrong claimed optima out of twenty. The measurement calls sixteen models from four providers (Anthropic, Z.AI, DeepSeek and eleven open-weight models through Ollama on one 8 GB laptop GPU), 320 calls with every row complete in a 640-record ledger, and scores each answer through the executable, structural and property layers with a single failure taxonomy held equal to the classifier. The workbench shows fourteen methods in four groups, a Duality tab that is a four-part optimality certificate, a sidebar that diagnoses the selected case from the ledger, model matrices, a cap-sensitivity table, and caveats computed from the ledger rather than written.

Key Performance Indicators

KPIBaselineResultImpact
Faithfulness, not compilationThe published evaluations score whether the model ranExecutable, structural and property layers over a solver, with a four-part duality certificate; the best local model, phi4, is faithful on 5 of 20 casesA model is chosen on what it got right, not on what it returned
The cap decides the reasoning modelsA single number per modelDeepSeek-V4-Pro is faithful on 2 of 20 at an 8192-token output cap and on 11 of 20 at 32768; the cap-sensitivity table and the at-cap counts are publishedA rank without its cap is not a result
A measurement that corrected itselfThe first report read a gap of +0.000It meant no measurement; the report now distinguishes the two, and the 0.07.001 release re-derived report, attempts and rescore after CI found the artifacts covered 551 calls against a 640-record ledgerEvery rate on the site is recounted from the committed ledger in CI

Architecture

enunciado pipeline

enunciado pipeline

Technology Stack

Python planteo copela Pyomo Ollama TypeScript React Vite KaTeX

Application Screenshots

Enunciado, How Faithfully Language Models Turn a Problem Statement into a Solvable Model
Enunciado, How Faithfully Language Models Turn a Problem Statement into a Solvable Model