Porvenir, Learned World Models for Mineral Processing, Measured Against Holding the Last Value
A forecaster maps history to a future; a world model takes a proposed action sequence as an argument and returns a distribution over futures, which is the single difference that makes what-if questions, planning and regime detection expressible. Porvenir asks three questions of mineral-processing records and answers each with a measurement: whether a learned latent can track process state the plant does not instrument, whether aleatoric and epistemic uncertainty can be separated and made separately actionable, and how far imagination drifts from reality when a plan is executed. Eleven rungs plus a zero-shot yardstick, from persistence and seasonal naive through VARX, GRU-D, DeepAR, MQ-RNN, a probabilistic ensemble, a patch transformer and recurrent state-space models, every one scored per horizon against holding the last value. On the real iron-ore flotation record, seven of nine learned rungs do not beat persistence, and the app says so on the case. Lifecycle: live, re-baked with its numbers withdrawn once.
Business Context
For a plant the deliverable is not a forecast but an answer to what-if, with the two kinds of doubt kept apart: noise the plant cannot remove and ignorance more data would reduce. On the real flotation record the honest reading is on the case page: seven of nine learned rungs did not beat holding the last value at the hourly cadence, no action set beat persistence beyond the first step, and the case reports that the plant record does not identify controllable dynamics at that cadence. The product is a research instrument, not a plant control system; it issues no setpoint advice, and logged actions come from a closed loop, so observational data does not identify interventions without assumptions, which the footer states.
Strategic Value
Porvenir published its first numbers and then withdrew them: the 0.09.000 engine bake had a GRU-D rollout that never consumed the first future action, a biased CRPS estimator, an unseeded VARX and yardstick, and a subsampled yardstick with a mismatched aggregation, so every number from that bake was marked withdrawn and replaced by a re-bake, with the defects recorded in the product's findings. The deploy script drives a real browser against the live URL and refuses on an empty root or any console error, after a green deploy once shipped a blank page whose title and data index were both correct. A CC-BY preprint carries the method and results on Zenodo. Its sibling Fragua ensembles closed-form equations where Porvenir learns latent dynamics.
The Challenge
Plant historians hold years of process records, and the models trained on them are almost always forecasters: given the past, predict the next values. A forecaster cannot answer the question an operator actually has, which is what would happen if a different action were taken, because the action is not an argument of the model. A world model makes the action sequence an input and returns a distribution over futures, which turns what-if questions, planning and regime detection into things that can be asked at all. Whether such a model earns its complexity on real process data is an empirical question, and the bar it has to clear is embarrassingly low and rarely reported: hold the last value.
Our Approach
The engine is latentplant, a separately published package (PyPI, MIT) from its own repository; Porvenir is the product on the CAOS archetype under the full scientific contract. The ladder is ordered by what each rung can express: persistence (the bar), seasonal naive, VARX, GRU-D (does modelling a channel's staleness beat masking it), DeepAR, MQ-RNN (how much long-horizon error is compounding rather than ignorance), a probabilistic ensemble that separates aleatoric from epistemic uncertainty, a patch transformer, a diagonal S4, a recurrent state-space model with KL balancing and free bits, and an ensembled RSSM that adds a usable epistemic signal; a Chronos zero-shot yardstick is scored per case on the same windows. Cases are the CC0 iron-ore flotation table (737,453 rows at 20 s, aggregated to the hourly assay cadence), the Tennessee Eastman archive with a separate surprise artifact, and two in-house simulators with ground truth. Each case declares its expectation and what would refute it before it runs; every rung must beat persistence per horizon or the report says it did not. The browser imagination lane runs the exported ONNX with onnxruntime-web; the reference library resolves 81 identifiers against arXiv and doi.org.
Key Performance Indicators
| KPI | Baseline | Result | Impact |
|---|---|---|---|
| The bar is holding the last value | Learned models reported against each other | Every rung scored per horizon against persistence; on the real flotation record 7 of 9 learned rungs did not beat it at the hourly cadence | The case page says the record does not identify controllable dynamics at that cadence |
| Two kinds of doubt, kept apart | One error bar | A probabilistic ensemble and an ensembled recurrent state-space model separate aleatoric from epistemic uncertainty and expose the epistemic signal as actionable | Noise the plant cannot remove is told apart from ignorance more data would reduce |
| Numbers withdrawn, then re-baked | A first bake published as final | Four defects found in the 0.09.000 engine bake (an action never consumed, a biased CRPS estimator, unseeded rungs, a mismatched yardstick aggregation); every number from it marked withdrawn and replaced | The findings are part of the product |
Proprietary, source code not publicly available
Architecture
porvenir pipeline
Technology Stack
Application Screenshots

