Encode the state once. Decide everything in parallel.
System One is Qwen3.5-4B fine-tuned to answer typed questions with a calibrated probability
distribution over exactly the options you allow — an 80% is meant to be right about 80% of the time.
The inference path is what makes many of those decisions cheap: the state is prefilled once, every
(question, option) pair branches from that cache in one batched forward, and nothing is generated, so the
output cannot leave your schema. The instruct model on the same base answers the same questions as JSON,
one token per forward, with no probabilities at all.
Pick an example or write your own state and decisions, then press Run.
Both 4B models run live on ZeroGPU; the first call on a fresh worker is slower.