savanty vs NL4Opt, Chat LLMs and Hand-Written Solver Models
Where savanty fits: an LLM-to-Clingo pipeline compared with asking a chatbot, NL4Opt-style research systems, and writing an OR-Tools or MiniZinc model yourself.
Correction, 7 October 2026: an earlier version of this article said
savanty compiles problems to MiniZinc, OR-Tools CP-SAT and Z3, that
MiniZinc is its default backend, that it returns a proof of optimality,
and it showed a Python object model (Problem, Variable, Constraint)
that does not exist. None of that was true. savanty has one solver,
Clingo, and one Python entry point, solve_optimization_problem. The
article has been rewritten from the source repository and the product
documentation.
The short version
There are four common ways to get from an English description of a constraint problem to an answer. You can ask a language model for the answer. You can ask a language model to write solver code and run it yourself. You can use one of the research systems that translate word problems into a formal optimisation model. Or you can write the model yourself.
savanty is a fifth arrangement, and it is narrower than any of the four. It accepts English, has a language model translate it into Answer Set Programming (ASP), and hands that program to the Clingo solver, which either returns an assignment that satisfies every constraint in the program or proves that no such assignment exists. This post positions savanty against each alternative, including the places where it loses. It contains no benchmark table, because we have not published one. Where a comparison would depend on a number, we say so and point at how you would measure it.
What savanty is
savanty is an MIT-licensed Python package (Python 3.10+, pip install savanty) with a CLI (savanty -p "...") and an optional FastAPI
service (savanty --web, OpenAPI docs at /docs). Language-model
calls are orchestrated with DSPy
typed signatures. The default model is GPT-4o; any OpenAI-compatible
endpoint works, including Ollama Cloud.
A request goes through a fixed sequence. A suitability check
classifies the problem and, if ASP is the wrong tool, returns
not_suitable with a suggested_tool such as scipy, cvxpy or pandas.
A gap-identification step returns clarifying questions when the
description is missing entities or counts. The model then analyses the
problem and generates a program as strict JSON (facts, rules and an
optional #minimize or #maximize directive) over a single canonical
decision relation, assign(Var, Value). Requirements are integrity
constraints, rules that begin with :-. Clingo runs the program and
the outcome is classified as ok, syntax_error, unsat or empty.
Anything other than ok goes to a typed repair loop; for unsat,
the runtime first computes a minimal unsatisfiable core by deletion
filtering, so the model is told exactly which constraints conflict.
The public API is one function:
from savanty import solve_optimization_problem
result = solve_optimization_problem("""
Schedule 4 nurses (Alice, Bob, Carol, Dave) for morning/evening
shifts over 5 days. Each shift needs 1 nurse. No nurse works two
shifts on the same day. Max 4 shifts per person.
""")
print(result.solution) # parsed assign/2 atoms
print(result.asp_code) # the generated encoding, for inspection
There is no backend selection. There is no intermediate representation that targets several solvers. The language model writes ASP; Clingo solves it.
Five routes from English to an answer
| Route | Who writes the formal model | Solver | What is guaranteed | Characteristic failure |
|---|---|---|---|---|
| Ask an LLM for the answer | Nobody | None | Nothing; the output is plausible text | A confident schedule that breaks a constraint, with no signal that it does |
| Chat LLM writes OR-Tools code; you run it | The LLM, as free-form Python | Usually CP-SAT | The solution satisfies the constraints in the generated code | An INFEASIBLE exit with no hint of which constraint is wrong; omissions buried in code you must read |
| NL4Opt-style research systems | A model, into an LP or MILP formulation | LP / MILP solvers | Solver guarantees for the generated formulation | Mis-extracted entities, coefficients or constraint directions |
| Hand-written model (OR-Tools, MiniZinc, Z3, Gurobi) | A trained modeller | Whatever you choose | Solver guarantees for the model; faithfulness is the modeller’s job | Modelling time, and the modeller’s own mistakes |
| savanty | The LLM, constrained to assign(Var, Value) | Clingo | A returned assignment satisfies every integrity constraint in the encoding; UNSAT is a proof for that encoding | A constraint quietly omitted during translation |
The rightmost column is the one to read. Every route that involves a solver inherits the solver’s guarantee for whatever model reached it, and none of them guarantees that the model matches what you meant. The routes differ in who writes the model, how readable it is, and what happens when the solver rejects it.
savanty vs asking a language model for the answer
This is the easiest comparison and the least interesting. A language model asked to produce a roster samples a plausible one. Small problems often come back correct, because the most plausible output happens to be valid. Larger ones come back looking equally confident with a double-booking somewhere in the middle, and nothing in the output distinguishes the two cases. The model is not running a search; it has no internal step that rejects an invalid assignment.
savanty never asks the model for a solution. It asks for a program, and Clingo does the search. If you only need a rough draft that a human will check line by line anyway, asking a chatbot directly is cheaper. If anything downstream will treat the answer as correct, it is the wrong tool.
savanty vs a chat LLM writing OR-Tools code
This is the closer comparison, because both routes use the same class of model and both end in a real solver. The difference is the loop around the model. In the chat-and-paste workflow, the model writes Python against OR-Tools, you run it, and if it fails you paste the stack trace back. When the generated model is infeasible, the script typically reports that no solution was found and nothing more, unless someone wrote assumption tracking. The loop runs in your head and your terminal history.
savanty makes that loop part of the program. Failures are typed: a
parse error is returned verbatim, an unsatisfiable program comes back
with its minimal core, and a program that emits no assign/2 atoms is
flagged as a contract violation. On an unsatisfiable core, the repair
prompt asks the model to decide between “one of these constraints
misformalises the problem” and “the problem as stated is genuinely
infeasible”, and in the second case to leave the encoding alone so the
infeasibility is reported faithfully rather than repaired away. The
suitability check also stops the model from writing solver code for a
problem that is really continuous or statistical.
Chat-and-paste is still the better choice when you are solving a problem once, can read OR-Tools Python fluently, or specifically want CP-SAT or one of OR-Tools’ specialised solvers.
savanty vs NL4Opt-style research
The research line on translating natural language into optimisation models is real and worth knowing. The NL4Opt competition at NeurIPS 2022 split the task into two sub-tasks: recognising the entities in a linear-programming word problem, and generating a logical form of the problem that can be converted into a solver’s input format. OptiMUS is an LLM-based agent that formulates mixed-integer linear programmes from natural-language descriptions, writes and debugs solver code, and checks the validity of generated solutions. Logic-LM is the closest relative in spirit: an LLM translates a problem into a symbolic formulation, a deterministic solver performs the inference, and a self-refinement module uses the solver’s error messages to revise the formulation.
savanty differs from these in three ways. First, the problem class:
NL4Opt and OptiMUS are about LP and MILP, much of which sits on the
continuous side of savanty’s suitability check and would be redirected
to scipy or cvxpy. savanty is deliberately confined to discrete,
finite-domain problems. Second, the target language: ASP with a single
decision relation, rather than an algebraic formulation. Third, the
repair signal. savanty ships a Logic-LM-style generic repair mode,
ASPRepairGeneric, in which the model sees only the raw Clingo
message, specifically as a baseline to compare against its typed-core
repair. The repository includes a benchmark harness for that
comparison.
What we are not claiming: we have not published results for savanty on NL4Opt, NLP4LP or any other public dataset, and we are not asserting that it outperforms any of these systems. If that comparison matters to you, the harness is the place to start, on your own problem family.
savanty vs writing the model yourself
If you can already model, you probably do not need savanty. OR-Tools gives you CP-SAT, with first-class interval, no-overlap and cumulative constraints, plus a dedicated routing library. MiniZinc is a solver-independent modelling language, so a model written once can be run on several backends; savanty offers no solver choice at all. Z3 is an SMT solver, the natural fit when the constraints involve arithmetic theories or verification conditions. Gurobi is a commercial MILP solver for large mixed-integer problems. Each is far more capable than a natural-language front end, and each exposes search controls that savanty does not.
savanty’s case is narrower: the person who understands the problem
cannot write any of those, the problem is ad hoc rather than a
long-lived model run millions of times, and the result needs to be
auditable. The generated ASP comes back as result.asp_code, and
because every decision is one assign/2 atom and every requirement is
one :- rule, the encoding is short enough for a non-specialist to
skim for a missing constraint.
What “guaranteed” covers
Clingo is sound and complete over finite domains. Sound: every returned
answer set satisfies every rule in the program, including every
integrity constraint. Complete: if an answer set exists and the solve
timeout (configurable, 120 seconds by default) is not hit, Clingo finds
one, and an UNSAT verdict means none exists. That is a property of
the solver applied to the program it was given.
It is not a guarantee that the program says what you said. The model can omit a constraint, add a spurious one, or turn “up to four” into “exactly four”. The repair loop catches syntax errors, unsatisfiable programs and empty outputs; it cannot catch a quiet omission, because an under-constrained program is perfectly consistent. The same is true of a hand-written CP-SAT model: the solver guarantees the model, the modeller guarantees the intent. The difference is who the modeller is.
When to use which
- Ask a chatbot for a rough draft that a person will verify by hand.
- Chat-and-paste OR-Tools for a one-off problem, if you read Python solver code fluently.
- Write the model yourself in OR-Tools, MiniZinc, Z3 or Gurobi for routing, heavy interval or cumulative constraints, large MILPs, or a stable model you will solve many times.
- cvxpy or scipy for continuous optimisation; savanty will send you there itself.
- savanty for discrete problems such as shift scheduling, task assignment, timetabling, graph colouring and logic puzzles, described by someone who cannot write a solver model, where an inspectable encoding and a faithful “no solution exists” matter.
Where savanty loses
- It cannot detect a constraint that was never translated.
- It has one solver and no solver-level tuning.
- ASP encodes routing and resource-cumulative problems awkwardly.
- Every request pays for a language-model call before the solver runs, so it does not suit tight latency budgets.
- Continuous, statistical and simulation problems are out of scope by design.
- It has no published accuracy figures on public NL-to-optimisation benchmarks.
What to read next
- From English to Clingo: where the translation breaks: the translation problem itself, and how to read an unsat core
- Better Rankings with Fewer Comparisons: the bandit ranking companion
- savanty on GitHub
- savanty product site: quickstart, FAQ and comparisons
- Clingo: the ASP solver savanty uses