x_t
A question comes off the stream
No label comes with it. Everything the agent knows about questions like this one is in its memory of the questions before.
Published at the Conference on Language Modeling (COLM) 2026 · Paper #358
Environment-Driven Dynamic Policies for Continual LLM Improvement
An LLM agent that answers one task after another tends to make the same kind of mistake over and over. Memory helps, but retrieving a similar past question only replays an example; it never says what to do differently. DRPG adds the missing half: a second model reads what the agent got right and wrong, and writes a handful of rules, a policy, that the agent follows on the next question. The policy is rewritten at every step, from a history that keeps growing.
Both policies below were generated by llama-4-scout at step 100 of a real run, from its own correct and incorrect answers up to that point. The difference between them is the paper’s main finding about when this mechanism pays off.
Every rule here travels. Schema validation and join discipline apply to the next SQL query whatever it asks about, which is why this run gained 8.5 points over Self-StreamICL on Spider.
Nothing is trained. The agent keeps a memory of everything it has answered, tagged with the environment’s verdict, and two retrievals run against it: one feeds the agent examples it got right, the other feeds the policy generator an equal number of successes and failures to compare.
# at time step t R_t = retrieve(x_t, correct cases) # examples for the agent R'_t = retrieve(x_t, correct) ∪ retrieve(x_t, wrong) # k/2 each, contrastive P_t = policy_generator(R'_t) # up to five rules ŷ_t = agent(x_t, R_t, P_t) fb_t = environment(x_t, ŷ_t) # correct: 1 or 0 memory.add(x_t, ŷ_t, fb_t)
Two LLM calls per question instead of one, and two retrievals instead of one. On mistral-medium × Spider that comes to 1,676 input and 143 output tokens per question, against 585 and 17 for Self-StreamICL: about three times the tokens, for a method that needs no gradient updates and no labelled training data.
The loop above, run once on a Spider question. Press play, or step through by hand: each stage lights up the part of the diagram that is working and shows what actually travels along the arrow. The question and the retrieved cases are illustrative; the five rules are the ones the generator wrote at step 100.
x_t
No label comes with it. Everything the agent knows about questions like this one is in its memory of the questions before.
R_t = retrieve(x_t, correct cases)
The question is embedded and matched against memory, restricted to entries the environment marked correct. These become the agent’s few-shot examples.
R'_t = retrieve(x_t, correct) ∪ retrieve(x_t, wrong)
Half successes, half failures, on questions like this one. The generator is never told why an answer failed; it has to work that out from the pair.
P_t = policy_generator(R'_t)
A separate call, and possibly a separate model, reads the contrastive set and writes down what separates the successes from the failures. The rules are rewritten from scratch every step; nothing is carried over.
ŷ_t = agent(x_t, R_t, P_t)
The question, the retrieved examples and the policy go into one prompt. The join through concert and the group-before-sort are what rules 2 and 3 ask for, and what the two failed cases were missing.
fb_t = environment(x_t, ŷ_t)
The query runs against the database and its result is compared with the gold result. The feedback is a 1 or a 0 with no explanation attached, and it is the only supervision DRPG ever gets.
memory.add(x_t, ŷ_t, fb_t)
Question, answer and verdict are stored together. At t + 1 the same loop runs on a memory one entry larger, and the policy is written again from whatever that retrieval turns up.
Next: t + 1, with this entry now available to both retrievals.
Spider · concert_singer · question t
Which stadium has hosted the most concerts? Give its name and capacity.
Show the stadium name and the number of concerts in each stadium.
SELECT T2.name, COUNT(*) FROM concert AS T1 JOIN stadium AS T2 ON T1.stadium_id = T2.stadium_id GROUP BY T1.stadium_idWhat are the names and capacities of stadiums that held concerts in 2014 or after?
SELECT T2.name, T2.capacity FROM concert AS T1 JOIN stadium AS T2 ON T1.stadium_id = T2.stadium_id WHERE T1.year >= 2014How many concerts were held in 2014 or 2015?
SELECT COUNT(*) FROM concert WHERE year = 2014 OR year = 2015Show the stadium name and the number of concerts in each stadium.
SELECT T2.name, COUNT(*) FROM concert AS T1 JOIN stadium AS T2 ON T1.stadium_id = T2.stadium_id GROUP BY T1.stadium_idShow name, country and age for all singers, oldest first.
SELECT name, country, age FROM singer ORDER BY age DESCShow the name and capacity of stadiums that hosted a concert in 2014.
SELECT name, capacity FROM stadium WHERE year = 2014How many concerts were held in the stadium with the largest capacity?
SELECT COUNT(*) FROM concert AS T1 JOIN stadium AS T2 ON T1.stadium_id = T2.stadium_id ORDER BY T2.capacity DESC LIMIT 1SELECT T2.name, T2.capacity FROM concert AS T1 JOIN stadium AS T2 ON T1.stadium_id = T2.stadium_id GROUP BY T1.stadium_id ORDER BY COUNT(*) DESC LIMIT 1
fb_t = 1 ✓ execution result matches
memory · n entries → n + 1
✓ How many concerts were held in 2014 or 2015?
✗ How many concerts were held in the stadium with the largest capacity?
✓ Which stadium has hosted the most concerts? Give its name and capacity.
Seven LLMs across three families (Gemini, Llama and Mistral) on six streaming benchmarks. The counts below are how many of the seven models improved under DRPG against each baseline.
| DRPG against | Spider | CoSQL | BIRD | HotpotQA | DDXPlus | DS-1000 |
|---|---|---|---|---|---|---|
| Zero-shot | 7/7 | 6/7 | 7/7 | 5/7 | 7/7 | 3/7 |
| Self-Refine | 7/7 | 6/7 | 7/7 | 5/7 | 6/7 | 4/7 |
| Self-StreamICL | 7/7 | 6/7 | 5/7 | 5/7 | 3/7 | 3/7 |
Text-to-SQL is where errors are structurally regular, so five rules cover a lot of ground. DS-1000 spans seven Python libraries whose APIs share no common failure mode, and DDXPlus has 49 diagnoses; there the policy collapses into a list of special cases and plain instance retrieval stays ahead. The paper reports full per-cell scores, significance tests, and ablations on retrieval strategy, policy continuity and the environment feedback itself.
The code release contains the DRPG agent, the six benchmark environments, the configs behind every table, and a walkthrough for running it on a model of your own.
@inproceedings{chang2026drpg,
title = {Smarter by the Moment: Environment-Driven Dynamic Policies for
Continual LLM Improvement},
author = {Chang, Ting-Wei and Chen, Po-Chun and Huang, Hen-Hsen and
Chen, Hsin-Hsi},
booktitle = {Conference on Language Modeling (COLM)},
year = {2026}
}