Фабрика задач для терминального агента: воспроизводим рекурсивное усложнение

Фабрика задач для терминального агента: воспроизводим рекурсивное усложнение
Temporary fallback cover; replace in editorial pass.

A task factory for terminal agents: reproducing recursive task extension, one verified level at a time.

Level: advanced · Reading time: 75 minutes · Updated:

Writing hard, checkable tasks for a terminal agent by hand is slow: each one needs an environment, a reference solution and tests that agree with each other. Fully synthetic tasks are cheap, but the instruction, the tests and the container often drift apart. That gives you tasks that can't be solved, or tests that pass for no real reason. This guide builds the middle path. You start from one task you have already verified. A model proposes a harder version, writes the new solution and tests, and the result has to pass a set of mechanical gates in a clean container before it's accepted. The accepted task then becomes the seed for the next level. At the end you measure how much your agent's success rate falls with each level, and how many candidate tasks the gates rejected along the way.

1. The problem in one paragraph

To train or evaluate a terminal agent you need tasks with three parts that agree: an instruction the agent reads, an environment (a container image with files and tools), and an automated checker (tests run after the agent finishes). People write good tasks but few of them. A language model writes many tasks, but its mistakes show up in predictable places: a test checks something the instruction never asked for, a required tool is missing from the image, the reference solution "passes" only because the tests are empty, or the solution was never actually run. Recursive extension helps because each new task inherits a base that already works. You only need to verify the difference, which is much easier than verifying a whole new task.

2. What you'll have at the end

3. Prerequisites

4. Task format

Each task is a directory. The agent sees only instruction.md and whatever the Dockerfile puts into the image. Tests and the reference solution are mounted after the agent finishes and are never baked into the image.

tasks/
  logreport-d0/
    task.json            # id, parent, depth, requirement ids, superseded tests
    instruction.md       # what the agent reads
    Dockerfile           # environment; must not COPY tests/ or solution.sh
    env/                 # files copied into the image (fixtures)
    solution.sh          # reference ("oracle") solution
    tests/
      test_outputs.py    # pytest; each test tagged with a requirement id

task.json for the seed:

{
  "id": "logreport-d0",
  "parent": null,
  "depth": 0,
  "requirements": {
    "R1": "Read every *.log file in /app/logs",
    "R2": "Write /app/report.json mapping HTTP status code (string) to count (int)"
  },
  "superseded_tests": []
}

The requirement map does real work. Gate G5 uses it to check that the instruction and the tests describe the same thing, and that is exactly where synthetic tasks tend to drift.

5. The concrete case: a seed task

We use a deliberately small seed so that every level can be read by a person. The point of the case is the mechanism. The difficulty of the seed doesn't matter here.

instruction.md (depth 0)

Access logs in combined format are in /app/logs/*.log.
Produce /app/report.json: a JSON object whose keys are HTTP status
codes as strings and whose values are the number of requests with
that status, summed across all files. [R1] [R2]

Dockerfile

FROM python:3.12-slim@sha256:<pin-a-digest-you-have-pulled>
RUN pip install --no-cache-dir pytest==8.3.3
WORKDIR /app
COPY env/logs /app/logs

Pin the base image by digest. Without that, "clean environment" quietly changes underneath you between runs, and drops you see across depths may really come from image drift.

solution.sh

#!/usr/bin/env bash
set -euo pipefail
python3 - <<'PY'
import json, glob, re, collections
pat = re.compile(r'"\S+ \S+ \S+" (\d{3}) ')
c = collections.Counter()
for f in sorted(glob.glob("/app/logs/*.log")):
    with open(f, encoding="utf-8", errors="replace") as fh:
        for line in fh:
            m = pat.search(line)
            if m:
                c[m.group(1)] += 1
json.dump(dict(c), open("/app/report.json", "w"), sort_keys=True)
PY

tests/test_outputs.py

import json, pathlib, pytest

REPORT = pathlib.Path("/app/report.json")

@pytest.mark.req("R2")
def test_report_exists_and_is_object():
    data = json.loads(REPORT.read_text())
    assert isinstance(data, dict)
    assert all(isinstance(k, str) and isinstance(v, int) for k, v in data.items())

@pytest.mark.req("R1", "R2")
def test_counts_match_fixture():
    # Expected counts are computed once from env/logs by a human and frozen here.
    expected = {"200": 7, "301": 1, "404": 2, "500": 1}
    assert json.loads(REPORT.read_text()) == expected

The expected counts above belong to our fixture. Compute your own from your own env/logs and freeze them in the test file. If the expected values are calculated at test time with the same logic as the solution, the test only proves that the code agrees with itself.

Register the marker in tests/conftest.py:

def pytest_configure(config):
    config.addinivalue_line("markers", "req(*ids): requirement ids covered by the test")

Before going further, check the seed by hand with the same validator you'll use for its children (section 7). A seed that fails its own gates will pass the defect on to every task built from it.

6. Extension step

The extender gets the full parent task and has to return one child task as JSON. The child adds one new capability and keeps every parent requirement unless it explicitly supersedes one. Each step changes as little as possible, so when a child fails a gate you know which addition caused it.

prompts/extend.txt

You are extending a verified terminal task. You receive the parent
instruction, Dockerfile, env file listing, reference solution, tests
and requirement map.

Produce ONE harder child task that:
- adds exactly one new capability that a competent engineer would need
  real terminal work to implement (new input format, new constraint,
  new output, error handling, performance bound);
- keeps every parent requirement unless listed in "superseded_tests"
  with a justification;
- is solvable offline inside the given image plus any packages you add
  to the Dockerfile with pinned versions;
- has deterministic tests with frozen expected values; never compute
  expected values with the same code as the solution;
- tags every test with @pytest.mark.req(...) using requirement ids, and
  mentions every requirement id in the instruction as [Rn].

Return only JSON with keys:
instruction_md, dockerfile, env_files (path -> content, text only),
solution_sh, tests_py, requirements (id -> text),
superseded_tests (list of test names), rationale.

extend.py

import json, os, pathlib, shlex, subprocess, sys

def read_task(d: pathlib.Path) -> dict:
    env = {str(p.relative_to(d)): p.read_text(errors="replace")
           for p in (d / "env").rglob("*") if p.is_file()}
    return {
        "task": json.loads((d / "task.json").read_text()),
        "instruction_md": (d / "instruction.md").read_text(),
        "dockerfile": (d / "Dockerfile").read_text(),
        "env_files": env,
        "solution_sh": (d / "solution.sh").read_text(),
        "tests_py": (d / "tests" / "test_outputs.py").read_text(),
    }

def extend(parent: pathlib.Path, child: pathlib.Path) -> None:
    prompt = pathlib.Path("prompts/extend.txt").read_text()
    payload = prompt + "\n\nPARENT:\n" + json.dumps(read_task(parent), indent=1)
    out = subprocess.run(shlex.split(os.environ["EXTENDER_CMD"]),
                         input=payload, capture_output=True, text=True,
                         timeout=900, check=True).stdout
    start, end = out.find("{"), out.rfind("}")
    spec = json.loads(out[start:end + 1])

    p = json.loads((parent / "task.json").read_text())
    child.mkdir(parents=True)
    (child / "tests").mkdir()
    for rel, content in spec["env_files"].items():
        dst = child / rel
        if not dst.resolve().is_relative_to(child.resolve()):
            raise ValueError(f"env path escapes task dir: {rel}")
        dst.parent.mkdir(parents=True, exist_ok=True)
        dst.write_text(content)
    (child / "instruction.md").write_text(spec["instruction_md"])
    (child / "Dockerfile").write_text(spec["dockerfile"])
    (child / "solution.sh").write_text(spec["solution_sh"])
    (child / "tests" / "test_outputs.py").write_text(spec["tests_py"])
    (child / "tests" / "conftest.py").write_text(
        (parent / "tests" / "conftest.py").read_text())
    (child / "task.json").write_text(json.dumps({
        "id": child.name, "parent": p["id"], "depth": p["depth"] + 1,
        "requirements": spec["requirements"],
        "superseded_tests": spec.get("superseded_tests", []),
        "rationale": spec.get("rationale", ""),
    }, indent=2))

if __name__ == "__main__":
    extend(pathlib.Path(sys.argv[1]), pathlib.Path(sys.argv[2]))

Note: the env files are rewritten by the model. If the parent fixtures are large or binary, copy them from the parent and let the model add files only. Otherwise a model that is told to keep the fixtures may still quietly change them.

7. The seven gates

An extension is accepted only if it passes every gate. Every gate runs in a fresh container built from the child's own Dockerfile, with --network none. A gate that needs the network to pass is a broken task, not an unlucky run.

GateCheckDefect it catches
G0 BuildImage builds with --no-cache; Dockerfile does not COPY tests/ or solution.shMissing tools, unpinned installs, test leakage into the image
G1 NoveltyParent solution fails the child tests"Extensions" that add nothing
G2 SolvableChild solution passes the child tests in 3 of 3 fresh runsUnsolvable tasks, flaky tests
G3 RegressionInherited parent tests (minus superseded ones) pass against the child solution in the child imageSilent loss of earlier requirements
G4 Non-vacuousA no-op solution (true) fails the child testsTests that pass on an untouched environment
G5 AlignmentEvery requirement id appears in the instruction and in at least one test marker; every test has a marker; no unknown idsInstruction–test drift
G6 LeakageInstruction does not contain expected literal values from the tests; image filesystem contains no test fileTasks answerable by copying

G1 and G4 are the gates people most often leave out, and they catch the most common defect of synthetic tasks: tests that don't actually tell a correct solution apart from a wrong one.

validate.py

import ast, hashlib, json, pathlib, re, shutil, subprocess, sys, tempfile

DOCKER_RUN = ["docker", "run", "--rm", "--network", "none",
              "--memory", "2g", "--cpus", "2", "--pids-limit", "512"]

def image_tag(task: pathlib.Path) -> str:
    h = hashlib.sha256()
    for p in sorted([task / "Dockerfile", *(task / "env").rglob("*")]):
        if p.is_file():
            h.update(str(p.relative_to(task)).encode()); h.update(p.read_bytes())
    return f"taskfactory/{task.name}:{h.hexdigest()[:12]}"

def build(task):
    df = (task / "Dockerfile").read_text()
    if re.search(r"^\s*(COPY|ADD)\s+.*(tests|solution\.sh)", df, re.M):
        raise GateError("G0", "Dockerfile copies tests or solution")
    tag = image_tag(task)
    r = subprocess.run(["docker", "build", "--no-cache", "-q", "-t", tag, str(task)],
                       capture_output=True, text=True, timeout=1800)
    if r.returncode:
        raise GateError("G0", r.stderr[-2000:])
    return tag

def run(tag, solution: pathlib.Path, tests_dir: pathlib.Path, deselect=()):
    """Run solution, then tests, in one fresh container. Returns (passed, junit_text)."""
    with tempfile.TemporaryDirectory() as tmp:
        out = pathlib.Path(tmp)
        args = " ".join(f"--deselect /tests/test_outputs.py::{t}" for t in deselect)
        script = (f"bash /solution.sh >/out/solution.log 2>&1; "
                  f"pytest -q -p no:cacheprovider {args} --junitxml=/out/junit.xml /tests "
                  f">/out/pytest.log 2>&1")
        r = subprocess.run(DOCKER_RUN + [
            "-v", f"{solution.resolve()}:/solution.sh:ro",
            "-v", f"{tests_dir.resolve()}:/tests:ro",
            "-v", f"{out}:/out", tag, "bash", "-c", script],
            capture_output=True, text=True, timeout=1200)
        junit = (out / "junit.xml").read_text() if (out / "junit.xml").exists() else ""
        return r.returncode == 0, junit

class GateError(Exception):
    def __init__(self, gate, detail): super().__init__(f"{gate}: {detail}"); self.gate = gate

def markers(tests_py: str):
    """Map test function name -> set of requirement ids from @pytest.mark.req(...)."""
    tree, found = ast.parse(tests_py), {}
    for node in ast.walk(tree):
        if isinstance(node, ast.FunctionDef) and node.name.startswith("test_"):
            ids = set()
            for d in node.decorator_list:
                if (isinstance(d, ast.Call) and getattr(d.func, "attr", "") == "req"):
                    ids |= {a.value for a in d.args if isinstance(a, ast.Constant)}
            found[node.name] = ids
    return found

def validate(child: pathlib.Path, parent: pathlib.Path | None) -> dict:
    meta = json.loads((child / "task.json").read_text())
    tests_py = (child / "tests" / "test_outputs.py").read_text()
    instr = (child / "instruction.md").read_text()
    report = {"task": meta["id"], "gates": {}}

    # G5 alignment (cheap, run first)
    req = set(meta["requirements"])
    tmap = markers(tests_py)
    in_instr = set(re.findall(r"\[(R\d+)\]", instr))
    tagged = set().union(*tmap.values()) if tmap else set()
    problems = []
    if req - in_instr: problems.append(f"not in instruction: {sorted(req - in_instr)}")
    if req - tagged: problems.append(f"no test: {sorted(req - tagged)}")
    if tagged - req: problems.append(f"unknown ids in tests: {sorted(tagged - req)}")
    if [t for t, ids in tmap.items() if not ids]: problems.append("untagged tests")
    if problems: raise GateError("G5", "; ".join(problems))
    report["gates"]["G5"] = "pass"

    # G6 leakage, part 1: long literals from tests must not appear in the instruction
    lits = {n.value for n in ast.walk(ast.parse(tests_py))
            if isinstance(n, ast.Constant) and isinstance(n.value, str) and len(n.value) >= 12}
    leaked = [l for l in lits if l in instr and not l.startswith("/")]
    if leaked: raise GateError("G6", f"instruction contains test literals: {leaked[:3]}")

    tag = build(child); report["gates"]["G0"] = "pass"; report["image"] = tag

    # G6 leakage, part 2: nothing named like tests inside the image
    r = subprocess.run(DOCKER_RUN + [tag, "bash", "-c",
        "find / -xdev -name 'test_outputs.py' -o -xdev -name 'solution.sh' 2>/dev/null"],
        capture_output=True, text=True, timeout=300)
    if r.stdout.strip(): raise GateError("G6", f"found in image: {r.stdout.strip()[:200]}")
    report["gates"]["G6"] = "pass"

    tests = child / "tests"
    # G2 solvable, 3 fresh runs
    for i in range(3):
        ok, junit = run(tag, child / "solution.sh", tests)
        if not ok: raise GateError("G2", f"reference failed on run {i+1}")
    report["gates"]["G2"] = "pass"

    # G4 non-vacuous
    with tempfile.NamedTemporaryFile("w", suffix=".sh", delete=False) as f:
        f.write("#!/usr/bin/env bash\ntrue\n"); noop = pathlib.Path(f.name)
    ok, _ = run(tag, noop, tests)
    if ok: raise GateError("G4", "no-op solution passes")
    report["gates"]["G4"] = "pass"

    if parent is not None:
        # G1 novelty: parent solution must fail child tests
        ok, _ = run(tag, parent / "solution.sh", tests)
        if ok: raise GateError("G1", "parent solution already passes child tests")
        report["gates"]["G1"] = "pass"

        # G3 regression: parent tests (minus superseded) pass with child solution
        ok, _ = run(tag, child / "solution.sh", parent / "tests",
                    deselect=meta.get("superseded_tests", []))
        if not ok: raise GateError("G3", "child solution breaks inherited tests")
        report["gates"]["G3"] = "pass"
    return report

if __name__ == "__main__":
    child = pathlib.Path(sys.argv[1])
    parent = pathlib.Path(sys.argv[2]) if len(sys.argv) > 2 else None
    try:
        print(json.dumps(validate(child, parent), indent=2))
    except GateError as e:
        print(json.dumps({"task": child.name, "rejected": e.gate, "detail": str(e)}))
        sys.exit(1)

Things to know about this code:

8. Recursive growth loop

Each accepted child becomes the parent of the next attempt. A rejected child is logged and retried, up to a fixed number of attempts per level. If a level runs out of attempts, the lineage stops there. You don't skip ahead from an unvalidated task.

grow.py

import json, pathlib, shutil, subprocess, sys, time
from extend import extend
from validate import validate, GateError

def grow(seed: pathlib.Path, max_depth=3, attempts=4, log="runs/grow.jsonl"):
    pathlib.Path(log).parent.mkdir(exist_ok=True)
    parent = seed
    for depth in range(1, max_depth + 1):
        accepted = None
        for a in range(attempts):
            child = seed.parent / f"{seed.name.rsplit('-d', 1)[0]}-d{depth}-a{a}"
            shutil.rmtree(child, ignore_errors=True)
            rec = {"ts": time.time(), "parent": parent.name, "child": child.name,
                   "depth": depth, "attempt": a}
            try:
                extend(parent, child)
                rec |= validate(child, parent); rec["status"] = "accepted"
                accepted = child
            except GateError as e:
                rec |= {"status": "rejected", "gate": e.gate, "detail": str(e)[:500]}
            except Exception as e:  # malformed JSON, timeouts, path errors
                rec |= {"status": "error", "detail": repr(e)[:500]}
            with open(log, "a") as fh: fh.write(json.dumps(rec) + "\n")
            if accepted: break
        if not accepted:
            print(f"lineage stopped at depth {depth}"); return
        parent = accepted

if __name__ == "__main__":
    grow(pathlib.Path(sys.argv[1]))

Run it for several seeds, not one. A single lineage tells you about one chain of model decisions. To say anything about "recursive extension" in general, you need several independent lineages, each with its own seed.

export EXTENDER_CMD="your-model-cli --max-tokens 16000"
python validate.py tasks/logreport-d0          # the seed must pass its own gates
for s in tasks/*-d0; do python grow.py "$s"; done
jq -r 'select(.status!="accepted") | [.depth,.gate // "error"] | @tsv' runs/grow.jsonl \
  | sort | uniq -c

What a lineage for our seed might look like (illustrative)

Here is the kind of chain a model may propose. We wrote it ourselves as an example. It is not output from a run:

  1. d1: rotated logs access.log.1.gz, access.log.2.gz must also be counted [R3]. Fixtures gain gzip files, and the parent solution undercounts, so G1 fires as intended.
  2. d2: requests from IPs listed in /app/config/healthcheck_ips.txt (one per line, # comments allowed) are excluded [R4]. Parent test test_counts_match_fixture becomes superseded because its frozen counts change.
  3. d3: also write /app/report_hourly.csv with header hour_utc,status,count, timestamps normalized from mixed offsets to UTC, rows sorted [R5]. A correct solution needs timezone handling, CSV formatting and the earlier filters all together.

Notice what happens at d2. The superseded test is legitimate, but the replacement test has to cover R1 and R2 again with new frozen values. Otherwise G5 still passes (R1 might be covered by another test), and part of the base behaviour is no longer checked. Review the superseded list together with the new test set.

9. Measuring the success drop

Agent contract

$AGENT_CMD receives the image tag and the instruction path. It must start a container from that image (network policy is your call, but write it down), let the agent work, and then save the final container filesystem as a new image that the tests can run against. A simple, honest way to do this is with docker commit:

#!/usr/bin/env bash
# agent_episode.sh IMAGE INSTRUCTION OUT_TAG
set -euo pipefail
cid=$(docker run -d --network none "$1" sleep infinity)
your-agent --container "$cid" --instruction "$2" --max-steps 60 --timeout 1800 || true
docker commit "$cid" "$3" >/dev/null
docker rm -f "$cid" >/dev/null

The tests then run against OUT_TAG with a no-op solution. The agent's work is the solution. Keep the step budget, timeout and agent model identical across depths. If you change any of them between depths, you are measuring that change and not the task difficulty.

measure.py

import json, math, pathlib, subprocess, sys, collections
from validate import run, image_tag

NOOP = pathlib.Path("noop.sh"); NOOP.write_text("#!/usr/bin/env bash\ntrue\n")

def wilson(k, n, z=1.96):
    if n == 0: return (float("nan"),) * 3
    p = k / n; d = 1 + z*z/n
    c = (p + z*z/(2*n)) / d
    h = z * math.sqrt(p*(1-p)/n + z*z/(4*n*n)) / d
    return p, max(0, c - h), min(1, c + h)

def episode(task, i):
    tag = image_tag(task); out = f"{tag}-ep{i}".replace(":", "-")
    subprocess.run(["bash", "agent_episode.sh", tag, str(task / "instruction.md"), out],
                   check=True, timeout=3600)
    ok, _ = run(out, NOOP, task / "tests")
    subprocess.run(["docker", "rmi", "-f", out], capture_output=True)
    return ok

def main(tasks_glob="tasks/*", n=20, log="runs/measure.jsonl"):
    res = collections.defaultdict(list)
    for task in sorted(pathlib.Path().glob(tasks_glob)):
        meta = json.loads((task / "task.json").read_text())
        if meta.get("status") == "rejected": continue
        for i in range(n):
            ok = episode(task, i)
            res[(meta["id"], meta["depth"])].append(ok)
            with open(log, "a") as fh:
                fh.write(json.dumps({"task": meta["id"], "depth": meta["depth"],
                                     "episode": i, "pass": ok}) + "\n")
    by_depth = collections.defaultdict(lambda: [0, 0])
    for (_, d), oks in res.items():
        by_depth[d][0] += sum(oks); by_depth[d][1] += len(oks)
    for d in sorted(by_depth):
        k, m = by_depth[d]; p, lo, hi = wilson(k, m)
        print(f"depth {d}: {k}/{m} = {p:.2f}  95% CI [{lo:.2f}, {hi:.2f}]")

if __name__ == "__main__":
    main(*sys.argv[1:2])

measure.py only scans directories that are accepted tasks. In practice, move rejected attempts out of tasks/ (or mark them) when grow.py finishes, so they never get into the measurement.

How to read the numbers

Results template (fill in from your runs)

DepthAccepted tasksAttemptsRejections by gateEpisodesPass rate [95% CI]
0—————
1—————
2—————
3—————

Also record: extender model and version, agent model and version, base image digest, step budget, timeout, network policy, N, and the commit of your factory code. Without these, nobody can reproduce your table, including you three months from now.

10. Verifying the factory itself

The gates guard the tasks. Something has to guard the gates. Before you trust any measurement, run these checks:

  1. Planted defects. Take an accepted child and break it in known ways: empty the tests to assert True (G4 must fire), make the Dockerfile COPY tests /tests (G0 must fire), remove an [R3] tag from the instruction (G5 must fire), copy the parent's tests unchanged (G1 must fire), insert random.random() < 0.5 into an assertion (G2 should fire most of the time). Write these as a small test suite for validate.py and run it on every change.
  2. Human audit sample. For each depth, read at least a few accepted tasks end to end: is the instruction sufficient to pass the tests without guessing? Track the share where the answer is "no". The gates check consistency, not fairness. A task can pass all seven gates and still depend on an unstated output format.
  3. Superseded-test cap. Reject a child automatically if it supersedes more than, say, half of the inherited tests. Pick the threshold yourself and write it down. A child that throws out most of its inheritance is a new task, not an extension, and it skips the cheap differential verification the whole approach relies on.
  4. Re-run from scratch. Delete all built images and re-run validate.py on every accepted task. Any task that now fails was accepted because of cache or timing, and must be removed.

11. Failure cases you should expect

12. Limitations

13. Short checklist

  1. Seed task with requirement ids, frozen expected values, pinned base image, and it passes validate.py.
  2. Extender prompt with strict JSON output; fixtures copied from the parent, additions only.
  3. Seven gates in fresh containers with no network; planted-defect suite for the validator.
  4. Several independent lineages, a fixed number of attempts per level, every rejection logged with its gate.
  5. Fixed agent settings across depths, N episodes per task, Wilson intervals, paired per-lineage drops, gate yield.
  6. Human audit sample per depth, superseded-test cap, held-out lineages for evaluation.
  7. Record models, versions, digests, budgets and code commit next to every results table.

More hands-on material: all guides and the glossary of terms.

We publish what works for us—and implement the same solutions for your business. We design AI automation, Telegram bots, chats, and AI agents for real-world processes. Discuss your project →