Строгость или исправление: тестируем парсер аргументов LLM-инструментов

Strict or repair: testing an argument parser for LLM tools.
A model asked to refund $19.99 sends "amount_cents": "1999". The data is right; the envelope is wrong — a string where the schema says integer. A strict parser rejects it and burns a round trip. A "helpful" parser that runs int(float(x)) on everything accepts it, and on the next call it also accepts "19.99" as 19 cents, silently. This article builds the layer in between: a small, tested gate that repairs representation errors with exactly one possible meaning, rejects everything with two or more, and records every repair so you can see how often your model needs help.
By the end you will have four files — a parser, a tool spec, a pytest suite and a replay script — that you can run locally in a few minutes and then adapt to your own tools. Nothing here depends on a particular model provider.
What this article does and does not claim
- Verified by you, not by us: the code and tests below are complete and intended to pass on Python 3.10+. We do not publish a pass count or timings; run them and trust your own output.
- Examples, not incidents: the refund tool, order IDs and recorded calls are illustrative. There is no customer data here, and no measured error rates for any specific model.
- No external benchmarks: where we say a pattern is common, read that as "a pattern you should test for", not as a statistic.
The problem: right data, wrong envelope
In tool calling, an LLM emits a tool name and a JSON object of arguments, and your code executes the tool. Even with a JSON Schema attached to the tool definition, what arrives at your server is not guaranteed to conform. Unless your provider enforces structured output with constrained decoding — and even then, if you accept calls from more than one model or from replayed logs — you will see things like:
- numbers as strings:
"1999"; - integral floats:
1999.0; - booleans as strings:
"true","False"; - enum values in the wrong case:
"Damaged"; - a single string where an array is expected:
"fragile"instead of["fragile"]; - the whole payload wrapped in a Markdown code fence, or JSON-encoded twice;
nullfor an optional field instead of omitting it.
And next to those, values that only look like the same kind of mistake:
"19.99"for a field in cents — dollars or a typo?"1,999"— one thousand nine hundred ninety-nine, or a European decimal?"03/04/2026"— March 4 or April 3?"yes"or1for a boolean that triggers a customer email;"fragile, gift"— one tag or two?"amount": 1999instead of"amount_cents"— a typo with an obvious fix that you should still not apply automatically.
The difference between the two lists is the whole design: repair when the value has exactly one reasonable interpretation; reject when it has more than one, even if one of them is very likely. A rejection costs one round trip in the agent loop. A wrong repair costs whatever the tool does — here, money.
The concrete case: issue_refund
We use one tool throughout. It has the field types that cause most trouble in practice: an ID with a format, an integer with bounds, a boolean with side effects, an enum, a list and a date.
| Field | Type | Rule |
|---|---|---|
order_id | string | required, matches ord_[A-Za-z0-9]{8} |
amount_cents | integer | required, 1 to 50000 |
notify_customer | boolean | optional, default false |
reason | enum | required: damaged, late, duplicate, other |
tags | array of strings | optional, default [] |
effective_date | date | optional, ISO YYYY-MM-DD only |
The policy, written down before the code
Write the policy as a table first. If you cannot justify a repair in one sentence, it is not unambiguous.
| Input | Decision | Why |
|---|---|---|
Payload in a ```json fence | Repair: strip_code_fence | The fence is formatting; content is unchanged. |
| Payload JSON-encoded twice | Repair once: unwrap_double_encoded | One level is a known serialization slip. Three levels is a bug elsewhere — reject. |
| Prose around JSON | Reject: invalid_json | Extracting "the JSON part" means guessing which part. |
| Duplicate keys | Reject: duplicate_key | Standard parsers keep the last value silently. Which one did the model mean? |
NaN, Infinity | Reject: non_finite_number | Not valid JSON; Python's parser accepts them by default. |
"1999" for an integer | Repair: numeric_string_to_int | Plain base-10 digits have one meaning. |
1999.0 for an integer | Repair: integral_float_to_int | No fractional part; no information lost. |
19.99, "19.99", "1,999", "1e3", "01999", " 1999" | Reject | Unit confusion, locale, notation or truncation — more than one reading. |
true for an integer | Reject: type | In Python bool is a subclass of int; an unguarded check lets true become 1. |
"true" / "FALSE" for a boolean | Repair: string_to_bool | The literal words, case-insensitively, have one meaning. |
"yes", 1, "0" for a boolean | Reject | Each is a convention, not a value. |
"Damaged" for an enum | Repair if exactly one case-insensitive match | Unique match is unambiguous; two matches are not. |
"fragile" for an array | Repair: scalar_to_list | One item, one list. |
"fragile, gift" or '["fragile"]' for an array | Reject: ambiguous_list | Splitting or parsing a string is a guess. |
12345678 for a string ID | Reject: type | Converting numbers to IDs loses leading zeros and invites confusion with amounts. |
null for an optional field | Repair: null_to_absent | Treated as "not provided"; the default applies. |
null or missing for a required field | Reject: missing | Never invent required values. |
Unknown field amount | Reject with a hint | Suggest amount_cents to the model; do not rename silently. |
"03/04/2026" for a date | Reject: ambiguous_date | Day/month order depends on locale. |
Two extra rules make the policy operable:
- Every repair is recorded with path, rule, before and after values. A repair you cannot see is the same as a silent coercion.
- Any field can opt out of repairs with
"strict": True. For a high-risk field you may prefer "the model must get it exactly right" even for"1999".
Step 1. Set up the project
mkdir toolargs-lab && cd toolargs-lab
python3 -m venv .venv
. .venv/bin/activate
python -m pip install pytest
python --version # expect 3.10 or newer
The parser uses only the standard library. pytest is needed for the tests only. The module is named toolargs.py rather than anything with "argparse" in it, to avoid shadowing the standard library module.
Step 2. The parser: toolargs.py
The structure is simple: decode the payload into a dict, reject unknown fields, then run each declared field through a type-specific coercer that either returns a clean value (optionally appending a Repair) or raises ArgumentError with a stable machine-readable code.
"""Argument gate for LLM tool calls: repair unambiguous representation errors, reject the rest."""
from __future__ import annotations
import copy
import difflib
import json
import re
from dataclasses import dataclass, field
from datetime import date
from typing import Any
class ArgumentError(ValueError):
def __init__(self, path: str, code: str, message: str):
super().__init__(f"{path}: {code}: {message}")
self.path = path
self.code = code
self.message = message
@dataclass(frozen=True)
class Repair:
path: str
rule: str
before: Any
after: Any
@dataclass
class Parsed:
args: dict[str, Any]
repairs: list[Repair] = field(default_factory=list)
FENCE_RE = re.compile(r"\A```(?:json)?[ \t]*\n(.*)\n```\s*\Z", re.DOTALL)
INT_RE = re.compile(r"-?(?:0|[1-9][0-9]*)")
ISO_DATE_RE = re.compile(r"[0-9]{4}-[0-9]{2}-[0-9]{2}")
LIST_SEPARATORS = (",", ";", "\n")
# --- payload decoding -------------------------------------------------------
def _no_duplicates(pairs):
obj = {}
for key, value in pairs:
if key in obj:
raise ArgumentError(f"$.{key}", "duplicate_key", "key appears more than once")
obj[key] = value
return obj
def _reject_constant(name):
raise ArgumentError("$", "non_finite_number", f"{name} is not allowed")
def _loads(text: str) -> Any:
return json.loads(text, object_pairs_hook=_no_duplicates, parse_constant=_reject_constant)
def decode_payload(raw: Any, repairs: list[Repair]) -> dict[str, Any]:
if isinstance(raw, dict):
return raw
if not isinstance(raw, str):
raise ArgumentError("$", "not_object", f"expected JSON object, got {type(raw).__name__}")
text = raw.strip()
fence = FENCE_RE.match(text)
if fence:
repairs.append(Repair("$", "strip_code_fence", raw, fence.group(1)))
text = fence.group(1)
try:
value = _loads(text)
except json.JSONDecodeError as err:
raise ArgumentError("$", "invalid_json", err.msg) from None
if isinstance(value, str):
# Exactly one level of double encoding is repaired; deeper nesting is rejected below.
try:
inner = _loads(value)
except json.JSONDecodeError:
raise ArgumentError("$", "not_object", "payload is a JSON string") from None
if isinstance(inner, dict):
repairs.append(Repair("$", "unwrap_double_encoded", value, inner))
value = inner
if not isinstance(value, dict):
raise ArgumentError("$", "not_object", f"expected JSON object, got {type(value).__name__}")
return value
# --- field coercers -----------------------------------------------------------
def _type_error(path: str, expected: str, value: Any) -> ArgumentError:
return ArgumentError(path, "type", f"expected {expected}, got {type(value).__name__}")
def coerce_string(path, value, spec, repairs):
if not isinstance(value, str):
raise _type_error(path, "string", value)
pattern = spec.get("pattern")
if pattern and not re.fullmatch(pattern, value):
raise ArgumentError(path, "pattern", f"{value!r} does not match {pattern}")
return value
def coerce_integer(path, value, spec, repairs):
if isinstance(value, bool):
raise _type_error(path, "integer", value)
if isinstance(value, int):
out = value
elif isinstance(value, float):
if not value.is_integer():
raise ArgumentError(path, "lossy_number", f"{value!r} has a fractional part")
out = int(value)
repairs.append(Repair(path, "integral_float_to_int", value, out))
elif isinstance(value, str):
if not INT_RE.fullmatch(value):
raise ArgumentError(path, "ambiguous_integer", f"{value!r} is not a plain base-10 integer")
out = int(value)
repairs.append(Repair(path, "numeric_string_to_int", value, out))
else:
raise _type_error(path, "integer", value)
low, high = spec.get("minimum"), spec.get("maximum")
if (low is not None and out < low) or (high is not None and out > high):
raise ArgumentError(path, "out_of_range", f"{out} is outside [{low}, {high}]")
return out
def coerce_boolean(path, value, spec, repairs):
if isinstance(value, bool):
return value
if isinstance(value, str):
lowered = value.lower()
if lowered in ("true", "false"):
out = lowered == "true"
repairs.append(Repair(path, "string_to_bool", value, out))
return out
raise ArgumentError(path, "ambiguous_boolean", f"{value!r} is not 'true' or 'false'")
raise _type_error(path, "boolean", value)
def coerce_enum(path, value, spec, repairs):
if not isinstance(value, str):
raise _type_error(path, "string enum", value)
values = spec["values"]
if value in values:
return value
matches = [v for v in values if v.casefold() == value.casefold()]
if len(matches) == 1:
repairs.append(Repair(path, "enum_casefold", value, matches[0]))
return matches[0]
raise ArgumentError(path, "enum", f"{value!r} is not one of {values}")
def coerce_string_list(path, value, spec, repairs):
if isinstance(value, list):
for i, item in enumerate(value):
if not isinstance(item, str):
raise _type_error(f"{path}[{i}]", "string", item)
return list(value)
if isinstance(value, str):
if (not value.strip()
or any(sep in value for sep in LIST_SEPARATORS)
or value.lstrip().startswith("[")):
raise ArgumentError(path, "ambiguous_list", f"{value!r} cannot be turned into a list safely")
repairs.append(Repair(path, "scalar_to_list", value, [value]))
return [value]
raise _type_error(path, "array of strings", value)
def coerce_date(path, value, spec, repairs):
if not isinstance(value, str):
raise _type_error(path, "date string", value)
if not ISO_DATE_RE.fullmatch(value):
raise ArgumentError(path, "ambiguous_date", f"{value!r} is not YYYY-MM-DD")
try:
return date.fromisoformat(value).isoformat()
except ValueError:
raise ArgumentError(path, "invalid_date", f"{value!r} is not a real calendar date") from None
COERCERS = {
"string": coerce_string,
"integer": coerce_integer,
"boolean": coerce_boolean,
"enum": coerce_enum,
"array_string": coerce_string_list,
"date": coerce_date,
}
# --- entry points -------------------------------------------------------------
def parse_arguments(raw: Any, spec: dict[str, dict[str, Any]]) -> Parsed:
repairs: list[Repair] = []
data = decode_payload(raw, repairs)
unknown = sorted(set(data) - set(spec))
if unknown:
name = unknown[0]
hint = difflib.get_close_matches(name, list(spec), n=1)
message = "unknown field" + (f"; did you mean {hint[0]!r}?" if hint else "")
raise ArgumentError(f"$.{name}", "unknown_field", message)
out: dict[str, Any] = {}
for name, fspec in spec.items():
path = f"$.{name}"
if name not in data or data[name] is None:
if fspec.get("required"):
raise ArgumentError(path, "missing", "required field is absent or null")
if name in data:
repairs.append(Repair(path, "null_to_absent", None, fspec.get("default")))
if "default" in fspec:
out[name] = copy.deepcopy(fspec["default"])
continue
before = len(repairs)
out[name] = COERCERS[fspec["type"]](path, data[name], fspec, repairs)
if fspec.get("strict") and len(repairs) > before:
rule = repairs[before].rule
raise ArgumentError(path, "repair_forbidden", f"{rule} is not allowed for this field")
return Parsed(out, repairs)
def error_for_model(err: ArgumentError) -> str:
"""Tool result to send back to the model so it can correct the call."""
return json.dumps(
{"ok": False, "error": err.code, "path": err.path, "message": err.message},
ensure_ascii=False,
)
Points worth reading twice:
object_pairs_hook=_no_duplicates— without it,{"amount_cents": 1999, "amount_cents": 199999}parses to the larger value with no warning.parse_constant=_reject_constant— Python'sjsonacceptsNaNandInfinityby default; this turns them into a rejection.- The
isinstance(value, bool)check comes beforeisinstance(value, int). Swap them andtruebecomes a 1-cent refund. - The unknown-field hint goes back to the model. The parser itself never renames
amounttoamount_cents. - Error codes are stable strings. Tests, dashboards and the model's correction prompt all depend on them; treat renaming a code as a breaking change.
Step 3. The tool spec: refund_spec.py
SPEC = {
"order_id": {"type": "string", "required": True, "pattern": r"ord_[A-Za-z0-9]{8}"},
"amount_cents": {"type": "integer", "required": True, "minimum": 1, "maximum": 50000},
"notify_customer": {"type": "boolean", "default": False},
"reason": {"type": "enum", "required": True,
"values": ["damaged", "late", "duplicate", "other"]},
"tags": {"type": "array_string", "default": []},
"effective_date": {"type": "date"},
}
This spec is deliberately smaller than full JSON Schema. If you already publish a JSON Schema to the model, generate this dict from it (or extend the parser to read your schema directly) so that the model's contract and the server's gate cannot drift apart. Keeping two hand-written copies is how fields get renamed in one place only.
Step 4. The tests: test_toolargs.py
The suite has four groups: a clean call, repairs that must happen, rejections that must happen, and properties that must hold across all repairs. The last group is the one that catches regressions when someone "just adds one more coercion".
import json
import pytest
from refund_spec import SPEC
from toolargs import ArgumentError, Repair, error_for_model, parse_arguments
BASE = {"order_id": "ord_A1b2C3d4", "amount_cents": 1999, "reason": "damaged"}
def payload(**overrides):
return json.dumps({**BASE, **overrides})
def rejects(raw, spec=SPEC):
with pytest.raises(ArgumentError) as exc:
parse_arguments(raw, spec)
return exc.value
# --- clean call ---------------------------------------------------------------
def test_clean_call_has_no_repairs():
result = parse_arguments(payload(), SPEC)
assert result.args == {
"order_id": "ord_A1b2C3d4",
"amount_cents": 1999,
"notify_customer": False,
"reason": "damaged",
"tags": [],
}
assert result.repairs == []
def test_defaults_are_not_shared_between_calls():
first = parse_arguments(payload(), SPEC)
first.args["tags"].append("mutated")
second = parse_arguments(payload(), SPEC)
assert second.args["tags"] == []
# --- repairs ------------------------------------------------------------------
REPAIR_CASES = [
("amount_cents", "1999", 1999, "numeric_string_to_int"),
("amount_cents", 1999.0, 1999, "integral_float_to_int"),
("notify_customer", "true", True, "string_to_bool"),
("notify_customer", "FALSE", False, "string_to_bool"),
("reason", "Damaged", "damaged", "enum_casefold"),
("tags", "fragile", ["fragile"], "scalar_to_list"),
("notify_customer", None, False, "null_to_absent"),
]
@pytest.mark.parametrize("field,value,expected,rule", REPAIR_CASES)
def test_unambiguous_values_are_repaired(field, value, expected, rule):
result = parse_arguments(payload(**{field: value}), SPEC)
assert result.args[field] == expected
assert type(result.args[field]) is type(expected)
assert [r.rule for r in result.repairs] == [rule]
assert result.repairs[0].path == f"$.{field}"
def test_repair_records_before_and_after():
result = parse_arguments(payload(amount_cents="1999"), SPEC)
assert result.repairs == [Repair("$.amount_cents", "numeric_string_to_int", "1999", 1999)]
def test_code_fence_is_stripped():
raw = "```json\n" + json.dumps(BASE) + "\n```"
result = parse_arguments(raw, SPEC)
assert result.args["amount_cents"] == 1999
assert [r.rule for r in result.repairs] == ["strip_code_fence"]
def test_double_encoding_is_unwrapped_once():
result = parse_arguments(json.dumps(json.dumps(BASE)), SPEC)
assert result.args["order_id"] == "ord_A1b2C3d4"
assert [r.rule for r in result.repairs] == ["unwrap_double_encoded"]
def test_triple_encoding_is_rejected():
raw = json.dumps(json.dumps(json.dumps(BASE)))
assert rejects(raw).code == "not_object"
def test_null_optional_without_default_is_dropped():
result = parse_arguments(payload(effective_date=None), SPEC)
assert "effective_date" not in result.args
assert [r.rule for r in result.repairs] == ["null_to_absent"]
# --- rejections ---------------------------------------------------------------
@pytest.mark.parametrize("field,value,code", [
("amount_cents", "19.99", "ambiguous_integer"),
("amount_cents", "1,999", "ambiguous_integer"),
("amount_cents", "1e3", "ambiguous_integer"),
("amount_cents", " 1999", "ambiguous_integer"),
("amount_cents", "01999", "ambiguous_integer"),
("amount_cents", 19.99, "lossy_number"),
("amount_cents", True, "type"),
("amount_cents", 0, "out_of_range"),
("amount_cents", 50001, "out_of_range"),
("notify_customer", "yes", "ambiguous_boolean"),
("notify_customer", 1, "type"),
("order_id", 12345678, "type"),
("order_id", "ORD_A1b2C3d4", "pattern"),
("reason", "broken", "enum"),
("reason", None, "missing"),
("tags", "fragile, gift", "ambiguous_list"),
("tags", '["fragile"]', "ambiguous_list"),
("effective_date", "03/04/2026", "ambiguous_date"),
("effective_date", "2026-02-30", "invalid_date"),
])
def test_ambiguous_values_are_rejected(field, value, code):
err = rejects(payload(**{field: value}))
assert err.code == code
assert err.path == f"$.{field}"
def test_non_string_list_item_reports_exact_path():
err = rejects(payload(tags=["fragile", 3]))
assert (err.code, err.path) == ("type", "$.tags[1]")
def test_missing_required_field():
data = dict(BASE)
del data["reason"]
err = rejects(json.dumps(data))
assert (err.code, err.path) == ("missing", "$.reason")
def test_unknown_field_gets_a_hint_but_is_not_renamed():
data = {**BASE, "amount": 1999}
err = rejects(json.dumps(data))
assert err.code == "unknown_field"
assert "amount_cents" in err.message
def test_duplicate_keys_are_rejected():
raw = '{"order_id":"ord_A1b2C3d4","amount_cents":1999,"amount_cents":49999,"reason":"late"}'
assert rejects(raw).code == "duplicate_key"
def test_nan_is_rejected():
raw = '{"order_id":"ord_A1b2C3d4","amount_cents":NaN,"reason":"late"}'
assert rejects(raw).code == "non_finite_number"
@pytest.mark.parametrize("raw,code", [
("not json", "invalid_json"),
('{"order_id": "ord_A1b2C3d4"', "invalid_json"),
("Here are the arguments: " + json.dumps(BASE), "invalid_json"),
("[]", "not_object"),
(42, "not_object"),
])
def test_bad_payloads(raw, code):
assert rejects(raw).code == code
# --- strict fields ------------------------------------------------------------
STRICT_SPEC = {**SPEC, "amount_cents": {**SPEC["amount_cents"], "strict": True}}
def test_strict_field_refuses_repairs():
err = rejects(payload(amount_cents="1999"), STRICT_SPEC)
assert (err.code, err.path) == ("repair_forbidden", "$.amount_cents")
def test_strict_field_accepts_exact_values():
assert parse_arguments(payload(), STRICT_SPEC).args["amount_cents"] == 1999
# --- properties ---------------------------------------------------------------
@pytest.mark.parametrize("field,value,expected,rule", REPAIR_CASES)
def test_repair_is_idempotent(field, value, expected, rule):
first = parse_arguments(payload(**{field: value}), SPEC)
second = parse_arguments(json.dumps(first.args), SPEC)
assert second.args == first.args
assert second.repairs == []
def test_enum_values_are_casefold_unique():
for name, fspec in SPEC.items():
if fspec["type"] == "enum":
folded = [v.casefold() for v in fspec["values"]]
assert len(set(folded)) == len(folded), name
def test_error_message_for_model_is_json():
err = rejects(payload(amount_cents="19.99"))
body = json.loads(error_for_model(err))
assert body == {
"ok": False,
"error": "ambiguous_integer",
"path": "$.amount_cents",
"message": err.message,
}
The idempotence property is the most useful guard you can add: once repaired, a value must be canonical. If a future change makes the parser produce output that it would itself repair again, the coercion is not a normalization — it is a transformation, and it deserves a design review.
Step 5. Run and verify
python -m pytest -q
Expected: every test passes, no failures and no errors. If something fails, read the assertion before changing code — the most common causes when adapting this to your own tools are:
- a coercer checked
intbeforebool; - a new enum contains two values that differ only in case (the lint test catches this);
- a "repair" that is not idempotent, for example trimming whitespace in one place and not another.
To see what pytest actually checked, list the collected cases:
python -m pytest --collect-only -q
Step 6. Replay recorded calls
Unit tests prove the policy. Replay tells you how the policy meets real traffic: how many calls were clean, how many needed a repair, and which rejections dominate. Create replay.py:
import collections
import json
import sys
from refund_spec import SPEC
from toolargs import ArgumentError, parse_arguments
counts = collections.Counter()
with open(sys.argv[1], encoding="utf-8") as fh:
for line in fh:
record = json.loads(line)
try:
parsed = parse_arguments(record["arguments"], SPEC)
except ArgumentError as err:
counts[f"reject:{err.code}"] += 1
continue
counts["ok_repaired" if parsed.repairs else "ok_clean"] += 1
for repair in parsed.repairs:
counts[f"repair:{repair.rule}"] += 1
for key, value in sorted(counts.items()):
print(f"{value:6d} {key}")
Create a small sample file, calls.jsonl. These five lines are synthetic; in your project they would come from your tool-call logs (redacted — see limitations).
{"id": 1, "arguments": "{\"order_id\":\"ord_A1b2C3d4\",\"amount_cents\":1999,\"reason\":\"late\"}"}
{"id": 2, "arguments": "{\"order_id\":\"ord_A1b2C3d4\",\"amount_cents\":\"1999\",\"reason\":\"late\"}"}
{"id": 3, "arguments": "{\"order_id\":\"ord_A1b2C3d4\",\"amount_cents\":\"19.99\",\"reason\":\"late\"}"}
{"id": 4, "arguments": "```json\n{\"order_id\":\"ord_A1b2C3d4\",\"amount_cents\":1999,\"reason\":\"late\"}\n```"}
{"id": 5, "arguments": {"order_id": "ord_A1b2C3d4", "amount": 1999, "reason": "late"}}
Line 5 holds arguments as an already-decoded object, which is what many SDKs hand you. Run:
python replay.py calls.jsonl
Working through the policy by hand, the output should be:
1 ok_clean
2 ok_repaired
1 reject:ambiguous_integer
1 reject:unknown_field
1 repair:numeric_string_to_int
1 repair:strip_code_fence
If your output differs, the parser and this article disagree — check which line moved before trusting either.
On real logs, the numbers to watch are the repair rate per rule and the rejection rate per code, broken down by model version and prompt version. A sudden rise in numeric_string_to_int after a prompt change is a regression in the prompt, even though every call still "works". That is the practical value of logging repairs instead of hiding them.
Step 7. Wire it into the agent loop
A sketch — adapt the names to your framework. The essential shape: parse before executing, return rejections to the model as the tool result, log repairs, and cap retries.
from toolargs import ArgumentError, error_for_model, parse_arguments
MAX_ARGUMENT_RETRIES = 2
def handle_tool_call(call, spec, execute, log, attempt):
try:
parsed = parse_arguments(call.arguments, spec)
except ArgumentError as err:
log.warning("tool_args_rejected", extra={"tool": call.name, "code": err.code, "path": err.path})
if attempt >= MAX_ARGUMENT_RETRIES:
raise # stop the loop; escalate to a human or fail the task
return error_for_model(err) # goes back to the model as the tool result
for repair in parsed.repairs:
log.info("tool_args_repaired", extra={"tool": call.name, "path": repair.path, "rule": repair.rule})
return execute(**parsed.args)
Three decisions to make explicitly:
- Retry budget. Two corrections per call is a reasonable starting point; a model that cannot fix
"19.99"after an explicit error message will rarely fix it on attempt five. - What goes into logs. The sketch logs path and rule, not
before/after. Raw values can contain personal data; log them only through your redaction layer. - Side effects after repair. Repairs are safe by construction, but for a non-idempotent tool like a refund you still want an idempotency key and, above a threshold, human confirmation. The parser decides whether the arguments are well-formed, not whether the action is wise.
Failure cases to test before you trust it
- The over-eager helper. Someone adds
float(value)as a fallback incoerce_integer"to reduce retries". The rejection test for"19.99"must fail. If it does not, your test table has a hole. - The silent rename. A teammate maps
amount→amount_centsbecause the model keeps getting it wrong. Better fix: rename the field in the schema or improve its description; the repair log shows you exactly which field confuses the model. - Enum collision. Adding
"Other"next to"other"makes case-insensitive matching ambiguous. The parser rejects such inputs at runtime, and the lint test fails at CI time. - Provider drift. An SDK update starts delivering arguments as a dict instead of a string, or vice versa.
decode_payloadhandles both, but duplicate-key detection only works on strings — when you receive a dict, that check has already been lost upstream. - Unit confusion that type checks.
"amount_cents": 19is a valid integer and passes. No representation-level parser can tell 19 cents from "the model meant $19". See below.
Limitations
- Representation, not meaning. This layer catches envelope errors. A well-formed call with the wrong order ID or wrong amount passes. Semantic checks — "does this order exist, is the amount ≤ what was paid" — belong in the tool itself.
- Not a security boundary on its own. Arguments influenced by prompt injection can be perfectly well-formed. Strict parsing reduces the attack surface (no surprising coercions to exploit) but authorization and limits are separate guardrails.
- First error only. The parser stops at the first problem. That keeps the code simple, but a call with three bad fields takes up to three round trips. If that shows up in your replay numbers, collect all errors and return them together.
- Narrow schema dialect. Nested objects,
oneOf, numeric strings for decimals and string formats beyond date are not covered. Extend by adding coercers with the same contract: return a canonical value or raise a coded error, and add both a repair and a reject test. - Strictness is a trade-off you measure. Rejecting
" 1999"is a deliberate choice; you may decide trimming whitespace is unambiguous for your fields. Make that change with a test and a rule name, not with a quiet.strip().
Checklist
- The policy table exists and every repair has a one-sentence justification.
python -m pytest -qpasses locally and in CI.- The idempotence test covers every repair rule.
- High-risk fields are marked
strictor have an explicit reason not to be. - Repairs and rejections are logged by rule and code, without raw personal data.
- Replay runs on a redacted sample of real calls after every model or prompt change.
- The retry budget is capped and exhaustion escalates instead of looping.
Next: browse more step-by-step material in the guides, and look up any unfamiliar term in the glossary.