The bug report said the export feature “sometimes just doesn’t work.” No error in the logs, no failed request, nothing in the alerting. We eventually found it: about 2% of the time, the model’s JSON response had a single trailing comma or an unescaped quote inside a string field, our parser threw, and a try/catch three layers up silently swallowed it and returned an empty result. Nobody had written that catch block to hide failures — it was there for a genuinely unrelated edge case — but it had been quietly eating a real, recurring parsing failure for weeks.

“Ask the model to return JSON” and “the model returns valid, schema-conforming JSON” are two different claims, and the gap between them is where a surprising number of AI feature bugs live. This is what actually breaks, the two layers we built to stop trusting raw output, and what we measure now to know how often it’s really happening.

Three ways structured output fails, in order of how sneaky they are

The obvious failure — the model ignores the instruction entirely and writes a paragraph of prose — is the easiest to catch, because it fails loudly and immediately. The other two are worse specifically because they look like success until something downstream chokes on them.

  • Prose wrapped around valid JSON. “Here’s the JSON you asked for:” followed by a perfectly valid object. Passes a human eyeball check, fails JSON.parse outright.
  • Subtly malformed JSON. A trailing comma, a smart quote instead of a straight one copied from an example in the prompt, an unescaped newline inside a string. Fails parsing, but only on some fraction of responses, which makes it feel random and is genuinely hard to reproduce on demand.
  • Schema drift. The JSON parses fine but a field is missing, renamed, or the wrong type — a confidence score returned as the string "high" instead of a number, an array field returned as a single object when there’s only one result. This is the worst one, because it doesn’t throw anywhere. It just produces wrong behavior three function calls downstream.

We had all three in production before we went looking for them. Schema drift was the one that actually caused the export bug above — not the trailing comma itself, but a version of it where a numeric field occasionally came back as a numeric string, which our parser accepted but a later calculation silently treated as NaN.

Layer one: constrained decoding, not a polite request

Asking nicely for JSON in the prompt text is the weakest version of this. Most model providers now offer a real structured-output mode — tool/function-calling with a JSON schema, or a dedicated strict JSON mode — that constrains what tokens the model is even allowed to generate next, rather than hoping it follows instructions. This eliminates the first failure mode (prose wrapping) almost entirely and cuts the second (malformed syntax) sharply, because the decoder itself won’t emit a token that breaks the grammar.

javascript
const response = await anthropic.messages.create({
  model: 'claude-opus-5',
  tools: [{
    name: 'extract_ticket_summary',
    description: 'Extract a structured summary of a support ticket',
    input_schema: {
      type: 'object',
      properties: {
        category: { type: 'string', enum: ['refund', 'shipping', 'billing', 'other'] },
        urgency: { type: 'integer', minimum: 1, maximum: 5 },
        summary: { type: 'string' },
      },
      required: ['category', 'urgency', 'summary'],
    },
  }],
  tool_choice: { type: 'tool', name: 'extract_ticket_summary' },
  messages: [{ role: 'user', content: ticketText }],
});
// The tool_use block's input is decoded against the schema's grammar directly —
// not free text the model happens to format as JSON

This alone cut our malformed-syntax rate from roughly 4% of responses to under 0.1%. It did not touch schema drift at all — a constrained decoder guarantees the shape is syntactically valid JSON matching the schema’s types, but an enum field is still just a string as far as raw decoding is concerned, and a description that’s technically satisfied by a useless one-word answer is still “valid.”

Layer two: validate, then repair, don’t just retry blind

Every response still gets validated against the schema in application code before anything downstream touches it — never trust the provider’s structured mode as the only gate, because it constrains syntax and type, not semantic correctness, and providers’ guarantees here can vary by model and mode. On a validation failure, the naive fix is to just retry the whole request and hope for a better roll. The better fix is a repair loop: send the model its own broken output plus the specific validation error, and ask it to fix only that.

javascript
async function extractWithRepair(ticketText, maxAttempts = 2) {
  let lastOutput, lastError;

  for (let attempt = 0; attempt <= maxAttempts; attempt++) {
    const raw = attempt === 0
      ? await callModel(ticketText)
      : await repairCall(lastOutput, lastError);

    const result = TicketSummarySchema.safeParse(raw);
    if (result.success) return result.data;

    lastOutput = raw;
    lastError = result.error.format();
  }

  throw new StructuredOutputError('exhausted repair attempts', lastError);
}

async function repairCall(brokenOutput, validationError) {
  return callModel(`Your previous output failed validation:
${JSON.stringify(brokenOutput)}

Validation error: ${JSON.stringify(validationError)}

Return corrected JSON matching the schema exactly. Fix only what the error names.`);
}

The repair call is cheap relative to a fresh attempt — it's usually a small, targeted fix, not a full re-generation — and it succeeds far more often than a blind retry, because the model is correcting a specific named error instead of rolling the dice on the whole task again. We cap it at two repair attempts; past that, the failure almost always turns out to be a genuinely ambiguous input rather than a fixable formatting slip, and it's cheaper to route to a fallback than to keep paying for repair calls that won't land.

Constrained decoding fixes the syntax. Validation plus a targeted repair call fixes the semantics. Neither one alone gets you to a number you can actually trust in production.

What we monitor now that we didn't before

The number that matters isn't "did it eventually produce valid output" — it's the breakdown of how it got there, because each stage costs something different and a shift in the ratios is an early warning sign on its own.

  • First-try valid rate — passed schema validation with zero repair calls. This is your real baseline quality, and it's the number that should move when you change a prompt or a model version.
  • Repaired rate — needed one or more repair calls but eventually passed. Healthy in small amounts; a rising trend here usually means the schema grew more complex than the prompt's examples cover.
  • Exhausted rate — failed even after repair attempts and fell back. This is the number that should page someone if it moves, because it means real requests are silently degrading to a fallback path.

Watching only a single aggregate "success rate" hides the difference between a system that's healthy-but-paying-a-repair-tax and one where the exhausted bucket is quietly growing behind an unchanged headline number.

Key takeaways

  • "The model returns JSON" and "the model returns schema-conforming JSON" are different claims — treat every model response as untrusted input to your own validation layer, the same way you'd treat a third-party API response.
  • Constrained decoding (tool/function calling with a schema, or a strict JSON mode) eliminates prose-wrapping and sharply reduces malformed syntax, but does not guarantee semantic correctness or schema drift.
  • Always validate against your schema in application code, even when using a provider's structured-output mode — treat the mode as a strong assist, not the only gate.
  • On a validation failure, send the model its own broken output plus the specific error and ask for a targeted fix, rather than retrying the whole request blind. It's cheaper and lands more often.
  • Cap repair attempts (we use two) — past that point the failure is usually a genuinely ambiguous input, not a fixable formatting slip, and it's cheaper to route to a fallback.
  • Track first-try-valid, repaired, and exhausted rates separately, not one aggregate success number — each stage costs something different and a shift in the ratio is an early warning a single number hides.

Frequently asked questions

Does constrained decoding (tool calling / JSON mode) make schema validation unnecessary?

No — it guarantees syntactic validity and correct types for a well-specified schema, but not semantic correctness. A required string field can still come back as an empty string, or a summary field can technically satisfy the schema while being useless. Validate the content, not just the shape.

How much slower is a repair call compared to just accepting the bad output?

A repair call adds real latency — typically comparable to the original call, sometimes faster since the correction is narrower — so it's not free. We accept that cost because the alternative is either a silent data-quality bug or a user-facing failure, both of which cost more than the extra round trip.

What if the schema itself is the problem, not the model's output?

Worth checking before blaming the model — an ambiguous or overly rigid schema (a field that should be optional but is marked required, an enum missing a real-world case) produces failures that look like model unreliability but are actually a spec bug. A rising repaired-rate trend after a schema change is usually this.

Should every AI feature use constrained tool-calling output instead of prompting for JSON in plain text?

For anything a program will parse and act on, yes — there's little reason not to once your provider supports it. For output meant purely for a human to read, structured mode is unnecessary overhead; save it for the responses your own code actually consumes.

What do you do with requests that land in the exhausted bucket?

Route to a defined fallback rather than failing silently — for us that's usually a simpler, more constrained extraction path or a human review queue, depending on the feature. The important part is that it's a deliberate, monitored path, not a bare try/catch that happens to swallow the error three layers up.

Add a response

Your email address will not be published. Required fields are marked *