How to Make Sure LLM Output Is Valid JSON

Yongkang Zou · 2026-01-10 · research · LLM Infrastructure

How to Make Sure LLM Output Is Valid JSON

If you've built anything with LLMs, you've hit this: you ask for JSON, you get JSON-ish. A missing bracket, an extra comma, a markdown code fence wrapping the response, or the model deciding to "explain" its output before the actual JSON.

This is a solved problem in 2026. Here's what actually works, ranked from worst to best.

The Three Levels

The structured output maturity model breaks down neatly into three levels:

graph LR
    L1[Level 1: Prompt Engineering] --> L2[Level 2: Function Calling]
    L2 --> L3[Level 3: Constrained Decoding]

Level 1: Prompt Engineering (80-95% reliable)

The naive approach. You write "Return JSON with fields: name, email, score" in your prompt. It works most of the time. Then it doesn't.

Failure modes:

  • Model wraps output in ```json fences
  • Adds conversational text before or after the JSON
  • Produces syntactically valid JSON that doesn't match your schema
  • Trailing commas, missing quotes, truncated output on long responses

People "fix" this with regex, string slicing, and retry loops. Don't. This is tech debt that scales linearly with the number of prompts you maintain.

Level 2: Function Calling / Tool Use (95-99% reliable)

You define a function schema and force the model to "call" it. The schema acts as a strong hint.

OpenAI calls this function calling with strict mode. Anthropic uses tool_use with tool_choice. Gemini has response_schema in GenerationConfig.

This works well but the schema functions as a hint, not a hard constraint. The model can still produce invalid JSON in edge cases, especially with deeply nested schemas or when the context is long.

Level 3: Constrained Decoding (100% schema-valid)

This is the real solution. The model literally cannot produce invalid JSON.

How it works: a finite state machine tracks your position in the JSON schema at every token generation step. Only tokens that lead to a valid next state get probability mass. Invalid tokens are masked to zero before sampling.

sequenceDiagram
    participant Model as LLM
    participant FSM as Schema FSM
    participant Out as Output
    Model->>FSM: Next token candidates
    FSM->>FSM: Check valid transitions
    FSM->>Model: Mask invalid tokens
    Model->>Out: Sample from valid only
    Note right of Out: Guaranteed schema-valid

OpenAI and Gemini both implement this natively. JSONSchemaBench tested 10K real-world schemas across six frameworks and found that constrained decoding achieves 100% schema validity, while prompt-based approaches fail on 5-20% of complex schemas.

Provider Comparison (2026)

ProviderMethodConstrained DecodingSchema ValidityPydantic/Zod Native
OpenAINative Structured OutputYes100%Yes (.parse())
AnthropicTool Use + tool_choicePartial99%+Manual
Geminiresponse_schemaYes100%Manual
Cohereresponse_formatYes100%Manual
Open-source (vLLM, llama.cpp)Outlines / XGrammarYes100%Yes

Note: Anthropic's tool_use approach is extremely reliable in practice (99%+) but technically doesn't use full constrained decoding on the response body. For most applications, this is fine.

What to Actually Use

OpenAI (Python)

from pydantic import BaseModel
from openai import OpenAI

class Analysis(BaseModel):
    sentiment: str  # "positive" | "negative" | "neutral"
    confidence: float
    topics: list[str]

client = OpenAI()
result = client.beta.chat.completions.parse(
    model="gpt-4o",
    messages=[{"role": "user", "content": "Analyze: Great product!"}],
    response_format=Analysis,
)
# result.choices[0].message.parsed is a validated Analysis object

Anthropic (Python)

import anthropic, json

client = anthropic.Anthropic()
response = client.messages.create(
    model="claude-sonnet-4-20250514",
    max_tokens=1024,
    tools=[{
        "name": "analyze",
        "description": "Analyze sentiment",
        "input_schema": {
            "type": "object",
            "properties": {
                "sentiment": {"type": "string", "enum": ["positive", "negative", "neutral"]},
                "confidence": {"type": "number"},
                "topics": {"type": "array", "items": {"type": "string"}}
            },
            "required": ["sentiment", "confidence", "topics"]
        }
    }],
    tool_choice={"type": "tool", "name": "analyze"},
    messages=[{"role": "user", "content": "Analyze: Great product!"}]
)
data = response.content[0].input  # dict matching schema

Gemini (REST)

{
  "generationConfig": {
    "responseMimeType": "application/json",
    "responseJsonSchema": {
      "type": "object",
      "properties": {
        "sentiment": {"type": "string"},
        "confidence": {"type": "number"},
        "topics": {"type": "array", "items": {"type": "string"}}
      },
      "required": ["sentiment", "confidence", "topics"]
    }
  }
}

The Validation Layer You Still Need

Even with 100% schema validity, you still need business logic validation. A schema says "confidence is a number." It doesn't say "confidence is between 0 and 1." A schema says "sentiment is a string." It doesn't say the string makes sense.

flowchart LR
    LLM[LLM Output] --> CD[Constrained Decoding]
    CD --> BV[Business Validation]
    BV -->|pass| USE[Use in app]
    BV -->|fail| RETRY[Retry or fallback]

Use Pydantic (Python) or Zod (TypeScript) as your second validation layer. Define your schema once, use it for both the LLM constraint and the runtime validation.

Decision Tree

  • Using OpenAI or Gemini? Use native structured output with JSON Schema. You get 100% schema validity for free.
  • Using Anthropic? Use tool_use with tool_choice forced to your schema tool. 99%+ reliable, clean API.
  • Using open-source models? Use Outlines or XGrammar for constrained decoding. Both integrate with vLLM and llama.cpp.
  • Need a library that works across providers? Instructor wraps OpenAI, Anthropic, Gemini, and others with a unified Pydantic-first API.
  • Still using regex to parse LLM JSON? Stop. It's 2026.

This is a reference note for structured output patterns I use across projects. The Gemini approach is what powers the AI blog generation on yongkang.dev (see responseJsonSchema in the generate-draft endpoint).

References