How to Make Sure LLM Output Is Valid JSON
Yongkang Zou · 2026-01-10 · research · LLM Infrastructure
How to Make Sure LLM Output Is Valid JSON
If you've built anything with LLMs, you've hit this: you ask for JSON, you get JSON-ish. A missing bracket, an extra comma, a markdown code fence wrapping the response, or the model deciding to "explain" its output before the actual JSON.
This is a solved problem in 2026. Here's what actually works, ranked from worst to best.
The Three Levels
The structured output maturity model breaks down neatly into three levels:
graph LR
L1[Level 1: Prompt Engineering] --> L2[Level 2: Function Calling]
L2 --> L3[Level 3: Constrained Decoding]
Level 1: Prompt Engineering (80-95% reliable)
The naive approach. You write "Return JSON with fields: name, email, score" in your prompt. It works most of the time. Then it doesn't.
Failure modes:
- Model wraps output in
```jsonfences - Adds conversational text before or after the JSON
- Produces syntactically valid JSON that doesn't match your schema
- Trailing commas, missing quotes, truncated output on long responses
People "fix" this with regex, string slicing, and retry loops. Don't. This is tech debt that scales linearly with the number of prompts you maintain.
Level 2: Function Calling / Tool Use (95-99% reliable)
You define a function schema and force the model to "call" it. The schema acts as a strong hint.
OpenAI calls this function calling with strict mode. Anthropic uses tool_use with tool_choice. Gemini has response_schema in GenerationConfig.
This works well but the schema functions as a hint, not a hard constraint. The model can still produce invalid JSON in edge cases, especially with deeply nested schemas or when the context is long.
Level 3: Constrained Decoding (100% schema-valid)
This is the real solution. The model literally cannot produce invalid JSON.
How it works: a finite state machine tracks your position in the JSON schema at every token generation step. Only tokens that lead to a valid next state get probability mass. Invalid tokens are masked to zero before sampling.
sequenceDiagram
participant Model as LLM
participant FSM as Schema FSM
participant Out as Output
Model->>FSM: Next token candidates
FSM->>FSM: Check valid transitions
FSM->>Model: Mask invalid tokens
Model->>Out: Sample from valid only
Note right of Out: Guaranteed schema-valid
OpenAI and Gemini both implement this natively. JSONSchemaBench tested 10K real-world schemas across six frameworks and found that constrained decoding achieves 100% schema validity, while prompt-based approaches fail on 5-20% of complex schemas.
Provider Comparison (2026)
| Provider | Method | Constrained Decoding | Schema Validity | Pydantic/Zod Native |
|---|---|---|---|---|
| OpenAI | Native Structured Output | Yes | 100% | Yes (.parse()) |
| Anthropic | Tool Use + tool_choice | Partial | 99%+ | Manual |
| Gemini | response_schema | Yes | 100% | Manual |
| Cohere | response_format | Yes | 100% | Manual |
| Open-source (vLLM, llama.cpp) | Outlines / XGrammar | Yes | 100% | Yes |
Note: Anthropic's tool_use approach is extremely reliable in practice (99%+) but technically doesn't use full constrained decoding on the response body. For most applications, this is fine.
What to Actually Use
OpenAI (Python)
from pydantic import BaseModel
from openai import OpenAI
class Analysis(BaseModel):
sentiment: str # "positive" | "negative" | "neutral"
confidence: float
topics: list[str]
client = OpenAI()
result = client.beta.chat.completions.parse(
model="gpt-4o",
messages=[{"role": "user", "content": "Analyze: Great product!"}],
response_format=Analysis,
)
# result.choices[0].message.parsed is a validated Analysis object
Anthropic (Python)
import anthropic, json
client = anthropic.Anthropic()
response = client.messages.create(
model="claude-sonnet-4-20250514",
max_tokens=1024,
tools=[{
"name": "analyze",
"description": "Analyze sentiment",
"input_schema": {
"type": "object",
"properties": {
"sentiment": {"type": "string", "enum": ["positive", "negative", "neutral"]},
"confidence": {"type": "number"},
"topics": {"type": "array", "items": {"type": "string"}}
},
"required": ["sentiment", "confidence", "topics"]
}
}],
tool_choice={"type": "tool", "name": "analyze"},
messages=[{"role": "user", "content": "Analyze: Great product!"}]
)
data = response.content[0].input # dict matching schema
Gemini (REST)
{
"generationConfig": {
"responseMimeType": "application/json",
"responseJsonSchema": {
"type": "object",
"properties": {
"sentiment": {"type": "string"},
"confidence": {"type": "number"},
"topics": {"type": "array", "items": {"type": "string"}}
},
"required": ["sentiment", "confidence", "topics"]
}
}
}
The Validation Layer You Still Need
Even with 100% schema validity, you still need business logic validation. A schema says "confidence is a number." It doesn't say "confidence is between 0 and 1." A schema says "sentiment is a string." It doesn't say the string makes sense.
flowchart LR
LLM[LLM Output] --> CD[Constrained Decoding]
CD --> BV[Business Validation]
BV -->|pass| USE[Use in app]
BV -->|fail| RETRY[Retry or fallback]
Use Pydantic (Python) or Zod (TypeScript) as your second validation layer. Define your schema once, use it for both the LLM constraint and the runtime validation.
Decision Tree
- Using OpenAI or Gemini? Use native structured output with JSON Schema. You get 100% schema validity for free.
- Using Anthropic? Use tool_use with tool_choice forced to your schema tool. 99%+ reliable, clean API.
- Using open-source models? Use Outlines or XGrammar for constrained decoding. Both integrate with vLLM and llama.cpp.
- Need a library that works across providers? Instructor wraps OpenAI, Anthropic, Gemini, and others with a unified Pydantic-first API.
- Still using regex to parse LLM JSON? Stop. It's 2026.
This is a reference note for structured output patterns I use across projects. The Gemini approach is what powers the AI blog generation on yongkang.dev (see responseJsonSchema in the generate-draft endpoint).
References
- LLM Structured Output in 2026: Stop Parsing JSON with Regex — DEV Community
- OpenAI Structured Outputs — OpenAI Docs
- Anthropic Tool Use — Claude Docs
- Constrained Decoding: Grammar-Guided Generation — Michael Brenndoerfer
- JSONSchemaBench: Generating Structured Outputs from LLMs — arXiv
- The Guide to Structured Outputs and Function Calling — Agenta
- Instructor: Structured LLM Outputs — GitHub
- Outlines: Structured Text Generation — GitHub