⚡
agent-tool-parser
Deterministic Extraction Engine Zero Runtime Dependencies

Structured Tool Call Parser for
Autonomous AI Agents

Extract, normalize, and repair heterogeneous tool calls from large language model text streams. Sub-0.05ms deterministic execution across standard OpenAI, Claude XML, DeepSeek DSML, Qwen, and Llama formats.

$ pip install agent-tool-parser
Try Live Playground ↓
< 0.05ms
Average Latency
0 Deps
Pure Standard Library
Multi-Tool
Parallel Invocation
MIT
Free Commercial Use

Interactive Evaluation Playground

Evaluate how heterogeneous or truncated model completions are normalized to canonical ToolCall structures.

Model Output Stream (Raw Input)
Normalized ToolCall Output 0.03ms

        

Architectural Challenges in Heterogeneous Tool Calling

Client-side agent runtimes encounter structural inconsistencies when interfacing with diverse cloud APIs and local inference engines.

1

Content Stream Delimiters

Inference endpoints (such as OpenRouter, DeepSeek API, or vLLM) frequently return empty tool_calls arrays while embedding raw XML, DSML, or proprietary delimiter tokens directly within message.content.

2

Stream Interruption & Truncation

When model generation terminates at token limits or network interruptions occur, standard JSON deserialization fails. The parser evaluates bracket stacks to close and recover partial JSON payloads deterministically.

3

Code Escaping & CDATA Blocks

Serializing shell scripts, unified diffs, or Python functions within JSON string literals frequently results in escape violations. The parser extracts unescaped CDATA blocks and multi-line script bodies without string corruption.

Architectural Integration Patterns

Reference implementations for integrating agent-tool-parser into agent loops, gateway proxies, and parallel execution pipelines.

Pattern 1

Autonomous Agent Loop (ReAct / Plan-and-Execute)

from agent_tool_parser import try_parse_tool_calls

def run_agent_turn(messages: list[dict], client, model_name: str):
    response = client.chat.completions.create(
        model=model_name,
        messages=messages,
        temperature=0.0,
    )
    raw_content = response.choices[0].message.content or ""

    # Resilient extraction across DeepSeek, Claude, Qwen, Llama & non-standard JSON
    tool_calls = try_parse_tool_calls(
        raw_content,
        allowed_tools={"read_file", "write_file", "bash", "search"},
        tool_aliases={"view": "read_file", "sh": "bash", "cmd": "bash"}
    )

    if not tool_calls:
        return {"action": "reply", "content": raw_content}

    # Execute all tools (in parallel or sequentially)
    results = []
    for call in tool_calls:
        output = execute_tool(call.name, call.args)
        results.append({"tool": call.name, "output": output})

    return {"action": "continue", "results": results}
Pattern 2

API Gateway & Router Normalization Middleware

from agent_tool_parser import parse_tool_calls, safe_json_loads, ToolError

def normalize_response(response_message):
    """Normalizes response structures when inference endpoints emit markup in message content."""
    # 1. Return native tool_calls if already structured by provider SDK
    if getattr(response_message, "tool_calls", None):
        return [
            {"name": tc.function.name, "args": safe_json_loads(tc.function.arguments)}
            for tc in response_message.tool_calls
        ]

    # 2. Extract structured invocations from text content stream
    content = getattr(response_message, "content", "") or ""
    try:
        calls = parse_tool_calls(content)
        return [{"name": c.name, "args": c.args} for c in calls]
    except ToolError:
        return []
Pattern 3

Concurrent Multi-Tool Dispatch (AsyncIO)

import asyncio
from agent_tool_parser import parse_tool_calls

async def handle_llm_turn(model_output: str):
    # Parse multiple parallel tool calls from one generation step
    calls = parse_tool_calls(model_output)

    # Execute all tools concurrently in parallel
    tasks = [async_run_tool(call.name, call.args) for call in calls]
    tool_results = await asyncio.gather(*tasks, return_exceptions=True)

    return list(zip([c.name for c in calls], tool_results))

Comparative Technical Evaluation

Objective comparison of common parsing, repair, and extraction approaches in AI agent runtimes.

Evaluation Criterion agent-tool-parser json-repair LangChain Core vLLM Server-Side
Runtime Dependencies 0 (Pure Python) 0 > 30 packages Torch, CUDA (GBs)
Average Latency < 0.05 ms ~ 0.15 ms ~ 1.20 ms N/A (Server-side)
Reasoning Stripping (<think>) ✓ Automatic ✗ Crashes Partial ✗ No
DeepSeek DSML & Native Tokens ✓ Full Support ✗ No ✗ No ✓ Yes
Anthropic Claude XML + CDATA ✓ Full Support ✗ No Partial ✗ No
Llama 3.1 Python AST Expressions ✓ Safe AST (No eval) ✗ No ✗ No ✗ No
Truncated Streaming Tail Repair ✓ Bracket-Stack Auto-close ✓ Yes ✗ Crashes ✗ No
Multi-Tool Calling ✓ Any syntax ✗ No JSON only ✗ No

Architectural Boundaries: Post-Parsing vs. Server Alternatives

System trade-offs across execution boundaries, sampler integration, and dependency stacks.

Architectural Dimension Client-Side Post-Parsing Logit-Level Constrained Decoding Secondary Model Pass
Execution Boundary Client Runtime Inference Engine Sampler Secondary Endpoint / Worker
Inference Engine Access Agnostic (HTTP/REST text streams) Requires direct logit & sampler access Requires secondary model call
Runtime Dependencies Python Standard Library PyTorch, CUDA, Triton PyTorch, Transformers / API
Provider API Compatibility ✓ Any provider returning text ✗ Incompatible with third-party APIs ✓ Via extra API round-trip
Grammar Enforcement Deterministic text parsing & repair FSM logit masking during token sampling Probabilistic prompting pass
Format Scope XML, DSML, Markdown JSON, AST, CDATA Strict JSON / EBNF Grammars Model training distribution

Ecosystem & Architecture Fit

Built for the reality of multi-agent frameworks, high-throughput engines, and fleet mission controls.

🛰️

Mission Control & Fleet Ops

Mission Control • Paperclip AI • Dify

When managing fleets of background agents, a single unparsed tool call causes worker crashes, poisoned queues, and blown token budgets. agent-tool-parser acts as an ingress sanitization layer before tasks hit approval gates.

👥

Multi-Agent Frameworks

CrewAI • AutoGen (AG2) • LangGraph • PydanticAI

Local models (Ollama/vLLM) frequently emit <think> tokens or parameter aliases that trigger fatal ValidationErrors in CrewAI or PydanticAI. Our interceptor normalizes them smoothly.

⚡

High-Throughput Serving

SGLang • vLLM • Ollama • llama.cpp

Strict constrained decoding (XGrammar / outlines) degrades reasoning quality in models like DeepSeek R1. With agent-tool-parser, keep token generation unconstrained while extracting tool calls in <0.05ms client-side.

🔀

AI Gateways & Routers

LiteLLM • Portkey • OpenRouter

Multi-model proxies often fail to extract vendor-specific XML or DSML out of message.content into standard tool_calls arrays. Our parser acts as a zero-dependency normalization hook.

💻

Coding & Terminal Agents

Aider • Cline • OpenHands (OpenDevin)

Multi-line Python scripts and Bash commands constantly break standard JSON escaping. Full support for XML tags, CDATA unwrapping, and code blocks ensures lossless code execution.

🛠️

Minimalist & Code Agents

Smolagents • Custom ReAct Loops

Hugging Face built CodeAgent because JSON tool-calling was notoriously fragile. agent-tool-parser gives developers the reliability of code agents without requiring sandboxed Python execution environments.

API Quick Reference

Simple, predictable functional interfaces with configurable class extensions.

parse_tool_calls(text, ...) -> list[ToolCall]

Parses an LLM response string and returns all structured ToolCall objects found. Raises ToolError if no valid tool calls could be extracted.

calls = parse_tool_calls(raw_text)
for call in calls:
    print(call.name, call.args)
try_parse_tool_calls(text, ...) -> list[ToolCall]

Safe non-raising variant. Returns an empty list [] if the model response only contained conversational text without any tool calls.

calls = try_parse_tool_calls(raw_text)
if not calls:
    return "No tool needed, answer directly"
ToolParser(allowed_tools, tool_aliases, ...)

Reusable, configurable parser instance. Enforces tool whitelists, normalizes custom aliases (e.g. view -> read_file), and handles argument mapping.

parser = ToolParser(
    allowed_tools=["read_file", "bash"],
    tool_aliases={"terminal": "bash"}
)
call = parser.parse(raw_text)
safe_json_loads(s, default=None) -> Any

Stand-alone resilient JSON loader. Auto-repairs single quotes, trailing commas, and converts Python literals (True/False/None) without throwing exceptions.

from agent_tool_parser import safe_json_loads

data = safe_json_loads("{'path': 'a.py', 'overwrite': True,}")
# {'path': 'a.py', 'overwrite': True}