For AI agents: a documentation index is available at https://docs.parallel.ai/llms.txt. The full text of all docs is at https://docs.parallel.ai/llms-full.txt. You may also fetch any page as Markdown by appending
Processors are the engines that execute Task Runs. The choice of Processor determines the performance profile and reasoning behavior used. Any Task Run can be executed on any Processor.
Each processor varies in performance characteristics and supported features. Use the table below to compare processors.
.md to its URL or sending Accept: text/markdown.For Search and Extract, use /v1/search and /v1/extract for all new integrations and agent tool calls. Use /v1beta/search and /v1beta/extract only when maintaining an existing integration that already uses them. The request bodies changed too: on /v1/extract, full_content, fetch_policy and excerpt_settings go inside advanced_settings, and sending the v1beta top-level excerpts, full_content and fetch_policy to /v1/extract returns a 422. Do not substitute /v1 for the documented FindAll or Ingest endpoint paths.About “Max Fields” (the
~ symbol): The max fields column shows approximate limits because actual capacity depends on field complexity. Simple fields like dates or booleans use less capacity than complex fields requiring extensive research. A task with 5 complex analytical fields may require more processing than one with 15 simple lookup fields. Use these numbers as guidelines. If you’re near the limit and seeing quality issues, try a higher-tier processor.See Pricing for processor costs and all API rates.
Observed Latency
The table below shows median (p50) and 90th percentile (p90) execution times measured in production for standard processors across all customer workloads. Your results will vary with task complexity, output schema size, and input difficulty.Execution Time vs Queue Time
A Task Run moves through three states: queued → running → completed on success, or transitions to failed from either queued or running. The latencies above measure only therunning phase—the time a processor actively spends executing your task. Time spent in the queue is not included.
Queue time varies with your workload. Runs execute concurrently, but processing capacity is finite: when a large burst of runs is submitted at once, runs beyond the available capacity wait in the queue until capacity frees up. End-to-end time (creation to completion) can therefore exceed the execution time ranges above.
Latency-Sensitive Workflows
The Task API is designed for asynchronous, accuracy-critical research: runs are queued, executed, and retrieved when complete. If your use case is latency-sensitive—an interactive application, a subagent call, or any agentic research workflow where a caller is actively waiting on results—use the Responses API instead. It is purpose-built for low-latency requests.Examples
Processors can be used flexibly depending on the scope and structure of your task. The examples below show how to:- Use a single processor (like
lite,base,core,pro, orultra) to handle specific types of input and reasoning depth. - Chain processors together to combine fast lookups with deeper synthesis.
Sample Task for each Processor
Multi-Processor Workflows
You can combine processors in sequence to support more advanced workflows. Start by retrieving basic information withbase:
core to generate detailed background information: