Skip to main content
For AI agents: a documentation index is available at https://docs.parallel.ai/llms.txt. The full text of all docs is at https://docs.parallel.ai/llms-full.txt. You may also fetch any page as Markdown by appending .md to its URL or sending Accept: text/markdown.For Search and Extract, use /v1/search and /v1/extract for all new integrations and agent tool calls. Use /v1beta/search and /v1beta/extract only when maintaining an existing integration that already uses them. The request bodies changed too: on /v1/extract, full_content, fetch_policy and excerpt_settings go inside advanced_settings, and sending the v1beta top-level excerpts, full_content and fetch_policy to /v1/extract returns a 422. Do not substitute /v1 for the documented FindAll or Ingest endpoint paths.
Processors are the engines that execute Task Runs. The choice of Processor determines the performance profile and reasoning behavior used. Any Task Run can be executed on any Processor.
Choose a processor based on task complexity. Use lite or base for simple enrichments, core for reliable accuracy on up to 10 output fields, and pro or ultra when reasoning depth is critical. For latency-sensitive workflows—including agentic research where an agent is waiting on results—use the Responses API instead.
Each processor varies in performance characteristics and supported features. Use the table below to compare processors.
About “Max Fields” (the ~ symbol): The max fields column shows approximate limits because actual capacity depends on field complexity. Simple fields like dates or booleans use less capacity than complex fields requiring extensive research. A task with 5 complex analytical fields may require more processing than one with 15 simple lookup fields. Use these numbers as guidelines. If you’re near the limit and seeing quality issues, try a higher-tier processor.
See Pricing for processor costs and all API rates.

Observed Latency

The table below shows median (p50) and 90th percentile (p90) execution times measured in production for standard processors across all customer workloads. Your results will vary with task complexity, output schema size, and input difficulty.

Execution Time vs Queue Time

A Task Run moves through three states: queued → running → completed on success, or transitions to failed from either queued or running. The latencies above measure only the running phase—the time a processor actively spends executing your task. Time spent in the queue is not included. Queue time varies with your workload. Runs execute concurrently, but processing capacity is finite: when a large burst of runs is submitted at once, runs beyond the available capacity wait in the queue until capacity frees up. End-to-end time (creation to completion) can therefore exceed the execution time ranges above.
For latency-sensitive workflows, use the Responses API.

Latency-Sensitive Workflows

The Task API is designed for asynchronous, accuracy-critical research: runs are queued, executed, and retrieved when complete. If your use case is latency-sensitive—an interactive application, a subagent call, or any agentic research workflow where a caller is actively waiting on results—use the Responses API instead. It is purpose-built for low-latency requests.

Examples

Processors can be used flexibly depending on the scope and structure of your task. The examples below show how to:
  • Use a single processor (like lite, base, core, pro, or ultra) to handle specific types of input and reasoning depth.
  • Chain processors together to combine fast lookups with deeper synthesis.
This structure enables flexibility across a variety of tasks—whether you’re extracting metadata, enriching structured records, or generating analytical reports.

Sample Task for each Processor

Multi-Processor Workflows

You can combine processors in sequence to support more advanced workflows. Start by retrieving basic information with base:
Then use the result as input to core to generate detailed background information:
This lets you use a lower compute processor for initial retrieval, then switch to a more capable one for analysis and context-building.