← All skills

Research · Web

Exa Contents

You already have the URLs; this is the other half. POST /contents in raw HTTP — extracted text, highlights, summaries, subpages, freshness-controlled crawling and per-URL status, with no SDK in the way.

Skill name
exa-contents
Triggers on
Call Exa Contents directly with cURL or raw HTTP. Use when an agent already has URLs and needs POST /contents without an SDK for extracted text, highlights, summaries, links, image links, subpages, freshness-controlled crawling, or per-URL status handling.
Read time
6 min · Markdown · free to use and edit
Download .md

Use it in your assistant

Claude Code — drop the file in your skills folder and it loads on the next session. Use ~/.claude/skills for every project, or .claude/skills inside a repo to keep it to that project.

mkdir -p ~/.claude/skills/exa-contents
curl -L https://growsteady.io/skills/exa-contents/download -o ~/.claude/skills/exa-contents/SKILL.md

Claude apps (web and desktop) — Settings → Capabilities → Skills → add a skill. Upload the file as SKILL.md inside a folder named exa-contents (zip the folder if an archive is asked for).

No install— paste the file into a Claude Project's custom instructions with “Copy as prompt”. Same behaviour, scoped to that project.

Requires API key: Get one at https://dashboard.exa.ai/api-keys Header: x-api-key: $EXA_API_KEY

Use POST https://api.exa.ai/contents when the agent already knows the URLs and needs clean, LLM-ready extraction without running a new search. Start with one content mode: highlights for compact agent context, text for broad page context, or summary for Exa-side compression.

Quick Start (cURL)

Basic text extraction

curl -sS -X POST "https://api.exa.ai/contents" \
  -H "Content-Type: application/json" \
  -H "x-api-key: $EXA_API_KEY" \
  -d '{
    "urls": ["https://example.com"],
    "text": true
  }'

Highlights with freshness control

curl -sS -X POST "https://api.exa.ai/contents" \
  -H "Content-Type: application/json" \
  -H "x-api-key: $EXA_API_KEY" \
  -d '{
    "urls": ["https://arxiv.org/abs/2307.06435"],
    "highlights": {
      "query": "methodology and results"
    },
    "maxAgeHours": 24,
    "livecrawlTimeout": 12000
  }'

Endpoint

POST https://api.exa.ai/contents

Authentication: x-api-key: <API_KEY> header. Exa also accepts Authorization: Bearer <API_KEY>, but prefer x-api-key in cURL examples for consistency.

Use this endpoint for known-URL extraction. If the agent needs discovery or ranking first, use POST /search.

Parameters

Core request parameters

ParameterTypeRequiredDefaultDescription
urlsstring[]Yes-URLs to extract content from. Use this for known URLs.
textboolean or objectNo-Return full page text as markdown. Object form supports maxCharacters, includeHtmlTags, verbosity, includeSections, and excludeSections.
highlightsboolean or objectNo-Return key excerpts. Prefer true for agent workflows unless a custom focus is needed.
summaryboolean or objectNo-Return per-page LLM summaries. Use when the caller wants Exa-side compression or structured extraction.
maxAgeHoursintegerNo-Freshness control. 0 always live crawls; -1 uses cache only; omit for default cache-first behavior with crawl fallback.
livecrawlTimeoutintegerNo10000Timeout for live crawling in milliseconds. Use 10000 to 15000 for most freshness-sensitive calls.
subpagesintegerNo0Number of linked subpages to crawl from each URL.
subpageTargetstring or string[]No-Terms used to prioritize which subpages matter, such as ["api", "reference", "pricing"].
extras.linksintegerNo0Number of links to extract from each page.
extras.imageLinksintegerNo0Number of image URLs to extract from each page.
compliancestringNo-Enterprise-only compliance mode, such as hipaa, when enabled for the account.

Text object options

ParameterTypeDefaultDescription
maxCharactersinteger-Character limit for returned text. Use this instead of tokensNum.
includeHtmlTagsbooleanfalsePreserve HTML tags in output.
verbositystringcompactcompact, standard, or full. Pair fresh section-aware extraction with maxAgeHours: 0.
includeSectionsstring[]-Only include selected sections: header, navigation, banner, body, sidebar, footer, metadata.
excludeSectionsstring[]-Exclude selected sections from the same section list.

Highlights object options

Prefer highlights: true for the highest-quality default. Only use object form when the agent needs a custom focus or budget.

ParameterTypeDefaultDescription
querystring-Custom query guiding which excerpts are returned.
maxCharactersinteger-Cap highlight characters per URL. Omit unless the caller has a strict budget.

Summary object options

ParameterTypeDefaultDescription
querystring-Custom query for the summary.
schemaobject-JSON Schema for structured per-page summaries.

Content Modes

On /contents, text, highlights, and summary are top-level request fields.

ModeBest forNotes
textDeep analysis and broad page contextUse maxCharacters to keep payloads bounded.
highlightsAgent workflows and factual lookupsMost token-efficient default. Excerpts are grounded in the source page.
summaryCompression or structured per-page extractionAdds Exa-side synthesis per page.

Avoid requesting multiple modes unless the caller truly needs multiple views of the same page.

Freshness and Crawling

Use maxAgeHours as the normative freshness control.

ValueBehavior
omittedUse default cache-first behavior with crawl fallback when needed.
positive integerUse cache if it is less than N hours old, otherwise live crawl.
0Always live crawl. Highest freshness, higher latency.
-1Cache only. Fastest, but fails if no cached content exists.

Set livecrawlTimeout when live crawling should not block past a fixed budget.

Subpages and extras

Use subpages and subpageTarget when linked pages matter.

curl -sS -X POST "https://api.exa.ai/contents" \
  -H "Content-Type: application/json" \
  -H "x-api-key: $EXA_API_KEY" \
  -d '{
    "urls": ["https://docs.example.com"],
    "text": {
      "maxCharacters": 5000
    },
    "subpages": 10,
    "subpageTarget": ["api", "reference", "guide"],
    "extras": {
      "links": 10,
      "imageLinks": 5
    }
  }'

Start with subpages around 5 to 10, then increase only when the caller needs broader site coverage.

Response Fields and Statuses

Inspect statuses even when the HTTP status is 200. The endpoint can succeed for one URL and fail for another in the same request.

curl -sS -X POST "https://api.exa.ai/contents" \
  -H "Content-Type: application/json" \
  -H "x-api-key: $EXA_API_KEY" \
  -d '{
    "urls": ["https://example.com", "https://failed-url.example.com"],
    "highlights": true
  }' | jq '{results, statuses}'
FieldTypeDescription
requestIdstringUnique request identifier.
resultsarrayExtracted content result objects.
results[].titlestringPage title.
results[].urlstringPage URL.
results[].publishedDatestring or nullEstimated publication date when available.
results[].authorstring or nullAuthor when available.
results[].textstringReturned when text is requested.
results[].highlightsstring[]Returned when highlights is requested.
results[].highlightScoresnumber[]Similarity scores for highlights.
results[].summarystringReturned when summary is requested.
results[].subpagesarrayNested result objects from subpage crawling.
results[].extras.linksstring[]Extracted links when requested.
statusesarrayPer-URL success or error states. Always inspect this field.
statuses[].idstringRequested URL.
statuses[].statusstringsuccess or error.
statuses[].error.tagstringError type for failed URLs.
statuses[].error.httpStatusCodeinteger or nullHTTP code associated with a per-URL failure.
costDollars.totalnumberTotal request cost when returned.

Common per-URL error tags include CRAWL_NOT_FOUND, CRAWL_TIMEOUT, CRAWL_LIVECRAWL_TIMEOUT, SOURCE_NOT_AVAILABLE, UNSUPPORTED_URL, and CRAWL_UNKNOWN_ERROR.

Critical Pitfalls

  • Keep text, highlights, and summary at the top level on /contents.
  • Do not wrap extraction options in a contents object; that nesting belongs to /search.
  • Do not assume HTTP 200 means every URL succeeded; inspect statuses.
  • Do not send stream: true; /contents is not a streaming endpoint.
  • Do not send tokensNum; use text.maxCharacters to cap extracted text.
  • Do not use useAutoprompt, numSentences, highlightsPerUrl, or older livecrawl string values in new requests.
  • Prefer maxAgeHours for freshness and pair it with livecrawlTimeout when crawl latency matters.
  • Use subpageTarget with subpages; otherwise subpage selection is best effort.
  • Pick one of highlights, text, or summary by default. Stack modes only when the caller truly needs multiple views of each page.