Skip to content
URL TO MARKDOWN API

URL to Markdown API

Turn any web page into clean, LLM-ready markdown in one call. Boilerplate stripped, structure preserved, image alt text kept — and token-aware chunking that returns embeddings-ready segments without a parsing library.

/v1/extract renders the page in a real browser, dissolves nav, ads, and cookie chrome, and returns clean markdown with structure intact. Set chunk: true and you get back segments split on heading boundaries, packed to your token budget with configurable overlap — ready to embed.

Messy page in, clean chunks out

POSTextract.sh
curl https://api.renderkit.tech/v1/extract \
  -H "x-api-key: $RENDERKIT_API_KEY" \
  -H "content-type: application/json" \
  -d '{ "url": "https://example.com", "chunk": true, "chunk_max_tokens": 512 }'

Built for RAG and agents

  • Boilerplate stripping with structure and heading hierarchy preserved
  • Token-aware chunking on heading boundaries with chunkoverlaptokens
  • Image alt-text extraction kept inline
  • Pipe meta.chunks straight into pgvector or Pinecone — no recursive splitter to tune
  • Pair with a screenshot for agents that need visual + text context, on one key

Single-URL precision

RenderKit is deliberately single-URL — for recursive whole-site crawling, see RenderKit vs Firecrawl. For the agent surface, the MCP server exposes extraction as a native tool. More on token-aware chunking.

Common questions

POST to /v1/extract with the URL. RenderKit returns clean markdown; add chunk: true for token-aware segments ready to embed.

Instead of splitting on a fixed character count, RenderKit segments markdown on heading boundaries and packs each chunk up to your chunk_max_tokens budget with overlap, so embeddings never capture half-thoughts.

No — extraction is single-URL by design. For recursive crawling with depth controls, run a crawler alongside RenderKit.

Ship your first render in 60 seconds.

500 free renders to start. No card. Cancel by ignoring it.

Get your free key$npm i renderkit-sdk