URL to Markdown API
Turn any web page into clean, LLM-ready markdown in one call. Boilerplate stripped, structure preserved, image alt text kept — and token-aware chunking that returns embeddings-ready segments without a parsing library.
/v1/extract renders the page in a real browser, dissolves nav, ads, and cookie chrome, and returns clean markdown with structure intact. Set chunk: true and you get back segments split on heading boundaries, packed to your token budget with configurable overlap — ready to embed.
Messy page in, clean chunks out
curl https://api.renderkit.tech/v1/extract \
-H "x-api-key: $RENDERKIT_API_KEY" \
-H "content-type: application/json" \
-d '{ "url": "https://example.com", "chunk": true, "chunk_max_tokens": 512 }'Built for RAG and agents
- Boilerplate stripping with structure and heading hierarchy preserved
- Token-aware chunking on heading boundaries with chunkoverlaptokens
- Image alt-text extraction kept inline
- Pipe meta.chunks straight into pgvector or Pinecone — no recursive splitter to tune
- Pair with a screenshot for agents that need visual + text context, on one key
Single-URL precision
RenderKit is deliberately single-URL — for recursive whole-site crawling, see RenderKit vs Firecrawl. For the agent surface, the MCP server exposes extraction as a native tool. More on token-aware chunking.
Common questions
POST to /v1/extract with the URL. RenderKit returns clean markdown; add chunk: true for token-aware segments ready to embed.
Instead of splitting on a fixed character count, RenderKit segments markdown on heading boundaries and packs each chunk up to your chunk_max_tokens budget with overlap, so embeddings never capture half-thoughts.
No — extraction is single-URL by design. For recursive crawling with depth controls, run a crawler alongside RenderKit.
Ship your first render in 60 seconds.
500 free renders to start. No card. Cancel by ignoring it.