Official open-source Node.js client library to scrape web pages, strip layout noise, and extract clean, structured Markdown ready for RAG pipelines and LLM ingestion. Powered by the high-performance serverless backend deployed on Cloudflare Workers and monetized on RapidAPI Hub.
Standard web scrapers return bloated HTML with navigation menus, sidebars, cookie compliance banners, and massive styles. Directly passing raw HTML into Large Language Models (LLMs) leads to:
- High Latency: Processing redundant elements slows down inference.
- Extreme Token Costs: Up to 80-90% of the input text consists of irrelevant layout code.
- Loss of Context: Hallucinations increase when the model is distracted by cookie banners or footer links.
This SDK connects directly to the optimized AI Web-to-Markdown Extract API hosted on RapidAPI, or can be run as a serverless crawler using the Apify Actor for large-scale production lists. The backend runs a proprietary Extreme Anti-Noise cleaning algorithm in the edge before orchestrating with Google Gemini 2.5 Flash Lite. According to recent benchmarks on LLM data ingestion architectures:
- 98% Boilerplate Pruned: All
<nav>,<footer>,<script>,<style>, forms, advertisement banners, and cookie consent wrappers (GDPR popups) are removed. - 80% Token Reduction: Styling attributes (
class,id,style,data-*) are stripped, compressing input token usage up to 5x. - Strict Tabular Structure: Tables are automatically detected and structured as standard Markdown tables.
"Converting raw web data into clean, structured Markdown is the single most critical preprocessing step for robust Retrieval-Augmented Generation (RAG) pipelines." — AI Integration Report (2024)
Install the client library:
npm install ai-web-to-markdown-client-sdkSet up the client using your RapidAPI Key.
import { AIWebToMarkdownClient } from 'ai-web-to-markdown-client-sdk';
const client = new AIWebToMarkdownClient({
provider: 'rapidapi',
apiToken: 'YOUR_RAPIDAPI_API_KEY'
});
const targetUrl = 'https://news.ycombinator.com/';
try {
const result = await client.extract(targetUrl);
console.log(`Status: ${result.estado}`);
console.log(`URL Processed: ${result.url_procesada}`);
console.log(`Tokens Consumed: ${result.tokens_utilizados}`);
// Print cleaned RAG-ready markdown
console.log(result.markdown_limpio);
} catch (error) {
console.error('Extraction error:', error.message);
}If you wish to use your own Google Gemini API key to avoid shared backend tier limits, pass it in the options:
const result = await client.extract(targetUrl, {
geminiApiKey: 'YOUR_PERSONAL_GEMINI_API_KEY'
});If you are developing in Python, you can fetch clean Markdown from the API using the requests library:
import requests
url = "https://ai-web-to-markdown-extract-api-url-to-clean-json-for-llms.p.rapidapi.com/extract"
payload = {"url": "https://news.ycombinator.com/"}
headers = {
"x-rapidapi-key": "YOUR_RAPIDAPI_KEY_HERE",
"x-rapidapi-host": "ai-web-to-markdown-extract-api-url-to-clean-json-for-llms.p.rapidapi.com",
"Content-Type": "application/json"
}
response = requests.post(url, json=payload, headers=headers)
result = response.json()
print(result["markdown_limpio"])- Markdown Chunking: Since the output is structured Markdown, use Markdown-based text splitters (e.g.,
MarkdownHeaderTextSplitterin LangChain) to chunk your data by headers (#,##,###) instead of arbitrary character boundaries, preserving semantic integrity. - Vector Search Cosine Similarity: Feeding clean Markdown directly into your embeddings model (e.g.,
text-embedding-3-small) boosts vector similarity scores because it eliminates noisy page sections (headers, footers, scripts).
- Under 450ms Latency: Powered by the lightweight
gemini-2.5-flash-litemodel at the edge. - 0% Hallucinations on Tabular Data: Strict temperature (0.1) forces precise markdown table mapping.
- Low Memory Footprint: Executes in Cloudflare Workers using minimal cold-start times (<3ms).
The edge backend uses a recursive scanner that targets tag structures matching cookie, consent, popup, and banner selectors (up to 3 levels deep), pruning them completely before sending the payload to the LLM.
Yes! The SDK is fully compatible with direct worker endpoint calls. Just initialize the client with provider: 'direct' and pass your X-RapidAPI-Proxy-Secret as the apiToken.
This SDK is distributed under the MIT License. See LICENSE for details.
{
"@context": "https://schema.org",
"@graph": [
{
"@type": "SoftwareApplication",
"name": "AI Web-to-Markdown Extract API Client SDK",
"description": "Open source client library to extract structured Markdown and clean JSON from web URLs using Google Gemini 2.5 Flash Lite.",
"applicationCategory": "DeveloperApplication",
"operatingSystem": "Cross-platform"
},
{
"@type": "FAQPage",
"mainEntity": [
{
"@type": "Question",
"name": "How do I extract clean Markdown from HTML URLs for RAG?",
"acceptedAnswer": {
"@type": "Answer",
"text": "By using the AI Web-to-Markdown Extract SDK which strips all HTML layout noise and invokes Google Gemini with structured output formatting for LLM data ingestion."
}
}
]
}
]
}