Skip to content

divyanshu-iitian/SearchForge

Repository files navigation

SearchForge

Open-source web search API and MCP server for LLMs, agents, and RAG.

CI License: MIT Node.js MCP Website Official MCP Registry

One local gateway. Intent-aware search. Clean Markdown. REST, MCP, CLI, and TypeScript.

Website · Quick start · Free tools · MCP · API · Design


SearchForge is a free, open-source web search API and MCP server for LLMs, AI agents, and retrieval-augmented generation (RAG) pipelines. It provides a predictable retrieval layer without forcing every project to integrate a paid search vendor. SearchForge routes each query to the right source, isolates provider failures, deduplicates URLs, fuses rankings, and turns public pages into LLM-ready Markdown.

It does not generate answers, hide citations, scrape public SearXNG instances, or send telemetry.

What you get

Capability Default source Cost / credentials
auto Intent-routed source mix No key by default
web Wikipedia; optional private SearXNG No key / self-hosted
code GitHub repository search No key; token optional
academic Crossref works and DOI metadata No key
community Hacker News via Algolia No key, community service
read_url Jina Reader No key, currently rate-limited

SearchForge starts with all no-key adapters enabled. auto is the default and routes code, research, and current/community intent to relevant sources while retaining a web fallback. A GitHub token only raises the public API quota, and Brave remains an optional keyed backend. Broad, independent web metasearch is provided by the included SearXNG stack.

Quick start

Try it without cloning

npx --yes --package github:divyanshu-iitian/SearchForge \
  searchforge search "latest open-source agent frameworks"

The first run downloads and builds the package from GitHub. Searches use intent-aware auto routing unless you select a category.

Zero-key local CLI

git clone https://github.com/divyanshu-iitian/SearchForge.git
cd SearchForge
npm install
npm run build

node dist/cli.js search "latest open-source agent frameworks"
node dist/cli.js search "retrieval augmented generation" --category academic
node dist/cli.js search "local LLM tooling" --category community
node dist/cli.js read "https://example.com"
node dist/cli.js doctor

Full web search with private SearXNG

docker compose up --build
curl -s http://localhost:3000/v1/search \
  -H "content-type: application/json" \
  -d '{"query":"open source vector databases","category":"web","limit":5}'

This starts SearchForge on port 3000 and a private, JSON-enabled SearXNG on port 8080. Before exposing the stack, change the SearXNG secret, set SEARCHFORGE_API_KEY, and terminate TLS at a trusted proxy.

Free tools

Search by capability

searchforge search "latest open-source agent frameworks"
searchforge search "browser agent" --category code
searchforge search "semantic reranking" --category academic --json
searchforge search "Show HN search engine" --category community

The default auto category detects code, academic, and current/community signals and queries the matching source families alongside the web fallback. Explicit categories prevent irrelevant providers from being queried. An explicit providers list overrides category routing, which is useful for evaluations.

Read a URL as Markdown

searchforge read "https://example.com/article"

read_url accepts public HTTP(S) URLs only. Credentials, localhost, private IP literals, and non-web protocols are rejected. Responses are size-bounded, timed out, and cached.

Diagnose the whole retrieval path

searchforge doctor

Doctor performs real, bounded probes and reports each provider's access tier, capability, latency, and error. A failed source produces degraded, not a misleading all-or-nothing status.

MCP

SearchForge exposes three stdio tools:

  • web_search — routed, citation-ready structured search
  • read_url — clean Markdown from a public URL
  • search_status — live capability and latency report
{
  "mcpServers": {
    "searchforge": {
      "command": "node",
      "args": ["/absolute/path/to/SearchForge/dist/mcp.js"],
      "env": {
        "SEARCHFORGE_SEARXNG_URL": "http://localhost:8080"
      }
    }
  }
}

The search and status tools return MCP structured content as well as readable text.

Run the MCP server straight from GitHub without a clone:

{
  "mcpServers": {
    "searchforge": {
      "command": "npx",
      "args": [
        "--yes",
        "--package",
        "github:divyanshu-iitian/SearchForge",
        "searchforge-mcp"
      ]
    }
  }
}

SearchForge is also published in the official MCP Registry as io.github.divyanshu-iitian/searchforge. To run the registry-backed OCI image directly from any MCP client that supports a Docker command:

{
  "mcpServers": {
    "searchforge": {
      "command": "docker",
      "args": [
        "run",
        "--rm",
        "-i",
        "ghcr.io/divyanshu-iitian/searchforge-mcp:0.2.0"
      ]
    }
  }
}

REST API

Search

POST /v1/search
Content-Type: application/json

{
  "query": "open source reranking models",
  "category": "academic",
  "limit": 8,
  "language": "en",
  "freshness": "month",
  "safeSearch": "moderate"
}
{
  "schemaVersion": "1.0",
  "query": "open source reranking models",
  "category": "academic",
  "results": [
    {
      "title": "Example work",
      "url": "https://doi.org/10.0000/example",
      "snippet": "Authors · Publisher · journal-article",
      "source": "crossref",
      "sources": ["crossref"],
      "score": 0.016393
    }
  ],
  "providers": [
    {
      "provider": "crossref",
      "ok": true,
      "latencyMs": 241,
      "resultCount": 8
    }
  ],
  "tookMs": 243,
  "cached": false
}

Read

POST /v1/read
Content-Type: application/json

{"url":"https://example.com/article"}

Other endpoints:

GET /healthz       Process liveness
GET /v1/providers  Configured capabilities and access tiers
GET /v1/doctor     Live dependency health

See the full OpenAPI contract.

TypeScript SDK

import {
  CrossrefProvider,
  GithubProvider,
  JinaReader,
  SearchForge,
} from "searchforge-rag";

const forge = new SearchForge({
  providers: [new GithubProvider(), new CrossrefProvider()],
  reader: new JinaReader(),
  timeoutMs: 8_000,
});

const evidence = await forge.search({
  query: "agentic retrieval",
  category: "academic",
  limit: 10,
});

const page = await forge.read("https://example.com/research");

Until an npm release is published:

npm install github:divyanshu-iitian/SearchForge

Provider details

Provider Capability Access Enabled
SearXNG Web Self-hosted, no vendor fee SEARCHFORGE_SEARXNG_URL
Wikipedia Web knowledge fallback No key Always
GitHub Code repositories No key; 60 unauthenticated REST requests/hour, search has tighter limits Always
Crossref Academic metadata No key; mailto recommended Always
HN Algolia Community No key; community-operated availability Always
Jina Reader URL to Markdown No key; documented no-key quota currently 20 RPM Always
Brave Search Web API key BRAVE_SEARCH_API_KEY

SearchForge intentionally does not configure public SearXNG instances. They often disable JSON or limit automated traffic; the Docker stack is the stable free path.

How it works

Agent / RAG / MCP client
           |
      validate + route
           |
  +--------+---------+-----------+
  |        |         |           |
 web      code    academic   community       read_url
  |        |         |           |              |
SearXNG  GitHub   Crossref   Hacker News    Jina Reader
Wikipedia
  +--------+---------+-----------+
           |
 normalize -> canonicalize -> deduplicate -> reciprocal-rank fusion
           |
 versioned evidence + provenance + per-source health

Each idempotent provider call has its own abortable timeout. One outage cannot erase healthy results. Tracking parameters are removed before deduplication, and every contributing provider remains in sources.

This capability-first design is inspired by Agent Reach. Agent Reach helps an agent operate many upstream tools directly; SearchForge complements that approach with one stable, embeddable retrieval API for RAG applications.

Configuration

Variable Default Purpose
SEARCHFORGE_SEARXNG_URL unset Private SearXNG base URL
GITHUB_TOKEN unset Optional GitHub quota increase
CROSSREF_MAILTO unset Crossref polite-pool identity
BRAVE_SEARCH_API_KEY unset Optional Brave backend
SEARCHFORGE_API_KEY unset REST bearer or x-api-key
SEARCHFORGE_PORT 3000 REST port
SEARCHFORGE_HOST 127.0.0.1 Bind address
SEARCHFORGE_TIMEOUT_MS 8000 Per-dependency timeout
SEARCHFORGE_CACHE_TTL_MS 300000 In-memory cache TTL
SEARCHFORGE_CACHE_MAX_ENTRIES 500 Cache entry bound
SEARCHFORGE_RATE_LIMIT 60 Requests/client/minute

Production boundary

  • Set an API key before binding to a public interface.
  • Search results and page content are untrusted input; delimit them and apply prompt-injection defenses.
  • The built-in cache and rate limiter are process-local. Use shared infrastructure for multiple replicas.
  • Provider bodies and credentials are excluded from surfaced errors.
  • healthz proves the process is alive; /v1/doctor checks dependencies.

See SECURITY.md, CONTRIBUTING.md, and CHANGELOG.md.

Principles

  1. Evidence over generated answers
  2. Free and self-hosted paths before vendor lock-in
  3. Partial results over total failure
  4. Honest capability and quota reporting
  5. Stable contracts and explicit provenance
  6. No telemetry by default

License

MIT © Divyanshu.

If SearchForge helps your agent, star the repository and share your integration in Discussions.

Releases

Packages

Contributors

Languages