spicrawlspicrawlDocs

Agent frameworks

Give an agent built with the OpenAI Agents SDK, Vercel AI SDK, LangChain, CrewAI, Google ADK, Mastra or smolagents the 25 spicrawl_* tools by connecting to the hosted Spicrawl MCP server.

Every framework below connects to the hosted MCP server and hands the model its 25 spicrawl_* tools: spicrawl_scrape, the spicrawl_batch_* tools, sessions, usage and docs search. Each tool call runs under your API key, with the same scopes, limits and credits as a direct API call.

Call the REST API directly instead when your code already knows which URL to fetch and no model decides. A plain function that wraps POST /v1/scrape costs no tool-schema tokens and is easier to test. Best practices has a reference fetch_page tool for that.

Each example runs one task: read https://example.com/pricing as markdown and list the plans. Set SPICRAWL_API_KEY and your model provider's key (OPENAI_API_KEY in these examples) before you run them.

SettingValue
URLhttps://mcp.spicrawl.com/mcp
TransportStreamable HTTP
HeaderAuthorization: Bearer $SPICRAWL_API_KEY
AuthAPI key only. There is no OAuth flow.

These APIs move quickly. The TypeScript examples were checked against ai 7.0 with @ai-sdk/mcp 2.0 and @mastra/mcp 2.1. If an import fails, check the framework's own MCP docs for the current name.

OpenAI Agents SDK

Install openai-agents. MCPServerStreamableHttp takes the URL and headers in params. create_static_tool_filter limits which tools the model sees.

openai_agents_example.py
import asyncio
import os

from agents import Agent, Runner
from agents.mcp import MCPServerStreamableHttp, create_static_tool_filter


async def main() -> None:
    async with MCPServerStreamableHttp(
        name="spicrawl",
        params={
            "url": "https://mcp.spicrawl.com/mcp",
            "headers": {"Authorization": f"Bearer {os.environ['SPICRAWL_API_KEY']}"},
        },
        cache_tools_list=True,
        tool_filter=create_static_tool_filter(
            allowed_tool_names=["spicrawl_scrape", "spicrawl_batch_submit"]
        ),
    ) as server:
        agent = Agent(
            name="Page reader",
            instructions="Use the Spicrawl tools to read web pages.",
            mcp_servers=[server],
        )
        result = await Runner.run(
            agent,
            "Read https://example.com/pricing as markdown and list the plans.",
        )
        print(result.final_output)


asyncio.run(main())

Vercel AI SDK

Install ai, @ai-sdk/mcp and a provider such as @ai-sdk/openai. In AI SDK 7 the client is createMCPClient from @ai-sdk/mcp (older releases exported it as experimental_createMCPClient from ai). The transport type for Streamable HTTP is "http". generateText stops after one step by default, so set stopWhen to let the model call a tool and then answer.

ai-sdk-example.ts
import { createMCPClient } from "@ai-sdk/mcp";
import { openai } from "@ai-sdk/openai";
import { generateText, isStepCount } from "ai";

const mcp = await createMCPClient({
  transport: {
    type: "http",
    url: "https://mcp.spicrawl.com/mcp",
    headers: { Authorization: `Bearer ${process.env.SPICRAWL_API_KEY}` },
  },
});

try {
  const all = await mcp.tools();
  const keep = ["spicrawl_scrape", "spicrawl_batch_submit"];
  const tools = Object.fromEntries(
    Object.entries(all).filter(([name]) => keep.includes(name)),
  );

  const { text } = await generateText({
    model: openai("gpt-5"),
    tools,
    stopWhen: isStepCount(5),
    prompt: "Read https://example.com/pricing as markdown and list the plans.",
  });
  console.log(text);
} finally {
  await mcp.close();
}

The AI SDK requires Node.js 22 or later. Close the client when the run ends, as above.

LangChain and LangGraph

Install langchain-mcp-adapters and langchain[openai]. MultiServerMCPClient takes one entry per server; "transport": "streamable_http" and headers configure the connection. create_agent builds a LangGraph agent from the tools.

langchain_example.py
import asyncio
import os

from langchain.agents import create_agent
from langchain_mcp_adapters.client import MultiServerMCPClient


async def main() -> None:
    client = MultiServerMCPClient(
        {
            "spicrawl": {
                "transport": "streamable_http",
                "url": "https://mcp.spicrawl.com/mcp",
                "headers": {"Authorization": f"Bearer {os.environ['SPICRAWL_API_KEY']}"},
            }
        }
    )
    keep = {"spicrawl_scrape", "spicrawl_batch_submit"}
    tools = [t for t in await client.get_tools() if t.name in keep]

    agent = create_agent("openai:gpt-4.1", tools)
    result = await agent.ainvoke(
        {"messages": "Read https://example.com/pricing as markdown and list the plans."}
    )
    print(result["messages"][-1].content)


asyncio.run(main())

By default each tool call opens its own MCP session. That is fine for spicrawl_scrape, which is stateless. See the adapter's README for client.session("spicrawl") if you want one session for a whole run.

CrewAI

CrewAI's current way to attach MCP servers is the mcps= field on an agent. MCPServerHTTP takes the URL and headers and uses Streamable HTTP by default (streamable=True). create_static_tool_filter limits the tools.

crewai_example.py
import os

from crewai import Agent, Crew, Task
from crewai.mcp import MCPServerHTTP
from crewai.mcp.filters import create_static_tool_filter

reader = Agent(
    role="Page reader",
    goal="Read web pages with the Spicrawl tools and report what they say",
    backstory="You fetch pages as markdown and summarize them accurately.",
    mcps=[
        MCPServerHTTP(
            url="https://mcp.spicrawl.com/mcp",
            headers={"Authorization": f"Bearer {os.environ['SPICRAWL_API_KEY']}"},
            streamable=True,
            cache_tools_list=True,
            tool_filter=create_static_tool_filter(
                allowed_tool_names=["spicrawl_scrape", "spicrawl_batch_submit"]
            ),
        )
    ],
)

task = Task(
    description="Read https://example.com/pricing as markdown and list the plans.",
    expected_output="A list of the plan names on the page.",
    agent=reader,
)

result = Crew(agents=[reader], tasks=[task]).kickoff()
print(result)

Google ADK

Install google-adk. McpToolset takes StreamableHTTPConnectionParams with a headers dict. tool_filter accepts a list of tool names. The example uses Gemini, so set GOOGLE_API_KEY.

adk_example.py
import asyncio
import os

from google.adk.agents import LlmAgent
from google.adk.runners import InMemoryRunner
from google.adk.tools.mcp_tool import McpToolset, StreamableHTTPConnectionParams
from google.genai import types


async def main() -> None:
    toolset = McpToolset(
        connection_params=StreamableHTTPConnectionParams(
            url="https://mcp.spicrawl.com/mcp",
            headers={"Authorization": f"Bearer {os.environ['SPICRAWL_API_KEY']}"},
        ),
        tool_filter=["spicrawl_scrape", "spicrawl_batch_submit"],
    )
    agent = LlmAgent(
        name="page_reader",
        model="gemini-2.5-flash",
        instruction="Use the Spicrawl tools to read web pages.",
        tools=[toolset],
    )
    runner = InMemoryRunner(agent=agent, app_name="spicrawl_example")
    session = await runner.session_service.create_session(
        app_name="spicrawl_example", user_id="user"
    )
    message = types.Content(
        role="user",
        parts=[types.Part(text="Read https://example.com/pricing as markdown and list the plans.")],
    )
    try:
        async for event in runner.run_async(
            user_id="user", session_id=session.id, new_message=message
        ):
            if event.is_final_response() and event.content:
                print(event.content.parts[0].text)
    finally:
        await toolset.close()


asyncio.run(main())

Mastra

Install @mastra/mcp and @mastra/core. MCPClient takes a servers map; each entry has a url (a URL object) and requestInit.headers. listTools() prefixes each tool name with the server key, so the model sees spicrawl_spicrawl_scrape.

mastra-example.ts
import { Agent } from "@mastra/core/agent";
import { MCPClient } from "@mastra/mcp";

const mcp = new MCPClient({
  servers: {
    spicrawl: {
      url: new URL("https://mcp.spicrawl.com/mcp"),
      requestInit: {
        headers: { Authorization: `Bearer ${process.env.SPICRAWL_API_KEY}` },
      },
    },
  },
});

const agent = new Agent({
  id: "page-reader",
  name: "Page reader",
  instructions: "Use the Spicrawl tools to read web pages.",
  model: "openai/gpt-5",
  tools: await mcp.listTools(),
});

try {
  const result = await agent.generate(
    "Read https://example.com/pricing as markdown and list the plans.",
  );
  console.log(result.text);
} finally {
  await mcp.disconnect();
}

To limit the tools, filter the object listTools() returns by key before you pass it to the agent, as in the AI SDK example.

Hugging Face smolagents

Install smolagents[mcp]. MCPClient takes a dict with url and transport; "streamable-http" is the default. The dict's other keys, headers included, go to the MCP SDK's Streamable HTTP client. The example uses Hugging Face Inference Providers, so set HF_TOKEN.

smolagents_example.py
import os

from smolagents import CodeAgent, InferenceClientModel, MCPClient

server_parameters = {
    "url": "https://mcp.spicrawl.com/mcp",
    "transport": "streamable-http",
    "headers": {"Authorization": f"Bearer {os.environ['SPICRAWL_API_KEY']}"},
}

with MCPClient(server_parameters, structured_output=True) as tools:
    tools = [t for t in tools if t.name in {"spicrawl_scrape", "spicrawl_batch_submit"}]
    agent = CodeAgent(tools=tools, model=InferenceClientModel())
    print(agent.run("Read https://example.com/pricing as markdown and list the plans."))

Troubleshooting

SymptomCauseFix
HTTP 401, or Unauthorized: send your Spicrawl API key, when the client connectsThe header is missing, has a literal ${...}, or carries an unknown, revoked or expired key.Build the header in code from SPICRAWL_API_KEY, as above, and check the variable is set in the process that runs the agent. Test with curl -i https://mcp.spicrawl.com/mcp -H "Authorization: Bearer $SPICRAWL_API_KEY".
The client starts an OAuth flow, or its MCP option has no place for headersThe client supports only OAuth. The Spicrawl server accepts API keys only.Use the framework's headers option (every example above has one), or a different connector.
HTTP 503 with Retry-After on connectThe MCP server could not reach the API to check your key.Retry after the given seconds. Your key is fine.
A misspelt argument error such as js_render on spicrawl_scrapeThe tool uses render, not the API's js_render, and format, not response_format.Tell the model the tool's argument names, or keep the tool list short so it reads the schema. See MCP server.
Prompt cost is high, or the model picks the wrong toolAll 25 tool schemas are sent on every model call.Keep only spicrawl_scrape and, for many URLs, spicrawl_batch_submit, spicrawl_batch_status and spicrawl_batch_results. The examples filter by name; OpenAI Agents SDK, CrewAI and Google ADK have a built-in tool_filter, and the others filter the returned tool list.
ERR::AUTH::INSUFFICIENT_SCOPE from a spicrawl_usage* toolThe key lacks the read scope.Grant read in the dashboard, or leave those tools out.
The run ends after one step with no answer (AI SDK)generateText stops after one step by default.Set stopWhen: isStepCount(5) or higher.

Tool errors reach the model as MCP tool errors with the API's code. Failed calls cost 0 credits. See How errors reach the agent and Errors.

On this page