Agent frameworks
Give an agent built with the OpenAI Agents SDK, Vercel AI SDK, LangChain, CrewAI, Google ADK, Mastra or smolagents the 25 spicrawl_* tools by connecting to the hosted Spicrawl MCP server.
Every framework below connects to the hosted MCP server and hands the model its 25 spicrawl_* tools: spicrawl_scrape, the spicrawl_batch_* tools, sessions, usage and docs search. Each tool call runs under your API key, with the same scopes, limits and credits as a direct API call.
Call the REST API directly instead when your code already knows which URL to fetch and no model decides. A plain function that wraps POST /v1/scrape costs no tool-schema tokens and is easier to test. Best practices has a reference fetch_page tool for that.
Each example runs one task: read https://example.com/pricing as markdown and list the plans. Set SPICRAWL_API_KEY and your model provider's key (OPENAI_API_KEY in these examples) before you run them.
| Setting | Value |
|---|---|
| URL | https://mcp.spicrawl.com/mcp |
| Transport | Streamable HTTP |
| Header | Authorization: Bearer $SPICRAWL_API_KEY |
| Auth | API key only. There is no OAuth flow. |
These APIs move quickly. The TypeScript examples were checked against ai 7.0 with @ai-sdk/mcp 2.0 and @mastra/mcp 2.1. If an import fails, check the framework's own MCP docs for the current name.
OpenAI Agents SDK
Install openai-agents. MCPServerStreamableHttp takes the URL and headers in params. create_static_tool_filter limits which tools the model sees.
import asyncio
import os
from agents import Agent, Runner
from agents.mcp import MCPServerStreamableHttp, create_static_tool_filter
async def main() -> None:
async with MCPServerStreamableHttp(
name="spicrawl",
params={
"url": "https://mcp.spicrawl.com/mcp",
"headers": {"Authorization": f"Bearer {os.environ['SPICRAWL_API_KEY']}"},
},
cache_tools_list=True,
tool_filter=create_static_tool_filter(
allowed_tool_names=["spicrawl_scrape", "spicrawl_batch_submit"]
),
) as server:
agent = Agent(
name="Page reader",
instructions="Use the Spicrawl tools to read web pages.",
mcp_servers=[server],
)
result = await Runner.run(
agent,
"Read https://example.com/pricing as markdown and list the plans.",
)
print(result.final_output)
asyncio.run(main())Vercel AI SDK
Install ai, @ai-sdk/mcp and a provider such as @ai-sdk/openai. In AI SDK 7 the client is createMCPClient from @ai-sdk/mcp (older releases exported it as experimental_createMCPClient from ai). The transport type for Streamable HTTP is "http". generateText stops after one step by default, so set stopWhen to let the model call a tool and then answer.
import { createMCPClient } from "@ai-sdk/mcp";
import { openai } from "@ai-sdk/openai";
import { generateText, isStepCount } from "ai";
const mcp = await createMCPClient({
transport: {
type: "http",
url: "https://mcp.spicrawl.com/mcp",
headers: { Authorization: `Bearer ${process.env.SPICRAWL_API_KEY}` },
},
});
try {
const all = await mcp.tools();
const keep = ["spicrawl_scrape", "spicrawl_batch_submit"];
const tools = Object.fromEntries(
Object.entries(all).filter(([name]) => keep.includes(name)),
);
const { text } = await generateText({
model: openai("gpt-5"),
tools,
stopWhen: isStepCount(5),
prompt: "Read https://example.com/pricing as markdown and list the plans.",
});
console.log(text);
} finally {
await mcp.close();
}The AI SDK requires Node.js 22 or later. Close the client when the run ends, as above.
LangChain and LangGraph
Install langchain-mcp-adapters and langchain[openai]. MultiServerMCPClient takes one entry per server; "transport": "streamable_http" and headers configure the connection. create_agent builds a LangGraph agent from the tools.
import asyncio
import os
from langchain.agents import create_agent
from langchain_mcp_adapters.client import MultiServerMCPClient
async def main() -> None:
client = MultiServerMCPClient(
{
"spicrawl": {
"transport": "streamable_http",
"url": "https://mcp.spicrawl.com/mcp",
"headers": {"Authorization": f"Bearer {os.environ['SPICRAWL_API_KEY']}"},
}
}
)
keep = {"spicrawl_scrape", "spicrawl_batch_submit"}
tools = [t for t in await client.get_tools() if t.name in keep]
agent = create_agent("openai:gpt-4.1", tools)
result = await agent.ainvoke(
{"messages": "Read https://example.com/pricing as markdown and list the plans."}
)
print(result["messages"][-1].content)
asyncio.run(main())By default each tool call opens its own MCP session. That is fine for spicrawl_scrape, which is stateless. See the adapter's README for client.session("spicrawl") if you want one session for a whole run.
CrewAI
CrewAI's current way to attach MCP servers is the mcps= field on an agent. MCPServerHTTP takes the URL and headers and uses Streamable HTTP by default (streamable=True). create_static_tool_filter limits the tools.
import os
from crewai import Agent, Crew, Task
from crewai.mcp import MCPServerHTTP
from crewai.mcp.filters import create_static_tool_filter
reader = Agent(
role="Page reader",
goal="Read web pages with the Spicrawl tools and report what they say",
backstory="You fetch pages as markdown and summarize them accurately.",
mcps=[
MCPServerHTTP(
url="https://mcp.spicrawl.com/mcp",
headers={"Authorization": f"Bearer {os.environ['SPICRAWL_API_KEY']}"},
streamable=True,
cache_tools_list=True,
tool_filter=create_static_tool_filter(
allowed_tool_names=["spicrawl_scrape", "spicrawl_batch_submit"]
),
)
],
)
task = Task(
description="Read https://example.com/pricing as markdown and list the plans.",
expected_output="A list of the plan names on the page.",
agent=reader,
)
result = Crew(agents=[reader], tasks=[task]).kickoff()
print(result)Google ADK
Install google-adk. McpToolset takes StreamableHTTPConnectionParams with a headers dict. tool_filter accepts a list of tool names. The example uses Gemini, so set GOOGLE_API_KEY.
import asyncio
import os
from google.adk.agents import LlmAgent
from google.adk.runners import InMemoryRunner
from google.adk.tools.mcp_tool import McpToolset, StreamableHTTPConnectionParams
from google.genai import types
async def main() -> None:
toolset = McpToolset(
connection_params=StreamableHTTPConnectionParams(
url="https://mcp.spicrawl.com/mcp",
headers={"Authorization": f"Bearer {os.environ['SPICRAWL_API_KEY']}"},
),
tool_filter=["spicrawl_scrape", "spicrawl_batch_submit"],
)
agent = LlmAgent(
name="page_reader",
model="gemini-2.5-flash",
instruction="Use the Spicrawl tools to read web pages.",
tools=[toolset],
)
runner = InMemoryRunner(agent=agent, app_name="spicrawl_example")
session = await runner.session_service.create_session(
app_name="spicrawl_example", user_id="user"
)
message = types.Content(
role="user",
parts=[types.Part(text="Read https://example.com/pricing as markdown and list the plans.")],
)
try:
async for event in runner.run_async(
user_id="user", session_id=session.id, new_message=message
):
if event.is_final_response() and event.content:
print(event.content.parts[0].text)
finally:
await toolset.close()
asyncio.run(main())Mastra
Install @mastra/mcp and @mastra/core. MCPClient takes a servers map; each entry has a url (a URL object) and requestInit.headers. listTools() prefixes each tool name with the server key, so the model sees spicrawl_spicrawl_scrape.
import { Agent } from "@mastra/core/agent";
import { MCPClient } from "@mastra/mcp";
const mcp = new MCPClient({
servers: {
spicrawl: {
url: new URL("https://mcp.spicrawl.com/mcp"),
requestInit: {
headers: { Authorization: `Bearer ${process.env.SPICRAWL_API_KEY}` },
},
},
},
});
const agent = new Agent({
id: "page-reader",
name: "Page reader",
instructions: "Use the Spicrawl tools to read web pages.",
model: "openai/gpt-5",
tools: await mcp.listTools(),
});
try {
const result = await agent.generate(
"Read https://example.com/pricing as markdown and list the plans.",
);
console.log(result.text);
} finally {
await mcp.disconnect();
}To limit the tools, filter the object listTools() returns by key before you pass it to the agent, as in the AI SDK example.
Hugging Face smolagents
Install smolagents[mcp]. MCPClient takes a dict with url and transport; "streamable-http" is the default. The dict's other keys, headers included, go to the MCP SDK's Streamable HTTP client. The example uses Hugging Face Inference Providers, so set HF_TOKEN.
import os
from smolagents import CodeAgent, InferenceClientModel, MCPClient
server_parameters = {
"url": "https://mcp.spicrawl.com/mcp",
"transport": "streamable-http",
"headers": {"Authorization": f"Bearer {os.environ['SPICRAWL_API_KEY']}"},
}
with MCPClient(server_parameters, structured_output=True) as tools:
tools = [t for t in tools if t.name in {"spicrawl_scrape", "spicrawl_batch_submit"}]
agent = CodeAgent(tools=tools, model=InferenceClientModel())
print(agent.run("Read https://example.com/pricing as markdown and list the plans."))Troubleshooting
| Symptom | Cause | Fix |
|---|---|---|
HTTP 401, or Unauthorized: send your Spicrawl API key, when the client connects | The header is missing, has a literal ${...}, or carries an unknown, revoked or expired key. | Build the header in code from SPICRAWL_API_KEY, as above, and check the variable is set in the process that runs the agent. Test with curl -i https://mcp.spicrawl.com/mcp -H "Authorization: Bearer $SPICRAWL_API_KEY". |
| The client starts an OAuth flow, or its MCP option has no place for headers | The client supports only OAuth. The Spicrawl server accepts API keys only. | Use the framework's headers option (every example above has one), or a different connector. |
HTTP 503 with Retry-After on connect | The MCP server could not reach the API to check your key. | Retry after the given seconds. Your key is fine. |
A misspelt argument error such as js_render on spicrawl_scrape | The tool uses render, not the API's js_render, and format, not response_format. | Tell the model the tool's argument names, or keep the tool list short so it reads the schema. See MCP server. |
| Prompt cost is high, or the model picks the wrong tool | All 25 tool schemas are sent on every model call. | Keep only spicrawl_scrape and, for many URLs, spicrawl_batch_submit, spicrawl_batch_status and spicrawl_batch_results. The examples filter by name; OpenAI Agents SDK, CrewAI and Google ADK have a built-in tool_filter, and the others filter the returned tool list. |
ERR::AUTH::INSUFFICIENT_SCOPE from a spicrawl_usage* tool | The key lacks the read scope. | Grant read in the dashboard, or leave those tools out. |
| The run ends after one step with no answer (AI SDK) | generateText stops after one step by default. | Set stopWhen: isStepCount(5) or higher. |
Tool errors reach the model as MCP tool errors with the API's code. Failed calls cost 0 credits. See How errors reach the agent and Errors.
AI app builders
Connect Spicrawl's hosted MCP server to Replit Agent, v0, Lovable and Bolt so the builder can read web pages, and call the Spicrawl API from the app it builds.
Markdown for LLMs
Turn any page or PDF into main-content markdown with response_format=markdown, and cut tokens with include_tags, exclude_tags and main_content_only.