← Blog

Build a Document Q&A Chatbot with the Claude API

2026-10-06 · 5 min read · SubToAPI Team

What you're actually building

A "document Q&A chatbot" is a system that takes a user's question, finds the relevant parts of one or more documents, and asks Claude to answer using only that retrieved context. It's not a single API call — it's a small pipeline: ingest documents, chunk and index them, retrieve relevant chunks at query time, build a prompt, call Claude, and stream the answer back to the user.

This guide walks through that pipeline end to end, with working code, so you can go from "I have some PDFs" to "I have a chat interface that answers questions about them" in an afternoon. We'll use plain retrieval-augmented generation (RAG) with Claude as the answering model, since that's the standard, reliable architecture for this use case.

The architecture in five steps

  1. Ingest — extract text from PDFs, markdown, HTML, or whatever source format you have.
  2. Chunk — split text into overlapping pieces small enough to fit in a prompt alongside the question.
  3. Embed and index — convert chunks to vectors and store them so you can retrieve by semantic similarity.
  4. Retrieve — at query time, find the top N chunks most relevant to the user's question.
  5. Generate — send the question plus retrieved chunks to Claude, with a system prompt that constrains it to answer from the provided context.

Steps 1–3 happen once per document (or whenever documents change). Steps 4–5 happen on every user question.

Chunking documents

Chunk size matters more than people expect. Too large and you waste context and dilute relevance; too small and you lose coherence. 500–1000 tokens with 10–15% overlap is a reasonable default for most prose documents.

function chunkText(text, chunkSize = 800, overlap = 100) {
  const words = text.split(/\s+/);
  const chunks = [];
  let i = 0;
  while (i < words.length) {
    chunks.push(words.slice(i, i + chunkSize).join(" "));
    i += chunkSize - overlap;
  }
  return chunks;
}

Store each chunk with metadata — document name, page number, section heading — so you can cite sources in the final answer.

Embedding and retrieval

You don't need Claude for embeddings; use whatever embedding model fits your budget and store vectors in a vector database (Pinecone, pgvector, Qdrant, or even an in-memory array for small datasets). At query time, embed the user's question and retrieve the top 3–8 chunks by cosine similarity.

async function retrieveChunks(question, index, topK = 5) {
  const queryVector = await embed(question);
  return index.query(queryVector, topK); // returns [{ text, source, score }]
}

Building the prompt

This is where most document Q&A bots go wrong. Put the retrieved chunks in the user message, not scattered through the conversation, and give Claude an explicit instruction to stick to the provided context and say when it doesn't know.

const systemPrompt = `You are a document Q&A assistant. Answer the user's
question using only the context provided below. If the answer isn't in the
context, say so clearly — do not guess. Cite the source document for each
claim you make.`;

function buildUserMessage(question, chunks) {
  const context = chunks
    .map((c, i) => `[${i + 1}] (${c.source}) ${c.text}`)
    .join("\n\n");
  return `Context:\n${context}\n\nQuestion: ${question}`;
}

Calling the model

Once you have the system prompt and context-stuffed user message, the API call itself is a standard message request. If you're calling Claude directly through SubToAPI, the request shape looks like this:

curl https://api.subtoapi.app/v1/messages \
  -H "Authorization: Bearer $SUBTOAPI_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "claude-sonnet",
    "max_tokens": 1024,
    "system": "You are a document Q&A assistant...",
    "messages": [
      { "role": "user", "content": "Context:\n[1] ... \n\nQuestion: What is the refund window?" }
    ]
  }'

See /docs/messages for the full request and response reference.

Streaming the answer

For a chatbot UI, you want tokens appearing as they're generated rather than waiting for the full response. Use the streaming endpoint and render chunks as they arrive — the setup is covered in /docs/streaming, and the pattern is the same regardless of document length: the retrieval step runs first and finishes before streaming starts, since you need the chunks in hand before you build the prompt.

const response = await fetch("https://api.subtoapi.app/v1/messages", {
  method: "POST",
  headers: {
    Authorization: `Bearer ${process.env.SUBTOAPI_KEY}`,
    "Content-Type": "application/json",
  },
  body: JSON.stringify({
    model: "claude-sonnet",
    max_tokens: 1024,
    stream: true,
    system: systemPrompt,
    messages: [{ role: "user", content: buildUserMessage(question, chunks) }],
  }),
});

for await (const chunk of response.body) {
  process.stdout.write(decoder.decode(chunk));
}

Handling multi-turn conversations

Real Q&A chatbots need follow-up questions ("what about for annual plans?") to work without re-explaining context. Two practical approaches:

The second approach keeps token usage lower on long conversations, which matters if you're running this at scale — usage adds up fast once dozens of users are asking questions against large document sets.

Where SubToAPI fits

If you're building this as a feature inside a product rather than a one-off script, you'll want a clean way to issue API keys per environment or per customer, track usage, and avoid wiring raw provider credentials into your app. SubToAPI turns your Claude access into a standard HTTPS API with sub_live_ application keys, so your document Q&A service calls /v1/messages the same way regardless of which account or team it belongs to. Usage metadata per key makes it straightforward to see which documents or customers are driving the most API calls — useful once you move past a prototype. Start with the quickstart at /docs/quickstart, check plans at /pricing, or just /signup and get a key.

questions

Do I need a vector database for a small document set? No. For a handful of documents under a few hundred pages, an in-memory array of embeddings with cosine similarity is fine. Move to a real vector database once retrieval latency or dataset size becomes a problem.

How many chunks should I send to Claude per question? Start with 3–5 chunks of 500–800 tokens each. Measure answer quality and adjust — too few chunks causes missed context, too many dilutes relevance and increases cost and latency.

Can Claude read a whole PDF directly instead of chunking it? For short documents that fit comfortably within the context window, yes — skip retrieval and send the full text. Chunking and retrieval become necessary once your document set exceeds what fits in a single prompt alongside the conversation.

Turn your Claude access into an HTTPS API

SubToAPI gives you application API keys, streaming, tool use and usage insights on top of your existing Claude access — set up in minutes.

Start free  Read the quickstart →