← Blog

Claude API Document QA Chatbot Tutorial

2026-10-04 · 5 min read · SubToAPI Team

What you'll build

This tutorial shows you how to build a chatbot that answers questions about a specific document — a PDF manual, a set of internal docs, a contract, a knowledge base article — using the Claude API. By the end you'll have a working Node.js script that takes a document, a user question, and returns a grounded answer with Claude doing the reasoning.

The core technique is simple and doesn't require a vector database for small-to-medium documents: you load the document text, pass it to Claude as context in the system prompt or user message, and let Claude's large context window do the retrieval work internally. For larger document sets you'll add chunking and a lightweight retrieval step before calling Claude. We'll cover both approaches.

Step 1: Prepare your document

Start by converting your source document into plain text. If it's a PDF, use a library like pdf-parse; if it's HTML, strip tags; if it's already markdown or .txt, you're set.

import fs from "fs";
import pdf from "pdf-parse";

const buffer = fs.readFileSync("manual.pdf");
const data = await pdf(buffer);
const documentText = data.text;

For documents under roughly 50,000 words, you can pass the whole thing directly to Claude in one request — modern Claude models handle very long context windows comfortably. For bigger corpora (multiple documents, large wikis), skip ahead to the chunking section.

Step 2: Single-document QA (no chunking needed)

The simplest and most reliable pattern for a single document under the context limit is to put the entire document in the system prompt and ask Claude to answer only based on it.

const systemPrompt = `You are a document QA assistant. Answer questions using ONLY the document below.
If the answer isn't in the document, say "I don't have that information in the document."

DOCUMENT:
${documentText}`;

const response = await fetch("https://api.subtoapi.app/v1/messages", {
  method: "POST",
  headers: {
    "Authorization": `Bearer ${process.env.SUBTOAPI_KEY}`,
    "Content-Type": "application/json"
  },
  body: JSON.stringify({
    model: "claude-sonnet-4-5",
    max_tokens: 1024,
    system: systemPrompt,
    messages: [
      { role: "user", content: "What is the warranty period for this product?" }
    ]
  })
});

const result = await response.json();
console.log(result.content[0].text);

This pattern works well because it forces Claude to ground every answer in the supplied text instead of relying on general knowledge, which cuts down hallucinations significantly. The instruction "say you don't know" is doing real work here — without it, models tend to guess.

Step 3: Chunking for large document sets

Once your content exceeds what comfortably fits in a single context window, or you're indexing dozens of documents, you need a retrieval step before calling Claude. The pattern is: chunk → embed → retrieve top-k chunks → pass only those chunks to Claude.

function chunkText(text, chunkSize = 1000, overlap = 200) {
  const chunks = [];
  let start = 0;
  while (start < text.length) {
    const end = start + chunkSize;
    chunks.push(text.slice(start, end));
    start += chunkSize - overlap;
  }
  return chunks;
}

const chunks = chunkText(documentText);

Embed each chunk with any embedding provider, store the vectors (an in-memory array is fine for a prototype, a vector DB like Pinecone or pgvector for production), and at query time retrieve the top 3–5 most relevant chunks by cosine similarity. Then build your Claude prompt from only those chunks:

const relevantChunks = retrieveTopChunks(userQuestion, chunks, 4);

const contextBlock = relevantChunks
  .map((c, i) => `[Excerpt ${i + 1}]\n${c}`)
  .join("\n\n");

const systemPrompt = `Answer the user's question using only the excerpts below. Cite which excerpt you used.

${contextBlock}`;

This keeps token usage low and lets you scale to document libraries far bigger than any single context window, while still letting Claude do the actual answer synthesis rather than a brittle keyword search.

Step 4: Streaming for a better chat UX

A document QA chatbot feels much more responsive when answers stream token by token instead of arriving all at once, especially for longer answers. The API supports server-sent events for this — see the streaming docs for the full event format.

const stream = await fetch("https://api.subtoapi.app/v1/messages", {
  method: "POST",
  headers: {
    "Authorization": `Bearer ${process.env.SUBTOAPI_KEY}`,
    "Content-Type": "application/json"
  },
  body: JSON.stringify({
    model: "claude-sonnet-4-5",
    max_tokens: 1024,
    stream: true,
    system: systemPrompt,
    messages: [{ role: "user", content: userQuestion }]
  })
});

for await (const chunk of stream.body) {
  process.stdout.write(chunk.toString());
}

Parse the SSE payload on the frontend to render tokens as they arrive — most chat UI libraries (or a simple EventSource-style reader) handle this with a few lines of code.

Step 5: Multi-turn conversations with document context

Real chatbots need to remember prior turns. Keep the document in the system prompt (it doesn't change per turn) and append each question/answer pair to the messages array:

const messages = [
  { role: "user", content: "What's covered under the warranty?" },
  { role: "assistant", content: "The warranty covers manufacturing defects for 12 months..." },
  { role: "user", content: "Does that include water damage?" }
];

Claude will use the conversation history plus the document context to resolve follow-up questions like "does that include..." correctly. Trim older turns once the conversation gets long to keep token usage predictable — see the messages documentation for request structure details.

Why route this through SubToAPI

If you're already paying for Claude access and want a clean HTTPS endpoint for this kind of app instead of managing SDK auth per environment, SubToAPI turns your existing Claude access into application API keys (sub_live_...) with streaming, usage metadata, and team seats built in. That's useful once your document QA chatbot moves from a prototype script into something your team or customers actually use — you get per-key usage tracking without building your own billing layer. Check the quickstart to get an API key in a couple of minutes, or compare plans if you're deciding between Solo and Team.

Common pitfalls to avoid

Questions

Do I need a vector database to build a document QA chatbot with Claude? No, not for small documents. If your content fits comfortably in the context window, pass it directly in the system prompt. A vector database becomes useful once you're indexing many documents or very large corpora.

How do I stop Claude from hallucinating answers not in the document? Explicitly instruct it to answer only from the supplied text and to say when it doesn't know. Combine this with retrieval of relevant chunks so irrelevant context doesn't lead it astray.

Can this chatbot handle PDFs, Word docs, and HTML pages? Yes — convert any format to plain text first using a parsing library (pdf-parse, mammoth, or an HTML stripper), then feed the resulting text through the same chunking and prompting pipeline.

Turn your Claude access into an HTTPS API

SubToAPI gives you application API keys, streaming, tool use and usage insights on top of your existing Claude access — set up in minutes.

Start free  Read the quickstart →