cd ../blog
GenAIAugust 23, 202612 min read

RAG on Amazon Bedrock: The Architecture That Survives Production

A practical walkthrough of building Retrieval-Augmented Generation on Amazon Bedrock — Knowledge Bases, vector store selection, chunking strategy, hybrid retrieval and reranking, Guardrails, evaluation, Terraform, and the cost model nobody shows you in the demo.

Roger Vasconcelos
Roger Vasconcelos
AWS DevOps Architect
RAG on Amazon Bedrock: The Architecture That Survives Production

Every organization I talk to wants the same thing from generative AI: answers grounded in their data — the runbooks, the contracts, the support tickets, the ten years of Confluence nobody has read since 2019. And almost every one of them starts from the same wrong assumption: that this is a model problem.

It isn't. A foundation model that has never seen your incident postmortems will confidently invent them. Fine-tuning teaches style and format far better than it teaches facts, and it goes stale the moment someone edits a document. Retrieval-Augmented Generation (RAG) fixes the actual problem: it retrieves the relevant passages from your corpus at query time and puts them in front of the model as context, so the answer is generated from your documents rather than from the model's memory of the internet.

Amazon Bedrock makes the happy path genuinely easy — you can have a working knowledge base in an afternoon. The gap between that afternoon and something you would put in front of customers is where this article lives.

What Bedrock actually gives you

Bedrock is not one service, it is a set of components you can adopt independently:

  • Model access — Anthropic Claude, Amazon Nova and Titan, Meta Llama, Mistral, Cohere and others behind a single API, with no GPU capacity to manage.
  • Knowledge Bases — the managed RAG pipeline: it reads a data source, chunks it, embeds it, writes vectors to a vector store, and exposes Retrieve and RetrieveAndGenerate APIs.
  • Guardrails — content filters, denied topics, PII redaction, and a contextual grounding check that scores whether the answer is actually supported by the retrieved passages.
  • Evaluations — LLM-as-a-judge scoring for both retrieval quality and end-to-end response quality.
  • Agents — multi-step orchestration when a single retrieval pass isn't enough.
  • The important design decision is how much of the pipeline you hand to Knowledge Bases. Managed ingestion is the right default: it removes an entire class of undifferentiated plumbing. But it is opinionated, and once your chunking or metadata needs get specific, you will want the Retrieve API with your own generation call rather than the fully managed RetrieveAndGenerate. Build so that swap is a one-line change.

    The reference architecture

    Two paths, and they fail for different reasons — keep them separate in your head and in your code.

    Ingestion path (asynchronous, batch): source documents land in S3 → an ingestion job chunks each document → each chunk is embedded → vectors plus metadata are written to the vector store. This runs on document change, not on user request.

    Query path (synchronous, latency-critical): user question → embed the question → vector search with metadata filters → optional rerank → assemble prompt with the top passages → generate → apply guardrails → return the answer with citations.

    Almost every "the AI is hallucinating" complaint in production is really a retrieval failure on the query path, or a chunking failure on the ingestion path. The model is usually the least broken part of the system.

    Choosing the vector store

    Bedrock Knowledge Bases can back onto several stores, and this choice is harder to reverse than any other, so make it deliberately.

  • OpenSearch Serverless — the default, and the one Bedrock will create for you. Hybrid search (vector plus BM25 keyword) is built in, which matters more than most teams expect. The catch is the minimum OCU billing: an idle collection is not a cheap collection, so it's a poor fit for a low-traffic internal proof of concept.
  • Aurora PostgreSQL Serverless v2 with pgvector — the pragmatic choice when you already run Postgres. Your vectors sit next to your relational data, you get real SQL for metadata filtering, and Serverless v2 scales down when nobody is asking questions. Index tuning (HNSW parameters, maintenance_work_mem during index builds) is on you.
  • Amazon Neptune Analytics — for GraphRAG, when relationships between entities carry as much meaning as the text itself. Powerful and narrow; don't reach for it first.
  • Pinecone, MongoDB Atlas, Redis — valid when the team already operates one of them. The integration is real, but you are adding a vendor to your data path and your compliance review.
  • Start with Aurora pgvector for cost-sensitive internal workloads and OpenSearch Serverless when hybrid search quality and scale are the priority.

    Chunking is the highest-leverage decision you will make

    If you change one thing after reading this, change your chunking. Retrieval can only return what chunking created. A chunk that splits a procedure in half will never answer a question about that procedure, no matter how good the embedding model is.

    Bedrock offers several strategies:

  • Fixed-size — N tokens with an overlap percentage. Fast, predictable, and blunt. Fine for homogeneous prose, bad for structured documents.
  • Hierarchical — small child chunks for precise matching, larger parent chunks returned as context. This is the strategy most teams should be using and most teams have never enabled. You match on specificity and generate on completeness.
  • Semantic — splits on embedding-similarity boundaries, so chunks follow topic shifts rather than token counts. Higher ingestion cost, better coherence on long-form documents.
  • No chunking — one document, one vector. Correct when your "documents" are already atomic: FAQ entries, product records, ticket summaries.
  • Custom via Lambda — a transformation function you own. This is the escape hatch for structured content, and it's where the real wins are: keeping a Markdown table intact, prepending the section heading path to every chunk, splitting a contract on clause boundaries.
  • A concrete rule that has paid for itself on every engagement: give every chunk its provenance in the text, not just in the metadata. Prefixing a chunk with Document: Incident Response Runbook > Section: Escalation > Page 12 measurably improves both retrieval and the model's ability to cite correctly.

    Embeddings: pick dimensions on purpose

    Titan Text Embeddings V2 supports 1024, 512, and 256 dimensions. Cohere Embed offers English and multilingual variants. The instinct is to take the largest vector available; resist it. Dimension count drives vector store size, index memory, and query latency, and on many corpora 512 dimensions retrieves within noise of 1024 while costing meaningfully less to store and search.

    Two rules that are not negotiable:

  • The same model must embed both documents and queries. Mixed embedding spaces produce retrieval that looks plausible and is quietly random.
  • Changing the embedding model means a full re-index. Treat the model ID as a versioned property of the index, not a config value someone can bump in a PR.
  • Retrieval: filters, hybrid, rerank

    Top-k similarity search alone is a demo, not a product. Three additions turn it into something that holds up.

    Metadata filtering is the highest-value and least-used feature. Attach metadata at ingestion (tenant, document type, effective date, classification) and filter at query time. This is also how multi-tenant isolation gets enforced — at the retrieval layer, never by asking the model nicely.

    import boto3
    
    agent = boto3.client("bedrock-agent-runtime")
    
    response = agent.retrieve(
        knowledgeBaseId="KB1234ABCD",
        retrievalQuery={"text": "What is the escalation path for a Sev1 database incident?"},
        retrievalConfiguration={
            "vectorSearchConfiguration": {
                "numberOfResults": 20,
                "overrideSearchType": "HYBRID",
                "filter": {
                    "andAll": [
                        {"equals": {"key": "tenant_id", "value": "acme-corp"}},
                        {"equals": {"key": "doc_type", "value": "runbook"}},
                    ]
                },
            }
        },
    )
    
    for result in response["retrievalResults"]:
        print(result["score"], result["location"], result["content"]["text"][:120])
    

    Hybrid search combines semantic similarity with keyword matching. Pure vector search is weak exactly where enterprise corpora are strong: error codes, part numbers, acronyms, product SKUs. ARN and Sev1 and ORA-01555 are keyword problems, not semantic ones.

    Reranking is the cheapest quality upgrade available. Retrieve 20–25 candidates, run them through a reranker (Amazon Rerank or Cohere Rerank in Bedrock), and pass only the top 5 to the model. A cross-encoder reranker reads the query and passage together, which a bi-encoder embedding search fundamentally cannot do. Expect a real precision jump for a small amount of added latency, and a lower generation bill because you stopped stuffing the prompt with near-misses.

    Generation, with citations

    For the managed path, one call does retrieval and generation together and returns citations you can render as source links:

    response = agent.retrieve_and_generate(
        input={"text": "What is the escalation path for a Sev1 database incident?"},
        retrieveAndGenerateConfiguration={
            "type": "KNOWLEDGE_BASE",
            "knowledgeBaseConfiguration": {
                "knowledgeBaseId": "KB1234ABCD",
                "modelArn": "arn:aws:bedrock:us-east-1::foundation-model/anthropic.claude-sonnet-4-20250514-v1:0",
                "retrievalConfiguration": {
                    "vectorSearchConfiguration": {"numberOfResults": 8}
                },
                "generationConfiguration": {
                    "guardrailConfiguration": {
                        "guardrailId": "gr-abc123",
                        "guardrailVersion": "1",
                    }
                },
            },
        },
    )
    
    print(response["output"]["text"])
    for citation in response["citations"]:
        for ref in citation["retrievedReferences"]:
            print("source:", ref["location"])
    

    Always render the citations. Not because it looks thorough, but because a grounded answer with a link a user can open is falsifiable, and an ungrounded one isn't. Users calibrate their trust on the sources, and your support team gets a reproduction path when an answer is wrong.

    Guardrails and grounding

    Bedrock Guardrails handles the obvious content and PII cases, but the feature that matters for RAG is the contextual grounding check. It scores two things independently: whether the response is supported by the retrieved passages (grounding), and whether it is actually relevant to the question (relevance). Set thresholds, and on a low score return "I don't have that in the knowledge base" instead of the model's best guess.

    That refusal path is a feature, not a failure. A system that admits ignorance 5% of the time earns the trust that makes the other 95% usable.

    Evaluate before you ship, and keep evaluating

    "It looked good when I tried it" is not an evaluation. Build a golden set of 50–200 real questions with known-correct source documents — pull them from your actual support queue, not from imagination — and measure the two halves separately:

  • Retrieval quality: is the correct passage in the top k? Track recall@k and MRR. If it isn't retrieved, no model can save the answer.
  • Generation quality: faithfulness to the retrieved context, relevance, completeness. Bedrock Evaluations will score these with an LLM judge, or run your own harness in CI.
  • Wire the golden set into CI. Chunking changes, embedding model swaps, and prompt edits all silently regress retrieval, and without a scored baseline you will find out from a user.

    Infrastructure as code

    Click-ops a knowledge base once to learn the shape of it, then delete it and write it down. The Terraform provider covers Knowledge Bases and data sources:

    resource "aws_bedrockagent_knowledge_base" "docs" {
      name     = "neuralops-docs-kb"
      role_arn = aws_iam_role.bedrock_kb.arn
    
      knowledge_base_configuration {
        type = "VECTOR"
        vector_knowledge_base_configuration {
          embedding_model_arn = "arn:aws:bedrock:us-east-1::foundation-model/amazon.titan-embed-text-v2:0"
        }
      }
    
      storage_configuration {
        type = "RDS"
        rds_configuration {
          resource_arn           = aws_rds_cluster.vectors.arn
          credentials_secret_arn = aws_secretsmanager_secret.vectors.arn
          database_name          = "bedrock_vectors"
          table_name             = "bedrock_integration.bedrock_kb"
          field_mapping {
            primary_key_field = "id"
            vector_field      = "embedding"
            text_field        = "chunks"
            metadata_field    = "metadata"
          }
        }
      }
    }
    
    resource "aws_bedrockagent_data_source" "s3_docs" {
      knowledge_base_id = aws_bedrockagent_knowledge_base.docs.id
      name              = "s3-corpus"
    
      data_source_configuration {
        type = "S3"
        s3_configuration {
          bucket_arn = aws_s3_bucket.corpus.arn
        }
      }
    
      vector_ingestion_configuration {
        chunking_configuration {
          chunking_strategy = "HIERARCHICAL"
          hierarchical_chunking_configuration {
            overlap_tokens = 60
            level_configuration {
              max_tokens = 1500
            }
            level_configuration {
              max_tokens = 300
            }
          }
        }
      }
    }
    

    Then trigger ingestion from your pipeline rather than the console — an S3 event or a scheduled job, whichever matches how your corpus actually changes:

    aws bedrock-agent start-ingestion-job \
      --knowledge-base-id KB1234ABCD \
      --data-source-id DS5678EFGH
    

    Ingestion is incremental: unchanged documents are skipped, so re-running on a schedule is cheap and keeps the index honest.

    The cost model

    RAG costs show up in four places, and only one of them is the part everyone budgets for:

  • Embeddings at ingestion — per token, paid once per chunk per re-index. Genuinely small for most corpora, and re-indexing the whole corpus for an embedding model change is the line item that surprises people.
  • Vector store — usually the largest recurring cost, and it accrues whether or not anyone asks a question. OpenSearch Serverless bills on OCU minimums; Aurora Serverless v2 scales down but never to zero if you need warm queries.
  • Query-time embedding and reranking — per request, small per unit, meaningful at volume.
  • Generation — per input and output token. Input dominates in RAG, because you are shipping retrieved passages on every call.
  • Two levers move the bill more than model choice does: retrieve fewer, better chunks (this is what reranking buys you) and cache aggressively. A surprising share of enterprise questions are near-duplicates — a semantic cache on the query embedding, and prompt caching for stable system instructions, cut real spend without touching answer quality.

    What goes wrong in production

    Ranked by how often I've seen it:

  • Chunking that ignores document structure — tables shredded mid-row, procedures split across chunks, headings orphaned from their content.
  • No metadata, therefore no filtering — everything is retrieved from everything, and multi-tenant isolation becomes a prompt instruction, which is to say it doesn't exist.
  • Stale indexes — the source document was updated, the vector wasn't. Users lose trust in a system that confidently quotes a superseded policy far faster than in one that says it doesn't know.
  • No evaluation baseline — every change is a coin flip, and regressions are reported by customers.
  • Skipping hybrid search — then wondering why the system can't find a document by its error code.
  • Over-retrieving — 20 chunks stuffed into the prompt, the relevant one buried in the middle, and a bill to match.
  • Treating it as a project, not a product — a corpus is a living thing. Someone has to own ingestion freshness, evaluation, and the feedback loop from wrong answers back into chunking and content fixes.
  • Where to start

    If you're beginning next week: pick one corpus with a real audience, use hierarchical chunking, back it with the vector store your team already knows how to operate, turn on hybrid search and reranking from day one, write fifty golden questions before you write the UI, and put a grounding threshold in front of every answer.

    The technology is no longer the hard part. Bedrock removes the infrastructure work almost entirely. What's left is the discipline — document hygiene, retrieval evaluation, and the willingness to let the system say "I don't know."


    Building a RAG platform on AWS, or trying to get one that already exists to behave? Let's talk about your architecture.

    AWSAmazon BedrockRAGGenAIOpenSearchTerraformLLM

    Share this article

    Roger Vasconcelos

    Roger Vasconcelos

    AWS DevOps Architect

    Senior AWS DevOps Architect with 10 AWS certifications and 20+ years of experience, running a boutique practice augmented by a fleet of AI agents.

    Learn more about me

    Need Help with Your Cloud Infrastructure?

    Let's discuss how I can help you implement these best practices in your organization.

    Get in Touch