Every organization I talk to wants the same thing from generative AI: answers grounded in their data — the runbooks, the contracts, the support tickets, the ten years of Confluence nobody has read since 2019. And almost every one of them starts from the same wrong assumption: that this is a model problem.
It isn't. A foundation model that has never seen your incident postmortems will confidently invent them. Fine-tuning teaches style and format far better than it teaches facts, and it goes stale the moment someone edits a document. Retrieval-Augmented Generation (RAG) fixes the actual problem: it retrieves the relevant passages from your corpus at query time and puts them in front of the model as context, so the answer is generated from your documents rather than from the model's memory of the internet.
Amazon Bedrock makes the happy path genuinely easy — you can have a working knowledge base in an afternoon. The gap between that afternoon and something you would put in front of customers is where this article lives.
What Bedrock actually gives you
Bedrock is not one service, it is a set of components you can adopt independently:
Retrieve and RetrieveAndGenerate APIs.The important design decision is how much of the pipeline you hand to Knowledge Bases. Managed ingestion is the right default: it removes an entire class of undifferentiated plumbing. But it is opinionated, and once your chunking or metadata needs get specific, you will want the Retrieve API with your own generation call rather than the fully managed RetrieveAndGenerate. Build so that swap is a one-line change.
The reference architecture
Two paths, and they fail for different reasons — keep them separate in your head and in your code.
Ingestion path (asynchronous, batch): source documents land in S3 → an ingestion job chunks each document → each chunk is embedded → vectors plus metadata are written to the vector store. This runs on document change, not on user request.
Query path (synchronous, latency-critical): user question → embed the question → vector search with metadata filters → optional rerank → assemble prompt with the top passages → generate → apply guardrails → return the answer with citations.
Almost every "the AI is hallucinating" complaint in production is really a retrieval failure on the query path, or a chunking failure on the ingestion path. The model is usually the least broken part of the system.
Choosing the vector store
Bedrock Knowledge Bases can back onto several stores, and this choice is harder to reverse than any other, so make it deliberately.
maintenance_work_mem during index builds) is on you.Start with Aurora pgvector for cost-sensitive internal workloads and OpenSearch Serverless when hybrid search quality and scale are the priority.
Chunking is the highest-leverage decision you will make
If you change one thing after reading this, change your chunking. Retrieval can only return what chunking created. A chunk that splits a procedure in half will never answer a question about that procedure, no matter how good the embedding model is.
Bedrock offers several strategies:
A concrete rule that has paid for itself on every engagement: give every chunk its provenance in the text, not just in the metadata. Prefixing a chunk with Document: Incident Response Runbook > Section: Escalation > Page 12 measurably improves both retrieval and the model's ability to cite correctly.
Embeddings: pick dimensions on purpose
Titan Text Embeddings V2 supports 1024, 512, and 256 dimensions. Cohere Embed offers English and multilingual variants. The instinct is to take the largest vector available; resist it. Dimension count drives vector store size, index memory, and query latency, and on many corpora 512 dimensions retrieves within noise of 1024 while costing meaningfully less to store and search.
Two rules that are not negotiable:
Retrieval: filters, hybrid, rerank
Top-k similarity search alone is a demo, not a product. Three additions turn it into something that holds up.
Metadata filtering is the highest-value and least-used feature. Attach metadata at ingestion (tenant, document type, effective date, classification) and filter at query time. This is also how multi-tenant isolation gets enforced — at the retrieval layer, never by asking the model nicely.
import boto3
agent = boto3.client("bedrock-agent-runtime")
response = agent.retrieve(
knowledgeBaseId="KB1234ABCD",
retrievalQuery={"text": "What is the escalation path for a Sev1 database incident?"},
retrievalConfiguration={
"vectorSearchConfiguration": {
"numberOfResults": 20,
"overrideSearchType": "HYBRID",
"filter": {
"andAll": [
{"equals": {"key": "tenant_id", "value": "acme-corp"}},
{"equals": {"key": "doc_type", "value": "runbook"}},
]
},
}
},
)
for result in response["retrievalResults"]:
print(result["score"], result["location"], result["content"]["text"][:120])
Hybrid search combines semantic similarity with keyword matching. Pure vector search is weak exactly where enterprise corpora are strong: error codes, part numbers, acronyms, product SKUs. ARN and Sev1 and ORA-01555 are keyword problems, not semantic ones.
Reranking is the cheapest quality upgrade available. Retrieve 20–25 candidates, run them through a reranker (Amazon Rerank or Cohere Rerank in Bedrock), and pass only the top 5 to the model. A cross-encoder reranker reads the query and passage together, which a bi-encoder embedding search fundamentally cannot do. Expect a real precision jump for a small amount of added latency, and a lower generation bill because you stopped stuffing the prompt with near-misses.
Generation, with citations
For the managed path, one call does retrieval and generation together and returns citations you can render as source links:
response = agent.retrieve_and_generate(
input={"text": "What is the escalation path for a Sev1 database incident?"},
retrieveAndGenerateConfiguration={
"type": "KNOWLEDGE_BASE",
"knowledgeBaseConfiguration": {
"knowledgeBaseId": "KB1234ABCD",
"modelArn": "arn:aws:bedrock:us-east-1::foundation-model/anthropic.claude-sonnet-4-20250514-v1:0",
"retrievalConfiguration": {
"vectorSearchConfiguration": {"numberOfResults": 8}
},
"generationConfiguration": {
"guardrailConfiguration": {
"guardrailId": "gr-abc123",
"guardrailVersion": "1",
}
},
},
},
)
print(response["output"]["text"])
for citation in response["citations"]:
for ref in citation["retrievedReferences"]:
print("source:", ref["location"])
Always render the citations. Not because it looks thorough, but because a grounded answer with a link a user can open is falsifiable, and an ungrounded one isn't. Users calibrate their trust on the sources, and your support team gets a reproduction path when an answer is wrong.
Guardrails and grounding
Bedrock Guardrails handles the obvious content and PII cases, but the feature that matters for RAG is the contextual grounding check. It scores two things independently: whether the response is supported by the retrieved passages (grounding), and whether it is actually relevant to the question (relevance). Set thresholds, and on a low score return "I don't have that in the knowledge base" instead of the model's best guess.
That refusal path is a feature, not a failure. A system that admits ignorance 5% of the time earns the trust that makes the other 95% usable.
Evaluate before you ship, and keep evaluating
"It looked good when I tried it" is not an evaluation. Build a golden set of 50–200 real questions with known-correct source documents — pull them from your actual support queue, not from imagination — and measure the two halves separately:
Wire the golden set into CI. Chunking changes, embedding model swaps, and prompt edits all silently regress retrieval, and without a scored baseline you will find out from a user.
Infrastructure as code
Click-ops a knowledge base once to learn the shape of it, then delete it and write it down. The Terraform provider covers Knowledge Bases and data sources:
resource "aws_bedrockagent_knowledge_base" "docs" {
name = "neuralops-docs-kb"
role_arn = aws_iam_role.bedrock_kb.arn
knowledge_base_configuration {
type = "VECTOR"
vector_knowledge_base_configuration {
embedding_model_arn = "arn:aws:bedrock:us-east-1::foundation-model/amazon.titan-embed-text-v2:0"
}
}
storage_configuration {
type = "RDS"
rds_configuration {
resource_arn = aws_rds_cluster.vectors.arn
credentials_secret_arn = aws_secretsmanager_secret.vectors.arn
database_name = "bedrock_vectors"
table_name = "bedrock_integration.bedrock_kb"
field_mapping {
primary_key_field = "id"
vector_field = "embedding"
text_field = "chunks"
metadata_field = "metadata"
}
}
}
}
resource "aws_bedrockagent_data_source" "s3_docs" {
knowledge_base_id = aws_bedrockagent_knowledge_base.docs.id
name = "s3-corpus"
data_source_configuration {
type = "S3"
s3_configuration {
bucket_arn = aws_s3_bucket.corpus.arn
}
}
vector_ingestion_configuration {
chunking_configuration {
chunking_strategy = "HIERARCHICAL"
hierarchical_chunking_configuration {
overlap_tokens = 60
level_configuration {
max_tokens = 1500
}
level_configuration {
max_tokens = 300
}
}
}
}
}
Then trigger ingestion from your pipeline rather than the console — an S3 event or a scheduled job, whichever matches how your corpus actually changes:
aws bedrock-agent start-ingestion-job \
--knowledge-base-id KB1234ABCD \
--data-source-id DS5678EFGH
Ingestion is incremental: unchanged documents are skipped, so re-running on a schedule is cheap and keeps the index honest.
The cost model
RAG costs show up in four places, and only one of them is the part everyone budgets for:
Two levers move the bill more than model choice does: retrieve fewer, better chunks (this is what reranking buys you) and cache aggressively. A surprising share of enterprise questions are near-duplicates — a semantic cache on the query embedding, and prompt caching for stable system instructions, cut real spend without touching answer quality.
What goes wrong in production
Ranked by how often I've seen it:
Where to start
If you're beginning next week: pick one corpus with a real audience, use hierarchical chunking, back it with the vector store your team already knows how to operate, turn on hybrid search and reranking from day one, write fifty golden questions before you write the UI, and put a grounding threshold in front of every answer.
The technology is no longer the hard part. Bedrock removes the infrastructure work almost entirely. What's left is the discipline — document hygiene, retrieval evaluation, and the willingness to let the system say "I don't know."
Building a RAG platform on AWS, or trying to get one that already exists to behave? Let's talk about your architecture.

