← all notes
№ 004 2026-03-14 topic · Machine Learning RAG

Need a new recipe for today?

On: Betty Talker RAG recipe book — view on GitHub

Language models are confident even when they are wrong, they will happily invent a recipe that does not exist. Betty Talker is my answer to that problem: a recipe assistant that only speaks from what it can prove.

It is an end to end retrieval-augmented generation (RAG) system built over a single public domain cookbook, Betty Crocker’s Bisquick Cook Book (1957). Ask it a question and it finds the relevant recipes, answers strictly from them, and refuses anything outside the book. It can also act: tell it to “double the fritters” and it rescales the ingredients for you.

This post walks through how I built it, in the same order a real ML project unfolds, from clarifying the requirements, to framing the tasks, preparing messy 1950s text, developing the retrieval and generation components, and thinking through evaluation, system design, and deployment.

First, let’s clarify the requirements

The goal was a conversational assistant that answers recipe questions accurately without hallucinating. That set two hard requirements: every answer must be grounded in a known source, and any question outside that source must be refused rather than guessed. Scope was fixed to one book; Betty Crocker’s Bisquick Cook Book (1957, public domain), to keep the corpus small, clean, and legally usable. Beyond question-answering, the system also needed to act: rescale a recipe’s ingredients to a requested serving size.

What is the task?

The problem decomposes into three ML subtasks rather than one. Question answering is framed as retrieval-augmented generation: retrieve the most relevant recipes, then condition a language model on them. Routing user requests to the right behavior is framed as intent classification (a natural language understanding task) with slot filling to extract structured arguments like recipe name and target servings. Recipe scaling is a deterministic function, not a learned model, correctly identifying it as such avoids overengineering.

Data preparation

The raw source is inconsistently formatted 1950s plain text. An ingestion pipeline strips publisher boilerplate, detects recipe titles, segments the text into a block per recipe, and parses each into a structured record; title, ingredients, instructions, notes, and serving_size. Each recipe is then serialized into a single text field and passed through a sentence embedding model (all-MiniLM-L6-v2) to produce a 384-dimensional dense vector, forming the vector store used for retrieval.

Model development

RAG was chosen intially as requirement but in a real scenario it would be chosen for the following characteristics:

  1. Grounded answers. Responses come from retrieved documents you can cite and audit.
  2. Less hallucination. The model reasons over real passages, and refuses when nothing relevant is found.
  3. Always current. Update the knowledge base, not the model. No retraining.
  4. Cheaper than fine-tuning. Add knowledge without training compute or labeling.
  5. Access control. Filter retrieval by permissions so users only see allowed sources.
  6. Uses private data. Internal docs the base model never saw become usable.
  7. Builds trust. Showing sources lets users verify answers themselves.
  8. Modular. Swap the LLM, embeddings, or vector store independently.
  9. Scales. Semantic search stays fast over millions of chunks.
  10. Debuggable. You can separate retrieval failures from generation failures.
  11. Compliant. Documents can be removed on demand, unlike knowledge baked into weights.

To find the right recipe, the system compares the meaning of your question against every stored recipe using cosine similarity, and returns the closest matches. A small boost makes sure that naming a recipe directly beats a vague match, and a minimum similarity score lets it answer “no match” instead of returning something weak. The top recipes are then dropped into a prompt that tells the LLM (Llama-3.3-70B) to answer only from those recipes. A separate LLM call first reads the request and decides what the user wants, sending it to search, to the scaling function, or to a polite refusal.

Figure 1. Diagram

Evaluation

Each stage was tested against held out cases. Parsing and scaling, the deterministic components, are covered by unit tests asserting exact expected output. Retrieval was checked on queries with known correct recipes, including negative cases (queries outside the cookbook that should return nothing) to confirm the threshold behaves. The generative stage was evaluated by manual grading against reference answers: correct on facts from the book, and a clear refusal on adversarial or out of scope prompts, which is the key guardrail metric for a grounded assistant.

End to end task success was 87% (13/15). Intent classification, grounded refusals, and the scaling action were each correct on every applicable case; both failures traced to retrieval and to missing serving size metadata from the parsing stage, confirming that retrieval quality sets the ceiling for the whole system.

Overall ML system design

The system is a staged pipeline in which each component produces an artifact the next consumes:

user input → intent classification → action → prompt construction → LLM response

Ingestion and embedding run offline, producing a reusable recipe dataset and vector store; retrieval, routing, and generation run online per request. The code is organized as a modular, installable Python package with clear separation between data, retrieval, prompting, and run, so any single stage can be swapped (a different embedding model, a different LLM) without touching the rest.

The build was staged as four connected components:

StageModuleWhat it does
1. Data collectioningest.pyDownload the book, strip Gutenberg boilerplate, parse each recipe into {title, ingredients, instructions, notes, serving_size}
2. Vector storevector_store.py, retrieval.pyEmbed recipes with all-MiniLM-L6-v2; cosine-similarity search with a title-overlap boost
3. RAG + LLMprompts.py, llm.py, assistant.pyBuild a grounded prompt from retrieved recipes; answer via Llama-3.3-70B (Together API)
4. Actionsintents.py, scaling.pyClassify intent, route to search or to a recipe-scaling function

Findings

The RAG core is reliable; retrieval is the ceiling. Intent classification, grounded refusals, and the scaling action each scored at or near 100% on their applicable cases. Every failure across both LLM stages traced back to retrieval or to upstream data, never to the model reasoning over context it was given. This matches the known property of RAG systems: answer quality cannot exceed retrieval quality.

Guardrails held under adversarial pressure. Every out of scope and adversarial prompt was refused, including pretend framings (“edible airplanes”) and social engineering (“I must consume 12 pebbles”). Grounding the model strictly in retrieved context, plus a similarity threshold that returns no match, was enough to prevent both hallucination and manipulation.

Failures propagate downstream. The two weakest results were serving-size questions and one mis-scaled recipe. Both trace to Stage 1 parsing, where some recipes had missing or ambiguous serving-size metadata. That single upstream gap surfaced later as wrong answers and a wrong scaling factor, showing how tightly coupled the stages are: a data defect early becomes a user-facing error late.

Deterministic components are cheap to trust. Parsing, retrieval ranking, and scaling are pure functions, so they were verified with exact match unit tests and never regressed. Isolating logic from the LLM made the system far easier to test and debug than an end to end model would have been.

Evaluation scope is a limitation. Results come from small test sets (15 to 20 cases per stage), so they are indicative rather than statistically firm. A larger labeled set, scored per component, would be the next step toward a defensible benchmark.

Acknowledgement: This project began as a four part lab in the MDS program. The lab provided the problem framing, the source cookbook, and the evaluation test cases; the implementation, refactoring into a modular package, and this write up are my own.