[8] Vaswani and colleagues. Attention Is All You Need, 2017. Foundational transformer research. Figure 1 is a simplified modern decoder-only teaching diagram, not a reproduction of the paper's encoder-decoder architecture.
[9–10] Hugging Face Transformers. Official guidance on fine-tuning and cache strategies. Consult model-specific documentation because architecture, cache support and hardware behaviour vary.
[11] NVIDIA serving guidance. Explains prefill and decode resource demands. Used for the performance discussion, with workload-dependent qualifications. It is vendor documentation, not a vendor-neutral hardware benchmark.
[12] pgvector. Official project documentation for vector search in PostgreSQL, including indexing and filtering behaviour.
[13–14] Orchestration and retrieval examples. Official documentation illustrating framework roles. These products are representative choices, not prerequisites for the labs or a recommendation to adopt a particular stack.
[13] LangGraph and LangChain context
[15–18] Local model tooling. Official documentation and repositories used to check the roles described in chapter 24. Confirm model support, licensing and data behaviour for the exact configuration you choose.
Product names throughout the guide identify examples of categories. Features, service terms, model availability and prices can change independently of the underlying concepts. The durable skills are choosing the right boundary, testing the outcome and understanding which part of the system failed.