Applications of Foundation Models

Domain 3: Applications of Foundation Models

65 practice questions for Domain 3 of the AWS Certified AI Practitioner (AIF-C01) exam, which makes up 28% of its scored content. Your answers count towards one score and one timer for the whole exam.

Domain 3: Applications of Foundation Models

28% of scored content · 65 practice questions

110. A support team wants a foundation model to adopt a specific tone and answer format without changing model weights. Which technique should they use first?

Answer and explanation

Answer: C. Tone and output format are usually controllable through instructions and a handful of in-prompt examples, which changes no weights, costs nothing to train, and can be iterated in minutes. Pre-training from scratch requires enormous data and expense and is never the first response to a formatting need. Continued pre-training injects broad domain knowledge and also modifies weights. Quantization reduces memory and latency by lowering numeric precision; it does not shape style.

111. After prompt engineering proves insufficient, a firm needs a model to reliably use specialized legal terminology from its own historical filings. It has several thousand curated prompt-and-response pairs. Which approach is most appropriate?

Answer and explanation

Answer: D. Several thousand curated prompt-and-response pairs is the classic input to supervised fine-tuning, which adapts the model's weights so the specialized behaviour becomes intrinsic rather than re-supplied on every call. Packing hundreds of examples into each prompt would exhaust the context window and inflate per-request cost. Top-p controls sampling diversity, not domain vocabulary. A larger instance changes throughput and latency, not what the model knows.

112. Which evaluation approach best measures whether a summarization model's output is acceptable to end users?

Answer and explanation

Answer: C. Summarization quality is partly subjective, so automated similarity metrics catch regressions cheaply at scale while human review of sampled outputs judges faithfulness, usefulness, and tone that metrics miss; using both is standard practice. Training loss measures fit to training data and correlates poorly with perceived quality. Parameter count is a property of the model, not evidence about its output. Latency is an operational measure that says nothing about whether a summary is good.

113. Which sequence describes a Retrieval Augmented Generation request?

Answer and explanation

Answer: B. RAG converts the query into an embedding, retrieves semantically similar passages, places them in the prompt as context, and generates an answer grounded in that context. Fine-tuning on a single query is not a runtime operation. Generating first and finding supporting sources afterwards is post-hoc justification, not grounding. Translation is a separate task unrelated to retrieval.

114. A company must give a foundation model access to its internal wiki so answers reflect current company policy, and the wiki changes weekly. Which approach is most appropriate?

Answer and explanation

Answer: D. RAG retrieves current content at query time, so refreshing the index keeps answers current without any retraining. Weekly fine-tuning is expensive, slow, and still leaves the model stale between runs. Temperature affects randomness, not knowledge. A larger model still has no knowledge of private content it never saw.

115. Which prompting technique asks a model to work through intermediate reasoning steps before giving a final answer?

Answer and explanation

Answer: D. Chain-of-thought prompting instructs the model to lay out intermediate steps, which measurably improves accuracy on multi-step reasoning tasks. Zero-shot prompting supplies an instruction with neither examples nor a reasoning structure. Negative prompting specifies what to avoid, commonly in image generation. Prompt caching is a cost and latency optimisation.

116. An agent built on a foundation model must look up a live order status before answering. What capability makes this possible?

Answer and explanation

Answer: D. Tool calling lets the model emit a structured request that the application executes against a real API, returning the result for the model to use. A larger context window holds more text but retrieves nothing. More output tokens allow a longer answer. A temperature of zero makes output deterministic without granting access to live data.

117. Which evaluation metric is commonly used to compare a generated summary against reference summaries?

Answer and explanation

Answer: D. ROUGE measures overlap between generated and reference text and is a standard automatic metric for summarization. Root mean squared error applies to continuous numeric predictions. Precision at k evaluates a retrieval ranking rather than generated text. Training cross-entropy measures fit to training data rather than output quality.

118. A team finds that a foundation model's answers are accurate but far too long for a chat interface. What is the cheapest effective fix?

Answer and explanation

Answer: B. Response length responds directly to instruction and to the output token ceiling, both of which are configuration changes costing nothing. Fine-tuning is an expensive way to enforce a formatting preference. A smaller model changes quality as well as length and may not be shorter. Index size affects what is retrieved, not how verbose the answer is.

119. When choosing between two foundation models for a production feature, which combination of factors should be compared? (Select TWO.)

Answer and explanation

Answer: C, E. A defensible choice rests on measured quality against a representative dataset together with the cost and latency the workload will actually incur. Provider headquarters and model naming are irrelevant. Parameter count alone correlates loosely with task performance and ignores cost and speed entirely.

120. Documents are split into chunks before being embedded for retrieval. Why does chunk size matter?

Answer and explanation

Answer: C. An embedding summarises the whole chunk, so an oversized chunk covering several topics matches broadly rather than precisely, while an undersized chunk can separate a fact from the context that makes it meaningful. Temperature, user limits, and encryption are unrelated to how content is segmented.

121. Which Amazon Bedrock capability lets a developer connect a foundation model to a managed retrieval index over their own documents?

Answer and explanation

Answer: C. Knowledge Bases handles ingestion, chunking, embedding, and retrieval over customer documents and connects the result to a model for RAG. Guardrails applies content filters to prompts and responses. Ground Truth labels training data. DataBrew performs visual data preparation for analytics.

122. A customer service assistant must answer only from approved company documents and refuse anything else. Which combination best enforces this? (Select TWO.)

Answer and explanation

Answer: C, E. Retrieval restricts what the model can see to the approved corpus, and guardrails with denied topics enforce refusal at the service layer rather than relying on the model's cooperation alone. Higher temperature increases variability and makes off-topic drift more likely. Removing the system prompt discards the instruction that shapes behaviour. Fine-tuning on unrelated data broadens rather than narrows the model's range.

123. What is the main advantage of prompt engineering over fine-tuning for a new use case?

Answer and explanation

Answer: D. Prompt engineering is the cheapest and fastest lever because it requires no training run and can be revised immediately, which is why it is the first thing to try. Fine-tuning often achieves higher accuracy on a narrow task. Prompting supplies context per request rather than teaching the model anything permanently. Evaluation remains necessary whichever technique is used.

124. A model must classify support tickets into twelve categories, and the team has 40 labelled examples per category. Which approach should be tried first?

Answer and explanation

Answer: B. With a small labelled set, few-shot prompting establishes a baseline in minutes and often suffices, and it tells you whether fine-tuning is worth the cost. Fine-tuning on 480 examples may help but should be justified against a measured baseline. Pre-training from scratch is never appropriate at this scale. Maximising the context window increases cost without addressing classification quality.

125. An application must return a structured JSON object from a foundation model for downstream processing. What is the most reliable practice?

Answer and explanation

Answer: A. A clearly specified schema substantially improves conformance, and programmatic validation with a retry path guarantees malformed output never reaches downstream systems. Assuming conformance fails eventually and silently. Higher temperature increases format variance. Regular expressions are brittle against exactly the structural variation they are meant to absorb.

126. Which use case is a poor fit for a generative foundation model?

Answer and explanation

Answer: C. Deterministic arithmetic with a legal obligation to be exact belongs in code or a calculation engine, not a probabilistic text generator. Drafting, summarizing, and grounded question answering are all mainstream generative use cases where a human reviews or the answer is grounded in retrieved sources.

127. Which capability allows a foundation model application to break a request into steps and call external systems to complete it?

Answer and explanation

Answer: A. An agent plans a sequence of steps and invokes defined actions or tools to accomplish a goal. Context window and token limits govern how much text is processed. Temperature affects sampling randomness.

128. A company must decide between retrieval augmented generation and fine-tuning. Which requirement points to retrieval?

Answer and explanation

Answer: A. Changing content and source citation both require supplying documents at query time, which is what retrieval does. Style and fixed formatting are behaviours better addressed by prompting or fine-tuning. Retrieval typically increases rather than reduces per-request cost because it adds context tokens.

129. Which prompting approach provides instructions with no examples and relies on the model's general capability?

Answer and explanation

Answer: A. Zero-shot prompting gives the instruction alone. Few-shot supplies demonstrations. Chain-of-thought asks for intermediate reasoning steps. Retrieval augmentation supplies external context rather than examples.

130. Which evaluation approach is most appropriate for judging whether generated customer emails are appropriate in tone?

Answer and explanation

Answer: A. Tone is a subjective quality best judged by people against an explicit rubric, with automated metrics as a supplementary signal. Similarity to references penalises acceptable variation. Confidence scores describe the model's certainty rather than appropriateness. Token count measures length.

131. An application must prevent a customer service assistant from discussing topics outside its remit. Which mechanism enforces this at the service layer?

Answer and explanation

Answer: C. Denied topics in a guardrail are evaluated by the service on input and output and block the interaction independently of the model's cooperation. A system prompt is advisory. Temperature and context window do not restrict subject matter.

132. Which factor most affects the quality of answers in a retrieval augmented generation system?

Answer and explanation

Answer: B. If retrieval does not surface the relevant passage, no model can answer correctly from it, which makes retrieval quality the dominant factor. Parameter count and temperature matter less than whether the right context was supplied. Implementation language is irrelevant.

133. A team wants to process 500,000 documents overnight through a foundation model at the lowest cost. Which approach fits?

Answer and explanation

Answer: C. Batch inference is designed for large asynchronous workloads and is priced below real-time invocation. A loop of real-time calls pays real-time pricing and risks throttling. Streaming delivers responses incrementally without changing the pricing model. Provisioned throughput reserves capacity continuously for a workload that runs once.

134. Which describes a prompt template and why it is useful?

Answer and explanation

Answer: A. A template fixes the instruction structure while allowing the variable parts to change, which makes behaviour consistent and easier to govern and version. A cache stores results. Fine-tuning modifies weights. Token limits constrain size.

135. An image generation prompt produces images containing an unwanted element. Which technique addresses this directly?

Answer and explanation

Answer: B. Negative prompts specify content to exclude and are supported by many image generation models. Temperature affects variability. Resolution affects image size. Retrieval applies to text grounding rather than image generation.

136. Which consideration should determine whether a generative feature needs human review before its output reaches a customer?

Answer and explanation

Answer: D. The need for review follows from how much harm an error would cause, which is a risk judgement about the context. Cost, implementation language, and user count do not determine the severity of an error.

137. An application must answer questions from a company's internal policy documents using retrieval augmented generation. (Place the steps in the correct order.)

Answer and explanation

Answer: A → D → B → C. Chunking and embedding must happen before anything can be stored, and the overlap is what stops a fact being split across a boundary. The index is populated next, with source metadata retained so answers can cite their origin. At query time the question is embedded and matched against the index. Generation comes last and is grounded in whatever was retrieved, which is why a retrieval failure cannot be repaired by the model.

138. A retail assistant must report the current status of a customer's order. Which approach should it use?

Answer and explanation

Answer: A. Live data must be fetched at the moment of the request, which is what tool use provides. Fine-tuning on historical records teaches patterns rather than current facts and would answer from stale memory. A daily index refresh leaves order status up to a day old, which is unacceptable for a status question. A system prompt is fixed and cannot carry every customer's orders.

139. An assistant must answer from a returns policy document and cite the passage it used. Which approach should it use?

Answer and explanation

Answer: B. Citation requires the answer to be grounded in a retrievable passage whose source can be named, which retrieval provides and fine-tuning does not. Fine-tuning embeds the policy in weights with no source to cite and must be repeated at each revision. Placing the whole document in every prompt wastes tokens on requests that do not concern returns. General training knowledge would not contain this company's policy at all.

140. An assistant must never discuss competitor products, and the restriction must be enforced rather than requested. Which control should be applied?

Answer and explanation

Answer: A. A denied topic is evaluated by the service on both the request and the response and blocks the interaction regardless of whether the model cooperates. A prompt instruction is advisory and is the first thing a prompt injection targets. Removing names from the index does not prevent the model discussing competitors from its general training knowledge. Weekly review detects breaches after customers have already seen them.

141. Which service allows prompts to be versioned and managed centrally rather than embedded in application code?

Answer and explanation

Answer: A. Prompt Management stores prompts as versioned templates an application references at invocation. Guardrails filter content. Knowledge Bases provide retrieval. Model Evaluation compares model output quality.

142. Which metric best indicates whether an AI assistant is meeting its business objective?

Answer and explanation

Answer: A. Task completion rate measures whether users achieve what they came to do, which is the business objective. Token counts and latency are cost and performance measures. A benchmark score describes the model rather than the application's outcome.

143. Which metric expresses the cost efficiency of a generative AI application from a business perspective?

Answer and explanation

Answer: C. Cost per interaction ties spend to a unit of business value. Total tokens and invocation counts measure volume without relating it to value. Parameter count is a model specification.

144. Which customization approach transfers the behaviour of a large model into a smaller one to reduce inference cost?

Answer and explanation

Answer: A. Distillation trains a smaller student model on a larger teacher's outputs so it approximates the teacher at lower cost. Retrieval supplies external context. In-context learning uses examples in the prompt. Continued pre-training adds domain knowledge to the existing model.

145. Which factor should influence the choice of foundation model for an application?

Answer and explanation

Answer: A. Model selection weighs capability, latency, and cost against the requirement. Parameter count and recency are not fitness criteria.

146. What does Retrieval Augmented Generation add to a foundation model application?

Answer and explanation

Answer: C. RAG supplies retrieved context at query time. It does not change the model's parameters, its inference speed, or translate output.

147. Why would an application use a vector database alongside a foundation model?

Answer and explanation

Answer: A. A vector database stores embeddings for similarity retrieval. Weights live with the model. Response caching and error logging use different stores.

148. What is an AI agent in the context of a foundation model application?

Answer and explanation

Answer: C. An agent plans and acts through tools. A chat interface, an edge deployment, and a dashboard are different components.

149. What does zero-shot prompting mean?

Answer and explanation

Answer: D. Zero-shot provides no examples. Few-shot provides examples. Fine-tuning trains on examples. Reasoning explanation is chain-of-thought prompting.

150. What is the purpose of few-shot prompting?

Answer and explanation

Answer: A. Few-shot supplies examples that demonstrate the pattern. It increases rather than reduces tokens, does not train the model, and does not limit response length.

151. What is a prompt injection attack?

Answer and explanation

Answer: D. Prompt injection manipulates the model through crafted input. Overloading is denial of service. Weight extraction and training data corruption are different attacks.

152. Which practice reduces the risk of a prompt injection succeeding?

Answer and explanation

Answer: C. Separation of untrusted input and limited permissions together reduce both likelihood and impact. A longer prompt is still part of the context an injection targets. Temperature adds randomness. Removing the system prompt loses the safe behaviour it defines.

153. What does fine-tuning a foundation model involve?

Answer and explanation

Answer: D. Fine-tuning continues training on curated data. Retrieval, prompting, and deployment do not change the model's weights.

154. When is continued pre-training preferred over fine-tuning?

Answer and explanation

Answer: D. Continued pre-training injects broad domain knowledge from unlabelled text. Format adherence suits fine-tuning. Changing documents and citation both suit retrieval.

155. Which data is required for supervised fine-tuning of a foundation model?

Answer and explanation

Answer: C. Supervised fine-tuning requires prompt and completion pairs. Unlabelled text suits continued pre-training. A vector index supports retrieval. Metrics evaluate rather than train.

156. Why is human evaluation used alongside automated metrics for generative AI?

Answer and explanation

Answer: B. Human judgement captures subjective qualities automated metrics approximate poorly. Automated metrics scale better and are faster. Labelled reference data often does exist.

157. Which service provides managed evaluation of foundation models on a dataset?

Answer and explanation

Answer: D. Bedrock model evaluation scores models against a dataset. CloudWatch monitors operations. Config records configuration. Comprehend analyses text.

158. Which evaluation approach is appropriate when no reference answer exists for a generated response?

Answer and explanation

Answer: C. Rubric-based scoring assesses open-ended output without a reference. Exact match against a prompt is meaningless. Token count measures length. Training data is not available for comparison.

159. Which consideration applies when deciding how much context to supply to a foundation model?

Answer and explanation

Answer: A. Context costs tokens and excess content can dilute relevance. More is not always better, context drives cost, and filling the window wastes it.

160. Which design choice reduces the latency a user perceives from a generative AI application?

Answer and explanation

Answer: C. Streaming shortens time to first output. More tokens and more context both lengthen generation. Temperature affects variation.

161. Which consideration applies when selecting a model for a multimodal application?

Answer and explanation

Answer: A. Multimodal use requires a model accepting the needed input types. Parameter count, recency, and context window do not determine modality support.

162. What does chain-of-thought prompting ask the model to do?

Answer and explanation

Answer: A. Chain-of-thought elicits intermediate reasoning. Alternatives, brevity, and retrieval are different techniques.

163. What is the purpose of a negative prompt?

Answer and explanation

Answer: A. A negative prompt specifies what to avoid. Temperature, token count, and invocation are unrelated.

164. Which practice improves the consistency of a prompt's results across requests?

Answer and explanation

Answer: D. A fixed versioned template produces consistent behaviour. Free phrasing, higher temperature, and varying examples all increase variation.

165. What does a benchmark dataset provide when evaluating a foundation model?

Answer and explanation

Answer: C. A benchmark enables like-for-like comparison. It does not guarantee production behaviour, supply application prompts, or measure cost.

166. Why should a foundation model be evaluated on the organization's own data as well as on public benchmarks?

Answer and explanation

Answer: D. Benchmarks may not represent the actual task. They are not necessarily outdated, do measure quality, and are generally reproducible.

167. Which metric family applies to evaluating a model's summarization output against reference summaries?

Answer and explanation

Answer: D. ROUGE measures overlap with reference summaries. Classification and regression metrics apply to different output types, and latency is operational.

168. What does a judge model contribute to evaluating generative output?

Answer and explanation

Answer: D. A judge model scores at scale and must be calibrated against human ratings. It does not guarantee agreement, remove the need for human input, or measure cost.

169. Which consideration applies when interpreting an evaluation score?

Answer and explanation

Answer: D. An evaluation score is only as predictive as its dataset is representative. Scores are not comparable across datasets and a single number cannot capture every dimension.

170. Which approach establishes whether a fine-tuned model is better than the base model?

Answer and explanation

Answer: A. A fair comparison uses identical data, metrics, and settings. Published benchmarks measure different tasks, training loss is not comparable, and parameter count is unchanged by fine-tuning.

171. What does a business-objective alignment metric measure for a generative AI application?

Answer and explanation

Answer: A. Alignment metrics tie the application to its intended business outcome. Grammar, latency, and model recency are quality and operational attributes.

172. Which evaluation approach detects that a model's behaviour has changed after an update?

Answer and explanation

Answer: C. Fixed-input comparison against recorded results detects behavioural change. Release notes describe intent, and latency and request counts are operational.

173. Which consideration applies when choosing between a larger and a smaller foundation model for an application?

Answer and explanation

Answer: C. Model size trades capability against cost and latency, and the requirement decides. Larger is not always better, smaller is not always worse, and size does affect both cost and latency.

174. Which technique supplies a model with the format of the expected answer without describing it in words?

Answer and explanation

Answer: D. Worked examples demonstrate the format directly. Temperature affects variation, output tokens affect length, and tone instructions address style.