55 practice questions for Domain 2 of the AWS Certified AI Practitioner (AIF-C01) exam, which makes up 24% of its scored content. Your answers count towards one score and one timer for the whole exam.
Domain 2: Fundamentals of Generative AI
55. An engineer must reduce hallucinated answers from a foundation model that answers questions about internal HR policy documents. Which approach addresses this most directly?
Answer and explanation
Answer: A. Hallucination on internal content occurs because the model was never trained on those documents; RAG retrieves the relevant passages at inference time and grounds the response in them, so the answer cites material the model can actually see. Raising temperature increases randomness and makes hallucination more likely. Raising the output token limit permits longer answers without improving factual grounding. A smaller context window reduces the room available for retrieved context, working against the goal.
56. Which statement correctly describes tokens in the context of a large language model?
Answer and explanation
Answer: A. Models decompose text into tokens — commonly sub-word fragments rather than whole words — and both context limits and consumption pricing on services such as Amazon Bedrock are expressed in input and output tokens. Weights are learned parameters, a distinct concept. Tokens in this sense are unrelated to authentication or encryption keys despite the shared word. GPU allocation is an infrastructure concern that does not define a token.
57. A company wants to use a foundation model through a fully managed API, choosing among models from several providers without provisioning inference infrastructure. Which AWS service fits?
Answer and explanation
Answer: C. Bedrock is the managed service that exposes foundation models from multiple providers behind a single serverless API, with no infrastructure to provision. SageMaker AI training jobs build or fine-tune models and still require configuring compute. GPU-backed EC2 instances mean managing servers, drivers, and scaling directly. Glue is a serverless ETL service for data integration and has no foundation model inference capability.
58. What does the context window of a large language model determine?
Answer and explanation
Answer: D. The context window bounds how many tokens the model can attend to in a single request, covering both the prompt and the generated response. Concurrency is a serving capacity concern. Models are stateless between requests and retain nothing unless it is resupplied in the prompt. Server memory is infrastructure and distinct from the model's token limit.
59. Raising the temperature parameter on a foundation model has what effect?
Answer and explanation
Answer: B. Temperature scales the sampling distribution, so a higher value makes less probable tokens more likely to be chosen and the output more varied. A lower value moves toward deterministic output. Temperature does not change how many tokens are processed. Increasing randomness tends to reduce factual reliability rather than improve it.
60. What is an embedding in the context of generative AI?
Answer and explanation
Answer: A. An embedding maps content into a vector space where semantically similar items sit close together, which is what makes similarity search possible. Model weights are learned parameters, a different artifact entirely. Embeddings are not an encryption mechanism. A cached response is an application-level optimisation unrelated to vector representation.
61. Which statement correctly describes the relationship between pre-training and fine-tuning?
Answer and explanation
Answer: D. A foundation model is pre-trained on a broad corpus at great expense, and fine-tuning then adjusts its weights on a much smaller task-specific dataset. The order is fixed: pre-training precedes fine-tuning. The two are distinct processes with different data scales and costs. Pre-training is normally done by the model provider, not by each customer.
62. A company wants to generate images from text descriptions. Which class of model is designed for this?
Answer and explanation
Answer: A. Diffusion models generate images by iteratively denoising random noise conditioned on a text prompt, and they underpin most current text-to-image systems. Forecasting models predict future values in a sequence. Gradient boosted trees and linear regression operate on tabular features and produce numeric or class predictions, not images.
63. Which AWS service allows a developer to experiment with prompts against several foundation models in a browser without writing code?
Answer and explanation
Answer: C. The Bedrock console provides playgrounds for text, chat, and image models where prompts and inference parameters can be tried interactively. CloudShell is a browser-based command line for AWS APIs. Ground Truth is a data labelling service. CodeBuild compiles and tests source code.
64. What is a hallucination in the output of a large language model?
Answer and explanation
Answer: B. A hallucination is plausible-sounding generated content that is not grounded in fact or in any supplied source, and its fluency is precisely what makes it dangerous. A refusal is a guardrail behaviour rather than a factual error. Exceeding the token limit truncates output. An endpoint error is an operational failure, not a content problem.
65. Which factor most directly drives the cost of using a foundation model through Amazon Bedrock on-demand pricing?
Answer and explanation
Answer: B. On-demand Bedrock pricing is consumption-based and charged per thousand input and output tokens, so token volume is the dominant cost driver. IAM user count, S3 storage, and Availability Zone usage are unrelated to model invocation charges.
66. What is the role of a vector database in a generative AI application?
Answer and explanation
Answer: D. A vector database indexes embeddings so a query embedding can be matched against stored content by semantic similarity, which is the retrieval step in a RAG system. Model weights live with the model service. Response caching is a separate optimisation using ordinary key-value storage. Authentication tokens belong in an identity service.
67. Which characteristic distinguishes a foundation model from a traditional task-specific machine learning model?
Answer and explanation
Answer: C. Foundation models are trained once on very broad data and then adapted through prompting, retrieval, or fine-tuning to many tasks, which is what distinguishes them from a model built for a single purpose. They require vastly more data to pre-train, not less. They can be customized through several mechanisms. They are commonly consumed as managed services rather than run on customer hardware.
68. A prompt includes three worked examples of the desired input and output before the actual request. What technique is this?
Answer and explanation
Answer: A. Supplying a small number of demonstrations inside the prompt is few-shot prompting, and it changes no model weights. Zero-shot prompting gives instructions with no examples. Fine-tuning and continued pre-training both modify weights through a training process rather than through prompt content.
69. What is the purpose of a system prompt in a generative AI application?
Answer and explanation
Answer: C. A system prompt carries the standing instructions and constraints that frame every user turn. Conversation history is separate and must be resupplied per request. Authentication is handled by IAM and application identity. Region selection is an API call parameter, not part of the prompt.
70. Which inference parameter limits sampling to the smallest set of tokens whose cumulative probability reaches a threshold?
Answer and explanation
Answer: B. Top-p restricts sampling to the smallest group of candidate tokens whose probabilities sum to the specified value. Top-k restricts to a fixed count of candidates regardless of their probabilities. Temperature rescales the distribution. Maximum tokens caps response length.
71. A model accepts both an image and a text question in the same request. What is such a model called?
Answer and explanation
Answer: D. Multimodal models accept and reason across more than one input type, such as images together with text. A diffusion model generates images from noise. An embedding model converts content into vectors. A classification model assigns a class label.
72. Which statement about model size and cost is generally accurate for foundation models?
Answer and explanation
Answer: D. Capability generally rises with size but so do price per token and latency, which is why routing simple tasks to smaller models is a common optimization. Larger models do not consume fewer tokens for the same task. Size and cost are related. Smaller models are not universally better in quality.
73. What does streaming a model response provide?
Answer and explanation
Answer: B. Streaming sends output incrementally so the user sees text almost immediately, though total generation time is unchanged. It does not reduce token consumption. Models remain stateless between requests regardless of streaming.
74. Which customization options does Amazon Bedrock offer for adapting a foundation model to a company's needs?
Answer and explanation
Answer: D. Bedrock supports fine-tuning with labelled prompt and response pairs and continued pre-training on unlabelled domain data. Source code and parameter count are properties of the provider's model. Retraining a base model from scratch is not a customer customization option.
75. An application sets maximum tokens to 100 and receives responses that stop mid-sentence. What is happening?
Answer and explanation
Answer: D. A maximum output token limit truncates generation when reached, which produces responses cut off mid-sentence. Comprehension failure produces irrelevant rather than truncated output. Low temperature makes output deterministic. Exceeding the context window with the prompt produces an error rather than a truncated completion.
76. Which statement best describes why two identical prompts can produce different responses?
Answer and explanation
Answer: B. Generation samples from a probability distribution, so variation is expected unless temperature and related parameters are set for deterministic output. Models are not retrained between requests. Region does not change the sampling behaviour. Tokenization of identical text is deterministic.
77. Which best describes what happens when a conversation exceeds a model's context window?
Answer and explanation
Answer: B. The context window is a hard limit, so an application must truncate or summarize older turns to stay within it. Models retain nothing outside the supplied context. Dropping earlier content can materially affect quality. The window is fixed per model and does not extend on demand.
78. A company is adapting a foundation model to its own domain and wants to spend the least it can at each stage. (Place the techniques in the order they should be attempted.)
Answer and explanation
Answer: B → D → C → A. Prompt refinement changes no weights, costs nothing to train, and frequently suffices, which is why it is always attempted first. Retrieval adds current proprietary knowledge without training. Fine-tuning adapts behaviour and format at the cost of a training run and curated data. Continued pre-training is the most expensive and is reserved for injecting broad domain knowledge the model lacks entirely. Each rung is attempted only when the one below has been shown insufficient.
79. How does the token-based pricing model of a foundation model affect the cost of an application?
Answer and explanation
Answer: C. Token-based pricing charges for input and output tokens, so a longer prompt or a longer response both increase cost. There is no fixed monthly charge for on-demand invocation. Request count alone ignores that requests vary enormously in length. Input tokens are charged as well as output tokens.
80. What does context engineering refer to in a foundation model application?
Answer and explanation
Answer: C. Context engineering is the practice of selecting, ordering, and formatting what goes into the context window so the model has what it needs and little else. Window size is a model property. Training on domain data is fine-tuning or continued pre-training. Response caching is a cost and latency optimization.
81. What is the purpose of the Model Context Protocol in an agentic AI system?
Answer and explanation
Answer: D. Model Context Protocol standardises how agents connect to external tools and data, so a tool exposed once can be consumed by any compatible client. Prompt compression, fine-tuning, and transport encryption are unrelated concerns.
82. In a multi-agent system, what is the purpose of memory management?
Answer and explanation
Answer: D. Memory management in an agentic system is about retaining and retrieving relevant context across a task, since models are stateless between calls. It does not refer to the compute environment's RAM, to weight placement, or to response caching.
83. Which AWS offering provides managed infrastructure for deploying and operating AI agents at scale?
Answer and explanation
Answer: A. Bedrock AgentCore provides managed infrastructure for running agents in production. Knowledge Bases provide retrieval for RAG. JumpStart offers pre-trained models and solution templates. Comprehend performs natural language processing tasks.
84. Which offering helps developers build agents with an open-source framework on AWS?
Answer and explanation
Answer: B. Strands Agents is an open-source framework for building agents that run on AWS. Guardrails filter content. Model Evaluation compares model quality. Textract extracts text and structure from documents.
85. Which AWS offering is an agentic development environment that assists developers as they build applications?
Answer and explanation
Answer: A. Kiro is an agentic development environment for building software with AI assistance. Amazon Quick serves business intelligence and analytics. Comprehend analyses text. Polly converts text to speech.
86. Which family of foundation models is developed by AWS and available through Amazon Bedrock?
Answer and explanation
Answer: D. Amazon Nova is AWS's own family of foundation models offered through Bedrock. Comprehend, Personalize, and Rekognition are purpose-built AI services rather than foundation model families.
87. What is a token in the context of a foundation model?
Answer and explanation
Answer: A. A token is the unit of text a model processes. Credentials, capacity, and licences are unrelated meanings of the word.
88. What does a foundation model's context window determine?
Answer and explanation
Answer: B. The context window bounds the text considered in one request. Throughput, parameter count, and training duration are separate properties.
89. Which architecture underlies most modern large language models?
Answer and explanation
Answer: C. Transformers underlie modern large language models. Convolutional networks are common in vision. Tree ensembles and linear regression are classical approaches.
90. What is an embedding?
Answer and explanation
Answer: A. An embedding is a vector representation capturing semantic similarity. Compression, encryption, and summarization are different operations.
91. What does the temperature parameter control in a foundation model's output?
Answer and explanation
Answer: A. Temperature governs sampling randomness. Speed, output length, and context size are controlled separately.
92. What is a hallucination in the context of generative AI?
Answer and explanation
Answer: D. A hallucination is plausible but unsupported content. Grammatical errors, length, and repetition are different defects.
93. Which characteristic of generative AI makes its output difficult to test with exact-match assertions?
Answer and explanation
Answer: D. Nondeterminism means the same input may yield different output, so exact-match testing fails. Latency, cost, and context limits are real constraints but do not affect determinism.
94. Which business metric would indicate that a generative AI customer service assistant is delivering value?
Answer and explanation
Answer: C. Resolution without escalation measures business outcome. Token counts, model counts, and knowledge base size describe activity and configuration.
95. Which AWS service provides pre-trained foundation models accessible through a managed API?
Answer and explanation
Answer: B. Bedrock provides foundation models through a managed API. SageMaker AI supports building and training custom models. Comprehend analyses text. Kendra provides enterprise search.
96. Which capability does Amazon SageMaker JumpStart provide?
Answer and explanation
Answer: B. JumpStart offers pre-built models and solutions to start from. Vector storage, guardrails, and billing analysis are provided by other services.
97. Which advantage does using a managed foundation model service provide over hosting a model yourself?
Answer and explanation
Answer: D. A managed service removes infrastructure management. Accuracy depends on the model rather than the hosting. Prompt engineering is still required. Inference is still charged.
98. Which limitation of foundation models affects their use for questions about recent events?
Answer and explanation
Answer: C. A training cutoff bounds what the model knows, which retrieval addresses. Input length, determinism, and multilingual capability are different properties.
99. Which advantage does generative AI offer over a rule-based system for drafting customer responses?
Answer and explanation
Answer: B. Generative models handle variation without enumerated rules. They do not guarantee correctness, and both evaluation and human review remain necessary.
100. Which business metric would indicate a generative AI document summarization tool is delivering value?
Answer and explanation
Answer: D. Time saved reaching a decision measures the business outcome. Volume, length, and corpus size describe activity.
101. Which limitation should be considered when a generative AI system is used for a decision with legal consequences?
Answer and explanation
Answer: A. Probabilistic output requires human accountability for consequential decisions. Document length, regulated use, and custom training are not the limitation at issue.
102. Which factor makes a generative AI use case more likely to succeed?
Answer and explanation
Answer: B. Tolerance for variation with human review on consequential output suits generative systems. Exact repeatability, zero error tolerance, and full autonomy all work against their characteristics.
103. Which consideration applies to the cost of a generative AI application at scale?
Answer and explanation
Answer: D. Token-based pricing means both input and output length drive cost. Fixed pricing, user counts, and parameter counts do not describe how the charge accrues.
104. Which capability of generative AI supports a customer service application?
Answer and explanation
Answer: C. Responding appropriately to varied phrasing is the capability. Factual guarantees do not exist, a knowledge base grounds the answers, and quality monitoring remains necessary.
105. Which Amazon Bedrock capability stores an organization's documents so a model can retrieve from them?
Answer and explanation
Answer: D. Knowledge Bases manage document ingestion and retrieval. Guardrails filter content, Prompt Management versions prompts, and model evaluation scores output.
106. Which Amazon Bedrock capability lets a model call external systems to complete a task?
Answer and explanation
Answer: B. Agents plan steps and call tools. Knowledge Bases supply retrieval, Guardrails filter, and evaluation scores quality.
107. Which cost consideration applies to Amazon Bedrock provisioned throughput?
Answer and explanation
Answer: A. Provisioned throughput reserves capacity billed continuously. Per-request billing is on-demand invocation. There is no general free tier of this kind, and billing does not follow model count.
108. Which advantage does Amazon Bedrock offer over hosting a foundation model on your own instances?
Answer and explanation
Answer: D. Bedrock removes infrastructure management. Accuracy depends on the model, prompt engineering is still required, and inference is charged.
109. Which capability does Amazon SageMaker AI provide that Amazon Bedrock does not?
Answer and explanation
Answer: C. SageMaker AI supports the full custom model lifecycle. Bedrock provides foundation model access, guardrails, and knowledge bases.