AI Practitioner
AWS AI services, SageMaker and Bedrock
The pre-trained Amazon AI services, Amazon SageMaker, Amazon Bedrock and SageMaker Clarify — which service solves which problem, and how they differ.
Foundational notes for AWS Certified AI Practitioner (AIF-C01). Unlike the AI Business Strategist credential, this one does assess AWS AI services — so the vocabulary pages and the service pages carry equal weight.
4 topics, 14 study points. Everything here is exam-oriented: each point is a fact or a distinction that AIF-C01 items are built on. Test yourself against the practice exam once you can explain a section without re-reading it.
1. Pre-Trained Amazon AI Services
AWS offers a portfolio of fully managed, pre-trained AI services that expose powerful ML capabilities through simple API calls, requiring no ML expertise or model training. These services are purpose-built for specific modalities and tasks, and they are the right tool when you need to add a specific capability (speech-to-text, image analysis, NLP) without building or managing models yourself.
Amazon Polly is a text-to-speech service that converts written text into lifelike spoken audio in dozens of languages and voices. It uses deep learning to produce natural-sounding speech and supports SSML (Speech Synthesis Markup Language) for fine-grained control over pronunciation, pauses, and emphasis. Polly is used for building voice interfaces, accessibility features, audio content generation, and notification systems. Amazon Transcribe performs the reverse: it converts audio speech recordings into text using deep learning. It supports real-time streaming transcription and batch transcription, speaker diarization (distinguishing between multiple speakers), custom vocabularies for domain-specific terminology, and automatic punctuation.
Amazon Comprehend is a Natural Language Processing (NLP) service that extracts insights and relationships from unstructured text without any ML expertise. It can identify entities (people, places, organizations, dates), detect the sentiment of text (positive, negative, neutral, mixed), classify topics, extract key phrases, and detect the language of the text. Amazon Lex is the service for building conversational interfaces — chatbots and voice assistants — for applications using both voice and text. It provides automatic speech recognition (ASR), natural language understanding (NLU), and dialog management to enable multi-turn conversations. Lex is the technology behind Amazon Alexa.
Amazon Rekognition is a computer vision service for analyzing images and videos. It can detect objects, scenes, activities, and inappropriate content; recognize faces and compare them against a database; read text in images (OCR); and analyze video for activities and person tracking. It requires no ML expertise and processes both real-time streams and stored media. Amazon Textract goes beyond simple OCR to intelligently extract structured data from scanned documents — it identifies forms (key-value pairs), tables, and signatures even in complex layouts like tax forms, medical records, and financial documents. The critical exam distinction: Comprehend finds insights in text; Transcribe converts speech to text; Lex builds conversational bots; Rekognition analyzes images and videos; Textract extracts structured data from scanned documents.
2. Amazon SageMaker
Amazon SageMaker is AWS’s fully managed, end-to-end machine learning platform. It covers the complete ML lifecycle: data preparation, model training, hyperparameter tuning, model evaluation, deployment, and monitoring. SageMaker provides ML engineers and data scientists with a unified environment (SageMaker Studio) plus a suite of purpose-built tools for each phase of the ML workflow. Understanding which SageMaker tool handles which phase is a key exam competency.
SageMaker Data Wrangler is the data preparation tool. It offers 300+ pre-built data transformations for cleaning, transforming, and normalizing tabular data, time-series data, and images — all without writing code. Data Wrangler integrates with S3, Redshift, Athena, and other data sources and exports prepared datasets directly to SageMaker training pipelines. SageMaker Canvas is the no-code ML solution for business users who need to build predictions without writing code. It uses a point-and-click interface powered by AutoML under the hood — you import data, select the target column, and Canvas automatically trains, evaluates, and deploys a model. SageMaker Ground Truth is the human-in-the-loop annotation service. It provides tools to build labeled datasets through a combination of automated labeling (using active learning to auto-label high-confidence examples) and human annotation via Amazon Mechanical Turk, private workforces, or third-party vendors.
SageMaker JumpStart is an ML hub providing access to hundreds of pre-trained foundation models, open-source algorithms, and solution templates. It allows you to evaluate, compare, and deploy FMs (like Llama 2, Falcon, and others) with one click and fine-tune them on your own data with minimal code. SageMaker Feature Store is a purpose-built repository for storing, sharing, and managing ML features across teams and projects. Features are computed once and stored in Feature Store, where they can be retrieved for both training (offline store) and real-time inference (online store), ensuring training-serving consistency. SageMaker Automatic Model Tuning (AMT) automates hyperparameter optimization by running multiple training jobs with different hyperparameter combinations and identifying the combination that maximizes model performance. MLflow integration in SageMaker provides experiment tracking — logging parameters, metrics, and artifacts from iterative training runs for comparison and reproducibility.
SageMaker Model Dashboard provides a centralized view of all models in your account — their deployment status, whether they are serving inference endpoints or running batch transform jobs, and their real-time performance metrics (CPU, GPU, disk, memory utilization). SageMaker Model Monitor continuously monitors deployed model endpoints for data quality drift (incoming feature distributions diverging from training distributions), model quality drift (prediction accuracy degrading over time), bias drift (demographic fairness metrics changing), and feature attribution drift (SHAP values shifting). When drift is detected, Model Monitor generates alerts that can trigger CloudWatch alarms or kick off retraining pipelines.
3. Amazon Bedrock
Amazon Bedrock is AWS’s fully managed service for accessing, customizing, and deploying foundation models from multiple providers through a single, serverless API. Instead of building and managing your own model infrastructure, Bedrock gives you on-demand access to high-performance FMs from Anthropic (Claude), Meta (Llama), Mistral AI, Stability AI, Cohere, and Amazon (Titan) — all without managing a single server. You pay per token processed (input + output), making it cost-effective for variable workloads.
Bedrock supports both fine-tuning and continued pre-training for compatible models. Fine-tuning in Bedrock requires labeled training data uploaded to S3 in a specific JSONL format; Bedrock handles the training infrastructure and returns a fine-tuned model version that you can deploy to an endpoint. Continued pre-training in Bedrock uses unlabeled domain-specific data to broaden a model’s knowledge — the ideal pattern for domain adaptation (medical, legal, financial) where you want the model to internalize industry terminology and reasoning patterns without targeting a specific task.
Bedrock Guardrails allow you to implement safety and compliance policies across all your Bedrock-powered applications from a centralized configuration. Guardrails can filter harmful content (violence, hate speech, sexual content), block topic areas that are off-limits for your application, redact sensitive personal information (PII), and restrict the model’s responses to specific topics. Amazon Q Business is a ready-to-use enterprise generative AI assistant built on Bedrock. It connects to your organization’s data sources via pre-built connectors (SharePoint, Confluence, S3, Salesforce, Jira) and enables employees to ask natural language questions that are answered from your own internal knowledge — with source citations and access controls that respect the permissions of the underlying data source.
4. SageMaker Clarify & Model Explainability
Amazon SageMaker Clarify is a service that addresses two critical aspects of responsible ML: bias detection and model explainability. It helps ML teams understand why a model makes the predictions it does, identify whether the model treats different demographic groups unfairly, and generate the documentation needed for regulatory and audit purposes. Clarify is model-agnostic — it works with any ML model framework.
For bias detection, Clarify analyzes your training dataset and model predictions to identify statistical imbalances in how different demographic groups (age, gender, ethnicity) are represented in the data and treated by the model. You specify “sensitive features” that should not drive predictions, and Clarify computes bias metrics before training (data bias) and after (model bias). This is critical for applications where biased predictions have legal or ethical consequences — loan approvals, hiring decisions, healthcare triage. SageMaker Ground Truth provides the human feedback loop that helps correct biased or incorrect labels in training data before they propagate into the model.
For model explainability, Clarify uses SHAP (SHapley Additive exPlanations) — a mathematically rigorous method from cooperative game theory that assigns each input feature a contribution score for a specific prediction. SHAP values provide local explanations: they tell you, for this specific loan rejection, how much each feature (income, credit score, zip code) contributed to or against the model’s decision. Partial Dependence Plots (PDPs) provide global explanations: they show the marginal effect of varying a single feature across its full range while holding all other features constant, revealing the overall relationship between a feature and the model’s output across the entire dataset. Local (SHAP) explanations are used for individual decision auditing; global (PDP) explanations are used for understanding overall model behavior.
Where to go next
- Back to the AWS Certified AI Practitioner overview.
- Look up any service you could not name in the AWS services glossary.
- Sit the 80-item practice exam once two or three note pages are solid.
Last updated Sep 18, 2026