AI Practitioner
AI and machine-learning foundations
Learning paradigms, deterministic versus probabilistic models, train/validation/test splits, bias and variance, evaluation metrics, embeddings and transfer learning.
Foundational notes for AWS Certified AI Practitioner (AIF-C01). Unlike the AI Business Strategist credential, this one does assess AWS AI services — so the vocabulary pages and the service pages carry equal weight.
6 topics, 20 study points. Everything here is exam-oriented: each point is a fact or a distinction that AIF-C01 items are built on. Test yourself against the practice exam once you can explain a section without re-reading it.
1. Types of Machine Learning
Machine learning is the discipline of training software systems to make predictions or decisions from data, rather than programming them with explicit rules. There are five core learning paradigms, and understanding which one applies to a given scenario is a foundational skill for the AI Practitioner exam. The right approach depends on the type of data you have (labeled or not), the task you are solving (prediction, generation, pattern discovery), and the feedback mechanism available during training.
Supervised learning trains models on labeled data — datasets where each input has a known, annotated output. The model learns to map inputs to outputs by minimizing the error between its predictions and the true labels. Classic supervised learning algorithms include linear regression (predicting continuous values, like a house price based on location and size) and neural networks (learning complex non-linear mappings, like recognizing handwritten digits). Decision trees and random forests are also supervised learning models. The defining characteristic is that a human has already done the work of labeling what the correct output should be for each training example.
Unsupervised learning trains on unlabeled data — the model receives inputs but no pre-defined outputs and must discover patterns on its own. Clustering algorithms (like k-means) group similar data points together, which is how recommendation systems identify customer segments or how network traffic analysis identifies anomalous patterns. Association rule learning (like the Apriori algorithm) discovers relationships between items — the classic example being market basket analysis: “customers who buy coffee also tend to buy milk.” Neither clustering nor association rule learning requires labels, making them applicable to massive unstructured datasets.
Semi-supervised learning combines both approaches: a small amount of labeled data and a large amount of unlabeled data are used together. The model first trains partially on the labeled data, then uses that partial knowledge to generate pseudo-labels for the unlabeled data, and finally retrains on the combined set. This is cost-effective for tasks where labeling every example would be prohibitively expensive — for example, labeling sentiment across millions of customer support tickets or categorizing a vast document library where only a small sample has been manually classified.
Self-supervised learning is the paradigm used to train Foundation Models. The model receives vast amounts of completely raw, unlabeled data and generates its own labels from the data structure itself. For text, a common self-supervised task is “predict the next word given all the previous words” — no human labeling is required because the correct answer is already in the text. This allows training on internet-scale datasets. Reinforcement learning is a fifth paradigm where an agent takes actions in an environment, receives reward or penalty signals, and learns to maximize cumulative reward over time. AWS DeepRacer is an accessible example: an autonomous model car learns to navigate a physical track by receiving a reward each time it stays on course.
2. Deterministic vs Probabilistic Models
An important conceptual distinction in ML is whether a model is deterministic or probabilistic — or a combination of both. This affects how you interpret model outputs, how you handle uncertainty in predictions, and which evaluation strategies are appropriate.
Deterministic models always produce the same output for the same input. Their behavior is fully predictable and repeatable. Decision trees are the canonical example: given the same input features, a decision tree follows the exact same path through its branches and always arrives at the same leaf node output. Rule-based systems, most classical algorithms, and lookup tables are also deterministic. Deterministic models are easy to audit and explain, which makes them attractive for regulated industries where you need to demonstrate why a specific decision was made.
Probabilistic models produce a distribution of possible outcomes rather than a single fixed answer. They quantify uncertainty and model randomness explicitly. Bayesian Networks represent probabilistic relationships between variables and provide probability estimates for different outcomes given observed evidence — for example, estimating the probability that a patient has a disease given certain symptoms. Naive Bayes classifiers, Gaussian Mixture Models, and Hidden Markov Models are also probabilistic. Many modern ML models — including neural networks — are technically a mix of both: their architecture is deterministic (the same weights always produce the same softmax logits), but the sampling process during generation introduces randomness (controlled by inference parameters like Temperature).
3. Training, Validation & Test Sets
A fundamental practice in supervised machine learning is splitting your labeled dataset into distinct subsets for different phases of model development. Using the same data for training and evaluation leads to overly optimistic performance estimates — the model may simply memorize the training examples rather than learning generalizable patterns. Proper data splitting is how you get an honest assessment of how your model will perform on real-world, unseen data.
The training set is the largest portion of your dataset and is the only data the model is allowed to learn from. The model’s weights, parameters, and internal representations are adjusted based entirely on training data through the optimization process (typically gradient descent). The validation set (also called the development or dev set) is used during the training process to tune hyperparameters — settings that control the training process itself, like learning rate, regularization strength, and model architecture choices. You evaluate the model on the validation set periodically during training to detect overfitting early and to compare different model configurations. Importantly, validation sets are optional: for small datasets, cross-validation is often used instead, which repeatedly partitions the training data into training and validation folds.
The test set is held out entirely and only used once — after all training and hyperparameter tuning is complete. It provides an unbiased estimate of how well the model will generalize to truly new, unseen data in production. The test set must never be used during training or validation, because any exposure to it would compromise the evaluation. A common mistake is “leaking” the test set into model selection decisions, which leads to over-optimistic performance estimates. The sequence is always: train on training data → tune using validation data → evaluate final model once on test data.
4. Bias, Variance & Model Fit
The bias-variance tradeoff is one of the most fundamental concepts in machine learning theory. It describes a fundamental tension in model design: models that are too simple fail to capture the patterns in the data (high bias), while models that are too complex learn the noise in the training data and fail to generalize (high variance). Finding the right balance is central to building models that perform well both on training data and on real-world inputs they have never seen.
Bias refers to error introduced by oversimplified assumptions in the model. A model with high bias underfits the data — it fails to capture the underlying structure and performs poorly even on the training set. Imagine trying to fit a straight line to data that follows a curve: the line will consistently miss the pattern regardless of how much data you show it. High-bias models are inflexible. Variance refers to error introduced by the model’s sensitivity to small fluctuations in the training data. A model with high variance overfits — it learns the training data so thoroughly (including its noise and random fluctuations) that it fails to generalize to new examples. A highly complex neural network trained on a small dataset might achieve 99% accuracy on training data but only 60% on the test set.
Several techniques address overfitting (high variance). Cross-validation repeatedly trains and evaluates on different partitions of the data, ensuring the model is not just memorizing one specific data split. Regularization (L1 and L2) adds a penalty term to the loss function that discourages large model weights, forcing the model to learn simpler, more general representations. L1 regularization (Lasso) drives some weights to exactly zero (feature selection). L2 regularization (Ridge) shrinks weights toward zero without eliminating them. Dropout randomly disables neurons during training, preventing the network from becoming over-reliant on any single pathway. Pruning simplifies decision trees by removing branches that contribute little predictive power. Hyperparameter tuning — adjusting learning rates, regularization coefficients, model depth, and batch sizes — is the practical lever for improving generalization in production ML pipelines.
5. Model Evaluation Metrics
Choosing the right evaluation metric is as important as choosing the right model architecture. Different metrics emphasize different types of errors, and the appropriate metric depends entirely on the business problem and the relative cost of different mistakes. The exam frequently presents medical or safety scenarios precisely because these domains make the cost asymmetry of errors very concrete.
For classification problems, the confusion matrix is the foundation of all evaluation. It tabulates four outcomes: True Positives (correctly predicted positives), True Negatives (correctly predicted negatives), False Positives (predicted positive but actually negative — “false alarms”), and False Negatives (predicted negative but actually positive — “missed detections”). Precision measures how trustworthy positive predictions are: of all the cases the model flagged as positive, what fraction were actually positive? High precision means few false alarms. Recall (also called sensitivity) measures how comprehensive the model is: of all actual positive cases, what fraction did the model correctly identify? High recall means few missed detections. The F1-Score is the harmonic mean of Precision and Recall, providing a single balanced metric. In medical diagnosis, recall is typically prioritized (you cannot afford to miss a disease) even at the cost of lower precision (some healthy patients will be flagged for further testing).
For regression problems (predicting continuous values), different metrics apply. Mean Absolute Error (MAE) is the average of the absolute differences between predictions and actual values — it is easy to interpret (in the same units as the target variable) and treats all errors equally. Root Mean Squared Error (RMSE) squares the differences before averaging (then takes the square root), which means large errors are penalized more heavily than small ones. RMSE is sensitive to outliers, making it appropriate when large prediction errors are particularly costly. The Pearson correlation coefficient (R) measures the linear relationship between predicted and actual values — R² (R-squared) indicates what proportion of the variance in the target variable is explained by the model, ranging from 0 (no explanatory power) to 1 (perfect prediction).
6. Embeddings & Transfer Learning
Embedding models are algorithms trained to convert high-dimensional data (words, sentences, images) into dense numerical vectors in a multi-dimensional space. These dense vector representations (embeddings) capture semantic meaning and relationships — words or concepts with similar meanings are mapped to nearby points in the embedding space, while unrelated concepts are mapped far apart. Embeddings are the bridge between raw human-readable data and the numerical operations that machine learning models perform internally.
Word2Vec is an early and influential embedding model that creates static vector representations based on word co-occurrence patterns in a corpus. Its limitation is that each word has exactly one vector, regardless of context — “bank” has the same embedding whether it refers to a financial institution or a riverbank. BERT (Bidirectional Encoder Representations from Transformers) addressed this with contextual embeddings: BERT generates a different embedding for a word depending on the surrounding text, enabling it to capture the nuanced meanings of polysemous words. Unlike Word2Vec (which only looks forward), BERT processes context from both directions simultaneously. PCA (Principal Component Analysis) and SVD (Singular Value Decomposition) are dimensionality reduction techniques from linear algebra — they can reduce the dimensions of embedding spaces but do not understand contextual language semantics the way neural embedding models do.
Transfer learning is the practice of taking a model that was already trained on one task (the source task) and adapting it to perform a different but related task (the target task), rather than training a new model from scratch. This dramatically reduces the data and compute required for the target task, since the model already has learned general features. The process involves three steps. First, select a pre-trained model whose source task is sufficiently related to your target task — a model pre-trained on general English text is a strong starting point for a domain-specific text classification task. Second, configure the model: you typically freeze the weights of the early layers (which capture general low-level features) and remove the final task-specific layer, then add new layers designed for your target task. Third, train the model on target task data, adjusting only the new layers (or fine-tuning all layers with a small learning rate). Hyperparameters — learning rate, regularization, dropout rates — are adjusted during this phase to optimize performance on the target task.
Where to go next
- Back to the AWS Certified AI Practitioner overview.
- Look up any service you could not name in the AWS services glossary.
- Sit the 80-item practice exam once two or three note pages are solid.
Last updated Sep 18, 2026