# Storage, deployment and provisioning

Storage and data management, deployment strategies, Lambda deployment configurations, advanced CloudFormation and Service Catalog — how infrastructure gets shipped and rolled back.

Operations notes for **AWS Certified CloudOps Engineer – Associate (SOA-C03)**, the credential formerly called SysOps Administrator – Associate. This is the operator exam: monitoring, remediation, automation, cost control and governance, tested from the point of view of the person on call.

**5 topics, 17 study points.** Everything here is exam-oriented: each point is a fact or a distinction that SOA-C03 items are built on. Test yourself against the [practice exam](/aws/practice-exam/) once you can explain a section without re-reading it.

## 1. Storage & Data Management

Amazon EBS (Elastic Block Store) provides persistent block storage volumes for EC2 instances. Unlike instance store (ephemeral storage physically attached to the host), EBS volumes exist independently of the instance lifecycle — they survive instance stops and reboots, and can be detached from one instance and reattached to another. EBS volumes come in four families: gp2 and gp3 (General Purpose SSD) balance cost and performance for most workloads; io1 and io2 (Provisioned IOPS SSD) deliver consistent, high IOPS for I/O-intensive databases; st1 (Throughput Optimized HDD) provides high sequential throughput for big data, log processing, and ETL workloads; sc1 (Cold HDD) is the lowest-cost option for infrequently accessed data.

EBS Snapshots are point-in-time backups of EBS volumes stored in Amazon S3 (in AWS-managed buckets — not your own S3 bucket). Snapshots are incremental: only the blocks that have changed since the last snapshot are saved, which minimizes storage costs. You can create a new EBS volume from any snapshot, in any AZ within the same region, at any size equal to or larger than the original. Snapshots can be copied across regions, enabling you to use them for disaster recovery or to migrate data between regions. For consistent snapshots of EBS volumes attached to running instances, use Amazon Data Lifecycle Manager (DLM) to automate snapshot creation and retention policies, or quiesce the application and flush the file system cache before taking the snapshot.

S3 lifecycle policies automate the transition of objects between storage classes as they age. A typical lifecycle policy might keep objects in S3 Standard for 30 days (active access), transition them to S3 Standard-IA after 30 days (occasional access), move them to S3 Glacier after 90 days (archival), and delete them after 365 days. Lifecycle rules can apply to the entire bucket or be scoped to specific prefixes or object tags. Combining lifecycle policies with versioning allows you to expire old versions of objects automatically, preventing storage costs from accumulating indefinitely in a versioned bucket.

AWS Backup is a centralized, policy-driven service for automating data backups across AWS services. Instead of managing separate backup configurations for RDS, EBS, DynamoDB, EFS, FSx, and Storage Gateway independently, AWS Backup lets you define a single backup plan that covers all of them. A backup plan specifies the backup frequency, retention period, and backup vault (the S3-backed storage location for backups). AWS Backup supports cross-region and cross-account backup copies, making it the standard tool for multi-region disaster recovery strategies that require consistent, auditable backup policies across an entire AWS Organization.

## 2. Deployment & Provisioning

AWS CloudFormation enables Infrastructure as Code (IaC) — you define your entire AWS infrastructure in declarative JSON or YAML templates, and CloudFormation provisions, updates, and deletes resources in the correct order. A CloudFormation Stack is a collection of AWS resources managed together as a single unit. When you update a stack, CloudFormation calculates what changed and applies updates in place or replaces resources as needed. ChangeSets allow you to preview exactly which resources will be added, modified, or deleted before you execute an update — a critical safety mechanism for production changes. Nested Stacks break large templates into reusable modules, with a parent stack referencing child stacks as resources.

AWS Systems Manager (SSM) is the operational hub for managing EC2 instances and on-premises servers at scale. Run Command executes shell scripts or PowerShell commands across any number of managed instances simultaneously, with full output logging — eliminating the need to SSH into individual instances. Patch Manager automates OS and application patching, with configurable patch baselines and maintenance windows to control when patches are applied. Parameter Store provides a secure, hierarchical store for configuration data and secrets — you can store plain-text parameters (database hostnames, feature flags) or SecureString parameters encrypted with KMS (passwords, API keys). Session Manager provides browser-based and CLI-based shell access to instances without opening port 22 or maintaining bastion hosts, with every session logged to CloudWatch and S3.

AWS Config is a service for recording, auditing, and evaluating the configuration of your AWS resources. Config continuously records configuration changes (a "configuration item" is created every time a resource is modified) and stores this history so you can answer the question "what did this resource look like at 3 PM last Tuesday?" Config Rules evaluate whether your resources comply with your organization's policies — for example, a rule that ensures all EC2 instances are in a VPC, all S3 buckets have server-side encryption enabled, or no security group allows inbound access on port 22 from 0.0.0.0/0. Config is reactive: it identifies non-compliant resources after they are created or modified, and can trigger automatic remediation via Systems Manager Automation documents.

AWS Organizations provides centralized governance and management for multiple AWS accounts. Service Control Policies (SCPs) are permission boundaries that apply to all accounts and users within an organizational unit — even the root user of a member account cannot exceed what an SCP allows. SCPs are ideal for preventing entire account types from using services outside their mandate (for example, preventing development accounts from launching resources in production regions). Consolidated Billing aggregates the usage from all accounts for volume discounts and simplifies the payment process to a single invoice. AWS Trusted Advisor evaluates your AWS environment against best practices in five categories: Cost Optimization, Performance, Security, Fault Tolerance, and Service Limits. It surfaces actionable recommendations — identifying underutilized EC2 instances, security groups with unrestricted access, and resources approaching service limits.

## 3. Lambda Deployment Configurations

AWS Lambda supports traffic-shifting deployment configurations through AWS CodeDeploy integration, enabling gradual, controlled rollouts of new Lambda function versions. Instead of instantly switching all traffic from the old version to the new version (which provides no safety net), traffic-shifting deployments allow you to route a percentage of invocations to the new version while monitoring for errors, and automatically roll back if something goes wrong. This pattern is essential for production Lambda functions where a bad deployment could immediately affect all users.

The three deployment configuration types differ in how quickly they shift traffic. Canary deployments split traffic in exactly two steps: first, a configurable percentage (for example, 10%) is shifted to the new version for a configurable time period (for example, 10 minutes). If no alarms trigger during that period, the remaining 90% of traffic is shifted immediately. If an alarm fires, CodeDeploy automatically rolls back to the original version. Canary deployments are the most conservative option — they limit initial exposure to a small slice of traffic before committing to a full rollout.

Linear deployments shift traffic in equal increments at equal time intervals until 100% is reached. For example, Linear10PercentEvery10Minutes shifts 10% of traffic every 10 minutes — after 100 minutes, the full migration is complete. This provides more gradual exposure than Canary and gives you a longer observation window with alarms monitoring each incremental step. All-at-once deployments shift 100% of traffic immediately to the new version — this is the fastest option but provides no staged rollout safety. All-at-once is appropriate for non-production environments (development, testing) or for functions where instant rollback via Lambda versioning aliases is sufficient protection.

## 4. AWS CloudFormation (Advanced)

CloudFormation provides several mechanisms for handling complex provisioning scenarios. The cfn-signal helper script is used to signal CloudFormation whether the bootstrapping of an EC2 instance was successful. By default, CloudFormation marks a stack resource as CREATE_COMPLETE as soon as the EC2 instance starts — it has no visibility into whether your user data script installed software successfully. Using cfn-signal with a CreationPolicy (specifying how many success signals to wait for, and a timeout), CloudFormation waits to mark the resource complete until the instance itself signals success. If the script fails or the timeout expires without enough signals, CloudFormation marks the resource as CREATE_FAILED and rolls back the stack.

DeletionPolicy controls what happens to a resource when its CloudFormation stack is deleted. The default behavior (Delete) destroys the resource — appropriate for ephemeral resources like EC2 instances. Retain keeps the resource alive after stack deletion — useful for S3 buckets containing important data or RDS instances you want to keep even after decommissioning a stack. Snapshot creates a final snapshot before deletion — critical for stateful resources like RDS databases and EBS volumes, ensuring you have a restore point even after the infrastructure is decommissioned. Setting DeletionPolicy: Snapshot on RDS resources is a best practice for any production database managed by CloudFormation.

CloudFormation Drift Detection identifies resources in your stack that have been modified outside of CloudFormation — for example, someone manually changed a Security Group rule in the console. When drift is detected, CloudFormation shows you the actual versus expected configuration for each drifted resource. This is essential for maintaining configuration governance in environments where multiple teams have console access. Remediation options include updating the template to match the actual state (accepting the drift), reverting the resource to the template state (correcting the drift), or importing the resource back under CloudFormation management. Stack Policies protect resources from accidental updates during stack updates — you define a policy that explicitly allows or denies update actions on specific resources.

## 5. AWS Service Catalog

AWS Service Catalog allows organizations to create and manage catalogs of approved IT services — CloudFormation templates, Terraform configurations, AMIs, and other AWS resources — and make them available to end users through a self-service portal. Instead of giving developers direct access to CloudFormation or the full AWS console (which could result in non-standardized, non-compliant infrastructure), Service Catalog provides a curated menu of pre-approved, governance-compliant products. Developers choose from the catalog, provide required parameters, and Service Catalog provisions the infrastructure on their behalf using a pre-configured IAM role.

Service Catalog products are CloudFormation templates packaged with metadata: a product name, description, version history, and constraints. Portfolios are collections of products shared with specific IAM users, groups, or roles. Launch Constraints specify the IAM role that Service Catalog uses to provision the product — this means end users do not need direct CloudFormation or EC2 permissions; they only need Service Catalog permissions. The launch constraint role has exactly the permissions needed to provision the specific infrastructure in the template, following the principle of least privilege.

A critical operational distinction between Service Catalog and AWS Config: Service Catalog is a proactive governance mechanism. When a user provisions a product from Service Catalog, the resulting resources are guaranteed to be created according to the approved CloudFormation template — non-standard configurations simply cannot be provisioned because they are not in the catalog. AWS Config, by contrast, is reactive: it evaluates resources after they are created and flags non-compliant ones. Service Catalog ensures tagging requirements, approved instance types, and security configurations are enforced at provisioning time, while AWS Config continuously audits the runtime state of all resources. In practice, both are used together for comprehensive governance.

---

## Where to go next

- Back to the [AWS SysOps Associate overview](/aws/sysops-associate/).
- Look up any service you could not name in the [AWS services glossary](/aws/services-glossary/).
- Sit the [80-item practice exam](/aws/practice-exam/) once two or three note pages are solid.
