CloudOps Engineer – Associate
High availability, scaling and networking
Auto Scaling policies and lifecycle hooks, load balancing, multi-AZ patterns, VPC networking and the security controls an operator is expected to configure and troubleshoot.
Operations notes for AWS Certified CloudOps Engineer – Associate (SOA-C03), the credential formerly called SysOps Administrator – Associate. This is the operator exam: monitoring, remediation, automation, cost control and governance, tested from the point of view of the person on call.
2 topics, 8 study points. Everything here is exam-oriented: each point is a fact or a distinction that SOA-C03 items are built on. Test yourself against the practice exam once you can explain a section without re-reading it.
1. High Availability & Auto Scaling
Auto Scaling Groups (ASGs) automatically manage a fleet of EC2 instances, maintaining a desired capacity, scaling out when demand increases, and scaling in when demand decreases. An ASG is defined by a Launch Template (or the older Launch Configuration), which specifies the AMI, instance type, security groups, key pair, and user data for each instance. The ASG itself specifies the minimum, maximum, and desired number of instances, the Availability Zones to distribute instances across, and the health check configuration.
Auto Scaling supports three types of scaling policies. Simple scaling adds or removes a fixed number of instances in response to a single CloudWatch alarm, with a cooldown period before the next scaling action can occur. Step scaling improves on simple scaling by mapping different alarm thresholds to different scaling step sizes — a moderate CPU spike might add 2 instances while an extreme spike might add 5. Target Tracking scaling is the easiest and most modern approach: you specify a target value for a metric (for example, keep average CPU at 50%), and Auto Scaling automatically creates and manages the alarms and scaling actions to maintain that target. Target Tracking is the recommended default for most applications.
Health checks determine when Auto Scaling should replace an unhealthy instance. EC2 health checks evaluate System and Instance Status Checks — only gross failures (hardware or OS-level issues) trigger replacement. ELB health checks are more application-aware: the load balancer periodically sends HTTP requests to each instance and marks an instance unhealthy if it returns a non-2xx/3xx response or times out. Configuring your ASG to use ELB health checks (instead of the default EC2 checks) ensures that instances with application-level failures are replaced even if the OS appears healthy. When an instance is terminated by Auto Scaling, it follows a Termination Policy — by default, Auto Scaling terminates the instance in the AZ with the most instances, selecting the one with the oldest Launch Configuration.
Multi-AZ deployments are the primary mechanism for fault tolerance within a single AWS Region. For RDS, Multi-AZ creates a synchronous standby replica in a different AZ — if the primary instance fails, RDS automatically promotes the standby and updates the DNS endpoint to point to the new primary, typically within 1–2 minutes. For EC2 workloads behind an ASG and ELB, spreading instances across at least 3 AZs ensures that a single AZ failure removes at most one-third of your capacity while the load balancer automatically routes traffic to healthy instances in the surviving AZs. For ELB itself, the load balancer is managed by AWS across multiple AZs — you enable each AZ, and AWS handles the rest.
2. Networking & Security
Security Groups and Network ACLs (NACLs) are complementary layers of network security in a VPC. Security Groups operate at the instance level (or ENI level) and are stateful — if you allow an inbound connection, the response traffic is automatically allowed regardless of outbound rules. Security Groups only support allow rules; you cannot explicitly deny traffic to a specific IP range (you can only remove an allow rule). By default, all inbound traffic is denied and all outbound traffic is allowed. Changes to Security Group rules take effect immediately and apply to all existing connections as well as new ones.
NACLs operate at the subnet level and are stateless — both inbound and outbound rules must explicitly allow the traffic, including the ephemeral return ports (1024–65535) used for response traffic. NACLs support both allow and deny rules, which makes them the right tool for blocking specific IP addresses or CIDR ranges at scale. NACL rules are evaluated in numerical order — the lowest-numbered matching rule wins. A common NACL pattern is a broad allow rule at a high number (e.g., rule 100: allow all) with specific deny rules at lower numbers (e.g., rule 10: deny a specific attacker IP) — because rule 10 is evaluated first, the deny takes precedence.
VPC Flow Logs capture metadata about the IP traffic flowing through your VPC. You can enable Flow Logs at the VPC level (captures all traffic across all interfaces), subnet level (all interfaces in the subnet), or individual ENI level. Flow Logs record: source and destination IP addresses, source and destination ports, protocol, number of bytes, whether traffic was ACCEPTED or REJECTED, and timestamps. Flow Logs are invaluable for security investigations (tracing the source of unauthorized access attempts), troubleshooting connectivity issues (identifying rejected packets), and auditing network patterns. Flow Logs can be delivered to CloudWatch Logs for querying with Logs Insights or to S3 for long-term retention and analysis with Athena.
AWS WAF (Web Application Firewall) protects web applications from common exploits at Layer 7 — SQL injection, cross-site scripting (XSS), bad bots, and other OWASP Top 10 vulnerabilities. WAF rules can match on IP addresses, HTTP headers, URI strings, request body content, and geographic origin. WAF integrates with CloudFront, Application Load Balancers, API Gateway, and AppSync. AWS Shield provides DDoS protection at two tiers: Shield Standard is automatically included at no charge for all AWS customers and protects against the most common network and transport layer attacks (volumetric floods, SYN floods). Shield Advanced is a paid service offering enhanced detection and mitigation for sophisticated DDoS attacks, a 24/7 DDoS response team (DRT), and financial protection against scaling charges incurred during an attack.
Where to go next
- Back to the AWS SysOps Associate overview.
- Look up any service you could not name in the AWS services glossary.
- Sit the 80-item practice exam once two or three note pages are solid.
Last updated Sep 18, 2026