AWS ECS/Fargate Cheat Sheet
Task definitions, services, CLI commands, and networking modes for running containers on ECS with the Fargate launch type.
Fargate Task Definition
Minimal task definition requiring awsvpc networking and explicit CPU/memory.
{ "family": "web-app", "requiresCompatibilities": ["FARGATE"], "networkMode": "awsvpc", "cpu": "256", "memory": "512", "executionRoleArn": "arn:aws:iam::123456789012:role/ecsTaskExecutionRole", "containerDefinitions": [ { "name": "web", "image": "123456789012.dkr.ecr.us-east-1.amazonaws.com/web-app:latest", "portMappings": [{ "containerPort": 8080, "protocol": "tcp" }], "logConfiguration": { "logDriver": "awslogs", "options": { "awslogs-group": "/ecs/web-app", "awslogs-region": "us-east-1", "awslogs-stream-prefix": "web" } } } ]}
Register Task & Create Service (CLI)
Register a task definition, then run it as a load-balanced service.
aws ecs register-task-definition --cli-input-json file://task-def.jsonaws ecs create-service \ --cluster prod-cluster \ --service-name web-app \ --task-definition web-app \ --desired-count 3 \ --launch-type FARGATE \ --network-configuration "awsvpcConfiguration={subnets=[subnet-abc,subnet-def],securityGroups=[sg-123],assignPublicIp=DISABLED}" \ --load-balancers "targetGroupArn=arn:aws:elasticloadbalancing:...,containerName=web,containerPort=8080"
ECS Exec into a Running Container
Debug a live task with an interactive shell (requires enableExecuteCommand).
# Enable exec when creating/updating the serviceaws ecs update-service --cluster prod-cluster --service web-app --enable-execute-command# Open a shell in the running taskaws ecs execute-command \ --cluster prod-cluster \ --task arn:aws:ecs:us-east-1:123456789012:task/prod-cluster/abc123 \ --container web \ --command "/bin/sh" \ --interactive
Fargate Sizing & Networking
Key constraints to know before writing a task definition.
- CPU/memory pairs- fixed combos, e.g. 256 CPU units requires 512/1024/2048 MB memory
- awsvpc mode- required for Fargate; each task gets its own ENI and private IP
- Execution role- pulls images and writes logs (ecsTaskExecutionRole)
- Task role- grants the app itself permissions to call other AWS services
- assignPublicIp- ENABLED needed for public subnets without a NAT gateway
- Fargate Spot- up to 70% cheaper, interruptible; mix with capacityProviderStrategy
Sidecar Containers: App + Log Router + Init
A production task typically runs more than one container — here an app container depends on a FireLens log router and shares data via an ephemeral volume.
{ "family": "web-app", "requiresCompatibilities": ["FARGATE"], "networkMode": "awsvpc", "cpu": "512", "memory": "1024", "volumes": [{ "name": "shared-tmp" }], "containerDefinitions": [ { "name": "web", "image": "123456789012.dkr.ecr.us-east-1.amazonaws.com/web-app:latest", "essential": true, "dependsOn": [{ "containerName": "log-router", "condition": "START" }], "mountPoints": [{ "sourceVolume": "shared-tmp", "containerPath": "/tmp/shared" }], "logConfiguration": { "logDriver": "awsfirelens", "options": { "Name": "cloudwatch", "region": "us-east-1", "log_group_name": "/ecs/web-app" } } }, { "name": "log-router", "image": "amazon/aws-for-fluent-bit:stable", "essential": true, "firelensConfiguration": { "type": "fluentbit" } } ]}
ECS Service Connect for Service-to-Service DNS
Service Connect gives services a stable DNS name and built-in client-side load balancing/retries without running a separate service mesh.
{ "serviceConnectConfiguration": { "enabled": true, "namespace": "prod-namespace", "services": [ { "portName": "web-8080", "discoveryName": "web-app", "clientAliases": [{ "port": 8080, "dnsName": "web-app" }] } ], "logConfiguration": { "logDriver": "awslogs", "options": { "awslogs-group": "/ecs/service-connect", "awslogs-region": "us-east-1", "awslogs-stream-prefix": "sc" } } }}
Deployment Circuit Breaker with Automatic Rollback
Detect a failing rolling deployment (tasks stuck unhealthy) and auto-rollback to the last stable task definition instead of paging someone at 2am.
aws ecs update-service \ --cluster prod-cluster \ --service web-app \ --deployment-configuration "deploymentCircuitBreaker={enable=true,rollback=true},maximumPercent=200,minimumHealthyPercent=100" \ --health-check-grace-period-seconds 60
Target-Tracking Auto Scaling on CPU + Custom Metric
Register the service as a Application Auto Scaling target, then attach a target-tracking policy so Fargate task count follows load without manual step scaling rules.
aws application-autoscaling register-scalable-target \ --service-namespace ecs \ --resource-id service/prod-cluster/web-app \ --scalable-dimension ecs:service:DesiredCount \ --min-capacity 3 --max-capacity 50aws application-autoscaling put-scaling-policy \ --service-namespace ecs \ --resource-id service/prod-cluster/web-app \ --scalable-dimension ecs:service:DesiredCount \ --policy-name cpu-target-tracking \ --policy-type TargetTrackingScaling \ --target-tracking-scaling-policy-configuration '{ "TargetValue": 60.0, "PredefinedMetricSpecification": { "PredefinedMetricType": "ECSServiceAverageCPUUtilization" }, "ScaleOutCooldown": 30, "ScaleInCooldown": 120 }'
Advanced Operational Concerns
Beyond sizing and basic networking — what to plan for once Fargate services are running real production traffic.
- Task ENI cold-start latencyeach Fargate task provisions a fresh ENI on awsvpc mode, adding tens of seconds to task startup versus EC2 launch type — factor this into scale-out responsiveness
- Ephemeral storage20GB by default, configurable up to 200GB per task; useful for large temp files without an EFS mount, but it's wiped on task stop
- EFS volumes for shared stateFargate tasks can mount EFS directly for durable/shared storage across tasks, unlike ephemeral storage which is task-local
- Container health checks vs. ALB health checksdefine a HEALTHCHECK in the container definition for ECS-level restarts, separate from the ALB target group health check that controls routing
- Graceful shutdown (SIGTERM)ECS sends SIGTERM then SIGKILL after stopTimeout (default 30s); handle SIGTERM in-app to drain in-flight requests before the hard kill
- Bin packing vs. FargateFargate has no bin-packing control since AWS manages the underlying capacity — if tight cost optimization via bin packing matters, that's an EC2 launch-type tradeoff
Set a capacityProviderStrategy that blends FARGATE and FARGATE_SPOT (e.g. weight 1/3) instead of an all-or-nothing launch type — ECS will automatically replace interrupted Spot tasks while keeping a stable baseline on standard Fargate.