What you will learn
- What AWS Batch is and the problem it solves that Lambda cannot
- AWS Batch components — Job Definitions, Job Queues, and Compute Environments
- AWS Batch on Fargate vs EC2 and when each is right
- Spot Instances with Batch for 90% compute cost savings
- Amazon Outposts — running AWS infrastructure on-premises
- Outposts use cases — data residency, low latency to on-premises systems
- AWS AppFlow — no-code SaaS data integration
- AWS Amplify — full-stack serverless app deployment
- AWS Instance Scheduler — cost optimisation for non-production resources
Why this matters
Zerodha needs to run end-of-day settlement calculations across millions of trades every evening. The job takes 4 hours, uses 100 vCPUs, and runs once a day. Lambda cannot run for more than 15 minutes. EC2 running 24/7 would waste 20 hours of compute per day. AWS Batch spins up the right compute when the job arrives, runs it to completion however long it takes, and terminates the compute when done. At a Mumbai bank, customer data must legally stay within India. AWS Outposts delivers AWS services — EC2, EBS, RDS — running on physical hardware in their own data center. The data never leaves the building but the bank uses AWS APIs and tools they already know.
What is AWS Batch
AWS Batch is a fully managed service for running batch computing jobs at any scale. You define the job. Batch figures out the compute, schedules it, runs it, and cleans up.
The problem Batch solves:
Lambda: Maximum 15 minutes per execution Maximum 10 GB RAM Cannot run jobs that take hours Always-on EC2: Available at any time But you pay 24/7 even when no jobs are running Wasteful for jobs that run once per day or once per week AWS Batch: No time limit — jobs run until they finish Compute spun up when a job arrives, terminated when it finishes Pay only for the compute time the job used Handles queuing, scheduling, retries, and parallelisationAWS Batch Key Components
Job:
A unit of work. Could be a shell script, a Docker container, an executable, or an ECS task. Each job has:
- Job Definition (what to run — the container image and resource requirements)
- Input data location (usually S3)
- Output data location (usually S3)
Job Queue:
Jobs are submitted to a queue. Multiple queues can have different priorities. High-priority queue gets compute before low-priority queue.
Production settlement queue → high priority → runs firstDevelopment testing queue → low priority → runs when capacity availableCompute Environment:
Where jobs actually run. You define the instance types, min/max vCPUs, and whether to use On-Demand or Spot.
Managed Compute Environment: AWS Batch manages EC2 instances automatically Scales up when jobs arrive, scales down when idle You define: min vCPUs, max vCPUs, instance types Unmanaged Compute Environment: You manage the compute yourself Batch just schedules jobs onto your instancesJob flow:
Submit job to queue ↓Batch evaluates queue priority ↓Batch scales up Compute Environment ↓Job runs on EC2 or Fargate ↓Job finishes → results written to S3 ↓Compute scales down (if no more jobs)AWS Batch on Fargate vs EC2
Fargate:
No EC2 instances to manage — fully serverlessFast startup — no instance warmup neededWorks best for: many small jobs, unpredictable job sizes, simple workloadsLimitation: maximum 4 vCPUs and 30 GB RAM per jobEC2:
Full control of instance typeNo resource limits — use any instance size including GPUWorks best for: large jobs needing many CPUs or GPU, ML training, genomicsSpot Instances work seamlessly for massive cost savingsRememberIf a batch job needs more than 4 vCPUs or 30 GB RAM, it must run on EC2 — Fargate cannot support it. For GPU workloads (ML model training, video processing) you must use EC2 with GPU instance types.
Spot Instances with Batch
Batch integrates natively with Spot Instances. Define a Compute Environment that uses Spot, set a max Spot price, and Batch uses Spot when available.
Compute Environment with Spot: Tries to acquire Spot Instances If Spot is interrupted → Batch automatically retries on a new Spot Instance If no Spot available → Batch waits or falls back to On-Demand Spot savings on a 4-hour batch job using 50 m5.xlarge instances:On-Demand: $0.192 × 50 × 4 = $38.40Spot (70% discount): ~$11.52Multi-instance types in one Compute Environment:
Set allowed instance types: [m5.large, m5.xlarge, m4.large, r5.large]Batch picks whichever Spot type is cheapest at job timeMore flexibility = higher chance of finding cheap Spot capacityAmazon Outposts — AWS in Your Data Center
Outposts brings AWS hardware, services, and APIs to your on-premises location. The physical servers are AWS-managed rack hardware delivered to your data center.
Normal AWS region: Your app → AWS services over internet or Direct Connect Data leaves your building and enters AWS region AWS Outposts: AWS delivers physical rack hardware to your building AWS manages, patches, and updates the hardware remotely Your app → AWS services running ON-PREMISES Data never leaves your buildingServices available on Outposts:
EC2, EBS, ECS, EKS, RDS, ElastiCache, S3 (Outposts S3)Use cases:
Data residency requirements: Banking regulations require customer financial data to stay in India Cannot use AWS Mumbai region — data must stay in the bank's building Outposts delivers AWS services inside the building Low-latency applications: Factory floor automation needs sub-millisecond response Manufacturing equipment → Outposts EC2 → immediate response Cloud region is too far for this latency requirement Local data processing: Process data locally on Outposts, send only results to AWS region Useful for large data volumes where transferring raw data is too slowConnectivity:
Outposts requires a reliable Direct Connect or VPN connection to the AWS parent region. Outposts communicates with AWS for management, APIs, and access to non-Outposts services.
AWS AppFlow — No-Code SaaS Integration
AppFlow is a fully managed integration service that moves data between AWS services and SaaS applications — no code required.
Sources: Destinations:Salesforce → Amazon S3ServiceNow → Amazon RedshiftZendesk → SalesforceGoogle Analytics → SnowflakeSlack → any destinationSAP →Set up:
AppFlow → Create flowSource: Salesforce → select object: ContactsDestination: S3 bucketTrigger: daily at 2 AMTransformation: filter, map fields, mask sensitive dataCreate flow → runs automaticallyWhat AppFlow handles automatically:
OAuth authentication to the SaaSData transfer with automatic batchingField mapping between source and destinationData transformation and filteringEncryption in transitNo Lambda glue code. No custom ETL. Just configuration.
AWS Amplify — Full-Stack App Deployment
Amplify is a set of tools for building and deploying full-stack web and mobile applications on AWS. Frontend code (React, Vue, Angular, Next.js) plus backend (APIs, Auth, Storage) deployed together.
Connect GitHub repository → AmplifyAmplify automatically: Builds the frontend on every push Deploys to CloudFront globally Manages custom domain and SSL certificate Creates preview deployments for every pull request Backend features Amplify configures with simple commands: Authentication → Cognito User Pools API → AppSync GraphQL or API Gateway REST Storage → S3 bucket Database → DynamoDBUse Amplify when:
Frontend team wants to deploy without managing AWS infrastructureBuilding a web or mobile app with standard auth, API, and storage needsNeed automatic preview environments per pull requestInstance Scheduler — Turn Off What You Do Not Need
Instance Scheduler is an AWS solution (deployed as a CloudFormation stack) that automatically starts and stops EC2 and RDS instances on a schedule.
Development instances: Start: Monday-Friday 8 AM IST Stop: Monday-Friday 10 PM IST Weekend: off entirely Savings calculation: Normal: 720 hours/month billed With scheduler: ~280 hours/month billed Savings: 61% reduction in instance costDeployment:
AWS Solutions Library → "Instance Scheduler on AWS" → Launch in CloudFormationTag your instances: Schedule=office-hoursScheduler reads the tags → starts and stops automaticallyNo code. No Lambda to maintain. One CloudFormation stack for your entire account.
Hands-on Lab — AWS Batch Job and Amplify Exploration
Step 1 — Create a Batch Compute Environment
AWS Batch → Compute environments → CreateCompute environment type: ManagedName: devops-batch-ceInstance configuration: FargateMinimum vCPUs: 0 Maximum vCPUs: 256VPC: your VPC Subnets: private subnetsCreate compute environmentStep 2 — Create a Job Queue
AWS Batch → Job queues → CreateName: devops-batch-queuePriority: 1Connected compute environments: devops-batch-ceCreate job queueStep 3 — Create a Job Definition
AWS Batch → Job definitions → CreateName: devops-batch-jobPlatform: FargateExecution role: ecsTaskExecutionRole (create if needed)Container image: public.ecr.aws/amazonlinux/amazonlinux:latestCommand:echo,Processing batch job startedsleep,10echo,Batch job completevCPUs: 0.25 Memory: 512 MBCreate job definitionStep 4 — Submit and monitor a job
AWS Batch → Jobs → Submit new jobJob name: devops-test-jobJob definition: devops-batch-jobJob queue: devops-batch-queueSubmit Watch the job progress through states:SUBMITTED → PENDING → RUNNABLE → STARTING → RUNNING → SUCCEEDED Click the job → view CloudWatch LogsSee the echo output from your containerStep 5 — Explore Amplify
AWS Amplify → New app → Host web appReview GitHub integration optionsSee how branch-based deployments workNo need to create — just review the workflowStep 6 — Cleanup
AWS Batch → Jobs → verify job completedAWS Batch → Job queues → devops-batch-queue → Disable → DeleteAWS Batch → Compute environments → devops-batch-ce → Disable → DeleteAWS Batch → Job definitions → devops-batch-job → DeregisterCommon Mistakes to Avoid
Common MistakeUsing Lambda for jobs that run longer than 15 minutes. Lambda has a hard maximum of 15 minutes per execution. If your batch job — report generation, data transformation, model training — takes longer, Lambda silently fails with a timeout error. AWS Batch has no time limit. Use Batch for any job that might take more than 15 minutes.
Common MistakeUsing Outposts without a reliable Direct Connect connection. Outposts requires network connectivity back to the AWS parent region for management plane operations. If the connection is unstable, Outposts operational management degrades. AWS recommends Direct Connect, not VPN, for production Outposts deployments.
TipCombine AWS Batch with Spot Instances and multi-instance type selection for maximum cost efficiency. Set 5-10 compatible instance types in your Compute Environment. Batch picks whichever Spot type is cheapest at job submission time. For embarrassingly parallel workloads (processing millions of files independently) this combination gives you massive compute at 70-90% less than On-Demand pricing.