Skip to main content

AWS Batch, Outposts, and Hybrid Compute - Running Jobs and Extending AWS

Run large-scale batch jobs with AWS Batch, extend AWS infrastructure to your data center with Outposts, and integrate SaaS data with AppFlow.

What you will learn

  • What AWS Batch is and the problem it solves that Lambda cannot
  • AWS Batch components — Job Definitions, Job Queues, and Compute Environments
  • AWS Batch on Fargate vs EC2 and when each is right
  • Spot Instances with Batch for 90% compute cost savings
  • Amazon Outposts — running AWS infrastructure on-premises
  • Outposts use cases — data residency, low latency to on-premises systems
  • AWS AppFlow — no-code SaaS data integration
  • AWS Amplify — full-stack serverless app deployment
  • AWS Instance Scheduler — cost optimisation for non-production resources

Why this matters

Zerodha needs to run end-of-day settlement calculations across millions of trades every evening. The job takes 4 hours, uses 100 vCPUs, and runs once a day. Lambda cannot run for more than 15 minutes. EC2 running 24/7 would waste 20 hours of compute per day. AWS Batch spins up the right compute when the job arrives, runs it to completion however long it takes, and terminates the compute when done. At a Mumbai bank, customer data must legally stay within India. AWS Outposts delivers AWS services — EC2, EBS, RDS — running on physical hardware in their own data center. The data never leaves the building but the bank uses AWS APIs and tools they already know.

What is AWS Batch

AWS Batch is a fully managed service for running batch computing jobs at any scale. You define the job. Batch figures out the compute, schedules it, runs it, and cleans up.

The problem Batch solves:

TEXT
Lambda:
Maximum 15 minutes per execution
Maximum 10 GB RAM
Cannot run jobs that take hours
Always-on EC2:
Available at any time
But you pay 24/7 even when no jobs are running
Wasteful for jobs that run once per day or once per week
AWS Batch:
No time limit — jobs run until they finish
Compute spun up when a job arrives, terminated when it finishes
Pay only for the compute time the job used
Handles queuing, scheduling, retries, and parallelisation

AWS Batch Key Components

Job:

A unit of work. Could be a shell script, a Docker container, an executable, or an ECS task. Each job has:

  • Job Definition (what to run — the container image and resource requirements)
  • Input data location (usually S3)
  • Output data location (usually S3)

Job Queue:

Jobs are submitted to a queue. Multiple queues can have different priorities. High-priority queue gets compute before low-priority queue.

◈ DIAGRAM
Production settlement queue → high priority → runs first
Development testing queue → low priority → runs when capacity available

Compute Environment:

Where jobs actually run. You define the instance types, min/max vCPUs, and whether to use On-Demand or Spot.

TEXT
Managed Compute Environment:
AWS Batch manages EC2 instances automatically
Scales up when jobs arrive, scales down when idle
You define: min vCPUs, max vCPUs, instance types
Unmanaged Compute Environment:
You manage the compute yourself
Batch just schedules jobs onto your instances

Job flow:

◈ DIAGRAM
Submit job to queue
↓
Batch evaluates queue priority
↓
Batch scales up Compute Environment
↓
Job runs on EC2 or Fargate
↓
Job finishes → results written to S3
↓
Compute scales down (if no more jobs)

AWS Batch on Fargate vs EC2

Fargate:

TEXT
No EC2 instances to manage — fully serverless
Fast startup — no instance warmup needed
Works best for: many small jobs, unpredictable job sizes, simple workloads
Limitation: maximum 4 vCPUs and 30 GB RAM per job

EC2:

TEXT
Full control of instance type
No resource limits — use any instance size including GPU
Works best for: large jobs needing many CPUs or GPU, ML training, genomics
Spot Instances work seamlessly for massive cost savings
Remember

If a batch job needs more than 4 vCPUs or 30 GB RAM, it must run on EC2 — Fargate cannot support it. For GPU workloads (ML model training, video processing) you must use EC2 with GPU instance types.

Spot Instances with Batch

Batch integrates natively with Spot Instances. Define a Compute Environment that uses Spot, set a max Spot price, and Batch uses Spot when available.

◈ DIAGRAM
Compute Environment with Spot:
Tries to acquire Spot Instances
If Spot is interrupted → Batch automatically retries on a new Spot Instance
If no Spot available → Batch waits or falls back to On-Demand
Spot savings on a 4-hour batch job using 50 m5.xlarge instances:
On-Demand: $0.192 × 50 × 4 = $38.40
Spot (70% discount): ~$11.52

Multi-instance types in one Compute Environment:

TEXT
Set allowed instance types: [m5.large, m5.xlarge, m4.large, r5.large]
Batch picks whichever Spot type is cheapest at job time
More flexibility = higher chance of finding cheap Spot capacity

Amazon Outposts — AWS in Your Data Center

Outposts brings AWS hardware, services, and APIs to your on-premises location. The physical servers are AWS-managed rack hardware delivered to your data center.

◈ DIAGRAM
Normal AWS region:
Your app → AWS services over internet or Direct Connect
Data leaves your building and enters AWS region
AWS Outposts:
AWS delivers physical rack hardware to your building
AWS manages, patches, and updates the hardware remotely
Your app → AWS services running ON-PREMISES
Data never leaves your building

Services available on Outposts:

TEXT
EC2, EBS, ECS, EKS, RDS, ElastiCache, S3 (Outposts S3)

Use cases:

◈ DIAGRAM
Data residency requirements:
Banking regulations require customer financial data to stay in India
Cannot use AWS Mumbai region — data must stay in the bank's building
Outposts delivers AWS services inside the building
Low-latency applications:
Factory floor automation needs sub-millisecond response
Manufacturing equipment → Outposts EC2 → immediate response
Cloud region is too far for this latency requirement
Local data processing:
Process data locally on Outposts, send only results to AWS region
Useful for large data volumes where transferring raw data is too slow

Connectivity:

Outposts requires a reliable Direct Connect or VPN connection to the AWS parent region. Outposts communicates with AWS for management, APIs, and access to non-Outposts services.

AWS AppFlow — No-Code SaaS Integration

AppFlow is a fully managed integration service that moves data between AWS services and SaaS applications — no code required.

◈ DIAGRAM
Sources: Destinations:
Salesforce → Amazon S3
ServiceNow → Amazon Redshift
Zendesk → Salesforce
Google Analytics → Snowflake
Slack → any destination
SAP →

Set up:

◈ DIAGRAM
AppFlow → Create flow
Source: Salesforce → select object: Contacts
Destination: S3 bucket
Trigger: daily at 2 AM
Transformation: filter, map fields, mask sensitive data
Create flow → runs automatically

What AppFlow handles automatically:

TEXT
OAuth authentication to the SaaS
Data transfer with automatic batching
Field mapping between source and destination
Data transformation and filtering
Encryption in transit

No Lambda glue code. No custom ETL. Just configuration.

AWS Amplify — Full-Stack App Deployment

Amplify is a set of tools for building and deploying full-stack web and mobile applications on AWS. Frontend code (React, Vue, Angular, Next.js) plus backend (APIs, Auth, Storage) deployed together.

◈ DIAGRAM
Connect GitHub repository → Amplify
Amplify automatically:
Builds the frontend on every push
Deploys to CloudFront globally
Manages custom domain and SSL certificate
Creates preview deployments for every pull request
Backend features Amplify configures with simple commands:
Authentication → Cognito User Pools
API → AppSync GraphQL or API Gateway REST
Storage → S3 bucket
Database → DynamoDB

Use Amplify when:

TEXT
Frontend team wants to deploy without managing AWS infrastructure
Building a web or mobile app with standard auth, API, and storage needs
Need automatic preview environments per pull request

Instance Scheduler — Turn Off What You Do Not Need

Instance Scheduler is an AWS solution (deployed as a CloudFormation stack) that automatically starts and stops EC2 and RDS instances on a schedule.

TEXT
Development instances:
Start: Monday-Friday 8 AM IST
Stop: Monday-Friday 10 PM IST
Weekend: off entirely
Savings calculation:
Normal: 720 hours/month billed
With scheduler: ~280 hours/month billed
Savings: 61% reduction in instance cost

Deployment:

◈ DIAGRAM
AWS Solutions Library → "Instance Scheduler on AWS" → Launch in CloudFormation
Tag your instances: Schedule=office-hours
Scheduler reads the tags → starts and stops automatically

No code. No Lambda to maintain. One CloudFormation stack for your entire account.

Hands-on Lab — AWS Batch Job and Amplify Exploration

Step 1 — Create a Batch Compute Environment

◈ DIAGRAM
AWS Batch → Compute environments → Create
Compute environment type: Managed
Name: devops-batch-ce
Instance configuration: Fargate
Minimum vCPUs: 0 Maximum vCPUs: 256
VPC: your VPC Subnets: private subnets
Create compute environment

Step 2 — Create a Job Queue

◈ DIAGRAM
AWS Batch → Job queues → Create
Name: devops-batch-queue
Priority: 1
Connected compute environments: devops-batch-ce
Create job queue

Step 3 — Create a Job Definition

Bash
AWS Batch → Job definitions → Create
Name: devops-batch-job
Platform: Fargate
Execution role: ecsTaskExecutionRole (create if needed)
Container image: public.ecr.aws/amazonlinux/amazonlinux:latest
Command:
Bash
echo,Processing batch job started
sleep,10
echo,Batch job complete
TEXT
vCPUs: 0.25 Memory: 512 MB
Create job definition

Step 4 — Submit and monitor a job

Bash
AWS Batch → Jobs → Submit new job
Job name: devops-test-job
Job definition: devops-batch-job
Job queue: devops-batch-queue
Submit
Watch the job progress through states:
SUBMITTED → PENDING → RUNNABLE → STARTING → RUNNING → SUCCEEDED
Click the job → view CloudWatch Logs
See the echo output from your container

Step 5 — Explore Amplify

◈ DIAGRAM
AWS Amplify → New app → Host web app
Review GitHub integration options
See how branch-based deployments work
No need to create — just review the workflow

Step 6 — Cleanup

◈ DIAGRAM
AWS Batch → Jobs → verify job completed
AWS Batch → Job queues → devops-batch-queue → Disable → Delete
AWS Batch → Compute environments → devops-batch-ce → Disable → Delete
AWS Batch → Job definitions → devops-batch-job → Deregister

Common Mistakes to Avoid

Common Mistake

Using Lambda for jobs that run longer than 15 minutes. Lambda has a hard maximum of 15 minutes per execution. If your batch job — report generation, data transformation, model training — takes longer, Lambda silently fails with a timeout error. AWS Batch has no time limit. Use Batch for any job that might take more than 15 minutes.

Common Mistake

Using Outposts without a reliable Direct Connect connection. Outposts requires network connectivity back to the AWS parent region for management plane operations. If the connection is unstable, Outposts operational management degrades. AWS recommends Direct Connect, not VPN, for production Outposts deployments.

Tip

Combine AWS Batch with Spot Instances and multi-instance type selection for maximum cost efficiency. Set 5-10 compatible instance types in your Compute Environment. Batch picks whichever Spot type is cheapest at job submission time. For embarrassingly parallel workloads (processing millions of files independently) this combination gives you massive compute at 70-90% less than On-Demand pricing.

Resources

AWS Direct Connect vs Site-to-Site VPN Failover

AWS Direct Connect vs Site-to-Site VPN Failover

Direct Connect vs VPN isn't really either/or for production — it's a primary-plus-failover pattern. Here's how to design it, and when either/or is right.

5 min read•Aug 2026
Lambda vs Fargate vs EC2 Spot: The Cost Crossover

Lambda vs Fargate vs EC2 Spot: The Cost Crossover

Lambda vs Fargate vs EC2 Spot, at the crossover where Lambda stops being cheaper — 2026 pricing, invocation thresholds, and interruption math.

5 min read•Aug 2026
Secrets Manager vs Parameter Store vs Vault

Secrets Manager vs Parameter Store vs Vault

AWS Secrets Manager, Parameter Store, and HashiCorp Vault compared for 2026 - cost math, rotation, multi-cloud fit, and the Vault-to-OpenBao fork.

5 min read•Aug 2026
AWS VPC Security: Hardening Every Layer

AWS VPC Security: Hardening Every Layer

Most cloud security incidents start with a misconfigured VPC. Here's how to harden every layer — subnets, Security Groups, NACLs, and IAM — for production.

5 min read•Jul 2026
Event-Driven Architecture on AWS Explained

Event-Driven Architecture on AWS Explained

Event-driven architecture on AWS decouples services and absorbs traffic spikes using SQS, SNS, EventBridge, and Lambda — workflows that scale themselves.

5 min read•Jul 2026
S3 vs RDS vs DynamoDB: Choosing AWS Storage

S3 vs RDS vs DynamoDB: Choosing AWS Storage

Choosing S3, RDS, or DynamoDB wrong costs you in performance, cost, and scalability. Here is a practical decision guide based on your actual access patterns.

5 min read•Jul 2026
AWS Cost Optimisation: Cut Cloud Bills 40-60%

AWS Cost Optimisation: Cut Cloud Bills 40-60%

AWS bills surprise teams every month. Here are the 8 concrete actions that cut cloud spend by 40-60% without touching your application architecture.

5 min read•Jul 2026
EC2 vs Lambda vs Fargate: Choosing AWS Compute

EC2 vs Lambda vs Fargate: Choosing AWS Compute

EC2, Lambda, or Fargate — choosing the wrong AWS compute option costs you money and performance. Here is exactly when to use each one in production.

5 min read•Jul 2026

Explore More in AWS DevOps, Cost, and Machine Learning

All 6 Topics

Frequently Asked Questions

Is AWS Batch, Outposts, and Hybrid Compute - Running Jobs and Extending AWS free to learn on DevOps Network?

Yes - this topic, like everything on DevOps Network, is 100% free with no paywall or sign-up gate.

What does the AWS Batch, Outposts, and Hybrid Compute - Running Jobs and Extending AWS topic cover?

Run large-scale batch jobs with AWS Batch, extend AWS infrastructure to your data center with Outposts, and integrate SaaS data with AppFlow.