Skip to main content

AWS Machine Learning Services - AI Without Building Models

Add AI capabilities to applications using AWS pre-built ML services — Rekognition, Transcribe, Polly, Translate, Comprehend, Lex, Kendra, Textract, and SageMaker.

What you will learn

  • The two categories of AWS ML — pre-built services vs custom model training
  • Amazon Rekognition — image and video analysis without any ML knowledge
  • Amazon Transcribe — speech to text with PII redaction
  • Amazon Polly — text to speech in multiple languages and voices
  • Amazon Translate — real-time text translation
  • Amazon Comprehend — natural language understanding and sentiment analysis
  • Amazon Lex — building conversational chatbots (the engine behind Alexa)
  • Amazon Kendra — intelligent document search powered by ML
  • Amazon Textract — extracting text and data from scanned documents
  • Amazon Personalize — real-time personalised recommendations
  • Amazon SageMaker — when you need to train your own custom models

Why this matters

A Hotstar content moderation team used to manually review flagged videos — slow, expensive, and inconsistent. Rekognition now scans every uploaded video automatically, flags inappropriate content in seconds, and routes only genuinely borderline cases to human reviewers. At Razorpay, customer support used to search through PDFs manually for transaction details. Textract now extracts data from uploaded bank statements and Kendra makes that data searchable instantly. At Swiggy, the recommendation engine that suggests restaurants and dishes is powered by Amazon Personalize — trained on order history, updated in real time, without a single ML engineer maintaining it. These are not experimental projects. They are production systems built on AWS ML services that any developer can integrate in a day.

Two Categories of AWS ML

AWS ML services split into two clean categories:

TEXT
Pre-built services (AI Services):
Ready to use via API call
No ML knowledge needed
No model training
Pay per API call
Use when: the capability fits your use case out of the box
Custom model training (SageMaker):
Build and train your own model on your own data
Requires ML knowledge
More expensive and complex
Use when: pre-built services do not fit your specific problem

Start with pre-built services. Only reach for SageMaker when the pre-built services cannot solve your problem.

Amazon Rekognition — See What Is in Images and Videos

Rekognition analyses images and videos and tells you what it sees — objects, faces, text, activities, and whether content is appropriate.

What it can detect:

TEXT
Objects and scenes:
"This image contains: car, road, traffic light, pedestrian"
Confidence scores for each detection
Faces:
Detect faces in an image (how many, where)
Compare two faces: "are these the same person?" (similarity score)
Analyse attributes: approximate age, gender, emotions, glasses, smile
Search a face against a collection of known faces
Text in images:
"The sign in this photo reads: Exit 5, Mumbai Highway"
Useful for reading license plates, signs, printed forms
Content moderation:
"This image contains explicit content" with confidence score
Categories: Explicit nudity, violence, hate symbols, visually disturbing
Route flagged content to human reviewers automatically
Video analysis:
Same capabilities applied to video frames
Track people across frames
Detect when a specific activity occurs

Real use case at Hotstar:

◈ DIAGRAM
User uploads video
↓
Lambda triggers Rekognition video analysis
↓
Rekognition returns: content moderation labels with timestamps
↓
Confidence > 90%: auto-reject
Confidence 50-90%: send to human reviewer queue
Confidence < 50%: approve automatically

Amazon Transcribe — Speech to Text

Transcribe converts audio to text. Give it an audio or video file and it returns a text transcript with timestamps, speaker identification, and punctuation.

Key features:

TEXT
Automatic punctuation and formatting
Speaker identification (who said what — up to 10 speakers)
Custom vocabulary (teach it domain-specific words like "PhonePe", "UPI", "NEFT")
Multiple language support including Hindi
Medical transcription with clinical terminology

PII Redaction — important for compliance:

Transcribe can automatically detect and redact personally identifiable information from transcripts.

TEXT
Audio: "My name is Rahul Sharma, PAN number ABCDE1234F, mobile 9876543210"
Transcript without redaction: exact text above
Transcript with PII redaction: "My name is [NAME], PAN number [PAN], mobile [PHONE]"

Critical for call centre recordings, customer support audio, and any audio containing personal data.

Amazon Polly — Text to Speech

Polly converts text into lifelike speech. Give it text. Get back an audio file in dozens of voices, languages, and accents.

TEXT
Input: "Your order ORD-2024-001 from Swiggy has been delivered"
Output: MP3 audio file in the voice and language you chose

Voices:

TEXT
Standard voices: text-to-speech synthesis (good quality)
Neural voices: deep learning based (much more natural, human-like)
Both available in English (Indian), Hindi, and many other languages

SSML — Speech Synthesis Markup Language:

Control exactly how Polly speaks — pauses, emphasis, speaking rate, pitch.

TEXT
"Your balance is <emphasis>zero</emphasis>.
<break time="500ms"/>
Please add funds <prosody rate="slow">immediately</prosody>."

Use cases:

  • Voice responses in IVR systems
  • Accessibility features (read page content aloud)
  • Audio content from text articles
  • Language learning applications

Amazon Translate — Real-Time Translation

Translate converts text from one language to another. Supports 75+ languages.

TEXT
Input: "Your order has been dispatched" (English)
Output: "आपका ऑर्डर भेज दिया गया है" (Hindi)

Auto-detect source language: Translate identifies the source language automatically — you do not need to tell it what the input language is.

Use cases:

  • Translate user-generated content in real time (reviews, comments, messages)
  • Multilingual customer support (translate incoming messages, outgoing responses)
  • Localise application content for different regions

Amazon Comprehend — Understand Text

Comprehend reads text and extracts meaning from it — sentiment, key phrases, entities, language, and topics.

What Comprehend can do:

◈ DIAGRAM
Sentiment Analysis:
"Worst delivery experience ever, food was cold" → NEGATIVE (confidence: 97%)
"Amazing food, delivered in 20 minutes!" → POSITIVE (confidence: 99%)
Entity Recognition:
"Rahul Sharma from Mumbai placed an order on 15th January"
Entities: Rahul Sharma (PERSON), Mumbai (LOCATION), 15th January (DATE)
Key Phrase Extraction:
"The new restaurant on MG Road serves excellent biryani"
Key phrases: "new restaurant", "MG Road", "excellent biryani"
Language Detection:
Input: "Bonjour, comment allez-vous" → French (99.9% confidence)
Topic Modeling:
Analyse 10,000 customer support tickets
Discover: 40% about delivery delays, 25% about payment, 20% about food quality

Amazon Comprehend Medical:

A specialised version that understands medical text — clinical notes, prescriptions, discharge summaries. Extracts: medications, dosages, diagnoses, procedures, and Protected Health Information (PHI).

Amazon Lex — Conversational Chatbots

Lex is the same technology that powers Amazon Alexa. It understands natural language and maintains conversation context to handle multi-turn dialogues.

◈ DIAGRAM
User: "I want to order food"
Lex: "Sure! What restaurant?"
User: "Pizza from Domino's"
Lex: "Got it. What size?"
User: "Large"
Lex: "Your order: Large pizza from Domino's. Confirm?"
User: "Yes"
Lex: → triggers Lambda → places order → "Order placed!"

Key concepts:

TEXT
Intent: what the user wants to do (OrderFood, CheckBalance, CancelOrder)
Slot: information Lex needs to fulfil the intent (restaurant, size, quantity)
Fulfilment: Lambda function that runs when Lex has all the slots filled

Lex + Connect:

Amazon Connect is AWS's contact centre service. Lex powers the automated voice responses — the IVR bot that handles simple requests automatically before routing complex ones to human agents.

◈ DIAGRAM
Customer calls → Connect answers → Lex bot handles the conversation
"Check account balance" → Lex extracts intent → Lambda queries DB → speaks the balance
"Speak to agent" → Lex recognises → Connect routes to human agent

Kendra indexes your documents — PDFs, Word files, FAQs, web pages, databases — and makes them searchable with natural language questions.

TEXT
Without Kendra:
User searches: "what is the refund policy for damaged items"
Keyword search finds: every document containing those words
User must read through results to find the answer
With Kendra:
User asks: "what is the refund policy for damaged items"
Kendra reads all documents, understands the question
Returns: the specific answer extracted from the correct document

Connectors for: S3, SharePoint, Salesforce, Confluence, ServiceNow, RDS, OneDrive. Index once. Search everything.

Use case at Razorpay: legal and compliance team searches thousands of contracts, regulations, and internal policies. Kendra finds the exact clause they need in seconds.

Amazon Textract — Extract Data from Documents

Textract goes beyond simple OCR. It does not just read text — it understands the structure of documents and extracts data in a meaningful way.

◈ DIAGRAM
Scanned bank statement → Textract extracts:
Tables with values in correct columns
Form fields with key-value pairs: "Account Number: 123456789"
Signatures detected as signatures
Checkboxes identified as checked or unchecked

What Textract can extract:

TEXT
Text (basic OCR)
Tables (preserving row and column structure)
Forms (key-value pairs from structured forms)
Queries (ask specific questions: "what is the invoice total?")
Signatures

Use case: Razorpay KYC process. Users upload scanned Aadhaar card. Textract extracts name, date of birth, address, and Aadhaar number automatically. No manual data entry.

Amazon Personalize — Real-Time Recommendations

Personalize creates recommendation models from your user interaction data and returns real-time personalised recommendations via API call.

◈ DIAGRAM
Feed historical data:
User rahul-001 ordered biryani 5 times, burger twice, pizza once
User priya-002 ordered salads 8 times, wraps 4 times
Personalize trains a model on this data
↓
At runtime:
GET /recommendations?userId=rahul-001
Returns: ["Biryani Palace", "Hyderabadi Biryani House", "Mughal Kitchen"]
GET /recommendations?userId=priya-002
Returns: ["Green Bowl", "Salad Bar", "Freshii"]

No ML knowledge required. You provide the data. Personalize handles the model training, hosting, and serving.

Use cases: restaurant recommendations (Swiggy), product recommendations (e-commerce), content recommendations (Hotstar), "customers also bought" (retail).

Amazon SageMaker — Train Your Own Models

When pre-built services cannot solve your problem — your data is unique, your domain is specialised, accuracy requirements are very high — SageMaker lets you build and train custom ML models.

◈ DIAGRAM
Pre-built service cannot solve it:
Content moderation in a very specific domain
Fraud detection on your specific transaction patterns
Demand forecasting on your specific inventory data
SageMaker workflow:
Collect and prepare your training data in S3
Choose an algorithm or bring your own model code
Train on managed infrastructure (GPU instances spun up, used, and terminated)
Evaluate model accuracy
Deploy to an endpoint (real-time inference)
Your application calls the endpoint: send data → get prediction

SageMaker is significantly more complex and expensive than the pre-built services. It requires ML expertise. Use it only when pre-built services fall short.

Hands-on Lab — Rekognition and Comprehend in the Console

Step 1 — Test Rekognition image analysis

◈ DIAGRAM
AWS Console → Amazon Rekognition → Label detection (left menu)
Try with a sample image or upload your own
Click: Use your own image → upload a photo with identifiable objects
Results show: detected labels with confidence scores
Example: Person (99.8%), Bicycle (97.2%), Road (94.1%)

Step 2 — Test content moderation

◈ DIAGRAM
Rekognition → Content moderation
Upload an image (use a stock photo of something clearly safe)
Results: content moderation labels or "No unsafe content detected"

Step 3 — Test Comprehend sentiment analysis

◈ DIAGRAM
Amazon Comprehend → Real-time analysis → Input text
Paste a positive review:
"Swiggy delivered my food in 15 minutes, everything was hot and fresh. Amazing service!"
Sentiment: POSITIVE
Paste a negative review:
"Waited 2 hours for my order. Food was cold. Never ordering again."
Sentiment: NEGATIVE

Step 4 — Test entity recognition

◈ DIAGRAM
Comprehend → Real-time analysis → Entity recognition
Paste:
"Rahul Sharma from Bengaluru placed an order worth Rs 450 on 15th January 2024 using PhonePe."
Results:
Rahul Sharma → PERSON
Bengaluru → LOCATION
Rs 450 → QUANTITY
15th January 2024 → DATE
PhonePe → ORGANIZATION

Step 5 — Test Translate

◈ DIAGRAM
Amazon Translate → Real-time translation
Source language: Auto
Target language: Hindi
Input: "Your order has been delivered. Thank you for using our service."
Output: "आपका ऑर्डर डिलीवर कर दिया गया है। हमारी सेवा का उपयोग करने के लिए धन्यवाद।"

Step 6 — Test Textract

◈ DIAGRAM
Amazon Textract → Analyze document
Upload a scanned document or use the sample provided
Select: Tables and Forms
Analyze
See how Textract identifies table structure, form fields, and key-value pairs
Compare to what a simple image-to-text tool would produce — Textract understands structure

No cleanup needed — these are all console demos with no persistent resources created.

Common Mistakes to Avoid

Common Mistake

Jumping to SageMaker before trying pre-built services. SageMaker requires ML expertise, data preparation, model training time, and ongoing maintenance. Rekognition, Comprehend, Textract, and Personalize solve most common AI use cases via a simple API call. Try the pre-built service first. Use SageMaker only when the pre-built service genuinely cannot solve your problem.

Common Mistake

Using Rekognition facial recognition without understanding the compliance requirements. In some countries and jurisdictions, storing and comparing facial data requires explicit user consent and has strict regulatory requirements. Check your local laws before building systems that identify individuals by face.

Tip

All the pre-built AI services (Rekognition, Transcribe, Polly, Translate, Comprehend, Lex, Kendra, Textract, Personalize) can be combined. A customer support call → Transcribe converts audio to text → Comprehend extracts sentiment and key phrases → Translate converts to English for a global team → stored and searched by Kendra. Each service does one job and they chain together naturally.

Resources

AWS Direct Connect vs Site-to-Site VPN Failover

AWS Direct Connect vs Site-to-Site VPN Failover

Direct Connect vs VPN isn't really either/or for production — it's a primary-plus-failover pattern. Here's how to design it, and when either/or is right.

5 min read•Aug 2026
Lambda vs Fargate vs EC2 Spot: The Cost Crossover

Lambda vs Fargate vs EC2 Spot: The Cost Crossover

Lambda vs Fargate vs EC2 Spot, at the crossover where Lambda stops being cheaper — 2026 pricing, invocation thresholds, and interruption math.

5 min read•Aug 2026
Secrets Manager vs Parameter Store vs Vault

Secrets Manager vs Parameter Store vs Vault

AWS Secrets Manager, Parameter Store, and HashiCorp Vault compared for 2026 - cost math, rotation, multi-cloud fit, and the Vault-to-OpenBao fork.

5 min read•Aug 2026
AWS VPC Security: Hardening Every Layer

AWS VPC Security: Hardening Every Layer

Most cloud security incidents start with a misconfigured VPC. Here's how to harden every layer — subnets, Security Groups, NACLs, and IAM — for production.

5 min read•Jul 2026
Event-Driven Architecture on AWS Explained

Event-Driven Architecture on AWS Explained

Event-driven architecture on AWS decouples services and absorbs traffic spikes using SQS, SNS, EventBridge, and Lambda — workflows that scale themselves.

5 min read•Jul 2026
S3 vs RDS vs DynamoDB: Choosing AWS Storage

S3 vs RDS vs DynamoDB: Choosing AWS Storage

Choosing S3, RDS, or DynamoDB wrong costs you in performance, cost, and scalability. Here is a practical decision guide based on your actual access patterns.

5 min read•Jul 2026
AWS Cost Optimisation: Cut Cloud Bills 40-60%

AWS Cost Optimisation: Cut Cloud Bills 40-60%

AWS bills surprise teams every month. Here are the 8 concrete actions that cut cloud spend by 40-60% without touching your application architecture.

5 min read•Jul 2026
EC2 vs Lambda vs Fargate: Choosing AWS Compute

EC2 vs Lambda vs Fargate: Choosing AWS Compute

EC2, Lambda, or Fargate — choosing the wrong AWS compute option costs you money and performance. Here is exactly when to use each one in production.

5 min read•Jul 2026

Explore More in AWS DevOps, Cost, and Machine Learning

All 6 Topics

Frequently Asked Questions

Is AWS Machine Learning Services - AI Without Building Models free to learn on DevOps Network?

Yes - this topic, like everything on DevOps Network, is 100% free with no paywall or sign-up gate.

What does the AWS Machine Learning Services - AI Without Building Models topic cover?

Add AI capabilities to applications using AWS pre-built ML services — Rekognition, Transcribe, Polly, Translate, Comprehend, Lex, Kendra, Textract, and SageMaker.