What you will learn
- The two categories of AWS ML — pre-built services vs custom model training
- Amazon Rekognition — image and video analysis without any ML knowledge
- Amazon Transcribe — speech to text with PII redaction
- Amazon Polly — text to speech in multiple languages and voices
- Amazon Translate — real-time text translation
- Amazon Comprehend — natural language understanding and sentiment analysis
- Amazon Lex — building conversational chatbots (the engine behind Alexa)
- Amazon Kendra — intelligent document search powered by ML
- Amazon Textract — extracting text and data from scanned documents
- Amazon Personalize — real-time personalised recommendations
- Amazon SageMaker — when you need to train your own custom models
Why this matters
A Hotstar content moderation team used to manually review flagged videos — slow, expensive, and inconsistent. Rekognition now scans every uploaded video automatically, flags inappropriate content in seconds, and routes only genuinely borderline cases to human reviewers. At Razorpay, customer support used to search through PDFs manually for transaction details. Textract now extracts data from uploaded bank statements and Kendra makes that data searchable instantly. At Swiggy, the recommendation engine that suggests restaurants and dishes is powered by Amazon Personalize — trained on order history, updated in real time, without a single ML engineer maintaining it. These are not experimental projects. They are production systems built on AWS ML services that any developer can integrate in a day.
Two Categories of AWS ML
AWS ML services split into two clean categories:
Pre-built services (AI Services): Ready to use via API call No ML knowledge needed No model training Pay per API call Use when: the capability fits your use case out of the box Custom model training (SageMaker): Build and train your own model on your own data Requires ML knowledge More expensive and complex Use when: pre-built services do not fit your specific problemStart with pre-built services. Only reach for SageMaker when the pre-built services cannot solve your problem.
Amazon Rekognition — See What Is in Images and Videos
Rekognition analyses images and videos and tells you what it sees — objects, faces, text, activities, and whether content is appropriate.
What it can detect:
Objects and scenes: "This image contains: car, road, traffic light, pedestrian" Confidence scores for each detection Faces: Detect faces in an image (how many, where) Compare two faces: "are these the same person?" (similarity score) Analyse attributes: approximate age, gender, emotions, glasses, smile Search a face against a collection of known faces Text in images: "The sign in this photo reads: Exit 5, Mumbai Highway" Useful for reading license plates, signs, printed forms Content moderation: "This image contains explicit content" with confidence score Categories: Explicit nudity, violence, hate symbols, visually disturbing Route flagged content to human reviewers automatically Video analysis: Same capabilities applied to video frames Track people across frames Detect when a specific activity occursReal use case at Hotstar:
User uploads video ↓Lambda triggers Rekognition video analysis ↓Rekognition returns: content moderation labels with timestamps ↓Confidence > 90%: auto-rejectConfidence 50-90%: send to human reviewer queueConfidence < 50%: approve automaticallyAmazon Transcribe — Speech to Text
Transcribe converts audio to text. Give it an audio or video file and it returns a text transcript with timestamps, speaker identification, and punctuation.
Key features:
Automatic punctuation and formattingSpeaker identification (who said what — up to 10 speakers)Custom vocabulary (teach it domain-specific words like "PhonePe", "UPI", "NEFT")Multiple language support including HindiMedical transcription with clinical terminologyPII Redaction — important for compliance:
Transcribe can automatically detect and redact personally identifiable information from transcripts.
Audio: "My name is Rahul Sharma, PAN number ABCDE1234F, mobile 9876543210"Transcript without redaction: exact text aboveTranscript with PII redaction: "My name is [NAME], PAN number [PAN], mobile [PHONE]"Critical for call centre recordings, customer support audio, and any audio containing personal data.
Amazon Polly — Text to Speech
Polly converts text into lifelike speech. Give it text. Get back an audio file in dozens of voices, languages, and accents.
Input: "Your order ORD-2024-001 from Swiggy has been delivered"Output: MP3 audio file in the voice and language you choseVoices:
Standard voices: text-to-speech synthesis (good quality)Neural voices: deep learning based (much more natural, human-like)Both available in English (Indian), Hindi, and many other languagesSSML — Speech Synthesis Markup Language:
Control exactly how Polly speaks — pauses, emphasis, speaking rate, pitch.
"Your balance is <emphasis>zero</emphasis>. <break time="500ms"/> Please add funds <prosody rate="slow">immediately</prosody>."Use cases:
- Voice responses in IVR systems
- Accessibility features (read page content aloud)
- Audio content from text articles
- Language learning applications
Amazon Translate — Real-Time Translation
Translate converts text from one language to another. Supports 75+ languages.
Input: "Your order has been dispatched" (English)Output: "आपका ऑर्डर भेज दिया गया है" (Hindi)Auto-detect source language: Translate identifies the source language automatically — you do not need to tell it what the input language is.
Use cases:
- Translate user-generated content in real time (reviews, comments, messages)
- Multilingual customer support (translate incoming messages, outgoing responses)
- Localise application content for different regions
Amazon Comprehend — Understand Text
Comprehend reads text and extracts meaning from it — sentiment, key phrases, entities, language, and topics.
What Comprehend can do:
Sentiment Analysis: "Worst delivery experience ever, food was cold" → NEGATIVE (confidence: 97%) "Amazing food, delivered in 20 minutes!" → POSITIVE (confidence: 99%) Entity Recognition: "Rahul Sharma from Mumbai placed an order on 15th January" Entities: Rahul Sharma (PERSON), Mumbai (LOCATION), 15th January (DATE) Key Phrase Extraction: "The new restaurant on MG Road serves excellent biryani" Key phrases: "new restaurant", "MG Road", "excellent biryani" Language Detection: Input: "Bonjour, comment allez-vous" → French (99.9% confidence) Topic Modeling: Analyse 10,000 customer support tickets Discover: 40% about delivery delays, 25% about payment, 20% about food qualityAmazon Comprehend Medical:
A specialised version that understands medical text — clinical notes, prescriptions, discharge summaries. Extracts: medications, dosages, diagnoses, procedures, and Protected Health Information (PHI).
Amazon Lex — Conversational Chatbots
Lex is the same technology that powers Amazon Alexa. It understands natural language and maintains conversation context to handle multi-turn dialogues.
User: "I want to order food"Lex: "Sure! What restaurant?"User: "Pizza from Domino's"Lex: "Got it. What size?"User: "Large"Lex: "Your order: Large pizza from Domino's. Confirm?"User: "Yes"Lex: → triggers Lambda → places order → "Order placed!"Key concepts:
Intent: what the user wants to do (OrderFood, CheckBalance, CancelOrder)Slot: information Lex needs to fulfil the intent (restaurant, size, quantity)Fulfilment: Lambda function that runs when Lex has all the slots filledLex + Connect:
Amazon Connect is AWS's contact centre service. Lex powers the automated voice responses — the IVR bot that handles simple requests automatically before routing complex ones to human agents.
Customer calls → Connect answers → Lex bot handles the conversation"Check account balance" → Lex extracts intent → Lambda queries DB → speaks the balance"Speak to agent" → Lex recognises → Connect routes to human agentAmazon Kendra — Intelligent Document Search
Kendra indexes your documents — PDFs, Word files, FAQs, web pages, databases — and makes them searchable with natural language questions.
Without Kendra: User searches: "what is the refund policy for damaged items" Keyword search finds: every document containing those words User must read through results to find the answer With Kendra: User asks: "what is the refund policy for damaged items" Kendra reads all documents, understands the question Returns: the specific answer extracted from the correct documentConnectors for: S3, SharePoint, Salesforce, Confluence, ServiceNow, RDS, OneDrive. Index once. Search everything.
Use case at Razorpay: legal and compliance team searches thousands of contracts, regulations, and internal policies. Kendra finds the exact clause they need in seconds.
Amazon Textract — Extract Data from Documents
Textract goes beyond simple OCR. It does not just read text — it understands the structure of documents and extracts data in a meaningful way.
Scanned bank statement → Textract extracts: Tables with values in correct columns Form fields with key-value pairs: "Account Number: 123456789" Signatures detected as signatures Checkboxes identified as checked or uncheckedWhat Textract can extract:
Text (basic OCR)Tables (preserving row and column structure)Forms (key-value pairs from structured forms)Queries (ask specific questions: "what is the invoice total?")SignaturesUse case: Razorpay KYC process. Users upload scanned Aadhaar card. Textract extracts name, date of birth, address, and Aadhaar number automatically. No manual data entry.
Amazon Personalize — Real-Time Recommendations
Personalize creates recommendation models from your user interaction data and returns real-time personalised recommendations via API call.
Feed historical data: User rahul-001 ordered biryani 5 times, burger twice, pizza once User priya-002 ordered salads 8 times, wraps 4 times Personalize trains a model on this data ↓At runtime:GET /recommendations?userId=rahul-001Returns: ["Biryani Palace", "Hyderabadi Biryani House", "Mughal Kitchen"] GET /recommendations?userId=priya-002Returns: ["Green Bowl", "Salad Bar", "Freshii"]No ML knowledge required. You provide the data. Personalize handles the model training, hosting, and serving.
Use cases: restaurant recommendations (Swiggy), product recommendations (e-commerce), content recommendations (Hotstar), "customers also bought" (retail).
Amazon SageMaker — Train Your Own Models
When pre-built services cannot solve your problem — your data is unique, your domain is specialised, accuracy requirements are very high — SageMaker lets you build and train custom ML models.
Pre-built service cannot solve it: Content moderation in a very specific domain Fraud detection on your specific transaction patterns Demand forecasting on your specific inventory data SageMaker workflow: Collect and prepare your training data in S3 Choose an algorithm or bring your own model code Train on managed infrastructure (GPU instances spun up, used, and terminated) Evaluate model accuracy Deploy to an endpoint (real-time inference) Your application calls the endpoint: send data → get predictionSageMaker is significantly more complex and expensive than the pre-built services. It requires ML expertise. Use it only when pre-built services fall short.
Hands-on Lab — Rekognition and Comprehend in the Console
Step 1 — Test Rekognition image analysis
AWS Console → Amazon Rekognition → Label detection (left menu)Try with a sample image or upload your own Click: Use your own image → upload a photo with identifiable objectsResults show: detected labels with confidence scoresExample: Person (99.8%), Bicycle (97.2%), Road (94.1%)Step 2 — Test content moderation
Rekognition → Content moderationUpload an image (use a stock photo of something clearly safe)Results: content moderation labels or "No unsafe content detected"Step 3 — Test Comprehend sentiment analysis
Amazon Comprehend → Real-time analysis → Input text Paste a positive review:"Swiggy delivered my food in 15 minutes, everything was hot and fresh. Amazing service!"Sentiment: POSITIVE Paste a negative review:"Waited 2 hours for my order. Food was cold. Never ordering again."Sentiment: NEGATIVEStep 4 — Test entity recognition
Comprehend → Real-time analysis → Entity recognition Paste:"Rahul Sharma from Bengaluru placed an order worth Rs 450 on 15th January 2024 using PhonePe." Results:Rahul Sharma → PERSONBengaluru → LOCATIONRs 450 → QUANTITY15th January 2024 → DATEPhonePe → ORGANIZATIONStep 5 — Test Translate
Amazon Translate → Real-time translationSource language: AutoTarget language: Hindi Input: "Your order has been delivered. Thank you for using our service."Output: "आपका ऑर्डर डिलीवर कर दिया गया है। हमारी सेवा का उपयोग करने के लिए धन्यवाद।"Step 6 — Test Textract
Amazon Textract → Analyze documentUpload a scanned document or use the sample providedSelect: Tables and FormsAnalyze See how Textract identifies table structure, form fields, and key-value pairsCompare to what a simple image-to-text tool would produce — Textract understands structureNo cleanup needed — these are all console demos with no persistent resources created.
Common Mistakes to Avoid
Common MistakeJumping to SageMaker before trying pre-built services. SageMaker requires ML expertise, data preparation, model training time, and ongoing maintenance. Rekognition, Comprehend, Textract, and Personalize solve most common AI use cases via a simple API call. Try the pre-built service first. Use SageMaker only when the pre-built service genuinely cannot solve your problem.
Common MistakeUsing Rekognition facial recognition without understanding the compliance requirements. In some countries and jurisdictions, storing and comparing facial data requires explicit user consent and has strict regulatory requirements. Check your local laws before building systems that identify individuals by face.
TipAll the pre-built AI services (Rekognition, Transcribe, Polly, Translate, Comprehend, Lex, Kendra, Textract, Personalize) can be combined. A customer support call → Transcribe converts audio to text → Comprehend extracts sentiment and key phrases → Translate converts to English for a global team → stored and searched by Kendra. Each service does one job and they chain together naturally.