Skip to main content

Advanced AI Specialisations

Explore five AI engineering tracks: advanced agents, multimodal, computer vision, voice, and language AI, then build one starter project in depth.

~2.5 hours of reading, plus the chosen track's lab
8 Topics
Hands-on Scenarios

What You'll Learn

Choosing Your Specialisation Track

By now acme-assist answers questions from the help centre, issues refunds safely, and has a release gate.

Going Deeper on Agent Engineering

This track is for you if building the refund agent was the part of the roadmap you enjoyed.

Building Multimodal Document Extraction

Multimodal AI does not mean image generation.

Building Computer Vision Counting

Computer vision is the track for precise, structured, real-time visual understanding at a scale and cost general models are not built for.

Building Voice AI Pipelines

Voice AI is the track for spoken interfaces: phone support, in-app assistants, replacing IVR menus.

Building Language AI Systems

Language AI, often called NLP, is the track for teams handling huge volumes of text where an LLM call per item is too slow, too expensive, or...

Skills You'll Master

AI-AGENTSMULTIMODAL-AICOMPUTER-VISIONVOICE-AILANGUAGE-AI

Curriculum Index8 topics

Career Impact

Roles that use the skills in this module.

  • AI Engineer

  • MLOps Engineer

  • Platform Engineer

See how this is asked in interviews

Practice on the Coding Sheet

Not a software engineer sheet. Every problem comes from real DevOps, SRE, Platform and Cloud interviews, from your first script to a system you build yourself.

Open the Coding Sheet

Frequently Asked Questions

No. Pick one, build its starter, and go deep. Interviewers can tell one working system with handled edge cases from five demos that break on the first unusual input. Keep working knowledge of the others.

Choose by the kind of product and company you want to work on, not by what sounds impressive. Text classification at scale points to Language AI, phone support points to Voice AI, and warehouse cameras point to Computer Vision.

No. Every starter in this module runs on a laptop CPU for free. The optional go-further steps, such as fine-tuning a detector on a dense-shelf dataset, are the parts that benefit from a GPU, and each says so.

No. It is covered at awareness level only. It matters for specialised ML and research roles, and an AI engineer mainly needs to know what RLHF is and why chat models behave differently from raw pre-trained ones.