Running and Choosing AI Models: Local and Cloud
Learn to choose models for ops work by quality, latency, and cost, route cheap and strong models, and run local models with Ollama for sensitive logs.
What You'll Learn
Understanding Why Model Choice Is an Engineering Decision
Picking a model is a trade between quality, speed, cost, and where your data goes, and no single model wins all four.
Choosing a Model by Class, Not by Name
Choose a class of model first and a specific name second, because names and prices change every few months while classes stay useful.
Estimating Cost Properly
Real LLM cost needs input tokens and output tokens priced separately, because output tokens usually cost several times more per token.
Running Local Models with Ollama
Ollama runs open models on your own machine and exposes an OpenAI-compatible API, so the same llm.py code talks to local and hosted models.
Understanding Quantization
Quantization stores a model's weights in fewer bits so it uses less memory and often runs faster, at a small cost in quality.
Keeping Local Endpoints Safe
A local model endpoint is a service on your network, and Ollama has no authentication of its own.
Skills You'll Master
Curriculum Index9 topics
Understanding Why Model Choice Is an Engineering Decision
Picking a model is a trade between quality, speed, cost, and where your data goes, and no single model wins all four.
Choosing a Model by Class, Not by Name
Choose a class of model first and a specific name second, because names and prices change every few months while...
Estimating Cost Properly
Real LLM cost needs input tokens and output tokens priced separately, because output tokens usually cost several times...
Running Local Models with Ollama
Ollama runs open models on your own machine and exposes an OpenAI-compatible API, so the same llm.py code talks to...
Understanding Quantization
Quantization stores a model's weights in fewer bits so it uses less memory and often runs faster, at a small cost in...
Keeping Local Endpoints Safe
A local model endpoint is a service on your network, and Ollama has no authentication of its own.
Routing Cheap and Strong Models
A router sends easy, high-volume work to a cheap model and only the hard cases to a stronger one.
Evaluating Models on Your Own Alerts
Evaluate models on your own labelled alerts with validated JSON, because public benchmarks measure general knowledge...
Running the Hands-on Lab, Quick Reference, and Common Mistakes
This lab estimates cost, measures two quantizations, secures a gateway, and compares models on labelled alerts.
Career Impact
Roles that use the skills in this module.
- High Demand
AIOps Engineer
₹18L - ₹35L a year
- High Demand
AI Platform Engineer
₹20L - ₹40L a year
- Growing
MLOps Engineer
₹18L - ₹32L a year
Next Modules
Related Guides
Practice on the Coding Sheet
Not a software engineer sheet. Every problem comes from real DevOps, SRE, Platform and Cloud interviews, from your first script to a system you build yourself.
Open the Coding SheetFrequently Asked Questions
Use a local model when logs or alerts contain data that must not leave your network, when the environment is air-gapped, or when volume makes per-token pricing painful. Use a hosted model when you need the strongest reasoning and the data is safe to send. Many teams use both through a router.
It stores each weight in about 4 bits instead of 16, so the model needs roughly a quarter of the memory and often runs faster. The cost is a small quality loss that varies by model. Measure it on your own alerts rather than assuming.
Not directly. Ollama has no authentication, so anyone who can reach the port can use your hardware and send prompts. Keep it on localhost and put a gateway with an API key, TLS, and a firewall in front.
Count input and output tokens separately, because output tokens usually cost several times more per token. Multiply each by its own price and add them. Ignoring output tokens can understate the bill by more than half.