Skip to main content

Running and Choosing AI Models: Local and Cloud

Learn to choose models for ops work by quality, latency, and cost, route cheap and strong models, and run local models with Ollama for sensitive logs.

~3 hours
9 Topics
Hands-on Scenarios

What You'll Learn

Understanding Why Model Choice Is an Engineering Decision

Picking a model is a trade between quality, speed, cost, and where your data goes, and no single model wins all four.

Choosing a Model by Class, Not by Name

Choose a class of model first and a specific name second, because names and prices change every few months while classes stay useful.

Estimating Cost Properly

Real LLM cost needs input tokens and output tokens priced separately, because output tokens usually cost several times more per token.

Running Local Models with Ollama

Ollama runs open models on your own machine and exposes an OpenAI-compatible API, so the same llm.py code talks to local and hosted models.

Understanding Quantization

Quantization stores a model's weights in fewer bits so it uses less memory and often runs faster, at a small cost in quality.

Keeping Local Endpoints Safe

A local model endpoint is a service on your network, and Ollama has no authentication of its own.

Skills You'll Master

OLLAMALOCAL-LLMMODEL-SELECTIONLLM-COSTAIOPS

Curriculum Index9 topics

Career Impact

Roles that use the skills in this module.

  • AIOps Engineer

    ₹18L - ₹35L a year

    High Demand
  • AI Platform Engineer

    ₹20L - ₹40L a year

    High Demand
  • MLOps Engineer

    ₹18L - ₹32L a year

    Growing
See how this is asked in interviews

Practice on the Coding Sheet

Not a software engineer sheet. Every problem comes from real DevOps, SRE, Platform and Cloud interviews, from your first script to a system you build yourself.

Open the Coding Sheet

Frequently Asked Questions

Use a local model when logs or alerts contain data that must not leave your network, when the environment is air-gapped, or when volume makes per-token pricing painful. Use a hosted model when you need the strongest reasoning and the data is safe to send. Many teams use both through a router.

It stores each weight in about 4 bits instead of 16, so the model needs roughly a quarter of the memory and often runs faster. The cost is a small quality loss that varies by model. Measure it on your own alerts rather than assuming.

Not directly. Ollama has no authentication, so anyone who can reach the port can use your hardware and send prompts. Keep it on localhost and put a gateway with an API key, TLS, and a firewall in front.

Count input and output tokens separately, because output tokens usually cost several times more per token. Multiply each by its own price and add them. Ignoring output tokens can understate the bill by more than half.