Skip to main content

LLM

Browse all articles, tutorials, and guides about LLM

10posts

Posts

FinOps
|14 min read

Knowing What Your AI Feature Costs Before Finance Does

The model invoice has one line per model, and finance wants one line per feature. We ran three features on one key and recorded every usage block: a one-word alert label cost a third to a half of a full incident investigation, because of reasoning tokens nobody reads. Here is how to measure that per request and tag it per feature.

DevOps
|14 min read

Your Semantic Cache Answers the Question Next Door

We put a semantic cache in front of an ops assistant on DigitalOcean Serverless Inference and replayed 288 hand-labelled questions through six embedding models. At the threshold that cut the bill by about 30 percent, about one hit in three was the answer to a different question. Someone asking how to undo a pushed commit got the instructions for an unpushed one.

DevOps
|14 min read

We Measured the 200x Claim, and Got It Wrong Twice First

Last week we told you to measure a vendor claim on your own workload instead of repeating it. Then we got access, did exactly that, and produced two confident numbers that were both artefacts of our own bad method. Here is the real result, and the two mistakes, which are more useful than the result.

DevOps
|10 min read

Jev and the Classification Problem Hiding in Your LLM Bill

Routing, tagging, triage and extraction are classification wearing a chat interface. A new model class is arguing that point loudly, with numbers worth reading carefully. Here is how to tell whether the argument applies to your pipeline, and how to read a 200x claim before you repeat it.

DevOps
|8 min read

Swapping Across 25 Models With One Line

Choosing a model is usually a commitment: an SDK, a key, an integration. Through the gateway it is a string, so you can shop the whole catalog per task. And the catalog spans a 100x price range, which turns model choice into your biggest cost lever. Here is the swap, the price spread, and a real multi-model run.

DevOps
|8 min read

Per-Branch AI Endpoints: Isolating Model Spend Across Prod, Preview, and CI

When previews, CI, and production all call models with the same key, you cannot tell what a preview cost or notice a runaway test until the invoice. Because a Neon branch is its own deployment with a usage ledger that lives in the branch's Postgres, model spend is attributed and isolated per environment. I proved it: a CI branch spent tokens while production stayed flat.

DevOps
|8 min read

Model Fallback and Routing Without a Provider SDK Each

Models have outages, rate limits, and bad minutes. A resilient app falls back to another one, but building that across providers normally means a different SDK and error shape for each. Through one OpenAI-compatible gateway, fallback is a loop over model names. Here it is, tested against a real failure.

DevOps
|9 min read

One Key for Claude, GPT, and Gemini: the Gateway Pattern

Using three model providers usually means three API keys, three SDKs, and three billing relationships sprayed across your code. An AI gateway collapses that to one credential and one OpenAI-compatible endpoint. I proved it on a Neon Function: the same call answered by GPT, Claude, and Gemini.

DevOps
|8 min read

Your First Serverless LLM Call on DigitalOcean in 10 Minutes

DigitalOcean's Inference Engine gives you an OpenAI-compatible endpoint with pay-per-token pricing and no GPU to manage. Here is the fastest path from zero to a working call, with curl, Python, and Node, every snippet run against the live API.

DevOps
|10 min read

The US Government Pulled Two Frontier Models Overnight. The Real Lesson Is About Your Stack

On June 12, 2026, an export-control directive forced Anthropic to disable Claude Fable 5 and Mythos 5 for every user worldwide, three days after launch. The policy fight is interesting. The operational lesson for anyone building on a single model provider is more urgent.