-

Self-Hosted Inference: Choosing Between Ollama and vLLM
Ollama or vLLM? Depends whether you’re running a single toll booth or building a tollway. A practical framework for choosing self-hosted inference engines across…
-

Scaffolding MCP Servers with kmcp
Everyone’s shipping MCP servers built to run on a laptop. Running one as a managed service other things depend on is a different problem:…
-

Operationalizing
Getting things running was Post 2’s job. Keeping them secure, observable, and maintainable is a different problem entirely.
-

The Weight on One Side of the Wall
Almost all the work happens on the connected side. Almost none of it happens on the air-gapped side. That’s not an accident — that’s…
-

The Eccentricities of Air-Gapped AI
No frontier model. No API key. No internet connection to fall back on. Here’s why agentic AI is still worth building under those constraints…
