WRITING
Engineering notes, with the numbers and the limits.
Every number in these articles traces to a file in a public repository. Where a result was not measured, the article says so.
Case notes
One project per article: the decision, the evidence and the limit.
- 01My 27B model ran at 4.4 tokens per second. The GPU was fine.What memory-bandwidth arithmetic, speculative decoding, a paired benchmark and an audited grader taught me about serving an open model at interactive speed8 min
- 02The RAG chat that says "I can't answer" on purposeOne rule, gates that run before any model call, and what broke on the first live runs4 min
- 03The batch my ML pipeline refused to scoreA traceable MLOps pipeline over Brazilian education-finance data, the drift gate that earned its keep, and how to build the same gates7 min
- 04Ranking procurement notices without pretending to know what suppliers wantA transparent recommender, a time-split evaluation, and why public procurement history is not user propensity5 min
- 05FinOps starts with what you promise not to look atTurning a narrow cloud-metadata boundary into auditable cost findings, a report and a deck, with the contract, the rules, the commands and the limits6 min
- 06Idempotency before autoscaling: what an agent runtime needs firstRedis Streams, killed workers and a three-node kind cluster: the state contract, the evidence for it, the benchmark I refuse to over-read, and the commands to reproduce each claim6 min
- 07Observable RAG without unsupported answersA response contract with three outcomes, a bug that a failing test caught, a gateway client and a trace exporter that I tested against real services, and how to run all of it7 min
- 08Testing a privacy policy like softwareA stable organization key, a release gate that raises before it writes, a worked run you can reproduce, and an honest account of what a hashed identifier does not protect11 min
- 09Four billion rows on Kaggle, and the checks that caught what I got wrongReorganising eleven Brazilian public-data sources into 31 datasets: Parquet row groups, upload speed, a command that deletes files, and columns that should never have shipped5 min
- 10A portfolio where every claim has an evidence stateThree labels, a denied-content test that runs in CI, a static export that is checked like software, and why I stopped writing "production-ready"8 min
Run your own model
A tutorial series on serving, benchmarking and operating a 27B open model on one rented GPU.
- 11Serve a 27B open model on a rented GPU: a reproducible, OpenAI-compatible endpointPart 1 of 5. The arithmetic that predicts your speed, the exact steps to bring the endpoint up, and the checks that tell you it is right10 min
- 12Serving an abliterated model: what changes, what does not, and how to choose a checkpointPart 2 of 5. The mechanism behind abliteration, FP8 against NVFP4 on the same weights, how to verify a third-party checkpoint, and a checklist before you open the endpoint6 min
- 13Benchmark a model you host without fooling yourselfPart 3 of 5. Speed, context length, quality and significance, with the tests that caught a silent KV-cache failure and a grader that was wrong8 min
- 14Test a model's refusal behaviour without writing anything dangerousPart 4 of 5. A refusal probe for lawful, adult-oriented prompts, run against a self-hosted abliterated model and Sonnet 5, and an honest reading of what 37 prompts can show6 min
- 15Make a GPU deployment survive a power cycle: idempotent scripts, Terraform with no provider, and a CI/CD splitPart 5 of 5. How to rebuild the same endpoint on a different machine by changing one address, how to prove it with tests, and how to keep the secrets out of a public repository7 min
The public data stack
Postgres for search and BI, a read-only SQL agent, and dashboards inside a portfolio.
- 16One Postgres for full-text search, vector search and dashboardsA tutorial: load a public-data lake into pgvector, index it three ways, give each reader its own limits, and keep it fast on a small server6 min
- 17A read-only SQL agent for public data, and the ways it got the answer wrong firstA tutorial: let a language model write SQL over numeric tables, keep it on a short leash, show the query as evidence, and test it with real questions6 min
- 18Put real dashboards in a portfolio: Metabase public links, an nginx edge and a seed scriptA tutorial: stand up Metabase over a read-only Postgres, build dashboards from code, embed them safely in your site, and check that they actually draw5 min
- 19I swapped a managed embedding API for my own model. The index cost more recall than the model did.Re-embedding 1.25 million procurement notices with Qwen3-Embedding, then measuring it against Vertex on the same 30 questions5 min