Skip to content

WRITING

Engineering notes, with the numbers and the limits.

Every number in these articles traces to a file in a public repository. Where a result was not measured, the article says so.

Case notes

One project per article: the decision, the evidence and the limit.

  1. 01My 27B model ran at 4.4 tokens per second. The GPU was fine.What memory-bandwidth arithmetic, speculative decoding, a paired benchmark and an audited grader taught me about serving an open model at interactive speed8 min
  2. 02The RAG chat that says "I can't answer" on purposeOne rule, gates that run before any model call, and what broke on the first live runs4 min
  3. 03The batch my ML pipeline refused to scoreA traceable MLOps pipeline over Brazilian education-finance data, the drift gate that earned its keep, and how to build the same gates7 min
  4. 04Ranking procurement notices without pretending to know what suppliers wantA transparent recommender, a time-split evaluation, and why public procurement history is not user propensity5 min
  5. 05FinOps starts with what you promise not to look atTurning a narrow cloud-metadata boundary into auditable cost findings, a report and a deck, with the contract, the rules, the commands and the limits6 min
  6. 06Idempotency before autoscaling: what an agent runtime needs firstRedis Streams, killed workers and a three-node kind cluster: the state contract, the evidence for it, the benchmark I refuse to over-read, and the commands to reproduce each claim6 min
  7. 07Observable RAG without unsupported answersA response contract with three outcomes, a bug that a failing test caught, a gateway client and a trace exporter that I tested against real services, and how to run all of it7 min
  8. 08Testing a privacy policy like softwareA stable organization key, a release gate that raises before it writes, a worked run you can reproduce, and an honest account of what a hashed identifier does not protect11 min
  9. 09Four billion rows on Kaggle, and the checks that caught what I got wrongReorganising eleven Brazilian public-data sources into 31 datasets: Parquet row groups, upload speed, a command that deletes files, and columns that should never have shipped5 min
  10. 10A portfolio where every claim has an evidence stateThree labels, a denied-content test that runs in CI, a static export that is checked like software, and why I stopped writing "production-ready"8 min

Run your own model

A tutorial series on serving, benchmarking and operating a 27B open model on one rented GPU.

  1. 11Serve a 27B open model on a rented GPU: a reproducible, OpenAI-compatible endpointPart 1 of 5. The arithmetic that predicts your speed, the exact steps to bring the endpoint up, and the checks that tell you it is right10 min
  2. 12Serving an abliterated model: what changes, what does not, and how to choose a checkpointPart 2 of 5. The mechanism behind abliteration, FP8 against NVFP4 on the same weights, how to verify a third-party checkpoint, and a checklist before you open the endpoint6 min
  3. 13Benchmark a model you host without fooling yourselfPart 3 of 5. Speed, context length, quality and significance, with the tests that caught a silent KV-cache failure and a grader that was wrong8 min
  4. 14Test a model's refusal behaviour without writing anything dangerousPart 4 of 5. A refusal probe for lawful, adult-oriented prompts, run against a self-hosted abliterated model and Sonnet 5, and an honest reading of what 37 prompts can show6 min
  5. 15Make a GPU deployment survive a power cycle: idempotent scripts, Terraform with no provider, and a CI/CD splitPart 5 of 5. How to rebuild the same endpoint on a different machine by changing one address, how to prove it with tests, and how to keep the secrets out of a public repository7 min

The public data stack

Postgres for search and BI, a read-only SQL agent, and dashboards inside a portfolio.

  1. 16One Postgres for full-text search, vector search and dashboardsA tutorial: load a public-data lake into pgvector, index it three ways, give each reader its own limits, and keep it fast on a small server6 min
  2. 17A read-only SQL agent for public data, and the ways it got the answer wrong firstA tutorial: let a language model write SQL over numeric tables, keep it on a short leash, show the query as evidence, and test it with real questions6 min
  3. 18Put real dashboards in a portfolio: Metabase public links, an nginx edge and a seed scriptA tutorial: stand up Metabase over a read-only Postgres, build dashboards from code, embed them safely in your site, and check that they actually draw5 min
  4. 19I swapped a managed embedding API for my own model. The index cost more recall than the model did.Re-embedding 1.25 million procurement notices with Qwen3-Embedding, then measuring it against Vertex on the same 30 questions5 min