Skip to content
Works on My Machine

Archives

All the articles I've archived.

202625
October2
  • The four weeks I spent not doing my job

    Another semi-management story. On an AWS ProServe engagement, a big four partner had already owned the project for six months, and they wanted us nowhere near the real work. My first instinct was to push for a seat at the decision-making table, and that instinct was the problem. Here is what four weeks of deliberately not doing my job cost, the technical debt I watched pile up, and why the trust-building turned out to be the work itself.

  • The engineer I kept promoting anyway

    A management story instead of an architecture one. I found Andrew in our company's training lab and over four years promoted him from student to architect. He was reliable, independent, technically excellent, and great with peers and customers. He was also, to exactly one person, a persistent pain: me. The Friday-evening complaint calls, the mailing list he quietly unsubscribed from to needle me, the HR conversation that backfired, and the promotion committee that needed an architect I happened to be. How I learned to manage for the team's output and not for my own comfort.

September16
  • The Kubernetes cluster that had to survive the open ocean

    A pre-sale asked for Kubernetes in the server room of a cruise ship. The ship is a distraction. The real constraint is that the network becomes a scheduled resource, present in port and gone at sea, and that splits the system into a runtime that has to survive for days with no link home and a lifecycle of updates and image pulls that can only happen while the ship is connected. EKS Anywhere runs the whole cluster on the boat, which is why it fit and why keeping the control plane onshore would not have.

  • The private cloud with no requirements

    A government customer in the Gulf asked for a private cloud with Kubernetes and handed over no requirements. The design work turned out to be inventing the requirements: imagining the tenants who would show up and what they would ask for, then choosing technology against the two questions that stand in for the word best when there is no spec, whether it locks you in and whether it is already paid for.

  • The lift-and-shift that took six months

    On paper it was the simplest kind of cloud migration: move one third-party server to a VM in AWS, same software, newer version. It took six months. The move was never the problem. The software assumed an environment that has mostly disappeared, the vendor who alone was allowed to install it couldn't operate the OS it ran on, and everything an operable system needs (readable logs, a health signal, automatic recovery) had to be built from the outside, around a box we weren't allowed to open.

  • The 502 nobody could find

    A legacy app served through a chain of gateways was throwing intermittent 502s to hundreds of call-center agents. A CIO escalation had run for two months with nobody able to say which component was failing. I joined as the new architect, tried to trace it the way everyone else had, hit the same wall, and concluded the wall was the point: stop hunting the fault, start removing components until it has nowhere to hide. This is that story, the identity provider being misused as a gateway underneath it all, and the two-stage design that shipped.

  • The VPN I taught myself to run

    My first proper consultancy project, in 2012, was a VPN service called PrivateWiFi. Hundreds of servers provisioned by hand, an API layer being patched live on two boxes at once with no version control, and an authentication path scattered across regions that a single login had to cross oceans to complete. This is the story of how I went from manual deployments to a 5,000-line Perl script to Chef, hired a PHP developer I was not qualified to hire, built extra services on top of the VPN, and finally moved RADIUS onto each server to stop OpenVPN from freezing.

  • Reading what the customer actually wrote

    The classifier I wrote about earlier was the last step in a longer pipeline. Before you can route a support request you have to read it, and reading it means extracting text, stripping signatures, translating HTML, detecting sentiment. Every one of those leans on a managed AWS service, and every one of those services quietly does the wrong thing on real customer mail. Here are four traps, each with a screenshot or a workaround, and the through-line that connects them.

  • One day Peter came to the office

    War stories from my first real networking job, at a mobile carrier called CDMA Ukraine. A predecessor named Peter who kept a nation-wide core network entirely in his head, passwords you could guess in one try, eBGP run as an interior protocol for fun, a live server balanced on a three-legged stool, and a pigeon in the wall. No morals, mostly. Just vignettes.

  • The OVSDB ORM I wasn't allowed to open-source

    The seventh and last post in the CertaScale series, and the one I'm still not over. Driving OVN means talking to several databases over the OVSDB wire protocol by hand. So I built a full ORM in Go: every table as a typed object, references as pointers, and an ACL match compiler built on Reverse Polish Notation. I wanted to open-source it. The company said no, and it died with the company.

  • From 12 to 60 Gbps with DPDK, when we tried to become a 5G-edge box

    The sixth post in the CertaScale series. When private cloud proved a hard sell, the owners pivoted to 5G edge, and edge means line-rate on commodity hardware. Our kernel data path topped out around 12 Gbps on 100 Gbps NICs. Taking it out of the kernel with DPDK, hugepages, poll-mode drivers, and userspace-bound NICs got the same servers to roughly 60 Gbps.

  • Running VMs inside pods in 2016, before Kata existed

    The fifth post in the CertaScale series, and the one the earlier posts kept promising: running full virtual machines inside Kubernetes pods in 2016, before Kata Containers existed and before RuntimeClass was a thing. A pod is just a namespace boundary and a lifecycle. Put KVM inside it and a VM schedules, migrates, and dies like a pod.

  • A flat network, GCP-style, on early Kubernetes

    The fourth post in the CertaScale series, and the idea I'm proudest of: a flat overlay borrowed from how Google Cloud addresses machines. /32 addresses plus routes, an address operator handing out CRDs from a pool, and the reason we built all of it — moving a workload between nodes without it losing its IP.

  • Giving Kubernetes an enterprise network

    The third post in the CertaScale series. Our customer was building a private cloud that had to drop into any enterprise's network, so we spoke the language every enterprise already speaks and extended it to Kubernetes pods: VLANs, static and DHCP addressing, and RFC 4594 QoS.

  • Finding ovn-kubernetes when it was seven days old

    The second post in the CertaScale series. Handed a failed project and a title, I inherited a broken Vagrant file that created a single bridge, panic-googled my way to a week-old repo, learned OVS internals the hard way, and rewrote our side from Python to Go when the team changed.

  • The private cloud we built before it was cool

    From 2016 to 2019 I helped build a private cloud on top of very early Kubernetes, with a drag-and-drop canvas, VMs running inside pods before Kata existed, and an enterprise-grade network layer I led, from GCP-style flat overlays to a 12-to-60 Gbps DPDK data path. The company is gone. The story shouldn't be.

  • I can run AI models on my gaming PC now

    I bought a powerful GPU for gaming a couple of years ago. It turns out a Radeon 7900 XTX is also a perfectly good reason to run language models locally.

  • Designing agentic AI on Bedrock, a real-world journey

    A support-ticket classifier that started as one enormous prompt costing $145k a year, went through a pile of overbuilt architectures, and ended as a deterministic pipeline with no agent in it at all, costing $744. Here's every wrong turn and what actually broke.

August7
  • Shift-left, or catching bugs before they cost

    In 2015 I gave a talk about building better CI with Jenkins pipelines. Ten years and half a dozen CI tools later, the tooling has changed completely and the core idea hasn't moved an inch. Here's the shift-left setup I run today, and why it's the same lesson I was giving in a conference room a decade ago.

  • Quantum computing, from someone who has to stay skeptical

    I worked on a quantum computing platform. Here's the honest version of what quantum can and can't do today, and what it costs to run a real circuit yourself.

  • I rebuilt my CV as code

    For years I kept my CV in a Word document. Then I went job hunting and discovered the landscape had completely changed. So I turned my CV into a data-driven static site.

  • My 22cm cube that replaced the cloud

    What started as a Synology NAS replacement turned into a full personal server running 70+ containers. Here's what's inside and why.

  • Looking inside the disc your radiologist gives you

    Every time you get an X-ray, CT, or MRI, you walk away with a disc. Did you ever look at what's actually on it? I did, and found far more than printed pictures.

  • What my CPAP machine knows about me

    I snore. Loud. After getting an APAP machine, I found its SD card holds far more data than the companion app shows. So I built a viewer for it.

  • Making sense of my own DNA with OSGenome

    How an underwhelming Ancestry test sent me down a rabbit hole of raw DNA files, SNPedia, and an open-source project I ended up rebuilding to work with my own data.