Posts
All the articles I've posted.
The four weeks I spent not doing my job
Another semi-management story. On an AWS ProServe engagement, a big four partner had already owned the project for six months, and they wanted us nowhere near the real work. My first instinct was to push for a seat at the decision-making table, and that instinct was the problem. Here is what four weeks of deliberately not doing my job cost, the technical debt I watched pile up, and why the trust-building turned out to be the work itself.
The engineer I kept promoting anyway
A management story instead of an architecture one. I found Andrew in our company's training lab and over four years promoted him from student to architect. He was reliable, independent, technically excellent, and great with peers and customers. He was also, to exactly one person, a persistent pain: me. The Friday-evening complaint calls, the mailing list he quietly unsubscribed from to needle me, the HR conversation that backfired, and the promotion committee that needed an architect I happened to be. How I learned to manage for the team's output and not for my own comfort.
The Kubernetes cluster that had to survive the open ocean
A pre-sale asked for Kubernetes in the server room of a cruise ship. The ship is a distraction. The real constraint is that the network becomes a scheduled resource, present in port and gone at sea, and that splits the system into a runtime that has to survive for days with no link home and a lifecycle of updates and image pulls that can only happen while the ship is connected. EKS Anywhere runs the whole cluster on the boat, which is why it fit and why keeping the control plane onshore would not have.
The private cloud with no requirements
A government customer in the Gulf asked for a private cloud with Kubernetes and handed over no requirements. The design work turned out to be inventing the requirements: imagining the tenants who would show up and what they would ask for, then choosing technology against the two questions that stand in for the word best when there is no spec, whether it locks you in and whether it is already paid for.
The lift-and-shift that took six months
On paper it was the simplest kind of cloud migration: move one third-party server to a VM in AWS, same software, newer version. It took six months. The move was never the problem. The software assumed an environment that has mostly disappeared, the vendor who alone was allowed to install it couldn't operate the OS it ran on, and everything an operable system needs (readable logs, a health signal, automatic recovery) had to be built from the outside, around a box we weren't allowed to open.
The 502 nobody could find
A legacy app served through a chain of gateways was throwing intermittent 502s to hundreds of call-center agents. A CIO escalation had run for two months with nobody able to say which component was failing. I joined as the new architect, tried to trace it the way everyone else had, hit the same wall, and concluded the wall was the point: stop hunting the fault, start removing components until it has nowhere to hide. This is that story, the identity provider being misused as a gateway underneath it all, and the two-stage design that shipped.
The VPN I taught myself to run
My first proper consultancy project, in 2012, was a VPN service called PrivateWiFi. Hundreds of servers provisioned by hand, an API layer being patched live on two boxes at once with no version control, and an authentication path scattered across regions that a single login had to cross oceans to complete. This is the story of how I went from manual deployments to a 5,000-line Perl script to Chef, hired a PHP developer I was not qualified to hire, built extra services on top of the VPN, and finally moved RADIUS onto each server to stop OpenVPN from freezing.
Reading what the customer actually wrote
The classifier I wrote about earlier was the last step in a longer pipeline. Before you can route a support request you have to read it, and reading it means extracting text, stripping signatures, translating HTML, detecting sentiment. Every one of those leans on a managed AWS service, and every one of those services quietly does the wrong thing on real customer mail. Here are four traps, each with a screenshot or a workaround, and the through-line that connects them.
One day Peter came to the office
War stories from my first real networking job, at a mobile carrier called CDMA Ukraine. A predecessor named Peter who kept a nation-wide core network entirely in his head, passwords you could guess in one try, eBGP run as an interior protocol for fun, a live server balanced on a three-legged stool, and a pigeon in the wall. No morals, mostly. Just vignettes.
The OVSDB ORM I wasn't allowed to open-source
The seventh and last post in the CertaScale series, and the one I'm still not over. Driving OVN means talking to several databases over the OVSDB wire protocol by hand. So I built a full ORM in Go: every table as a typed object, references as pointers, and an ACL match compiler built on Reverse Polish Notation. I wanted to open-source it. The company said no, and it died with the company.