When all you have is Kubernetes everything is a nail

TLDR; People are now so obsessed with using Kubernetes that they just try to shove it everywhere they can even when it is really not the best choice. It is becoming increasingly annoying to deal with such practices.

Kubernetes is good and has many very useful features

Kubernetes is an amazing piece of software and it has standardized a lot of things in the tech industry.

Instead of rewriting something to emulate some of its capabilities like described here, you should always evaluate whether you need the features that come with kubernetes before coding it on your own once again.

Amongst them are :

All of these points are killers. I mean, those are all really useful features that are needed at some point or another and even if you don’t need them you can just ignore the noise and go with the flow as your project grows. It is understandable why kubernetes became so ubiquitous in tech industry over the years.

You can bend Kubernetes to do a lot of things

You can now do everything in kubernetes. Via the extensibility of the API you can manage certificates, manage applications lifecycle, you can run statefullset and even VM with kubevirt. It doesn’t matter what you want to achieve, it is possible to do it on kubernetes because of how far the tool is willing to let you go. You might encounter weird behaviour, you might struggle to make it work, you might have to rewrite everything after each update but when there is a will there is a way and people are more than willing to go all in into kubernetes. Entire startups are based on the extensibility of the kube API. At the end of the day it is just a bunch of containers bunched together on three or more VM. And with the right permission set and a little bit of coding you can do whatever you want. People now even use the kube api to run VM with things like kubervirt. This is even more crazy to me because while I have not used it directly I can’t imagine it is better than any commercial hypervisor out there and the provided API.

The ecosystem might be a little bit too much

Cost of the solution

Many teams do not host kubernetes themselves and rely on cloud providers to do so. But in using those services you are vulnerable to their aggressive pricing. For a company hosting one or two clusters it isn’t really a big expense but when time comes to cut cost you will end up putting everything on one cluster and this is really bad. If you compare the cost of an EKS to a simple beefy VM you are always coming up ahead with the VM.

Ressource allocation is also widly inneficient, it is efficient only when you have a cloud native application that can scale and let the cluster scale itself. The second you introduce statefulsets you are in for a lot of troubles to manage costs.

Cost of maintenance with a release every 6 months

This is one the main pain points of kubernetes. The release cycle is short and the support timeline is even shorter. There is a reason AWS and others charge you extra to keep an outdated version running. It moves fast and while there is generally no problems, if you are the kind of team to add extensions and operator everywhere, you will not be able to keep up to date with all the new versions. Not mentionning you need testing for each of your applications stack. So if you have a large number of clusters, the moment you finished one migration there is a new version out and you have to start all over again. This is, I think, a very frequent burn out motive in infra teams.

The contributors problems

While not a problem unique to kubernetes and more associated with open source in general, there is a lack of maintainers able to work on the kubernetes core implementation.

Notably stuff like ingress nginx that went out of date at the start of 2026

A lot of sidekick in the cnfc landscape are also propped up by venture capitalist and the money will eventually dry out. Relying on these stacks seems all the more risky when again, the other possible path in a lot of situation is to simply use a trusted linux server distribution like debian or rhel.

The operator pattern

Always in need of reinventing the wheel, people in this field decided that having a fully capable orchestrator in the name of kubernetes was not enough. They needed to add an orchestrator inside the orchestrator by using the operator pattern.

There is no software that needs this. All the tools you are looking for are already a part of kubernetes. Using operators is just adding even more friction to something that should be simple. I don’t need a container running indefinitely to cycle through the real application lifecycle and checking it out against the CRDs in the kube API. This is nuts. It does not solve anything, you still need someone to interact with the kube api at some point, adding CRDs and a container in the middle only complexifies this. Your application can handle a lot more than that by itself. Your code can run the migration by itself and set the living probes autonomously.

Same but not the same

A VM can be exported as a disk and used everywhere on any provider. Your cluster is a lot more difficult to save, backup and relaunch and even if you manage to do it…there are incompatible drivers, csi and other specific cloud software you need to have. So if trying to bootstrap the same applications in AKS and EKS you have to install the helm charts from the different providers, etc, etc…it is an endless pit of complication for what is supposed to be a unified API.

For a fair comparaison some capabilities in VM require guest tools installed but this is in no way the same amount of complexity and work compared to adapting every third party from your EKS to your AKS.

Tools that are out of their place notably helm

Speaking of installing third party software, you end up setting things up with helm chart everywhere. Helm chart that are themselves completely abused and used for configuration of a lot of software. How many helm chart really restrict you in the way you are doing things ? You can have thousands of line long values.yaml that allows you to modify every part of the manifest produced at the end. It comes to a point where even if using kubernetes I don’t want to even speak about helm charts. Give me the kubernetes manifests directly. Your service needs to be a statefulset so give me a manifest that do this, do not let me choose in the values file of your helm chart. Helm was only necessary because it is hard to install a lot of manifest easily and people wanted an esay peasy way to install third party applications in their clusters.

My personnal example

I used to work somewhere where they decided to host EVERYTHING with kubernetes. Only one of the two services that we were selling was really made with cloud in mind and integrated perfectly with this approach. And sure it was bliss, updates were rolled over and the service had selfhealing and integrated perfetcly in a cloud environment. For this first SAAS project it was pure perfection (minus some defects but nothing here is to blame on kubernetes).

For the other part of the infra we managed though…

We had to host a fully stateful VPN connected multi instance multi client high availibity legacy product in the cloud.

Of course Kubernetes was not the correct choice and despite my appeals, nobody wanted to change a thing. I don’t know if this is a case of “you can’t get fired for buying IBM” but it became a problem quickly.

Statefulness

The project being stateful already is a problem. We are from the get go using one instance for one client. It means a lot of instances on the cluster that we put in different namespaces. It also mean we use the cluster more like a hypervisor with containers trying to be VM. From this issue arose a lot of other issues.

For this reason alone we were already on the wrong foot and should have used VM.

Ressource usage

Ressources usage was really difficult to control, because each client got its node group but the different instances of the clients either prod or demo or staging didn’t have a proper node. And clients could request more sites to run while the pricing setup by sales only container informations about nodes. Some clients were over provisionned and some underprovisionned.

The solution was to give one node to one instance of the application but at this point just use a VM. In the meantime, demo bugs caused production outage on a software that was vital for some client’s operations.

Networking

One of the worst offender in the list. The application is supposed to run in a private subnet for it to work, but not any private subnet, the subnet of the client on its premises. It means they allocate us a part of a private subnet to do our shenanigans and allow some ip through their firewall to connect. It also means that we have at least one end to end connection per client on our main AWS VPC, meaning that we need to find a private subnet accomodating everyone at the same time. But these entreprises never had any intention of using the software when they first created their network years ago. We wound up with a weird solutions of using 100.64/10 but this is only a patchwork and not a real fix. By using this subnet and a lof of tricks in the VPC we managed to let anyone connect with private IP. But let me tell you, nobody wanted to touch the networking stack before, and now nobody dared speak about it after. The migrations took a long time and while it lasted to entire environment with consumming clusters needed to stay up. I don’t know if you have tried to organize migrations with clients that have working solutions but they don’t want to do it right now so it always spanned out months to migrate everyone.

Again this would have been so much easier if everything was separated in VM groups dedicated to clients. No common network would have meant no network problem to begin with and a migration at the clients pace.

Refactors

For all the reasons above and to try to have better grip on the upgrade path for our clients there were mutliple refactors. Please believe me when I tell you what a pain it was for the different refactors to happen. Between the endless rewriting of software, the gigantic tests to run to ensure everything could restart from IAC only, the different clients to migrate after the initial testing was done, the amount of preparational work to ensure everything went smoothly was enormous. And still I didn’t tell you the best of it, as networks and kubernetes cluster were all shared, we had to first create a new env and migrate everyone. Even clients where everything was working great and that we could have left alone needed to be migrated (with a downtime of course) because another client chose the wrong private subnet at home ten years ago.

We also refactored to try and separate the demo and production environnement due to ressource usage issues. It went great, we now had double the networks craziness and double the amount of kubernetes clusters. At this point it became an entire mental load to know what you were doing where. You had to concentrate for 5 minutes to even know what you were trying to achieve where in the infra.

Developpement

Adding new features was a nightmare as all of AWS services were intertwined with each other and with the mess of a network we had made. Every new flow of data had to go through hoops before reaching any kind of services and we had some pretty weird limitations. No flexibility on our part because we had tungled ourselves in our own web of dead ends. With each new fix, no solutions in sights and only more patchwork to make it appear as everything is fine.

There was also friction with dev team because they were developping with defaults for Windows Server and not whatever we were doing in the cloud, on kubernetes, in a linux container.

Availibility

Availibility was promised to the clients and the plan was to use AZ from AWS to provide for it through kubernetes. Once again this was never truly tested, the times it was tested it didn’t work but we all kept trying to pretend the availibility of this application was bullet-proof and that a datacenter from Amazon failing wasn’t going to stop the service.

We would have been better with hyperconverged VMs in this regard too.

All this for what ?

You want to hear the dreaded truth my team didn’t want to hear ?

All of this was for nothing.

The application itself was really simple with a front SPA that could be delivered with any web reverse proxy and the backend service consisting of 3 main Java services that were pretty well coded and maintained.

All of this would have fit in a simple VM and could have been intregrated with IAC either with a cloudinit config or a Nixos config. Kubernetes didn’t provide anything but breakages and complexity.

Being reasonable is part of your job

Now where does this left us ? We should aspire to choose as wisely as possible when trying to do a good job shouldn’t we ?

People should focus on the core functionnality of their applications. Sure if you are designing from the get go a scalable, cloud native application then go straight for Kubernetes. But for fuck sake, when porting a legacy system that is as stateful as can be, restrain and choose an hyperconverged VM. Ressource allocation will be easier, maintenance too. In the end a tool is just a tool, you CAN make it work but it is going to be be much easier to use the right tool.

I guess for me it is only hyperconverged VM for stateful application from now on, true cloud native applications are more than welcome on my kubernetes clusters though. I have had enough of trying to make things work in Kubernetes if they are not meant to be, hypers be damned.