Skip to content

Available for the next problem

Between the rack
and the cloud

I build the platform application teams deploy onto themselves, so shipping a service stops being a ticket.

Right now that means taking a Windows monolith apart into services a team can deploy itself, and running the multi-GPU inference behind it. I came up through hardware, hypervisors and Windows before containers were in my job description, which is why when a deployment fails I know which layer to open.

Production services deployed
100+Production services deployed
Endpoints under scripted management
500+Endpoints under scripted management
Multi-GPU inference clusters
5+Multi-GPU inference clusters

In infrastructure since January 2025, spent going up the stack rather than along it.

Nothing above works

if the layer below it is wrong

What I do

What I actually do

In production now

Open-weight models, serving real traffic

Multi-GPU, multi-model vLLM containers behind the company’s applications. Each model its own container, sharing a tensor-parallel GPU pool with an explicit memory fraction rather than a default that works right up until two models are resident.

One of these was blocked outright by RHEL 9 dependency conflicts. Rebuilding it on a custom Ubuntu base made it execute consistently everywhere instead of on one machine somebody had hand-fixed, which is the difference between a demo and a deployment.

Inference clusters in production
5+Inference clusters in production
Tensor-parallel, memory pinned
Multi-GPUTensor-parallel, memory pinned
Open-weight, self-hosted
Gemma · Qwen CoderOpen-weight, self-hosted
APPLICATION TIERGATEWAYINFERENCE CONTAINERSGPU POOLProduct APIsynchronousBatch workerqueuednginxTLS · routingper-model pathrate limitvLLM · gemmachat / summarisationvLLM · qwen-codercode assistancevLLM · instruct-4bclassificationGPU 0GPU 1GPU 2GPU 3tensor-parallelmem fraction pinned/metrics → Prometheuslatency · queue depth · GPU util
Fig. 1. Multi-model inference topology. Each container serves one open-weight model, shares a tensor-parallel GPU pool with an explicit memory fraction, and exports its own latency and queue-depth metrics.

Why me

Why the bottom-up route matters

Selected work

Work that changed something

Built independently

Built because I needed it

Also public: CyberSentinel (web vulnerability scanner, Python) and Proxmox on a laptop over Wi-Fi (a write-up). All repositories →

Experience

The record

Oct 2025 – present

Gandhinagar

System Engineer · Knovos

Infrastructure and systems engineering across a multi–data-centre cloud and on-premise environment.

  • Replaced a commercial monitoring product with a self-hosted multi-region platform built on Prometheus, Grafana, Alertmanager, Blackbox Exporter and nginx with TLS, eliminating the associated licence cost.
  • Instrumented 250+ Windows hosts with Windows Exporter and deployed Prometheus in Agent mode per environment, producing a multi-node architecture that scales beyond a single scrape target.
  • Delivered OCI infrastructure for multiple data centre migration projects: compute, VCN topology, Security Lists, NSGs and site-to-site IPSec VPN, provisioned entirely as Terraform.
  • Integrated LDAP and Postgres into the monitoring stack and implemented Teams and email alert routing, with dashboards covering both cloud and on-premise environments.

Jan 2025 – Oct 2025

Vadodara

System Administrator · WovV Technologies

Server administration, virtualisation, and network and endpoint security.

  • Administered CentOS, Ubuntu and Windows Server environments across office and cloud environments.
  • Managed identity and access through Active Directory, covering accounts, permissions and group policy.
  • Operated VMware ESXi hosts, including virtual machine lifecycle and resource optimisation.
  • Automated recurring administrative tasks with Bash and batch scripting, reducing manual handling and error rates.

Jun 2024 – Jan 2025

Vadodara

Customer Service Specialist · Patterns 247

Client operations support for US real-estate professionals, covering lead management and CRM administration.

  • Managed client data and daily operations across Salesforce and Propertybase.
  • Built and maintained real-estate websites using WordPress and no-code platforms.
  • Recognised as Best Performer of the Month, September 2024.

Certifications

  • Oracle Cloud Infrastructure Foundations Associate · Oracle
  • CyberOps Associate · Cisco
  • Cyber Security and Privacy · Certification
  • Certified Kubernetes Administrator · CNCFin progress

Education

  • Master of Computer Application, Computer Science · ITM Gwalior
  • Bachelor of Computer Application, Computer Science · Jiwaji University

Questions

Before you ask

The things people want to know before sending a first message.

That is a lot of claims for someone early in their career. Why believe them?

Check them rather than believe them. The page is built so you can. Every case study states whether I owned it, led it or contributed to it, because a page that is vague about that on its strongest claims is a page rounding up. The platform serving this page is mine end to end and open source, including its decision log and the incidents I caused, so you can read what broke and what I changed afterwards. The honest summary is that the ground covered is unusual for the time and the seniority is not: I have gone up the stack rather than along it, and I would rather you judge that from the record than from an adjective.

How do you actually use AI tooling?

Daily, and as an engineering tool rather than a novelty. Claude Code and Codex are how I read a codebase I have not seen before, turn an undocumented application into something with a working Dockerfile, and generate the tests and validation steps that prove a change before it ships. It compresses the typing and the research; it does not remove the reviewing, and the turnaround is short because verification is part of the task rather than a later phase.

What kind of problems do you want to be working on?

The ones where the technology is new enough that nobody has written the runbook yet. Getting open-weight models into production was that a year ago and is close to routine now. The next one is probably the platform underneath it: making AI-assisted teams able to ship without waiting on a person, which is a harder problem than the models are. I would rather be early to a layer and on the infrastructure side of it, where being wrong shows up as a broken deployment instead of a slide.

What does bottom-up actually mean, in practice?

That I did the layers in order. Physical hardware and hypervisors, then Windows and Linux administration, then applications, then containers and orchestration, each one on the job rather than from a course. It matters because most production failures are not inside a layer, they are between two, and diagnosing those needs somebody who has worked on both sides.

Are you comfortable with both Windows and Linux?

Yes, and that combination is less common than it should be. Active Directory, Group Policy and PowerShell on one side; systemd, containers and Bash on the other. A large part of the automation work has been making one managed fleet out of both, instead of two separate ones with a person in the middle.

Where are you based, and can you work outside India?

Gandhinagar, Gujarat, India. The current work is already spread across data centres and time zones, so remote is normal rather than an adjustment. For a role outside India I would need visa sponsorship.

Is any of the work public?

The platform serving this page is open source, including its decision log, the incidents and the fixes for them. Client work is not public, which is why the case studies above describe outcomes and architecture rather than internals.

Where this is going

AI will equip one engineer with the output of a team.

What decides whether any of it reaches a user is the layer underneath.

The bottleneck is moving. Writing the service stops being the hard part when a model can draft it, review it and test it. The hard part becomes everything that has always been hard: where it runs, what it talks to, who is allowed to reach it, what happens at three in the morning when it stops.

That layer does not get easier because the layer above it got faster. It gets busier. More services, deployed more often, by people who did not build the platform underneath them and should not have to.

So this is the work I want: building the floor that AI-accelerated teams stand on, and running the inference that makes them faster in the first place. I came up through hardware, hypervisors and operating systems to get here, which is the reason I know what the floor is made of.

Get in touch

Neeraj Jeswani

I work on the layer between an application and the infrastructure it runs on: turning a service into something a team can deploy for themselves, and servers into something nobody has to log into by hand. If that is the problem in front of you, email is the fastest way to reach me.

This page is served by a platform I designed, built and operate end to end. TLS, routing, monitoring and alerting on a single free-tier ARM instance. Two cores, and most of the work is making them behave like more.