Targeting Swiggy, Razorpay, PhonePe, Meesho, or one of the 200+ startups in Koramangala? Their DevOps interviews are different. They don't ask "What is Docker?" — they ask "Design a deployment system for 10 million daily orders with zero downtime." They want to know how you'd handle a production incident at 2 AM when the database is down and customers can't place orders. This training is built specifically for Koramangala's startup interview patterns: system design, incident response, on-call culture, and the fast-paced GitOps workflows that product companies use. 100% live online, no Sarjapur Road traffic required.

Swiggy's DevOps interview: "We have 50,000 delivery partners online. Database goes down. Walk me through your incident response." Razorpay's system design round: "Design a payment processing system that handles 10,000 transactions per second with 99.99% uptime." PhonePe's technical round: "Explain how you'd implement canary deployments for a critical payment API." These aren't textbook questions — they're real scenarios from our alumni who got offers. We practice these exact patterns every week: whiteboard system design, simulate production incidents, debug Kubernetes clusters under time pressure. By week 6, you'll be able to walk into a Koramangala startup interview and confidently design systems on a whiteboard while explaining your trade-offs.
Training designed for Swiggy, Razorpay, PhonePe interview patterns — system design, incident response, on-call culture, not just tools.
Koramangala startups ask: "Design a zero-downtime deployment for 10M users." We practice these scenarios weekly with real architecture diagrams.
Learn GitOps, feature flags, canary deployments, and chaos engineering — the tools Koramangala product companies actually use daily.
Startup interviews are brutal. Small batch means we can drill system design questions 1-on-1 until you can whiteboard confidently.
Where we have alumni at a startup and they are willing to refer, qualified students get internal referrals after completing the projects.
Simulate real production incidents: database failover, pod crashes, traffic spikes. Learn how to debug under pressure like Swiggy SREs do.
Everyone says Koramangala when you ask where the startups are. Then you ask which ones, and you get four logos.
Trouble is, there are four different kinds of company inside those two kilometres, and they run four different interviews. Preparing for the wrong one is how decent candidates lose rounds they should have taken.
Late-stage first. Swiggy, Razorpay, PhonePe. Platform teams that have existed for years, an SRE track with its own ladder, a loop that's been run across enough hires that there's a rubric sitting behind the questions. The system design round here is structured and somebody is scoring you against criteria you can't see. You'll spot these in the JD long before you apply: error budgets, SLOs, internal tool names nobody has bothered to expand, and separate postings for "platform engineer" and "SRE".
Growth stage next. Meesho tier. Two platform engineers holding up an engineering org of eighty, which is why the JD lists eleven tools and says "own it end to end" somewhere in the middle. Breadth wins here. Deep in one thing and blank everywhere else scores worse than solid across the board — the opposite of what the late-stage loop rewards, which is why the same CV does well at one and badly at the other.
Then seed and Series A. There are hundreds of these, more than the other three groups put together. Nobody owns infrastructure. The CTO has been doing it at night for a year. You'd be the first person hired to take it off them. Look for setup verbs: set up CI, set up monitoring, set up staging. Nobody is asking you to design for ten million orders a day. They'll ask what you'd do in your first thirty days, and at some point someone will ask whether you can stop the AWS bill doubling again.
The fourth group never gets written about. Koramangala's fintech and payments cluster interviews like a bank, and the sticker-covered laptops are misleading. RBI expectations on data localisation, audit trails and change management don't relax because you raised a Series A instead of holding a banking licence. So the round turns into approval workflows, segregation of duties, who signs off on a production change and where that's recorded. PCI-DSS anywhere in the description, an audit cycle mentioned in passing, sign-off language — that's this group, and general startup prep will not carry you through it.
Read the description, not the logo.
The page claims 40% of class time goes on system design. Fair enough. Here's one worked all the way through so you can decide whether that's worth anything.
"Design a food delivery order system for ten million orders a day."
Weak candidates start drawing. Strong ones spend four minutes asking questions, and the scoring has already begun by then.
Do the arithmetic out loud. Ten million a day is roughly 115 orders a second averaged out, except food delivery has no average — lunch and dinner carry nearly all of it. Peak is eight to ten times the mean, so call it a thousand writes a second at the top and design for that instead.
Then ask about reads, because that's where the load actually sits. Every customer with a live order is refreshing a tracking screen. Every delivery partner's phone is pushing GPS every few seconds. This isn't a ten-to-one read-write ratio, it's closer to a thousand to one, and that single question changes the shape of everything you're about to draw.
Budgets after that. What's acceptable latency on placing an order? Propose 500ms and be ready to defend it. A tracking update? Two to five seconds of staleness is fine, and saying so tells the interviewer you know not everything has to be fresh — a distinction a surprising number of candidates never make out loud.
Consistency. The order state machine has to be strongly consistent, because double-charging or double-dispatching is a money bug and it lands in someone's support queue within minutes. The ETA doesn't need it. Eventually consistent is fine and no user will ever notice.
Now draw. Gateway, order service, assignment, tracking, notifications. Orders in a relational store, because you want transactions wrapped around the state transitions. Location writes go somewhere built for that volume — a time-series store or a Redis geospatial index — and nowhere near the same database as the orders.
Pick one box and go deep. Take the order state machine, because that's where the hardest problem lives: idempotency. Payment gateway times out. Client retries. You now have to guarantee exactly one order and one charge exist, not two. Talk about idempotency keys, where they live, how long you keep them, and what happens when the key store itself is unavailable, which is the follow-up question.
Failure modes, which is where most candidates run out of clock because they spent eleven minutes drawing. Assignment service down: do you accept the order and queue it, or refuse at the door? Either answer is fine. Not having one isn't. Payment provider slow but not down is the harder case and much more common — have something for it.
Finish on a trade-off you made on purpose. "I'll accept orders during assignment outages and reconcile later, which means some get cancelled twenty minutes in, and I'd rather cancel than refuse." That's the sentence they remember.
"Fifty thousand delivery partners are online. The primary database goes down. Walk me through your response."
Nearly everyone opens with the fix. That's what the interviewer is waiting for, and it's the wrong first move.
Acknowledge the page and open a channel. Say this out loud rather than assuming it's implied. Ack within a minute, call a severity, name an incident commander even if that's you and you're on your own, start a log. In a real one, three engineers will start debugging in parallel and repeat each other's work for twenty minutes unless somebody is running it. Mention that and the interviewer knows you've sat through one.
Then scope, before you touch anything.
"The database is down" isn't a scope, it's a headline. Primary gone, or one replica? Are writes failing while reads still serve off a replica? Can partners still accept orders they've already been assigned, or have they only lost history? One region or all of them? How aggressive you're allowed to be depends entirely on the answers, and the two minutes you spend here are what stop you running the command that makes it worse.
Stabilise before you diagnose. That's the line that separates people who've held a pager from people who've read about holding one. Fail over. Shed non-critical load. Serve stale from cache. Flag off anything that writes. You are not trying to understand the failure yet — you're trying to get the affected number down, and if you promote a standby and never learn what killed the primary, that's still a good incident.
Comms on a clock. Every fifteen to twenty minutes, news or no news. The hard part is what you write when you know nothing, and there's a template for it: what's affected, what isn't, what you're doing, when the next update lands. "Order placement failing for partners in the south region. Login unaffected. Failing over to standby, next update 14:20." What you don't do is post "looking into it" and vanish for forty minutes. That's the most common real failure in real incidents and it isn't a technical one.
Root cause afterwards. Postmortem with nobody's name in it.
One thing to carry in: nobody is testing whether you know the failover syntax. They're testing whether you'll stay calm and keep people informed while it's on fire, because that's the part they can't teach you after you join.
Everything above is about getting the job. This part is about whether you should take it.
Rotations at Indian startups usually run weekly, primary and secondary, handover on a fixed day. None of that is the number that matters. The number that matters is how many people are in the rotation. Six to eight and your turn comes round every six to eight weeks, which is liveable. Three isn't a rotation, it's a burnout with a schedule attached, and you're on call for a third of your life. Ask that before you ask about pay.
Compensation is all over the place. Flat weekly allowance at some companies, comp-off at others, and a fair number of early-stage places pay nothing at all and treat it as part of the role. None of those is disqualifying by itself. What you don't want is to discover it in month two. So ask, and then ask the second question, which is whether it's written down anywhere or whether it's just what your manager currently does.
Ask what handover looks like as well. With no handover, every incident starts from zero and the person coming on has no idea a database has been throwing intermittent errors since Wednesday. Fifteen minutes and a written summary is normal. If the answer is "you just check the channel", that tells you how the rest of it runs.
Page volume is the honest signal. Zero to two actionable pages across a full week is healthy. Being woken most nights isn't a rough patch, it's an organisation that worked out alert fatigue was cheaper than fixing the alerts. Ask how many fired last month and how many were real. If they know the number, that's a good sign in itself. If nobody has any idea, that's also an answer.
Worth asking, roughly in this order: how many engineers are in the rotation. How many pages last month and how many were actionable. Whether there's a runbook for the top five alerts and when it was last updated. What happens after a bad night — comp-off, a late start, nothing. Who's allowed to page you and whether support can escalate straight through. Whether postmortems get written, and whether they're blameless in practice or only in the template.
Ask all of it. Candidates who do read as senior, not as difficult.
ArgoCD, Flux, feature flags, canary, blue-green. Five names on a syllabus, which tells you nothing until you know what each one removes and when it isn't worth having.
GitOps first. Normally a pipeline reaches into your cluster from outside and runs kubectl apply. GitOps flips that around: a controller sits inside the cluster, watches a git repo, and keeps comparing what's running against what the repo says should be running. Where they differ it either corrects the cluster or makes a lot of noise. Deploy credentials never leave the cluster, and git history becomes your audit trail for every production change.
Here's the failure it exists to prevent. Tell it this way in an interview — it's a story rather than a definition, and it lands better.
Two in the morning, checkout is down. Somebody edits the live deployment to raise a memory limit. It works, everyone goes back to bed, nobody writes it down. Three weeks later an unrelated pipeline run applies the manifest from the repo, the memory limit quietly goes back to what it was, and checkout falls over again at the next dinner peak. With a controller reconciling, that edit shows up as drift inside a minute and gets either reverted on the spot or surfaced loudly enough that someone has to go and commit it properly.
Repo structure does more for you than the choice between ArgoCD and Flux ever will. Split app code from deployment config so a config change doesn't kick off an image build. In the config repo, one base directory with the shared manifests and thin per-environment overlays that patch only what genuinely differs — replicas, resource limits, ingress hosts. If staging and prod are two full copies, they've already diverged and nobody has noticed yet.
Progressive delivery gets bundled in with this constantly. Separate idea. Feature flags split deploying code from releasing behaviour, so it ships dark and you switch it on for one percent of users. Canary moves a slice of real traffic onto the new version and watches error rates before moving more. Blue-green runs two complete environments and flips between them, which is easy to reason about and twice as expensive.
Then the honest bit. A five-person startup with one cluster and two services should not be running ArgoCD. The reconciliation loop, the repo split, the review overhead — all of it costs more than it saves at that size and a plain pipeline is the right call. Saying that in an interview, and being able to say roughly where it stops being true, does more for you than having the tool listed on your CV.
If you want the mechanism in depth — reconciliation, drift, Argo CD versus Flux, and how canary and blue-green actually get wired up — read GitOps explained.
Say "chaos engineering" and most people picture somebody killing production boxes at random to see what falls over. That's the definition doing the rounds. It's wrong, and repeating it in an interview will cost you.
It's an experiment. You start from a hypothesis, written precisely enough that it can come out false. Something like: kill one of three payment service pods under peak load, and checkout p99 stays under 800ms with the error rate under 0.1%. That's a claim about your system you currently believe and have never actually checked.
Blast radius before anything else. One pod, not one node. Staging first. If it ever goes near production, a thin traffic slice, an abort condition, and somebody actually watching the dashboard rather than watching Slack. Pick a window when the team is awake and the business is quiet, and make sure you can stop it in one command.
Then run it and measure. Hypothesis holds, you now have evidence for a specific behaviour instead of a belief about it. Hypothesis fails, you found the bug in daylight with the whole team watching rather than at three on a Friday morning. Both outcomes pay for the afternoon.
One prerequisite people skip: you need working observability before any of this is worth doing. If you can't see p99 latency and error rate per service on a dashboard, the experiment produces no data and you've just broken something for no reason. Get the metrics first. That's usually the real work, and chaos is what you do once it's in place.
Litmus and Chaos Mesh handle the injection. That's the easy part and it isn't the skill.
A caution on how you talk about it. Almost nobody in Indian startups runs chaos experiments as ongoing practice, so claim deep production experience and you'll get probed until it falls apart. Better: run one experiment on your own cluster, write down the hypothesis, the radius, what you measured and what actually happened, and bring that page in with you. Somebody who's run one honest experiment and can say what surprised them beats somebody who's read every Netflix engineering post.
That's what the interviews test. Here's what to have in hand before you sit in one. Five things, roughly in order of return.
A deployed thing with a URL a stranger can open. Not a repo — something running on a real domain with a certificate, however small it is. This clears the first filter, which is just: has this person ever pushed anything past localhost.
A GitOps repo where the cluster genuinely reconciles. Managed Kubernetes, ArgoCD or Flux, base and overlays, and a moment you can demonstrate where you change a number in git and watch the cluster follow it. Record thirty seconds of drift being corrected. That clip works harder in an interview than any line about it on your CV.
An incident write-up from something you broke on purpose. Fill the disk. Kill the database. Exhaust the connection pool and watch what it does downstream. Then write it up properly: timeline, impact, how you detected it, what you tried that didn't work, what did, what you'd change. Two pages. Hardly anyone walks in with one of these, and it's the most convincing thing on this list because it's proof you've been on the receiving end of your own decisions instead of reading about someone else's.
A load test with a before and an after. k6 or Locust, point it at your service, find where it breaks, change one thing, run it again. The throughput number isn't the point. The point is being able to say it fell over at 400 requests a second on connection pool exhaustion, you resized the pool and added a timeout, and now it falls over at 1,100 for an entirely different reason.
A design doc for something you built yourself. One page. Requirements, components, why that datastore, and a trade-offs section where you name what you gave up. It's the cheapest possible practice for the system design round above and it takes an afternoon.
Three and five are the rare ones, and both take an afternoon. Most people skip them because neither produces anything you can show off in a demo, which is exactly why having them puts you in a different pile.
Put all five behind one link. A single public repo with a README that points at the live URL, the write-up, the load test results and the design doc, in that order. Recruiters spend under a minute on this. Five scattered repos with no index gets you credit for none of it.
No salary figures here. Anything we published would be stale within a quarter and you'd have no way to check it against reality, so what follows is method instead.
The ESOP grant. It gets presented to you as a rupee number. It isn't one. Five things make it possible to value at all, and you can ask for all five in writing without anybody thinking less of you: how many options, the total shares outstanding rather than a percentage, the strike price, the vesting schedule and the cliff, and the exercise window after you leave.
Ask for the count, not the percentage. Percentages get rounded and they move every time the company raises.
That last item catches people. Ninety days post-departure means that if you resign at year three, you have three months to find the cash, buy the options, and pay tax on a gain that exists only on paper, for shares you may not be able to sell for years afterwards. Most people find this out at exit, which is too late to negotiate it.
Ask about buybacks too — has the company ever run one, and when. A company that has put a liquidity event in front of employees has proved the paper can turn into money. One that hasn't, hasn't. Neither is a reason to walk away. But in the second case, treat the grant as a lottery ticket and not as pay. Most grants end up worth nothing. Assume yours will and let yourself be wrong about it.
Startup band versus GCC band. Same title, and the capability centre usually offers a higher cash floor with less variance and a smaller equity component. The clean comparison: value the equity at zero, put cash against cash, then ask whether the option upside covers the gap you just measured. If it doesn't, the offer isn't competitive, however the grant is written up in the email.
The switch premium. Internal hikes get budgeted as a percentage of what you're already on. External offers get priced against what the market pays for the skill today. Those two numbers drift, and the gap compounds, which is why moving has historically paid better than staying and why someone six years in finds out they're behind a colleague who joined in March. Worth understanding before you sit down with a counter-offer.
Four people shouldn't enrol. Better said here than after we've taken the fee.
If you want a job guarantee. We don't have one, and the referral network isn't one wearing a different hat. A referral gets you the interview. It is not an offer, and what happens in the room after that is yours. If a guarantee is what decides it for you, there are institutes that will put one in writing — read what it actually commits them to before you sign anything.
If you've never used a terminal. This assumes you can move around a Linux filesystem, tail a log, edit a file over SSH and explain what a process is. If that isn't true yet, do three weeks of Linux fundamentals and then come back. Starting without it means ten weeks spent a step behind everyone else, and we'd be charging you to watch that happen.
If you want recordings. Live cohort, capped at fifteen, built around being interrupted and corrected while you're getting something wrong. If your schedule genuinely can't hold a fixed weekly slot, a video subscription is the better product and costs a fraction of this.
If you can't protect five hours a week outside class. The system design practice is the thing that gets you through the interview and it doesn't happen passively. Five hours, weekly, drawing and being pulled apart. Not five hours of watching. Without them this is ten weeks of interesting listening followed by the same interview outcome you're getting now, and we've had people finish the course that way. It's avoidable and it's the one failure mode that's entirely in your hands.
Koramangala startups like Swiggy, Razorpay, and PhonePe focus heavily on system design and incident response in interviews. We dedicate 40% of training time to system design practice (design payment systems, food delivery platforms, etc.), production incident simulations, and whiteboard problem-solving. You also learn GitOps, chaos engineering, and on-call workflows that product companies use — not just basic Docker and Kubernetes.
Where we have alumni at a company and they are willing to refer, yes. After completing all projects and passing our internal system design mock interview, qualified students get internal referrals through our alumni network. A referral gets you the interview; it is not an offer.
We cover: designing scalable food delivery systems (Swiggy pattern), payment processing with 99.99% uptime (Razorpay pattern), real-time order tracking (Dunzo pattern), and high-throughput API gateways. Each week includes whiteboard practice where you design systems and explain trade-offs — exactly what Koramangala startup interviews test.
Course fee is 35,000 (all inclusive). Classes run on Saturday and Sunday, 3 hours each, 100% live online, and a new batch starts every Sunday. Limited to 15 students for personalized system design coaching.
Join our Bangalore cohort. A new batch is starting soon.
In your free demo you will build a live CI/CD pipeline from scratch — not watch a slideshow. It takes 45 minutes and shows you exactly what the full course feels like.
45 minutes, live with the instructor. No payment required.
We call you back within 24 hours. No spam.
Our students work at