Microservices (Java/Spring) · میکروسرویس سنیورSenior ~49 دقیقه مطالعه~42 min read

کتابِ راهِ سنیور: نکاتِ واقعیِ پروداکشنThe Senior Playbook: Real-World Production Wisdom

تفاوتِ سنیور با میدلِ باتجربه در دانستنِ فرمول‌ها نیست، در قضاوت است؛ این فصل آنتی‌پترن‌های میکروسرویس، مشاهده‌پذیریِ واقعیِ پروداکشن، آن‌کال و ظرفیت و افتِ باوقار را طوری می‌آموزد که نیم‌سال تجربه را به قضاوتِ چندساله تبدیل کند.The gap between a senior and an experienced mid-level engineer is judgment, not formulas; this chapter teaches microservice anti-patterns, real production observability, on-call, capacity, and graceful degradation so half a year of experience turns into years of senior judgment.

پیش‌نیاز:Prerequisites: مدیریتِ داده: Database-per-Service، Saga، Outbox و CQRSData Management: DB-per-Service, Saga, Outbox, CQRSارتباطِ سرویس‌ها و مسیرِ کاملِ یک درخواستService Communication & the Full Request Path


بذار با یک حقیقتِ ناخوشایند شروع کنیم: بیشترِ چیزهایی که در توتوریال‌ها و کورس‌ها یاد گرفتی، در پروداکشن یا اشتباه است یا ناقص. نه به این خاطر که مدرّس دروغ گفته، بلکه چون توتوریال «مسیرِ خوش‌بینانه» (happy path) را نشان می‌دهد و پروداکشن جایی است که همه‌چیز هم‌زمان خراب می‌شود: شبکه کند می‌شود، دیتابیس تحتِ فشار قفل می‌کند، یک سرویسِ downstream تایم‌اوت می‌دهد، یک deploy نیمه‌کاره می‌ماند، و درست همان لحظه پیجِرت (pager) نصفِ شب زنگ می‌زند.

سنیور کسی نیست که همه‌ی الگوها را حفظ باشد. سنیور کسی است که وقتی تیم هیجان‌زده می‌گوید «بیا این را میکروسرویس کنیم»، می‌پرسد «چرا؟ چه مشکلی را حل می‌کند؟ هزینه‌اش را دیده‌ای؟» — و وقتی همه می‌گویند «مونولیت بد است»، آرام می‌گوید «مونولیتِ خوب از میکروسرویسِ بد بهتر است.»

این فصل قرار نیست بهت الگوی جدید یاد بدهد. قرار است بهت قضاوت بدهد: کِی چیزی خوب است، کِی همان چیز فاجعه است، و چطور این‌ها را در اینترویو و در جلسه‌ی طراحی بلند و روشن بیان کنی.

نقشه‌ی راه این فصل

۱) آنتی‌پترن‌ها — distributed monolith، chatty services، shared DB، nano-services، entity services، و «کِی میکروسرویس اشتباه است». ۲) نگرانی‌های پروداکشن — مشاهده‌پذیریِ واقعی (correlation ID، RED/USE، tracing)، مبانیِ on-call و incident، ظرفیت و back-pressure، افتِ باوقار (graceful degradation)، هزینه، و مهاجرتِ داده بینِ سرویس‌ها. ۳) اینترویوِ system design — چطور بلند فکر کنی. ۴) چک‌لیستِ آمادگیِ سنیور. همه با کدِ واقعیِ Spring Boot و دیاگرام.

نسخه‌ها (تیرماهِ ۱۴۰۵ / اواسط ۲۰۲۶)

مثال‌ها بر پایه‌ی Spring Boot 3.5.x (سریِ 3.5.0 اردیبهشتِ ۲۰۲۵ تا 3.5.7 مهرِ ۲۰۲۵) و Spring Cloud 2025.0.x «Northfields» هستند. سریِ بعدی، Spring Cloud 2025.1 «Oakwood» روی Spring Framework 7 و Spring Boot 4 نشسته. برای resilience از resilience4j-spring-boot3 نسخه‌ی 2.2.x و برای tracing از Micrometer Tracing با پُلِ Brave یا OpenTelemetry استفاده می‌کنیم. اصول این فصل مستقل از نسخه‌اند؛ نسخه‌ها فقط برای کدِ دقیق مهم‌اند.


بخش یک: آنتی‌پترن‌ها — جایی که تیم‌ها زمین می‌خورند

قبل از این‌که یاد بگیری چه کاری بکنی، باید بدانی چه کاری نکنی. هر آنتی‌پترنِ زیر را من در پروداکشنِ واقعی دیده‌ام که تیم را ماه‌ها عقب انداخته. اسمشان را یاد بگیر؛ چون در اینترویو وقتی می‌توانی یک تصمیمِ بد را با اسمِ درست نام‌گذاری کنی، سنیور به نظر می‌رسی.

۱) مونولیتِ توزیع‌شده (Distributed Monolith)

یک تیمِ فوتبال با پابندِ مشترک

تصور کن یازده بازیکن داری، ولی مچِ پای هر بازیکن با طنابی به بغل‌دستی‌اش بسته است. ظاهراً یازده نفرِ مستقل‌اند، اما هیچ‌کدام نمی‌تواند بدون این‌که بقیه هم‌زمان حرکت کنند، یک قدم بردارد. این دقیقاً distributed monolith است: به ظاهر چند سرویسِ جدا، ولی چنان به هم گره خورده‌اند که برای هر تغییر باید همه را با هم دیپلوی کنی.

مونولیتِ توزیع‌شده بدترین حالتِ ممکن است: هزینه‌ی عملیاتیِ میکروسرویس (شبکه، دیپلویِ جدا، دیباگِ توزیع‌شده) را می‌پردازی، ولی هیچ‌کدام از فوایدش (استقلالِ دیپلوی، مقیاس‌پذیریِ مستقل، ایزوله‌شدنِ خطا) را نمی‌گیری. نشانه‌ها:

  • برای رها کردن (release) یک فیچر، مجبوری چند سرویس را با یک ترتیبِ مشخص و هم‌زمان دیپلوی کنی.
  • تغییرِ یک فیلد در یک سرویس، بیلدِ سه سرویسِ دیگر را می‌شکند.
  • سرویس‌ها به صورتِ synchronous زنجیره‌ای همدیگر را صدا می‌زنند: A منتظرِ B، B منتظرِ C، C منتظرِ D.
  • همه از یک shared library مشترک برای مدل‌های داده استفاده می‌کنند و آپدیتِ آن library یعنی آپدیتِ همه.
پابندِ نسخه‌ی shared model

شایع‌ترین ریشه‌ی distributed monolith، یک ماژولِ مشترکِ common-domain است که تمام سرویس‌ها به آن وابسته‌اند. لحظه‌ای که Order و User و Product در یک jar مشترک تعریف شوند، دیگر هیچ سرویسی نمی‌تواند مدلش را مستقل تغییر دهد. اشتراکِ قرارداد (contract/schema/API) خوب است؛ اشتراکِ کدِ دامنه سم است. هر سرویس باید DTOهای خودش را داشته باشد، حتی اگر تکراری به نظر برسند. تکرارِ کنترل‌شده از اتصالِ (coupling) پنهان بهتر است.

راهِ درست: مرزها را حولِ قابلیت‌های کسب‌وکار (business capabilities) بکِش، نه حولِ لایه‌های فنی. ارتباطِ بینِ سرویس‌ها را تا حدِ ممکن asynchronous و مبتنی بر event کن تا زنجیره‌ی synchronous شکسته شود.

بیایید تفاوتِ توپولوژی را ببینیم.

نمودار: مونولیتِ توزیع‌شده (زنجیره‌ی همگام) در برابر سرویس‌های مستقلِ رویدادمحور — Distributed monolith vs. event-driven independence.

flowchart LR
  subgraph Bad[Distributed Monolith]
    A1[Order] --> B1[Inventory]
    B1 --> C1[Pricing]
    C1 --> D1[Notification]
  end
  subgraph Good[Event-Driven]
    A2[Order] -- OrderPlaced --> K[(Event Bus)]
    K --> B2[Inventory]
    K --> C2[Pricing]
    K --> D2[Notification]
  end

در سمتِ چپ، تاخیرِ کل برابرِ جمعِ تاخیرِ همه است و اگر Notification بیفتد، کلِ ثبتِ سفارش می‌افتد. در سمتِ راست، Order فقط یک event منتشر می‌کند و بلافاصله پاسخ می‌دهد؛ بقیه در زمانِ خودشان مصرف می‌کنند و خرابیِ یکی، بقیه را زمین نمی‌زند.

۲) سرویس‌های پرحرف (Chatty Services) — همان N+1، این‌بار روی شبکه

تلفنی که برای هر کلمه قطع و وصل می‌شود

فرض کن می‌خواهی یک جمله‌ی ده‌کلمه‌ای به دوستت بگویی، ولی به‌جای یک تماس، ده بار زنگ می‌زنی و هر بار فقط یک کلمه می‌گویی و قطع می‌کنی. هزینه‌ی «برقراریِ تماس» ده برابر می‌شود. سرویسِ chatty دقیقاً همین است: برای ساختنِ یک پاسخ، ده‌ها بار سرویسِ دیگر را صدا می‌زند.

در مونولیت، فراخوانیِ یک متد چند نانوثانیه است. روی شبکه، هر فراخوانی صدها میکروثانیه تا چند میلی‌ثانیه است — هزاران برابر گران‌تر، و هر کدام می‌تواند fail شود. کدِ زیر بی‌گناه به نظر می‌رسد ولی در پروداکشن قاتل است:

// آنتی‌پترن: N+1 روی شبکه
List<Order> orders = orderClient.findRecent(userId);      // ۱ فراخوانی شبکه
for (Order o : orders) {
    // برای هر سفارش، یک فراخوانی جدا — اگر ۵۰ سفارش باشد، ۵۰ round-trip
    Product p = productClient.getById(o.productId());     // N فراخوانی شبکه
    o.setProductName(p.name());
}

اگر هر فراخوانی ۲۰ms و لیست ۵۰ آیتم باشد، فقط همین حلقه ۱ ثانیه تاخیر اضافه می‌کند — و این خوش‌بینانه است. راهِ حل، طراحیِ API برای batch و پاسخِ درشت‌دانه (coarse-grained) است:

// درست: یک فراخوانی batch به‌جای N
List<Order> orders = orderClient.findRecent(userId);
Set<Long> ids = orders.stream().map(Order::productId).collect(toSet());
Map<Long, Product> products = productClient.getByIds(ids);   // ۱ فراخوانی
orders.forEach(o -> o.setProductName(products.get(o.productId()).name()));
قضاوتِ سنیور: مرزِ سرویس = مرزِ گفتگو

یک راهنمای عملی: اگر دو سرویس مدام و پرحجم با هم حرف می‌زنند (chatty)، احتمالاً نباید دو سرویس باشند. مرزِ درستِ سرویس جایی است که ترافیکِ بینِ دو طرف کم و درشت‌دانه باشد، و ترافیکِ درونِ هر طرف زیاد. این همان اصلِ «high cohesion, low coupling» است، ولی این‌بار با هزینه‌ی شبکه اندازه‌گیری می‌شود. اگر مجبوری برای یک عملیات ۵ سرویس را در یک ترنزکشن هماهنگ کنی، مرزهایت اشتباه است.

۳) دیتابیسِ مشترک (Shared Database)

یک دفترچه‌ی مشترک برای کلِ خانواده

تصور کن کلِ خانواده در یک دفترچه‌ی حساب می‌نویسند. تا وقتی همه مؤدب‌اند خوب است؛ ولی روزی یک نفر ستونِ «تاریخ» را عوض می‌کند و ناگهان محاسباتِ بقیه به‌هم می‌ریزد، بی‌آن‌که بداند کِی و چرا. دیتابیسِ مشترک بینِ میکروسرویس‌ها همین است: هیچ‌کس مالکِ schema نیست، پس هیچ‌کس نمی‌تواند بی‌ترس تغییرش دهد.

قانونِ طلاییِ میکروسرویس: هر سرویس، دیتابیسِ خودش. نه فقط جدولِ جدا، بلکه ترجیحاً schema یا instanceِ جدا، و هیچ سرویسی مستقیم به جدولِ سرویسِ دیگر دست نمی‌زند. چرا این‌قدر سخت‌گیری؟

  • استقلالِ تکامل: اگر سرویس A بتواند به جدولِ B کوئری بزند، آن‌گاه B هرگز نمی‌تواند ساختارِ جدولش را عوض کند بدونِ ترسِ شکستنِ A. schema به یک API عمومیِ ناخواسته تبدیل می‌شود — بدترین نوعِ coupling، چون پنهان است.
  • مالکیتِ داده: قوانینِ کسب‌وکار (invariants) باید در یک جا اعمال شوند. اگر دو سرویس هر دو در جدولِ account بنویسند، منطقِ «موجودی نباید منفی شود» در دو جا تکرار می‌شود و روزی یکی از آن‌ها فراموش می‌شود.
وسوسه‌ی JOIN

سخت‌ترین چیزی که هنگامِ جدا کردنِ دیتابیس‌ها از دست می‌دهی، JOIN است. تیم‌ها اغلب زیرِ فشارِ «فقط یک کوئریِ گزارش‌گیری» دوباره دیتابیس‌ها را به هم وصل می‌کنند و همه‌چیز خراب می‌شود. راهِ درست برای گزارش‌گیریِ چند-سرویسی، ساختنِ یک read model جدا (CQRS) است که با مصرفِ eventها پُر می‌شود، یا یک data warehouse/lake که سرویس‌ها داده را به آن push می‌کنند — نه JOIN مستقیم روی دیتابیسِ عملیاتیِ سرویس‌ها.

برای گزارش‌گیری بینِ سرویس‌ها، الگوی رایج این است که هر سرویس رویدادهایش را به یک استخرِ داده می‌فرستد. مثالِ کوئریِ گزارشی روی یک read model که با هر دو دیالکت کار می‌کند:

-- PostgreSQL: صفحه‌بندی روی read model
SELECT order_id, product_name, total
FROM   reporting.order_summary
WHERE  created_at >= now() - interval '7 days'
ORDER  BY total DESC
LIMIT  20 OFFSET 40;
-- Oracle: همان کوئری، صفحه‌بندیِ استاندارد
SELECT order_id, product_name, total
FROM   reporting.order_summary
WHERE  created_at >= SYSTIMESTAMP - INTERVAL '7' DAY
ORDER  BY total DESC
OFFSET 40 ROWS FETCH FIRST 20 ROWS ONLY;
تفاوتِ دیالکت که باید بلد باشی

صفحه‌بندی مهم‌ترین جای اختلاف است: PostgreSQL از LIMIT n OFFSET m استفاده می‌کند، ولی Oracle (از 12c به بعد) از OFFSET m ROWS FETCH FIRST n ROWS ONLY. روشِ قدیمیِ Oracle با ROWNUM هنوز در کدهای legacy فراوان است ولی خطاخیز است چون ROWNUM قبل از ORDER BY اعمال می‌شود. برای بازه‌های زمانی، PostgreSQL از now() - interval '7 days' و Oracle از SYSTIMESTAMP - INTERVAL '7' DAY استفاده می‌کند. اگر کدت باید روی هر دو اجرا شود، این تفاوت‌ها را در یک لایه‌ی abstraction (مثلاً query builder یا jOOQ) پنهان کن.

۴) نانو-سرویس‌ها (Nano-services) — وقتی زیادی ریز می‌کنی

نانو-سرویس یعنی سرویسی چنان کوچک که هزینه‌ی وجودش (فرآیندِ جدا، شبکه، دیپلوی، مانیتورینگ، دیتابیس) از ارزشی که تولید می‌کند بیشتر است. سرویسی که فقط یک تابع دارد — مثلاً «تبدیلِ ارز» یا «فرمت‌کردنِ تاریخ» — احتمالاً باید یک library باشد، نه یک سرویس.

هزینه‌ی پنهانِ هر سرویس

هر میکروسرویس یک «مالیاتِ ثابت» دارد که هیچ ربطی به منطقش ندارد: یک pipeline‌ی CI/CD، یک داشبوردِ مانیتورینگ، آلارم‌ها، on-call، مدیریتِ نسخه، health check، یک دیتابیس، و ذهنی که تیم باید صرفِ فهمیدنش کند. اگر ۲۰۰ نانو-سرویس داشته باشی، این مالیات تو را خفه می‌کند. یک تیمِ ۵ نفره نمی‌تواند ۴۰ سرویس را سالم نگه دارد. قاعده‌ی سرانگشتی: تعدادِ سرویس‌ها نباید خیلی از تعدادِ آدم‌های تیم بیشتر باشد.

۵) سرویس‌های موجودیت‌محور (Entity Services / CRUD-as-a-Service)

این یکی ظریف‌تر است. entity service یعنی سرویسی که حولِ یک جدول ساخته شده، نه حولِ یک قابلیت: «UserService»، «OrderService»، «ProductService» که هر کدام فقط CRUDِ یک entity را انجام می‌دهند و کلِ منطقِ کسب‌وکار در یک «orchestrator» بالای آن‌ها زندگی می‌کند.

مشکل: منطقِ کسب‌وکار در هیچ سرویسی نیست، بلکه در بینِ آن‌هاست. هر عملیاتِ واقعی به chatty callها و ترنزکشن‌های توزیع‌شده تبدیل می‌شود. این عملاً همان distributed monolith است با ماسکِ «تمیز».

مرز را حولِ رفتار بکِش، نه داده

به‌جای «UserService» که فقط CRUD می‌کند، به قابلیت فکر کن: «Onboarding»، «Billing»، «Fulfillment». یک سرویسِ خوب یک تصمیم می‌گیرد و یک invariant را محافظت می‌کند، نه این‌که فقط یک جدول را بخواند و بنویسد. اگر اسمِ سرویست با اسمِ یک جدول یکی است، زنگِ خطر است. اسمِ سرویس باید یک فعلِ کسب‌وکار را تداعی کند، نه یک اسمِ داده‌ای.

کِی میکروسرویس اصلاً اشتباه است؟

حقیقتی که کمتر گفته می‌شود

میکروسرویس یک ابزارِ سازمانی است، نه یک ابزارِ فنی. مشکلی که حل می‌کند این است: «چند تیم می‌خواهند مستقل از هم دیپلوی کنند بدونِ این‌که پای هم را لِه کنند.» اگر این مشکل را نداری، احتمالاً میکروسرویس مشکلت را حل نمی‌کند، فقط مشکلاتِ تازه می‌سازد.

میکروسرویس معمولاً اشتباه است وقتی:

  • تیم کوچک است (زیرِ ۱۵–۲۰ نفر): یک مونولیتِ خوش‌ساخت (modular monolith) با مرزهای داخلیِ تمیز بهت ۹۰٪ فوایدِ میکروسرویس را با ۱۰٪ هزینه می‌دهد.
  • دامنه هنوز ناپخته است: اگر هنوز نمی‌دانی مرزهای درست کجاست (و در استارتاپِ نوپا هرگز نمی‌دانی)، مرزهای اشتباهِ سرویس را در بتن می‌ریزی. جابه‌جا کردنِ مرز در مونولیت یک refactor است؛ در میکروسرویس یک پروژه‌ی چندماهه.
  • بلوغِ عملیاتی نداری: بدونِ CI/CD خودکار، مشاهده‌پذیریِ خوب، و فرهنگِ on-call، میکروسرویس یعنی تاریکیِ توزیع‌شده.
چرا «مونولیت اول» (monolith-first) اغلب توصیه‌ی درستی است؟

پاسخ: چون بزرگ‌ترین ریسکِ میکروسرویس، کشیدنِ مرزهای اشتباه است و تو در ابتدای کار کم‌ترین دانش را درباره‌ی مرزهای درست داری. مونولیتِ ماژولار بهت اجازه می‌دهد مرزها را ارزان جابه‌جا کنی تا زمانی که دامنه پخته شود و بفهمی کدام بخش‌ها واقعاً باید مستقل باشند. آن‌وقت می‌توانی همان ماژول‌ها را — که حالا مرزهای پایداری دارند — به سرویس بیرون بکِشی. مارتین فاولر این را «monolith-first» می‌نامد. من در اینترویو اضافه می‌کنم: «میکروسرویس یک بهینه‌سازیِ سازمانی است؛ تا وقتی درد سازمانی (چند تیم، تداخلِ دیپلوی) را حس نکرده‌ای، زود است.»


بخش دو: نگرانی‌های پروداکشن — جایی که سنیور واقعاً می‌درخشد

اینجاست که خطِ بینِ «کسی که کد می‌نویسد» و «کسی که سیستم را نگه می‌دارد» کشیده می‌شود. کد نوشتن آسان است؛ زنده نگه داشتنِ سیستم در ساعتِ ۳ صبح، وقتی نصفِ اطلاعات را نداری، هنر است.

مشاهده‌پذیری در عمل (Observability)

مشاهده‌پذیری یعنی توانِ پاسخ دادن به سؤالاتی که از قبل نمی‌دانستی قرار است بپرسی، فقط از روی خروجی‌های سیستم. سه ستونِ کلاسیک: logs (چه اتفاقی افتاد)، metrics (چقدر، چند تا، چه سرعتی)، traces (یک درخواستِ خاص از کجا به کجا رفت).

جعبه‌ی سیاهِ هواپیما

monitoring مثلِ چراغ‌های هشدارِ داشبوردِ خلبان است: از قبل می‌دانی چه چیزهایی ممکن است خراب شود و برایشان چراغ گذاشته‌ای. observability مثلِ جعبه‌ی سیاه است: بعد از حادثه، از روی داده‌های ضبط‌شده می‌توانی هر سؤالی بپرسی، حتی سؤالی که قبلِ سقوط به ذهنت نرسیده بود. در سیستمِ توزیع‌شده، به هر دو نیاز داری، ولی observability آن چیزی است که شبِ حادثه نجاتت می‌دهد.

Correlation ID و Distributed Tracing

بزرگ‌ترین دردِ دیباگ در سیستمِ توزیع‌شده این است: یک کاربر می‌گوید «سفارشم ثبت نشد»، و تو باید ردِ آن یک درخواست را در لاگِ ۸ سرویس پیدا کنی. بدونِ یک شناسه‌ی مشترک، این کارِ محال است.

correlation ID (یا trace ID) یک شناسه‌ی یکتاست که در لبه‌ی سیستم (gateway) تولید می‌شود و با هر فراخوانیِ بعدی — چه HTTP چه پیام در صف — همراه می‌شود. حالا با یک grep می‌توانی کلِ سفرِ آن درخواست را ببینی.

خبرِ خوب: در Spring Boot مدرن این کار تقریباً خودکار است. با Micrometer Tracing کافی است یک پُل (bridge) اضافه کنی و Boot خودش traceId و spanId را به MDC (Mapped Diagnostic Context) لاگ تزریق می‌کند و آن‌ها را روی مرزِ سرویس‌ها propagate می‌کند.

<!-- پُل به سیستمِ tracing: یا Brave/Zipkin یا OpenTelemetry -->
<dependency>
  <groupId>io.micrometer</groupId>
  <artifactId>micrometer-tracing-bridge-brave</artifactId>
</dependency>
<!-- exporter: مثلاً به Zipkin یا از طریق OTLP به Tempo/Jaeger -->
<dependency>
  <groupId>io.zipkin.reporter2</groupId>
  <artifactId>zipkin-reporter-brave</artifactId>
</dependency>

الگوی لاگ را طوری تنظیم کن که این شناسه‌ها دیده شوند:

# application.yml
logging:
  pattern:
    # [appName,traceId,spanId] — Boot این‌ها را از MDC می‌خواند
    level: "%5p [${spring.application.name:},%X{traceId:-},%X{spanId:-}]"
management:
  tracing:
    sampling:
      probability: 1.0    # در dev همه‌چیز؛ در prod معمولاً 0.05 تا 0.1
نمونه‌گیریِ ۱۰۰٪ در پروداکشن، اشتباهِ گران

probability: 1.0 یعنی هر درخواست trace شود. در dev عالی است، ولی در پروداکشنِ پرترافیک، این یعنی حجمِ عظیمی داده‌ی trace که هم پول (هزینه‌ی storage در Tempo/Datadog) و هم پهنای‌باند می‌سوزاند. در پروداکشن معمولاً ۵٪ تا ۱۰٪ نمونه‌گیری کافی است. اما یک نکته‌ی سنیوری: tail-based sampling بهتر است — یعنی به‌جای تصمیمِ تصادفی در ابتدا، همه را موقتاً نگه دار و فقط traceهایی را که کند بودند یا error داشتند نگه دار. این کار در سطحِ collector (مثلِ OpenTelemetry Collector) انجام می‌شود، نه در اپلیکیشن.

وقتی درخواست بینِ سرویس‌ها می‌رود، RestClient یا WebClient‌ای که Spring می‌سازد به‌طور خودکار هدرهای propagation (مثلِ traceparent در استانداردِ W3C) را حمل می‌کند — به شرطی که آن client را از bean کارخانه‌ایِ Spring بگیری، نه با new. اگر خودت دستی client بسازی، این زنجیره می‌شکند.

نمودار: یک درخواست با یک trace ID مشترک از میانِ چند سرویس عبور می‌کند — One trace ID flowing through the whole request.

sequenceDiagram
    participant C as Client
    participant G as Gateway
    participant O as OrderSvc
    participant P as PaymentSvc
    participant K as Kafka
    C->>G: POST /orders
    Note over G: generate traceId=abc123
    G->>O: traceparent: abc123
    O->>P: traceparent: abc123
    P-->>O: 200
    O->>K: OrderPlaced (traceId=abc123 in header)
    O-->>G: 201
    G-->>C: 201

نکته‌ی طلایی: propagation باید حتی از میانِ صفِ پیام هم عبور کند. وقتی OrderPlaced را در Kafka منتشر می‌کنی، traceId را در هدرِ پیام بگذار تا مصرف‌کننده بتواند trace را ادامه دهد. وگرنه ردِ درخواست دقیقاً همان‌جا که async می‌شود، گم می‌شود.

RED و USE — کدام متریک را نگاه کنم؟

تازه‌کار همه‌چیز را monitor می‌کند و در نتیجه هیچ‌چیز را نمی‌بیند. سنیور می‌داند کدام چند متریک واقعاً مهم‌اند. دو چارچوبِ استاندارد وجود دارد:

چارچوب برای چه سه متریک
RED (Tom Wilkie، Grafana) سلامتِ سرویس (نمای کاربر) Rate (نرخِ درخواست)، Errors (نرخِ خطا)، Duration (تاخیر/latency)
USE (Brendan Gregg) سلامتِ منابع (CPU, disk, pool) Utilization (بهره‌وری)، Saturation (اشباع/صف)، Errors (خطا)
Four Golden Signals (Google SRE) نمای کاربر Latency، Traffic، Errors، Saturation
چطور این‌ها را با هم استفاده کنی

RED را روی هر سرویس بگذار: «این سرویس چند درخواست در ثانیه می‌گیرد، چند درصد خطا می‌دهد، و کوانتایلِ p99 تاخیرش چند است؟» USE را روی هر منبع بگذار: «pool کانکشنِ دیتابیس چند درصد پُر است؟ صفِ thread چقدر طولانی شده؟». وقتی incident می‌شود، RED بهت می‌گوید کدام سرویس درد دارد، و USE بهت می‌گوید چرا (کدام منبع اشباع شده). این تقسیمِ کار سرعتِ تشخیصت را چند برابر می‌کند.

میانگین دروغ می‌گوید — همیشه به کوانتایل نگاه کن

اگر فقط میانگینِ latency را ببینی، فریب می‌خوری. یک سرویس ممکن است میانگینِ ۵۰ms داشته باشد ولی p99 آن ۳ ثانیه باشد — یعنی از هر ۱۰۰ کاربر، یکی تجربه‌ی فاجعه دارد. و در سیستمِ توزیع‌شده که یک درخواست ۱۰ سرویس را صدا می‌زند، احتمالِ برخورد با آن دُمِ کند به‌شدت بالا می‌رود. همیشه p95, p99, p99.9 را monitor کن، نه میانگین. این پدیده را «tail latency amplification» می‌نامند.

در Spring، Micrometer این متریک‌ها را تقریباً رایگان می‌دهد. با افزودنِ micrometer-registry-prometheus، اکچوایتور یک endpoint در /actuator/prometheus می‌سازد که همه‌ی متریک‌های HTTP (شاملِ http.server.requests با تگ‌های status و uri و کوانتایل‌ها) را بیرون می‌دهد:

management:
  endpoints:
    web:
      exposure:
        include: health,info,prometheus,metrics
  metrics:
    distribution:
      percentiles-histogram:
        http.server.requests: true   # هیستوگرام برای محاسبه‌ی کوانتایلِ سمتِ Prometheus
      slo:
        http.server.requests: 100ms,300ms,500ms

اضافه کردنِ یک متریکِ کسب‌وکار هم ساده است — و سنیور می‌داند که متریک‌های کسب‌وکار (مثلِ «نرخِ سفارشِ ناموفق») گاهی از متریک‌های فنی مهم‌ترند:

@Service
public class OrderService {
    private final Counter failedOrders;

    public OrderService(MeterRegistry registry) {
        this.failedOrders = Counter.builder("orders.failed")
            .description("Orders that failed to place")
            .tag("channel", "web")
            .register(registry);
    }

    public void place(Order order) {
        try {
            // ... business logic ...
        } catch (PaymentDeclinedException e) {
            failedOrders.increment();
            throw e;
        }
    }
}

مبانیِ on-call و مدیریتِ حادثه (Incident Management)

روزی می‌رسد که پیجِر زنگ می‌زند و تو مسئولی. این بخش را کمتر جایی به مهندس‌ها یاد می‌دهند، ولی همین بخش «سنیور» را از «برنامه‌نویسِ خوب» جدا می‌کند.

سلسله‌مراتبِ اولویت در حادثه

وقتی سیستم down است، ترتیبِ کارها این است: ۱) بازیابی (mitigate) — سرویس را به هر قیمتی برگردان، حتی با rollback یا کشتنِ یک فیچر. ۲) ارتباط (communicate) — به stakeholderها بگو چه خبر است. ۳) ریشه‌یابی (root cause) — این بعد از بازیابی است، نه هم‌زمان. تازه‌کارها وقتِ حادثه دنبالِ «چرا» می‌گردند؛ سنیور اول سیستم را برمی‌گرداند، بعد دنبالِ چرا می‌گردد. «بازیابی مقدم بر تشخیص است.»

مفاهیمی که باید بلد باشی:

  • SLI / SLO / Error Budget: SLI چیزی است که اندازه می‌گیری (مثلاً «درصدِ درخواست‌های زیرِ ۳۰۰ms»). SLO هدفی است که می‌خواهی به آن برسی (مثلاً «۹۹.۹٪ درخواست‌ها زیرِ ۳۰۰ms»). error budget بودجه‌ی خطایی است که از SLO به دست می‌آید: SLO برابرِ ۹۹.۹٪ یعنی ۰.۱٪ اجازه‌ی خطا، که در یک ماه تقریباً ۴۳ دقیقه downtimeِ مجاز است. این عدد فوق‌العاده قدرتمند است چون بحثِ «باثبات‌تر باشیم یا سریع‌تر فیچر بدهیم» را از دعوای احساسی به یک عددِ مشترک تبدیل می‌کند.
error budget، سلاحِ مذاکره‌ی سنیور

وقتی مدیرِ محصول برای سرعتِ بیشتر فشار می‌آورد و تیمِ زیرساخت برای پایداریِ بیشتر، error budget داور است. اگر بودجه‌ی خطا هنوز باقی مانده، یعنی می‌توانی ریسک کنی و سریع‌تر deploy کنی. اگر بودجه تمام شده (این ماه خیلی خراب بوده‌ای)، یعنی باید ترمز بکِشی و روی پایداری کار کنی. این مفهوم بحث‌های بی‌پایانِ «کیفیت در برابر سرعت» را به یک قاعده‌ی عینی تبدیل می‌کند. در اینترویوِ سنیور اگر این را بگویی، امتیازِ بزرگی می‌گیری.

  • Runbook: یک سندِ ساده که می‌گوید «اگر این آلارم زنگ زد، این کارها را بکن.» شبِ حادثه کسی حالِ فکرِ خلاقانه ندارد؛ runbook مغزِ ذخیره‌ات است.
  • Blameless postmortem: بعد از حادثه، جلسه‌ای بدونِ سرزنشِ فرد. هدف پیدا کردنِ نقصِ سیستم است، نه مقصر. چون اگر آدم‌ها بترسند، حادثه‌ها را پنهان می‌کنند و همان اشتباه تکرار می‌شود.
آلارمی که همیشه زنگ می‌زند، آلارم نیست

بدترین چیز برای on-call، «alert fatigue» است: وقتی آن‌قدر آلارمِ بی‌اهمیت زنگ می‌زند که مغز یاد می‌گیرد نادیده بگیرد — و روزی که آلارمِ واقعی می‌آید، آن هم نادیده گرفته می‌شود. قاعده: هر آلارم باید قابلِ اقدام (actionable) و متصل به یک SLO یا نشانه‌ی درد کاربر باشد. آلارم روی «CPU بالای ۸۰٪» بد است (شاید طبیعی باشد)؛ آلارم روی «نرخِ خطای ۵xx بالای ۱٪ برای ۵ دقیقه» خوب است. اگر برای یک آلارم نمی‌توانی runbook بنویسی، آن آلارم نباید کسی را از خواب بیدار کند.

ظرفیت و Back-pressure

قیفِ آب و ظرفِ سرریز

تصور کن آب را با شلنگ در قیف می‌ریزی. اگر سرعتِ ورودی از خروجی بیشتر شود، قیف پُر می‌شود و سرریز می‌کند. سیستمِ نرم‌افزاری هم همین است: اگر نرخِ درخواستِ ورودی از توانِ پردازشت بیشتر شود، صف‌ها پُر می‌شوند، حافظه تمام می‌شود، و سیستم به جای «کند شدن»، «می‌میرد». back-pressure یعنی مکانیزمی که به بالادست می‌گوید «آهسته‌تر، من دارم پُر می‌شوم» — به‌جای این‌که بی‌صدا غرق شود.

خطای مرگبارِ رایج: صفِ نامحدود (unbounded queue). یک ExecutorService با صفِ بی‌نهایت زیرِ فشار به آرامی حافظه را می‌بلعد تا OutOfMemoryError و مرگ. راهِ درست، صفِ محدود با سیاستِ ردِ روشن است:

// درست: pool محدود + صفِ محدود + سیاستِ رد
ThreadPoolExecutor executor = new ThreadPoolExecutor(
    10, 20,                          // core / max threads
    60, TimeUnit.SECONDS,
    new ArrayBlockingQueue<>(500),   // صفِ محدود — نه LinkedBlockingQueue بی‌نهایت
    new ThreadPoolExecutor.CallerRunsPolicy()  // back-pressure: صف‌کننده خودش اجرا می‌کند
);

CallerRunsPolicy یک ترفندِ ظریف است: وقتی صف پُر می‌شود، thread‌ای که کار را submit کرده مجبور می‌شود خودش آن را اجرا کند. این عملاً thread پذیرنده را کند می‌کند و به‌طور طبیعی نرخِ ورودی را کم می‌کند — یک back-pressure ساده و مؤثر.

در لایه‌ی سرویس، bulkhead و rate limiter ابزارِ اصلی‌اند. bulkhead تعدادِ فراخوانیِ هم‌زمان به یک dependency را محدود می‌کند تا یک dependencyِ کند، کلِ thread poolِ سرویس را نبلعد:

@Bulkhead(name = "inventory", type = Bulkhead.Type.SEMAPHORE)
@RateLimiter(name = "inventory")
@CircuitBreaker(name = "inventory", fallbackMethod = "fallbackStock")
public StockLevel checkStock(String sku) {
    return inventoryClient.getStock(sku);
}

public StockLevel fallbackStock(String sku, Throwable t) {
    // افتِ باوقار: مقدارِ محافظه‌کارانه به‌جای خطا
    return StockLevel.unknown(sku);
}
resilience4j:
  bulkhead:
    instances:
      inventory:
        max-concurrent-calls: 25       # حداکثر ۲۵ فراخوانیِ هم‌زمان
  ratelimiter:
    instances:
      inventory:
        limit-for-period: 100          # ۱۰۰ فراخوانی
        limit-refresh-period: 1s       # در هر ثانیه
        timeout-duration: 0            # اگر سهمیه تمام شد، فوراً رد کن (fail-fast)
  circuitbreaker:
    instances:
      inventory:
        sliding-window-size: 20
        failure-rate-threshold: 50     # اگر ۵۰٪ فراخوانی‌ها خطا داد، مدار را باز کن
        wait-duration-in-open-state: 10s
تایم‌اوت‌ها را زنجیره‌ای بچین، وگرنه بی‌فایده‌اند

یک اشتباهِ کلاسیک: سرویس A تایم‌اوتِ ۱۰ ثانیه دارد، ولی B که A صدایش می‌زند تایم‌اوتِ ۳۰ ثانیه دارد. نتیجه: A منتظرِ B نمی‌ماند و قطع می‌کند، ولی B هنوز کارِ بی‌فایده می‌کند و منابع را می‌سوزاند (کارِ یتیم/orphaned work). قاعده: تایم‌اوتِ بالادست باید از تایم‌اوتِ پایین‌دست کوتاه‌تر باشد و هر لایه بودجه‌ی زمانی‌اش را به لایه‌ی بعد منتقل کند (deadline propagation). بدونِ تایم‌اوتِ درست، circuit breaker هم بی‌فایده است، چون thread قبل از این‌که مدار متوجه شود، بلاک شده.

نمودار: چرخه‌ی حالتِ circuit breaker — Closed → Open → Half-Open lifecycle.

stateDiagram-v2
    [*] --> Closed
    Closed --> Open: failure rate > threshold
    Open --> HalfOpen: after wait duration
    HalfOpen --> Closed: trial calls succeed
    HalfOpen --> Open: trial calls fail
    note right of Open
        fail fast, run fallback
    end note

افتِ باوقار (Graceful Degradation)

آسانسورِ خراب، پله‌ی سالم

اگر آسانسورِ یک ساختمان خراب شود، ساختمان تعطیل نمی‌شود؛ مردم از پله استفاده می‌کنند. تجربه بدتر است، ولی ساختمان کار می‌کند. graceful degradation یعنی وقتی بخشی از سیستم می‌افتد، به‌جای مرگِ کامل، با کیفیتِ پایین‌تر سرویس بدهی. سنیور همیشه می‌پرسد: «اگر این dependency بیفتد، آیا کاربر همچنان می‌تواند کارِ اصلی‌اش را انجام دهد؟»

مثال‌های واقعی:

  • سرویسِ «پیشنهادهای شخصی» افتاد؟ به‌جای خطا، لیستِ پرفروش‌ها را نشان بده.
  • سرویسِ قیمت‌گذاریِ لحظه‌ای کند شد؟ قیمتِ cache‌شده‌ی چند دقیقه‌پیش را نشان بده و بنری بزن «قیمت‌ها ممکن است تغییر کند».
  • سرویسِ تصاویر افتاد؟ placeholder نشان بده، ولی دکمه‌ی خرید را زنده نگه دار.
هر dependency را «بحرانی» یا «غیربحرانی» برچسب بزن

یک تمرینِ سنیوری: برای هر سرویس، dependencyهایش را به دو دسته تقسیم کن. بحرانی (اگر بیفتد، عملیاتِ اصلی غیرممکن است — مثلِ دیتابیسِ سفارش) و غیربحرانی (اگر بیفتد، فقط تجربه بدتر می‌شود — مثلِ سرویسِ توصیه). برای dependencyهای غیربحرانی، هرگز نباید کلِ درخواست fail شود. این برچسب‌گذاری را در طراحی انجام بده، نه بعد از اولین حادثه. یک fallback روی هر فراخوانیِ غیربحرانی، تفاوتِ بینِ «سایت کند است» و «سایت down است» را رقم می‌زند.

هزینه (Cost) — بُعدی که مهندس‌ها فراموش می‌کنند

سنیور می‌داند که هر تصمیمِ معماری یک برچسبِ قیمت دارد و کسی باید آن را بپردازد. میکروسرویس ذاتاً گران‌تر از مونولیت است:

  • هزینه‌ی شبکه: هر فراخوانیِ بینِ سرویس، پول است (به‌خصوص عبورِ ترافیک بینِ availability zoneها در ابر که هزینه‌ی egress دارد).
  • هزینه‌ی تکرار: هر سرویس CPU/RAM/دیتابیسِ خودش را می‌خواهد. ۲۰ سرویسِ کوچک اغلب گران‌تر از یک مونولیتِ متراکم‌اند چون هر کدام یک baseline مصرف دارند (JVM حداقلِ heap، sidecar، و…).
  • هزینه‌ی مشاهده‌پذیری: این جای غافلگیریِ بزرگ است. حجمِ لاگ و trace و metric در سیستمِ توزیع‌شده انفجاری است، و ابزارهای APM (مثلِ Datadog) بر اساسِ حجم پول می‌گیرند. دیده‌ام تیمی که صورت‌حسابِ observabilityاش از صورت‌حسابِ computeاش بیشتر شده.
قاتلِ خاموشِ صورت‌حساب: log و trace بی‌حساب

سه منبعِ اصلیِ هزینه‌ی پنهان: (۱) نمونه‌گیریِ ۱۰۰٪ trace، (۲) لاگِ سطحِ DEBUG در پروداکشن، (۳) cardinalityِ بالای متریک‌ها — مثلاً گذاشتنِ userId به‌عنوان tag روی یک متریک. هر مقدارِ یکتای tag یک time-series جدا می‌سازد؛ اگر userId (میلیون‌ها مقدار) را tag کنی، سیستمِ متریکت منفجر می‌شود (cardinality explosion) و هزینه‌ات سر به فلک می‌کشد. tagها باید کم‌کاردینالیتی باشند: status, region, endpoint — نه شناسه‌های یکتا.

مهاجرتِ داده بینِ سرویس‌ها (Data Migration)

سخت‌ترین کارِ عملیاتی در دنیای میکروسرویس این است: می‌خواهی یک جدول یا مالکیتِ یک بخش از داده را از سرویسِ A به سرویسِ B منتقل کنی — بدونِ downtime. این جایی است که خیلی از سنیورهای تازه هم زمین می‌خورند.

نکته‌ی طلایی: هرگز big-bang نکن

هرگز داده را در یک شبِ خاموشیِ بزرگ منتقل نکن («جمعه شب down می‌کنیم، migration می‌زنیم، شنبه بالا می‌آییم»). این تقریباً همیشه بد پیش می‌رود و rollback ندارد. الگوی درست، مهاجرتِ تدریجی و برگشت‌پذیر است. مشهورترین الگو expand–migrate–contract (یا parallel-change) است.

الگوی expand/contract برای تغییرِ schema بدونِ downtime سه فاز دارد:

  1. Expand — ساختارِ جدید را در کنارِ قدیمی اضافه کن (ستونِ جدید، جدولِ جدید). هنوز چیزی حذف نمی‌شود. کد به هر دو می‌نویسد (dual-write) ولی هنوز از قدیمی می‌خواند.
  2. Migrate — داده‌ی موجود را به تدریج (backfill) به ساختارِ جدید کپی کن، و کد را عوض کن تا از جدید بخواند. حالا هم قدیمی و هم جدید به‌روزند.
  3. Contract — وقتی مطمئن شدی همه‌چیز از جدید می‌خواند و می‌نویسد، ساختارِ قدیمی را حذف کن.

هر فاز به‌تنهایی deploy و rollback می‌شود. کلیدِ ماجرا این است که هیچ deploy‌ای به‌طور هم‌زمان هم schema و هم کد را به شکلِ ناسازگار عوض نکند.

برای انتقالِ داده بینِ دو سرویسِ مختلف (نه فقط schema)، الگوی رایج Change Data Capture (CDC) یا Outbox است: سرویسِ منبع رویدادهای تغییرِ داده را منتشر می‌کند، سرویسِ مقصد آن‌ها را مصرف و state خودش را می‌سازد، تا وقتی هم‌تراز شدند و بتوانی ترافیک را سوییچ کنی.

نمودار: مهاجرتِ مالکیتِ داده با dual-write و backfill — Zero-downtime ownership move via dual-write + backfill.

flowchart TD
    subgraph Phase1[Expand: dual-write]
        W[Write path] --> Old[(Old Store)]
        W --> New[(New Store)]
    end
    subgraph Phase2[Migrate: backfill + read new]
        B[Backfill job] --> New
        R[Read path] --> New
    end
    subgraph Phase3[Contract: drop old]
        R2[All traffic] --> New
    end
    Phase1 --> Phase2 --> Phase3

مثالِ SQL برای فازِ backfill که باید idempotent و batch باشد (تا جدولِ زنده را قفل نکند):

-- PostgreSQL: upsert idempotent در batch
INSERT INTO new_service.customer (id, email, tier)
SELECT id, email, tier FROM old_service.customer
WHERE  id BETWEEN :lo AND :hi
ON CONFLICT (id) DO UPDATE
   SET email = EXCLUDED.email, tier = EXCLUDED.tier;
-- Oracle: همان idempotency با MERGE
MERGE INTO new_service.customer t
USING (SELECT id, email, tier FROM old_service.customer
       WHERE id BETWEEN :lo AND :hi) s
ON (t.id = s.id)
WHEN MATCHED THEN UPDATE SET t.email = s.email, t.tier = s.tier
WHEN NOT MATCHED THEN INSERT (id, email, tier)
     VALUES (s.id, s.email, s.tier);
upsert: ON CONFLICT در برابر MERGE

PostgreSQL برای upsert از INSERT ... ON CONFLICT ... DO UPDATE استفاده می‌کند که مختصر و اتمیک است. Oracle معادلِ مستقیم ندارد و از MERGE استفاده می‌کند (که در PostgreSQL هم از نسخه‌ی ۱۵ به بعد هست، ولی ON CONFLICT رایج‌تر و روان‌تر است). backfill را همیشه در batchهای کوچک با BETWEEN :lo AND :hi بزن و بینِ batchها مکث کن، وگرنه یک UPDATE روی میلیون‌ها سطر، replica lag و قفل ایجاد می‌کند و همان چیزی را که می‌خواستی نجات بدهی، down می‌کنی.

dual-write بدونِ تضمین، منبعِ ناسازگاری است

dual-write (نوشتنِ هم‌زمان به دو store) یک تله دارد: اگر نوشتن به store اول موفق شود ولی به دومی fail، دو منبعِ داده ناسازگار می‌شوند و هیچ ترنزکشنِ اتمیکی بینشان نیست (چون دو سیستمِ جدا هستند). راهِ امن این است که یکی از دو نوشتن را از طریقِ outbox انجام بدهی: در همان ترنزکشنِ دیتابیسِ اول، یک ردیف در جدولِ outbox بنویس، و یک relay آن را به store دوم منتقل کند. این‌طور اتمیک بودن حفظ می‌شود. dual-write مستقیم فقط وقتی قابل‌قبول است که یک reconciliation job مرتب ناسازگاری‌ها را پیدا و تصحیح کند.


بخش سه: چطور در اینترویوِ system design بلند فکر کنی

اینترویوِ طراحیِ سیستم، دانشِ تو را نمی‌سنجد؛ قضاوت و فرآیندِ فکریِ تو را می‌سنجد. اینترویوئر می‌خواهد ببیند که وقتی مسئله مبهم است، چطور آن را روشن، بخش‌بندی، و trade-off می‌کنی. سکوت و رفتنِ مستقیم سرِ راه‌حل، بدترین کار است.

چارچوبِ ۶ مرحله‌ای برای بلند فکر کردن

۱) روشن‌سازی (clarify) — سؤال بپرس؛ محدوده را مشخص کن. ۲) تخمینِ مقیاس (scale) — چند کاربر، چند QPS، چقدر داده؟ ۳) APIها و مدلِ داده — قراردادها را تعریف کن. ۴) طراحیِ سطحِ بالا (high-level) — یک دیاگرامِ جعبه‌ای بکِش. ۵) عمیق شدن (deep dive) — یک بخشِ سخت را باز کن. ۶) trade-offها و گلوگاه‌ها — بگو چه چیزی را قربانی کردی و چرا. کلید: در تمامِ این مراحل بلند حرف بزن و فرض‌هایت را اعلام کن.

اولین جمله‌ات نباید نامِ یک تکنولوژی باشد

تازه‌کار می‌گوید «از Kafka و Redis و Cassandra استفاده می‌کنم». سنیور اول می‌پرسد «الگوی خواندن به نوشتن چطور است؟ چند کاربرِ هم‌زمان؟ آیا سازگاریِ قوی لازم است یا eventual کافی است؟». تکنولوژی پیامدِ نیازمندی‌هاست، نه شروعِ بحث. اگر اول تکنولوژی نام ببری، اینترویوئر فکر می‌کند فقط buzzword بلدی. اول مسئله را بفهم، بعد ابزار را از دلِ trade-off بیرون بکِش.

مثالِ عینی: اگر بپرسند «یک سرویسِ کوتاه‌کننده‌ی URL طراحی کن»، مسیرِ سنیور این است:

  • «چند URL در روز ساخته می‌شود؟ چند بار خوانده می‌شود؟» (معمولاً read خیلی بیشتر از write) → پس سیستم read-heavy است → cache و read replica مهم می‌شوند.
  • «طولِ کوتاه‌شده مهم است؟ لینک‌ها منقضی می‌شوند؟» → مدلِ داده و TTL روشن می‌شود.
  • «سازگاری: اگر دو نفر هم‌زمان همان کد را بگیرند چه؟» → استراتژیِ تولیدِ کلید (counter مرکزی در برابر hash).
دام: پیچیده کردنِ زودهنگام

اشتباهِ رایجِ کاندیدها این است که از همان اول معماریِ ۲۰-سرویسیِ فضایی می‌کِشند تا «سنیور» به نظر برسند. اثرِ برعکس دارد. سنیورِ واقعی با ساده‌ترین طراحیِ ممکن شروع می‌کند و فقط وقتی پیچیدگی اضافه می‌کند که یک نیازمندیِ مشخص (مقیاس، دسترس‌پذیری) آن را مجبور کند. جمله‌ی طلایی در اینترویو: «ساده‌ترین چیزی که کار می‌کند این است؛ حالا بیایید ببینیم کجا زیرِ فشار می‌شکند و فقط همان‌جا را پیچیده کنیم.»

چطور بین سازگاریِ قوی (strong) و نهایی (eventual) تصمیم می‌گیری؟

پاسخ: به پیامدِ کسب‌وکارِ داده‌ی کهنه نگاه می‌کنم. اگر داده‌ی کهنه منجر به خطای مالی یا نقضِ invariant شود — مثلِ موجودیِ حساب یا رزروِ صندلی — سازگاریِ قوی لازم است، حتی به قیمتِ تاخیر و دسترس‌پذیریِ کمتر. اگر داده‌ی کهنه فقط کمی آزاردهنده باشد — مثلِ تعدادِ لایکِ یک پست یا لیستِ توصیه‌ها — eventual consistency کاملاً کافی است و بهم مقیاس و دسترس‌پذیریِ بهتر می‌دهد. طبقِ قضیه‌ی CAP، در یک سیستمِ توزیع‌شده هنگامِ partition باید بین سازگاری و دسترس‌پذیری یکی را انتخاب کنی؛ ولی در عمل این یک انتخابِ باینری نیست، بلکه یک طیف است که به ازای هر نوع داده جدا تصمیم می‌گیری. یک سیستمِ واقعی معمولاً برای بخشِ پول strong و برای بخشِ اجتماعی eventual است.

یک درخواستِ کاربر کند شده و ۸ سرویس درگیرند. چطور ریشه‌یابی می‌کنی؟

پاسخ: اول distributed trace آن درخواست را با trace ID پیدا می‌کنم — این بلافاصله نشان می‌دهد کدام span بیشترین زمان را برده، پس نمی‌گردم بلکه می‌بینم گلوگاه کجاست. بعد روی آن سرویس، متریک‌های USE را نگاه می‌کنم: آیا pool کانکشنِ دیتابیس اشباع شده؟ صفِ thread طولانی شده؟ GC pause زیاد است؟ اگر تاخیر در یک فراخوانیِ downstream است، بررسی می‌کنم آیا آن downstream واقعاً کند است یا فقط retryهای ما داریم بارش را چند برابر می‌کنیم (retry storm). نکته‌ی سنیوری: همیشه به p99 نگاه می‌کنم نه میانگین، و مراقبِ tail latency amplification هستم — یک سرویسِ کمی کُند در یک زنجیره‌ی ۸-تایی می‌تواند p99 کل را منفجر کند.

تفاوتِ orchestration و choreography در saga چیست و کِی کدام را انتخاب می‌کنی؟

پاسخ: در orchestration، یک هماهنگ‌کننده‌ی مرکزی (orchestrator) گام‌های saga را صریحاً صدا می‌زند و در صورتِ خطا جبران (compensation) را مدیریت می‌کند — منطق در یک جا متمرکز و قابلِ ردیابی است، ولی آن orchestrator یک نقطه‌ی coupling می‌شود. در choreography، هر سرویس به eventها واکنش نشان می‌دهد و event بعدی را منتشر می‌کند — coupling کمتر است ولی جریانِ کلی در هیچ‌جا نوشته نشده و دنبال کردنش سخت است (اصطلاحاً «هیچ‌کس تصویرِ کامل را نمی‌بیند»). قاعده‌ی من: برای جریان‌های ساده با چند گام، choreography سبک‌تر است؛ برای جریان‌های پیچیده با شرط و جبران‌های زیاد، orchestration را انتخاب می‌کنم چون قابلیتِ observe و debug مهم‌تر از coupling کمتر می‌شود. و در هر دو حالت، هر گام باید idempotent باشد.

چطور exactly-once را در یک سیستمِ پیام‌محور تضمین می‌کنی؟

پاسخ: صادقانه: exactly-once delivery در سیستمِ توزیع‌شده عملاً غیرممکن است؛ چیزی که واقعاً به دست می‌آوری «at-least-once delivery + idempotent processing» است که اثرش مثلِ exactly-once است. یعنی می‌پذیرم پیام ممکن است چند بار برسد، ولی مصرف‌کننده را طوری می‌سازم که پردازشِ تکراری بی‌اثر باشد — با یک کلیدِ idempotency (مثلِ messageId) که در دیتابیس ذخیره می‌کنم و قبل از پردازش چک می‌کنم آیا قبلاً دیده‌امش. سمتِ تولید هم از الگوی transactional outbox استفاده می‌کنم تا نوشتنِ داده و انتشارِ event اتمیک باشند. اگر بگویم «بله با فلان تنظیم exactly-once می‌شود»، اینترویوئرِ باتجربه می‌فهمد که عمقِ موضوع را نگرفته‌ام.

سرویس‌ات به یک dependency وابسته است که گاهی ۵ ثانیه طول می‌کشد. چه می‌کنی؟

پاسخ: چند لایه دفاع: (۱) تایم‌اوتِ تهاجمی — مثلاً ۸۰۰ms، چون بهتر است سریع fail کنم تا thread را ۵ ثانیه بلاک کنم. (۲) circuit breaker — اگر نرخِ کندی/خطا از آستانه رد شد، مدار را باز کنم تا فشار از روی dependency برداشته شود و فرصتِ بهبود بدهم. (۳) bulkhead — تعدادِ فراخوانیِ هم‌زمان به این dependency را محدود کنم تا کندی‌اش کلِ thread poolِ من را نبلعد. (۴) fallback — اگر غیربحرانی است، یک پاسخِ cache‌شده یا پیش‌فرض بدهم (graceful degradation). (۵) در سطحِ معماری می‌پرسم: آیا این فراخوانی باید synchronous باشد؟ شاید بتوانم آن را async کنم و از حلقه‌ی درخواستِ کاربر خارجش کنم. ترتیبِ این‌ها مهم است: اول از همه timeout، چون بدونِ آن بقیه بی‌فایده‌اند.

یک deploy جدید نرخِ خطا را بالا برد، ولی فقط برای ۲٪ کاربران. چه اتفاقی افتاده و چطور واکنش می‌دهی؟

پاسخ: ۲٪ بودن قویاً به یک rollout تدریجی (canary) اشاره دارد — احتمالاً نسخه‌ی جدید روی بخشی از instanceها یا کاربران است و همان‌ها خطا می‌دهند. واکنش: اول بازیابی — rollback یا کم کردنِ درصدِ canary به صفر، نه ریشه‌یابی هم‌زمان. بعد trace و لاگِ همان ۲٪ را با فیلترِ نسخه بررسی می‌کنم. علتِ محتملِ چنین چیزی: یک تغییرِ ناسازگارِ schema که فقط برای دیتای خاصی می‌شکند، یا یک feature flag که برای بخشی روشن است، یا یک نودِ خراب. نکته‌ی سنیوری که اینجا می‌گویم: این دقیقاً چرا progressive delivery (canary + متریکِ خودکار) ارزش دارد — چون شعاعِ انفجار (blast radius) را از ۱۰۰٪ به ۲٪ کاهش داد و به من فرصتِ rollback داد قبل از فاجعه.

چطور تصمیم می‌گیری یک قابلیت را سرویسِ جدا کنی یا در سرویسِ موجود نگه داری؟

پاسخ: به چند محور نگاه می‌کنم: (۱) مرزِ تیم — آیا تیمِ جدایی مالکِ این قابلیت است و می‌خواهد مستقل دیپلوی کند؟ این قوی‌ترین دلیل است. (۲) پروفایلِ مقیاس‌پذیریِ متفاوت — آیا این بخش نیازِ مقیاسی کاملاً متفاوتی دارد (مثلاً پردازشِ تصویر که CPU-bound است در کنارِ APIِ سبک)؟ جدا کردن اجازه‌ی مقیاسِ مستقل می‌دهد. (۳) مرزِ تراکنشی — اگر این قابلیت با بقیه در یک ترنزکشنِ اتمیک گره خورده، جدا کردنش یعنی ساختنِ saga و پیچیدگی؛ شاید نیارزد. (۴) نرخِ تغییر — آیا این بخش خیلی بیشتر یا کم‌تر از بقیه تغییر می‌کند؟ اگر همه‌ی این‌ها می‌گویند «جدا»، جدا می‌کنم؛ اگر تردید دارم، آن را یک ماژولِ مرزبندی‌شده‌ی داخلِ همان سرویس نگه می‌دارم تا بعداً ارزان جدا شود. coupling کاذب بدتر از یک سرویسِ کمی بزرگ‌تر است.

چطور از یک مونولیت به میکروسرویس مهاجرت می‌کنی بدونِ توقفِ کسب‌وکار؟

پاسخ: با الگوی Strangler Fig (انجیرِ خفه‌کننده، نامِ فاولر). به‌جای بازنویسیِ big-bang که تقریباً همیشه شکست می‌خورد، یک gateway/proxy جلوی مونولیت می‌گذارم و قابلیت‌ها را یکی‌یکی به سرویس‌های جدید منتقل می‌کنم. proxy ترافیکِ هر قابلیتِ منتقل‌شده را به سرویسِ جدید و بقیه را به مونولیت می‌فرستد. به مرور، سرویس‌های جدید مونولیت را «خفه» می‌کنند تا چیزی نماند. کلید: از قابلیتی شروع می‌کنم که کم‌ترین coupling با بقیه دارد (میوه‌ی کم‌ارتفاع)، و مرزِ داده را با dual-write/CDC مدیریت می‌کنم. هرگز همه را با هم جدا نمی‌کنم و همیشه در هر گام قابلیتِ rollback دارم. مهاجرت باید یک سری قدمِ کوچکِ برگشت‌پذیر باشد، نه یک جهشِ بزرگ.

یعنی چه که «Kafka تضمینِ ترتیب می‌دهد»؟ این تضمین کجا می‌شکند؟

پاسخ: Kafka ترتیب را فقط درونِ یک partition تضمین می‌کند، نه در کلِ topic. پیام‌هایی که به یک partition می‌روند به همان ترتیبِ نوشتن خوانده می‌شوند. این یعنی برای حفظِ ترتیبِ رویدادهای مربوط به یک entity (مثلاً همه‌ی رویدادهای یک orderId)، باید آن‌ها را با یک partition key یکسان (همان orderId) بفرستم تا همه در یک partition بیفتند. جایی که می‌شکند: (۱) اگر key نگذاری، پیام‌ها round-robin پخش می‌شوند و ترتیب از بین می‌رود. (۲) اگر تعدادِ partitionها را عوض کنی، نگاشتِ key به partition تغییر می‌کند و ترتیبِ تاریخی به‌هم می‌ریزد. (۳) اگر مصرف‌کننده چند thread موازی داشته باشد، ممکن است پیام‌های یک partition را بی‌ترتیب پردازش کند. پس ترتیب یک تضمینِ محدود و مشروط است، نه مطلق.

سیستم‌ات باید ۱۰ برابر ترافیک را تحمل کند. از کجا شروع می‌کنی؟

پاسخ: اول اندازه‌گیری، نه حدس. با load test و متریک‌های فعلی می‌فهمم گلوگاهِ واقعی کجاست — تقریباً همیشه یک منبعِ مشخص است (اغلب دیتابیس، نه اپلیکیشن). بعد به ترتیبِ ارزان‌به‌گران: (۱) cache — اگر read-heavy است، بیشترین بازده را دارد. (۲) scale-out افقیِ سرویس‌های stateless — ساده است اگر واقعاً stateless باشند. (۳) read replica و اتصال pooling برای دیتابیس. (۴) اگر write گلوگاه است، sharding/partitioning یا جدا کردنِ writeها به یک صف (async). (۵) back-pressure و rate limiting تا زیرِ بار به‌جای فروپاشی، باوقار افت کند. نکته‌ی سنیوری: ۱۰ برابر شدنِ ترافیک اغلب یک گلوگاهِ غیرخطی را آشکار می‌کند که در بارِ عادی پنهان بود — مثلِ قفلِ دیتابیس یا connection pool. پس فرض نمی‌کنم چیزی که در ۱x کار می‌کرد در ۱۰x هم خطی می‌ماند.

چطور می‌فهمی که «زیادی» میکروسرویس داری؟

پاسخ: چند نشانه: (۱) برای اضافه کردنِ یک فیچرِ ساده مجبوری چند سرویس را هماهنگ تغییر و دیپلوی کنی (نشانه‌ی distributed monolith). (۲) بیشترِ وقتت صرفِ debug کردنِ ارتباطِ بینِ سرویس‌ها می‌شود تا منطقِ کسب‌وکار. (۳) تعدادِ سرویس‌ها خیلی بیشتر از تعدادِ آدم‌های تیم است و کسی تصویرِ کامل را نمی‌فهمد. (۴) سرویس‌هایی داری که فقط CRUDِ یک جدول‌اند (entity service). (۵) هزینه‌ی زیرساخت و observability بی‌تناسب با اندازه‌ی کسب‌وکار بالاست. راهِ درمان معمولاً ادغام است — چند سرویسِ chatty و به‌هم‌وابسته را دوباره در یک سرویس merge کن. یک اعترافِ سنیوری در اینترویو: «merge کردنِ سرویس‌ها هم یک مهارتِ معماری است، نه فقط splitting؛ و شجاعتِ بیشتری می‌خواهد چون خلافِ مُد است.»


بخش چهار: چک‌لیستِ آمادگیِ سنیور

قبل از این‌که یک سرویس را به پروداکشن بفرستی، این‌ها را از خودت بپرس. اگر جوابِ نصفشان «نمی‌دانم» است، هنوز آماده‌ی on-call نیستی.

مشاهده‌پذیری

  • آیا هر درخواست یک correlation/trace ID دارد که در همه‌ی لاگ‌ها و از میانِ صف‌ها propagate می‌شود؟
  • آیا متریک‌های RED (rate/errors/duration با p99) برای هر endpoint دارم؟
  • آیا داشبورد و آلارمِ قابلِ اقدام متصل به SLO دارم، نه آلارمِ نویزی؟

تاب‌آوری (Resilience)

  • آیا هر فراخوانیِ خارجی timeout، retry (با backoff و jitter)، و circuit breaker دارد؟
  • آیا تایم‌اوت‌ها زنجیره‌ای‌اند (بالادست کوتاه‌تر از پایین‌دست)؟
  • آیا برای هر dependencyِ غیربحرانی یک fallback دارم؟
  • آیا صف‌ها و pۆۆۆlها محدودند و back-pressure دارم؟

داده

  • آیا هر سرویس مالکِ دیتابیسِ خودش است و کسی به جدولِ دیگری دست نمی‌زند؟
  • آیا مصرف‌کننده‌های پیام idempotent‌اند (کلیدِ idempotency دارند)؟
  • آیا تغییراتِ schema با expand/contract و بدونِ downtime قابلِ اجرا هستند؟
  • آیا انتشارِ event اتمیک است (transactional outbox)، نه dual-write شکننده؟

عملیات

  • آیا runbook برای آلارم‌های اصلی دارم؟
  • آیا deploy تدریجی (canary/blue-green) با rollback خودکار است؟
  • آیا health check واقعی است (نه فقط «process زنده است» بلکه «به dependencyها می‌رسم»)؟
  • آیا هزینه‌ی این سرویس (compute + شبکه + observability) را می‌دانم؟
حقیقتِ نهاییِ سنیور بودن

سنیور بودن یعنی ساده نگه داشتنِ چیزها تا جای ممکن، و پیچیده کردن فقط وقتی که یک نیازِ واقعی مجبورت کند — و توانِ توضیحِ این‌که چرا. مهندسِ متوسط پیچیدگی اضافه می‌کند تا باهوش به نظر برسد؛ سنیور پیچیدگی کم می‌کند و شهامتِ گفتنِ «این را لازم نداریم» را دارد. بهترین معماری آن نیست که بیشترین الگو را دارد، بلکه آن است که تیم بتواند ساعتِ ۳ صبح بفهمدش و تعمیرش کند.

جمع‌بندی

سنیور بودن دانستنِ الگوهای بیشتر نیست، قضاوتِ بهتر است. آنتی‌پترن‌ها را بشناس — distributed monolith (هزینه بده، فایده نگیر)، chatty services (N+1 روی شبکه)، shared DB (coupling پنهان)، nano/entity services (ریزکردنِ بی‌جا) — و بدان که میکروسرویس یک ابزارِ سازمانی است که فقط وقتی چند تیم داری معنی می‌دهد. در پروداکشن: مشاهده‌پذیری را با correlation ID و RED/USE و کوانتایلِ p99 جدی بگیر (نه میانگین)؛ در حادثه اول بازیابی کن بعد ریشه‌یابی؛ با error budget بین سرعت و پایداری تعادل بساز؛ با timeout/circuit breaker/bulkhead و back-pressure از فروپاشی جلوگیری کن؛ dependencyها را بحرانی/غیربحرانی برچسب بزن و graceful degradation بساز؛ هزینه و cardinality را فراموش نکن؛ و داده را تدریجی و برگشت‌پذیر با expand/contract و outbox مهاجرت بده. در اینترویو بلند فکر کن، از ساده شروع کن، فرض‌هایت را اعلام کن، و trade-off را نام ببر. و همیشه یادت باشد: مونولیتِ خوب از میکروسرویسِ بد بهتر است.

Let's start with an uncomfortable truth: most of what you learned in tutorials and courses is either wrong or incomplete in production. Not because the instructor lied, but because a tutorial shows the happy path, and production is where everything breaks at once: the network gets slow, the database locks under pressure, a downstream service times out, a deploy stalls halfway, and right at that moment the pager goes off at 3 a.m.

A senior isn't someone who has memorized every pattern. A senior is the person who, when the team excitedly says "let's make this a microservice," asks "why? what problem does it solve? have you counted the cost?" — and when everyone says "monoliths are bad," calmly replies "a good monolith beats a bad microservice."

This chapter isn't going to teach you a new pattern. It's going to give you judgment: when something is good, when that same thing is a disaster, and how to articulate all of this loudly and clearly in an interview and in a design meeting.

This chapter's roadmap
  1. Anti-patterns — distributed monolith, chatty services, shared DB, nano-services, entity services, and "when microservices are wrong." 2) Production concerns — real observability (correlation IDs, RED/USE, tracing), on-call and incident basics, capacity and back-pressure, graceful degradation, cost, and cross-service data migration. 3) System-design interview — how to think out loud. 4) A senior readiness checklist. All with real Spring Boot code and diagrams.
Versions (mid-2026)

Examples target Spring Boot 3.5.x (the 3.5.0 line from May 2025 through 3.5.7 in Oct 2025) and Spring Cloud 2025.0.x "Northfields". The next line, Spring Cloud 2025.1 "Oakwood", sits on Spring Framework 7 and Spring Boot 4. For resilience we use resilience4j-spring-boot3 version 2.2.x, and for tracing Micrometer Tracing with either the Brave or the OpenTelemetry bridge. The principles here are version-independent; versions matter only for exact code.


Part One: Anti-patterns — where teams fall down

Before you learn what to do, you must know what not to do. I've seen each anti-pattern below set a real team back by months. Learn their names; because in an interview, when you can name a bad decision with the right term, you sound like a senior.

1) The Distributed Monolith

A soccer team roped together at the ankles

Imagine eleven players, but each player's ankle is tied to their neighbor's with rope. On paper they're eleven independent people, yet none can take a single step unless everyone moves at once. That's the distributed monolith: seemingly separate services, so tangled that every change forces you to deploy them all together.

The distributed monolith is the worst of both worlds: you pay the operational cost of microservices (network, separate deploys, distributed debugging) but reap none of the benefits (independent deployability, independent scaling, fault isolation). Symptoms:

  • To release one feature, you must deploy several services together in a specific order.
  • Changing one field in one service breaks the build of three others.
  • Services call each other in a synchronous chain: A waits on B, B waits on C, C waits on D.
  • Everyone depends on one shared library for data models, so updating that library means updating everyone.
The shared-model shackle

The most common root of a distributed monolith is a common-domain module every service depends on. The moment Order, User, and Product live in one shared jar, no service can evolve its model independently. Sharing a contract (schema/API) is good; sharing domain code is poison. Each service should own its own DTOs, even if they look duplicated. Controlled duplication beats hidden coupling.

The right way: draw boundaries around business capabilities, not technical layers. Make inter-service communication asynchronous and event-based wherever possible, to break the synchronous chain.

Let's see the difference in topology.

Diagram: distributed monolith (synchronous chain) vs. event-driven independence — تفاوت زنجیره‌ی همگام با استقلالِ رویدادمحور.

flowchart LR
  subgraph Bad[Distributed Monolith]
    A1[Order] --> B1[Inventory]
    B1 --> C1[Pricing]
    C1 --> D1[Notification]
  end
  subgraph Good[Event-Driven]
    A2[Order] -- OrderPlaced --> K[(Event Bus)]
    K --> B2[Inventory]
    K --> C2[Pricing]
    K --> D2[Notification]
  end

On the left, total latency is the sum of everyone's latency, and if Notification falls over, the whole order-placement falls over. On the right, Order just publishes an event and responds immediately; the others consume in their own time, and one failure doesn't take the rest down.

2) Chatty Services — the same N+1, now over the network

A phone call redialed for every word

Suppose you want to tell your friend a ten-word sentence, but instead of one call you dial ten times and say a single word each time, hanging up between. The "call setup" cost multiplies by ten. A chatty service is exactly this: to build one response, it calls another service dozens of times.

In a monolith, a method call is a few nanoseconds. Over the network, each call is hundreds of microseconds to several milliseconds — thousands of times more expensive, and each one can fail. The code below looks innocent but is a killer in production:

// Anti-pattern: N+1 over the network
List<Order> orders = orderClient.findRecent(userId);      // 1 network call
for (Order o : orders) {
    // one call per order — 50 orders means 50 round-trips
    Product p = productClient.getById(o.productId());     // N network calls
    o.setProductName(p.name());
}

If each call is 20ms and the list has 50 items, this loop alone adds 1 second of latency — and that's optimistic. The fix is to design APIs for batch and coarse-grained responses:

// Right: one batch call instead of N
List<Order> orders = orderClient.findRecent(userId);
Set<Long> ids = orders.stream().map(Order::productId).collect(toSet());
Map<Long, Product> products = productClient.getByIds(ids);   // 1 call
orders.forEach(o -> o.setProductName(products.get(o.productId()).name()));
Senior judgment: a service boundary is a conversation boundary

A practical heuristic: if two services talk constantly and voluminously (chatty), they probably shouldn't be two services. The right boundary is where cross-traffic is low and coarse-grained, and intra-traffic is high. That's "high cohesion, low coupling" again — but this time measured in network cost. If you must coordinate 5 services in one transaction for a single operation, your boundaries are wrong.

3) The Shared Database

One shared ledger for the whole family

Picture the whole family writing in a single account ledger. It's fine while everyone is polite; but one day someone changes the "date" column and suddenly everyone else's math breaks, without knowing when or why. A database shared between microservices is exactly this: nobody owns the schema, so nobody can change it without fear.

The golden rule of microservices: each service owns its database. Not just separate tables, but ideally a separate schema or instance, and no service ever touches another service's tables directly. Why so strict?

  • Independent evolution: if service A can query B's table, then B can never change that table's structure without fearing it breaks A. The schema becomes an unintended public API — the worst kind of coupling, because it's hidden.
  • Data ownership: business invariants must be enforced in one place. If two services both write to the account table, the "balance must not go negative" rule is duplicated in two places, and one day one of them forgets it.
The JOIN temptation

The hardest thing you lose when splitting databases is JOIN. Under pressure for "just one reporting query," teams often reconnect the databases and everything rots. The right way to do cross-service reporting is to build a separate read model (CQRS) populated by consuming events, or a data warehouse/lake that services push data into — not a direct JOIN against services' operational databases.

For cross-service reporting, the common pattern is that each service ships its events into a data pool. A reporting query against a read model, working in both dialects:

-- PostgreSQL: pagination over a read model
SELECT order_id, product_name, total
FROM   reporting.order_summary
WHERE  created_at >= now() - interval '7 days'
ORDER  BY total DESC
LIMIT  20 OFFSET 40;
-- Oracle: same query, standard pagination
SELECT order_id, product_name, total
FROM   reporting.order_summary
WHERE  created_at >= SYSTIMESTAMP - INTERVAL '7' DAY
ORDER  BY total DESC
OFFSET 40 ROWS FETCH FIRST 20 ROWS ONLY;
The dialect difference you must know

Pagination is the biggest divergence: PostgreSQL uses LIMIT n OFFSET m, while Oracle (12c and later) uses OFFSET m ROWS FETCH FIRST n ROWS ONLY. The old Oracle ROWNUM trick is still everywhere in legacy code but is error-prone because ROWNUM is applied before ORDER BY. For time ranges, PostgreSQL uses now() - interval '7 days' and Oracle uses SYSTIMESTAMP - INTERVAL '7' DAY. If your code must run on both, hide these differences behind an abstraction layer (a query builder or jOOQ).

4) Nano-services — when you slice too thin

A nano-service is one so small that the cost of its existence (a separate process, network, deploy, monitoring, database) exceeds the value it produces. A service that does one function — say "currency conversion" or "date formatting" — should probably be a library, not a service.

The hidden per-service tax

Every microservice carries a "fixed tax" that has nothing to do with its logic: a CI/CD pipeline, a monitoring dashboard, alarms, on-call, version management, health checks, a database, and mental space the team must spend understanding it. With 200 nano-services, that tax suffocates you. A team of 5 can't keep 40 services healthy. Rule of thumb: the number of services shouldn't wildly exceed the number of people on the team.

5) Entity Services (CRUD-as-a-Service)

This one is subtler. An entity service is built around a table, not a capability: "UserService," "OrderService," "ProductService," each doing only CRUD on one entity, while all the business logic lives in an "orchestrator" on top of them.

The problem: the business logic isn't in any service — it's between them. Every real operation becomes chatty calls and distributed transactions. This is effectively a distributed monolith wearing a "clean" mask.

Draw the boundary around behavior, not data

Instead of a "UserService" that only does CRUD, think about the capability: "Onboarding," "Billing," "Fulfillment." A good service makes a decision and protects an invariant; it doesn't just read and write a table. If your service's name matches a table's name, that's a warning bell. A service name should evoke a business verb, not a data noun.

When are microservices simply wrong?

A truth that's rarely stated

Microservices are an organizational tool, not a technical one. The problem they solve is: "multiple teams want to deploy independently without stepping on each other." If you don't have that problem, microservices probably won't solve your problem — they'll just create new ones.

Microservices are usually wrong when:

  • The team is small (under ~15–20 people): a well-built modular monolith with clean internal boundaries gives you 90% of the benefit at 10% of the cost.
  • The domain is still immature: if you don't yet know where the right boundaries are (and in an early startup you never do), you'll cast wrong service boundaries in concrete. Moving a boundary in a monolith is a refactor; in microservices it's a multi-month project.
  • You lack operational maturity: without automated CI/CD, good observability, and an on-call culture, microservices just mean distributed darkness.
Why is "monolith-first" often the right advice?

Answer: Because the biggest risk in microservices is drawing the wrong boundaries, and at the start you have the least knowledge about the right ones. A modular monolith lets you move boundaries cheaply until the domain matures and you understand which parts truly need to be independent. Then you can extract those same modules — which now have stable boundaries — into services. Martin Fowler calls this "monolith-first." In an interview I'd add: "Microservices are an organizational optimization; until you feel the organizational pain — multiple teams, deploy contention — it's premature."


Part Two: Production concerns — where a senior really shines

This is where the line is drawn between "someone who writes code" and "someone who keeps a system alive." Writing code is easy; keeping the system up at 3 a.m., when you're missing half the information, is an art.

Observability in practice

Observability is the ability to answer questions you didn't know in advance you'd need to ask, purely from the system's outputs. The three classic pillars: logs (what happened), metrics (how much, how many, how fast), traces (where one specific request went).

The aircraft black box

Monitoring is like a pilot's dashboard warning lights: you knew in advance what might break and put a light there for it. Observability is like the black box: after an incident, from recorded data you can ask any question, even one you never thought of before the crash. In a distributed system you need both, but observability is what saves you on the night of the incident.

Correlation IDs and distributed tracing

The single biggest debugging pain in a distributed system is this: a user says "my order didn't go through," and you must trace that one request across the logs of 8 services. Without a shared identifier, that's impossible.

A correlation ID (or trace ID) is a unique identifier generated at the edge of the system (the gateway) that travels with every subsequent call — HTTP or message-on-a-queue alike. Now a single grep shows you that request's entire journey.

The good news: in modern Spring Boot this is nearly automatic. With Micrometer Tracing, you just add a bridge, and Boot injects traceId and spanId into the logging MDC (Mapped Diagnostic Context) and propagates them across service boundaries.

<!-- Bridge to a tracing system: either Brave/Zipkin or OpenTelemetry -->
<dependency>
  <groupId>io.micrometer</groupId>
  <artifactId>micrometer-tracing-bridge-brave</artifactId>
</dependency>
<!-- exporter: e.g. to Zipkin, or via OTLP to Tempo/Jaeger -->
<dependency>
  <groupId>io.zipkin.reporter2</groupId>
  <artifactId>zipkin-reporter-brave</artifactId>
</dependency>

Configure the log pattern so these IDs are visible:

# application.yml
logging:
  pattern:
    # [appName,traceId,spanId] — Boot reads these from the MDC
    level: "%5p [${spring.application.name:},%X{traceId:-},%X{spanId:-}]"
management:
  tracing:
    sampling:
      probability: 1.0    # everything in dev; usually 0.05–0.1 in prod
100% sampling in production is an expensive mistake

probability: 1.0 traces every request. Great in dev, but in high-traffic production it means an enormous volume of trace data that burns both money (storage cost in Tempo/Datadog) and bandwidth. In production, 5%–10% sampling is usually enough. A senior note: tail-based sampling is better — instead of deciding randomly at the start, hold everything briefly and keep only traces that were slow or had errors. That's done at the collector level (like the OpenTelemetry Collector), not in the application.

When a request moves between services, the RestClient or WebClient that Spring builds automatically carries the propagation headers (like traceparent in the W3C standard) — provided you get that client from Spring's factory bean, not via new. If you build the client by hand, that chain breaks.

Diagram: one shared trace ID flowing through several services — یک trace ID مشترک از میانِ چند سرویس.

sequenceDiagram
    participant C as Client
    participant G as Gateway
    participant O as OrderSvc
    participant P as PaymentSvc
    participant K as Kafka
    C->>G: POST /orders
    Note over G: generate traceId=abc123
    G->>O: traceparent: abc123
    O->>P: traceparent: abc123
    P-->>O: 200
    O->>K: OrderPlaced (traceId=abc123 in header)
    O-->>G: 201
    G-->>C: 201

The golden point: propagation must survive even the message queue. When you publish OrderPlaced to Kafka, put the traceId in the message header so the consumer can continue the trace. Otherwise the request's trail is lost exactly where it goes async.

RED and USE — which metrics do I watch?

A beginner monitors everything and therefore sees nothing. A senior knows which few metrics truly matter. Two standard frameworks:

Framework For what The three metrics
RED (Tom Wilkie, Grafana) service health (user view) Rate, Errors, Duration (latency)
USE (Brendan Gregg) resource health (CPU, disk, pool) Utilization, Saturation (queue), Errors
Four Golden Signals (Google SRE) user view Latency, Traffic, Errors, Saturation
How to use them together

Put RED on each service: "how many requests per second does this service take, what percentage error, and what's its p99 latency quantile?" Put USE on each resource: "how full is the DB connection pool? how long has the thread queue grown?" When an incident hits, RED tells you which service is hurting, and USE tells you why (which resource saturated). This division of labor multiplies your diagnosis speed.

The average lies — always look at quantiles

If you only look at average latency, you'll be fooled. A service can have a 50ms average but a 3-second p99 — meaning one in every 100 users has a catastrophic experience. And in a distributed system where one request hits 10 services, the probability of hitting that slow tail rises sharply. Always monitor p95, p99, p99.9, not the average. This phenomenon is called "tail latency amplification."

In Spring, Micrometer gives these metrics almost for free. Adding micrometer-registry-prometheus creates an endpoint at /actuator/prometheus that exports all HTTP metrics (including http.server.requests with status/uri tags and quantiles):

management:
  endpoints:
    web:
      exposure:
        include: health,info,prometheus,metrics
  metrics:
    distribution:
      percentiles-histogram:
        http.server.requests: true   # histogram for Prometheus-side quantiles
      slo:
        http.server.requests: 100ms,300ms,500ms

Adding a business metric is just as easy — and a senior knows business metrics (like "failed-order rate") are sometimes more important than technical ones:

@Service
public class OrderService {
    private final Counter failedOrders;

    public OrderService(MeterRegistry registry) {
        this.failedOrders = Counter.builder("orders.failed")
            .description("Orders that failed to place")
            .tag("channel", "web")
            .register(registry);
    }

    public void place(Order order) {
        try {
            // ... business logic ...
        } catch (PaymentDeclinedException e) {
            failedOrders.increment();
            throw e;
        }
    }
}

On-call and incident management basics

A day comes when the pager rings and you're responsible. This is taught to engineers less often than anything, yet it's exactly the part that separates "senior" from "good programmer."

The priority order during an incident

When the system is down, the order of operations is: 1) Mitigate — restore service at any cost, even by rollback or killing a feature. 2) Communicate — tell stakeholders what's going on. 3) Root cause — this comes after mitigation, not alongside it. Beginners chase "why" during an incident; a senior restores the system first, then hunts the why. "Recovery precedes diagnosis."

Concepts you must know:

  • SLI / SLO / Error Budget: an SLI is what you measure (e.g. "percentage of requests under 300ms"). An SLO is the target you aim for (e.g. "99.9% of requests under 300ms"). The error budget is the failure allowance derived from the SLO: an SLO of 99.9% means 0.1% error allowed, which is roughly 43 minutes of permitted downtime per month. This number is extraordinarily powerful because it turns the "be more stable vs. ship faster" debate from an emotional argument into a shared number.
The error budget: a senior's negotiating weapon

When the product manager pushes for more speed and the platform team pushes for more stability, the error budget is the referee. If budget remains, you can take risks and deploy faster. If the budget is spent (you've been unreliable this month), pump the brakes and work on stability. This turns endless "quality vs. speed" debates into an objective rule. Say this in a senior interview and you score big.

  • Runbook: a simple document that says "if this alarm fires, do these things." Nobody feels creative on the night of an incident; the runbook is your backup brain.
  • Blameless postmortem: a post-incident meeting with no personal blame. The goal is to find the system's flaw, not a culprit. Because if people are afraid, they hide incidents and the same mistake repeats.
An alarm that always fires isn't an alarm

The worst thing for on-call is "alert fatigue": so many unimportant alarms fire that the brain learns to ignore them — and the day a real alarm comes, it too gets ignored. Rule: every alarm must be actionable and tied to an SLO or a sign of user pain. An alarm on "CPU over 80%" is bad (it might be normal); an alarm on "5xx error rate over 1% for 5 minutes" is good. If you can't write a runbook for an alarm, that alarm shouldn't wake anyone up.

Capacity and back-pressure

A funnel and the overflow bowl

Picture pouring water into a funnel with a hose. If the inflow rate exceeds the outflow, the funnel fills and overflows. A software system is the same: if incoming request rate exceeds your processing capacity, queues fill, memory runs out, and the system doesn't "slow down" — it "dies." Back-pressure is a mechanism that tells upstream "slower, I'm filling up" — instead of silently drowning.

A common fatal error: the unbounded queue. An ExecutorService with an infinite queue slowly devours memory under pressure until OutOfMemoryError and death. The right way is a bounded queue with an explicit rejection policy:

// Right: bounded pool + bounded queue + rejection policy
ThreadPoolExecutor executor = new ThreadPoolExecutor(
    10, 20,                          // core / max threads
    60, TimeUnit.SECONDS,
    new ArrayBlockingQueue<>(500),   // bounded queue — not an infinite LinkedBlockingQueue
    new ThreadPoolExecutor.CallerRunsPolicy()  // back-pressure: the submitter runs it
);

CallerRunsPolicy is a subtle trick: when the queue fills, the thread that submitted the work is forced to run it itself. This effectively slows the accepting thread and naturally reduces the inflow rate — a simple, effective back-pressure.

At the service layer, bulkhead and rate limiter are the main tools. A bulkhead caps concurrent calls to a dependency so that one slow dependency doesn't devour the whole service's thread pool:

@Bulkhead(name = "inventory", type = Bulkhead.Type.SEMAPHORE)
@RateLimiter(name = "inventory")
@CircuitBreaker(name = "inventory", fallbackMethod = "fallbackStock")
public StockLevel checkStock(String sku) {
    return inventoryClient.getStock(sku);
}

public StockLevel fallbackStock(String sku, Throwable t) {
    // graceful degradation: a conservative value instead of an error
    return StockLevel.unknown(sku);
}
resilience4j:
  bulkhead:
    instances:
      inventory:
        max-concurrent-calls: 25       # at most 25 concurrent calls
  ratelimiter:
    instances:
      inventory:
        limit-for-period: 100          # 100 calls
        limit-refresh-period: 1s       # per second
        timeout-duration: 0            # if quota is exhausted, reject immediately (fail-fast)
  circuitbreaker:
    instances:
      inventory:
        sliding-window-size: 20
        failure-rate-threshold: 50     # if 50% of calls fail, open the circuit
        wait-duration-in-open-state: 10s
Chain your timeouts, or they're useless

A classic mistake: service A has a 10-second timeout, but B, which A calls, has a 30-second timeout. The result: A gives up and disconnects, but B is still doing useless work and burning resources (orphaned work). Rule: the upstream timeout must be shorter than the downstream timeout, and each layer should propagate its time budget to the next (deadline propagation). Without correct timeouts, even a circuit breaker is useless, because the thread is blocked before the circuit notices.

Diagram: the circuit breaker's state lifecycle — چرخه‌ی حالتِ circuit breaker.

stateDiagram-v2
    [*] --> Closed
    Closed --> Open: failure rate > threshold
    Open --> HalfOpen: after wait duration
    HalfOpen --> Closed: trial calls succeed
    HalfOpen --> Open: trial calls fail
    note right of Open
        fail fast, run fallback
    end note

Graceful degradation

Broken elevator, working stairs

If a building's elevator breaks, the building doesn't shut down; people use the stairs. The experience is worse, but the building works. Graceful degradation means: when part of the system fails, instead of dying completely, you serve at lower quality. A senior always asks: "if this dependency fails, can the user still do their core task?"

Real examples:

  • The "personalized recommendations" service is down? Instead of an error, show best-sellers.
  • The real-time pricing service got slow? Show the cached price from a few minutes ago with a banner: "prices may change."
  • The image service is down? Show a placeholder, but keep the buy button alive.
Label every dependency "critical" or "non-critical"

A senior exercise: for each service, split its dependencies into two buckets. Critical (if it fails, the core operation is impossible — like the order database) and non-critical (if it fails, only the experience worsens — like the recommendation service). For non-critical dependencies, the whole request must never fail. Do this labeling at design time, not after your first incident. A fallback on every non-critical call is the difference between "the site is slow" and "the site is down."

Cost — the dimension engineers forget

A senior knows every architectural decision has a price tag, and someone must pay it. Microservices are inherently more expensive than a monolith:

  • Network cost: every inter-service call is money (especially cross-availability-zone traffic in the cloud, which incurs egress cost).
  • Duplication cost: each service wants its own CPU/RAM/database. 20 small services are often more expensive than one dense monolith, because each has a baseline footprint (JVM minimum heap, sidecar, etc.).
  • Observability cost: this is the big surprise. Log, trace, and metric volume in a distributed system is explosive, and APM tools (like Datadog) charge by volume. I've seen a team whose observability bill exceeded its compute bill.
The silent bill killer: unbounded logs and traces

Three main sources of hidden cost: (1) 100% trace sampling, (2) DEBUG-level logging in production, (3) high-cardinality metrics — e.g. putting userId as a tag on a metric. Every unique tag value creates a new time-series; if you tag userId (millions of values), your metrics system explodes (cardinality explosion) and your cost skyrockets. Tags must be low-cardinality: status, region, endpoint — never unique identifiers.

Cross-service data migration

The hardest operational task in the microservices world is this: you want to move a table, or ownership of some data, from service A to service B — with no downtime. This is where even fresh seniors fall down.

The golden rule: never big-bang

Never migrate data in one big blackout ("we'll take it down Friday night, run the migration, and be up Saturday"). That almost always goes wrong and has no rollback. The right pattern is a gradual, reversible migration. The most famous is expand–migrate–contract (a.k.a. parallel-change).

The expand/contract pattern for zero-downtime schema change has three phases:

  1. Expand — add the new structure alongside the old (new column, new table). Nothing is deleted yet. The code writes to both (dual-write) but still reads from the old.
  2. Migrate — gradually backfill existing data into the new structure, and switch the code to read from the new. Now both old and new are up to date.
  3. Contract — once you're sure everything reads and writes from the new, drop the old structure.

Each phase deploys and rolls back on its own. The key is that no single deploy changes both schema and code incompatibly at once.

For moving data between two different services (not just a schema), the common pattern is Change Data Capture (CDC) or Outbox: the source service publishes data-change events, the target service consumes them and builds its own state, until they're aligned and you can switch traffic.

Diagram: zero-downtime ownership move via dual-write + backfill — مهاجرتِ مالکیتِ داده با dual-write و backfill.

flowchart TD
    subgraph Phase1[Expand: dual-write]
        W[Write path] --> Old[(Old Store)]
        W --> New[(New Store)]
    end
    subgraph Phase2[Migrate: backfill + read new]
        B[Backfill job] --> New
        R[Read path] --> New
    end
    subgraph Phase3[Contract: drop old]
        R2[All traffic] --> New
    end
    Phase1 --> Phase2 --> Phase3

SQL for the backfill phase, which must be idempotent and batched (so it doesn't lock the live table):

-- PostgreSQL: idempotent upsert in batches
INSERT INTO new_service.customer (id, email, tier)
SELECT id, email, tier FROM old_service.customer
WHERE  id BETWEEN :lo AND :hi
ON CONFLICT (id) DO UPDATE
   SET email = EXCLUDED.email, tier = EXCLUDED.tier;
-- Oracle: same idempotency with MERGE
MERGE INTO new_service.customer t
USING (SELECT id, email, tier FROM old_service.customer
       WHERE id BETWEEN :lo AND :hi) s
ON (t.id = s.id)
WHEN MATCHED THEN UPDATE SET t.email = s.email, t.tier = s.tier
WHEN NOT MATCHED THEN INSERT (id, email, tier)
     VALUES (s.id, s.email, s.tier);
Upsert: ON CONFLICT vs. MERGE

PostgreSQL uses INSERT ... ON CONFLICT ... DO UPDATE for upsert, which is concise and atomic. Oracle has no direct equivalent and uses MERGE (which PostgreSQL also has since v15, but ON CONFLICT is more common and smoother). Always run the backfill in small batches with BETWEEN :lo AND :hi and pause between batches; otherwise one UPDATE over millions of rows creates replica lag and locks — you'd take down the very thing you were trying to save.

Dual-write without a guarantee is a source of inconsistency

Dual-write (writing to two stores at once) has a trap: if the write to the first store succeeds but the second fails, the two data sources become inconsistent, and there's no atomic transaction between them (they're separate systems). The safe way is to do one of the two writes via an outbox: in the same database transaction as the first write, insert a row into an outbox table, and a relay ships it to the second store. This preserves atomicity. Direct dual-write is only acceptable if a reconciliation job regularly finds and fixes inconsistencies.


Part Three: How to think out loud in a system-design interview

The system-design interview doesn't measure your knowledge; it measures your judgment and thought process. The interviewer wants to see how you clarify, decompose, and trade off an ambiguous problem. Going silent and jumping straight to a solution is the worst move.

A 6-step framework for thinking out loud
  1. Clarify — ask questions; scope the problem. 2) Estimate scale — how many users, how many QPS, how much data? 3) APIs and data model — define the contracts. 4) High-level design — draw a box diagram. 5) Deep dive — open up one hard part. 6) Trade-offs and bottlenecks — state what you sacrificed and why. Key: throughout all of this, talk out loud and state your assumptions.
Your first sentence must not be a technology name

A beginner says "I'll use Kafka and Redis and Cassandra." A senior first asks "what's the read-to-write ratio? how many concurrent users? do we need strong consistency, or is eventual enough?" Technology is the consequence of requirements, not the start of the discussion. If you name technology first, the interviewer thinks you only know buzzwords. Understand the problem first, then pull the tool out of the trade-off.

A concrete example: if asked to "design a URL shortener," the senior path is:

  • "How many URLs are created per day? How many reads?" (usually reads far exceed writes) → so the system is read-heavy → cache and read replicas matter.
  • "Does the short length matter? Do links expire?" → the data model and TTL clarify.
  • "Consistency: what if two people grab the same code at once?" → the key-generation strategy (central counter vs. hash).
The trap: premature complexity

A common candidate mistake is drawing a spacey 20-service architecture from the start to sound "senior." It has the opposite effect. A real senior starts with the simplest possible design and adds complexity only when a specific requirement (scale, availability) forces it. The golden line in an interview: "The simplest thing that works is this; now let's see where it breaks under pressure and complicate only that part."

How do you decide between strong and eventual consistency?

Answer: I look at the business consequence of stale data. If stale data causes a financial error or violates an invariant — like an account balance or a seat reservation — strong consistency is required, even at the cost of higher latency and lower availability. If stale data is merely slightly annoying — like a post's like count or a recommendation list — eventual consistency is perfectly fine and gives me better scale and availability. Per the CAP theorem, in a distributed system during a partition you must choose between consistency and availability; but in practice it isn't a binary choice, it's a spectrum you decide per data type. A real system is usually strong for the money part and eventual for the social part.

A user request got slow and 8 services are involved. How do you root-cause it?

Answer: First I pull that request's distributed trace by trace ID — this immediately shows which span took the most time, so I don't search, I see where the bottleneck is. Then, on that service, I look at USE metrics: is the DB connection pool saturated? has the thread queue grown? are GC pauses high? If the latency is in a downstream call, I check whether that downstream is really slow, or whether our retries are multiplying its load (a retry storm). Senior note: I always look at p99, not the average, and I watch for tail latency amplification — one slightly-slow service in an 8-deep chain can blow up the overall p99.

What's the difference between orchestration and choreography in a saga, and when do you pick each?

Answer: In orchestration, a central coordinator explicitly calls the saga's steps and manages compensation on failure — logic is centralized and traceable, but that orchestrator becomes a coupling point. In choreography, each service reacts to events and publishes the next event — less coupling, but the overall flow isn't written down anywhere and is hard to follow ("nobody sees the full picture"). My rule: for simple flows with a few steps, choreography is lighter; for complex flows with many conditions and compensations, I choose orchestration because observability and debuggability matter more than reduced coupling. And in both cases, every step must be idempotent.

How do you guarantee exactly-once in a message-driven system?

Answer: Honestly: exactly-once delivery is effectively impossible in a distributed system; what you actually get is "at-least-once delivery + idempotent processing," whose effect is like exactly-once. That is, I accept that a message may arrive several times, but I build the consumer so that duplicate processing is a no-op — with an idempotency key (like the messageId) that I store in the database and check before processing to see if I've already handled it. On the producer side, I use the transactional outbox pattern so the data write and the event publish are atomic. If I said "yes, with such-and-such setting it becomes exactly-once," an experienced interviewer would know I hadn't grasped the depth of the topic.

Your service depends on something that sometimes takes 5 seconds. What do you do?

Answer: Several layers of defense: (1) An aggressive timeout — say 800ms, because it's better to fail fast than block a thread for 5 seconds. (2) A circuit breaker — if the slowness/error rate crosses a threshold, open the circuit to take pressure off the dependency and give it a chance to recover. (3) A bulkhead — cap concurrent calls to this dependency so its slowness doesn't devour my whole thread pool. (4) A fallback — if it's non-critical, return a cached or default response (graceful degradation). (5) At the architecture level I ask: does this call have to be synchronous? Maybe I can make it async and take it off the user's request path. The order matters: timeout first, because without it the rest are useless.

A new deploy raised the error rate, but only for 2% of users. What happened, and how do you react?

Answer: The 2% strongly points to a gradual rollout (canary) — the new version is probably on a subset of instances or users, and those are erroring. Reaction: first mitigate — roll back or dial the canary percentage to zero, not root-cause in parallel. Then I examine the trace and logs of those 2%, filtered by version. Likely causes: an incompatible schema change that only breaks for certain data, a feature flag on for a subset, or a bad node. The senior point I'd make: this is exactly why progressive delivery (canary + automated metrics) is worth it — it shrank the blast radius from 100% to 2% and gave me a chance to roll back before a disaster.

How do you decide whether to make a capability a separate service or keep it in an existing one?

Answer: I look at several axes: (1) Team boundary — is a separate team going to own this capability and want to deploy independently? That's the strongest reason. (2) A different scaling profile — does this part have a completely different scaling need (e.g. image processing that's CPU-bound alongside a light API)? Splitting allows independent scaling. (3) Transactional boundary — if this capability is knotted into one atomic transaction with the rest, splitting it means building a saga and complexity; maybe not worth it. (4) Rate of change — does this part change far more or far less than the rest? If all of these say "split," I split; if I'm unsure, I keep it as a bounded module inside the same service so it can be cheaply extracted later. False coupling is worse than a slightly larger service.

How do you migrate from a monolith to microservices without stopping the business?

Answer: With the Strangler Fig pattern (Fowler's name). Instead of a big-bang rewrite that almost always fails, I put a gateway/proxy in front of the monolith and move capabilities to new services one at a time. The proxy routes traffic for each migrated capability to the new service and everything else to the monolith. Over time, the new services "strangle" the monolith until nothing is left. Key: I start with the capability that has the least coupling with the rest (low-hanging fruit), and I manage the data boundary with dual-write/CDC. I never split everything at once, and I always keep rollback capability at each step. A migration should be a series of small reversible steps, not one big leap.

What does "Kafka guarantees ordering" actually mean? Where does that guarantee break?

Answer: Kafka guarantees ordering only within a single partition, not across the whole topic. Messages that go to one partition are read in write order. So to preserve the order of events for one entity (say all events of an orderId), I must send them with the same partition key (that orderId) so they all land in one partition. Where it breaks: (1) if you set no key, messages are spread round-robin and ordering is lost. (2) if you change the number of partitions, the key-to-partition mapping changes and historical ordering is disrupted. (3) if the consumer runs multiple parallel threads, it may process one partition's messages out of order. So ordering is a limited, conditional guarantee, not an absolute one.

Your system must handle 10x the traffic. Where do you start?

Answer: First measure, don't guess. With a load test and current metrics I find where the real bottleneck is — it's almost always one specific resource (often the database, not the app). Then, cheapest to most expensive: (1) cache — if it's read-heavy, this gives the most return. (2) Horizontal scale-out of stateless services — easy if they're truly stateless. (3) Read replicas and connection pooling for the database. (4) If writes are the bottleneck, sharding/partitioning or moving writes onto a queue (async). (5) Back-pressure and rate limiting so that under load the system degrades gracefully instead of collapsing. Senior note: a 10x traffic increase often reveals a non-linear bottleneck that was hidden at normal load — like a database lock or connection pool. So I don't assume what worked at 1x stays linear at 10x.

How do you know you have "too many" microservices?

Answer: Several signs: (1) to add a simple feature you must change and deploy several services in coordination (a distributed-monolith sign). (2) you spend most of your time debugging inter-service communication rather than business logic. (3) the number of services wildly exceeds the number of people on the team and nobody understands the full picture. (4) you have services that are just CRUD on one table (entity services). (5) infrastructure and observability cost is disproportionately high for the size of the business. The cure is usually consolidation — merge several chatty, interdependent services back into one. A senior confession in an interview: "merging services is also an architectural skill, not just splitting; and it takes more courage because it goes against the fashion."


Part Four: The senior readiness checklist

Before you ship a service to production, ask yourself these. If the answer to half of them is "I don't know," you're not ready for on-call yet.

Observability

  • Does every request have a correlation/trace ID that propagates through all logs and across queues?
  • Do I have RED metrics (rate/errors/duration with p99) for each endpoint?
  • Do I have dashboards and actionable alarms tied to an SLO, not noisy alarms?

Resilience

  • Does every external call have a timeout, retry (with backoff and jitter), and circuit breaker?
  • Are timeouts chained (upstream shorter than downstream)?
  • Do I have a fallback for each non-critical dependency?
  • Are queues and pools bounded, and do I have back-pressure?

Data

  • Does each service own its own database, and does no one touch another's tables?
  • Are message consumers idempotent (do they have an idempotency key)?
  • Can schema changes run via expand/contract with no downtime?
  • Is event publishing atomic (transactional outbox), not fragile dual-write?

Operations

  • Do I have runbooks for the main alarms?
  • Is deployment gradual (canary/blue-green) with automatic rollback?
  • Is the health check real (not just "the process is alive" but "I can reach my dependencies")?
  • Do I know this service's cost (compute + network + observability)?
The final truth of being senior

Being senior means keeping things as simple as possible, and complicating only when a real need forces you — and being able to explain why. The average engineer adds complexity to look smart; the senior removes complexity and has the courage to say "we don't need this." The best architecture isn't the one with the most patterns; it's the one the team can understand and fix at 3 a.m.

Wrap-up

Being senior isn't knowing more patterns, it's better judgment. Know the anti-patterns — distributed monolith (pay the cost, get no benefit), chatty services (N+1 over the network), shared DB (hidden coupling), nano/entity services (needless slicing) — and know that microservices are an organizational tool that only makes sense when you have multiple teams. In production: take observability seriously with correlation IDs, RED/USE, and p99 quantiles (not averages); during an incident mitigate first, then root-cause; balance speed and stability with an error budget; prevent collapse with timeout/circuit breaker/bulkhead and back-pressure; label dependencies critical/non-critical and build graceful degradation; don't forget cost and cardinality; and migrate data gradually and reversibly with expand/contract and outbox. In an interview, think out loud, start simple, state your assumptions, and name the trade-off. And always remember: a good monolith beats a bad microservice.