Craft & Process · مهارت و فرایند پایهBeginner ~71 دقیقه مطالعه~68 min read
Agile، Scrum، UML و مهندسی نیازمندیAgile, Scrum, UML & Requirements Engineering
از مقایسهی چرخههای عمر نرمافزار و خوانش انتقادی بیانیهی Agile تا Scrum کامل (نقشها، مصنوعات، رویدادها، DoD و DoR)، برش عمودی و INVEST، تخمین و تلهی velocity، معیارهای DORA و جریان، Kanban و WIP، کار عملی با Jira و JQL، مهندسی نیازمندی با EARS و اولویتبندی، نمودارهای UML و مدل C4 با مثال Mermaid، ADR و runbook، و فرهنگ code review — بههمراه نحوهی تعریف باورپذیر این تجربه در مصاحبه.From comparing software lifecycles and reading the Agile Manifesto critically to Scrum in full (accountabilities, artifacts, events, DoD and DoR), vertical slicing and INVEST, estimation and the velocity trap, DORA and flow metrics, Kanban and WIP limits, hands-on Jira and JQL, requirements engineering with EARS and prioritisation, the UML diagrams that pay off plus the C4 model in Mermaid, ADRs and runbooks, and code review culture — plus how to describe all of it convincingly in an interview.
تقریباً هر آگهی شغلی بکاند یک بند دارد که مهندسها با نگاه سریع از رویش رد میشوند: «آشنایی با Agile و Scrum»، «تسلط بر Jira»، «توانایی تحلیل و مستندسازی نیازمندیها»، «آشنایی با UML». بعد در مصاحبه همان بند تبدیل میشود به سؤالهایی که جواب دادنشان از یک سؤال concurrency سختتر است، چون هیچوقت کسی آنها را به تو یاد نداده — فقط سالها در جلسههایی نشستهای که کسی توضیح نداده چرا برگزار میشوند.
فرق مهندس میانی و ارشد در همین لایه بیشتر از لایهی کد دیده میشود. هر دو میتوانند یک endpoint بنویسند. اما وقتی یک درخواست مبهم از سمت کسبوکار میرسد، یکی میپرسد «چه چیزی را کد بزنم؟» و دیگری میپرسد «مسئلهی واقعی چیست، چطور میفهمیم حل شده، کدام بخشش را میشود همین هفته تحویل داد، و چه چیزی را قربانی میکنیم؟» اینها «کار مدیر» نیست؛ سازمانهای بالغ انتظار دارند خود مهندس نیازمندی را نقد کند، story را برش بزند، معیار پذیرش بنویسد، تصمیمش را مستند کند و در review هم کد و هم آدم را بهتر کند.
اول چرخههای عمر تولید نرمافزار را مقایسه میکنیم و میبینیم Agile از دل چه دردی بیرون آمد؛ بعد بیانیهی Agile را با نگاه انتقادی میخوانیم. سپس Scrum را کامل باز میکنیم: نقشها، مصنوعات، رویدادها و اینکه هرکدام چطور خراب میشوند، بهعلاوهی Definition of Done و Ready. بعد backlog (user story، INVEST، معیار پذیرش، برش عمودی)، تخمین (story point، planning poker، velocity و سوءاستفاده از آن) و معیارهایی که واقعاً معنی دارند (DORA و flow metrics). بعد Kanban و جریان کار و نگاهی کوتاه به مقیاسپذیری. سپس Jira بهصورت عملی: issue type، workflow، board، JQL و اتصال commit و PR به تیکت. بعد مهندسی نیازمندی: کشف نیاز، FR و NFR، نوشتن نیازمندی آزمونپذیر با EARS، traceability، اولویتبندی و مهار scope creep. بعد UML — فقط نمودارهایی که برای بکاند ارزش دارند، هرکدام با مثال Mermaid — و مدل C4 و نگاهی به BPMN. بعد مستنداتی که دوام میآورد (ADR، README، runbook، diagrams-as-code)، فرهنگ code review، و در پایان: چطور این تجربه را در مصاحبه باورپذیر بگویی.
۱. چرا «فرایند» یک مهارت مهندسی است
برای ساختن یک آلاچیق در حیاط فرایند لازم نداری. اما اگر بیست نفر روی یک ساختمان کار کنند، لحظهای که لولهکش نداند برقکار کجا کانال گذاشته، دیوار را سوراخ میکند. فرایند یعنی همان قراردادهای حداقلی: چه کسی تصمیم میگیرد، کجا نوشته میشود، از کجا میفهمیم کاری تمام شده. نرمافزار بدتر است، چون ساختمان دیده میشود و نرمافزار نه — تنها راه دیدنش همین مصنوعات فرایندی است.
سه دلیل مهندسی — نه مدیریتی — برای اهمیت فرایند: هزینهی اشتباه در نیازمندی تصاعدی است (سوءتفاهمی که در تحلیل با یک سؤال حل میشد، در production به migration داده و از دست رفتن اعتماد ختم میشود)؛ کار نرمافزاری ذاتاً غیرقطعی است (داری چیزی میسازی که قبلاً نساختهای، پس باید بازخورد کوتاه بگیری و مسیر را اصلاح کنی)؛ و حافظهی سازمانی در آدمها نمیماند (آدمها میروند؛ تیکتها، ADRها، runbookها و تستها میمانند).
۲. چرخهی عمر تولید نرمافزار: از آبشار تا Agile
SDLC یعنی ترتیبی که نرمافزار از ایده تا بازنشستگی طی میکند: نیازمندی، طراحی، پیادهسازی، تست، تحویل، نگهداری. مدلها فقط در نحوهی چیدن و تکرار این مراحل فرق دارند.
۲.۱ آبشار (Waterfall)
مراحل را یکبار و پشتسرهم انجام بده و تا فاز قبلی امضا نشده وارد فاز بعد نشو. نکتهی تاریخی جالب: مقالهی معروفی که این نمودار را کشید، خودش این مدل تکگذر را پرریسک توصیف کرد و پیشنهاد داد حداقل دو بار اجرا شود — ولی نمودار ماند و هشدارش گم شد.
مسیر یکطرفهی آبشار — the one-way waterfall flow.
flowchart LR
A[Requirements] --> B[Design]
B --> C[Implementation]
C --> D[Verification]
D --> E[Deployment]
E --> F[Maintenance]
مشکل بنیادی، زمانبندی بازخورد است: اولین لحظهای که کاربر واقعی نرمافزار را میبیند، بعد از صرف تمام بودجه است.
۲.۲ مدل V و مدل تکراری-افزایشی
مدل V همان آبشار است با یک افزودهی ارزشمند: هر سطح طراحی، جفت تستی دارد که همانجا نوشته میشود (نیازمندی کاربر ← تست پذیرش، طراحی معماری ← تست یکپارچگی، طراحی ماژول ← تست واحد). این ایده حتی در Agile هم فوقالعاده مفید است و در فصل «دیسیپلین تست» عمیقتر آمده.
دو کلمه که مدام قاطی میشوند: Incremental یعنی محصول را تکهتکه تحویل بده (اول ورود، بعد سبد خرید، بعد پرداخت)؛ Iterative یعنی همان تکه را چند بار بهتر کن (اول پرداخت ساده، بعد با retry، بعد با reconciliation). Agile هر دو را با هم استفاده میکند.
حلقهی بازخورد تکراری — the iterative feedback loop.
flowchart LR
P[Plan a thin slice] --> B[Build]
B --> R[Release to real users]
R --> M[Measure and learn]
M --> P
۲.۳ مقایسهی صادقانه
| معیار | آبشار / V | تکراری-افزایشی | Agile (Scrum/Kanban) |
|---|---|---|---|
| فرض دربارهی نیازمندی | از ابتدا معلوم و پایدار | تا حد زیادی معلوم | مبهم و در حال تغییر |
| اولین بازخورد واقعی | انتهای پروژه | انتهای هر iteration بزرگ | هر یک تا چهار هفته |
| برخورد با تغییر | فرایند رسمی change request | قابل جذب با هزینه | ورودی طبیعی فرایند |
| نقطهی قوت | قرارداد، ممیزی، سیستم ایمنیبحرانی | پروژهی بزرگ با دامنهی روشن | محصول و بازار متغیر |
| نقطهی ضعف | ریسک متمرکز در انتها | چرخهی هنوز بلند | بدون دیسیپلین فنی به هرجومرج میرسد |
| «تمامشده» یعنی | امضای فاز | تحویل iteration | افزایش قابل انتشار (DoD) |
Agile وقتی برنده است که بازخورد ارزان و تغییر محتمل باشد. اگر بازخورد گران باشد (سختافزار، firmware دستگاه پزشکی، قرارداد ثابتقیمت با جریمه، الزام ممیزی)، برنامهریزی جلوتر ارزش خودش را دارد. در مصاحبه، کسی که این مرز را میکشد از کسی که فقط «Agile خوب است» میگوید فرسنگها جلوتر است.
۳. بیانیهی Agile، با خواندن انتقادی
بیانیه (۲۰۰۱) چهار جملهی مقایسهای دارد، و نکتهای که اکثراً جا میاندازند در جملهی پایانی است: «در حالی که به موارد سمت راست ارزش قائلیم، به موارد سمت چپ بیشتر ارزش میدهیم.» یعنی سمت راست بیارزش نیست.
| به این بیشتر ارزش میدهیم | تا به این | ترجمهی عملی |
|---|---|---|
| افراد و تعاملات | فرایندها و ابزارها | Jira جای گفتوگو را نمیگیرد؛ ولی گفتوگوی بدون رد کتبی هم گم میشود |
| نرمافزار کارکننده | مستندات جامع | مستند کم و زنده، نه «هیچ مستندی» |
| همکاری با مشتری | مذاکره بر سر قرارداد | زودتر نشان بده تا بحث کمتر شود |
| پاسخ به تغییر | پیروی از برنامه | برنامه داشته باش، ولی به آن وفادار نباش |
محتوای واقعی در دوازده اصل پشت بیانیه است. چهار اصلی که مهندسها بیشتر نادیده میگیرند: «تحویل مکرر نرمافزار کارکننده با ترجیح بازهی کوتاهتر»؛ «توجه مستمر به برتری فنی و طراحی خوب، چابکی را افزایش میدهد» (یعنی کیفیت فنی جزء تعریف Agile است، نه قربانی سرعت)؛ «سادگی — هنر بیشینهکردن کاری که انجام نمیشود — ضروری است»؛ و «باید بتوانند سرعت ثابتی را بهطور نامحدود حفظ کنند» (یعنی اضافهکاری مزمن ضدAgile است).
علائم بالینی: daily هست ولی گزارش وضعیت به مدیر شده؛ retro برگزار میشود ولی هیچ اقدامی از آن اجرا نمیشود؛ story point به ساعت تبدیل شده و با آن «تعهد» تیم را میسنجند؛ «افزایش قابل انتشار» تا سه ماه بعد منتشر نمیشود. قاعدهی تشخیص: اگر حذف یک جلسه هیچ تصمیمی را تغییر نمیدهد، آن جلسه بخشی از فرایند نیست، فقط تشریفات است.
تفاوت اصلی زمانبندی بازخورد و برخورد با تغییر است، نه وجود یا نبود مستندات. آبشار فرض میکند نیازمندی از ابتدا معلوم است و ریسک را به انتهای پروژه منتقل میکند؛ Agile فرض میکند نیازمندی حین ساخت کشف میشود و ریسک را با چرخههای کوتاه زودتر آشکار میکند.
بعد نشان بده مطلقگرا نیستی: آبشار وقتی معقول است که تغییر گران یا ممنوع باشد و بازخورد زودهنگام ممکن نباشد — قرارداد با دامنه و قیمت ثابت، الزام رگولاتوری، سیستم ایمنیبحرانی، یا یکپارچهسازی با سامانهای که سالی دو بار پنجرهی انتشار دارد. در عمل بیشتر سازمانها ترکیبی کار میکنند: برنامهی کلان مرحلهای، تحویل داخلی تکراری.
۴. Scrum، کامل و بیرحمانه
Scrum چارچوبی سبک برای تولید محصولات پیچیده است، مبتنی بر empiricism (تجربهگرایی) و lean thinking. سه ستون تجربهگرایی: Transparency (کار و وضعیتش برای تصمیمگیرنده دیده شود)، Inspection (مرتب نگاه کن)، Adaptation (بر اساس آنچه دیدی تغییر بده). هر رویداد Scrum یکی از اینها را نهادینه میکند؛ رویدادی که هیچکدام را انجام نمیدهد باید حذف شود. پنج ارزش هم شعار نیستند: «Courage» یعنی جرئت گفتن «این story آماده نیست» جلوی مدیر.
یک sprint کامل از backlog تا افزایش — one full Scrum sprint.
flowchart LR
PB[(Product Backlog)] --> SP[Sprint Planning]
SP --> SB[(Sprint Backlog)]
SB --> DEV[Sprint work + Daily Scrum]
DEV --> INC[Increment meeting the DoD]
INC --> SR[Sprint Review]
SR --> PB
INC --> RETRO[Sprint Retrospective]
RETRO --> SP
۴.۱ نقشها — که راهنما آنها را accountability مینامد
راهنمای ۲۰۲۰ عمداً کلمهی «role» را کنار گذاشت: اینها عنوان شغلی نیستند، پاسخگوییاند. تیم Scrum یک واحد است، معمولاً ده نفر یا کمتر، بدون زیرتیم.
| پاسخگویی | پاسخگوی چیست | تصمیم نهایی دربارهی | حالت خرابِ رایج |
|---|---|---|---|
| Product Owner | بیشینهکردن ارزش محصول و مدیریت Product Backlog | ترتیب backlog، محتوای آیتمها، Product Goal | «تحویلگیرندهی سفارش» میشود: فقط خواستهها را منتقل میکند بدون نهگفتن |
| Scrum Master | استقرار Scrum و اثربخشی تیم | چگونگی برگزاری رویدادها، برداشتن موانع | تبدیل به مدیر پروژه میشود؛ بهجای تسهیل، کار تخصیص میدهد |
| Developers | ساختن افزایش قابل استفاده و کیفیت آن | چگونگی انجام کار، تخمین، Sprint Backlog | تخمین را از بیرون میپذیرد؛ کیفیت را با فشار زمانی معامله میکند |
اگر Scrum Master تسکها را تخصیص میدهد و در daily گزارش میگیرد، یک مدیر پروژه با اسم تازه داری و خودسازماندهی تیم مرده است. سنجهی موفقیت او این است که با نبودنش هیچ اتفاق بدی نیفتد.
PO پاسخگوی ارزش و اولویت است — «چه چیزی و به چه ترتیبی». او یک نفر است نه یک کمیته، و اختیار «نه» گفتن دارد؛ اگر ندارد، عملاً PO نیست. Scrum Master پاسخگوی اثربخشی فرایند و تیم است و هیچ اقتدار مستقیمی بر محتوای محصول یا تخصیص کار ندارد. Business Analyst در Scrum نقش رسمی نیست؛ در عمل کسی است که در کشف و دقیقکردن نیازمندی کمک میکند.
جملهای که خوب مینشیند: «PO تصمیم میگیرد چه چیزی ارزش دارد، Developers تصمیم میگیرند چطور ساخته شود، Scrum Master مطمئن میشود این گفتوگو سالم بماند.» اگر در تیمت یک نفر هر سه کلاه را داشته، صادق باش و بگو چه مشکلی ایجاد کرد.
۴.۲ مصنوعات و تعهدهایشان
هر مصنوع یک «تعهد» دارد که شفافیت را تضمین میکند: Product Backlog فهرست مرتب و همیشهزندهی هرچه ممکن است لازم باشد، و تعهدش Product Goal است (هدف بلندمدت تیم). Sprint Backlog شامل چراییِ sprint، آیتمهای انتخابشده و برنامهی رسیدن به آنهاست؛ این مصنوع مال Developers است و تعهدش Sprint Goal یعنی یک هدف واحد برای کل sprint. Increment پلهای ملموس بهسوی Product Goal است و تعهدش Definition of Done؛ جملهی صریح راهنما: کاری که DoD را برآورده نکند، اصلاً بخشی از افزایش نیست.
اکثر تیمها sprint را با «۱۴ تیکت نامرتبط» شروع میکنند، یعنی هیچ هدفی ندارند که بشود دربارهاش مذاکره کرد. با Sprint Goal، وقتی کار فوری وارد میشود سؤال ساده است: «آیا این به هدف کمک میکند؟» اگر نه، یا هدف آگاهانه عوض میشود یا کار صبر میکند. یک Sprint Goal خوب نتیجهمحور است — «کاربر بتواند با کارت ذخیرهشده پرداخت کند»، نه «۸ تیکت پرداخت را ببندیم».
۴.۳ رویدادها: تایمباکس، هدف واقعی، حالت خراب
تایمباکسها برای sprint یکماههاند؛ برای sprint کوتاهتر معمولاً کوتاهتر میشوند.
| رویداد | تایمباکس | ستون تجربهگرایی | هدف واقعی | حالت خرابِ رایج |
|---|---|---|---|---|
| Sprint | یک ماه یا کمتر، طول ثابت | ظرف بقیه | آهنگ منظم و محدودکردن ریسک | تمدید sprint برای «تمامکردن کارها» |
| Sprint Planning | حداکثر ۸ ساعت | تطبیق | ساختن Sprint Goal و برنامهی رسیدن به آن | جلسهی تخصیص تیکت بدون هدف |
| Daily Scrum | ۱۵ دقیقه، هر روز کاری | بازرسی | بازبرنامهریزی روزانه توسط Developers | گزارش وضعیت به مدیر؛ سه سؤال طوطیوار |
| Sprint Review | حداکثر ۴ ساعت | بازرسی + تطبیق | بررسی افزایش با ذینفعان و تنظیم backlog | نمایش اسلاید بهجای نرمافزار کارکننده |
| Sprint Retrospective | حداکثر ۳ ساعت | تطبیق | بهبود فرایند، تعامل و ابزار | جلسهی شکایت بدون اقدام پیگیریشده |
Sprint Planning سه سؤال را جواب میدهد: چرا این sprint ارزشمند است، چه چیزی میشود انجام داد، و چطور. اگر بیشتر وقتش صرف چانهزنی بر سر عدد میشود، ریشهی مشکل در planning نیست: backlog بهاندازهی کافی refine نشده — و refinement یک فعالیت مستمر در طول sprint است، نه رویداد رسمی.
Daily Scrum جلسهی گزارش نیست؛ بازبرنامهریزی ۱۵ دقیقهای برای Developers است. سؤال درست «دیروز چه کردی» نیست، بلکه «با آنچه دیروز فهمیدیم، امروز چطور به Sprint Goal نزدیکتر میشویم و چه چیزی جلویمان را گرفته؟» روش خوب: بهجای دور زدن آدمها، دور زدن کارها روی برد از راست به چپ.
Sprint Review جلسهی امضا نیست؛ جلسهی کاری با ذینفعان است که در آن افزایش واقعی نشان داده و backlog همانجا تنظیم میشود. Retrospective الگوی مؤثری دارد: داده جمع کن (نه فقط احساس) ← یک یا دو مشکل انتخاب کن ← ریشهیابی کن ← حداکثر یک یا دو اقدام با صاحب مشخص تعریف کن ← آن اقدام را وارد Sprint Backlog کن تا واقعاً انجام شود. یک اقدام که واقعاً تمام شود، بیشتر از هفت اقدام فراموششده تیم را جلو میبرد؛ و اقدام را مثل یک آیتم کاری با معیار پذیرش بنویس: «تا پایان sprint، اجرای تستهای integration در CI زیر ۱۰ دقیقه باشد» — نه «بهتر تست بنویسیم».
لحظهای که کسی با اقتدار ارزیابی عملکرد بپرسد «چرا این هنوز تمام نشده؟»، آدمها بهجای حقیقت («گیر کردهام»، «تخمینم اشتباه بود») شروع میکنند به گزارشدادن مصونسازانه. آنوقت شفافیت — ستون اول Scrum — از بین میرود و کل سیستم بازخورد کور میشود.
اول شفافسازی، نه قهرمانبازی: همان روزی که فهمیدی — نه روز آخر — در daily مطرح کن. بعد گزینهها را روی میز بگذار: کاهش دامنهی همان هدف با PO (کدام بخش بیشترین ارزش را دارد)؛ حذف آیتمهای غیرمرتبط با هدف از Sprint Backlog؛ بازتعریف Sprint Goal اگر واقعیت عوض شده؛ و در موارد نادر لغو sprint — که طبق راهنما فقط PO اختیارش را دارد و وقتی معنی دارد که هدف منسوخ شده باشد.
چیزی که نباید بگویی: «اضافهکاری میکنیم» یا «تستها را حذف میکنیم». جملهی پایانی خوب: «تاریخ و کیفیت را ثابت نگه میدارم و دامنه را متغیر میکنم؛ دامنه تنها متغیری است که میشود بیخطر کوچکش کرد.»
۴.۴ Definition of Done و Definition of Ready
DoD یک قرارداد کیفیت مشترک است: فهرستی که هر آیتم باید همهاش را داشته باشد تا تمامشده حساب شود. نباید آرزو باشد؛ باید در هر sprint واقعاً رعایت شود.
## Definition of Done — Payments squad
- [ ] معیارهای پذیرش برآورده و توسط یک نفر دیگر بررسی شده
- [ ] کد review شده و حداقل یک approve دارد
- [ ] تست واحد برای منطق جدید + تست integration برای مسیر اصلی
- [ ] pipeline سبز است (build، تست، lint، اسکن وابستگیها)
- [ ] migration پایگاهداده backward-compatible است و rollback تست شده
- [ ] لاگ و متریک مسیر جدید در داشبورد دیده میشود
- [ ] مستند OpenAPI و در صورت نیاز runbook بهروز شده
- [ ] پشت feature flag روی staging منتشر شده
- [ ] هر میانبر عمدی بهصورت تیکت ثبت شده
DoR فهرست شرایطی است که یک آیتم باید داشته باشد تا وارد sprint شود: هدف روشن، معیار پذیرش، وابستگیهای مشخص، دادهی تست در دسترس.
DoR تمرین رایج و مفیدی است ولی جزئی از Scrum نیست، و خطرش این است که به یک آبشار کوچک تبدیل شود: «تا PO این ۹ فیلد را پر نکند شروع نمیکنیم.» DoR را راهنمای گفتوگو نگه دار نه دروازهی قراردادی؛ جواب یک آیتم مبهم معمولاً یک spike کوتاه است، نه رد کردنش.
جواب قوی مشخص و فنی است: DoD را خود Developers تعریف میکنند (اگر سازمان استاندارد حداقلی دارد، تیم فقط میتواند سختگیرانهترش کند نه شلترش)، جایی دیده میشود و در review به آن ارجاع داده میشود. بعد یکی دو بند غیرواضح از DoD واقعیات را بگو — مثلاً «migration باید backward-compatible باشد چون deploy ما rolling است» یا «هر endpoint جدید باید متریک latency و error rate داشته باشد» — و بگو چه اتفاقی باعث شد آن بند اضافه شود. داستان ریشهدار همیشه از فهرست کلی باورپذیرتر است.
۵. Backlog: از epic تا یک آیتم قابل ساخت
User story یک قالب یادآوری برای گفتوگو است، نه قالب مستندسازی: «بهعنوان یک <نقش> میخواهم <قابلیت> تا <ارزشی که به دست میآورم>». مهمترین بخش، «تا ...» است چون چرایی را نگه میدارد و به تیم اجازه میدهد راهحل ارزانتری پیشنهاد دهد. سه بخش یک story را با 3C به یاد میآورند: Card، Conversation، Confirmation.
«بهعنوان یک توسعهدهنده میخواهم Hibernate را ارتقا دهم تا کد بهتر باشد» story نیست؛ یک کار فنی است و اشکالی ندارد همانطور نوشته شود. قاعده: اگر ذینفع نهایی کاربر است story بنویس؛ اگر ذینفع خود سیستم است، آیتم فنی با توجیه ریسک/هزینه بنویس: «ارتقای X برای بستن فلان CVE؛ در صورت انجامندادن از پشتیبانی خارج میشویم».
۵.۱ INVEST
| حرف | معنی | بوی بد وقتی نقض شود |
|---|---|---|
| Independent | تا حد امکان مستقل | «این تا وقتی آن ۳ تا تمام نشود شروع نمیشود» |
| Negotiable | جزئیات قابل مذاکره | story با ۴۰ خط مشخصات از پیش قفلشده |
| Valuable | برای کاربر یا کسبوکار ارزش دارد | «فقط لایهی repository را میسازیم» |
| Estimable | تیم میتواند اندازهاش را حدس بزند | «معلوم نیست» ← به spike نیاز داری |
| Small | در یک sprint تمام میشود | آیتمی که دو sprint باز مانده |
| Testable | معیار پذیرش عینی دارد | «باید سریع باشد» |
۵.۲ معیار پذیرش
سبک فهرستی برای موارد ساده کافی است؛ سبک Gherkin وقتی میارزد که رفتار شرطی داری و میخواهی مستقیم به تست خودکار وصل شود. نوشتنِ این معیارها قبل از کدزدن، معمولاً یکی دو حالت مرزی را همانجا آشکار میکند (پرداخت تکراری؟ کارت منقضی؟ timeout درگاه؟) و ارزانترین باگگیری ممکن است؛ تکنیکهای طراحی تستکیس در فصل «دیسیپلین تست» آمده است.
Feature: Pay with a saved card
Scenario: Successful payment with a valid saved card
Given a customer with a saved card ending in 4242
And an open order of 250000 IRR
When the customer confirms payment with that card
Then the order status becomes PAID
And a payment receipt is stored with the gateway reference
Scenario: Payment declined by the issuer
Given a customer with a saved card that the issuer will decline
When the customer confirms payment with that card
Then the order status stays PENDING_PAYMENT
And no receipt is stored
۵.۳ برش (slicing): مهمترین مهارت backlog
Epic یعنی آیتمی بزرگ که در یک sprint جا نمیشود. مهمترین قاعده: برش عمودی بزن، نه افقی.
تفاوت برش افقی و عمودی — horizontal versus vertical slicing.
flowchart TD
subgraph H[Horizontal slicing - nothing usable until the end]
H1[Sprint 1: all DB tables] --> H2[Sprint 2: all services] --> H3[Sprint 3: all UI]
end
subgraph V[Vertical slicing - usable output every sprint]
V1[Slice 1: pay with new card, happy path] --> V2[Slice 2: saved cards] --> V3[Slice 3: refunds]
end
الگوهای عملی برش (خلاصهی SPIDR و چند مورد دیگر): Spike برای وقتی ندانستن مانع است — کاری زمانمحدود با خروجی تصمیم، نه کد؛ Path فقط یک مسیر (پرداخت با کارت جدید) و بقیه بعداً؛ Interface اول یک رابط مینیمال؛ Data اول یک زیرمجموعه (فقط ریال، بعد چندارزی)؛ Rules اول قواعد ساده (بدون تخفیف، بعد کوپن و سقف)؛ حالت خطا بعداً (اول مسیر خوشبینانه با یک خطای عمومی)؛ و دستی قبل از خودکار (اول گزارش دستی، بعد زمانبندیشده).
اول معیارت را بگو: هر برش باید بهتنهایی قابل انتشار و قابل مشاهده برای کاربر باشد. بعد مثال عینی بزن. برای «افزودن پرداخت آنلاین»: برش اول = پرداخت یک سفارش با کارت جدید، فقط ریال، فقط مسیر موفق، بدون ذخیرهی کارت و بدون استرداد، پشت feature flag و فقط برای کاربران داخلی. این تکه واقعاً پول جابهجا میکند و بازخورد واقعی میدهد. برشهای بعدی: مدیریت خطا و timeout، کارت ذخیرهشده، استرداد، مغایرتگیری.
بعد ریسک را نشان بده: اولین برش را طوری انتخاب میکنم که بزرگترین ناشناخته را زود لمس کند (یکپارچگی با درگاه)، نه سادهترین بخش را. برشزدن فقط برای کوچککردن نیست، برای زودتر آشکارکردن ریسک است.
۶. تخمین: story point، planning poker و تلهی velocity
اگر بپرسم «بلندکردن این وزنه چند دقیقه طول میکشد؟» جواب به آدم بستگی دارد. اگر بپرسم «چند کیلو است؟» جواب برای همه یکسان است. Story point همان کیلوگرم است: اندازهی کار (حجم + پیچیدگی + عدمقطعیت)، نه زمان انجامش.
دلیل ترجیح point بر ساعت: ساعت وابسته به فرد است و همیشه به بحث «تو سریعتری» میرسد؛ انسان در تخمین مطلق بد و در مقایسهی نسبی خوب است («این تقریباً دو برابر آن یکی است»)؛ و ساعت بهسرعت به تعهد قراردادی تبدیل میشود. مقیاس رایج شبهفیبوناچی است (۱، ۲، ۳، ۵، ۸، ۱۳، ۲۰، ۴۰، ۱۰۰ و «؟») و فاصلهی زیادشوندهاش عمدی است: هرچه آیتم بزرگتر، دقت کمتر.
Planning poker یعنی همه همزمان کارتشان را رو میکنند تا لنگرِ نظر اولین نفر ایجاد نشود. ارزش واقعیاش عدد نیست، اختلاف است: وقتی یکی ۲ میدهد و یکی ۱۳، دو تصور کاملاً متفاوت از کار وجود دارد و همان بحث مهمترین خروجی جلسه است.
لحظهای که جملهی «هر point یعنی ۴ ساعت» گفته شود، story point مرده است: ساعت با یک لفافه. نتیجهی قابل پیشبینی این است که تیم برای «رسیدن به تعهد» تخمینها را تورم میدهد یا کیفیت را میبرد. اگر مدیریت زمان میخواهد، جواب درست پیشبینی آماری بر پایهی throughput است، نه تبدیل واحد.
Velocity یعنی مجموع pointهای تمامشده در sprint، و تنها کاربرد مشروعش ظرفیتسنجی همان تیم برای برنامهریزی است. اگر KPI شود، دقیقاً همان چیزی که میسنجی خراب میشود (قانون گودهارت): تخمینها بزرگ میشوند، کارها ناقص Done اعلام میشوند و کیفیت قربانی میشود.
روش بالغتر، پیشبینی بهجای تعهد است: throughput (تعداد آیتم تمامشده در هفته) و شبیهسازی مونتکارلو روی دادهی گذشته که به تو اجازه میدهد بگویی «با احتمال ۸۵٪ بین ۶ تا ۹ هفته» — جملهای صادقانه و قابل دفاع، برخلاف «۷ هفته». اگر آیتمها را کوچک و هماندازه نگه داری، شمردن آیتمها تقریباً بهخوبی جمعزدن pointها کار میکند.
با تفاوت مفهومی شروع کن: point اندازهی نسبی کار است، ساعت مدتزمان است و به فرد و وقفهها وابسته؛ تیم روی اندازه راحتتر توافق میکند تا روی مدت. بعد صادق باش: point فقط ابزار است و اگر آیتمها کوچک و یکنواخت باشند، شمردن آیتمها و پیشبینی با throughput اغلب دقیقتر و کمهزینهتر است.
اگر پرسیدند «پس چطور به کسبوکار تاریخ میدهی؟» بگو با بازهی احتمالاتی بر پایهی دادهی تاریخی و با بهروزرسانی مکرر، نه با یک تاریخ قطعی در ابتدا. و اضافه کن که velocity را هرگز بهعنوان معیار عملکرد بین تیمها گزارش نمیکنی و چرا.
۷. معیارهایی که واقعاً مهماند
Burndown کار باقیمانده را در برابر زمان نشان میدهد؛ burnup کار انجامشده و کل دامنه را جدا نشان میدهد و دقیقاً به همین دلیل بهتر است: اگر دامنه وسط کار بزرگ شود، در burnup بهصورت بالا رفتن خط سقف دیده میشود، در حالی که burndown فقط «تیم کند است» را نشان میدهد.
| دسته | معیار | چه چیزی را واقعاً میگوید | چطور سوءاستفاده میشود |
|---|---|---|---|
| تحویل (DORA) | Deployment Frequency | هر چند وقت یکبار تغییر به production میرسد | deployهای بیمعنا برای بهبود عدد |
| تحویل (DORA) | Lead Time for Changes | از commit تا اجرا در production | اندازهگیری از زمان merge برای زیبا شدن عدد |
| پایداری (DORA) | Change Failure Rate | چند درصد تغییرها به خرابی منجر میشود | ثبتنکردن incidentهای کوچک |
| پایداری (DORA) | Failed Deployment Recovery Time | بعد از خرابی چقدر طول میکشد تا سرویس برگردد | بستن incident قبل از رفع واقعی |
| جریان | Cycle Time | از شروع کار روی آیتم تا تمامشدنش | «شروع» را دیر ثبت کردن |
| جریان | WIP | چند کار همزمان باز است | نادیدهگرفتن کارهای خارج از برد |
| جریان | Work Item Age | آیتمهای در جریان چقدر پیر شدهاند | استفادهنکردن از آن در daily |
| کیفیت | نرخ فرار نقص، تعداد rollback | چقدر مشکل به کاربر میرسد | شمردن باگ بهعنوان معیار عملکرد فرد |
اول تله را خنثی کن: «هیچ عدد واحدی بهرهوری مهندسی را نشان نمیدهد و هر معیاری که به ارزیابی فردی وصل شود خراب میشود.» بعد چارچوب بده: چهار معیار DORA (deployment frequency، lead time for changes، change failure rate، failed deployment recovery time)، بهعلاوهی معیارهای جریان مثل cycle time و WIP، و در نهایت معیارهای نتیجهی محصول که تنها چیزی است که واقعاً اهمیت دارد.
سپس یک مثال بزن: «cycle time ما ۱۱ روز بود؛ با نگاه به cumulative flow دیدیم بیشترش انتظار برای code review است نه کدزدن. با محدودکردن WIP و قاعدهی «review قبل از شروع کار جدید»، به ۴ روز رسید.» این نشان میدهد معیار را برای پیدا کردن گلوگاه استفاده میکنی، نه برای قضاوت آدمها.
۸. Kanban و مدیریت جریان
Kanban روشی برای بهبود جریان کار موجود است، بدون تحمیل نقش یا رویداد جدید. راهنمای رسمی سه تمرین تعریف میکند: تعریف و مصورسازی جریان کار، مدیریت فعال آیتمها در جریان، و بهبود جریان کار.
Definition of Workflow (DoW) حداقل باید مشخص کند: آیتم کاری چیست، نقطهی «شروع» و «پایان» کجاست، چه وضعیتهایی بین آنهاست، WIP چطور کنترل میشود، سیاست صریح هر ستون چیست، و SLE (Service Level Expectation) چقدر است — یعنی «۸۵٪ آیتمها ظرف ۷ روز تمام میشوند». چهار معیار جریانی که راهنما اجباری میداند: WIP، Throughput، Work Item Age و Cycle Time.
بردی با محدودیت WIP — a Kanban board with WIP limits.
flowchart LR
A[Backlog] --> B["Ready (3)"]
B --> C["In Progress (2)"]
C --> D["In Review (2)"]
D --> E["Ready to Deploy (3)"]
E --> F[Done]
قانون لیتل برای یک سیستم پایدار میگوید میانگین زمان چرخه ≈ میانگین کار در جریان ÷ میانگین throughput. یعنی اگر throughput ثابت باشد و WIP را نصف کنی، cycle time نصف میشود — کار سریعتر انجام نمیشود، فقط کمتر منتظر میماند.
Cumulative Flow Diagram تعداد آیتمها در هر وضعیت را در طول زمان بهصورت نوارهای انباشته نشان میدهد. خواندنش ساده است: هر نواری که مدام پهنتر میشود یک گلوگاه است — کار سریعتر از خروج، وارد آن وضعیت میشود. معمولاً آن نوار «In Review» است.
| معیار | Scrum مناسبتر است وقتی | Kanban مناسبتر است وقتی |
|---|---|---|
| ماهیت کار | قابل برنامهریزی برای یک بازه | ورود غیرقابل پیشبینی (پشتیبانی، حادثه) |
| نیاز به تمرکز | هدف مشترک sprint ارزش دارد | اولویتها روزانه تغییر میکند |
| اندازهی آیتم | متغیر، نیازمند برنامهریزی | کوچک و نسبتاً یکنواخت |
| بلوغ تیم | تیم به ساختار و آهنگ نیاز دارد | تیم بالغ که فقط جریان را بهینه میکند |
| خطر اصلی | تبدیل sprint به مینیآبشار دوهفتهای | نامرئیشدن اولویت و بیپایانشدن کار |
در عمل بسیاری از تیمها ترکیب میکنند («Scrumban»): آهنگ و رویدادهای Scrum را نگه میدارند ولی برد را با محدودیت WIP اداره میکنند و بهجای تعهد sprint، جریان را میسنجند. این کاملاً مشروع است، به شرط آگاهانه بودن.
جواب ضعیف: «Scrum، sprint دوهفتهای.» جواب قوی انتخاب را به ماهیت کار وصل میکند: «تیم محصول ما Scrum بود چون میتوانستیم دو هفته را حول یک هدف قفل کنیم؛ ولی تیم on-call با Kanban کار میکرد چون ۷۰٪ ورودیاش غیرقابل پیشبینی بود و تعهد sprint هر هفته میشکست.» اگر تیمت ترکیبی بود همان را بگو — نشانهی بلوغ است، نه بینظمی. و یک عدد اضافه کن: «بعد از گذاشتن محدودیت WIP روی ستون review، cycle time از ۹ روز به ۵ روز رسید.»
۹. مقیاسپذیری: SAFe، LeSS و Nexus در یک نگاه
وقتی چند تیم روی یک محصول کار میکنند، مسئلهی جدید هماهنگی وابستگیها، یکپارچگی مکرر و اولویتگذاری مشترک است. SAFe سنگینترین و پرجزئیاتترین پاسخ است؛ مفاهیمی که باید بشناسی: ART (Agile Release Train، مجموعهای از تیمها که با هم منتشر میکنند)، PI Planning (برنامهریزی مشترک برای یک بازهی چندهفتهای) و WSJF برای اولویتبندی. LeSS حداقلگراست: همان Scrum با یک Product Owner و یک Product Backlog برای همهی تیمها و رویدادهای مشترک — فلسفهاش «مقیاسدادن با حذف، نه با افزودن» است. Nexus لایهی نازکی روی Scrum برای ۳ تا ۹ تیم است با تمرکز بر رفع وابستگیها و یک افزایش یکپارچه.
Squad و Tribe و Chapter و Guild توصیف یک مقطع زمانی از یک شرکت خاص بود، نه روشی قابل کپی — و خود نویسندگانش بعدها هشدار دادند که آن را بهعنوان الگو نفروشند. کپیکردن نمودار سازمانی یک شرکت دیگر، فرهنگ و محدودیتهای آن را با خودش نمیآورد. همچنین دربارهی نسخهها محتاط باش: بگو «تجربهی من با SAFe 6.0 بود» بهجای ادعای آشنایی با آخرین نسخهی روز؛ این چارچوبها مرتب بهروز میشوند.
۱۰. کار عملی با Jira
انتظار واقعی این نیست که مدیر Jira باشی؛ این است که بتوانی تیکت خوب بنویسی، برد را بخوانی، فیلتر بسازی و کدت را به تیکت وصل کنی.
۱۰.۱ انواع issue و workflow
سلسلهمراتب استاندارد از بالا به پایین: Epic ← Story / Task / Bug ← Sub-task (در نسخههای premium سطح بالاتری مثل Initiative هم قابل تعریف است). Epic چتری برای چند آیتم مرتبط است، Story تغییری با ارزش کاربری، Task کاری بدون ارزش مستقیم کاربری (ارتقا، پیکربندی، مهاجرت)، Bug انحراف از رفتار مورد انتظار، و Sub-task تقسیم داخلی برای هماهنگی روزانه.
Workflow یعنی مجموعهی وضعیتها و گذارهای مجاز بین آنها. هر status به یک status category تعلق دارد — To Do، In Progress یا Done — و برد و گزارشها بر همین اساس کار میکنند.
یک workflow واقعی و ساده — a realistic Jira workflow.
stateDiagram-v2
[*] --> Backlog
Backlog --> Ready: refined and accepted
Ready --> InProgress: developer starts
InProgress --> InReview: pull request opened
InReview --> InProgress: changes requested
InReview --> Testing: merged to main
Testing --> InProgress: defect found
Testing --> Done: acceptance criteria verified
Done --> [*]
تیمها برای «شفافیت» ستون اضافه میکنند و نتیجه این میشود که آیتمها روزها در «Ready for QA» میمانند و کسی مسئولش نیست. قاعده: برای هر ستون باید یک سیاست صریح بنویسی (چه وقت وارد میشود، چه وقت خارج، چه کسی مسئول است) و ترجیحاً یک محدودیت WIP بگذاری. اگر نمیتوانی، آن ستون را حذف کن.
در Jira Cloud دو نوع پروژه وجود دارد: team-managed (تنظیمات مستقل و ساده، مدیریتشده توسط خود تیم) و company-managed (طرحهای مشترک workflow و field در کل سازمان، انعطاف کمتر ولی یکپارچگی بیشتر). همچنین Atlassian فیلد قدیمی Epic Link را با فیلد یکپارچهی parent جایگزین کرده است؛ در JQLهای جدید از parent استفاده کن و بدان که رفتار عملگرهای منفی مثل != با این تغییر میتواند نتایج متفاوتی بدهد.
۱۰.۲ JQL
ساختار پایه field operator value است، با ترکیب AND / OR / NOT و در انتها ORDER BY.
project = PAY AND sprint in openSprints() AND assignee = currentUser() ORDER BY rank ASC
project = PAY AND statusCategory != Done AND priority in (Highest, High) ORDER BY created ASC
project = PAY AND status changed to "In Progress" after -7d ORDER BY updated DESC
project = PAY AND parent = PAY-1200 AND type in (Story, Bug)
| نیاز | عبارت |
|---|---|
| کارهای باز خودم | assignee = currentUser() AND statusCategory != Done |
| آیتمهای sprint جاری | sprint in openSprints() |
| آیتمهای sprintهای بسته | sprint in closedSprints() |
| زیرمجموعهی یک epic | parent = PAY-1200 |
| اعضای یک گروه | assignee in membersOf("payments-team") |
| ساختهشده از ابتدای هفته | created >= startOfWeek() |
| سررسید تا پایان امروز | due <= endOfDay() |
| لینکشده به یک تیکت | issue in linkedIssues(PAY-1200) |
| در نسخهی منتشرنشده | fixVersion in unreleasedVersions(PAY) |
| تغییر وضعیت در یک بازه | status changed to Done during (startOfMonth(), endOfMonth()) |
| متن آزاد | text ~ "timeout" |
| بدون تخمین | "Story Points" is EMPTY |
| راکد بیش از ۵ روز | statusCategory = "In Progress" AND updated <= -5d |
۱۰.۳ اتصال کد به تیکت
کلید تیکت را در نام شاخه، پیام commit و عنوان PR بگذار تا ابزارها خودشان همه چیز را وصل کنند.
# نام شاخه شامل کلید تیکت
git switch -c PAY-1234-saved-card-payment
# پیام commit با کلید تیکت در ابتدا
git commit -m "PAY-1234 add saved-card payment flow"
# smart commit: ثبت زمان، کامنت و انتقال وضعیت در یک commit
git commit -m "PAY-1234 #time 3h 30m #comment gateway retry added #review"
نکات smart commit: کلید تیکت باید قالب استاندارد داشته باشد (دو حرف بزرگ یا بیشتر، خط تیره، عدد — مثل PAY-1234) و در ابتدای پیام بیاید؛ #comment متن بعد از خود را کامنت میکند؛ #time زمان صرفشده را ثبت میکند؛ و برای انتقال وضعیت، بخش اول نام گذار کافی است (#review برای گذاری به نام «review changes»). این قابلیت باید در سطح سازمان فعال و به یک ابزار متصل به Jira وابسته باشد.
۱۰.۴ یک تیکت خوب چه شکلی است
# PAY-1234 — Pay an order with a saved card
## Context
۳۸٪ کاربران در مرحلهی وارد کردن اطلاعات کارت رها میکنند (داشبورد funnel، ۳۰ روز اخیر).
## Goal / Value
کاهش رها کردن پرداخت با حذف نیاز به تایپ مجدد کارت برای مشتریان بازگشتی.
## Scope
- انتخاب کارت ذخیرهشده در صفحهی پرداخت و پرداخت با token کارت
خارج از دامنه: افزودن کارت جدید در همین صفحه، کیف پول، پرداخت اقساطی.
## Acceptance criteria
1. کاربر با حداقل یک کارت ذخیرهشده، فهرست کارتها را با چهار رقم آخر میبیند.
2. پرداخت موفق ← وضعیت سفارش PAID و رسید با شناسهی درگاه ذخیره میشود.
3. رد شدن توسط بانک ← وضعیت سفارش تغییر نمیکند و پیام قابل فهم نمایش داده میشود.
4. timeout درگاه ← تراکنش idempotent قابل تکرار است و پرداخت دوباره کسر نمیشود.
## Non-functional
- p95 زمان پاسخ endpoint پرداخت زیر ۸۰۰ms در بار عادی
- هیچ دادهی کامل کارت در لاگ نوشته نشود
## Links
طراحی: <link> · سند درگاه: <link> · وابسته به PAY-1187
سه ویژگی که این تیکت را خوب میکند: چرایی با عدد، مرز صریح دامنه (چه چیزی خارج است)، و معیار پذیرش آزمونپذیر شامل حالتهای خطا.
جواب ضعیف: «تیکتها را میبستم.» جواب قوی سه لایه دارد. اول مدل داده: انواع issue و سلسلهمراتب، workflow و status categoryها، و چرا برد تیم شما همان ستونها را داشت. دوم جستوجو: یکی دو JQL واقعی که خودت استفاده میکردی، مثلاً project = PAY AND statusCategory != Done AND updated <= -5d ORDER BY updated ASC برای پیدا کردن کارهای راکد. سوم اتصال به مهندسی: نامگذاری شاخه با کلید تیکت، ارجاع در commit و PR، و انتقال خودکار وضعیت هنگام merge.
اگر تجربهات با ابزار دیگری بوده، همین ساختار را بگو و اضافه کن که مفاهیم منتقل میشوند؛ چیزی که سنجیده میشود ابزار نیست، دیسیپلین ردیابی کار است.
۱۱. مهندسی نیازمندی
مهندسی نیازمندی یعنی فرایند کشف، تحلیل، مستندسازی، اعتبارسنجی و مدیریت آنچه سیستم باید انجام دهد. برای یک مهندس بکاند این «کار تحلیلگر» نیست؛ مهارتی است که تعیین میکند دو هفتهی بعدیات هدر میرود یا نه.
۱۱.۱ کشف نیاز (Elicitation)
نیازمندی «جمعآوری» نمیشود، کشف میشود — چون ذینفعان معمولاً راهحل میگویند نه مسئله. تکنیکهای اصلی: مصاحبه با سؤال باز («فرایند فعلیات را از اول تا آخر توضیح بده») و مقاومت در برابر پریدن به راهحل؛ مشاهدهی میدانی که تقریباً همیشه با آنچه گفته شده فرق دارد (فایل اکسل مخفی، راه دور زدن سیستم)؛ کارگاه و event storming برای کشف جمعی فرایند و رویدادهای دامنه (به فصل DDD وصل میشود)؛ تحلیل داده و لاگ سیستم موجود که اغلب حقیقت را بهتر از حرفها میگوید؛ پروتوتایپ بهعنوان ارزانترین راه تبدیل «نمیدانم چه میخواهم» به بازخورد مشخص؛ و ۵ چرا برای رسیدن از خواسته به نیاز.
اگر بپرسی «با آن فایل چه میکنی؟» ممکن است جواب باشد «هر صبح بازش میکنم و ردیفهای قرمز را به مدیرم ایمیل میکنم». نیاز واقعی اکسل نیست: یک هشدار روزانه است. راهحل درست شاید یک ایمیل زمانبندیشده باشد که یک روز کار دارد، نه یک ماژول گزارشسازی. قاعده: خواسته را بگیر، ولی تا کارِ پشتش را نفهمیدهای شروع نکن.
۱۱.۲ نیازمندی کارکردی و غیرکارکردی
FR میگوید سیستم چه کاری انجام دهد؛ NFR میگوید چقدر خوب انجامش دهد — و اینجاست که پروژهها میمیرند، چون NFRها معمولاً نانوشتهاند و بعد از انتشار کشف میشوند. استاندارد ISO/IEC 25010 (ویرایش ۲۰۲۳) نه ویژگی کیفیت تعریف میکند که فهرست وارسی خوبی برای NFR است:
| ویژگی کیفیت | سؤالی که باید بپرسی |
|---|---|
| Functional suitability | آیا کار درست را کامل و دقیق انجام میدهد؟ |
| Performance efficiency | زمان پاسخ، توان عملیاتی و مصرف منابع در چه حدی؟ |
| Compatibility | با چه سیستمها و نسخههایی باید کنار بیاید؟ |
| Interaction capability | کاربر چقدر راحت و در دسترس با آن کار میکند؟ |
| Reliability | در دسترس بودن، تحمل خطا، بازیابی |
| Security | محرمانگی، یکپارچگی، عدم انکار، پاسخگویی، مقاومت |
| Maintainability | ماژولاریتی، آزمونپذیری، قابلیت اصلاح |
| Flexibility | سازگاری، نصبپذیری، جایگزینی، مقیاسپذیری |
| Safety | جلوگیری از آسیب ناخواسته، fail-safe، هشدار خطر |
«باید سریع باشد» یک نیازمندی نیست، یک آرزوست. نیازمندی این است: «p95 زمان پاسخ endpoint جستوجو زیر ۳۰۰ms با ۵۰۰ درخواست بر ثانیه روی دادهی ۱۰ میلیون ردیفی». تفاوت این دو جمله، تفاوت یک ایندکس ساده با یک بازطراحی کامل ذخیرهسازی است. اگر NFR را قبل از طراحی نگیری، معماری را کور انتخاب کردهای.
۱۱.۳ نوشتن نیازمندی آزمونپذیر
استاندارد ISO/IEC/IEEE 29148 ویژگیهای نیازمندی خوب را فهرست میکند: necessary، appropriate، unambiguous، complete، singular، feasible، verifiable، correct، conforming و traceable. دو مورد بیشترین نقض را دارند: singular (کلمهی «و» معمولاً نشانهی دو نیازمندی است) و verifiable (اگر نتوانی تستی طراحی کنی که نقضش را نشان دهد، نیازمندی نیست).
قالب کمشناخته ولی فوقالعاده مفید EARS است (Easy Approach to Requirements Syntax) با پنج الگو:
Ubiquitous: The <system> shall <response>.
Event-driven: WHEN <trigger>, the <system> shall <response>.
State-driven: WHILE <state>, the <system> shall <response>.
Unwanted: IF <unwanted condition>, THEN the <system> shall <response>.
Optional: WHERE <feature is included>, the <system> shall <response>.
WHEN a payment authorisation request times out after 30 seconds,
the payment service shall retry the authorisation with the same idempotency key
at most twice, and shall mark the payment as UNKNOWN if all attempts time out.
IF the account balance is insufficient,
THEN the payment service shall reject the request with error code INSUFFICIENT_FUNDS
and shall not create a ledger entry.
قدرت EARS این است که ابهام را بدون نیاز به زبان صوری از بین میبرد و هر جمله تقریباً مستقیم به یک تست تبدیل میشود.
۱۱.۴ اولویتبندی
| روش | چطور کار میکند | بهترین کاربرد | ضعف |
|---|---|---|---|
| MoSCoW | Must / Should / Could / Won't برای یک بازهی مشخص | مذاکره با ذینفعان روی دامنهی یک انتشار | همه چیز Must میشود اگر سقف نگذاری |
| RICE | (Reach × Impact × Confidence) ÷ Effort | مقایسهی ایدههای محصولی با داده | اعداد ورودی خودشان حدساند |
| Kano | تفکیک ویژگیهای پایه، عملکردی و مسرورکننده | تصمیم دربارهی «چه چیزی تمایز میسازد» | نیاز به نظرسنجی کاربر |
| WSJF | ارزش تأخیرافتاده ÷ اندازهی کار | صفبندی در محیط چندتیمی (SAFe) | نیاز به تخمین چند مؤلفه |
| ارزش/ریسک | ابتدا آیتمهای پرارزش و پرریسک | کاهش زودهنگام ریسک فنی | مستقیماً ارزش کاربر را نمیسنجد |
جزئیات RICE: Reach تعداد افراد اثرپذیر در یک بازهی مشخص؛ Impact با مقیاس گسسته (۳ = بسیار زیاد، ۲ = زیاد، ۱ = متوسط، ۰٫۵ = کم، ۰٫۲۵ = خیلی کم)؛ Confidence درصد اطمینان (۱۰۰٪ دادهی محکم، ۸۰٪ شواهد خوب، ۵۰٪ حدس آگاهانه)؛ Effort برآورد نفر-ماه شامل همهی نقشها.
قاعدهی عملی رایج: مجموع تلاش آیتمهای Must نباید بیشتر از حدود ۶۰٪ ظرفیت بازه باشد. اگر همه چیز Must است، یعنی هیچ اولویتبندیای انجام نشده و اولین اتفاق غیرمنتظره کل برنامه را میشکند. سؤال جادویی برای شکستن جلسهی «همه چیز حیاتی است»: «اگر فقط یکی از اینها را میتوانستیم تحویل دهیم، کدام؟»
۱۱.۵ Traceability
Traceability یعنی بتوانی از هر نیازمندی به آیتم کاری، به کد و به تستِ تأییدکنندهاش برسی و برعکس. در محیطهای رگولهشده اجباری است و در بقیه هم وقتی میخواهی بفهمی «اگر این را حذف کنیم چه میشکند؟» نجاتدهنده است. نسخهی سبکش: کلید تیکت را در نام تست بگذار و یک جدول ساده برای پوشش نگه دار. پرسوجوی زیر نیازمندیهایی را پیدا میکند که هیچ تست قبولشدهای ندارند:
SELECT r.req_key,
r.title,
count(t.id) FILTER (WHERE t.last_result = 'PASS') AS passing_tests
FROM requirement r
LEFT JOIN requirement_test rt ON rt.req_id = r.id
LEFT JOIN test_case t ON t.id = rt.test_id
WHERE r.status = 'APPROVED'
GROUP BY r.req_key, r.title
HAVING count(t.id) FILTER (WHERE t.last_result = 'PASS') = 0
ORDER BY r.req_key;SELECT r.req_key,
r.title,
COUNT(CASE WHEN t.last_result = 'PASS' THEN 1 END) AS passing_tests
FROM requirement r
LEFT JOIN requirement_test rt ON rt.req_id = r.id
LEFT JOIN test_case t ON t.id = rt.test_id
WHERE r.status = 'APPROVED'
GROUP BY r.req_key, r.title
HAVING COUNT(CASE WHEN t.last_result = 'PASS' THEN 1 END) = 0
ORDER BY r.req_key;PostgreSQL از FILTER (WHERE ...) روی توابع تجمعی پشتیبانی میکند که خواناتر است؛ Oracle آن را ندارد و باید از COUNT(CASE WHEN ... THEN 1 END) استفاده کنی. هر دو در HAVING مجازند. جزئیات بیشتر در فصل «تفاوتهای Oracle و PostgreSQL» آمده است.
۱۱.۶ مدیریت تغییر، scope creep و امکانسنجی
Scope creep یعنی رشد بیسروصدای دامنه بدون تغییر در زمان یا ظرفیت — معمولاً از راه جملهی «فقط یک چیز کوچک هم اضافه کن». تفاوتش با تغییر مشروع این است که تغییر مشروع آگاهانه است و چیزی جایش کنار گذاشته میشود. سه ابزار عملی: ثبت بهجای رد کردن («این را بهعنوان آیتم ثبت میکنم و با PO اولویتش را مشخص میکنیم»)؛ قاعدهی مبادله (هر ورودی جدید به sprint جاری باید چیزی هماندازه را بیرون کند، و تصمیمش با PO است نه با درخواستکننده)؛ و مرز صریح در تیکت (بخش «خارج از دامنه» بیشتر از هر جلسهای جلوی خزش را میگیرد).
قبل از تعهد به یک قابلیت، چهار بعد امکانسنجی را بررسی کن: فنی (با معماری و مهارت فعلی شدنی است؟)، عملیاتی (میشود پشتیبانی و مانیتورش کرد؟)، اقتصادی (ارزشش به هزینهاش میارزد؟) و قانونی (دادهی شخصی، نگهداری، ممیزی). ابزار مهندسی این کار spike است: بررسی زمانمحدود با سؤال مشخص و خروجی نوشتاری.
«حالا که اینجاییم بیا این را هم refactor کنیم»، «بیا این را generic بنویسیم برای بعداً». این نوع خزش نه ثبت میشود و نه دیده میشود، ولی تخمینها را میشکند و PR را غیرقابل review میکند. راه درست: refactor لازم برای همین کار را انجام بده، بقیه را تیکت کن. بحث تفصیلی این مرز در فصل کد تمیز و بدهی فنی آمده است.
نشان بده که میدانی این جمله قابل پیادهسازی نیست و میدانی چطور قابل پیادهسازیاش کنی. سؤالهایی که میپرسم: کدام عملیات؟ برای چه کسی و در چه سناریویی؟ میانگین یا صدک ۹۵؟ زیر چه باری و با چه حجم دادهای؟ «بهاندازهی کافی سریع» یعنی چه و مبنایش چیست (کاربر منتظر میماند یا فرایند پسزمینه است)؟ اگر از این عدد فراتر رفتیم پیامدش چیست؟
بعد آن را به نیازمندی آزمونپذیر تبدیل میکنم: «صدک ۹۵ زمان پاسخ GET /orders زیر ۴۰۰ms در بار ۳۰۰ درخواست بر ثانیه با ۵ میلیون سفارش» — و بلافاصله میگویم چطور اثباتش میکنم: تست بار در محیط شبیه production و یک هشدار روی همان متریک. جملهی پایانی: «هر NFR باید یک عدد، یک شرط بار و یک روش اندازهگیری داشته باشد، وگرنه فقط یک آرزوست.»
۱۲. UML برای مهندس بکاند: فقط چیزی که ارزشش را دارد
UML بیش از ده نوع نمودار دارد و در عمل یک مهندس بکاند از پنجتای آنها استفاده میکند. هدف نه «مستندسازی کامل» است و نه «تولید کد از مدل»؛ هدف رساندن یک ایده به آدم دیگر با کمترین ابهام است.
| نمودار | چه چیزی را نشان میدهد | کِی واقعاً ارزش دارد | معادل در Mermaid |
|---|---|---|---|
| Class | ساختار ایستا: کلاسها و روابط | مدلسازی دامنه | classDiagram |
| Sequence | تعامل زمانی بین اجزا | مسیر یک درخواست بین سرویسها و نقاط خطا | sequenceDiagram |
| State machine | چرخهی حیات یک موجودیت | سفارش، پرداخت، اشتراک | stateDiagram-v2 |
| Activity | جریان کار با شرط و موازیسازی | فرایند کسبوکار یا الگوریتم پیچیده | flowchart |
| Component | ماژولها و رابطهای بینشان | نقشهی سیستم برای عضو جدید تیم | flowchart + subgraph |
| Deployment | نگاشت به محیط اجرا | توپولوژی استقرار و مرزهای امنیتی | flowchart + subgraph |
| Use case | بازیگران و اهداف سیستم | دامنهبندی اولیه با ذینفعان غیرفنی | flowchart |
۱۲.۱ Class diagram
نمادهای روابط که باید بشناسی: <|-- ارثبری، *-- ترکیب (جزء بدون کل نمیماند)، o-- تجمیع (جزء مستقل است)، --> انجمن، ..> وابستگی، ..|> پیادهسازی رابط. نشانههای دسترسی: + عمومی، - خصوصی، # محافظتشده.
مدل دامنهی یک سفارش — an order domain model.
classDiagram
class Order {
+UUID id
+OrderStatus status
+Money total()
+void addLine(Product p, int qty)
}
class OrderLine {
+int quantity
+Money unitPrice
}
class Customer {
+UUID id
+String email
}
class Payment {
+UUID id
+Money amount
+PaymentStatus status
}
class PaymentMethod {
<<interface>>
+authorize(Money amount) AuthResult
}
class SavedCard
Customer "1" --> "0..*" Order : places
Order "1" *-- "1..*" OrderLine : contains
Order "1" --> "0..*" Payment : paid by
Payment ..> PaymentMethod : uses
SavedCard ..|> PaymentMethod
۱۲.۲ Sequence diagram — پرکاربردترین نمودار برای بکاند
مسیر یک پرداخت شامل timeout و تلاش مجدد — a payment flow with a gateway timeout.
sequenceDiagram
autonumber
participant C as Client
participant API as Order API
participant PS as Payment Service
participant GW as Card Gateway
participant DB as Ledger DB
C->>API: POST /orders/{id}/pay (Idempotency-Key)
API->>PS: authorize(orderId, cardToken, key)
PS->>DB: insert payment attempt (PENDING, key)
PS->>GW: authorize request
GW--xPS: timeout after 30s
PS->>GW: retry with same key
GW-->>PS: approved (auth ref)
PS->>DB: update payment to AUTHORIZED
PS-->>API: AuthResult(approved)
API-->>C: 200 OK (order PAID)
این نمودار در یک طراحی یا review بیشتر از دو صفحه متن ارزش دارد، چون ترتیب، مرزهای شبکه و نقاط شکست را همزمان نشان میدهد. برای مسیرهای خطا حتماً یک نسخهی جداگانه بکش؛ مسیر خوشبینانه معمولاً بدیهی است.
۱۲.۳ State machine
هر موجودیتی که فیلد status دارد در واقع یک ماشین حالت است. کشیدنش سه چیز را آشکار میکند: گذارهای غیرمجاز، وضعیتهای بنبست، و وضعیتهای تکراری که واقعاً یکی هستند.
چرخهی حیات یک پرداخت — the payment lifecycle.
stateDiagram-v2
[*] --> Pending
Pending --> Authorized: gateway approves
Pending --> Declined: gateway declines
Pending --> Unknown: timeout, result not known
Unknown --> Authorized: reconciliation finds a match
Unknown --> Declined: reconciliation finds no charge
Authorized --> Captured: funds captured
Captured --> Refunded: refund requested
Declined --> [*]
Refunded --> [*]
Captured --> [*]
بزرگترین اشتباه در طراحی ماشین حالت سیستمهای توزیعشده، فرض کردن دنیای دودویی موفق/ناموفق است. وقتی درخواست timeout میشود، تو نمیدانی طرف مقابل چه کرد. اگر وضعیت UNKNOWN نداشته باشی، کد مجبور است یکی از دو دروغ را بگوید و نتیجهاش یا کسر دوبارهی پول است یا کالای تحویلشده بدون پرداخت. کشیدن ماشین حالت، دقیقاً همین حفره را قبل از کدزدن نشان میدهد.
۱۲.۴ Activity، Component، Deployment و Use case
Activity معادل عملیاش در Mermaid یک flowchart با گرههای تصمیم است؛ برای فرایندهای شرطی مفید است.
فرایند تأیید یک سفارش — order approval activity.
flowchart TD
S([Order submitted]) --> V{Stock available?}
V -- no --> B[Backorder and notify customer]
V -- yes --> R{Risk score acceptable?}
R -- no --> M[Send to manual review]
R -- yes --> P[Authorize payment]
P --> Q{Authorized?}
Q -- no --> F[Mark order as payment failed]
Q -- yes --> A[Reserve stock and confirm order]
A --> E([Order confirmed])
Component نشان میدهد سیستم از چه قطعاتی ساخته شده و از چه رابطهایی حرف میزنند؛ بهترین کاربردش صفحهی اول README یک سرویس است.
اجزای یک سرویس سفارش — components of an order service.
flowchart LR
subgraph OrderService[Order Service]
API[REST API layer]
APP[Application services]
DOM[Domain model]
OUT[Outbox publisher]
REP[JPA repositories]
end
API --> APP --> DOM
APP --> REP
APP --> OUT
REP --> DB[(PostgreSQL)]
OUT --> MQ[[Kafka topic: order-events]]
API --> PAY[Payment Service REST]
Deployment نگاشت نرمافزار به زیرساخت است: چه چیزی کجا اجرا میشود و مرزهای شبکه کجاست.
توپولوژی استقرار — a deployment topology.
flowchart TB
subgraph Edge[Public zone]
LB[Load balancer / TLS termination]
end
subgraph K8s[Kubernetes cluster - private zone]
ING[Ingress controller]
P1[order-service pods x3]
P2[payment-service pods x2]
end
subgraph Data[Data zone]
PG[(PostgreSQL primary + replica)]
RD[(Redis cache)]
KF[[Kafka cluster]]
end
LB --> ING --> P1
ING --> P2
P1 --> PG
P1 --> RD
P1 --> KF
P2 --> PG
Use case سادهترین و سیاسیترین نمودار است: با ذینفع غیرفنی مینشینی و بازیگران و اهدافشان را میکشی تا مرز سیستم روشن شود.
بازیگران و اهداف — actors and goals.
flowchart LR
CU([Customer]) --> UC1[Place an order]
CU --> UC2[Pay for an order]
CU --> UC3[Track order status]
OP([Operations staff]) --> UC4[Issue a refund]
OP --> UC5[Resolve a stuck payment]
UC2 --> GW([Card gateway])
UC4 --> GW
اگر سازمانی برای هر story نمودار «تحویل» بگیرد، دو اتفاق میافتد: نمودارها بلافاصله بعد از merge منسوخ میشوند، و آدمها یاد میگیرند نمودار را بعد از نوشتن کد بسازند تا فرم پر شود. نمودار وقتی ارزش دارد که قبل از تصمیم کشیده شود تا در تصمیم اثر بگذارد، یا برای توضیح چیزی پیچیده و پایدار نگهداری شود. بقیهاش را روی وایتبرد بکش و عکس بگیر.
جواب باید عملی و مقتصد باشد: «برای هر تغییری که بیش از یک سرویس را درگیر میکند یک sequence diagram میکشم و در توضیح PR میگذارم، چون ترتیب و نقاط timeout را روشن میکند. برای هر موجودیتی که فیلد status دارد یک state machine میکشم و بهصورت Mermaid در همان مخزن نگه میدارم تا کد و نمودار با هم review شوند. برای معماری کلی بهجای UML از C4 استفاده میکنم — سطح context و container معمولاً کافی است. class diagram را فقط برای مدل دامنه میکشم.»
بعد بگو چه چیزی را عمداً نمیکشی: نمودارهایی که IDE میتواند تولید کند، و هر نموداری که کسی متعهد بهروزرسانیاش نیست. جملهی پایانی: «نمودار بیارزش، نمودار غلط است؛ و نمودار غلط از نبودن نمودار بدتر است.»
۱۳. مدل C4: جایگزین مدرن برای نمودار معماری
مشکل نمودارهای معماری معمول این است که همه چیز را در یک تصویر میریزند و هیچکس نمیداند یک مستطیل یعنی سرور، فرایند، ماژول یا تیم. C4 این را با چهار سطح بزرگنمایی حل میکند، مثل نقشهای که از کشور تا خیابان zoom میکند: System Context (سیستم ما بهعنوان یک جعبه، بهعلاوهی کاربران و سیستمهای بیرونی — مخاطبش همه، حتی غیرفنیها)؛ Container (واحدهای قابل اجرا/استقرار داخل سیستم: یک اپلیکیشن سمت سرور، یک SPA، یک پایگاهداده، یک message broker — توجه کن «container» اینجا مفهوم معماری است و لزوماً یعنی Docker نیست)؛ Component (بلوکهای داخل یک container و مسئولیتهایشان)؛ و Code (جزئیات کلاسها که معمولاً به IDE سپرده میشود). بهعلاوه نمودارهای مکمل: System Landscape، Dynamic و Deployment.
سطح یک C4 — C4 level 1, system context.
flowchart TB
U([Customer]) -->|places and pays for orders| S[Order Platform]
OPS([Operations staff]) -->|handles refunds| S
S -->|authorises payments| GW[Card Gateway - external]
S -->|sends receipts| MAIL[Email provider - external]
S -->|reports| BI[Data warehouse - internal]
سطح دو C4 — C4 level 2, containers.
flowchart TB
U([Customer]) --> SPA[Web SPA - TypeScript]
SPA -->|JSON over HTTPS| API[Order API - Spring Boot]
API -->|JDBC| DB[(Order DB - PostgreSQL)]
API -->|publishes events| K[[Kafka]]
API -->|REST| PAY[Payment Service - Spring Boot]
PAY -->|JDBC| PDB[(Payment DB - PostgreSQL)]
K --> PROJ[Reporting projector - Spring Boot]
PROJ --> WH[(Reporting store)]
C4 را میشود بهصورت متن نگه داشت (diagrams-as-code) — مثلاً با Structurizr DSL:
workspace "Order Platform" {
model {
customer = person "Customer"
platform = softwareSystem "Order Platform" {
spa = container "Web SPA" "TypeScript"
api = container "Order API" "Spring Boot"
db = container "Order DB" "PostgreSQL"
}
gateway = softwareSystem "Card Gateway" "External"
customer -> spa "Uses"
spa -> api "Calls" "JSON/HTTPS"
api -> db "Reads from and writes to" "JDBC"
api -> gateway "Authorises payments"
}
views {
systemContext platform "Context" {
include *
autoLayout lr
}
container platform "Containers" {
include *
autoLayout lr
}
theme default
}
}
۱۴. BPMN در یک نگاه
BPMN نمادگذاری استاندارد برای مدلسازی فرایندهای کسبوکار است. تفاوت اصلیاش با activity diagram این است که علاوه بر خواندهشدن توسط انسان، قابل اجرا است: موتورهای فرایند میتوانند مستقیماً همان مدل را اجرا کنند. عناصری که باید بشناسی: pool و lane (بازیگران و مسئولیتها)، task (کار انسانی یا سرویسی)، gateway (شرط انحصاری، موازی، رویدادی)، event (شروع، پایان، تایمر، پیام، خطا) و sequence flow. وقتی فرایندی چند روز طول میکشد، تأیید انسانی دارد یا باید جبران (compensation) داشته باشد، مدلسازی BPMN و اجرای آن با یک موتور فرایند منطقیتر از پیادهسازی دستی با جدول وضعیت است. فصل «موتور فرایند و گردش کار» این موضوع را کامل با مثال اجرایی پوشش داده؛ اینجا کافی است بدانی چه زمانی سراغش بروی.
۱۵. مستنداتی که دوام میآورد
مستند بد، مستندی است که کسی بهروزش نمیکند. راهحل «مستند بیشتر» نیست؛ مستند کمتر، نزدیکتر به کد، و با عمر مشخص است.
۱۵.۱ ADR — ثبت تصمیم معماری
ADR یک فایل کوتاه است که یک تصمیم مهم و چراییِ آن را ثبت میکند. ارزشش این نیست که تصمیم را توضیح دهد؛ این است که گزینههای رد شده و دلایلشان را حفظ کند تا یک سال بعد کسی همان بحث را از صفر تکرار نکند. قالب متداول MADR (نسخهی ۴) با نامگذاری NNNN-title-with-dashes.md در پوشهی docs/decisions/:
---
status: accepted
date: 2026-03-11
decision-makers: payments team
consulted: platform team, security
informed: product
---
# 0007 — Use an outbox table for publishing payment events
## Context and Problem Statement
سرویس پرداخت باید پس از هر تغییر وضعیت رویداد منتشر کند. نوشتن همزمان در پایگاهداده و
message broker اتمیک نیست؛ در صورت خرابی بین دو عمل، یا رویداد گم میشود یا رویدادی
منتشر میشود که تراکنشش commit نشده.
## Decision Drivers
* از دست نرفتن رویداد در صورت خرابی سرویس؛ بدون تراکنش توزیعشده؛ تأخیر زیر ۵ ثانیه قابل قبول
## Considered Options
* نوشتن مستقیم در broker پس از commit
* تراکنش توزیعشده (XA) بین پایگاهداده و broker
* جدول outbox + publisher جداگانه
* change data capture روی جدول پرداخت
## Decision Outcome
گزینهی «جدول outbox + publisher» انتخاب شد، چون اتمیک بودن را با همان تراکنش
پایگاهداده تضمین میکند و به زیرساخت جدیدی نیاز ندارد.
### Consequences
* خوب: هیچ رویدادی گم نمیشود؛ انتشار at-least-once با کلید idempotency در مصرفکننده
* بد: یک جدول و یک فرایند پسزمینهی دیگر برای نگهداری، و تأخیر وابسته به بازهی polling
### Confirmation
تست integration که خرابی سرویس بین commit و انتشار را شبیهسازی میکند، بهعلاوهی
هشدار روی طول صف outbox.
وضعیتهای متداول: proposed، accepted، rejected، deprecated و superseded by ADR-NNNN. قاعدهی طلایی: ADR را ویرایش نکن، جایگزینش کن — تاریخچهی تصمیمها همان چیزی است که ارزش دارد.
اگر تصمیم برگرداندنش گران است یا بیش از یک تیم را تحت تأثیر قرار میدهد، ADR بنویس. انتخاب کتابخانهی لاگ معمولاً نه؛ مدل consistency، مرزبندی سرویسها، فرمت پیامها، استراتژی احراز هویت یا انتخاب پایگاهداده حتماً بله. یک ADR خوب در ۱۵ دقیقه نوشته میشود و سالها جواب میدهد؛ اگر نوشتنش یک روز طول بکشد، داری گزارش مینویسی نه ADR.
۱۵.۲ README، runbook و diagrams-as-code
README برای کسی است که میخواهد کد را اجرا و تغییر دهد: این سرویس چه میکند و مالکش کیست، چطور لوکال بالا میآید (ترجیحاً یک دستور)، وابستگیهای بیرونی، چطور تست اجرا میشود، یک نمودار معماری، و لینک به ADRها. Runbook برای کسی است که ساعت سه بامداد pager خورده: عملیاتی، مرحلهبهمرحله، بدون تئوری.
# Runbook — payment-service
## Health and dashboards
سلامت: GET /actuator/health · داشبورد: <link> · هشدارها: <link>
## Alert: payment_outbox_lag_high
معنی: صف outbox بیش از ۵ دقیقه عقب است؛ رویدادها منتشر نمیشوند.
1. بررسی کن publisher در حال اجراست: kubectl get pods -l app=payment-outbox
2. لاگ خطای اتصال به broker را ببین: kubectl logs deploy/payment-outbox --since=15m
3. اگر broker در دسترس نیست، تیم پلتفرم را خبر کن و صبر کن — داده گم نمیشود.
4. اگر publisher در crash loop است، به نسخهی قبل برگرد: <دستور rollback>
5. اگر صف بیش از ۱۰۰هزار رکورد دارد، مقیاس افقی publisher را افزایش بده.
تشدید: اگر ظرف ۳۰ دقیقه رفع نشد، on-call تیم پرداخت را صدا بزن.
مستند API باید از خود کد یا از قرارداد تولید شود، نه دستی (جزئیات OpenAPI در فصل «طراحی API» آمده است)؛ نکتهی فرایندی این است که تولید و انتشارش را به pipeline بسپار تا هیچوقت عقب نماند:
name: docs
on:
push:
branches: [main]
jobs:
publish:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: actions/setup-java@v4
with:
distribution: temurin
java-version: '21'
- name: Generate OpenAPI document
run: ./mvnw -B verify -Dtest=OpenApiExportTest
- name: Upload API docs
uses: actions/upload-artifact@v4
with:
name: openapi
path: target/openapi.json
بهترین جواب یک ADR واقعی است: مسئله چه بود، چه گزینههایی بررسی شد، چه معیارهایی داشتی (نه فقط «بهتر بود»)، چه چیزی را انتخاب کردی و چه چیزی را از دست دادی. اشاره به پیامد منفی معمولاً نقطهی برندهی جواب است: نشان میدهد تصمیم را واقعی میبینی نه تبلیغاتی.
اگر تیمت ADR نداشت، صادق باش و بگو کجا نوشتی (توضیح PR، صفحهی ویکی) و چه مشکلی از نبودش دیدی — مثلاً «شش ماه بعد همان بحث دوباره شروع شد چون هیچکس دلیل اولیه را نمیدانست» — و بعد بگو الان چطور انجامش میدهی.
۱۶. فرهنگ code review
Code review دو هدف دارد که معمولاً یکیشان فراموش میشود: بالا بردن کیفیت کد، و انتشار دانش در تیم. اگر فقط اولی را دنبال کنی، به یک دروازهی کنترل کیفیت آزاردهنده میرسی.
| اولویت | چه چیزی را بررسی میکنی | نمونه سؤال |
|---|---|---|
| ۱ | درستی و مسئله | آیا واقعاً مسئلهی تیکت را حل میکند؟ حالتهای مرزی؟ |
| ۲ | امنیت و داده | ورودی اعتبارسنجی شده؟ دادهی حساس در لاگ؟ کنترل دسترسی؟ |
| ۳ | همزمانی و قابلیت اطمینان | idempotency، retry، timeout، رفتار در خرابی جزئی |
| ۴ | طراحی و مرزها | آیا در لایهی درست است؟ انتزاع نشتی ندارد؟ |
| ۵ | تست | آیا تستها رفتار را میسنجند یا فقط پوشش میسازند؟ |
| ۶ | عملیات | لاگ، متریک، migration، سازگاری با نسخهی قبل |
| ۷ | خوانایی | نامها، اندازهی توابع، کامنتهایی که «چرا» را میگویند |
| ۸ | سبک | باید خودکار باشد، نه بحث انسانی |
هرچه پایینتر میروی، ارزش بازخورد انسانی کمتر میشود؛ سبک و فرمت را به formatter و linter بسپار (در فصل کد تمیز) تا انسانها وقتشان را روی طراحی و ریسک بگذارند.
برای بازخوردی که کارساز باشد: سطح را مشخص کن — یک قرارداد ساده و رایج، برچسبگذاری کامنتهاست: blocking: (بدون رفعش merge نمیکنم)، suggestion: (بهتر است ولی اختیاری)، nit: (سلیقه)، question: (توضیح بده)، praise: (این کار خوب بود). دربارهی کد حرف بزن نه آدم («این متد سه مسئولیت دارد» نه «تو همیشه متدهای بزرگ مینویسی»). چرا را بگو («این کوئری در حلقه اجرا میشود و به N+1 منجر میشود؛ با یک fetch join حل میشود» بهجای «بهینه نیست»). و اگر بحثی به دور سوم رسید، یک تماس پنجدقیقهای ده کامنت را صرفهجویی میکند. گیرندهی بازخورد هم مسئولیت دارد: هر کامنت را یا اعمال کن یا جواب بده.
کیفیت review با اندازهی تغییر بهشدت افت میکند؛ بالای چند صد خط، بازبین دیگر واقعاً نمیخواند. اگر PRت بزرگ است مشکل بازبین نیست، مشکل برش کار است. راهحلها: برش عمودی کوچکتر، جدا کردن refactor خالص از تغییر رفتار در دو PR (این یکی بهتنهایی نیمی از مشکل را حل میکند)، و feature flag برای merge زودهنگام کد ناتماماماایمن.
اول مطمئن شو مخالفتت مبتنی بر معیار است نه سلیقه — درستی، امنیت، کارایی سنجیدهشده، یا هزینهی نگهداری. بعد اختلاف را طبقهبندی کن: اگر سلیقه است، صریح بگو nit و رد شو. اگر ریسک واقعی است، آن را با یک سناریوی مشخص بیان کن («اگر دو درخواست همزمان بیایند این کد موجودی را دوبار کم میکند؛ این تست نشانش میدهد»). نشاندادن با تست یا داده، بحث را تمام میکند.
اگر باز هم توافق نشد: از کامنت به گفتوگو برو، و اگر لازم شد نظر سوم بگیر یا تصمیم را در ADR ثبت کن. و مهمترین بخش جواب: گاهی مخالفت را کنار میگذاری («disagree and commit») چون هزینهی توقف کار از هزینهی آن تفاوت بیشتر است — و اگر بعداً درست از آب درآمد، بدون سرزنش اصلاحش میکنی. ترکیب استدلال فنی و پختگی اجتماعی دقیقاً چیزی است که مصاحبهگر دنبالش است.
۱۷. چطور تجربهی فرایندیات را در مصاحبه بگویی
بیشتر کاندیداها اینجا میبازند، نه چون تجربه ندارند، بلکه چون تجربهشان را با کلمات کلیشهای تعریف میکنند. سه قاعده: از کلیگویی به روایت مشخص برو («Agile کار میکردیم» هیچ اطلاعاتی ندارد؛ «sprint دوهفتهای، DoD شامل تست integration و بهروزرسانی OpenAPI، و از یک retro تصمیم گرفتیم PRها را زیر ۴۰۰ خط نگه داریم» یک تصویر واقعی میسازد)؛ عدد بگو (cycle time، تعداد deploy در هفته، اندازهی تیم، درصد کار برنامهریزینشده)؛ و از ساختار STAR استفاده کن (موقعیت، وظیفه، اقدام، نتیجه) با یک جملهی «چه چیزی را متفاوت انجام میدادم».
گفتن «فرایند ما ضعیف بود» ضعف نیست؛ نتوانستن تشخیص اینکه چرا ضعیف بود، ضعف است. جواب قوی این شکلی است: «ما اسماً Scrum بودیم ولی هر sprint نصف ظرفیت به کار فوری میرفت و Sprint Goal معنی نداشت. شروع کردم به ثبت هر کار برنامهریزینشده روی برد؛ بعد از یک ماه داده نشان داد ۴۰٪ ظرفیت صرف پشتیبانی میشود. با همان داده توافق کردیم یک نفر بهصورت چرخشی نگهبان پشتیبانی باشد و بقیه محافظت شوند؛ cycle time از ۱۲ روز به ۶ روز رسید.» این جواب میگوید تو مشاهده میکنی، داده جمع میکنی و بدون اقتدار رسمی تغییر ایجاد میکنی.
یک روایت خطی بده، نه فهرست اصطلاحات: «یک نیاز از سمت محصول با یک مسئله و یک عدد شروع میشد. در refinement آن را میشکستیم و معیار پذیرش مینوشتیم؛ اگر ناشناختهی فنی داشت یک spike دوروزه تعریف میکردیم. در planning یک Sprint Goal انتخاب میکردیم. توسعه روی شاخهی کوتاهعمر با کلید تیکت، PR کوچک با حداقل یک approve، pipeline شامل تست واحد و integration با Testcontainers، اسکن وابستگی و ساخت image. merge به main یعنی deploy خودکار به staging؛ انتشار به production روزانه و پشت feature flag. بعد از انتشار داشبورد و هشدارها را نگاه میکردیم و اگر مشکلی بود flag را میبستیم. هر دو هفته retro با حداکثر دو اقدام که خودشان وارد sprint بعد میشدند.»
این جواب همزمان چند چیز را ثابت میکند: مالکیت end-to-end، آشنایی با CI/CD، توجه به قابلیت مشاهده، و اینکه فرایند برایت ابزار است نه تشریفات.
دو تله: غر زدن («مدیریت بد بود») و بیتفاوتی («همه چیز خوب بود»). جواب خوب یک مشکل واقعی را نام میبرد، ریشهاش را تحلیل میکند و راهحلی با معیار موفقیت ارائه میدهد.
نمونه: «کارها با سرعت شروع میشدند ولی جریان کند بود، چون هرکس دو سه کار باز داشت و PRها روزها منتظر میماندند. تغییر پیشنهادی من: محدودیت WIP روی ستونهای In Progress و In Review، بهعلاوهی قاعدهی «قبل از شروع کار جدید، review معلق را انجام بده». معیار موفقیت: کاهش صدک ۸۵ cycle time و کاهش سن متوسط آیتمهای باز.» توجه کن که راهحل ارزان، قابل اندازهگیری و بدون سرزنش کسی است.
مدلهای چرخهی عمر با «زمانبندی بازخورد» از هم متمایز میشوند: آبشار ریسک را به انتها میبرد، Agile آن را در چرخههای کوتاه پخش میکند — و هیچکدام همیشه درست نیستند. Scrum یک چارچوب تجربهگراست با سه پاسخگویی (PO، SM، Developers)، سه مصنوع با تعهدهایشان (Product Goal، Sprint Goal، Definition of Done) و پنج رویداد که هرکدام یک ستون تجربهگرایی را عملی میکنند؛ هر رویدادی که تصمیمی تولید نکند تشریفات است. Backlog با user story و INVEST و معیار پذیرش زنده میماند و مهمترین مهارتش برش عمودی است. تخمین با story point برای برنامهریزی است نه تعهد؛ velocity معیار بهرهوری نیست و پیشبینی احتمالاتی با throughput صادقانهتر است. معیارهای واقعی، چهار معیار DORA بهعلاوهی معیارهای جریان (WIP، cycle time، work item age) هستند و همیشه جفتی — یکی سرعت، یکی پایداری. Kanban با محدودکردن WIP و قانون لیتل جریان را کوتاه میکند و در کارهای غیرقابل پیشبینی بر Scrum برتری دارد. در Jira آنچه واقعاً سنجیده میشود تیکت خوب، workflow ساده، JQL کاربردی و اتصال commit/PR به تیکت است. در مهندسی نیازمندی، نیاز را کشف کن نه جمعآوری؛ NFR را با عدد و شرط بار بنویس (ISO 25010 فهرست وارسی خوبی است)، از EARS برای بیابهامنویسی استفاده کن، با MoSCoW و RICE اولویت بده و scope creep را با ثبت و مبادله مهار کن. از UML فقط پنج نمودار پرکاربرد را نگه دار، برای معماری از C4 استفاده کن و همه را بهصورت diagrams-as-code کنار کد بگذار. تصمیمهای گرانبرگشت را در ADR ثبت کن و runbook بنویس برای ساعت سه بامداد. در code review اول درستی و ریسک، آخر سبک؛ PR کوچک، بازخورد برچسبدار و بدون تأخیر. و در مصاحبه، بهجای نام بردن از متدولوژیها، یک روایت مشخص با عدد و یک تأمل صادقانه ارائه کن.
Almost every backend job ad contains a clause engineers skim past: "familiar with Agile and Scrum", "proficient with Jira", "able to analyse and document requirements", "knowledge of UML". Then, in the interview, that clause turns into questions that are harder to answer than a tricky concurrency problem — because nobody ever taught you this material. You just sat in meetings for years while nobody explained why they existed.
The gap between a mid-level and a senior engineer shows up more in this layer than in the code layer. Both can write an endpoint. But when a vague request arrives from the business, one asks "what should I code?" and the other asks "what is the real problem, how will we know it is solved, which part can ship this week, and what are we giving up?" None of this is "the manager's job". Mature organisations expect the engineer to challenge a requirement, slice a story, write acceptance criteria, record the decision, and improve both the code and the person in a review.
First we compare software lifecycles and see what pain Agile grew out of, then read the Agile Manifesto critically. Next, Scrum in full: accountabilities, artifacts, events and exactly how each one goes wrong, plus Definition of Done and Ready. Then the backlog (user stories, INVEST, acceptance criteria, vertical slicing), estimation (story points, planning poker, velocity and its misuse), and the metrics that actually mean something (DORA and flow metrics). Then Kanban and flow, and a short look at scaling. Then Jira in practice: issue types, workflows, boards, JQL and linking commits and PRs to issues. Then requirements engineering: elicitation, functional versus non-functional requirements, writing testable requirements with EARS, traceability, prioritisation and containing scope creep. Then UML — only the diagrams that pay off for a backend engineer, each with a Mermaid example — plus the C4 model and a brief look at BPMN. Then documentation that survives (ADRs, README, runbooks, diagrams-as-code), code review culture, and finally: how to describe this experience convincingly in an interview.
1. Why "process" is an engineering skill
You do not need a process to build a shed in your garden. But if twenty people work on one building, the moment the plumber does not know where the electrician ran the conduit, a wall gets drilled through. Process is just the minimum set of agreements: who decides, where it is written down, and how we know something is finished. Software is worse, because a building is visible and software is not — those process artifacts are the only way to see it.
Three engineering — not managerial — reasons process matters. The cost of a requirements mistake grows exponentially: a misunderstanding that one question in analysis would have resolved turns into a data migration and lost trust once it reaches production. Software work is inherently uncertain: you are building something you have not built before, so you need short feedback and course correction. And organisational memory does not live in people: people leave; tickets, ADRs, runbooks and tests remain.
2. The software development lifecycle: from waterfall to Agile
SDLC is the order in which software moves from idea to retirement: requirements, design, implementation, testing, delivery, maintenance. The models differ only in how those phases are arranged and how often they repeat.
2.1 Waterfall
Do each phase once, in order, and do not enter the next phase until the previous one is signed off. A nice historical detail: the famous paper that first drew this diagram described the single-pass model as risky and recommended running it at least twice — the diagram survived, the warning did not.
The one-way waterfall flow — مسیر یکطرفهی آبشار.
flowchart LR
A[Requirements] --> B[Design]
B --> C[Implementation]
C --> D[Verification]
D --> E[Deployment]
E --> F[Maintenance]
The fundamental problem is feedback timing: the first moment a real user sees the software is after the entire budget has been spent.
2.2 The V-model and iterative-incremental development
The V-model is waterfall with one valuable addition: every design level has a matching test level written at the same time (user requirements → acceptance tests, architectural design → integration tests, module design → unit tests). Even in an Agile team the idea that every level of decision has a corresponding level of testing is extremely useful; the testing chapter covers it in depth.
Two words that constantly get mixed up: incremental means you deliver the product in pieces (first login, then the cart, then payment); iterative means you improve the same piece repeatedly (first a simple payment, then with retries, then with reconciliation). Agile uses both at once.
The iterative feedback loop — حلقهی بازخورد تکراری.
flowchart LR
P[Plan a thin slice] --> B[Build]
B --> R[Release to real users]
R --> M[Measure and learn]
M --> P
2.3 An honest comparison
| Criterion | Waterfall / V | Iterative-incremental | Agile (Scrum/Kanban) |
|---|---|---|---|
| Assumption about requirements | known and stable up front | largely known | vague and changing |
| First real feedback | end of the project | end of each large iteration | every one to four weeks |
| Handling change | formal change-request process | absorbable at a cost | change is a normal input |
| Strength | contracts, audits, safety-critical systems | large projects with a clear scope | products and volatile markets |
| Weakness | risk concentrated at the end | still a long cycle | degenerates without engineering discipline |
| "Done" means | a signed-off phase | an iteration delivery | a releasable increment (DoD) |
Agile wins when feedback is cheap and change is likely. When feedback is expensive — hardware, medical-device firmware, a fixed-price contract with penalties, a regulatory audit requirement — planning further ahead earns its keep. In an interview, the person who can draw that line is far ahead of the person who only says "Agile is good".
3. The Agile Manifesto, read critically
The manifesto (2001) has four comparative statements, and most people skip the closing line: "That is, while there is value in the items on the right, we value the items on the left more." The right-hand side is not worthless.
| We value this more | than this | practical translation |
|---|---|---|
| Individuals and interactions | processes and tools | Jira does not replace a conversation; but a conversation with no written trace is lost |
| Working software | comprehensive documentation | less documentation that is alive, not "no documentation" |
| Customer collaboration | contract negotiation | show it earlier so there is less to argue about |
| Responding to change | following a plan | have a plan, but do not be loyal to it |
The real content sits in the twelve principles behind the manifesto. Four that engineers most often ignore: "Deliver working software frequently... with a preference to the shorter timescale"; "Continuous attention to technical excellence and good design enhances agility" (technical quality is part of the definition of agile, not a sacrifice to speed); "Simplicity — the art of maximizing the amount of work not done — is essential"; and "should be able to maintain a constant pace indefinitely" (chronic overtime is anti-Agile).
Clinical symptoms: the daily exists but has become a status report to a manager; retrospectives happen but no action from them is ever done; story points have become hours and are used to measure the team's "commitment"; the "releasable increment" is not actually released for three months. A simple diagnostic: if removing a meeting would change no decision, that meeting is not part of your process — it is ceremony.
The core difference is feedback timing and the treatment of change, not the presence or absence of documentation. Waterfall assumes requirements are known up front and pushes risk to the end of the project; Agile assumes requirements are discovered while building and surfaces risk earlier in short cycles.
Then show that you are not dogmatic: waterfall is reasonable when change is expensive or forbidden and early feedback is impossible — fixed scope and price, regulatory sign-off, safety-critical systems, or integrating with a platform that has two release windows a year. In practice most organisations are hybrid: a phased macro plan with iterative internal delivery.
4. Scrum, in full and without mercy
Scrum is a lightweight framework for building complex products, founded on empiricism and lean thinking. The three pillars of empiricism are transparency (the work and its state must be visible to the people who decide), inspection (look at it often) and adaptation (change based on what you saw). Every Scrum event institutionalises one of these; an event that does none of them should be deleted. The five values are not slogans either: "courage" means daring to say "this story is not ready" or "that estimate is unrealistic" in front of a manager.
One full Scrum sprint — یک sprint کامل.
flowchart LR
PB[(Product Backlog)] --> SP[Sprint Planning]
SP --> SB[(Sprint Backlog)]
SB --> DEV[Sprint work + Daily Scrum]
DEV --> INC[Increment meeting the DoD]
INC --> SR[Sprint Review]
SR --> PB
INC --> RETRO[Sprint Retrospective]
RETRO --> SP
4.1 The roles — which the Guide calls accountabilities
The 2020 Guide deliberately dropped the word "role" in favour of accountability: these are not job titles, they are responsibilities. The Scrum Team is one unit, typically ten or fewer people, with no sub-teams.
| Accountability | Accountable for | Has the final say on | Common failure mode |
|---|---|---|---|
| Product Owner | maximising product value and managing the Product Backlog | backlog order, item content, the Product Goal | becomes an "order taker": relays wishes without prioritising or saying no |
| Scrum Master | establishing Scrum and team effectiveness | how events run, removing impediments | turns into a project manager; assigns work instead of facilitating |
| Developers | creating a usable increment and its quality | how the work is done, estimates, the Sprint Backlog | accepts estimates imposed from outside; trades quality for schedule |
If the Scrum Master assigns tasks and collects status in the daily, you have a project manager with a new title and you have killed the team's self-management. The measure of their success is that nothing bad happens when they are away.
The PO is accountable for value and priority — "what, and in what order". They are one person, not a committee, and they have the authority to say no; without that authority they are not really a PO. The Scrum Master is accountable for process and team effectiveness and has no direct authority over product content or work assignment. Business Analyst is not a Scrum role at all; in practice it is whoever helps discover and sharpen requirements.
A line that lands well: "The PO decides what is valuable, the Developers decide how it gets built, the Scrum Master makes sure that conversation stays healthy." If one person wore all three hats on your team, say so honestly and describe the problem it caused.
4.2 Artifacts and their commitments
Every artifact carries a "commitment" that makes it transparent. The Product Backlog is the ordered, always-living list of everything that might be needed, and its commitment is the Product Goal (the team's long-term objective). The Sprint Backlog is the why of the sprint plus the selected items plus the plan to deliver them; it belongs to the Developers, and its commitment is the Sprint Goal, a single objective for the whole sprint. The Increment is a concrete step toward the Product Goal, and its commitment is the Definition of Done — the Guide is blunt about this: work that does not meet the DoD is not part of the increment at all.
Most teams start a sprint with "fourteen unrelated tickets", which means there is no goal to negotiate about. With a Sprint Goal, when urgent work arrives the question becomes simple: "does this help the goal?" If not, either the goal changes deliberately or the work waits. A good Sprint Goal is outcome-shaped — "a customer can pay with a saved card" — not "close eight payment tickets".
4.3 Events: timebox, real purpose, failure mode
Timeboxes below are for a one-month sprint; shorter sprints usually get shorter events.
| Event | Timebox | Empiricism pillar | Real purpose | Common failure mode |
|---|---|---|---|---|
| Sprint | one month or less, fixed length | container for the rest | a steady cadence that bounds risk | extending the sprint "to finish things" |
| Sprint Planning | max 8 hours | adaptation | craft the Sprint Goal and a plan to reach it | a ticket-assignment meeting with no goal |
| Daily Scrum | 15 minutes, every working day | inspection | daily re-planning by the Developers | status report to a manager; three parroted questions |
| Sprint Review | max 4 hours | inspection + adaptation | inspect the increment with stakeholders and adjust the backlog | slides instead of working software |
| Sprint Retrospective | max 3 hours | adaptation | improve process, interactions and tools | a complaint session with no tracked action |
Sprint Planning answers three questions: why is this sprint valuable, what can be done, and how. If most of it is spent haggling over numbers, the problem is not planning — the backlog was not refined enough, and refinement is an ongoing activity during the sprint, not a formal event.
The Daily Scrum is not a reporting meeting; it is a 15-minute re-planning session for the Developers. The right question is not "what did you do yesterday" but "given what we learned yesterday, how do we get closer to the Sprint Goal today, and what is blocking us?" A good technique: instead of walking the people, walk the board right to left — the work closest to done first.
The Sprint Review is not a sign-off; it is a working session with stakeholders where a real increment is shown and the backlog is adjusted on the spot. The Retrospective has an effective pattern: gather data (not just feelings) → pick one or two problems → find the root cause → define at most one or two actions with a named owner → put that action into the Sprint Backlog so it actually happens.
The moment someone with performance-review authority asks "why isn't this done yet?", people start reporting defensively instead of saying the truth ("I'm stuck", "my estimate was wrong"). Transparency — Scrum's first pillar — disappears, and the whole feedback system goes blind.
Transparency first, heroics never: raise it the day you realise it — not on the last day — in the daily. Then put options on the table: reduce the scope of the goal with the PO (which part carries the most value); drop items unrelated to the goal from the Sprint Backlog; redefine the Sprint Goal if reality changed; and, rarely, cancel the sprint — which per the Guide only the PO may do, and only when the goal has become obsolete.
What you must not say: "we'll work overtime" or "we'll skip the tests". A good closing line: "I hold the date and the quality fixed and let scope be the variable, because scope is the only variable you can shrink safely."
4.4 Definition of Done and Definition of Ready
The DoD is a shared quality contract: the list every item must fully satisfy to count as finished. It must not be aspirational; it has to be genuinely met every sprint.
## Definition of Done — Payments squad
- [ ] acceptance criteria met and verified by someone else
- [ ] code reviewed with at least one approval
- [ ] unit tests for new logic + integration test for the main path
- [ ] pipeline green (build, tests, lint, dependency scan)
- [ ] database migration is backward-compatible and rollback tested
- [ ] logs and metrics for the new path visible on the dashboard
- [ ] OpenAPI document and, if needed, the runbook updated
- [ ] deployed to staging behind a feature flag
- [ ] every deliberate shortcut captured as a ticket
The DoR is the list of conditions an item should meet before entering a sprint: a clear goal, acceptance criteria, known dependencies, available test data.
DoR is a common and useful practice, but it is not part of Scrum, and its risk is that it turns into a miniature waterfall: "we will not start until the PO fills in these nine fields." Keep it as a conversation guide, not a contractual gate; the answer to a vague item is usually a short spike, not a rejection.
A strong answer is specific and technical: the Developers define the DoD (if the organisation has a minimum standard, the team may only make it stricter, never looser), it is visible somewhere, and it is referenced during review. Then quote one or two non-obvious lines from your real DoD — "migrations must be backward-compatible because we deploy with rolling updates", or "every new endpoint needs latency and error-rate metrics" — and say what incident caused that line to be added. A story with roots always beats a generic list.
5. The backlog: from epic to a buildable item
A user story is a reminder to have a conversation, not a documentation template: "As a
"As a developer I want to upgrade Hibernate so the code is better" is not a story; it is technical work, and there is nothing wrong with writing it as such. The rule: if the beneficiary is a user, write a story; if the beneficiary is the system, write a technical item with a risk/cost justification — "upgrade X to close CVE-such-and-such; if we skip it we fall out of support".
5.1 INVEST
| Letter | Meaning | The smell when violated |
|---|---|---|
| Independent | as independent of others as possible | "this cannot start until those three finish" |
| Negotiable | details open to discussion | a story with forty lines of locked-down spec |
| Valuable | valuable to a user or the business | "we'll just build the repository layer" |
| Estimable | the team can guess its size | "no idea" → you need a spike |
| Small | fits inside one sprint | an item still open after two sprints |
| Testable | has objective acceptance criteria | "it should be fast" |
5.2 Acceptance criteria
A bullet list is enough for simple cases; Gherkin earns its keep when behaviour is conditional and you want it wired straight into automated tests:
Feature: Pay with a saved card
Scenario: Successful payment with a valid saved card
Given a customer with a saved card ending in 4242
And an open order of 250000 IRR
When the customer confirms payment with that card
Then the order status becomes PAID
And a payment receipt is stored with the gateway reference
Scenario: Payment declined by the issuer
Given a customer with a saved card that the issuer will decline
When the customer confirms payment with that card
Then the order status stays PENDING_PAYMENT
And no receipt is stored
5.3 Slicing: the most important backlog skill
An epic is an item too big for one sprint. The most important rule: slice vertically, not horizontally.
Horizontal versus vertical slicing — تفاوت برش افقی و عمودی.
flowchart TD
subgraph H[Horizontal slicing - nothing usable until the end]
H1[Sprint 1: all DB tables] --> H2[Sprint 2: all services] --> H3[Sprint 3: all UI]
end
subgraph V[Vertical slicing - usable output every sprint]
V1[Slice 1: pay with new card, happy path] --> V2[Slice 2: saved cards] --> V3[Slice 3: refunds]
end
Practical slicing patterns (the SPIDR set plus a couple more): spike when not knowing is the blocker — a timeboxed piece of work whose output is a decision, not code; path — build one path only (pay with a new card) and defer the rest; interface — a minimal interface first; data — one data subset first (one currency, then multi-currency); rules — simple rules first (no discounts, then coupons and caps); error handling later (happy path with one generic error first); and manual before automated (a manual report first, scheduled later).
State your criterion first: every slice must be independently releasable and visible to a user. Then give a concrete example. For "add online payments": slice one is paying for an order with a new card, one currency, happy path only, no saved cards, no refunds, behind a feature flag and limited to internal users. That slice genuinely moves money and produces real feedback. Later slices: error handling and gateway timeouts, saved cards, refunds, reconciliation.
Then show risk thinking: I choose the first slice so it touches the biggest unknown early (the gateway integration), not the easiest part. Slicing is not only about making things smaller — it is about surfacing risk sooner.
6. Estimation: story points, planning poker and the velocity trap
If I ask "how many minutes will it take to lift this weight?", the answer depends on the person. If I ask "how many kilograms is it?", the answer is the same for everyone. A story point is the kilogram: the size of the work (volume + complexity + uncertainty), not the time to do it.
Why points beat hours: hours are person-dependent and always end in "you are faster than me"; humans are bad at absolute estimation and good at relative comparison ("this is about twice that one"); and hours turn into a contractual commitment almost instantly. The common scale is pseudo-Fibonacci (1, 2, 3, 5, 8, 13, 20, 40, 100 plus "?") and the widening gaps are deliberate: the bigger the item, the lower the precision.
Planning poker means everyone reveals their card simultaneously so nobody anchors on the first (or loudest) opinion. Its real value is not the number but the spread: when one person says 2 and another says 13, two very different mental models of the work exist, and that discussion is the most valuable output of the session.
The moment "a point means four hours" is said out loud in an organisation, the story point is dead: it is an hour in a wrapper. The predictable result is that the team inflates estimates to "meet the commitment", or cuts quality. If management wants dates, the correct answer is statistical forecasting from throughput, not a unit conversion.
Velocity is the sum of points completed in a sprint, and its only legitimate use is capacity planning for that same team. Turn it into a KPI and Goodhart's law does its work: estimates inflate, items get declared done prematurely, and quality is the first casualty.
The more mature approach is forecasting instead of committing: throughput (items finished per week) plus Monte Carlo simulation over historical data, which lets you say "85% confidence it lands between six and nine weeks" — an honest, defensible sentence, unlike "seven weeks". If you keep items small and reasonably uniform, counting items works about as well as summing points.
Start with the conceptual difference: a point is the relative size of the work; an hour is a duration and depends on the person, interruptions and meetings. Teams agree on size far more easily than on duration. Then be honest: points are just a tool, and if items are small and uniform, counting items and forecasting from throughput is often more accurate and cheaper.
If they ask "then how do you give the business a date?", say: a probabilistic range from historical data, updated frequently — not a single hard date at the start. Add that you never report velocity as a cross-team performance metric, and explain why.
7. The metrics that actually matter
Burndown shows remaining work against time; burnup shows completed work and total scope as separate lines, and that is exactly why it is better: if scope grows mid-flight, a burnup shows the ceiling line rising, while a burndown just makes it look like the team is slow.
| Category | Metric | What it really tells you | How it gets gamed |
|---|---|---|---|
| Delivery (DORA) | Deployment Frequency | how often change reaches production | meaningless deploys to move the number |
| Delivery (DORA) | Lead Time for Changes | from commit to running in production | measuring from merge time to flatter the number |
| Stability (DORA) | Change Failure Rate | what share of changes cause a failure | not recording small incidents |
| Stability (DORA) | Failed Deployment Recovery Time | how long service restoration takes after a failure | closing the incident before it is really fixed |
| Flow | Cycle Time | from starting an item to finishing it | recording "start" late |
| Flow | WIP | how many things are open at once | ignoring work that never hits the board |
| Flow | Work Item Age | how old the in-progress items are | never using it in the daily |
| Quality | defect escape rate, rollback count | how much reaches the user broken | counting bugs as an individual performance metric |
Always pick metrics in pairs — one speed metric and one stability metric, for example cycle time together with change failure rate. Measure only speed and the team sacrifices quality; measure only quality and the team becomes slow and defensive. And always at team level, never per person.
Defuse the trap first: "no single number captures engineering productivity, and any metric attached to individual performance review will be gamed." Then give a framework: the four DORA metrics (deployment frequency, lead time for changes, change failure rate, failed deployment recovery time), plus flow metrics such as cycle time and WIP, and finally product outcome metrics, which are the only ones that ultimately matter.
Then give an example: "our cycle time was eleven days; the cumulative flow diagram showed most of it was waiting for code review, not coding. Limiting WIP and adopting 'clear the review queue before starting new work' brought it to four days." That shows you use metrics to find bottlenecks, not to judge people.
8. Kanban and managing flow
Kanban is a way to improve an existing workflow without imposing new roles or events. The official guide defines three practices: defining and visualising the workflow, actively managing items in the workflow, and improving the workflow.
The Definition of Workflow (DoW) must at minimum specify: what a work item is, where "started" and "finished" are, the states in between, how WIP is controlled, the explicit policies for each state, and the SLE (Service Level Expectation) — for example "85% of items finish within seven days". The four flow metrics the guide makes mandatory are WIP, throughput, work item age and cycle time.
A Kanban board with WIP limits — بردی با محدودیت WIP.
flowchart LR
A[Backlog] --> B["Ready (3)"]
B --> C["In Progress (2)"]
C --> D["In Review (2)"]
D --> E["Ready to Deploy (3)"]
E --> F[Done]
Little's law for a stable system says average cycle time ≈ average work in progress ÷ average throughput. So if throughput is constant and you halve WIP, cycle time halves — the work does not get faster, it just waits less. This is the highway analogy: a motorway does not gain capacity as it fills; past a point, more cars means everyone is slower. When each person has three open items, nothing finishes, everyone is context-switching, and the manager sees "everyone is busy and nothing ships".
A cumulative flow diagram shows the number of items in each state over time as stacked bands. Reading it is easy: any band that keeps widening is a bottleneck — work is entering that state faster than it leaves. It is usually "In Review".
| Criterion | Scrum fits better when | Kanban fits better when |
|---|---|---|
| Nature of the work | plannable for a fixed period | unpredictable arrival (support, incidents) |
| Need for focus | a shared sprint goal is valuable | priorities change daily |
| Item size | variable, needs planning | small and fairly uniform |
| Team maturity | the team needs structure and cadence | a mature team just optimising flow |
| Main risk | the sprint becomes a two-week mini-waterfall | priority becomes invisible, work becomes endless |
In practice many teams blend the two ("Scrumban"): they keep Scrum's cadence and events but run the board with WIP limits and measure flow instead of sprint commitment. That is entirely legitimate as long as it is deliberate.
A weak answer: "Scrum, two-week sprints." A strong one ties the choice to the nature of the work: "our product team ran Scrum because we could lock two weeks around one goal; the on-call team ran Kanban because 70% of its input was unpredictable and a sprint commitment broke every week." If your team was a hybrid, say so — that is maturity, not disorder. And add a number: "after putting a WIP limit on the review column, cycle time dropped from nine days to five."
9. Scaling: SAFe, LeSS and Nexus at a glance
When several teams work on one product, the new problem is coordinating dependencies, integrating frequently and prioritising jointly. SAFe is the heaviest and most prescriptive answer; the concepts you should recognise are ART (Agile Release Train, a set of teams that release together), PI Planning (a shared planning event for a multi-week increment) and WSJF for prioritisation. LeSS is minimalist: plain Scrum with one Product Owner and one Product Backlog across all teams and shared events — its philosophy is "scale by removing, not by adding". Nexus is a thin layer on top of Scrum for three to nine teams, focused on resolving dependencies and producing one integrated increment.
Squads, tribes, chapters and guilds described one company at one point in time, not a copyable method — and its own authors later warned people not to sell it as a template. Copying another company's org chart does not import its culture or constraints. Be careful with versions too: say "my experience was with SAFe 6.0" rather than claiming familiarity with the very latest release; these frameworks are revised regularly.
10. Working with Jira in practice
Nobody expects you to be a Jira administrator. The expectation is that you can write a good ticket, read a board, build a filter and connect your code to an issue.
10.1 Issue types and workflows
The standard hierarchy, top to bottom, is Epic → Story / Task / Bug → Sub-task (premium plans allow a higher level such as Initiative). An epic is an umbrella for related items serving a larger goal; a story is a change with user value; a task is work with no direct user value (an upgrade, configuration, a migration); a bug is a deviation from expected behaviour; a sub-task is an internal split for daily coordination.
A workflow is the set of statuses and the transitions allowed between them. Every status belongs to a status category — To Do, In Progress or Done — and boards and reports are driven by that.
A realistic Jira workflow — یک workflow واقعی.
stateDiagram-v2
[*] --> Backlog
Backlog --> Ready: refined and accepted
Ready --> InProgress: developer starts
InProgress --> InReview: pull request opened
InReview --> InProgress: changes requested
InReview --> Testing: merged to main
Testing --> InProgress: defect found
Testing --> Done: acceptance criteria verified
Done --> [*]
Teams add columns "for transparency" and the result is that items sit for days in "Ready for QA" with nobody accountable. The rule: every column needs an explicit policy (when work enters, when it leaves, who owns it) and preferably a WIP limit. If you cannot write that policy, delete the column.
Jira Cloud has two project types: team-managed (independent, simpler configuration owned by the team) and company-managed (shared workflow and field schemes across the organisation — less flexible, more consistent). Atlassian has also replaced the old Epic Link field with the unified parent field; use parent in new JQL, and be aware that negative operators such as != can return different results after that change.
10.2 JQL
The basic shape is field operator value, combined with AND / OR / NOT and closed with ORDER BY.
project = PAY AND sprint in openSprints() AND assignee = currentUser() ORDER BY rank ASC
project = PAY AND statusCategory != Done AND priority in (Highest, High) ORDER BY created ASC
project = PAY AND status changed to "In Progress" after -7d ORDER BY updated DESC
project = PAY AND parent = PAY-1200 AND type in (Story, Bug)
| Need | Expression |
|---|---|
| my open work | assignee = currentUser() AND statusCategory != Done |
| items in the current sprint | sprint in openSprints() |
| items from closed sprints | sprint in closedSprints() |
| children of an epic | parent = PAY-1200 |
| members of a group | assignee in membersOf("payments-team") |
| created since the start of the week | created >= startOfWeek() |
| due by end of today | due <= endOfDay() |
| linked to a given issue | issue in linkedIssues(PAY-1200) |
| in an unreleased version | fixVersion in unreleasedVersions(PAY) |
| status changed within a window | status changed to Done during (startOfMonth(), endOfMonth()) |
| free text search | text ~ "timeout" |
| unestimated | "Story Points" is EMPTY |
| stalled more than five days | statusCategory = "In Progress" AND updated <= -5d |
That last row is the work item age view from the Kanban section, and it is a far better daily-standup topic than walking around the people. Save it as a filter and put it on a dashboard.
10.3 Linking code to issues
Put the issue key in the branch name, the commit message and the PR title, and the tooling connects everything for you.
# branch name carries the issue key
git switch -c PAY-1234-saved-card-payment
# commit message with the issue key first
git commit -m "PAY-1234 add saved-card payment flow"
# smart commit: log time, add a comment and transition in one commit
git commit -m "PAY-1234 #time 3h 30m #comment gateway retry added #review"
Smart commit rules worth knowing: the issue key must use the default format (two or more uppercase letters, a hyphen, a number — such as PAY-1234) and appear at the start of the message; #comment turns the text after it into a comment; #time logs work; and for a transition, the first word of the transition name is enough (#review for a transition called "review changes"). The feature must be enabled organisation-wide and depends on a Jira-connected tool.
10.4 What a good ticket looks like
# PAY-1234 — Pay an order with a saved card
## Context
38% of users abandon at the card-entry step (funnel dashboard, last 30 days).
## Goal / Value
Reduce payment abandonment by removing card re-entry for returning customers.
## Scope
- Choose a saved card on the checkout page and pay with the card token
Out of scope: adding a new card here, wallets, instalment payments.
## Acceptance criteria
1. A customer with at least one saved card sees the list with the last four digits.
2. Successful payment -> order becomes PAID and a receipt with the gateway reference is stored.
3. Issuer decline -> order status is unchanged and a comprehensible message is shown.
4. Gateway timeout -> the transaction is idempotently retryable and the customer is not charged twice.
## Non-functional
- p95 latency of the payment endpoint under 800ms at normal load
- no full card data written to logs
## Links
Design: <link> · Gateway spec: <link> · Depends on PAY-1187
Three things make this ticket good: a why backed by a number, an explicit scope boundary (what is out), and testable acceptance criteria that include failure modes.
A weak answer: "I closed tickets." A strong one has three layers. First the data model: issue types and the hierarchy, workflows and status categories, and why your team's board had the columns it had. Second search: one or two real JQL queries you actually used, for example project = PAY AND statusCategory != Done AND updated <= -5d ORDER BY updated ASC to find stalled work. Third the link to engineering: branch naming with the issue key, referencing it in commits and PRs, and automatic transitions on merge.
If your experience was with another tool, use the same structure and note that the concepts transfer; what is being tested is not the tool but the discipline of tracking work.
11. Requirements engineering
Requirements engineering is the process of discovering, analysing, documenting, validating and managing what a system must do. For a backend engineer this is not "the analyst's job"; it is the skill that decides whether your next two weeks are wasted.
11.1 Elicitation
Requirements are not "gathered", they are discovered — because stakeholders describe solutions, not problems. The main techniques: interviews with open questions ("walk me through your current process end to end") and resisting the jump to a solution; field observation, which almost always differs from what was described (the hidden spreadsheet, the workaround); workshops and event storming for collective discovery of the process and domain events (this connects to the DDD chapter); analysing the existing system's data and logs, which usually tells the truth better than people do; prototypes, the cheapest way to convert "I don't know what I want" into concrete feedback; and five whys to get from a want to a need.
If you ask "what do you do with that file?", the answer may be "every morning I open it and email the red rows to my manager". The real need is not a spreadsheet: it is a daily alert. The right solution might be a scheduled email that takes a day to build, not a reporting module. The rule: take the want, but do not start until you understand the job behind it.
11.2 Functional and non-functional requirements
An FR says what the system does; an NFR says how well it does it — and that is where projects die, because NFRs are usually unwritten and discovered after release. ISO/IEC 25010 (2023 revision) defines nine quality characteristics, which make an excellent NFR checklist:
| Quality characteristic | The question to ask |
|---|---|
| Functional suitability | does it do the right thing, completely and correctly? |
| Performance efficiency | what response time, throughput and resource usage? |
| Compatibility | which systems and versions must it coexist with? |
| Interaction capability | how usable and accessible is it? |
| Reliability | availability, fault tolerance, recoverability |
| Security | confidentiality, integrity, non-repudiation, accountability, resistance |
| Maintainability | modularity, testability, modifiability |
| Flexibility | adaptability, installability, replaceability, scalability |
| Safety | preventing unintended harm, fail-safe behaviour, hazard warning |
"It should be fast" is not a requirement, it is a wish. This is a requirement: "p95 latency of the search endpoint under 300ms at 500 requests per second over a 10-million-row dataset." The difference between those two sentences is the difference between adding an index and redesigning your storage. If you do not capture NFRs before design, you have chosen your architecture blind.
11.3 Writing testable requirements
ISO/IEC/IEEE 29148 lists the characteristics of a good requirement: necessary, appropriate, unambiguous, complete, singular, feasible, verifiable, correct, conforming and traceable. Two are violated most often: singular (the word "and" usually means you have two requirements) and verifiable (if you cannot design a test that would show it violated, it is not a requirement).
An under-appreciated but excellent template is EARS (Easy Approach to Requirements Syntax), with five patterns:
Ubiquitous: The <system> shall <response>.
Event-driven: WHEN <trigger>, the <system> shall <response>.
State-driven: WHILE <state>, the <system> shall <response>.
Unwanted: IF <unwanted condition>, THEN the <system> shall <response>.
Optional: WHERE <feature is included>, the <system> shall <response>.
WHEN a payment authorisation request times out after 30 seconds,
the payment service shall retry the authorisation with the same idempotency key
at most twice, and shall mark the payment as UNKNOWN if all attempts time out.
IF the account balance is insufficient,
THEN the payment service shall reject the request with error code INSUFFICIENT_FUNDS
and shall not create a ledger entry.
The power of EARS is that it removes ambiguity without requiring a formal language, and each sentence converts almost directly into a test.
11.4 Prioritisation
| Method | How it works | Best for | Weakness |
|---|---|---|---|
| MoSCoW | Must / Should / Could / Won't for a defined period | negotiating release scope with stakeholders | everything becomes Must unless you cap it |
| RICE | (Reach × Impact × Confidence) ÷ Effort | comparing product ideas with data | the inputs are themselves guesses |
| Kano | separating basic, performance and delight features | deciding what actually differentiates | needs user research |
| WSJF | cost of delay ÷ job size | queueing in a multi-team setting (SAFe) | requires several component estimates |
| Value/risk | highest value and highest risk first | reducing technical risk early | does not directly measure user value |
RICE in detail: reach is the number of people affected in a defined period; impact uses a discrete scale (3 = massive, 2 = high, 1 = medium, 0.5 = low, 0.25 = minimal); confidence is a percentage (100% solid data, 80% good evidence, 50% informed guess); effort is person-months across all roles.
A common practical rule: the Must items should not exceed roughly 60% of the period's capacity, leaving room for Should, Could and the unknown. If everything is a Must, no prioritisation has happened and the first surprise breaks the whole plan. The magic question to break a "it's all critical" meeting: "if we could deliver only one of these, which one?"
11.5 Traceability
Traceability means you can walk from any requirement to the work item, to the code, to the test that verifies it — and back. It is mandatory in regulated environments and a lifesaver everywhere else when you need to answer "what breaks if we delete this?". The lightweight version: put the issue key in the test name and keep a simple coverage table. The query below finds approved requirements with no passing test:
SELECT r.req_key,
r.title,
count(t.id) FILTER (WHERE t.last_result = 'PASS') AS passing_tests
FROM requirement r
LEFT JOIN requirement_test rt ON rt.req_id = r.id
LEFT JOIN test_case t ON t.id = rt.test_id
WHERE r.status = 'APPROVED'
GROUP BY r.req_key, r.title
HAVING count(t.id) FILTER (WHERE t.last_result = 'PASS') = 0
ORDER BY r.req_key;SELECT r.req_key,
r.title,
COUNT(CASE WHEN t.last_result = 'PASS' THEN 1 END) AS passing_tests
FROM requirement r
LEFT JOIN requirement_test rt ON rt.req_id = r.id
LEFT JOIN test_case t ON t.id = rt.test_id
WHERE r.status = 'APPROVED'
GROUP BY r.req_key, r.title
HAVING COUNT(CASE WHEN t.last_result = 'PASS' THEN 1 END) = 0
ORDER BY r.req_key;PostgreSQL supports FILTER (WHERE ...) on aggregate functions, which reads better; Oracle does not have it, so you use COUNT(CASE WHEN ... THEN 1 END). Both are legal inside HAVING. The Oracle/PostgreSQL dialects chapter covers more of these differences.
11.6 Change management, scope creep and feasibility
Scope creep is the silent growth of scope with no change in time or capacity — usually arriving as "could you just add one small thing". It differs from legitimate change in that legitimate change is deliberate and something else is set aside for it. Three practical tools: capture instead of refusing ("I'll log that as an item and we'll prioritise it with the PO"); the trade rule (anything new entering the current sprint must push something of similar size out, and that decision belongs to the PO, not to the requester); and an explicit boundary in the ticket (an "out of scope" section prevents more creep than any meeting).
Before committing to a feature, check four feasibility dimensions: technical (possible with the current architecture and skills?), operational (can it be supported and monitored?), economic (is the value worth the cost?) and legal (personal data, retention, audit). The engineering tool for this is a spike: timeboxed investigation with a specific question and a written outcome.
"While we're in here, let's refactor this too", "let's make this generic for later". This kind of creep is never logged and never seen, but it breaks estimates and makes the PR unreviewable. The right approach: do the refactoring this change requires, and log the rest as tickets. The clean-code chapter discusses that boundary and technical debt in detail.
Show that you know the sentence is not implementable and that you know how to make it implementable. The questions I ask: which operation? for whom, in which scenario? mean or 95th percentile? under what load and what data volume? what does "fast enough" mean and what is that based on (is a user waiting, or is it a background job)? and what is the consequence of exceeding it?
Then I convert it into a testable requirement: "p95 latency of GET /orders under 400ms at 300 requests per second with five million orders" — and immediately say how I would prove it: a load test in a production-like environment plus an alert on the same metric. Closing line: "every NFR needs a number, a load condition and a measurement method, otherwise it is a wish."
12. UML for a backend engineer: only what pays off
UML has more than ten diagram types and a backend engineer really uses five. The goal is neither "complete documentation" nor "generating code from a model"; the goal is transferring an idea to another human with minimal ambiguity.
| Diagram | What it shows | When it genuinely pays off | Mermaid equivalent |
|---|---|---|---|
| Class | static structure: classes and relationships | domain modelling | classDiagram |
| Sequence | interaction over time | a request path across services and its failure points | sequenceDiagram |
| State machine | the lifecycle of an entity | orders, payments, subscriptions | stateDiagram-v2 |
| Activity | a workflow with conditions and parallelism | a business process or complex algorithm | flowchart |
| Component | modules and the interfaces between them | a system map for a new team member | flowchart + subgraph |
| Deployment | mapping to the runtime environment | deployment topology and security boundaries | flowchart + subgraph |
| Use case | actors and system goals | initial scoping with non-technical stakeholders | flowchart |
12.1 Class diagram
The relationship notation to know: <|-- inheritance, *-- composition (the part cannot exist without the whole), o-- aggregation (the part is independent), --> association, ..> dependency, ..|> interface realisation. Visibility markers: + public, - private, # protected.
An order domain model — مدل دامنهی یک سفارش.
classDiagram
class Order {
+UUID id
+OrderStatus status
+Money total()
+void addLine(Product p, int qty)
}
class OrderLine {
+int quantity
+Money unitPrice
}
class Customer {
+UUID id
+String email
}
class Payment {
+UUID id
+Money amount
+PaymentStatus status
}
class PaymentMethod {
<<interface>>
+authorize(Money amount) AuthResult
}
class SavedCard
Customer "1" --> "0..*" Order : places
Order "1" *-- "1..*" OrderLine : contains
Order "1" --> "0..*" Payment : paid by
Payment ..> PaymentMethod : uses
SavedCard ..|> PaymentMethod
Draw class diagrams for the domain, not for all your code. Nobody reads a sixty-class diagram and nobody keeps it current. If your IDE can generate the diagram automatically, drawing it is probably not worth it; if it requires human judgement — which entities exist, which are aggregate roots, what belongs to what — it is. Aggregates and bounded contexts are covered in the DDD chapter.
12.2 Sequence diagram — the highest-value diagram for backend work
A payment flow with a gateway timeout — مسیر پرداخت با timeout.
sequenceDiagram
autonumber
participant C as Client
participant API as Order API
participant PS as Payment Service
participant GW as Card Gateway
participant DB as Ledger DB
C->>API: POST /orders/{id}/pay (Idempotency-Key)
API->>PS: authorize(orderId, cardToken, key)
PS->>DB: insert payment attempt (PENDING, key)
PS->>GW: authorize request
GW--xPS: timeout after 30s
PS->>GW: retry with same key
GW-->>PS: approved (auth ref)
PS->>DB: update payment to AUTHORIZED
PS-->>API: AuthResult(approved)
API-->>C: 200 OK (order PAID)
In a design discussion or a review this diagram is worth two pages of prose, because it shows ordering, network boundaries and failure points at the same time. Always draw a separate version for the error paths; the happy path is usually obvious.
12.3 State machine
Any entity with a status field is already a state machine. Drawing it exposes three things: illegal transitions, dead-end states, and duplicate states that are really the same state.
The payment lifecycle — چرخهی حیات پرداخت.
stateDiagram-v2
[*] --> Pending
Pending --> Authorized: gateway approves
Pending --> Declined: gateway declines
Pending --> Unknown: timeout, result not known
Unknown --> Authorized: reconciliation finds a match
Unknown --> Declined: reconciliation finds no charge
Authorized --> Captured: funds captured
Captured --> Refunded: refund requested
Declined --> [*]
Refunded --> [*]
Captured --> [*]
The biggest mistake in modelling distributed-system state machines is assuming a binary success/failure world. When a request times out you genuinely do not know what the other side did. Without an UNKNOWN state, the code is forced to pick one of two lies, and the result is either a double charge or goods shipped without payment. Drawing the state machine exposes exactly that hole before you write the code.
12.4 Activity, component, deployment and use case
Activity diagrams map to a Mermaid flowchart with decision nodes; they are useful for conditional processes.
Order approval activity — فرایند تأیید سفارش.
flowchart TD
S([Order submitted]) --> V{Stock available?}
V -- no --> B[Backorder and notify customer]
V -- yes --> R{Risk score acceptable?}
R -- no --> M[Send to manual review]
R -- yes --> P[Authorize payment]
P --> Q{Authorized?}
Q -- no --> F[Mark order as payment failed]
Q -- yes --> A[Reserve stock and confirm order]
A --> E([Order confirmed])
Component diagrams show what pieces the system is made of and which interfaces they talk over; their best home is the first page of a service README.
Components of an order service — اجزای سرویس سفارش.
flowchart LR
subgraph OrderService[Order Service]
API[REST API layer]
APP[Application services]
DOM[Domain model]
OUT[Outbox publisher]
REP[JPA repositories]
end
API --> APP --> DOM
APP --> REP
APP --> OUT
REP --> DB[(PostgreSQL)]
OUT --> MQ[[Kafka topic: order-events]]
API --> PAY[Payment Service REST]
Deployment diagrams map software to infrastructure: what runs where and where the network boundaries are.
A deployment topology — توپولوژی استقرار.
flowchart TB
subgraph Edge[Public zone]
LB[Load balancer / TLS termination]
end
subgraph K8s[Kubernetes cluster - private zone]
ING[Ingress controller]
P1[order-service pods x3]
P2[payment-service pods x2]
end
subgraph Data[Data zone]
PG[(PostgreSQL primary + replica)]
RD[(Redis cache)]
KF[[Kafka cluster]]
end
LB --> ING --> P1
ING --> P2
P1 --> PG
P1 --> RD
P1 --> KF
P2 --> PG
Use case diagrams are the simplest and most political: you sit with a non-technical stakeholder and draw the actors and their goals so the system boundary becomes explicit.
Actors and goals — بازیگران و اهداف.
flowchart LR
CU([Customer]) --> UC1[Place an order]
CU --> UC2[Pay for an order]
CU --> UC3[Track order status]
OP([Operations staff]) --> UC4[Issue a refund]
OP --> UC5[Resolve a stuck payment]
UC2 --> GW([Card gateway])
UC4 --> GW
If an organisation demands a diagram "deliverable" for every story, two things happen: the diagrams go stale immediately after merge, and people learn to draw them after writing the code just to fill the form. A diagram earns its keep when it is drawn before a decision so it influences the decision, or when it explains something complex and stable. Everything else belongs on a whiteboard photo.
The answer should be practical and frugal: "For any change that touches more than one service I draw a sequence diagram and put it in the PR description, because it makes ordering and timeout points explicit. For any entity with a status field I draw a state machine and keep it as Mermaid in the same repository so the code and the diagram get reviewed together. For overall architecture I use C4 rather than UML — context and container levels are usually enough. I draw class diagrams only for the domain model."
Then say what you deliberately do not draw: anything the IDE can generate, and any diagram nobody has committed to maintaining. A good closing line: "a worthless diagram is a wrong diagram, and a wrong diagram is worse than none."
13. The C4 model: the modern alternative for architecture diagrams
The problem with typical architecture diagrams is that everything ends up in one picture and nobody knows whether a rectangle is a server, a process, a module or a team. C4 solves this with four levels of zoom, like a map that zooms from a country to a street. System Context: our system as one box, plus users and external systems — the audience is everyone, including non-technical people. Container: the separately runnable or deployable units inside the system — a server-side application, a SPA, a database, a message broker (note that "container" here is an architectural concept and does not necessarily mean Docker). Component: the building blocks inside one container and their responsibilities. Code: class-level detail, usually left to the IDE. There are also supplementary views: system landscape, dynamic and deployment.
C4 level 1, system context — سطح یک C4.
flowchart TB
U([Customer]) -->|places and pays for orders| S[Order Platform]
OPS([Operations staff]) -->|handles refunds| S
S -->|authorises payments| GW[Card Gateway - external]
S -->|sends receipts| MAIL[Email provider - external]
S -->|reports| BI[Data warehouse - internal]
C4 level 2, containers — سطح دو C4.
flowchart TB
U([Customer]) --> SPA[Web SPA - TypeScript]
SPA -->|JSON over HTTPS| API[Order API - Spring Boot]
API -->|JDBC| DB[(Order DB - PostgreSQL)]
API -->|publishes events| K[[Kafka]]
API -->|REST| PAY[Payment Service - Spring Boot]
PAY -->|JDBC| PDB[(Payment DB - PostgreSQL)]
K --> PROJ[Reporting projector - Spring Boot]
PROJ --> WH[(Reporting store)]
C4 can be kept as text (diagrams-as-code), for example with the Structurizr DSL:
workspace "Order Platform" {
model {
customer = person "Customer"
platform = softwareSystem "Order Platform" {
spa = container "Web SPA" "TypeScript"
api = container "Order API" "Spring Boot"
db = container "Order DB" "PostgreSQL"
}
gateway = softwareSystem "Card Gateway" "External"
customer -> spa "Uses"
spa -> api "Calls" "JSON/HTTPS"
api -> db "Reads from and writes to" "JDBC"
api -> gateway "Authorises payments"
}
views {
systemContext platform "Context" {
include *
autoLayout lr
}
container platform "Containers" {
include *
autoLayout lr
}
theme default
}
}
The value of diagrams-as-code is not only version control: you keep one model and generate several views from it, so the views can never contradict each other — unlike drawing tools, where three inconsistent versions of the same architecture typically live in three different slide decks. If Structurizr is too heavy, Mermaid in the repository with one simple rule is enough: any architectural change updates the diagram in the same PR.
14. BPMN at a glance
BPMN is the standard notation for modelling business processes. Its key difference from an activity diagram is that besides being readable by humans it is executable: process engines can run the model directly. The elements to know are pools and lanes (actors and responsibilities), tasks (human or service work), gateways (exclusive, parallel, event-based), events (start, end, timer, message, error) and sequence flows. When a process runs for days, involves human approval, or needs compensation, modelling it in BPMN and running it on a process engine is far more sensible than hand-rolling it with a status column. The workflow-engine chapter covers this fully with runnable examples; here you only need to know when to reach for it.
15. Documentation that survives
Bad documentation is documentation nobody updates. The answer is not "more documentation" but less documentation, closer to the code, with a defined lifespan.
15.1 ADRs — recording architecture decisions
An ADR is a short file that records one important decision and why it was made. Its value is not explaining the decision; it is preserving the options that were rejected and the reasons, so nobody restarts the same argument a year later. The common template is MADR (version 4), with files named NNNN-title-with-dashes.md under docs/decisions/:
---
status: accepted
date: 2026-03-11
decision-makers: payments team
consulted: platform team, security
informed: product
---
# 0007 — Use an outbox table for publishing payment events
## Context and Problem Statement
The payment service must publish an event after every state change. Writing to the
database and the message broker at the same time is not atomic; if the service dies
between the two, we either lose an event or publish one whose transaction never committed.
## Decision Drivers
* no event may be lost if the service crashes; no distributed transaction; publish delay under 5s is acceptable
## Considered Options
* publish directly to the broker after commit
* a distributed (XA) transaction between database and broker
* an outbox table plus a separate publisher
* change data capture on the payment table
## Decision Outcome
Chosen: outbox table plus publisher, because it guarantees atomicity using the same
database transaction and requires no new infrastructure.
### Consequences
* Good: no event is lost; at-least-once delivery with an idempotency key on the consumer
* Bad: one more table and one more background process to operate, and publish latency
depends on the polling interval
### Confirmation
An integration test that kills the service between commit and publish and proves eventual
delivery, plus an alert on outbox queue depth.
The usual statuses are proposed, accepted, rejected, deprecated and superseded by ADR-NNNN. The golden rule: do not edit an ADR, supersede it — the history of decisions is the part with value.
If a decision is expensive to reverse or affects more than one team, write an ADR. Choosing a logging library, usually not. Consistency model, service boundaries, message formats, authentication strategy or database choice, definitely yes. A good ADR takes fifteen minutes to write and pays off for years; if it takes a day, you are writing a report, not an ADR.
15.2 README, runbooks and diagrams-as-code
A README is for someone who wants to run and change the code: what the service does and who owns it, how to start it locally (ideally one command), external dependencies, how to run tests, one architecture diagram, and links to the ADRs. A runbook is for someone paged at three in the morning: operational, step by step, no theory.
# Runbook — payment-service
## Health and dashboards
Health: GET /actuator/health · Dashboard: <link> · Alerts: <link>
## Alert: payment_outbox_lag_high
Meaning: the outbox queue is more than 5 minutes behind; events are not being published.
1. Check the publisher is running: kubectl get pods -l app=payment-outbox
2. Look for broker connection errors: kubectl logs deploy/payment-outbox --since=15m
3. If the broker is down, notify the platform team and wait — no data is lost.
4. If the publisher is crash-looping, roll back to the previous version: <rollback command>
5. If the queue exceeds 100k rows, scale the publisher horizontally.
Escalation: if unresolved within 30 minutes, page the payments on-call engineer.
API documentation should be generated from the code or the contract, not written by hand (the API design chapter covers OpenAPI); the process point here is to push generation and publication into the pipeline so it never lags:
name: docs
on:
push:
branches: [main]
jobs:
publish:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: actions/setup-java@v4
with:
distribution: temurin
java-version: '21'
- name: Generate OpenAPI document
run: ./mvnw -B verify -Dtest=OpenApiExportTest
- name: Upload API docs
uses: actions/upload-artifact@v4
with:
name: openapi
path: target/openapi.json
Any wiki page explaining "how to deploy" that has not been touched in six months is actively dangerous, because someone will trust it. Two rules: keep operational documentation next to the code so it is updated in the same PR, and give every document an owner and a review date — and when it is obsolete, delete it. Deleting wrong documentation is an improvement, not a loss.
The best answer is a real ADR: what the problem was, which options you evaluated, what criteria you used (not just "it was better"), what you chose and what you gave up. Mentioning the negative consequence is usually the winning move: it shows you see the decision realistically rather than as marketing.
If your team had no ADRs, be honest about where you wrote it instead (the PR description, a wiki page) and what problem their absence caused — "six months later the same argument restarted because nobody knew the original reasoning" — and then say how you do it now.
16. Code review culture
Code review has two goals and one is usually forgotten: raising code quality, and spreading knowledge through the team. Pursue only the first and you end up with an irritating quality gate.
| Priority | What you check | Sample question |
|---|---|---|
| 1 | correctness and the problem | does it actually solve the ticket? edge cases? |
| 2 | security and data | is input validated? sensitive data in logs? access control? |
| 3 | concurrency and reliability | idempotency, retries, timeouts, partial-failure behaviour |
| 4 | design and boundaries | is it in the right layer? any leaky abstraction? |
| 5 | tests | do the tests assert behaviour, or just create coverage? |
| 6 | operations | logs, metrics, migrations, backward compatibility |
| 7 | readability | names, function size, comments that explain "why" |
| 8 | style | should be automated, not a human argument |
The further down the list you go, the less human feedback is worth; leave style and formatting to a formatter and linter (see the clean-code chapter) so humans spend their attention on design and risk.
For feedback that lands: label the level — a simple and widely used convention is to prefix comments with blocking: (I will not merge without this), suggestion: (better, but optional), nit: (taste), question: (explain this to me) and praise: (this was well done). Talk about the code, not the person ("this method has three responsibilities", not "you always write huge methods"). Say why ("this query runs inside a loop and causes N+1; a fetch join fixes it" beats "not optimal"). If a thread reaches its third round, a five-minute call saves ten comments. And the receiver has duties too: apply every comment or answer it — an ignored comment teaches the reviewer to care less next time.
Review latency is the team's most expensive invisible waste: while a PR waits two days the author loses context, starts new work (WIP rises) and merge conflicts accumulate. A simple rule that dramatically reduces cycle time is clear the review queue before starting new work — Kanban's "stop starting, start finishing". Pair and mob programming are another form of review rather than a replacement: for complex, unfamiliar or high-stakes work (security, money, data migration) pairing is often faster and less risky than solo work plus a later review, while for routine work it rarely pays.
Review quality collapses as change size grows; past a few hundred lines the reviewer is no longer really reading. If your PR is huge, the problem is not the reviewer, it is how the work was sliced. Fixes: smaller vertical slices, separating pure refactoring from behaviour change into two PRs (this one alone solves half the problem), and feature flags so incomplete-but-safe code can merge early.
First make sure your objection is criteria-based rather than taste — correctness, security, measured performance, or maintenance cost. Then classify the disagreement: if it is taste, mark it nit explicitly and move on. If it is a real risk, express it as a concrete scenario ("if two requests arrive concurrently this decrements stock twice; here is a test that shows it"). Demonstrating with a test or data ends the argument.
If you still do not agree: move from comments to a conversation, and if necessary bring in a third opinion or record the decision in an ADR. And the most important part of the answer: sometimes you disagree and commit anyway, because the cost of blocking exceeds the cost of the difference — and if it later proves wrong, you fix it without blame. That combination of technical reasoning and social maturity is exactly what the interviewer is probing for.
17. How to describe your process experience in an interview
Most candidates lose points here, not because they lack experience but because they describe it in clichés. Three rules. Move from generalities to a specific narrative: "we did Agile" carries no information; "two-week sprints, a DoD that included an integration test and an OpenAPI update, and after one retrospective we agreed to keep PRs under 400 lines" paints a real picture. Give numbers: cycle time, deploys per week, team size, the percentage of unplanned work. Use STAR (situation, task, action, result) and always add one sentence about what you would do differently.
Saying "our process was weak" is not a weakness; being unable to say why it was weak is. A strong answer sounds like this: "we were nominally Scrum, but half of every sprint went to urgent work and the Sprint Goal was meaningless. I started logging every unplanned item on the board; after a month the data showed 40% of capacity went to support. With that data we agreed on a rotating support duty so the rest of the team was protected; cycle time went from twelve days to six." That answer says you observe, you gather data, and you create change without formal authority.
Give a linear narrative, not a list of buzzwords: "A need arrived from product with a problem statement and a number. In refinement we broke it down and wrote acceptance criteria; if it had a technical unknown we defined a two-day spike. In planning we picked a Sprint Goal and pulled items. Development happened on short-lived branches named with the issue key, small PRs with at least one approval, a pipeline running unit and integration tests with Testcontainers, dependency scanning and an image build. Merging to main auto-deployed to staging; production releases were daily and behind feature flags. After release we watched the dashboard and alerts, and closed the flag if anything looked wrong. Every two weeks a retrospective producing at most two actions, which themselves entered the next sprint."
That answer simultaneously proves end-to-end ownership, familiarity with CI/CD, attention to observability, and that process is a tool to you rather than ceremony.
Two traps: complaining ("management was bad") and indifference ("everything was fine"). A good answer names a real problem, analyses its root cause, and proposes a fix with a success measure.
For example: "work started fast but flowed slowly, because everyone had two or three open items and PRs waited for days. My proposed change: WIP limits on the In Progress and In Review columns, plus the rule 'do your pending reviews before starting new work'. The success measure: a drop in the 85th-percentile cycle time and in the average age of open items." Note that the fix is cheap, measurable, and blames nobody.
Lifecycle models differ mainly in feedback timing: waterfall pushes risk to the end, Agile spreads it across short cycles — and neither is universally right. Scrum is an empirical framework with three accountabilities (PO, Scrum Master, Developers), three artifacts with their commitments (Product Goal, Sprint Goal, Definition of Done) and five events, each enacting one pillar of empiricism; any event that produces no decision is ceremony. A backlog stays alive through user stories, INVEST and acceptance criteria, and its most important skill is vertical slicing. Story points serve planning, not commitment; velocity is not a productivity metric, and probabilistic forecasting from throughput is more honest. The metrics that matter are the four DORA keys plus flow metrics (WIP, cycle time, work item age), always chosen in pairs — one speed, one stability. Kanban shortens flow through WIP limits and Little's law and beats Scrum for unpredictable work. In Jira what is really being assessed is a good ticket, a simple workflow, useful JQL and linking commits and PRs to issues. In requirements engineering, discover rather than gather; write NFRs with a number and a load condition (ISO 25010 is a good checklist), use EARS to remove ambiguity, prioritise with MoSCoW and RICE, and contain scope creep by capturing and trading. Keep only the five UML diagrams that pay off, use C4 for architecture, and store everything as diagrams-as-code beside the source. Record expensive-to-reverse decisions in ADRs and write runbooks for three in the morning. In code review, correctness and risk first and style last; small PRs, labelled feedback, no waiting. And in the interview, instead of naming methodologies, give one specific narrative with numbers and one honest reflection.