Craft & Process · مهارت و فرایند متوسطIntermediate ~95 دقیقه مطالعه~86 min read

دیسیپلین تست: از Test Plan تا تست بارTesting Discipline: From Test Plans to Load Testing

از سطوح و انواع تست و تکنیک‌های طراحی تست‌کیس (افراز هم‌ارزی، مقادیر مرزی، جدول تصمیم، انتقال حالت، pairwise) تا Test Plan و traceability و گزارش نقص، هرم تست و تست قرارداد، مدیریت تست‌های flaky و داده‌ی تست، mutation testing، و تست غیرکارکردی به‌صورت عمیق با JMeter و Gatling و k6 — به‌علاوه‌ی TDD/BDD صادقانه و نحوه‌ی حرف‌زدن درباره‌ی تست در مصاحبه.From test levels and types and the design techniques that actually generate test cases (equivalence partitioning, boundary values, decision tables, state transition, pairwise) to test plans, traceability, defect reports, the pyramid versus the ice-cream cone, contract tests, flaky-test and test-data management, mutation testing, and non-functional testing in depth with JMeter, Gatling and k6 — plus honest TDD/BDD and how to talk about testing in an interview.

پیش‌نیاز:Prerequisites: تست: JUnit 5، Mockito، AssertJ و TestcontainersTesting: JUnit 5, Mockito, AssertJ, Testcontainers


تو احتمالاً بلدی @Test بنویسی. با Mockito mock می‌سازی، با Testcontainers یک Postgres واقعی بالا می‌آوری و تست integration می‌نویسی. این‌ها را فصل «تست» همین سایت پوشش داده است. اما آگهی‌های شغلی چیز دیگری هم می‌خواهند و همان‌جاست که اکثر مهندس‌های پنج‌ساله می‌لغزند: «آشنایی با سطوح و انواع تست»، «توانایی نوشتن Test Plan و Test Case»، «تجربه‌ی تست بار و کارایی».

این‌ها مکانیک نیستند، دیسیپلین هستند. فرق یک مهندس میانی و یک مهندس ارشد در تست این نیست که کدام یک verify() بلد است؛ این است که وقتی به او می‌گویند «این feature را تست کن»، یکی چند تست تصادفی می‌نویسد و دیگری در ده دقیقه فهرستی از موارد تست بیرون می‌دهد که هیچ حالت مهمی را جا نینداخته، می‌داند کدام‌شان را خودکار کند و کدام را نه، و می‌تواند بگوید «با این مجموعه تست، چه ریسکی هنوز باز است».

این فصل دقیقاً همان لایه است: چارچوب فکری، واژگان حرفه‌ای، تکنیک‌های تولید test case، مستنداتی که یک سازمان بالغ انتظار دارد، و تست غیرکارکردی — به‌ویژه تست بار — با ابزارهای واقعی و مثال قابل اجرا.

نقشه‌ی راه این فصل

اول واژگان پایه را می‌سازیم (خطا/نقص/خرابی، verification در برابر validation) و سطوح تست و انواع تست را مرزبندی می‌کنیم. بعد جعبه‌سیاه/سفید/خاکستری و هفت تکنیک طراحی تست با مثال حل‌شده. سپس مستندات: Test Plan، Test Case، ماتریس ردیابی، گزارش نقص، معیار خروج. بعد هرم تست در برابر مخروط بستنی و جای contract test. سپس استراتژی اتوماسیون: چه چیزی را خودکار کنیم، تست‌های flaky، داده و محیط تست، coverage و mutation testing. بعد بخش سنگین: تست غیرکارکردی — performance/load/stress/soak/spike، SLO و مدل بار، و مقایسه‌ی JMeter، Gatling، k6 با مثال اجراشدنی و روش خواندن نتایج. بعد تست امنیت (SAST/DAST/SCA) و دسترس‌پذیری. بعد TDD و BDD بدون شعار، با Cucumber جاوا. و آخر: چطور در مصاحبه درباره‌ی تست حرف بزنی و چک‌لیست قضاوت درباره‌ی یک test suite.


بخش ۰ — واژگانی که سِنیورها اشتباه به کار نمی‌برند

پرواز و چک‌لیست پیش از برخاست

خلبان قبل از هر پرواز چک‌لیست می‌خواند. نه چون فراموش‌کار است، بلکه چون در شرایط فشار، حافظه‌ی انسان بدترین ابزار جهان است. چک‌لیست یعنی «فکرکردن را از قبل انجام داده‌ایم». تست خوب هم دقیقاً همین است: تصمیم درباره‌ی این‌که چه چیزی باید درست کار کند، قبل از فشارِ لحظه‌ی تحویل گرفته می‌شود. اگر تست‌هایت را وقتی می‌نویسی که PR باز است و مدیر پشت سرت ایستاده، چیزی که نوشته‌ای «چک‌لیست» نیست، «آرامش‌بخش» است.

خطا، نقص، خرابی

سه کلمه که در محاوره یکی گرفته می‌شوند و در گفت‌وگوی حرفه‌ای سه چیزند:

  • error (mistake) — کار انسانی اشتباه. برنامه‌نویس <= را < نوشت.
  • defect (fault, bug) — نتیجه‌ی آن خطا در کد یا مستند. آن < اشتباه، همان‌جا در فایل نشسته است.
  • failure — وقتی آن نقص در اجرا خودش را نشان می‌دهد و رفتار سیستم با انتظار فرق می‌کند. کاربر با موجودی دقیقاً ۱۰۰۰۰ نمی‌تواند برداشت کند.

چرا مهم است؟ چون یک نقص می‌تواند سال‌ها بدون failure بماند (کد آن مسیر اجرا نشود) و چون همه‌ی failureها از نقص کد نمی‌آیند — شرایط محیطی (حافظه‌ی خراب، ساعت سیستم، قطع شبکه) هم failure می‌سازند.

verification در برابر validation

  • verification: «آیا محصول را درست ساختیم؟» یعنی مطابق مشخصات و طراحی است.
  • validation: «آیا محصول درست را ساختیم؟» یعنی نیاز واقعی کاربر را حل می‌کند.

می‌شود سیستمی داشت که صد درصد تست‌هایش سبز است (verification عالی) و هیچ‌کس از آن استفاده نمی‌کند (validation صفر). این تمایز، قلب بحث UAT است.

قضاوت ارشد — «کیفیت» عدد نیست، ریسکِ باقی‌مانده است — هدف تست، «اثبات بی‌نقصی» نیست؛ چون اثبات بی‌نقصی برای هر نرم‌افزار غیرتریویال ناممکن است (تعداد ورودی‌های ممکن بی‌نهایت است). هدف، کاهش ریسک تا سطح قابل قبول با هزینه‌ی منطقی است. هر وقت کسی پرسید «چقدر تست کافی است؟»، جواب حرفه‌ای این است: «تا وقتی ریسکِ باقی‌مانده از آستانه‌ی پذیرش کسب‌وکار پایین‌تر بیاید» — و بعد باید بتوانی آن ریسک را نام ببری.

چرا زودتر ارزان‌تر است

هرچه نقص دیرتر پیدا شود، گران‌تر است — نه به‌خاطر یک ضریب جادویی، بلکه به دلایل ساده: کد بیشتری روی آن ساخته شده، افراد بیشتری درگیر می‌شوند، و اگر به تولید رسیده باشد باید data fix و اطلاع‌رسانی و احتمالاً جبران خسارت هم بکنی. به همین دلیل shift-left (بردن تست به سمت چپِ خط زمانی: بازبینی نیازمندی، تست واحد، تست در CI) اقتصاد دارد، نه فقط زیبایی.

و در کنارش shift-right: بخشی از تست را نمی‌شود قبل از تولید انجام داد (رفتار کاربر واقعی، ترافیک واقعی). آن‌جا canary release، feature flag، synthetic monitoring و chaos engineering ابزار تو هستند — که با فصل observability و resilience گره می‌خورند.


بخش ۱ — سطوح تست: چه کسی چه چیزی را می‌بندد

سطح تست (test level) یعنی گروهی از فعالیت‌های تست که با هم مدیریت می‌شوند و معمولاً به یک مرحله از توسعه گره خورده‌اند. استاندارد رایج (ISTQB نسخه‌ی ۴) پنج سطح می‌شمارد:

سطح موضوع تست چه کسی معمولاً محیط چه چیزی را می‌گیرد
Component (unit) یک کلاس/تابع/ماژول به‌تنهایی توسعه‌دهنده لپ‌تاپ + CI منطق، شرط‌ها، مرزها
Component integration تعامل بین اجزای داخل یک سرویس توسعه‌دهنده CI با Testcontainers نگاشت، تراکنش، قرارداد داخلی
System کل سرویس/محصول از بیرون تیم توسعه/QA محیط شبیه تولید جریان end-to-end، رفتار غیرکارکردی
System integration تعامل چند سیستم/سرویس بیرونی QA/تیم یکپارچه‌سازی staging با sandbox شرکا پروتکل، نسخه، خطاهای شریک
Acceptance آمادگی برای تحویل و استفاده کسب‌وکار/کاربر/عملیات UAT/pre-prod «آیا این همانی است که خواستیم؟»

نمودار: نگاشت سطوح تست به سطوح طراحی در مدل V — Mapping test levels to design levels in the V-model.

flowchart LR
  R[Requirements] --> A[Architecture]
  A --> D[Detailed design]
  D --> C[Code]
  C --> UT[Component tests]
  UT --> IT[Component integration]
  IT --> ST[System tests]
  ST --> AT[Acceptance tests]
  R -.validated by.-> AT
  A -.verified by.-> ST
  D -.verified by.-> IT

زیرگونه‌های acceptance که در مصاحبه پرسیده می‌شود

  • UAT (User Acceptance Test) — کاربر/مالک محصول با سناریوهای واقعی کسب‌وکار.
  • OAT (Operational Acceptance Test) — تیم عملیات: آیا backup کار می‌کند؟ روال rollback؟ alertها؟ مستند runbook؟ این را تیم‌های جوان همیشه فراموش می‌کنند و بعد شب استقرار می‌فهمند.
  • Contractual / Regulatory acceptance — انطباق با قرارداد یا مقررات (مثلاً الزامات نگهداری لاگ، حریم خصوصی).
  • Alpha / Beta — alpha در محل سازنده با کاربران واقعی؛ beta در محیط خود کاربر.
تله‌ی مرزهای مبهم

شایع‌ترین بیماری در تیم‌های ایرانی و غیرایرانی یکی است: «تست واحد»ی که Spring context بالا می‌آورد و به دیتابیس می‌زند. اسمش unit است، رفتارش system است، سرعتش فاجعه است و وقتی می‌شکند نمی‌دانی مقصر منطق است یا محیط. مرز را با «چه چیزی می‌تواند این تست را بشکند» تعریف کن، نه با اسم فایل. اگر تغییر در schema دیتابیس می‌تواند تستت را قرمز کند، آن تست unit نیست.

تفاوت unit و integration test دقیقاً کجاست؟ مرزش را چطور تعیین می‌کنی؟

مرز را با وابستگی‌های خارج از فرایند (out-of-process) تعریف می‌کنم، نه با تعداد کلاس‌ها. تست واحد فقط کد را در همان process اجرا می‌کند، هیچ I/O واقعی ندارد (نه شبکه، نه دیسک، نه ساعت سیستم بدون کنترل)، در چند میلی‌ثانیه تمام می‌شود و قطعی است. اگر یک تست چند کلاس همکار را با هم صدا بزند ولی همچنان درون-فرایندی و قطعی باشد، من هنوز آن را unit می‌دانم — این همان مکتب «سوشیال/کلاسیک» است در برابر مکتب «solitary/mockist» که هر همکاری را mock می‌کند.

تست یکپارچه‌سازی جایی است که واقعاً از مرز فرایند رد می‌شویم: دیتابیس واقعی با Testcontainers، broker واقعی، فایل‌سیستم. ارزشش این است که چیزهایی را می‌گیرد که mock هرگز نمی‌گیرد — نگاشت ستون‌ها، رفتار isolation level، محدودیت‌های unique، سریال‌سازی پیام.

در عمل قانون من این است: منطق تصمیم‌گیری را در هسته‌ی خالص نگه دار و با تست واحدِ ارزان و انبوه پوشش بده؛ برای هر adapter (repository, client, listener) یک لایه‌ی نازک از تست یکپارچه‌سازی بنویس که فقط ترجمه را بررسی کند؛ و چند تست system برای جریان‌های حیاتی. این همان چیزی است که معماری hexagonal را از نظر تست ارزشمند می‌کند.


بخش ۲ — انواع تست: کارکردی، غیرکارکردی، ساختاری، و تست تغییر

سطح می‌گوید «کجا تست می‌کنیم»؛ نوع (test type) می‌گوید «دنبال چه ویژگی‌ای هستیم». چهار دسته:

  1. Functional — «چه کاری انجام می‌دهد؟» درستی محاسبه، قواعد کسب‌وکار، جریان‌ها.
  2. Non-functional — «چقدر خوب انجام می‌دهد؟» کارایی، امنیت، قابلیت اطمینان، قابلیت استفاده، سازگاری، قابلیت نگهداشت، قابلیت حمل. این‌ها همان مشخصه‌های کیفی ISO/IEC 25010 هستند.
  3. Structural (white-box) — «آیا ساختار کد پوشش داده شده؟» پوشش دستور، شاخه، مسیر، MC/DC.
  4. Change-relatedconfirmation test (تست تأیید: همان باگ واقعاً رفع شد؟) و regression test (تست بازگشتی: چیز دیگری خراب نشد؟).

ISO/IEC 25010 و نسخه‌ی ۲۰۲۳ — مدل کیفیت محصول ISO/IEC 25010 در بازنگری ۲۰۲۳ به ۹ مشخصه رسید و «Safety» به‌عنوان مشخصه‌ی مستقل اضافه شد (پیش‌تر ۸ مشخصه بود). وقتی در مصاحبه از «انواع تست غیرکارکردی» پرسیدند، بردن نام این مدل و شمردن چند مشخصه‌ی آن — performance efficiency، reliability، security، usability، compatibility، maintainability، portability، functional suitability، safety — نشان می‌دهد چارچوب داری، نه فهرست حفظی.

غیرکارکردی را به آخر پروژه هل ندهی — تست غیرکارکردی «فاز آخر» نیست. اگر معماری‌ات با ۵۰ کاربر همزمان کار می‌کند و SLO تو ۵۰۰۰ است، هیچ مقدار «تیونینگ در هفته‌ی آخر» نجاتت نمی‌دهد؛ باید مدل داده و الگوی دسترسی عوض شود. یک تست بار کوچک اما زودهنگام روی مسیر بحرانی، ارزشش از یک تست بار عظیم دو روز قبل از go-live بیشتر است.

فرق smoke، sanity و regression چیست؟

این سه، «انواع تست» به معنای فنی نیستند؛ بسته‌های اجرایی با هدف متفاوت هستند و همین را باید توضیح داد.

Smoke test یک بسته‌ی خیلی کوچک و سریع است که بعد از هر build اجرا می‌شود تا بگوید «آیا این نسخه اصلاً قابل تست است؟» — سرویس بالا می‌آید، health سبز است، ورود کار می‌کند، یک تراکنش ساده رد می‌شود. اگر smoke قرمز شود، اصلاً وارد تست عمیق نمی‌شویم و build را برمی‌گردانیم. عمق کم، عرض زیاد.

Sanity test برعکس است: باریک ولی عمیق. بعد از یک رفع باگ یا تغییر کوچک، فقط همان ناحیه و اطرافش را با دقت بررسی می‌کنیم تا مطمئن شویم منطق درست شده — بدون اجرای کل مجموعه.

Regression test بسته‌ی کاملی است که مطمئن می‌شود تغییر جدید، رفتار قبلی را نشکسته. چون بزرگ است، معمولاً بخشی از آن در هر PR و کاملش شبانه اجرا می‌شود؛ و انتخاب هوشمند (test impact analysis) بر اساس این‌که چه کدی تغییر کرده، زمان اجرا را چند برابر کم می‌کند.

نکته‌ای که معمولاً امتیاز می‌آورد: confirmation test با regression فرق دارد. confirmation یعنی «همان باگ گزارش‌شده واقعاً رفع شد؟»؛ regression یعنی «چیز دیگری خراب نشد؟». هر رفع باگ باید هر دو را داشته باشد و تست confirmation باید برای همیشه در suite بماند.


بخش ۳ — جعبه‌ی سیاه، سفید، خاکستری

ماشین لباسشویی

اگر فقط با دفترچه‌ی راهنما کار کنی و دکمه بزنی و نتیجه را ببینی، جعبه‌سیاه‌کاری. اگر پشت دستگاه را باز کنی و سیم‌ها و برد را دنبال کنی، جعبه‌سفید. اگر دفترچه دستت باشد ولی بدانی داخلش یک موتور اینورتر است و بر همان اساس آزمون بچینی، جعبه‌خاکستری.

رویکرد مبنا نقطه‌ی قوت نقطه‌ی کور
Black-box مشخصات/نیازمندی مستقل از پیاده‌سازی، refactor آن را نمی‌شکند شاخه‌های داخلی و مسیرهای خطای پنهان
White-box ساختار کد پوشش شاخه‌ها، کد مرده، شرط‌های مرکب نیازمندیِ جاافتاده را هرگز پیدا نمی‌کند
Grey-box مشخصات + دانش داخلی تست هوشمندانه‌ی مرزها، cache، ایندکس، صف وابستگی نسبی به پیاده‌سازی

نکته‌ی کلیدی که سِنیورها می‌گویند: پوشش صددرصدی کد، هیچ تضمینی درباره‌ی نیازمندی‌های نوشته‌نشده نمی‌دهد. اگر توسعه‌دهنده یادش رفته باشد حالت «موجودی منفی» را پیاده کند، هیچ ابزار پوشش کدی به تو نمی‌گوید که چیزی کم است. برای همین تکنیک‌های جعبه‌سیاه (بخش بعد) پایه‌ی کارند و white-box مکمل.


بخش ۴ — تکنیک‌های طراحی تست: چطور از «نیازمندی» به «فهرست تست» برسی

این بخش قلب فصل است. اگر فقط یک بخش را عمیق یاد بگیری، همین باشد؛ چون در مصاحبه معمولاً یک نیازمندی کوچک به تو می‌دهند و می‌گویند «موارد تست را بنویس».

۴٫۱ افراز هم‌ارزی (Equivalence Partitioning)

ایده: ورودی‌ها را به دسته‌هایی تقسیم کن که سیستم انتظار می‌رود با همه‌ی اعضای یک دسته یکسان رفتار کند. بعد از هر دسته یک نماینده تست کن. اگر یک نماینده باگ را نشان دهد، بقیه‌ی اعضای همان دسته هم می‌دادند؛ پس تست بیشتر، هزینه‌ی بیشتر بدون اطلاعات بیشتر.

مثال کاری — کارمزد انتقال وجه:

مبلغ انتقال بین ۱۰٬۰۰۰ تا ۵۰٬۰۰۰٬۰۰۰ ریال مجاز است. تا ۱٬۰۰۰٬۰۰۰ کارمزد ثابت ۵٬۰۰۰؛ از ۱٬۰۰۰٬۰۰۰ تا ۱۰٬۰۰۰٬۰۰۰ کارمزد ۰٫۱٪؛ بالای ۱۰٬۰۰۰٬۰۰۰ کارمزد ثابت ۲۵٬۰۰۰. مبلغ باید عدد صحیح باشد.

افرازها (هم معتبر هم نامعتبر — این نکته‌ای است که اکثر کاندیداها جا می‌اندازند):

# افراز نمونه معتبر؟
P1 مبلغ < 10,000 9,000 نامعتبر
P2 10,000 ≤ مبلغ ≤ 1,000,000 500,000 معتبر (کارمزد ثابت)
P3 1,000,000 < مبلغ ≤ 10,000,000 5,000,000 معتبر (درصدی)
P4 10,000,000 < مبلغ ≤ 50,000,000 30,000,000 معتبر (ثابت بالا)
P5 مبلغ > 50,000,000 60,000,000 نامعتبر
P6 مبلغ غیرعددی/اعشاری/منفی/خالی "abc"، 1000.5، -5، null نامعتبر
قضاوت ارشد — یک نامعتبر در هر تست، نه بیشتر

در تست‌های معتبر می‌توانی چند افراز را در یک test case ترکیب کنی (مثلاً مبلغ معتبر + ارز معتبر + کاربر فعال). اما برای افرازهای نامعتبر هر بار فقط یکی را نامعتبر کن. چرا؟ اگر هم مبلغ منفی بدهی و هم ارز ناشناخته، سیستم روی اولی خطا می‌دهد و تو هرگز نمی‌فهمی که اعتبارسنجی ارز اصلاً وجود دارد یا نه. این پدیده را fault masking (پوشاندن نقص) می‌گویند.

۴٫۲ تحلیل مقادیر مرزی (Boundary Value Analysis)

ایده: باگ‌ها روی مرزها لانه می‌کنند، چون < و <= یک کاراکتر فاصله دارند. برای هر مرز، مقادیر اطرافش را تست کن.

دو مکتب:

  • ۲-مقداری (2-value BVA): خودِ مرز و اولین مقدار خارج از آن. برای مرز ۱٬۰۰۰٬۰۰۰ یعنی 1,000,000 و 1,000,001.
  • ۳-مقداری (3-value BVA): یکی قبل، خودِ مرز، یکی بعد. یعنی 999,999, 1,000,000, 1,000,001. سخت‌گیرانه‌تر و برای منطق مالی/ایمنی توصیه می‌شود.

برای مثال بالا با ۳-مقداری، مقادیر تست: 9,999 / 10,000 / 10,001 و 999,999 / 1,000,000 / 1,000,001 و 9,999,999 / 10,000,000 / 10,000,001 و 49,999,999 / 50,000,000 / 50,000,001.

و همین را می‌شود مستقیم به یک تست پارامتری JUnit 5 تبدیل کرد:

@ParameterizedTest(name = "amount={0} -> fee={1}")
@CsvSource({
    "10_000,      5_000",
    "1_000_000,   5_000",
    "1_000_001,   1_000",     // 0.1% of 1,000,001 rounded down
    "10_000_000, 10_000",
    "10_000_001, 25_000",
    "50_000_000, 25_000"
})
void feeIsCorrectAtBoundaries(long amount, long expectedFee) {
    assertThat(feeCalculator.feeFor(amount)).isEqualTo(expectedFee);
}

@ParameterizedTest
@ValueSource(longs = {9_999L, 50_000_001L, -1L, 0L})
void amountsOutsideAllowedRangeAreRejected(long amount) {
    assertThatThrownBy(() -> feeCalculator.feeFor(amount))
        .isInstanceOf(AmountOutOfRangeException.class);
}
مرز فقط عدد نیست

مرزهایی که فراموش می‌شوند: طول رشته (۰، ۱، حداکثر، حداکثر+۱)، اندازه‌ی مجموعه (لیست خالی، تک‌عضوی، دقیقاً به اندازه‌ی page size، یکی بیشتر)، زمان (نیمه‌شب، آخرین روز ماه، ۲۹ فوریه، تغییر ساعت تابستانی، مرز timezone)، انواع عددی (Integer.MAX_VALUE, سرریز long, -0.0, NaN)، و یونیکد (emoji چهاربایتی که طول String را با تعداد کاراکترِ دیده‌شده متفاوت می‌کند). یک باگ کلاسیک: فیلد «حداکثر ۲۰۰ کاراکتر» که با ۲۰۰ emoji می‌شکند چون ستون دیتابیس VARCHAR(200) بایتی است.

۴٫۳ جدول تصمیم (Decision Table)

وقتی خروجی به ترکیب چند شرط بستگی دارد، افراز هم‌ارزی کافی نیست؛ باید ترکیب‌ها را سیستماتیک بچینی.

مثال کاری — اجازه‌ی برداشت: شرایط: (C1) حساب فعال است؟ (C2) موجودی کافی است؟ (C3) کاربر احراز دومرحله‌ای کرده؟ (C4) مبلغ بالای سقف روزانه است؟

قانون C1 فعال C2 موجودی C3 دومرحله‌ای C4 بالای سقف نتیجه
R1 N رد: حساب مسدود
R2 Y N رد: موجودی ناکافی
R3 Y Y N رد: نیاز به تأیید دومرحله‌ای
R4 Y Y Y Y نیاز به تأیید دستی
R5 Y Y Y N تأیید فوری

نکته‌ی مهم: علامت یعنی بی‌اهمیت (don't care)؛ جدول کامل ۲⁴=۱۶ ترکیب دارد ولی با ادغام قوانین به ۵ رسیدیم. این «فشرده‌سازی» جدول تصمیم است و همان چیزی است که مصاحبه‌گر دنبالش می‌گردد: نشان بده که می‌فهمی کدام ترکیب‌ها غیرقابل‌دسترس یا بی‌اثر هستند.

قضاوت ارشد — جدول تصمیم را وارد کد کن — اگر یک جدول تصمیم را کشیدی، همان جدول را به‌عنوان منبع حقیقت وارد تست کن. با @CsvSource هر ردیف یک قانون می‌شود و اگر فردا کسی قانونی اضافه کرد، دیف PR دقیقاً نشان می‌دهد کدام قانون تغییر کرده. این کار سند و تست را یکی می‌کند و از «مستندی که با کد هم‌گام نیست» جلوگیری می‌کند.

@ParameterizedTest(name = "R{index}: active={0} funds={1} mfa={2} overLimit={3} -> {4}")
@CsvSource({
    "false, true,  true,  false, ACCOUNT_BLOCKED",
    "true,  false, true,  false, INSUFFICIENT_FUNDS",
    "true,  true,  false, false, MFA_REQUIRED",
    "true,  true,  true,  true,  MANUAL_REVIEW",
    "true,  true,  true,  false, APPROVED"
})
void withdrawalDecisionTable(boolean active, boolean funds, boolean mfa,
                             boolean overLimit, Decision expected) {
    var request = new WithdrawalRequest(active, funds, mfa, overLimit);
    assertThat(policy.decide(request)).isEqualTo(expected);
}

۴٫۴ تست انتقال حالت (State Transition Testing)

وقتی سیستم حافظه دارد — یعنی پاسخ به یک رویداد به تاریخچه بستگی دارد — باید ماشین حالت بکشی.

نمودار: ماشین حالت یک سفارش و رویدادهای مجاز — Order state machine with the legal events.

stateDiagram-v2
  [*] --> CREATED
  CREATED --> PAID: pay
  CREATED --> CANCELLED: cancel
  PAID --> SHIPPED: ship
  PAID --> REFUNDED: refund
  SHIPPED --> DELIVERED: deliver
  SHIPPED --> RETURNED: return
  DELIVERED --> RETURNED: return
  RETURNED --> REFUNDED: refund
  CANCELLED --> [*]
  REFUNDED --> [*]

از این نمودار سه سطح پوشش بیرون می‌آید:

  • پوشش حالت (state coverage): هر حالت حداقل یک بار دیده شود. ضعیف‌ترین سطح.
  • پوشش انتقال / 0-switch: هر یال حداقل یک بار طی شود. این حداقلِ قابل‌قبول است.
  • 1-switch: هر جفتِ متوالی از انتقال‌ها (مثلاً pay سپس ship) طی شود. باگ‌های وابسته به تاریخچه را می‌گیرد.

و مهم‌تر از همه: انتقال‌های نامعتبر. جدول کامل حالت‌ها × رویدادها را بکش و خانه‌های خالی را تست کن — یعنی «رویداد ship در حالت CREATED چه می‌کند؟» جواب درست باید «خطای مشخص و بدون تغییر حالت» باشد، نه NullPointerException.

@ParameterizedTest
@EnumSource(OrderState.class)
void shipIsOnlyLegalFromPaid(OrderState from) {
    Order order = Order.inState(from);
    if (from == OrderState.PAID) {
        order.ship();
        assertThat(order.state()).isEqualTo(OrderState.SHIPPED);
    } else {
        assertThatThrownBy(order::ship)
            .isInstanceOf(IllegalStateTransitionException.class);
        assertThat(order.state()).isEqualTo(from);   // no side effect
    }
}

انتقال نامعتبر باید «بی‌اثر» باشد، نه فقط خطا بدهد — شایع‌ترین باگ ماشین‌های حالت این است که متد اول side effect می‌زند (رویداد منتشر می‌کند، فیلد را عوض می‌کند) و بعد شرط را چک می‌کند و exception می‌اندازد. اگر تراکنش دور کار نباشد، حالا حالت نیمه‌تغییریافته داری. همیشه در تست انتقال نامعتبر، علاوه بر exception، عدم تغییر حالت و عدم انتشار رویداد را هم assert کن.

۴٫۵ تست جفتی (Pairwise / All-Pairs)

وقتی چند پارامتر مستقل داری، ضرب دکارتی منفجر می‌شود. فرض کن: مرورگر (۴) × سیستم‌عامل (۳) × زبان (۵) × نوع پرداخت (۴) × وضعیت کاربر (۳) = ۷۲۰ ترکیب. تست همه‌شان غیرممکن است.

مشاهده‌ی تجربی: اکثریت قاطع نقص‌ها با تعامل یک یا دو پارامتر فعال می‌شوند، نه پنج پارامتر. پس اگر مجموعه‌ای بسازی که هر جفت مقدار از هر دو پارامتر حداقل یک بار کنار هم بیاید، با ~۲۰ ترکیب همان قدرت کشف را می‌گیری.

ابزار عملی: PICT (مایکروسافت، متن‌باز) یا ACTS (NIST). فایل ورودی PICT ساده است:

Browser:  Chrome, Firefox, Safari, Edge
OS:       Windows, macOS, Linux
Locale:   fa-IR, en-US, ar-SA, tr-TR, de-DE
Payment:  Card, Wallet, Gateway, COD
UserTier: Guest, Basic, Premium

IF [Payment] = "COD" THEN [UserTier] <> "Guest";
# تولید مجموعه‌ی جفتی (پیش‌فرض PICT همان order=2 است)
pict model.txt > pairs.tsv

# پوشش سه‌تایی برای مسیرهای پرریسک
pict model.txt /o:3 > triples.tsv
pairwise درمان همه‌چیز نیست

سه تله: (۱) اگر پارامترها واقعاً مستقل نباشند و منطق کسب‌وکار به ترکیب خاصی وابسته باشد، pairwise ممکن است دقیقاً آن ترکیب را تولید نکند — پس ترکیب‌های حیاتی را دستی pin کن. (۲) pairwise به تو نمی‌گوید نتیجه‌ی مورد انتظار چیست؛ فقط ورودی می‌سازد. اوراکل تست هنوز کار توست. (۳) خروجی pairwise بین دو اجرا می‌تواند فرق کند؛ مجموعه‌ی تولیدشده را در repo commit کن تا تست‌ها تکرارپذیر بمانند.

۴٫۶ حدس خطا و تست اکتشافی

Error guessing یعنی استفاده از تجربه برای حدس‌زدن جاهایی که معمولاً می‌شکند: رشته‌ی خالی، فاصله‌ی ابتدا/انتها، کاراکتر ' در ورودی، تاریخ ۱۴۰۴/۱۲/۳۰، فایل صفربایتی، دو بار کلیک روی «پرداخت»، بستن مرورگر وسط تراکنش، بازگشت با دکمه‌ی back بعد از موفقیت.

این را حرفه‌ای کن با fault attack list: فهرست مکتوب و نگهداری‌شده‌ی خطاهایی که در پروژه‌ات واقعاً رخ داده‌اند. هر بار incident جدید، یک سطر اضافه می‌شود. این ارزشمندترین مستند تست یک تیم بالغ است.

تست اکتشافی (exploratory) برادر ساختاریافته‌ی آن است: در قالب session-based test management، یک تایم‌باکس ۶۰ تا ۹۰ دقیقه‌ای با یک charter مشخص («بررسی رفتار سبد خرید وقتی موجودی انبار حین پرداخت صفر می‌شود») و یادداشت‌برداری همزمان. خروجی: نقص‌ها + سؤالات + ایده‌های تست خودکار.

۴٫۷ تست مثبت و منفی

  • مثبت (happy path): با ورودی معتبر، نتیجه‌ی درست بگیر.
  • منفی (negative / unhappy path): با ورودی نامعتبر یا شرایط بد، خطای درست و کنترل‌شده بگیر.

نسبت سالم در یک سرویس بالغ معمولاً به نفع منفی است — چون happy path یکی است و راه‌های شکست ده‌ها. اگر در PR کسی فقط تست happy path دیدی، این دقیقاً همان‌جایی است که در code review باید بایستی.

قضاوت ارشد — کدام تکنیک را کِی به کار ببری:

نوع نیازمندی تکنیک اول تکنیک مکمل
ورودی عددی/بازه‌ای افراز هم‌ارزی BVA سه‌مقداری
قواعد کسب‌وکار چندشرطی جدول تصمیم افراز روی هر شرط
موجودیت دارای چرخه‌ی حیات انتقال حالت تست انتقال نامعتبر
ماتریس پیکربندی/سازگاری pairwise pin کردن ترکیب‌های حیاتی
الگوریتم پیچیده با شاخه‌ی زیاد جعبه‌سیاه + پوشش شاخه mutation testing
ناحیه‌ی تاریک و بی‌مستند تست اکتشافی fault attack list
یک نیازمندی به تو می‌دهم: «رمز عبور باید بین ۸ تا ۶۴ کاراکتر باشد و حداقل یک عدد داشته باشد.» موارد تست را بگو.

اول تکنیک را نام می‌برم تا مصاحبه‌گر بداند تصادفی جواب نمی‌دهم: افراز هم‌ارزی روی طول و روی محتوا، به‌علاوه‌ی BVA سه‌مقداری روی مرزهای ۸ و ۶۴.

طول: ۷ (رد)، ۸ (قبول)، ۹ (قبول)، ۶۳ (قبول)، ۶۴ (قبول)، ۶۵ (رد)، صفر/خالی (رد)، null (رد و نه NPE).

محتوا: بدون عدد (رد)، دقیقاً یک عدد (قبول)، فقط عدد (قبول یا رد؟ اینجا سؤال می‌پرسم — نیازمندی مبهم است و همین سؤال‌پرسیدن بخشی از جواب است).

مرزهای پنهان: فاصله‌ی ابتدا و انتها آیا trim می‌شود؟ اگر بشود، «۸ فاصله + یک عدد» چه می‌شود؟ کاراکتر یونیکد چندبایتی: ۶۴ emoji یعنی ۶۴ کاراکتر یا ۲۵۶ بایت؟ اگر ذخیره‌سازی محدودیت بایتی دارد، این باگ تولید است. ارقام غیرلاتین (۱۲۳ فارسی) عدد حساب می‌شوند؟ Character.isDigit می‌گوید بله — آیا این همان چیزی است که کسب‌وکار می‌خواهد؟

منفی/امنیتی: پیام خطا نباید بگوید کدام قانون نقض شده به شکلی که به enumeration کمک کند؛ و رمز نباید در لاگ بیفتد. آخر هم می‌گویم این نیازمندی خودش بو می‌دهد: راهنماهای مدرن (NIST SP 800-63B) روی طول تأکید می‌کنند و ترکیب اجباری کاراکتر را توصیه نمی‌کنند — این را می‌گویم تا نشان دهم فقط تست نمی‌کنم، نیازمندی را هم نقد می‌کنم.


بخش ۵ — مستنداتی که یک سازمان بالغ انتظار دارد

اینجا جایی است که مهندس‌های «فقط کدنویس» کم می‌آورند. لازم نیست عاشق مستندسازی باشی؛ لازم است بدانی هر سند چه سؤالی را جواب می‌دهد و کمترین نسخه‌ی مفیدش چیست.

۵٫۱ Test Plan

Test Plan سندی است که می‌گوید در این پروژه/رهاسازی، چه چیزی، چگونه، توسط چه کسی، در چه محیطی، تا کِی و با چه معیار توقفی تست می‌شود. استاندارد قدیمی IEEE 829 و استاندارد جاری ISO/IEC/IEEE 29119-3 قالب می‌دهند، ولی هیچ‌کس از تو انتظار ۴۰ صفحه ندارد. اسکلت مفید:

بخش سؤالی که جواب می‌دهد مثال یک‌خطی
Scope: in / out چه چیزی تست می‌شود و مهم‌تر، چه چیزی نمی‌شود «migration داده‌ی تاریخی خارج از دامنه است»
Test items دقیقاً کدام نسخه/کامپوننت «payment-service 3.4.0، gateway 2.1.x»
Approach سطوح، انواع، تکنیک‌ها، درصد اتوماسیون «unit+integration خودکار، UAT دستی»
Environment محیط، داده، سرویس‌های بدل «staging با sandbox شرکا، داده‌ی ناشناس‌شده»
Entry criteria کِی شروع می‌کنیم «build سبز + smoke pass + محیط آماده»
Exit criteria کِی تمام است پایین‌تر توضیح می‌دهم
Risks & mitigation چه چیزی می‌تواند برنامه را خراب کند «sandbox شریک ناپایدار → mock جایگزین»
Roles چه کسی چه چیزی را امضا می‌کند «PO امضای UAT»
Schedule & effort زمان و نفر-روز
Deliverables چه چیزی تحویل می‌شود «گزارش اجرا، لیست نقص باز، گزارش تست بار»
قضاوت ارشد — بخش «out of scope» گران‌بهاترین بخش سند است

هر دعوایی که بعد از رهاسازی سر «چرا این تست نشده بود؟» درمی‌گیرد، ریشه‌اش در نبود یک فهرست صریحِ «تست نمی‌شود» است. یک صفحه‌ی ساده که بگوید «در این رهاسازی، عملکرد زیر بار بالای ۲۰۰۰ TPS تست نشده و ریسکش پذیرفته شده است — تأییدکننده: X» تو را از هزار جلسه نجات می‌دهد. این مهارت، مهارت مهندسی نیست، مهارت مدیریت ریسک است و دقیقاً همان چیزی است که سِنیور را می‌سازد.

۵٫۲ Test Case خوب

یک test case خوب باید توسط کسی که پروژه را نمی‌شناسد قابل اجرا باشد و نتیجه‌اش قابل داوری. اجزا:

فیلد توضیح مثال
ID شناسه‌ی پایدار و قابل ارجاع TC-PAY-014
Title یک جمله‌ی رفتاری «برداشت بالای سقف روزانه به صف تأیید دستی می‌رود»
Requirement ref ردیابی به نیازمندی REQ-PAY-7
Priority P1..P3 بر اساس ریسک P1
Preconditions حالت لازم قبل از شروع «کاربر فعال، موجودی ۵۰ میلیون، MFA فعال»
Test data داده‌ی دقیق «مبلغ = 30,000,000، مقصد = IR...»
Steps گام‌های عددی و بدون ابهام ۱) ورود ۲) درخواست برداشت ۳) ارسال
Expected result نتیجه‌ی قابل مشاهده HTTP 202 + وضعیت PENDING_REVIEW + رویداد ReviewRequested
Postconditions تمیزکاری/حالت نهایی «درخواست در صف بازبینی باقی می‌ماند»
«انتظار: درست کار کند» یعنی هیچ

سه ضدالگوی کشنده در نوشتن test case: (۱) نتیجه‌ی مورد انتظار مبهم («صفحه به‌درستی نمایش داده شود») — نتیجه باید قابل رد یا تأیید بدون بحث باشد. (۲) وابستگی زنجیره‌ای بین test caseها، طوری که TC-02 بدون اجرای TC-01 معنی ندارد؛ این هم اجرای موازی را می‌کشد هم دیباگ را. (۳) گنجاندن جزئیات UI شکننده در گام‌ها («روی دکمه‌ی سوم از چپ کلیک کن»). گام‌ها را در سطح قصد بنویس، نه مختصات پیکسل.

۵٫۳ ردیابی از نیازمندی تا تست (Traceability)

ماتریس ردیابی نیازمندی (RTM) جدولی است که هر نیازمندی را به تست‌هایی که آن را پوشش می‌دهند وصل می‌کند. سه سؤال را جواب می‌دهد که هیچ ابزار دیگری جواب نمی‌دهد:

  1. پوشش: کدام نیازمندی هیچ تستی ندارد؟ (سوراخ واقعی)
  2. اثر تغییر: اگر REQ-PAY-7 عوض شود، کدام تست‌ها باید بازبینی شوند؟
  3. گزارش وضعیت: «۹۲٪ نیازمندی‌های P1 پوشش داده و سبز است» — این جمله برای مدیر معنا دارد؛ «۸۴٪ line coverage» ندارد.

نمودار: زنجیره‌ی ردیابی از نیاز کسب‌وکار تا نقص — The traceability chain from business need to defect.

flowchart LR
  BR[Business need] --> REQ[Requirement / user story]
  REQ --> AC[Acceptance criteria]
  AC --> TC[Test cases]
  TC --> RUN[Test runs]
  RUN --> DEF[Defects]
  DEF -.reopens.-> REQ

در عمل لازم نیست Excel نگه داری. اگر شناسه‌ی نیازمندی را در نام تست یا در annotation بگذاری، ماتریس را از کد استخراج می‌کنی:

@Test
@Tag("REQ-PAY-7")
@DisplayName("REQ-PAY-7 | withdrawal above daily limit goes to manual review")
void withdrawalAboveDailyLimitGoesToManualReview() { /* ... */ }

و اگر داده‌ی اجرای تست را در دیتابیس نگه می‌داری (بسیاری از ابزارهای مدیریت تست همین کار را می‌کنند)، پیدا کردن نیازمندی‌های بی‌پوشش یک کوئری است:

SELECT r.req_id,
       r.title,
       COUNT(tc.tc_id) AS test_cases,
       COUNT(*) FILTER (WHERE tr.status = 'PASSED') AS passed
FROM requirement r
LEFT JOIN test_case tc ON tc.req_id = r.req_id
LEFT JOIN LATERAL (
    SELECT status
    FROM test_run
    WHERE test_run.tc_id = tc.tc_id
    ORDER BY executed_at DESC
    FETCH FIRST 1 ROW ONLY
) tr ON TRUE
WHERE r.priority = 'P1'
GROUP BY r.req_id, r.title
HAVING COUNT(tc.tc_id) = 0
    OR COUNT(*) FILTER (WHERE tr.status = 'PASSED') < COUNT(tc.tc_id)
ORDER BY r.req_id;
تفاوت دیالکت

FILTER (WHERE ...) روی تجمیع، نحو استاندارد SQL است که PostgreSQL پیاده کرده ولی Oracle ندارد؛ در Oracle معادل قابل‌حمل آن COUNT(CASE WHEN ... THEN 1 END) است. همچنین LEFT JOIN LATERAL ... ON TRUE در PostgreSQL معادل OUTER APPLY در Oracle (از 12c به بعد) است. جزئیات بیشتر در فصل «تفاوت دیالکت‌های Oracle و PostgreSQL».

۵٫۴ گزارش نقص که توسعه‌دهنده می‌تواند رویش کار کند

یک گزارش نقص بد، یعنی یک رفت‌وبرگشت سه‌روزه. یک گزارش خوب این‌ها را دارد:

  • عنوان: «چه چیزی، کجا، تحت چه شرایطی» در یک خط. نه «سایت کار نمی‌کند».
  • محیط: نسخه‌ی build، مرورگر/کلاینت، محیط (staging/prod)، شناسه‌ی مستأجر، زمان دقیق با timezone.
  • گام‌های بازتولید: حداقلِ گام‌هایی که باگ را می‌سازد. اگر توانستی، کوچک‌سازی کن.
  • نتیجه‌ی واقعی در برابر مورد انتظار.
  • شواهد: لاگ با traceId، اسکرین‌شات، پاسخ HTTP خام، نه «عکس مانیتور با موبایل».
  • نرخ بازتولید: ۱۰ از ۱۰ یا ۲ از ۱۰؟ این عدد برای باگ‌های همزمانی حیاتی است.
  • Severity و Priority — که پایین می‌آید.
  • workaround اگر وجود دارد.
severity و priority یکی نیستند و این تله‌ی محبوب مصاحبه است

Severity (شدت) یک قضاوت فنی است: اثر نقص روی سیستم. Priority (اولویت) یک قضاوت کسب‌وکاری است: چقدر زود باید رفع شود. این چهار ترکیب واقعی‌اند:

  • شدت بالا، اولویت بالا: پرداخت در تولید کار نمی‌کند. ← الان.
  • شدت بالا، اولویت پایین: crash در قابلیتی که فقط در گزارش سالانه استفاده می‌شود و ده ماه دیگر اجرا می‌شود.
  • شدت پایین، اولویت بالا: نام شرکت در صفحه‌ی اول غلط املایی دارد، فردا رونمایی است. سیستم سالم است، آبرو نه.
  • شدت پایین، اولویت پایین: تراز یک آیکون در صفحه‌ی تنظیمات.

اگر در مصاحبه فقط بگویی «severity یعنی شدت و priority یعنی اولویت»، امتیاز نمی‌گیری. مثال چهارخانه بزن.

۵٫۵ معیار خروج (Exit Criteria) و Definition of Done

«تست تمام شد» یعنی چه؟ اگر جوابت «وقت تمام شد» است، آن معیار خروج نیست، معیار تسلیم است. معیار خروجِ قابل دفاع ترکیبی است از:

  • پوشش: ۱۰۰٪ نیازمندی‌های P1 و P2 حداقل یک تست اجراشده دارند.
  • اجرا: ≥۹۵٪ test caseهای برنامه‌ریزی‌شده اجرا شده؛ ۱۰۰٪ P1 اجرا و pass.
  • نقص: صفر نقص باز با شدت Critical/High؛ نقص‌های Medium با تأیید مکتوب مالک محصول موکول شده‌اند.
  • غیرکارکردی: p95 زیر SLO تحت بار هدف؛ نرخ خطا زیر آستانه؛ اسکن امنیتی بدون یافته‌ی Critical.
  • آمادگی عملیاتی: runbook، alert، داشبورد و مسیر rollback تست شده.

بخش ۶ — هرم تست، مخروط بستنی، و جای تست قرارداد

کنترل کیفیت در کارخانه

در یک کارخانه‌ی خوب، هر قطعه هنگام ساخت اندازه‌گیری می‌شود (ارزان و فوری)، بعد زیرمجموعه‌ها مونتاژ و آزمایش می‌شوند، و در آخر چند دستگاه کامل تست عملکرد می‌شوند. تصور کن کارخانه‌ای که هیچ قطعه‌ای را نمی‌سنجد و فقط محصول نهایی را روشن می‌کند: وقتی چراغ روشن نمی‌شود، باید کل دستگاه را باز کند تا بفهمد کدام مقاومت سوخته است. آن کارخانه، «مخروط بستنی» دارد.

نمودار: هرم سالم در برابر مخروط بستنی — A healthy pyramid versus the ice-cream cone anti-pattern.

flowchart TB
  subgraph Pyramid["Healthy pyramid"]
    P1["E2E: few, slow, high value"] --> P2["Integration / contract: some"]
    P2 --> P3["Unit: many, fast, cheap"]
  end
  subgraph Cone["Ice-cream cone"]
    C1["Manual testing: huge"] --> C2["E2E UI: many"]
    C2 --> C3["Integration: few"]
    C3 --> C4["Unit: almost none"]
  end

منطق هرم اقتصادی است، نه ایدئولوژیک: هرچه بالاتر می‌روی، هر تست کندتر، شکننده‌تر، گران‌تر برای نگهداری و مبهم‌تر در تشخیص علت می‌شود. پس تعداد کمتری از آن بساز — ولی حذفش نکن، چون فقط بالای هرم است که می‌گوید «کل سیستم واقعاً کار می‌کند».

لایه سرعت هدف چه چیزی را ثابت می‌کند هزینه‌ی نگهداری
Unit < 10ms منطق، شاخه‌ها، مرزها کم
Integration (با Testcontainers) 0.1–2s نگاشت، SQL، سریال‌سازی، تراکنش متوسط
Contract < 1s سازگاری provider/consumer کم تا متوسط
System / API E2E 1–30s جریان واقعی درون یک سرویس یا چند سرویس بالا
UI E2E 5–60s مسیر بحرانی کاربر خیلی بالا
Manual / exploratory دقیقه‌ها ناشناخته‌ها، UX، ریسک‌های جدید ثابت و تکرارشونده

نشانه‌های بالینی «مخروط بستنی» — اگر این‌ها را دیدی، شکل هرمت وارونه است: build بیش از ۲۰ دقیقه طول می‌کشد و کسی منتظرش نمی‌ماند؛ برای فهمیدن علت شکست باید ویدئوی تست UI را ببینی؛ تیم یک «هفته‌ی تثبیت» قبل از هر رهاسازی دارد؛ بخش زیادی از باگ‌ها را QA دستی پیدا می‌کند نه CI؛ و جمله‌ی «دوباره بزن شاید سبز شد» در تیم عادی شده است.

تست قرارداد (Contract Testing) — لایه‌ی گم‌شده

مشکل microservices این است: تست واحدِ هر سرویس سبز است، ولی وقتی کنار هم می‌گذاری‌شان می‌شکنند، چون consumer فرض می‌کرد فیلد amount رشته است و provider آن را عدد کرد. راه‌حل ساده‌لوحانه «E2E همه‌ی سرویس‌ها با هم» است — کُند، شکننده، و نیازمند بالا آوردن کل دنیا در CI.

Contract test می‌گوید: قرارداد بین دو سرویس را در یک فایل قابل‌اجرا ثبت کن؛ سپس هر طرف را جداگانه در برابر همان قرارداد تست کن.

  • Consumer-driven (Pact): مصرف‌کننده انتظارش را می‌نویسد، خروجی یک pact file است؛ تولیدکننده در CI خودش آن را verify می‌کند. پیش‌شرط عملی: یک broker مشترک برای تبادل قراردادها.
  • Provider-driven (Spring Cloud Contract): تولیدکننده قرارداد را در قالب DSL می‌نویسد، از آن هم تست provider تولید می‌شود و هم stub که مصرف‌کننده در تست خود استفاده می‌کند.
// src/test/resources/contracts/shouldReturnAccountBalance.groovy
Contract.make {
    description "should return balance for an existing account"
    request {
        method GET()
        url "/api/accounts/ACC-1001/balance"
        headers { accept(applicationJson()) }
    }
    response {
        status OK()
        headers { contentType(applicationJson()) }
        body(
            accountId: "ACC-1001",
            balance: 250000,
            currency: "IRR"
        )
        bodyMatchers {
            jsonPath('$.balance', byType())
            jsonPath('$.currency', byRegex('[A-Z]{3}'))
        }
    }
}
<plugin>
  <groupId>org.springframework.cloud</groupId>
  <artifactId>spring-cloud-contract-maven-plugin</artifactId>
  <version>4.3.0</version>
  <extensions>true</extensions>
  <configuration>
    <baseClassForTests>com.example.contract.ContractBase</baseClassForTests>
    <testFramework>JUNIT5</testFramework>
  </configuration>
</plugin>

قضاوت ارشد — contract test جایگزین E2E نیست، جایگزین ۹۰٪ آن است — با contract test می‌فهمی «شکل پیام سازگار است». هنوز نمی‌فهمی «جریان کسب‌وکار درست است» (مثلاً آیا سفارش بعد از پرداخت واقعاً ارسال می‌شود). پس چند E2E روی مسیرهای درآمدزا نگه دار و بقیه‌ی سازگاری‌ها را به contract بسپار. سؤال کلیدی که در مصاحبه امتیاز می‌آورد: «آیا provider می‌تواند بدون شکستن هیچ consumer‌ای deploy کند؟ اگر جواب را نمی‌دانی، contract test نداری.»

چطور بدون بالا آوردن کل سیستم مطمئن می‌شوی تغییر در یک سرویس، سرویس دیگر را نمی‌شکند؟

با تست قرارداد. هر مصرف‌کننده انتظارات واقعی‌اش را به‌صورت قرارداد اجراشدنی ثبت می‌کند؛ آن قرارداد در یک broker منتشر می‌شود؛ و pipeline تولیدکننده در هر build آن را در برابر پیاده‌سازی واقعی خودش verify می‌کند. اگر فیلدی حذف یا نوعش عوض شود، build تولیدکننده قرمز می‌شود — قبل از deploy، بدون این‌که هیچ سرویس دیگری بالا بیاید.

نکته‌ی مهم‌تر، قاعده‌ی تغییر سازگار است: افزودن فیلد اختیاری سازگار است؛ حذف فیلد، تغییر نوع، تنگ‌تر کردن قیود، یا تغییر معنای یک مقدار ناسازگار است. برای تغییر ناسازگار از نسخه‌بندی یا الگوی expand-and-contract استفاده می‌کنم: اول فیلد جدید را اضافه می‌کنم، مصرف‌کننده‌ها مهاجرت می‌کنند، بعد فیلد قدیمی حذف می‌شود.

و برای این‌که این چرخه در عمل کار کند، به «can-i-deploy» نیاز دارم: گیتِ استقرار که می‌پرسد آیا نسخه‌ای که می‌خواهم deploy کنم با نسخه‌های در حال اجرای همه‌ی مصرف‌کننده‌ها verify شده است یا نه.


بخش ۷ — استراتژی اتوماسیون

۷٫۱ چه چیزی را خودکار کنیم و چه چیزی را نه

اتوماسیون یک سرمایه‌گذاری است: هزینه‌ی ساخت + هزینه‌ی نگهداری در برابر صرفه‌جویی در اجراهای بعدی. اگر تستی سالی دو بار اجرا می‌شود و هر بار سه دقیقه دستی وقت می‌برد، خودکارکردنش زیان است.

خودکار کن اگر… خودکار نکن اگر…
مکرر اجرا می‌شود (هر PR، هر شب) یک‌بارمصرف است یا نیازمندی هنوز شناور
نتیجه‌ی قطعی و ماشین‌خوان دارد داوری انسانی لازم دارد (زیبایی، UX، متن)
ریسک بالا / مسیر درآمدزا است ناحیه‌ی کم‌ریسک و کم‌استفاده است
داده و محیطش کنترل‌پذیر است به سرویس بیرونی غیرقابل‌کنترل وابسته است
دستی‌اش خسته‌کننده و خطاپذیر است اکتشافی است و ارزشش در خلاقیت انسان است
ضدالگو: «همه‌چیز را خودکار می‌کنیم»

تیم‌هایی که هدفشان «۱۰۰٪ اتوماسیون» است، معمولاً به یک suite بزرگِ کند و flaky می‌رسند که کسی به آن اعتماد ندارد — و بدترین حالت ممکن همین است: هزینه‌ی نگهداری را می‌دهی و سیگنال نمی‌گیری. یک suite کوچکِ قابل‌اعتماد، از یک suite بزرگِ مشکوک بی‌نهایت ارزشمندتر است. معیار سلامت را این بگیر: «وقتی build قرمز می‌شود، چند درصد مواقع واقعاً باگ بوده؟» اگر زیر ۹۰٪ است، مشکل تو کمبود تست نیست، اضافه‌ی نویز است.

۷٫۲ مدیریت تست‌های flaky

تست flaky تستی است که بدون تغییر کد، گاهی سبز و گاهی قرمز می‌شود. خطر اصلی‌اش خودِ شکست نیست؛ این است که تیم را به بی‌اعتنایی به قرمز عادت می‌دهد و روزی که یک باگ واقعی قرمز می‌شود، کسی جدی نمی‌گیرد.

علت ریشه‌ای نشانه درمان
وابستگی به زمان واقعی شکست حوالی نیمه‌شب یا در ماشین کُند تزریق Clock و استفاده از Clock.fixed(...)
همزمانی و race با -parallel می‌شکند همگام‌سازی صریح، Awaitility به‌جای Thread.sleep
ترتیب اجرا / حالت مشترک تنها یا با seed دیگر پاس می‌شود ایزوله‌سازی داده، ریست حالت static
انتظار ثابت روی UI/شبکه در CI شلوغ می‌شکند انتظار مبتنی بر شرط (explicit wait)
داده‌ی مشترک بین تست‌ها با اجرای موازی می‌شکند داده‌ی یکتا به‌ازای تست، schema جدا
منابع خارجی ناپایدار شکست‌های پراکنده و بی‌الگو test double در سطح مرز

فرایندی که کار می‌کند:

  1. اندازه‌گیری: نرخ flakiness را از تاریخچه‌ی CI استخراج کن (همان تست، همان commit، نتیجه‌ی متفاوت).
  2. قرنطینه‌ی زمان‌دار: تست flaky را از gate اصلی خارج کن ولی در job جداگانه اجرا کن، با مالک و مهلت. قرنطینه‌ی بی‌مهلت یعنی حذف.
  3. رفع ریشه‌ای، نه retry.
  4. بودجه: مثلاً «حداکثر ۵ تست در قرنطینه»؛ اگر پر شد، کار جدید متوقف تا پاک‌سازی.

retryFailedTests مسکّن است نه درمان — افزودن retry خودکار به تست‌ها وسوسه‌انگیز است و در کوتاه‌مدت build را سبز می‌کند. ولی دو ضرر دارد: باگ‌های واقعیِ نادر (مثل race در کد تولید، نه در تست) را پنهان می‌کند، و زمان اجرای بدترین‌حالت را چند برابر می‌کند. اگر مجبور به retry شدی، حداقل آن را گزارش کن: تستی که با retry پاس شده باید در گزارش با پرچم بیاید و روی داشبورد flakiness ثبت شود، وگرنه اطلاعات را دور ریخته‌ای.

۷٫۳ مدیریت داده‌ی تست

سه راهبرد، با ترتیبِ ترجیح:

  1. ساخت در خود تست (test data builder) — بهترین حالت. هر تست داده‌ی خودش را می‌سازد، پس مستقل و خواناست.
  2. fixture مشترک و کوچک — برای داده‌ی مرجع تغییرناپذیر (فهرست ارز، کد بانک‌ها).
  3. کپی از تولید (ناشناس‌شده) — فقط برای تست بار و مهاجرت داده. با ریسک انطباق (GDPR و معادل‌های محلی) و هزینه‌ی نگهداری بالا.

الگوی builder با مقادیر پیش‌فرض معقول، خوانایی تست را متحول می‌کند: هر تست فقط چیزی را می‌گوید که برایش مهم است.

public final class AccountBuilder {
    private String id = "ACC-" + UUID.randomUUID();
    private long balance = 1_000_000L;
    private boolean active = true;
    private boolean mfaEnabled = true;

    public static AccountBuilder anAccount() { return new AccountBuilder(); }

    public AccountBuilder withBalance(long balance) { this.balance = balance; return this; }
    public AccountBuilder inactive() { this.active = false; return this; }
    public AccountBuilder withoutMfa() { this.mfaEnabled = false; return this; }

    public Account build() { return new Account(id, balance, active, mfaEnabled); }
}

// خواناییِ حاصل: فقط تفاوتِ معنادار دیده می‌شود
var poorAccount = anAccount().withBalance(0).build();

برای ریست دیتابیس بین تست‌های یکپارچه‌سازی، TRUNCATE معمولاً از حذف رکوردبه‌رکورد سریع‌تر است — ولی نحو بازنشانی sequence بین دو دیالکت فرق می‌کند:

-- خالی کردن جداول و صفر کردن identity در یک دستور
TRUNCATE TABLE payment, account RESTART IDENTITY CASCADE;

داده‌ی تولید را «ناشناس» نکن، «مصنوعی» کن — حذف نام و شماره تلفن کافی نیست: ترکیب تاریخ تولد + کد پستی + تراکنش‌های اخیر معمولاً برای شناسایی مجدد کافی است. اگر مجبوری از داده‌ی واقعی استفاده کنی، از تکنیک‌های واقعی pseudonymisation (نگاشت پایدار با کلید جدا) استفاده کن، محیط را در همان سطح محرمانگی تولید نگه دار و دسترسی را لاگ کن. بهترین حالت این است که یک مولد داده‌ی مصنوعی بنویسی که توزیع آماری تولید را تقلید کند بدون این‌که یک رکورد واقعی داشته باشد.

۷٫۴ استراتژی محیط و بدل‌های تست

واژگان دقیق test double (طبقه‌بندی رایج) — در مصاحبه به کار می‌آید:

نوع تعریف مثال
Dummy فقط برای پر کردن پارامتر، هرگز استفاده نمی‌شود null-object
Stub پاسخ ازپیش‌تعیین‌شده می‌دهد when(repo.find(id)).thenReturn(acc)
Spy واقعی است ولی تعامل‌ها را ثبت می‌کند شمارش تعداد فراخوانی
Mock انتظارِ رفتاری دارد و آن را verify می‌کند verify(gateway).charge(...)
Fake پیاده‌سازی واقعی ولی ساده‌شده repository درون‌حافظه‌ای

و برای محیط:

  • hermetic / ephemeral: هر build محیط خودش را می‌سازد و نابود می‌کند (Testcontainers، namespace موقت Kubernetes). گران‌تر برای راه‌اندازی، ولی بدون تداخل و بدون drift.
  • shared staging: ارزان، ولی سه تیم همزمان رویش کار می‌کنند و شکست‌ها به هم آلوده می‌شوند.
  • سرویس شریک بیرونی: اگر sandbox پایدار دارد، از آن استفاده کن اما در gate اصلی نه؛ مسیر اصلی را با یک بدل قراردادمحور (WireMock/mock server) بزن و یک job شبانه در برابر sandbox واقعی اجرا کن تا drift را کشف کنی.

۷٫۵ اجرای تست در CI

جزئیات pipeline در فصل «CI/CD» است؛ اینجا فقط لایه‌بندیِ مربوط به تست:

<!-- surefire: تست‌های سریع واحد در فاز test -->
<plugin>
  <groupId>org.apache.maven.plugins</groupId>
  <artifactId>maven-surefire-plugin</artifactId>
  <configuration>
    <excludedGroups>integration,slow</excludedGroups>
    <parallel>classes</parallel>
    <threadCount>4</threadCount>
  </configuration>
</plugin>

<!-- failsafe: تست‌های یکپارچه‌سازی در فاز integration-test -->
<plugin>
  <groupId>org.apache.maven.plugins</groupId>
  <artifactId>maven-failsafe-plugin</artifactId>
  <configuration>
    <groups>integration</groups>
  </configuration>
  <executions>
    <execution>
      <goals>
        <goal>integration-test</goal>
        <goal>verify</goal>
      </goals>
    </execution>
  </executions>
</plugin>
# حلقه‌ی سریع توسعه‌دهنده: فقط unit
mvn -q test

# گیت ادغام: unit + integration، شکست در verify
mvn -q verify

# اجرای موازی JUnit 5 (فایل src/test/resources/junit-platform.properties)
# junit.jupiter.execution.parallel.enabled = true
# junit.jupiter.execution.parallel.mode.default = concurrent
قضاوت ارشد — لایه‌بندی بر اساس بودجه‌ی زمانی، نه بر اساس اسم

به‌جای بحث فلسفی «این unit است یا integration»، یک بودجه‌ی زمانی تعریف کن: مرحله‌ی pre-commit زیر ۹۰ ثانیه، گیت PR زیر ۱۰ دقیقه، pipeline شبانه بدون سقف. هر تست در سریع‌ترین مرحله‌ای می‌نشیند که از بودجه عبور نکند. این قاعده هم دعوا را تمام می‌کند هم رفتار درست را تشویق می‌کند: کسی که تست کندی می‌نویسد، خودش انگیزه‌ی سریع‌کردنش را پیدا می‌کند.

یک جریان رویدادمحور را چطور تست می‌کنی؟ سرویس یک پیام مصرف می‌کند، دیتابیس را به‌روز می‌کند و یک رویداد دیگر منتشر می‌کند.

آن را به سه لایه می‌شکنم تا هر لایه ارزان‌ترین تست ممکن را داشته باشد.

لایه‌ی منطق: تابع «رویداد ورودی + وضعیت فعلی → وضعیت جدید + رویدادهای خروجی» را خالص و بدون I/O نگه می‌دارم و با تست واحد سریع پوشش می‌دهم؛ اینجا افراز هم‌ارزی و انتقال حالت را به کار می‌برم.

لایه‌ی adapter: با یک broker واقعی در Testcontainers تست می‌کنم که deserialize، commit آفست، و انتشار رویداد خروجی درست کار می‌کند. اینجا مهم‌ترین چیز، تست رفتار در برابر خطا است: پیام بدشکل باید به DLQ برود نه این‌که مصرف‌کننده را در حلقه‌ی بی‌پایان بیندازد.

لایه‌ی قرارداد: schema پیام را با contract test یا registry قفل می‌کنم تا تغییر ناسازگار قبل از deploy گرفته شود.

سه نکته‌ی اختصاصی این دنیا که سِنیور را نشان می‌دهد: (۱) هرگز Thread.sleep ننویس — از انتظار شرطی (مثلاً Awaitility با timeout و poll interval) استفاده کن، وگرنه یا تست flaky می‌شود یا بی‌خودی کند. (۲) idempotency را عمداً تست کن: همان پیام را دو بار بفرست و assert کن که اثر جانبی یک بار اعمال شده؛ چون در تحویل at-least-once، تکرار قطعی است نه استثنا. (۳) ترتیب و کلید پارتیشن را تست کن: دو رویداد مرتبط با کلید یکسان باید ترتیبشان حفظ شود؛ این جایی است که باگ‌های تولید متولد می‌شوند. جزئیات پروتکل در فصل‌های Kafka و RabbitMQ آمده؛ اینجا فقط دیسیپلین تستش را می‌گویم.

۷٫۶ پوشش کد: سیگنال، نه هدف

coverage فقط می‌گوید کدام خط/شاخه اجرا شده — نه این‌که بررسی شده. تستی که هیچ assert ندارد هم پوشش تولید می‌کند.

انواع مفید:

  • line/statement coverage — ضعیف‌ترین.
  • branch coverage — هر شاخه‌ی if هر دو طرفش رفته باشد. حداقلِ معنادار.
  • MC/DC — برای هر شرط اتمی در یک شرط مرکب، نشان بده که به‌تنهایی نتیجه را عوض می‌کند. در حوزه‌های ایمنی‌بحرانی الزامی است.
<plugin>
  <groupId>org.jacoco</groupId>
  <artifactId>jacoco-maven-plugin</artifactId>
  <version>0.8.13</version>
  <executions>
    <execution><goals><goal>prepare-agent</goal></goals></execution>
    <execution>
      <id>check</id>
      <phase>verify</phase>
      <goals><goal>check</goal></goals>
      <configuration>
        <rules>
          <rule>
            <element>BUNDLE</element>
            <limits>
              <limit>
                <counter>BRANCH</counter>
                <value>COVEREDRATIO</value>
                <minimum>0.70</minimum>
              </limit>
            </limits>
          </rule>
        </rules>
      </configuration>
    </execution>
  </executions>
</plugin>
وقتی coverage هدف شود، معیار بودنش را از دست می‌دهد

اگر مدیریت بگوید «۸۰٪ اجباری»، تیم در یک بعدازظهر به آن می‌رسد — با تست‌هایی که همه‌چیز را صدا می‌زنند و هیچ‌چیز را assert نمی‌کنند، یا با @Generated روی کلاس‌های پر از منطق. عدد بالا می‌رود، کیفیت نه. استفاده‌ی درست از coverage دو چیز است: (۱) دیف پوشش روی کد جدید PR، (۲) پیدا کردن نواحی صفر که کسی نمی‌دانست تست ندارند. عدد کل پروژه، تقریباً بی‌معناست.

۷٫۷ Mutation testing — تستِ تست‌ها

اگر coverage می‌گوید «این خط اجرا شد»، mutation testing می‌پرسد: «اگر این خط را عمداً خراب کنم، آیا هیچ تستی قرمز می‌شود؟»

روش کار: ابزار در بایت‌کد جهش (mutant) ایجاد می‌کند — < را به <= تبدیل می‌کند، return true را به return false، فراخوانی یک متد را حذف می‌کند — و تست‌ها را اجرا می‌کند. اگر تستی شکست، جهش کشته شد (killed). اگر همه سبز ماندند، جهش زنده ماند (survived) و یعنی آن رفتار عملاً بی‌محافظ است.

mutation score = جهش‌های کشته‌شده ÷ کل جهش‌ها. این عدد بسیار صادق‌تر از coverage است.

<plugin>
  <groupId>org.pitest</groupId>
  <artifactId>pitest-maven</artifactId>
  <version>1.19.1</version>
  <configuration>
    <targetClasses><param>com.example.payment.domain.*</param></targetClasses>
    <targetTests><param>com.example.payment.domain.*Test</param></targetTests>
    <mutationThreshold>75</mutationThreshold>
    <timestampedReports>false</timestampedReports>
  </configuration>
  <dependencies>
    <dependency>
      <groupId>org.pitest</groupId>
      <artifactId>pitest-junit5-plugin</artifactId>
      <version>1.2.2</version>
    </dependency>
  </dependencies>
</plugin>
# اجرای کامل روی ماژول
mvn -q test-compile org.pitest:pitest-maven:mutationCoverage

# حالت افزایشی: فقط کدی که نسبت به شاخه‌ی مبنا تغییر کرده (مناسب PR)
mvn -q test-compile org.pitest:pitest-maven:scmMutationCoverage \
    -Dinclude=ADDED,MODIFIED
# گزارش HTML: target/pit-reports/index.html

قضاوت ارشد — mutation testing را روی کل پروژه اجرا نکن — PIT کُند است چون برای هر جهش تست‌ها را اجرا می‌کند. راهبرد عملی: آن را فقط روی هسته‌ی دامنه (منطق پول، قیمت‌گذاری، مجوز، محاسبات) و فقط روی کد تغییریافته در PR اجرا کن. حتی اگر هیچ آستانه‌ای اجباری نکنی، خواندن یک بار گزارش PIT روی هسته‌ی دامنه‌ات تجربه‌ی تکان‌دهنده‌ای است: معمولاً چند جهش زنده پیدا می‌شود که دقیقاً همان جاهایی است که در تولید ترسناک‌اند.

چرا ۹۰٪ code coverage تضمین کیفیت نیست؟

چون coverage یک معیار اجرا است، نه یک معیار راستی‌آزمایی. من می‌توانم فردا با نوشتن تست‌هایی که هر متد را صدا می‌زنند و هیچ assert ندارند، پوشش را به ۱۰۰٪ برسانم بدون این‌که یک باگ هم گرفته شود. حتی با assert هم، پوشش خط چیزی درباره‌ی مقادیر مرزی، ترکیب شرط‌ها یا نیازمندی‌های پیاده‌نشده نمی‌گوید — کدی که اصلاً نوشته نشده، صفر خط دارد و صفر خطِ پوشش‌نداده.

آنچه من نگاه می‌کنم: پوشش شاخه روی کد جدید در PR (نه عدد کل پروژه)، نواحی با پوشش صفر در هسته‌ی دامنه، و مهم‌تر از همه mutation score روی منطق حیاتی. mutation عملاً می‌پرسد «اگر کد را خراب کنم، آیا کسی می‌فهمد؟» و این همان سؤالی است که ما واقعاً می‌خواهیم جوابش را بدانیم.

در عمل من پوشش را به‌عنوان گیت نرم می‌گذارم — کاهش پوشش در PR باعث هشدار و بحث می‌شود نه شکست خودکار — و در عوض روی مرور کیفیت assertها در code review وقت می‌گذارم.

یک تست flaky در CI داری که تیم را کلافه کرده. چطور برخورد می‌کنی؟

اول اندازه‌گیری: از تاریخچه‌ی CI بیرون می‌کشم که چند درصد اجراها و روی چه commitهایی شکسته؛ اگر روی commit یکسان هم سبز و هم قرمز شده، قطعاً flaky است و نه رگرسیون.

بعد مهار سریع: تست را قرنطینه می‌کنم — از gate اصلی خارج، اما در یک job جداگانه همچنان اجرا می‌شود تا سیگنالش را از دست ندهیم. برایش issue با مالک و مهلت می‌سازم. اجازه نمی‌دهم قرنطینه به قبرستان تبدیل شود.

بعد ریشه‌یابی: با تکرار زیاد و تحت بار اجرایش می‌کنم، ترتیب اجرا را عوض می‌کنم، seed را تغییر می‌دهم. تقریباً همیشه یکی از این‌هاست: زمان (Instant.now() به‌جای Clock تزریق‌شده)، همزمانی (Thread.sleep به‌جای انتظار شرطی)، حالت مشترک بین تست‌ها، یا وابستگی به سرویس خارجی.

نکته‌ای که معمولاً امتیاز می‌گیرد: گاهی تست flaky در حال کشف یک باگ واقعی همزمانی در کد تولید است. پس قبل از این‌که تست را «تعمیر» کنم، مطمئن می‌شوم که مشکل در تست است نه در سیستم. و در سطح فرهنگی، «قرمز یعنی توقف» را زنده نگه می‌دارم؛ تیمی که به قرمز عادت کند، عملاً CI ندارد.


بخش ۸ — تست غیرکارکردی: کارایی، بار، فشار، دوام

پل و کامیون

یک پل «کار می‌کند» یعنی چه؟ اگر یک ماشین از رویش رد شود، کار می‌کند. اگر ۵۰۰ ماشین همزمان روی آن باشند چطور؟ اگر یک کامیون ۸۰ تنی رد شود؟ اگر ده سال بدون تعمیر زیر ترافیک روزانه بماند؟ اگر یک زلزله بیاید؟ هر کدام از این‌ها یک نوع تست است و جواب «کار می‌کند» برای همه‌شان یکی نیست.

۸٫۱ واژگان دقیق — این‌ها را تیم‌ها اشتباه به کار می‌برند

نوع تست سؤال شکل بار خروجی مورد انتظار
Performance (چتر کلی) سیستم چقدر سریع و چقدر ظرفیت دارد؟ متغیر مشخصه‌سازی رفتار
Load آیا زیر بار مورد انتظار SLO را نگه می‌دارد؟ بار هدف، ثابت p95/p99 و نرخ خطا در آستانه
Stress نقطه‌ی شکست کجاست و چطور می‌شکند؟ افزایش تا فروپاشی تخریب آبرومندانه، نه فروپاشی کامل
Soak / Endurance آیا در طول زمان تخریب می‌شود؟ بار متوسط، ۴–۷۲ ساعت بدون نشت حافظه/connection/دیسک
Spike واکنش به جهش ناگهانی؟ جهش سریع و بازگشت جذب یا throttle، و بازیابی
Capacity / Breakpoint حداکثر توان با SLO سالم؟ پله‌ای تا نقض SLO عدد ظرفیت برای ظرفیت‌سنجی
Scalability با دو برابر منبع چه می‌شود؟ بار ثابت، منابع متغیر ضریب مقیاس‌پذیری واقعی
Volume با داده‌ی بزرگ چه می‌شود؟ داده‌ی انبوه کوئری‌ها همچنان از ایندکس استفاده کنند

جهش، بازیابی را تست می‌کند نه فقط تحمل را — اکثر تیم‌ها در تست spike فقط نگاه می‌کنند که «آیا سیستم دوام آورد؟». سؤال مهم‌تر این است: بعد از عبور جهش، چقدر طول کشید تا به حالت عادی برگردد؟ سیستم‌هایی که صف‌های بی‌کران دارند ممکن است در لحظه‌ی جهش خطا ندهند ولی نیم‌ساعت بعد هنوز در حال پردازش درخواست‌های منقضی‌شده باشند. الگوی درست، رد کردن سریع (load shedding) و صف کران‌دار است — که در فصل resilience آمده.

۸٫۲ اول SLO، بعد تست

تست بار بدون معیار قبولی، فقط تولید نمودار رنگی است. پس قبل از هر چیز:

  • SLI (شاخص): چیزی که اندازه می‌گیری. مثال: «درصد درخواست‌های POST /payments که در کمتر از ۸۰۰ms پاسخ ۲xx/4xx می‌گیرند».
  • SLO (هدف): مقدار هدف روی آن شاخص در یک بازه. مثال: «۹۹٫۵٪ در هر ۳۰ روز».
  • SLA (توافق): تعهد قراردادی با جریمه. معمولاً شل‌تر از SLO داخلی.
  • error budget: ۱ منهای SLO. با ۹۹٫۵٪، بودجه‌ی خطای ماهانه حدود ۳ ساعت و ۳۶ دقیقه است.
میانگین دروغ می‌گوید؛ صدک‌ها را بخوان

اگر ۹۹ درخواست در ۱۰ms و یکی در ۵ ثانیه پاسخ بگیرد، میانگین ۶۰ms است و «عالی» به نظر می‌رسد — در حالی که یک کاربر واقعی ۵ ثانیه منتظر مانده. همیشه p50/p90/p95/p99 و max را گزارش کن. و دو نکته‌ی ظریف: (۱) صدک‌ها جمع‌پذیر نیستند؛ میانگین‌گرفتن از p95 چند نمونه‌ی زمانی، از نظر ریاضی بی‌معناست — باید از هیستوگرام‌های خام دوباره محاسبه کنی. (۲) در یک صفحه که ۲۰ فراخوانی سرویس دارد، احتمال این‌که کاربر حداقل یک بار به دُم p99 بخورد بسیار زیاد است؛ به همین دلیل p99 «موردی نادر» نیست، تجربه‌ی روزمره‌ی بخشی از کاربران است.

۸٫۳ مدل بار: باز در برابر بسته

این مفهوم، مرز بین یک تست بار اسباب‌بازی و یک تست بار معتبر است.

  • مدل بسته (closed model): تعداد ثابتی کاربر مجازی که هر کدام «درخواست بزن، پاسخ بگیر، کمی فکر کن، دوباره بزن» می‌کنند. اگر سیستم کند شود، نرخ ورود خودبه‌خود کم می‌شود. این مدل، سیستم‌های با تعداد کاربر محدود (مثل یک اپلیکیشن داخلی با ۲۰۰ کارمند) را خوب مدل می‌کند.
  • مدل باز (open model): کاربرها با یک نرخ ورود مستقل می‌آیند، صرف‌نظر از این‌که سیستم چقدر کند شده. این مدل، ترافیک اینترنتی واقعی را مدل می‌کند — دنیا منتظر تو نمی‌ماند.
هماهنگی حذف‌شده (Coordinated Omission)

این بزرگ‌ترین تله‌ی آماری تست بار است. در مدل بسته، وقتی سیستم گیر می‌کند، کاربر مجازی هم گیر می‌کند و درخواست‌های بعدی را اصلاً نمی‌فرستد. نتیجه: دقیقاً همان درخواست‌هایی که باید کُند ثبت می‌شدند، هرگز اندازه‌گیری نمی‌شوند و آمار تو خوش‌بینانه‌ی تخیلی می‌شود. راه‌حل: از executorهای مبتنی بر نرخ ورود استفاده کن — در k6 یعنی constant-arrival-rate یا ramping-arrival-rate، و در Gatling یعنی injectOpen(constantUsersPerSec(...)). اگر تست بار تو با «۵۰۰ کاربر مجازی» تعریف شده و هیچ نرخی در آن نیست، احتمالاً قربانی این تله شده‌ای.

قانون لیتل (Little's Law) ابزار ذهنی تو برای تبدیل این‌هاست:

L = λ × W

یعنی «تعداد درخواست‌های همزمان در سیستم = نرخ ورود × میانگین زمان اقامت». اگر ۲۰۰ درخواست بر ثانیه داری و میانگین پاسخ ۲۵۰ms است، به‌طور متوسط ۵۰ درخواست همزمان در حال پردازش‌اند — که مستقیماً به اندازه‌ی connection pool و thread pool تو ربط دارد. اگر pool تو ۲۰ است، صف تشکیل می‌شود و p99 منفجر می‌شود، حتی وقتی CPU بی‌کار است.

۸٫۴ مدل کردن بار واقعی

بار ساختگی («همه فقط GET / می‌زنند») نتیجه‌ی بی‌ارزش می‌دهد. مدل بار را از داده‌ی واقعی بساز: از لاگ دسترسی یا جدول رویدادها، توزیع endpointها و ساعت اوج را دربیاور.

SELECT endpoint,
       COUNT(*) AS hits,
       ROUND(100.0 * COUNT(*) / SUM(COUNT(*)) OVER (), 2) AS pct,
       ROUND(COUNT(*) / 3600.0, 2) AS rps_in_peak_hour
FROM access_log
WHERE ts >= TIMESTAMP '2026-08-05 12:00:00'
  AND ts <  TIMESTAMP '2026-08-05 13:00:00'
GROUP BY endpoint
ORDER BY hits DESC
FETCH FIRST 15 ROWS ONLY;

از خروجی، یک پروفایل بار بساز: مثلاً ۶۰٪ جست‌وجو، ۲۵٪ مشاهده‌ی جزئیات، ۱۰٪ افزودن به سبد، ۵٪ پرداخت — و همان نسبت‌ها را در سناریوی تست پیاده کن. به‌علاوه‌ی think time واقعی بین گام‌ها، چون بدون آن الگوی همزمانی‌ات کاملاً غیرواقعی می‌شود.

نمودار: توپولوژی یک تست بار درست، با نقاطی که باید همزمان مانیتور شوند — Load-test topology and the signals to watch at each hop.

flowchart LR
  LG[Load generators] --> LB[Load balancer]
  LB --> APP[App instances]
  APP --> DB[(Database)]
  APP --> CACHE[(Cache)]
  APP --> EXT[External / stubbed services]
  MON[Metrics: RED + USE] -.scrapes.- APP
  MON -.scrapes.- DB
  MON -.scrapes.- LG

۸٫۵ سه ابزار: JMeter، Gatling، k6

معیار Apache JMeter Gatling Grafana k6
زبان تست GUI + XML (.jmx) Java/Kotlin/Scala DSL JavaScript (ES modules)
مدل اجرا thread به‌ازای کاربر مجازی مبتنی بر Akka/غیرمسدودکننده Go + موتور JS (goja)
مصرف منابع به‌ازای VU بالا کم کم
بازبینی در Git دشوار (XML حجیم) عالی عالی
پروتکل‌ها بسیار وسیع (JDBC، JMS، LDAP، FTP…) HTTP، WebSocket، JMS، gRPC HTTP، WebSocket، gRPC، browser
نقطه‌ی قوت اکوسیستم افزونه، تیم‌های غیربرنامه‌نویس گزارش عالی، DSL تایپ‌دار برای تیم جاوا ادغام DevOps، thresholds به‌عنوان گیت CI
نقطه‌ی ضعف حافظه در VU بالا، دیف سخت جامعه‌ی کوچک‌تر از JMeter نیاز به JS، برخی قابلیت‌ها ابری
قضاوت ارشد — انتخاب ابزار را به مهارت تیم گره بزن

هر سه ابزار می‌توانند یک API را زیر بار ببرند. عامل تعیین‌کننده این است که چه کسی تست را نگه می‌دارد: اگر تیم جاوایی است و تست بار باید کنار کد زندگی کند، Gatling با Java DSL انتخاب طبیعی است. اگر تیم DevOps است و می‌خواهی تست بار یک مرحله‌ی معمولی در pipeline باشد با معیار قبولی خودکار، k6 با thresholds کمترین اصطکاک را دارد. اگر باید JDBC یا JMS یا پروتکل‌های قدیمی سازمانی را بار بزنی، JMeter هنوز بی‌رقیب است. جواب «کدام بهتر است؟» در مصاحبه هرگز نام یک ابزار نیست؛ معیار انتخاب است.

k6 — کمترین اصطکاک برای CI

k6 یک باینری Go است؛ تست را به JavaScript می‌نویسی و thresholds مستقیماً کد خروج فرایند را تعیین می‌کند — یعنی گیت CI مجانی است.

// payments-load.js
import http from 'k6/http';
import { check, sleep } from 'k6';
import { Trend, Rate } from 'k6/metrics';

const paymentLatency = new Trend('payment_latency', true);
const businessErrors = new Rate('business_errors');

export const options = {
  scenarios: {
    // مدل باز: نرخ ورود مستقل از کندی سیستم
    steady_traffic: {
      executor: 'ramping-arrival-rate',
      startRate: 20,
      timeUnit: '1s',
      preAllocatedVUs: 100,
      maxVUs: 500,
      stages: [
        { target: 20,  duration: '1m' },   // warm-up
        { target: 200, duration: '3m' },   // ramp to target
        { target: 200, duration: '10m' },  // steady state
        { target: 0,   duration: '1m' },   // ramp down
      ],
    },
  },
  thresholds: {
    'http_req_failed': ['rate<0.01'],
    'http_req_duration{name:createPayment}': ['p(95)<800', 'p(99)<2000'],
    'business_errors': ['rate<0.005'],
    'checks': ['rate>0.99'],
  },
};

export default function () {
  const payload = JSON.stringify({ accountId: 'ACC-1001', amount: 250000, currency: 'IRR' });
  const res = http.post('https://staging.example.internal/api/payments', payload, {
    headers: { 'Content-Type': 'application/json' },
    tags: { name: 'createPayment' },   // برچسب پایدار برای گروه‌بندی متریک
  });

  const ok = check(res, {
    'status is 201': (r) => r.status === 201,
    'has payment id': (r) => r.json('paymentId') !== undefined,
  });

  paymentLatency.add(res.timings.duration);
  businessErrors.add(!ok);
  sleep(Math.random() * 2 + 1);   // think time واقع‌گرایانه
}
# اجرای محلی
k6 run payments-load.js

# بازنویسی گزینه‌ها از خط فرمان (برای smoke سریع)
k6 run --vus 5 --duration 30s payments-load.js

# خروجی به فایل JSON برای تحلیل بعدی
k6 run --out json=results.json payments-load.js

# ساخت اسکلت یک اسکریپت جدید
k6 new payments-load.js

چرا tags مهم است — اگر URL شامل شناسه‌ی متغیر باشد (/api/payments/9f3a...)، k6 هر کدام را یک نام جدا می‌شمارد و متریک‌هایت خرد می‌شود. با tags: { name: 'createPayment' } همه را زیر یک نام پایدار جمع می‌کنی و می‌توانی threshold روی همان نام بگذاری — دقیقاً مثل http_req_duration{name:createPayment} بالا.

Gatling — وقتی تیم جاوایی است

Gatling از نسخه‌ی ۳٫۷ یک Java DSL کامل دارد؛ لازم نیست Scala بلد باشی. تست کنار کد پروژه در src/test/java زندگی می‌کند و با maven plugin اجرا می‌شود.

package simulations;

import io.gatling.javaapi.core.*;
import io.gatling.javaapi.http.*;

import java.time.Duration;

import static io.gatling.javaapi.core.CoreDsl.*;
import static io.gatling.javaapi.http.HttpDsl.*;

public class PaymentSimulation extends Simulation {

    HttpProtocolBuilder httpProtocol = http
        .baseUrl("https://staging.example.internal")
        .acceptHeader("application/json")
        .contentTypeHeader("application/json")
        .shareConnections();

    FeederBuilder<String> accounts = csv("accounts.csv").random();

    ScenarioBuilder createPayment = scenario("Create payment")
        .feed(accounts)
        .exec(
            http("createPayment")
                .post("/api/payments")
                .body(StringBody("""
                    {"accountId":"#{accountId}","amount":250000,"currency":"IRR"}
                    """))
                .check(status().is(201))
                .check(jsonPath("$.paymentId").saveAs("paymentId"))
        )
        .pause(Duration.ofSeconds(1), Duration.ofSeconds(3))   // think time
        .exec(
            http("getPayment")
                .get("/api/payments/#{paymentId}")
                .check(status().is(200))
        );

    {
        setUp(
            createPayment.injectOpen(          // مدل باز: نرخ ورود
                nothingFor(Duration.ofSeconds(5)),
                rampUsersPerSec(10).to(200).during(Duration.ofMinutes(3)),
                constantUsersPerSec(200).during(Duration.ofMinutes(10))
            )
        )
        .protocols(httpProtocol)
        .assertions(
            global().responseTime().percentile3().lt(800),   // p95 پیش‌فرض
            global().failedRequests().percent().lt(1.0),
            details("createPayment").responseTime().percentile(99.0).lt(2000)
        );
    }
}
# اجرا با maven plugin (io.gatling:gatling-maven-plugin)
mvn gatling:test -Dgatling.simulationClass=simulations.PaymentSimulation

# گزارش HTML در target/gatling/<simulation>-<timestamp>/index.html

مدل باز در برابر بسته در GatlinginjectOpen(...) مدل باز است: با rampUsersPerSec/constantUsersPerSec نرخ ورود را کنترل می‌کنی. injectClosed(...) مدل بسته است: با constantConcurrentUsers/rampConcurrentUsers تعداد کاربران همزمان را ثابت نگه می‌داری. برای APIهای عمومی تقریباً همیشه open درست است؛ closed را برای سیستم‌های با جمعیت کاربری محدود یا شبیه‌سازی یک استخر ثابت کارگر نگه دار.

JMeter — پوشش پروتکلی و اجرای بدون GUI

JMeter نسخه‌ی پایدار ۵٫۶٫۳ است و به Java 8+ نیاز دارد (Java 17+ توصیه می‌شود). قاعده‌ی طلایی: GUI فقط برای ساخت و دیباگ تست است؛ اجرای واقعی همیشه بدون GUI.

# اجرای بدون GUI + تولید داشبورد HTML
jmeter -n -t payments.jmx -l results.jtl -e -o report/

# بازنویسی پارامترها از خط فرمان (در .jmx با ${__P(threads,50)} خوانده می‌شود)
jmeter -n -t payments.jmx -Jthreads=200 -Jrampup=120 -Jduration=600 \
       -l results.jtl -e -o report/

# اجرای توزیع‌شده: کنترلر به چند ماشین سرور وصل می‌شود
jmeter -n -t payments.jmx -R 10.0.0.11,10.0.0.12 -l results.jtl

# تولید داشبورد از یک فایل نتیجه‌ی موجود
jmeter -g results.jtl -o report/

سه تنظیم JMeter که همه اشتباه می‌کنند — (۱) اجرا در حالت GUI برای تست واقعی — GUI خودش منابع می‌خورد و نتایج را تحریف می‌کند. (۲) گذاشتن listenerهای گرافیکی («View Results Tree») در پلن اجرا — این‌ها همه‌ی پاسخ‌ها را در حافظه نگه می‌دارند و در تست طولانی OOM می‌دهند؛ در اجرای بدون GUI حذف یا غیرفعالشان کن. (۳) نادیده گرفتن heap پیش‌فرض — برای بار بالا باید HEAP را در اسکریپت راه‌انداز افزایش دهی، وگرنه بازدارنده‌ی واقعی، GC خودِ JMeter می‌شود نه سیستم تحت تست.

۸٫۶ خواندن نتایج

چهار خانواده‌ی عدد، و ترتیب نگاه‌کردن به آن‌ها:

  1. Throughput — نرخ واقعی تکمیل‌شده (req/s). اگر نرخ واقعی از نرخ هدف کمتر است، یعنی یا load generator کم آورده یا سیستم اشباع شده.
  2. Latency — p50 برای «حس معمول»، p95/p99 برای «بدترین تجربه‌ی رایج»، max برای «چه اتفاق بدی افتاد». حتماً هیستوگرام یا نمودار زمانی ببین، نه فقط یک عدد.
  3. Error rate — و تفکیک نوع خطا: تفاوت ۵۰۳ (رد شدن به‌خاطر ظرفیت)، timeout کلاینت، و ۵۰۰ (باگ) زمین تا آسمان است.
  4. Saturation — منابع: CPU، حافظه، صف thread pool، اشغال connection pool، IO دیسک، پهنای باند. اینجا USE (Utilization, Saturation, Errors) برای منابع و RED (Rate, Errors, Duration) برای سرویس‌ها چارچوب استانداردند.

الگوی مهم برای تشخیص: تا نقطه‌ای، افزایش بار باعث افزایش throughput می‌شود و latency تقریباً ثابت می‌ماند. بعد از زانوی منحنی، throughput مسطح می‌شود ولی latency تصاعدی بالا می‌رود — این یعنی صف تشکیل شده. اگر بعد از زانو throughput کاهش پیدا کند، یعنی سیستم در حال فروپاشی است (thrashing، GC مداوم، timeout و retry طوفانی).

ده اشتباه رایج تست بار

(۱) اندازه‌گیری در محیطی که نصف تولید است و ضرب‌کردن خطی نتیجه. (۲) نادیده‌گرفتن warm-up: JIT هنوز کامپایل نکرده، cache سرد، connection pool خالی — دقایق اول را از آمار حذف کن. (۳) استفاده از یک کاربر/یک شناسه‌ی ثابت که همه‌چیز را در cache می‌نشاند و نتیجه‌ی خیالی می‌دهد. (۴) اشباع خودِ load generator (CPU یا پورت‌های ephemeral یا پهنای باند) و مقصر دانستن سیستم. (۵) نبود think time و ایجاد الگوی همزمانی غیرواقعی. (۶) فقط happy path و بدون هیچ خطایی در بار. (۷) نداشتن مانیتورینگ سمت سرور — تست بدون متریک سرور فقط می‌گوید «کند است» نه «چرا». (۸) دیتابیس تست با یک‌صدم حجم داده‌ی تولید (plan اجرایی کاملاً فرق می‌کند). (۹) سرویس‌های بیرونی واقعی که rate limit می‌خورند و نتیجه را خراب می‌کنند. (۱۰) نبود معیار قبولی از پیش تعریف‌شده، که باعث می‌شود بعد از تست دنبال روایتی بگردی که نتیجه را قابل قبول جلوه دهد.

چطور یک تست بار برای یک API پرداخت جدید طراحی می‌کنی؟

اول هدف و معیار قبولی: با کسب‌وکار روی SLO توافق می‌کنم، مثلاً «p95 زیر ۸۰۰ms و نرخ خطا زیر ۱٪ در بار اوج ۲۰۰ تراکنش بر ثانیه». بدون این عدد، تست فقط نمودار تولید می‌کند.

بعد مدل بار را از داده‌ی واقعی می‌سازم: توزیع endpointها و نسبت‌ها از لاگ ساعت اوج، ضریب رشد شش‌ماهه، و الگوی روزانه. think time واقعی می‌گذارم و از مدل باز (arrival rate) استفاده می‌کنم تا در تله‌ی coordinated omission نیفتم.

بعد محیط و داده: محیطی هم‌اندازه‌ی تولید یا با نسبت مستند و ثابت؛ دیتابیس با حجم داده‌ی نزدیک به تولید، چون plan اجرایی با جدول کوچک کاملاً فرق می‌کند؛ داده‌ی متنوع برای هر کاربر مجازی تا cache نتیجه را جعل نکند؛ و سرویس‌های شریک با stub قراردادمحور به‌علاوه‌ی تأخیر و نرخ خطای واقع‌گرایانه.

بعد اجرا: اول smoke با چند کاربر برای صحت اسکریپت، بعد warm-up، بعد پله‌ای تا بار هدف، بعد حالت پایدار حداقل ۱۵ دقیقه، بعد تست ظرفیت تا نقطه‌ی شکست، و در پایان یک soak شبانه برای نشت.

و در تمام مدت مانیتورینگ سمت سرور: RED برای سرویس، USE برای منابع، آمار GC، اشغال connection pool، و کوئری‌های کند. تحلیل را با یک گلوگاه مشخص و پیشنهاد اقدام تمام می‌کنم — نه با یک اسکرین‌شات از نمودار.

وسط تست بار، p99 منفجر شده ولی CPU فقط ۳۰٪ است. چه اتفاقی افتاده؟

CPU پایین با latency بالا تقریباً همیشه یعنی انتظار در صف، نه کمبود توان محاسباتی. مسیر بررسی من ترتیب دارد.

اول استخرهای محدود: connection pool دیتابیس. اگر pool کوچک باشد، threadها پشت گرفتن connection صف می‌کشند؛ متریک «زمان انتظار برای connection» این را فوراً نشان می‌دهد. قانون لیتل هم همین را می‌گوید: با نرخ ورود و زمان اقامت مشخص، حداقل اندازه‌ی pool معلوم است.

بعد قفل و انحصار: قفل روی یک ردیف داغ، SELECT ... FOR UPDATE روی یک شمارنده‌ی مشترک، یا یک synchronized در مسیر داغ. یک thread dump در اوج، این را در چند ثانیه نشان می‌دهد.

بعد وابستگی‌های کند: یک سرویس پایین‌دستی یا کوئری کند که thread را نگه داشته. اینجا نبودِ timeout و bulkhead باعث می‌شود یک وابستگی کند کل سرویس را قفل کند.

بعد GC و حافظه: مکث‌های طولانی، latency را بدون بارِ CPU پایدار بالا می‌برد.

و در آخر چیزهایی که همیشه فراموش می‌شوند: اشباع خود load generator، محدودیت تعداد پورت ephemeral، و مسدود شدن روی DNS یا TLS handshake. نکته‌ی سطح ارشد این است که CPU پایین معمولاً خبر خوب است: یعنی گلوگاه ساختاری است و با یک تغییر پیکربندی یا یک بهبود همزمانی قابل حل است، نه با اضافه‌کردن ماشین.


بخش ۹ — تست امنیت: حداقلی که از یک بک‌اند انتظار می‌رود

جزئیات آسیب‌پذیری‌ها در فصل «امنیت اپلیکیشن و OWASP» است؛ اینجا فقط جای هر ابزار در چرخه را می‌سازیم، چون همین را در مصاحبه می‌پرسند.

رویکرد چه چیزی را می‌بیند کِی اجرا می‌شود نقطه‌ی ضعف
SAST (تحلیل ایستا) کد منبع بدون اجرا هر PR مثبت کاذب زیاد، منطق زمان اجرا را نمی‌بیند
SCA (تحلیل ترکیب) وابستگی‌های شخص ثالث و CVEها هر build + شبانه CVE بدون بهره‌برداری واقعی؛ نیاز به triage
Secret scanning کلید/توکن در کد و تاریخچه pre-commit + CI الگومحور؛ رازهای غیرمعمول را رد می‌کند
DAST (تحلیل پویا) برنامه‌ی در حال اجرا از بیرون شبانه روی staging نیاز به محیط زنده؛ پوشش وابسته به crawl
IAST / RASP داخل runtime حین تست همراه تست‌های عملکردی نیاز به agent
Pen test / red team زنجیره‌ی حمله‌ی واقعی توسط انسان دوره‌ای گران و نقطه‌ای

نمودار: جای هر نوع تست امنیت در خط لوله — Where each security test type sits in the pipeline.

flowchart LR
  DEV[Commit] --> SEC1[Secret scan + SAST]
  SEC1 --> BUILD[Build]
  BUILD --> SCA[Dependency / SCA scan]
  SCA --> IMG[Container image scan]
  IMG --> DEPLOY[Deploy to staging]
  DEPLOY --> DAST[DAST baseline scan]
  DAST --> PROD[Production]
  PROD -.periodic.-> PEN[Penetration test]
# SCA وابستگی‌های Maven با OWASP Dependency-Check
mvn org.owasp:dependency-check-maven:12.1.3:check \
    -DfailBuildOnCVSS=7 -Dformats=HTML,SARIF

# اسکن ایمیج و فایل‌سیستم با Trivy
trivy image --severity HIGH,CRITICAL --exit-code 1 myapp:3.4.0
trivy fs --scanners vuln,secret,misconfig .

# اسکن پایه‌ی DAST با OWASP ZAP (نسخه‌ی containerized)
docker run --rm -v "$(pwd)":/zap/wrk/:rw \
  ghcr.io/zaproxy/zaproxy:stable zap-baseline.py \
  -t https://staging.example.internal -r zap-report.html

شکستن build روی هر CVE، تیم را به دور زدن گیت وادار می‌کند — اگر هر CVE با CVSS بالای ۷ باعث قرمزشدن build شود، ظرف دو هفته کسی یک flag اضافه می‌کند که اسکن را رد کند و امنیت شما عملاً صفر می‌شود. سیاست عملی: CVEهای قابل بهره‌برداری در مسیر واقعی کد build را می‌شکنند؛ بقیه به‌صورت issue با SLA زمانی ثبت می‌شوند. برای همین است که مفهوم VEX (سند اعلام این‌که یک CVE در محصول ما قابل بهره‌برداری نیست) و «تحلیل قابلیت دسترسی» اهمیت پیدا کرده. و راز واقعی: suppressions.xml باید تاریخ انقضا و دلیل داشته باشد، وگرنه به سطل زباله تبدیل می‌شود.


بخش ۱۰ — دسترس‌پذیری و قابلیت استفاده

دو نوع تست که بک‌اندی‌ها نادیده می‌گیرند و در سازمان‌های بزرگ (به‌ویژه بخش عمومی) الزام قانونی دارند.

  • Accessibility (a11y): آیا فرد با اسکرین‌ریدر، بدون ماوس، یا با کم‌بینایی می‌تواند کار کند؟ مرجع: WCAG 2.2 با سه سطح انطباق A/AA/AAA — هدف عملی صنعت معمولاً AA است.
  • Usability: آیا کاربر بدون آموزش می‌تواند کار را تمام کند؟ با آزمون کاربری روی چند کاربر واقعی سنجیده می‌شود، نه با نظر تیم.
# اسکن خودکار a11y روی یک صفحه با axe (متن‌باز)
npx @axe-core/cli https://staging.example.internal --tags wcag2a,wcag2aa

ابزار خودکار فقط بخشی از کار را می‌کند — تحلیل‌های صنعتی نشان می‌دهند ابزارهای خودکار تنها بخشی (تقریباً یک‌سوم تا نیمی) از مسائل WCAG را می‌گیرند — چیزهایی مثل نبود alt، کنتراست پایین یا نبود label. اما «آیا متن جایگزین معنادار است؟» یا «آیا ترتیب فوکوس منطقی است؟» قضاوت انسانی می‌خواهد. حداقل کاری که هر تیم باید بکند: یک بار پیمایش کامل فرم‌های حیاتی فقط با کیبورد. همین یک تمرین، بیشتر باگ پیدا می‌کند از یک ماه اسکن خودکار.


بخش ۱۱ — TDD و BDD، صادقانه

۱۱٫۱ حلقه‌ی قرمز-سبز-بازآرایی

کوه‌نوردی با طناب

TDD مثل بستن طناب قبل از هر حرکت است. هر بار کمی بالا می‌روی، طناب را محکم می‌کنی، بعد حرکت بعدی. کند به نظر می‌رسد تا لحظه‌ای که پایت سُر بخورد. بازآرایی بدون تست، یعنی کوه‌نوردی بدون طناب: سریع‌تر، تا وقتی که نیست.

نمودار: چرخه‌ی TDD و قانون هر مرحله — The TDD cycle and the rule of each step.

stateDiagram-v2
  [*] --> RED
  RED --> GREEN: write the simplest code that passes
  GREEN --> REFACTOR: tests stay green
  REFACTOR --> RED: next small behaviour
  note right of RED
    Test must fail for the RIGHT reason
  end note
  note right of REFACTOR
    Change structure, never behaviour
  end note

قواعدی که TDD واقعی را از «تست‌نویسی بعد از کد» جدا می‌کند:

  1. تست را قبل از کد بنویس و مطمئن شو به دلیل درست شکست می‌خورد (نه به‌خاطر خطای کامپایل یا typo).
  2. ساده‌ترین کدی که تست را سبز می‌کند بنویس، حتی اگر ابتدایی باشد.
  3. فقط وقتی سبز است بازآرایی کن — و در فاز بازآرایی هیچ رفتار جدیدی اضافه نکن.
  4. گام‌ها را کوچک نگه دار؛ اگر بیش از چند دقیقه قرمزی، گام را کوچک‌تر کن.
قضاوت ارشد — TDD کِی کمک می‌کند و کِی نه

کمک می‌کند وقتی: منطق مشخص و قابل بیان دارد (محاسبه، قواعد، پارسر، ماشین حالت)؛ نیازمندی نسبتاً پایدار است؛ یا داری باگی را رفع می‌کنی (اول تستی که باگ را بازتولید می‌کند). کمک نمی‌کند وقتی: هنوز داری کاوش می‌کنی و طراحی هر ساعت عوض می‌شود (spike بزن و بعد دور بریز)؛ کار اساساً یکپارچه‌سازی و پیکربندی است (تست اول در برابر یک API خارجی ناشناخته، عملاً حدس‌زدن است)؛ یا UI اکتشافی است. جواب بالغ در مصاحبه این نیست که «همیشه TDD»، این است که «برای منطق دامنه بله؛ برای لایه‌ی adapter، اول یک spike می‌زنم تا شکل API را بفهمم، بعد تستش را تثبیت می‌کنم».

تست‌هایی که پیاده‌سازی را قفل می‌کنند — اگر تست‌هایت هر همکاری داخلی را mock کنند و روی ترتیب فراخوانی‌ها verify بگذارند، به‌جای رفتار، ساختار را قفل کرده‌ای. نشانه: هر بازآرایی — حتی بدون تغییر رفتار — ده تست را می‌شکند. این دقیقاً برعکس هدف تست است. قاعده‌ی عملی: mock را برای مرزهای بیرونی (پرداخت، ایمیل، صف) نگه دار؛ برای همکارهای داخلی از شیء واقعی یا fake استفاده کن. تست باید بگوید «چه اتفاقی افتاد»، نه «چطور».

۱۱٫۲ BDD و Gherkin

BDD (Behaviour-Driven Development) قبل از این‌که یک ابزار باشد، یک گفت‌وگو است: کسب‌وکار، توسعه و تست با هم مثال‌های مشخص می‌سازند («three amigos»). زبان Gherkin آن مثال‌ها را در قالبی می‌نویسد که هم انسان بخواند هم ماشین اجرا کند.

# src/test/resources/features/withdrawal.feature
# language: en
Feature: Cash withdrawal limits
  As a bank customer
  I want withdrawals above my daily limit to be reviewed
  So that fraudulent transfers can be stopped

  Background:
    Given an active account "ACC-1001" with balance 50000000 IRR
    And two-factor authentication is enabled for "ACC-1001"

  Scenario: Withdrawal within the daily limit is approved immediately
    When a withdrawal of 5000000 IRR is requested from "ACC-1001"
    Then the withdrawal is approved
    And the balance of "ACC-1001" becomes 45000000 IRR

  Scenario Outline: Withdrawals above the daily limit need manual review
    When a withdrawal of <amount> IRR is requested from "ACC-1001"
    Then the withdrawal status is "<status>"

    Examples:
      | amount   | status         |
      | 20000000 | PENDING_REVIEW |
      | 20000001 | PENDING_REVIEW |
      | 19999999 | APPROVED       |

پیاده‌سازی گام‌ها در جاوا با Cucumber-JVM:

package com.example.acceptance;

import io.cucumber.java.en.Given;
import io.cucumber.java.en.When;
import io.cucumber.java.en.Then;
import static org.assertj.core.api.Assertions.assertThat;

public class WithdrawalSteps {

    private final WithdrawalService service;   // تزریق‌شده با cucumber-spring
    private WithdrawalResult result;

    public WithdrawalSteps(WithdrawalService service) { this.service = service; }

    @Given("an active account {string} with balance {long} IRR")
    public void anActiveAccount(String accountId, long balance) {
        service.createAccount(accountId, balance);
    }

    @Given("two-factor authentication is enabled for {string}")
    public void mfaEnabled(String accountId) {
        service.enableMfa(accountId);
    }

    @When("a withdrawal of {long} IRR is requested from {string}")
    public void requestWithdrawal(long amount, String accountId) {
        result = service.withdraw(accountId, amount);
    }

    @Then("the withdrawal status is {string}")
    public void statusIs(String expected) {
        assertThat(result.status().name()).isEqualTo(expected);
    }
}
package com.example.acceptance;

import org.junit.platform.suite.api.ConfigurationParameter;
import org.junit.platform.suite.api.IncludeEngines;
import org.junit.platform.suite.api.SelectClasspathResource;
import org.junit.platform.suite.api.Suite;

import static io.cucumber.junit.platform.engine.Constants.GLUE_PROPERTY_NAME;
import static io.cucumber.junit.platform.engine.Constants.PLUGIN_PROPERTY_NAME;

@Suite
@IncludeEngines("cucumber")
@SelectClasspathResource("features")
@ConfigurationParameter(key = GLUE_PROPERTY_NAME, value = "com.example.acceptance")
@ConfigurationParameter(key = PLUGIN_PROPERTY_NAME,
                        value = "pretty, html:target/cucumber-report.html")
class RunAcceptanceTest { }
<dependency>
  <groupId>io.cucumber</groupId>
  <artifactId>cucumber-java</artifactId>
  <version>7.22.1</version>
  <scope>test</scope>
</dependency>
<dependency>
  <groupId>io.cucumber</groupId>
  <artifactId>cucumber-junit-platform-engine</artifactId>
  <version>7.22.1</version>
  <scope>test</scope>
</dependency>
<dependency>
  <groupId>org.junit.platform</groupId>
  <artifactId>junit-platform-suite</artifactId>
  <scope>test</scope>
</dependency>
Gherkin که به اسکریپت UI تبدیل شده، بدترین هر دو دنیاست

سناریوی بد: «Given کاربر روی دکمه‌ی ورود کلیک می‌کند / And در فیلد اول نام کاربری می‌نویسد / And روی دکمه‌ی سوم کلیک می‌کند». این نه برای کسب‌وکار خواندنی است نه برای مهندس قابل نگهداری، و با هر تغییر کوچک UI می‌شکند. سناریو را در سطح رفتار کسب‌وکار بنویس و جزئیات تعامل را در کد گام پنهان کن. آزمون ساده: اگر یک نفر از تیم کسب‌وکار سناریو را بخواند و بگوید «خب که چه؟»، سناریو غلط نوشته شده.

هزینه‌ی Cucumber را وقتی بپرداز که خریدارش وجود دارد — Cucumber یک لایه‌ی غیرمستقیم اضافه می‌کند: بین feature و کد، یک نگاشت regex هست که باید نگهداری شود. این هزینه فقط وقتی می‌ارزد که واقعاً کسی غیر از توسعه‌دهنده‌ها آن feature fileها را می‌خواند یا می‌نویسد. اگر تنها خواننده‌ی feature تیم مهندسی است، همان تست با JUnit و AssertJ خواناتر، سریع‌تر و قابل دیباگ‌تر است. جواب سطح ارشد در مصاحبه دقیقاً همین است: «BDD را به‌عنوان روش گفت‌وگو همیشه استفاده می‌کنم؛ Cucumber را فقط وقتی که مخاطب غیرفنی واقعی دارد.»

Living documentation یعنی خروجی اجرای همین سناریوها (گزارش HTML) به‌عنوان مستند رسمی رفتار سیستم منتشر شود. مزیتش این است که هرگز کهنه نمی‌شود: اگر رفتار عوض شود و سند به‌روز نشود، تست قرمز می‌شود. این تنها نوع مستندی است که خودش را صادق نگه می‌دارد.

فرق TDD و BDD چیست؟

TDD یک تکنیک توسعه است در سطح توسعه‌دهنده: قرمز، سبز، بازآرایی؛ هدفش بازخورد سریع و طراحی بهتر است. BDD یک رویکرد همکاری است در سطح تیم و کسب‌وکار: قبل از کد، با مثال‌های مشخص روی معنای رفتار توافق می‌کنی و همان مثال‌ها به تست اجراشدنی تبدیل می‌شوند.

از نظر تکنیکی، BDD در واقع TDD است که واژگانش عوض شده تا روی رفتار تمرکز کند نه روی متد؛ به همین دلیل به آن ATDD یا specification by example هم می‌گویند. یکی جای دیگری را نمی‌گیرد: در پروژه‌ی واقعی من BDD را در سطح پذیرش برای چند جریان کلیدی به کار می‌برم و TDD را در سطح واحد برای منطق دامنه.

و صادقانه: بزرگ‌ترین ارزش BDD در خود ابزار نیست، در گفت‌وگوی سه‌جانبه قبل از کدنویسی است. تیم‌هایی که Cucumber را نصب می‌کنند ولی آن گفت‌وگو را ندارند، فقط یک لایه‌ی regex به تست‌هایشان اضافه کرده‌اند.

وقتی می‌گویی «کد را بازآرایی کردم»، از کجا مطمئنی چیزی نشکسته؟

اول، تعریف را دقیق می‌کنم: بازآرایی یعنی تغییر ساختار داخلی بدون تغییر رفتار قابل مشاهده. اگر رفتار عوض شود، آن بازآرایی نیست، تغییر است — و تست‌ها باید همراهش عوض شوند.

ابزار من: مجموعه‌ی تستی که رفتار را از بیرونِ واحدِ در حال تغییر پوشش می‌دهد. اگر تست‌ها به جزئیات داخلی چسبیده باشند، شبکه‌ی ایمنی ندارم، دارم همان کد را دو بار می‌نویسم. پس قبل از یک بازآرایی بزرگ، اول تست‌های سطح بالاتر (characterization test) می‌نویسم که رفتار فعلی را — حتی اگر عجیب باشد — ثبت کنند.

برای کد قدیمیِ بدون تست، از تکنیک seam استفاده می‌کنم: نقطه‌ای که می‌توانم بدون تغییر رفتار، وابستگی را تزریق‌پذیر کنم و تست بنویسم. و برای اطمینان از کیفیت خودِ آن شبکه‌ی ایمنی، یک بار PIT را روی همان بسته اجرا می‌کنم؛ اگر جهش‌های زیادی زنده بمانند، تست‌هایم برای بازآرایی کافی نیستند.

آخرین لایه: بازآرایی را در commitهای کوچک و جدا از تغییر رفتار انجام می‌دهم، تا اگر چیزی در تولید خراب شد، revert سریع و بی‌ابهام باشد.

یک باگ در تولید گزارش شده. اولین کاری که می‌کنی چیست؟

اولویت اول مهار است نه ریشه‌یابی: وسعت اثر را می‌سنجم (چند کاربر، چه مبلغی، از کِی)، و اگر لازم باشد با feature flag یا rollback جلوی خون‌ریزی را می‌گیرم. ریشه‌یابی بعد از تثبیت.

بعد بازتولید: از لاگ‌ها با traceId و از داده‌ی واقعی، کوچک‌ترین حالتی را می‌سازم که باگ را نشان می‌دهد. تا وقتی نتوانم بازتولید کنم، هر «رفعی» حدس است.

بعد تست شکست‌خورده اول: یک تست خودکار در پایین‌ترین سطح ممکن می‌نویسم که دقیقاً همان باگ را بازتولید کند و قرمز باشد. این هم اثبات می‌کند که فهمیده‌ام مشکل چیست، هم برای همیشه تبدیل به تست بازگشتی می‌شود.

بعد رفع و تأیید: کد را اصلاح می‌کنم تا آن تست سبز شود (confirmation test)، و مجموعه‌ی رگرسیون را اجرا می‌کنم.

و در پایان پرسش سیستمی: چرا این باگ از همه‌ی لایه‌ها رد شد؟ آیا تکنیک طراحی تست کم داشتیم (مثلاً مقدار مرزی)؟ آیا در مرور کد قابل دیدن بود؟ آیا مانیتورینگ باید زودتر هشدار می‌داد؟ خروجی یک postmortem بدون سرزنش، معمولاً یک تغییر در فرایند است نه یک خط کد.


بخش ۱۲ — چطور در مصاحبه درباره‌ی تست حرف بزنی

سه اشتباه رایج: (۱) شمردن اسم ابزار به‌جای توضیح تصمیم؛ (۲) ادعای «همیشه TDD» که با یک سؤال پیگیر فرو می‌ریزد؛ (۳) نداشتن هیچ عددی.

چارچوبی که جواب می‌دهد وقتی می‌پرسند «چطور X را تست می‌کنی؟»:

  1. ریسک — چه چیزی اگر خراب شود بیشترین آسیب را دارد؟ از همان‌جا شروع کن.
  2. سطح — کدام تست در کدام لایه ارزان‌ترین جواب را می‌دهد.
  3. تکنیک — نام ببر: افراز هم‌ارزی، مقدار مرزی، جدول تصمیم، انتقال حالت.
  4. داده و محیط — داده از کجا، سرویس بیرونی چطور بدل می‌شود.
  5. معیار قبولی — کِی می‌گویی تمام است.
  6. آنچه تست نمی‌کنم — و چرا این ریسک قابل قبول است.

عددهایی که یک سِنیور درباره‌ی کیفیت می‌آورد (و با معیارهای DORA هم‌خانواده‌اند): نرخ نقص فراری (باگ‌هایی که به تولید رسیدند)، change failure rate، زمان بازیابی (MTTR)، زمان اجرای pipeline، نرخ flakiness، و پوشش نیازمندی‌های P1. اگر بتوانی بگویی «pipeline از ۴۰ به ۹ دقیقه رسید و نرخ نقص فراری در سه ماه نصف شد»، دیگر لازم نیست کسی را قانع کنی که تست بلدی.


بخش ۱۳ — چک‌لیست قضاوت درباره‌ی یک test suite

روز اولی که وارد یک کد جدید می‌شوی، این‌ها را بررسی کن. جواب‌ها سلامت مهندسی تیم را بهتر از هر مستندی نشان می‌دهند:

# پرسش نشانه‌ی سلامت
۱ تست‌ها روی لپ‌تاپ تازه با یک دستور اجرا می‌شوند؟ mvn verify کافی است
۲ تست واحد چقدر طول می‌کشد؟ کل suite واحد زیر ۹۰ ثانیه
۳ تست‌ها به هم وابسته‌اند؟ با ترتیب تصادفی و موازی هم سبز
۴ assertها معنادارند یا فقط «not null»؟ assert روی رفتار مشخص
۵ تست منفی و مرزی وجود دارد؟ نسبت به نفع مسیرهای شکست
۶ mock کجاست؟ فقط روی مرزهای بیرونی
۷ داده‌ی تست چطور ساخته می‌شود؟ builder، نه SQL دستی تکراری
۸ تست flaky شناخته‌شده هست؟ فهرست قرنطینه با مالک و مهلت
۹ قرمز شدن CI چه معنایی دارد؟ جلوی merge را می‌گیرد و کسی reruns نمی‌زند
۱۰ نام تست‌ها رفتار را می‌گویند؟ shouldRejectWithdrawalAboveLimit نه test1
۱۱ آیا تست غیرکارکردی وجود دارد؟ حداقل یک تست بار خودکار روی مسیر بحرانی
۱۲ آخرین باگ تولید، تست بازگشتی دارد؟ هر incident یک تست جدید ساخته

جدول مرجع فرمان‌ها

کار فرمان
فقط تست‌های واحد mvn -q test
واحد + یکپارچه‌سازی mvn -q verify
یک تست خاص mvn test -Dtest=PaymentServiceTest#shouldRejectAboveLimit
تست‌های یک تگ mvn test -Dgroups=integration
گزارش پوشش JaCoCo mvn verify سپس target/site/jacoco/index.html
mutation testing mvn test-compile org.pitest:pitest-maven:mutationCoverage
mutation فقط روی تغییرات mvn org.pitest:pitest-maven:scmMutationCoverage -Dinclude=ADDED,MODIFIED
اجرای Cucumber mvn test -Dtest=RunAcceptanceTest
Gatling mvn gatling:test -Dgatling.simulationClass=simulations.PaymentSimulation
k6 محلی k6 run payments-load.js
k6 با بازنویسی گزینه k6 run --vus 50 --duration 5m payments-load.js
JMeter بدون GUI + گزارش jmeter -n -t plan.jmx -l out.jtl -e -o report/
JMeter توزیع‌شده jmeter -n -t plan.jmx -R host1,host2 -l out.jtl
اسکن وابستگی mvn org.owasp:dependency-check-maven:check -DfailBuildOnCVSS=7
اسکن ایمیج trivy image --severity HIGH,CRITICAL myapp:tag
DAST پایه zap-baseline.py -t https://staging.example.internal -r report.html
تولید مجموعه‌ی pairwise pict model.txt > pairs.tsv
جمع‌بندی

تست، «نوشتن @Test» نیست؛ یک دیسیپلین مهندسی است. سطح می‌گوید کجا تست می‌کنی و نوع می‌گوید دنبال چه ویژگی‌ای هستی. تکنیک‌های طراحی — افراز هم‌ارزی، مقدار مرزی، جدول تصمیم، انتقال حالت، pairwise، حدس خطا — همان چیزی هستند که از یک نیازمندی مبهم، فهرست موارد تستِ کامل می‌سازند. مستندات (Test Plan با «out of scope» صریح، test case با نتیجه‌ی مورد انتظار غیرقابل‌بحث، ماتریس ردیابی، گزارش نقص با تمایز severity/priority، و معیار خروج) چیزی است که سازمان بالغ از تو می‌خواهد. هرم را اقتصادی نگه دار، جای E2E اضافی را با contract test پر کن، flaky را قرنطینه‌ی مهلت‌دار کن، coverage را سیگنال بدان و با mutation testing بسنجش. در تست غیرکارکردی، اول SLO، بعد مدل بار باز برای فرار از coordinated omission، بعد ابزار — k6 برای گیت CI، Gatling برای تیم جاوا، JMeter برای پروتکل‌های سازمانی — و نتیجه را با throughput، صدک‌ها، نرخ خطا و اشباع بخوان. امنیت را با SAST/SCA/DAST در جای درست خط لوله بگذار و دسترس‌پذیری را فقط به ابزار خودکار نسپار. TDD را برای منطق دامنه به کار ببر و Cucumber را فقط وقتی که مخاطب غیرفنی واقعی دارد. و در مصاحبه، به‌جای اسم ابزار، ریسک، تصمیم، معیار قبولی و آنچه عمداً تست نکردی را بگو — این چیزی است که سِنیور را از میانی جدا می‌کند.

You can probably write an @Test. You mock with Mockito, spin up a real Postgres with Testcontainers, write an integration test. This site's "testing" chapter covers that. But job postings ask for something else, and that is exactly where most five-year engineers slip: "familiar with test levels and types", "able to write a Test Plan and Test Cases", "experience with performance and load testing".

Those are not mechanics, they are discipline. The difference between a mid-level engineer and a senior one is not who knows verify(). It is that when both are told "test this feature", one writes a few arbitrary tests and the other produces, in ten minutes, a list of test conditions that misses nothing important, knows which of them to automate and which not to, and can state out loud: "with this suite, here is the risk that remains open."

This chapter is exactly that layer: the mental framework, the professional vocabulary, the techniques that generate test cases, the documentation a mature organisation expects, and non-functional testing — load testing above all — with real tools and runnable examples.

Roadmap for this chapter

First we build the base vocabulary (error/defect/failure, verification versus validation) and draw the boundaries between test levels and test types. Then black-box/white-box/grey-box and seven test design techniques with worked examples. Then documentation: Test Plan, Test Case, traceability matrix, defect report, exit criteria. Then the test pyramid versus the ice-cream cone and where contract tests fit. Then automation strategy: what to automate, flaky tests, test data and environments, coverage and mutation testing. Then the heavy part: non-functional testing — performance/load/stress/soak/spike, SLOs and workload models, and a comparison of JMeter, Gatling and k6 with a runnable example each plus how to read the results. Then security testing (SAST/DAST/SCA) and accessibility. Then TDD and BDD without the slogans, with a Java Cucumber example. And finally: how to talk about testing in an interview, and a checklist for judging a codebase's test suite.


Part 0 — The vocabulary seniors do not get wrong

The pre-flight checklist

A pilot reads a checklist before every flight. Not because they are forgetful, but because under pressure human memory is the worst instrument in the world. A checklist means "the thinking was done in advance". Good testing is the same: the decision about what must work is taken before the pressure of delivery day. If you write your tests while the PR is open and your manager is standing behind you, what you wrote is not a checklist, it is a comfort blanket.

Error, defect, failure

Three words that everyday speech merges and professional conversation keeps apart:

  • error (mistake) — the human act. The developer typed < where <= belonged.
  • defect (fault, bug) — the result of that error, sitting in the code or the document. That wrong < is right there in the file.
  • failure — when the defect manifests during execution and behaviour deviates from expectation. A user with a balance of exactly 10,000 cannot withdraw.

Why does it matter? Because a defect can sit for years without producing a failure (that path is never executed), and because not every failure comes from a code defect — environmental conditions (bad memory, a wrong system clock, a network partition) produce failures too.

Verification versus validation

  • Verification: "did we build the product right?" — conformance to specification and design.
  • Validation: "did we build the right product?" — does it solve the real user need.

You can have a system that is 100% green (perfect verification) and that nobody uses (zero validation). This distinction is the heart of the UAT conversation.

Senior judgment — "quality" is not a number, it is residual risk — The goal of testing is not to prove the absence of defects; for any non-trivial program that is impossible (the input space is effectively infinite). The goal is reducing risk to an acceptable level at a sane cost. Whenever someone asks "how much testing is enough?", the professional answer is: "until the residual risk falls below the business's acceptance threshold" — and then you must be able to name that risk.

Why earlier is cheaper

The later a defect is found, the more it costs — not because of some magic multiplier, but for plain reasons: more code has been built on top of it, more people get involved, and if it reached production you also owe a data fix, a customer notification and possibly compensation. That is why shift-left (pushing testing leftward on the timeline: requirement reviews, unit tests, tests in CI) is economics, not aesthetics.

And alongside it, shift-right: some testing genuinely cannot happen before production (real user behaviour, real traffic). There, canary releases, feature flags, synthetic monitoring and chaos engineering are your tools — which ties into the observability and resilience chapters.


Part 1 — Test levels: who catches what

A test level is a group of test activities managed together, usually tied to a stage of development. The common standard (ISTQB v4) counts five:

Level What is under test Who typically Environment What it catches
Component (unit) one class/function/module alone developer laptop + CI logic, conditions, boundaries
Component integration interaction between parts inside one service developer CI with Testcontainers mapping, transactions, internal contracts
System the whole service/product from outside dev team / QA production-like end-to-end flows, non-functional behaviour
System integration interaction across systems/external services QA / integration team staging with partner sandboxes protocol, versions, partner errors
Acceptance readiness for delivery and use business / users / ops UAT / pre-prod "is this what we actually asked for?"

Diagram: mapping test levels to design levels in the V-model — نمودار: نگاشت سطوح تست به سطوح طراحی در مدل V.

flowchart LR
  R[Requirements] --> A[Architecture]
  A --> D[Detailed design]
  D --> C[Code]
  C --> UT[Component tests]
  UT --> IT[Component integration]
  IT --> ST[System tests]
  ST --> AT[Acceptance tests]
  R -.validated by.-> AT
  A -.verified by.-> ST
  D -.verified by.-> IT

The acceptance sub-types interviewers ask about

  • UAT (User Acceptance Test) — the user or product owner runs real business scenarios.
  • OAT (Operational Acceptance Test) — the ops team: does backup work? Is there a rollback path? Are alerts wired? Is the runbook written? Junior teams forget this one and discover it at 2 a.m. on release night.
  • Contractual / regulatory acceptance — conformance to a contract or regulation (log retention, privacy obligations).
  • Alpha / beta — alpha at the producer's site with real users; beta in the user's own environment.
The blurred-boundary trap

The most common disease across teams everywhere is the same: a "unit test" that boots a Spring context and hits a database. It is named unit, behaves like a system test, is catastrophically slow, and when it breaks you cannot tell whether the logic or the environment is guilty. Define the boundary by "what can break this test", not by the filename. If a database schema change can turn your test red, that test is not a unit test.

Where exactly is the line between a unit test and an integration test, and how do you draw it?

I draw it by out-of-process dependencies, not by class count. A unit test runs code inside the same process, does no real I/O (no network, no disk, no uncontrolled system clock), finishes in milliseconds and is deterministic. If a test exercises several collaborating classes but stays in-process and deterministic, I still call it a unit test — that is the "sociable/classicist" school, as opposed to the "solitary/mockist" school that mocks every collaborator.

An integration test is where we genuinely cross a process boundary: a real database via Testcontainers, a real broker, the filesystem. Its value is catching what a mock never can — column mappings, isolation-level behaviour, unique constraints, message serialisation.

In practice my rule is: keep decision logic in a pure core and cover it with cheap, plentiful unit tests; write a thin layer of integration tests per adapter (repository, client, listener) that only checks the translation; and keep a handful of system tests for critical flows. That is precisely what makes hexagonal architecture valuable from a testing point of view.


Part 2 — Test types: functional, non-functional, structural, change-related

The level says where you test; the type says which property you are hunting. Four families:

  1. Functional — "what does it do?" Correct calculations, business rules, flows.
  2. Non-functional — "how well does it do it?" Performance, security, reliability, usability, compatibility, maintainability, portability. These are the quality characteristics of ISO/IEC 25010.
  3. Structural (white-box) — "is the code structure exercised?" Statement, branch, path, MC/DC coverage.
  4. Change-relatedconfirmation testing (did that specific bug actually get fixed?) and regression testing (did anything else break?).

ISO/IEC 25010 and the 2023 revision — The product quality model of ISO/IEC 25010 was revised in 2023 and now has 9 characteristics, with "Safety" promoted to a top-level characteristic (previously there were 8). When an interviewer asks about "types of non-functional testing", naming this model and listing a few of its characteristics — performance efficiency, reliability, security, usability, compatibility, maintainability, portability, functional suitability, safety — shows you have a framework rather than a memorised list.

Do not push non-functional testing to the end of the project — Non-functional testing is not a "final phase". If your architecture works with 50 concurrent users and your SLO is 5,000, no amount of last-week tuning will save you; the data model and access patterns have to change. A small but early load test on the critical path is worth more than a giant one two days before go-live.

What is the difference between smoke, sanity and regression testing?

These three are not "test types" in the technical sense; they are execution suites with different purposes, and that is what you should explain.

A smoke test is a very small, very fast suite run after every build to answer "is this build even testable?" — the service starts, health is green, login works, one simple transaction goes through. If smoke is red we do not proceed to deeper testing; we reject the build. Broad but shallow.

A sanity test is the inverse: narrow but deep. After a bug fix or a small change, we carefully check that area and its immediate neighbourhood to confirm the logic is now right, without running the whole suite.

A regression test is the full suite that confirms the new change did not break existing behaviour. Because it is large, a subset usually runs per PR and the full set runs nightly; test impact analysis — selecting tests based on which code actually changed — cuts execution time dramatically.

The point that usually earns credit: confirmation testing is not regression testing. Confirmation asks "is the reported bug actually fixed?"; regression asks "did anything else break?". Every bug fix should have both, and the confirmation test should stay in the suite forever.


Part 3 — Black-box, white-box, grey-box

The washing machine

If you work only from the manual, press buttons and observe results, you are black-box testing. If you open the back and trace the wiring and the control board, that is white-box. If you hold the manual but happen to know there is an inverter motor inside and design your trials accordingly, that is grey-box.

Approach Based on Strength Blind spot
Black-box specification / requirements independent of implementation, survives refactoring hidden internal branches and error paths
White-box code structure branch coverage, dead code, compound conditions never finds a missing requirement
Grey-box spec plus internal knowledge smart probing of caches, indexes, queues, boundaries partially coupled to implementation

The key point seniors make: 100% code coverage guarantees nothing about requirements that were never written. If the developer forgot to implement the "negative balance" case, no coverage tool will tell you something is missing. That is why black-box techniques (next part) are the foundation and white-box is the complement.


Part 4 — Test design techniques: getting from a requirement to a list of tests

This part is the heart of the chapter. If you learn only one section deeply, make it this one — because in an interview you will typically be handed a small requirement and asked to "write the test cases".

4.1 Equivalence partitioning

The idea: split the inputs into groups the system is expected to treat identically, then test one representative per group. If one representative reveals a defect, so would the others; more tests from the same partition cost money and add no information.

Worked example — a transfer fee:

Transfer amounts from 10,000 to 50,000,000 are allowed. Up to 1,000,000 the fee is a flat 5,000; above 1,000,000 and up to 10,000,000 the fee is 0.1%; above 10,000,000 the fee is a flat 25,000. The amount must be an integer.

Partitions — both valid and invalid, and it is the invalid ones most candidates forget:

# Partition Sample Valid?
P1 amount < 10,000 9,000 invalid
P2 10,000 ≤ amount ≤ 1,000,000 500,000 valid (flat fee)
P3 1,000,000 < amount ≤ 10,000,000 5,000,000 valid (percentage)
P4 10,000,000 < amount ≤ 50,000,000 30,000,000 valid (upper flat fee)
P5 amount > 50,000,000 60,000,000 invalid
P6 non-numeric / decimal / negative / empty "abc", 1000.5, -5, null invalid
Senior judgment — one invalid value per test, never two

For valid partitions you may combine several in one test case (valid amount + valid currency + active user). But for invalid partitions, make exactly one thing invalid at a time. Why? If you send both a negative amount and an unknown currency, the system rejects on the first and you never learn whether currency validation exists at all. That phenomenon is called fault masking.

4.2 Boundary value analysis

The idea: bugs nest on boundaries, because < and <= are one character apart. For every boundary, test the values around it.

Two schools:

  • 2-value BVA: the boundary itself and the first value outside it. For the 1,000,000 boundary: 1,000,000 and 1,000,001.
  • 3-value BVA: one before, the boundary, one after — 999,999, 1,000,000, 1,000,001. Stricter, and recommended for financial or safety logic.

For the example above with 3-value BVA the test values are 9,999 / 10,000 / 10,001, 999,999 / 1,000,000 / 1,000,001, 9,999,999 / 10,000,000 / 10,000,001 and 49,999,999 / 50,000,000 / 50,000,001.

And that maps straight onto a JUnit 5 parameterised test:

@ParameterizedTest(name = "amount={0} -> fee={1}")
@CsvSource({
    "10_000,      5_000",
    "1_000_000,   5_000",
    "1_000_001,   1_000",     // 0.1% of 1,000,001 rounded down
    "10_000_000, 10_000",
    "10_000_001, 25_000",
    "50_000_000, 25_000"
})
void feeIsCorrectAtBoundaries(long amount, long expectedFee) {
    assertThat(feeCalculator.feeFor(amount)).isEqualTo(expectedFee);
}

@ParameterizedTest
@ValueSource(longs = {9_999L, 50_000_001L, -1L, 0L})
void amountsOutsideAllowedRangeAreRejected(long amount) {
    assertThatThrownBy(() -> feeCalculator.feeFor(amount))
        .isInstanceOf(AmountOutOfRangeException.class);
}
Boundaries are not only numbers

The boundaries people forget: string length (0, 1, max, max+1), collection size (empty, single element, exactly page size, one more), time (midnight, last day of month, 29 February, daylight-saving transitions, timezone edges), numeric types (Integer.MAX_VALUE, long overflow, -0.0, NaN), and Unicode (a four-byte emoji makes String length differ from the number of visible characters). A classic production bug: a "max 200 characters" field that explodes on 200 emoji because the database column is a byte-based VARCHAR(200).

4.3 Decision tables

When the outcome depends on a combination of conditions, equivalence partitioning is not enough; you must lay out the combinations systematically.

Worked example — withdrawal authorisation: Conditions: (C1) is the account active? (C2) are funds sufficient? (C3) has the user completed two-factor authentication? (C4) is the amount above the daily limit?

Rule C1 active C2 funds C3 MFA C4 over limit Outcome
R1 N reject: account blocked
R2 Y N reject: insufficient funds
R3 Y Y N reject: MFA required
R4 Y Y Y Y manual review
R5 Y Y Y N approve immediately

Important detail: means don't care; the full table has 2⁴ = 16 combinations and rule collapsing brought us to 5. That collapsing is the point, and it is what an interviewer is looking for: show that you understand which combinations are unreachable or irrelevant.

Senior judgment — put the decision table into the code — If you drew a decision table, feed that same table into the test as the source of truth. With @CsvSource, each row becomes one rule, and if someone adds a rule tomorrow the PR diff shows exactly which rule changed. This fuses document and test and prevents the classic "documentation that drifted from the code".

@ParameterizedTest(name = "R{index}: active={0} funds={1} mfa={2} overLimit={3} -> {4}")
@CsvSource({
    "false, true,  true,  false, ACCOUNT_BLOCKED",
    "true,  false, true,  false, INSUFFICIENT_FUNDS",
    "true,  true,  false, false, MFA_REQUIRED",
    "true,  true,  true,  true,  MANUAL_REVIEW",
    "true,  true,  true,  false, APPROVED"
})
void withdrawalDecisionTable(boolean active, boolean funds, boolean mfa,
                             boolean overLimit, Decision expected) {
    var request = new WithdrawalRequest(active, funds, mfa, overLimit);
    assertThat(policy.decide(request)).isEqualTo(expected);
}

4.4 State transition testing

When the system has memory — the response to an event depends on history — draw the state machine.

Diagram: an order state machine with the legal events — نمودار: ماشین حالت یک سفارش و رویدادهای مجاز.

stateDiagram-v2
  [*] --> CREATED
  CREATED --> PAID: pay
  CREATED --> CANCELLED: cancel
  PAID --> SHIPPED: ship
  PAID --> REFUNDED: refund
  SHIPPED --> DELIVERED: deliver
  SHIPPED --> RETURNED: return
  DELIVERED --> RETURNED: return
  RETURNED --> REFUNDED: refund
  CANCELLED --> [*]
  REFUNDED --> [*]

Three coverage levels come out of that diagram:

  • State coverage: every state visited at least once. The weakest level.
  • Transition coverage / 0-switch: every edge traversed at least once. This is the acceptable minimum.
  • 1-switch: every consecutive pair of transitions (e.g. pay then ship) traversed. Catches history-dependent bugs.

And most importantly: the invalid transitions. Draw the full states × events grid and test the empty cells — "what does ship do in state CREATED?" The right answer is a specific error with no state change, not a NullPointerException.

@ParameterizedTest
@EnumSource(OrderState.class)
void shipIsOnlyLegalFromPaid(OrderState from) {
    Order order = Order.inState(from);
    if (from == OrderState.PAID) {
        order.ship();
        assertThat(order.state()).isEqualTo(OrderState.SHIPPED);
    } else {
        assertThatThrownBy(order::ship)
            .isInstanceOf(IllegalStateTransitionException.class);
        assertThat(order.state()).isEqualTo(from);   // no side effect
    }
}

An invalid transition must be inert, not merely throw — The most common state-machine bug is a method that performs a side effect first (publishes an event, mutates a field) and then checks the guard and throws. Without a surrounding transaction you now hold a half-mutated state. In every invalid-transition test, assert not only the exception but also that the state did not change and no event was published.

4.5 Pairwise (all-pairs) testing

With several independent parameters the Cartesian product explodes. Suppose browser (4) × OS (3) × locale (5) × payment type (4) × user tier (3) = 720 combinations. Testing them all is impossible.

The empirical observation: the overwhelming majority of defects are triggered by the interaction of one or two parameters, not five. So if you build a set where every pair of values from every two parameters appears together at least once, roughly 20 combinations give you nearly the same detection power.

Practical tooling: PICT (Microsoft, open source) or ACTS (NIST). A PICT model file is plain text:

Browser:  Chrome, Firefox, Safari, Edge
OS:       Windows, macOS, Linux
Locale:   fa-IR, en-US, ar-SA, tr-TR, de-DE
Payment:  Card, Wallet, Gateway, COD
UserTier: Guest, Basic, Premium

IF [Payment] = "COD" THEN [UserTier] <> "Guest";
# generate the pairwise set (PICT defaults to order=2)
pict model.txt > pairs.tsv

# three-way coverage for the high-risk paths
pict model.txt /o:3 > triples.tsv
Pairwise is not a cure-all

Three traps: (1) if the parameters are not truly independent and business logic depends on one specific combination, pairwise may simply never generate it — so pin the critical combinations manually. (2) Pairwise tells you nothing about the expected result; it only produces inputs. The test oracle is still your job. (3) The generated set can differ between runs; commit the generated set to the repository so your tests stay reproducible.

4.6 Error guessing and exploratory testing

Error guessing means using experience to guess where things usually break: empty string, leading/trailing whitespace, an apostrophe in the input, the last day of a leap year, a zero-byte file, double-clicking "pay", closing the browser mid-transaction, hitting back after a successful submit.

Make it professional with a fault attack list: a written, maintained list of failures that have actually happened in your project. Every incident adds a row. This is the single most valuable test document a mature team owns.

Exploratory testing is its structured sibling: under session-based test management, a 60–90 minute timebox with an explicit charter ("investigate cart behaviour when stock drops to zero during checkout") and simultaneous note-taking. Output: defects, questions, and ideas for automated tests.

4.7 Positive and negative testing

  • Positive (happy path): valid input, correct result.
  • Negative (unhappy path): invalid input or bad conditions produce the correct, controlled error.

In a mature service the healthy ratio usually favours the negative cases — there is one happy path and dozens of ways to fail. If a PR contains only happy-path tests, that is exactly where you should stop in code review.

Senior judgment — which technique when:

Requirement shape First technique Complement
Numeric / range input equivalence partitioning 3-value BVA
Multi-condition business rules decision table partitioning per condition
Entity with a lifecycle state transition invalid-transition tests
Configuration / compatibility matrix pairwise pin critical combinations
Complex algorithm, many branches black-box plus branch coverage mutation testing
Dark, undocumented area exploratory testing fault attack list
Here is a requirement: "a password must be 8 to 64 characters and contain at least one digit." Give me the test cases.

First I name the techniques so the interviewer knows this is not improvisation: equivalence partitioning on length and on content, plus 3-value BVA on the 8 and 64 boundaries.

Length: 7 (reject), 8 (accept), 9 (accept), 63 (accept), 64 (accept), 65 (reject), zero/empty (reject), null (reject, and not an NPE).

Content: no digit (reject), exactly one digit (accept), digits only (accept or reject? Here I ask — the requirement is ambiguous, and asking is part of the answer).

Hidden boundaries: is leading and trailing whitespace trimmed? If so, what happens with "8 spaces plus one digit"? Multi-byte Unicode: are 64 emoji 64 characters or 256 bytes? If storage has a byte limit, that is a production bug. Do non-Latin digits count as digits? Character.isDigit says yes — is that what the business meant?

Negative and security angles: the error message must not reveal which rule was violated in a way that helps enumeration, and the password must never reach the logs. I close by noting the requirement itself smells: modern guidance (NIST SP 800-63B) emphasises length and discourages mandatory composition rules — I say that to show I do not only test requirements, I critique them.


Part 5 — The documentation a mature organisation expects

This is where "just a coder" engineers come up short. You do not have to love documentation; you do have to know which question each document answers and what its smallest useful version is.

5.1 The Test Plan

A Test Plan states, for this project or release, what is tested, how, by whom, in which environment, by when, and with what stopping criteria. The old IEEE 829 and the current ISO/IEC/IEEE 29119-3 give you templates, but nobody expects 40 pages. The useful skeleton:

Section Question it answers One-line example
Scope: in / out what is tested and, more importantly, what is not "historical data migration is out of scope"
Test items exactly which version/component "payment-service 3.4.0, gateway 2.1.x"
Approach levels, types, techniques, automation share "unit+integration automated, UAT manual"
Environment environment, data, test doubles "staging with partner sandbox, anonymised data"
Entry criteria when we start "green build + smoke pass + environment ready"
Exit criteria when we are done expanded below
Risks & mitigation what could wreck the plan "unstable partner sandbox → fall back to mocks"
Roles who signs off on what "PO signs off UAT"
Schedule & effort time and person-days
Deliverables what is handed over "execution report, open defect list, load-test report"
Senior judgment — the "out of scope" section is the most valuable page

Every post-release argument that starts with "why wasn't this tested?" is rooted in the absence of an explicit "not tested" list. One plain page saying "in this release, behaviour above 2,000 TPS was not tested and that risk is accepted — approved by X" saves you a thousand meetings. That skill is not an engineering skill, it is risk management, and it is exactly what makes someone senior.

5.2 A good test case

A good test case must be executable by someone who does not know the project, and its outcome must be judgeable. The anatomy:

Field Meaning Example
ID stable, referenceable identifier TC-PAY-014
Title one behavioural sentence "withdrawal above the daily limit goes to manual review"
Requirement ref traceability REQ-PAY-7
Priority P1..P3 by risk P1
Preconditions required state before starting "active user, balance 50,000,000, MFA enabled"
Test data exact data "amount = 30,000,000, destination = ..."
Steps numbered and unambiguous 1) log in 2) request withdrawal 3) submit
Expected result an observable outcome HTTP 202 + status PENDING_REVIEW + ReviewRequested event
Postconditions cleanup / final state "request remains in the review queue"
"Expected: it works correctly" means nothing

Three fatal anti-patterns in test-case writing: (1) a vague expected result ("the page displays correctly") — the outcome must be pass/fail without argument. (2) Chained dependencies where TC-02 is meaningless unless TC-01 ran first; that kills both parallel execution and debugging. (3) Baking brittle UI details into the steps ("click the third button from the left"). Write steps at the level of intent, not pixel coordinates.

5.3 Traceability from requirement to test

A requirements traceability matrix (RTM) links each requirement to the tests that cover it. It answers three questions nothing else does:

  1. Coverage: which requirement has no test at all? (a real hole)
  2. Change impact: if REQ-PAY-7 changes, which tests must be revisited?
  3. Status reporting: "92% of P1 requirements are covered and green" means something to a manager; "84% line coverage" does not.

Diagram: the traceability chain from business need to defect — نمودار: زنجیره‌ی ردیابی از نیاز کسب‌وکار تا نقص.

flowchart LR
  BR[Business need] --> REQ[Requirement / user story]
  REQ --> AC[Acceptance criteria]
  AC --> TC[Test cases]
  TC --> RUN[Test runs]
  RUN --> DEF[Defects]
  DEF -.reopens.-> REQ

In practice you do not need a spreadsheet. Put the requirement id in the test name or an annotation and derive the matrix from code:

@Test
@Tag("REQ-PAY-7")
@DisplayName("REQ-PAY-7 | withdrawal above daily limit goes to manual review")
void withdrawalAboveDailyLimitGoesToManualReview() { /* ... */ }

And if execution data lives in a database (which most test-management tools do), finding uncovered requirements is one query:

SELECT r.req_id,
       r.title,
       COUNT(tc.tc_id) AS test_cases,
       COUNT(*) FILTER (WHERE tr.status = 'PASSED') AS passed
FROM requirement r
LEFT JOIN test_case tc ON tc.req_id = r.req_id
LEFT JOIN LATERAL (
    SELECT status
    FROM test_run
    WHERE test_run.tc_id = tc.tc_id
    ORDER BY executed_at DESC
    FETCH FIRST 1 ROW ONLY
) tr ON TRUE
WHERE r.priority = 'P1'
GROUP BY r.req_id, r.title
HAVING COUNT(tc.tc_id) = 0
    OR COUNT(*) FILTER (WHERE tr.status = 'PASSED') < COUNT(tc.tc_id)
ORDER BY r.req_id;
Dialect difference

The FILTER (WHERE ...) aggregate clause is standard SQL implemented by PostgreSQL; Oracle does not have it, and the portable equivalent there is COUNT(CASE WHEN ... THEN 1 END). Likewise LEFT JOIN LATERAL ... ON TRUE in PostgreSQL corresponds to OUTER APPLY in Oracle (12c and later). More detail lives in the Oracle-versus-PostgreSQL dialects chapter.

5.4 A defect report a developer can act on

A bad defect report costs three days of ping-pong. A good one contains:

  • Title: what, where, under which conditions — in one line. Not "the site is broken".
  • Environment: build version, browser/client, environment (staging/prod), tenant id, exact time with timezone.
  • Reproduction steps: the minimal steps that produce the bug. Minimise them if you can.
  • Actual versus expected result.
  • Evidence: logs with the traceId, screenshots, the raw HTTP response — not a phone photo of a monitor.
  • Reproduction rate: 10 out of 10, or 2 out of 10? That number is critical for concurrency bugs.
  • Severity and priority — see below.
  • Workaround, if one exists.
Severity and priority are not the same thing, and this is a favourite interview trap

Severity is a technical judgment: the impact of the defect on the system. Priority is a business judgment: how soon it must be fixed. All four combinations are real:

  • High severity, high priority: payments are down in production. Fix now.
  • High severity, low priority: a crash in a feature used only by the annual report, which runs in ten months.
  • Low severity, high priority: the company name is misspelled on the landing page and the launch is tomorrow. The system is fine; the reputation is not.
  • Low severity, low priority: a misaligned icon on the settings screen.

If you answer only "severity means how bad and priority means how urgent", you get no credit. Give the four-quadrant example.

5.5 Exit criteria and Definition of Done

What does "testing is finished" mean? If your answer is "we ran out of time", that is not an exit criterion, that is surrender. A defensible exit criterion combines:

  • Coverage: 100% of P1 and P2 requirements have at least one executed test.
  • Execution: ≥95% of planned test cases executed; 100% of P1 executed and passing.
  • Defects: zero open Critical/High defects; Medium defects deferred with written product-owner approval.
  • Non-functional: p95 under the SLO at target load; error rate below threshold; security scan free of Critical findings.
  • Operational readiness: runbook, alerts, dashboard and a tested rollback path.

Part 6 — The pyramid, the ice-cream cone, and where contract tests fit

Quality control in a factory

In a good factory each part is measured as it is made (cheap and immediate), then subassemblies are assembled and tested, and finally a few complete units get a functional test. Now picture a factory that measures nothing and only switches on the finished product: when the lamp does not light, it must disassemble the entire machine to find the burnt resistor. That factory has an "ice-cream cone".

Diagram: a healthy pyramid versus the ice-cream cone anti-pattern — نمودار: هرم سالم در برابر مخروط بستنی.

flowchart TB
  subgraph Pyramid["Healthy pyramid"]
    P1["E2E: few, slow, high value"] --> P2["Integration / contract: some"]
    P2 --> P3["Unit: many, fast, cheap"]
  end
  subgraph Cone["Ice-cream cone"]
    C1["Manual testing: huge"] --> C2["E2E UI: many"]
    C2 --> C3["Integration: few"]
    C3 --> C4["Unit: almost none"]
  end

The pyramid's logic is economic, not ideological: the higher you go, the slower, more brittle, more expensive to maintain and vaguer at pinpointing causes each test becomes. So build fewer of them — but do not delete them, because only the top of the pyramid tells you the whole system actually works.

Layer Target speed What it proves Maintenance cost
Unit < 10ms logic, branches, boundaries low
Integration (Testcontainers) 0.1–2s mapping, SQL, serialisation, transactions medium
Contract < 1s provider/consumer compatibility low to medium
System / API E2E 1–30s real flow within or across services high
UI E2E 5–60s critical user journeys very high
Manual / exploratory minutes unknowns, UX, new risks recurring

Clinical symptoms of the ice-cream cone — If you see these, your pyramid is upside down: the build takes over 20 minutes and nobody waits for it; diagnosing a failure requires watching a UI test video; the team has a "stabilisation week" before every release; most bugs are found by manual QA rather than CI; and "just re-run it, maybe it goes green" has become a normal sentence.

Contract testing — the missing layer

The microservices problem: each service's unit tests are green, yet putting them together breaks, because the consumer assumed amount was a string and the provider made it a number. The naive answer is "E2E with all services running" — slow, brittle, and requiring the whole world to boot in CI.

Contract testing says: record the contract between two services in an executable file, then test each side separately against that contract.

  • Consumer-driven (Pact): the consumer writes its expectations, producing a pact file; the provider verifies it in its own CI. The practical prerequisite is a shared broker to exchange contracts.
  • Provider-driven (Spring Cloud Contract): the provider writes the contract in a DSL, which generates both the provider test and the stub the consumer uses in its own tests.
// src/test/resources/contracts/shouldReturnAccountBalance.groovy
Contract.make {
    description "should return balance for an existing account"
    request {
        method GET()
        url "/api/accounts/ACC-1001/balance"
        headers { accept(applicationJson()) }
    }
    response {
        status OK()
        headers { contentType(applicationJson()) }
        body(
            accountId: "ACC-1001",
            balance: 250000,
            currency: "IRR"
        )
        bodyMatchers {
            jsonPath('$.balance', byType())
            jsonPath('$.currency', byRegex('[A-Z]{3}'))
        }
    }
}
<plugin>
  <groupId>org.springframework.cloud</groupId>
  <artifactId>spring-cloud-contract-maven-plugin</artifactId>
  <version>4.3.0</version>
  <extensions>true</extensions>
  <configuration>
    <baseClassForTests>com.example.contract.ContractBase</baseClassForTests>
    <testFramework>JUNIT5</testFramework>
  </configuration>
</plugin>

Senior judgment — contract tests do not replace E2E, they replace 90% of it — Contract tests tell you the message shapes are compatible. They still do not tell you the business flow is right (does an order actually ship after payment?). So keep a few E2E tests on the revenue paths and delegate the rest of the compatibility surface to contracts. The question that scores points in an interview: "can the provider deploy without breaking any consumer? If you do not know the answer, you do not have contract tests."

How do you make sure a change in one service does not break another, without booting the whole system?

With contract testing. Each consumer records its real expectations as an executable contract; that contract is published to a broker; and the provider's pipeline verifies it against its real implementation on every build. If a field is removed or retyped, the provider's build turns red — before deployment, without any other service running.

The more important part is the compatible-change rule: adding an optional field is compatible; removing a field, changing a type, tightening a constraint or changing the meaning of a value is not. For an incompatible change I use versioning or the expand-and-contract pattern: add the new field first, let consumers migrate, then remove the old one.

And for the loop to work in practice I need a "can-i-deploy" gate: a deployment check that asks whether the version I am about to ship has been verified against the versions currently running for every consumer.


Part 7 — Automation strategy

7.1 What to automate and what not to

Automation is an investment: build cost plus maintenance cost versus the saving on future runs. If a test runs twice a year and takes three manual minutes, automating it is a loss.

Automate it if… Do not automate it if…
it runs often (every PR, nightly) it is one-off, or the requirement is still fluid
it has a deterministic, machine-readable outcome it needs human judgment (aesthetics, UX, copy)
it covers high risk or a revenue path it is a low-risk, rarely used corner
its data and environment are controllable it depends on an uncontrollable external service
doing it by hand is tedious and error-prone it is exploratory and its value is human creativity
The anti-pattern: "we automate everything"

Teams whose goal is "100% automation" usually end up with a large, slow, flaky suite nobody trusts — the worst possible outcome, because you pay the maintenance and get no signal. A small trustworthy suite beats a large suspicious one every single time. Use this health metric: "when the build goes red, what fraction of the time is it a real bug?" If it is under 90%, your problem is not missing tests, it is excess noise.

7.2 Managing flaky tests

A flaky test passes sometimes and fails other times with no code change. The real danger is not the failure itself; it is that the team learns to ignore red, and the day a genuine bug turns the build red, nobody reacts.

Root cause Symptom Cure
Dependence on wall-clock time fails around midnight or on slow machines inject a Clock, use Clock.fixed(...)
Concurrency and races fails with -parallel explicit synchronisation, Awaitility instead of Thread.sleep
Execution order / shared state passes alone or with a different seed isolate data, reset static state
Fixed waits on UI or network fails on a busy CI agent condition-based (explicit) waits
Shared data between tests fails under parallel execution unique data per test, separate schema
Unstable external resources scattered, patternless failures test double at the boundary

The process that works:

  1. Measure: derive the flakiness rate from CI history (same test, same commit, different result).
  2. Time-boxed quarantine: pull the flaky test out of the main gate but keep running it in a separate job, with an owner and a deadline. Quarantine without a deadline is deletion.
  3. Fix the root cause, not with a retry.
  4. Budget: for example "at most 5 tests in quarantine"; when it fills, new work stops until it is cleaned.

retryFailedTests is a painkiller, not a cure — Adding automatic retries is tempting and turns the build green in the short term. It has two costs: it hides rare real bugs (races in production code, not in the test) and it multiplies the worst-case run time. If you must retry, at least report it: a test that passed only on retry should be flagged in the report and recorded on a flakiness dashboard, otherwise you have thrown information away.

7.3 Test data management

Three strategies, in order of preference:

  1. Built inside the test (test data builder) — the best case. Each test creates its own data, so it is independent and readable.
  2. A small shared fixture — for immutable reference data (currency list, bank codes).
  3. A copy of production (anonymised) — only for load testing and data migration. It carries compliance risk (GDPR and local equivalents) and a high maintenance cost.

The builder pattern with sensible defaults transforms test readability: each test states only what matters to it.

public final class AccountBuilder {
    private String id = "ACC-" + UUID.randomUUID();
    private long balance = 1_000_000L;
    private boolean active = true;
    private boolean mfaEnabled = true;

    public static AccountBuilder anAccount() { return new AccountBuilder(); }

    public AccountBuilder withBalance(long balance) { this.balance = balance; return this; }
    public AccountBuilder inactive() { this.active = false; return this; }
    public AccountBuilder withoutMfa() { this.mfaEnabled = false; return this; }

    public Account build() { return new Account(id, balance, active, mfaEnabled); }
}

// the resulting readability: only the meaningful difference is visible
var poorAccount = anAccount().withBalance(0).build();

To reset the database between integration tests, TRUNCATE is usually faster than row-by-row deletion — but the identity-reset syntax differs between the dialects:

-- empty the tables and reset identity in one statement
TRUNCATE TABLE payment, account RESTART IDENTITY CASCADE;

Do not "anonymise" production data — synthesise it — Stripping names and phone numbers is not enough: date of birth plus postcode plus recent transactions is usually sufficient for re-identification. If you must use real data, use genuine pseudonymisation (a stable mapping with a separately held key), keep the environment at the same confidentiality level as production, and log every access. The best option is to write a synthetic data generator that mimics production's statistical distribution without containing a single real record.

7.4 Environment strategy and test doubles

The precise test double vocabulary (the common taxonomy) — useful in interviews:

Type Definition Example
Dummy fills a parameter slot, never used a null object
Stub returns canned answers when(repo.find(id)).thenReturn(acc)
Spy real object that records interactions counting invocations
Mock carries behavioural expectations and verifies them verify(gateway).charge(...)
Fake a real but simplified implementation in-memory repository

And for environments:

  • Hermetic / ephemeral: each build creates and destroys its own environment (Testcontainers, a temporary Kubernetes namespace). More expensive to set up, but free of interference and drift.
  • Shared staging: cheap, but three teams work on it simultaneously and failures contaminate each other.
  • External partner services: if a stable sandbox exists, use it — but not in the main gate. Run the main path against a contract-based double (WireMock or a mock server) and run a nightly job against the real sandbox to detect drift.

7.5 Running tests in CI

Pipeline detail belongs to the CI/CD chapter; here we only cover the test layering:

<!-- surefire: fast unit tests in the test phase -->
<plugin>
  <groupId>org.apache.maven.plugins</groupId>
  <artifactId>maven-surefire-plugin</artifactId>
  <configuration>
    <excludedGroups>integration,slow</excludedGroups>
    <parallel>classes</parallel>
    <threadCount>4</threadCount>
  </configuration>
</plugin>

<!-- failsafe: integration tests in the integration-test phase -->
<plugin>
  <groupId>org.apache.maven.plugins</groupId>
  <artifactId>maven-failsafe-plugin</artifactId>
  <configuration>
    <groups>integration</groups>
  </configuration>
  <executions>
    <execution>
      <goals>
        <goal>integration-test</goal>
        <goal>verify</goal>
      </goals>
    </execution>
  </executions>
</plugin>
# the fast developer loop: unit only
mvn -q test

# the merge gate: unit + integration, failing at verify
mvn -q verify

# JUnit 5 parallel execution (src/test/resources/junit-platform.properties)
# junit.jupiter.execution.parallel.enabled = true
# junit.jupiter.execution.parallel.mode.default = concurrent
Senior judgment — layer by time budget, not by label

Instead of the philosophical argument "is this a unit or an integration test", define a time budget: the pre-commit stage under 90 seconds, the PR gate under 10 minutes, the nightly pipeline uncapped. Every test lives in the fastest stage it fits into without blowing the budget. That rule ends the argument and rewards the right behaviour: whoever writes a slow test now has a personal incentive to make it fast.

How do you test an event-driven flow? The service consumes a message, updates the database and publishes another event.

I break it into three layers so each gets the cheapest possible test.

The logic layer: I keep the function "incoming event + current state → new state + outgoing events" pure and I/O-free, and cover it with fast unit tests; here I apply equivalence partitioning and state transition techniques.

The adapter layer: with a real broker in Testcontainers I test that deserialisation, offset commit and outgoing publication work. The most important thing here is failure behaviour: a malformed message must go to the DLQ, not put the consumer into an infinite redelivery loop.

The contract layer: I lock the message schema with a contract test or a schema registry so an incompatible change is caught before deployment.

Three specifics of this world that mark a senior: (1) never write Thread.sleep — use conditional waiting (for example Awaitility with a timeout and a poll interval), otherwise the test is either flaky or needlessly slow. (2) Test idempotency deliberately: send the same message twice and assert the side effect happened once; with at-least-once delivery, duplication is a certainty, not an exception. (3) Test ordering and partition keys: two related events with the same key must retain their order; this is where production bugs are born. Protocol detail lives in the Kafka and RabbitMQ chapters; here I only cover the testing discipline.

7.6 Code coverage: a signal, not a target

Coverage only tells you which lines or branches were executed — not that they were checked. A test with no assertion still produces coverage.

The useful flavours:

  • Line/statement coverage — the weakest.
  • Branch coverage — both sides of every if were taken. The meaningful minimum.
  • MC/DC — for each atomic condition inside a compound condition, show it independently affects the outcome. Mandatory in safety-critical domains.
<plugin>
  <groupId>org.jacoco</groupId>
  <artifactId>jacoco-maven-plugin</artifactId>
  <version>0.8.13</version>
  <executions>
    <execution><goals><goal>prepare-agent</goal></goals></execution>
    <execution>
      <id>check</id>
      <phase>verify</phase>
      <goals><goal>check</goal></goals>
      <configuration>
        <rules>
          <rule>
            <element>BUNDLE</element>
            <limits>
              <limit>
                <counter>BRANCH</counter>
                <value>COVEREDRATIO</value>
                <minimum>0.70</minimum>
              </limit>
            </limits>
          </rule>
        </rules>
      </configuration>
    </execution>
  </executions>
</plugin>
When coverage becomes a target, it stops being a measure

If management mandates "80% minimum", the team reaches it in one afternoon — with tests that call everything and assert nothing, or with @Generated sprinkled on classes full of logic. The number goes up; the quality does not. There are exactly two correct uses of coverage: (1) the coverage diff on new code in a PR, and (2) finding zero-coverage areas nobody realised were untested. The project-wide total is nearly meaningless.

7.7 Mutation testing — testing your tests

If coverage says "this line executed", mutation testing asks "if I deliberately break this line, does any test go red?"

How it works: the tool creates mutants in the bytecode — turning < into <=, return true into return false, removing a method call — and runs the tests. If a test fails, the mutant was killed. If everything stays green, the mutant survived, which means that behaviour is effectively unguarded.

Mutation score = killed mutants ÷ total mutants. It is a far more honest number than coverage.

<plugin>
  <groupId>org.pitest</groupId>
  <artifactId>pitest-maven</artifactId>
  <version>1.19.1</version>
  <configuration>
    <targetClasses><param>com.example.payment.domain.*</param></targetClasses>
    <targetTests><param>com.example.payment.domain.*Test</param></targetTests>
    <mutationThreshold>75</mutationThreshold>
    <timestampedReports>false</timestampedReports>
  </configuration>
  <dependencies>
    <dependency>
      <groupId>org.pitest</groupId>
      <artifactId>pitest-junit5-plugin</artifactId>
      <version>1.2.2</version>
    </dependency>
  </dependencies>
</plugin>
# full run over the module
mvn -q test-compile org.pitest:pitest-maven:mutationCoverage

# incremental mode: only code changed against the base branch (ideal for PRs)
mvn -q test-compile org.pitest:pitest-maven:scmMutationCoverage \
    -Dinclude=ADDED,MODIFIED
# HTML report: target/pit-reports/index.html

Senior judgment — do not run mutation testing over the whole project — PIT is slow because it re-runs tests per mutant. The practical strategy is to run it only on the domain core (money, pricing, authorisation, calculations) and only on the code changed in the PR. Even if you enforce no threshold at all, reading one PIT report over your domain core is a sobering experience: it usually surfaces a handful of surviving mutants in exactly the places that are scary in production.

Why is 90% code coverage not a guarantee of quality?

Because coverage measures execution, not verification. Tomorrow I could push coverage to 100% by writing tests that call every method and assert nothing, without catching a single bug. Even with assertions, line coverage says nothing about boundary values, condition combinations or unimplemented requirements — code that was never written has zero lines and therefore zero uncovered lines.

What I actually look at: branch coverage on new code in a PR (not the project total), zero-coverage areas inside the domain core, and above all the mutation score on critical logic. Mutation testing effectively asks "if I break the code, does anyone notice?" — which is the question we really want answered.

In practice I make coverage a soft gate: a drop in a PR raises a warning and a conversation rather than an automatic failure, and I spend the saved effort on reviewing the quality of assertions in code review.

You have a flaky test in CI that is driving the team mad. How do you handle it?

First measure: from CI history I extract how often and on which commits it failed; if the same commit produced both green and red, it is definitively flaky and not a regression.

Then contain quickly: I quarantine the test — out of the main gate, but still running in a separate job so we do not lose its signal. I file an issue with an owner and a deadline. I do not let quarantine become a graveyard.

Then root cause: I run it many times, under load, with a different execution order and a different seed. It is almost always one of these: time (Instant.now() instead of an injected Clock), concurrency (Thread.sleep instead of conditional waiting), shared state between tests, or a dependency on an external service.

The point that usually scores: sometimes a flaky test is actually discovering a real concurrency bug in production code. So before I "fix" the test, I make sure the problem is in the test and not in the system. And culturally I keep "red means stop" alive; a team that gets used to red effectively has no CI.


Part 8 — Non-functional testing: performance, load, stress, endurance

The bridge and the truck

What does it mean that a bridge "works"? If one car crosses it, it works. What about 500 cars at once? What about an 80-tonne truck? What about ten years of daily traffic without maintenance? What about an earthquake? Each of those is a different kind of test, and "it works" does not mean the same thing for any two of them.

8.1 Precise vocabulary — teams routinely misuse these

Test type Question Load shape Expected outcome
Performance (umbrella) how fast and how much capacity? varies characterising behaviour
Load does it hold the SLO under expected load? target load, steady p95/p99 and error rate within threshold
Stress where is the breaking point, and how does it break? ramp to collapse graceful degradation, not collapse
Soak / endurance does it degrade over time? moderate load, 4–72 hours no memory/connection/disk leak
Spike how does it react to a sudden surge? fast surge then return absorb or throttle, and recover
Capacity / breakpoint maximum throughput with a healthy SLO? stepped until SLO breach a capacity number for planning
Scalability what happens with twice the resources? fixed load, varying resources the real scaling factor
Volume what happens with large data? bulk data queries still use their indexes

A spike test measures recovery, not just survival — Most teams look only at "did the system survive?". The more important question is: after the spike passes, how long until it returns to normal? Systems with unbounded queues may show no errors during the surge yet still be chewing through expired requests half an hour later. The correct pattern is fast rejection (load shedding) plus bounded queues — covered in the resilience chapter.

8.2 SLO first, then the test

A load test without a pass criterion is just colourful chart generation. So before anything else:

  • SLI (indicator): what you measure. Example: "the percentage of POST /payments requests that receive a 2xx/4xx response in under 800ms".
  • SLO (objective): the target value for that indicator over a window. Example: "99.5% over any 30 days".
  • SLA (agreement): the contractual commitment with penalties. Usually looser than the internal SLO.
  • Error budget: 1 minus the SLO. At 99.5%, the monthly budget is roughly 3 hours 36 minutes.
The average lies; read the percentiles

If 99 requests take 10ms and one takes 5 seconds, the average is 60ms and looks excellent — while a real user waited five seconds. Always report p50/p90/p95/p99 and max. Two subtleties: (1) percentiles are not additive; averaging the p95 of several time buckets is mathematically meaningless — you must recompute from the raw histograms. (2) On a page that makes 20 service calls, the probability that a user hits the p99 tail at least once is very high; p99 is not a "rare edge case", it is the daily experience of a slice of your users.

8.3 The workload model: open versus closed

This concept is the line between a toy load test and a credible one.

  • Closed model: a fixed number of virtual users, each looping "send request, get response, think, send again". If the system slows down, the arrival rate automatically drops. This models systems with a bounded user population well (an internal app with 200 employees).
  • Open model: users arrive at an arrival rate independent of how slow the system has become. This models real internet traffic — the world does not wait for you.
Coordinated omission

This is the biggest statistical trap in load testing. In a closed model, when the system stalls the virtual user stalls too and simply never sends the following requests. The result: exactly the requests that should have been recorded as slow are never measured at all, and your statistics become fantasy-grade optimistic. The fix is to use arrival-rate executors — in k6 that means constant-arrival-rate or ramping-arrival-rate, and in Gatling injectOpen(constantUsersPerSec(...)). If your load test is defined as "500 virtual users" with no rate anywhere in it, you have probably fallen into this trap.

Little's Law is your mental tool for converting between these:

L = λ × W

That is, "requests concurrently in the system = arrival rate × average time in system". At 200 requests per second with a 250ms average response, about 50 requests are in flight on average — which maps directly onto your connection-pool and thread-pool sizing. If the pool holds 20, a queue forms and p99 explodes even while the CPU sits idle.

8.4 Modelling real load

Synthetic load ("everyone just hits GET /") produces worthless results. Build the workload model from real data: derive the endpoint distribution and the peak hour from access logs or an events table.

SELECT endpoint,
       COUNT(*) AS hits,
       ROUND(100.0 * COUNT(*) / SUM(COUNT(*)) OVER (), 2) AS pct,
       ROUND(COUNT(*) / 3600.0, 2) AS rps_in_peak_hour
FROM access_log
WHERE ts >= TIMESTAMP '2026-08-05 12:00:00'
  AND ts <  TIMESTAMP '2026-08-05 13:00:00'
GROUP BY endpoint
ORDER BY hits DESC
FETCH FIRST 15 ROWS ONLY;

From the output, build a load profile: for example 60% search, 25% detail view, 10% add-to-cart, 5% checkout — then implement those same ratios in the scenario. Add realistic think time between steps, without which your concurrency pattern is entirely unrealistic.

Diagram: load-test topology and the signals to watch at each hop — نمودار: توپولوژی تست بار و سیگنال‌هایی که در هر گره باید دید.

flowchart LR
  LG[Load generators] --> LB[Load balancer]
  LB --> APP[App instances]
  APP --> DB[(Database)]
  APP --> CACHE[(Cache)]
  APP --> EXT[External / stubbed services]
  MON[Metrics: RED + USE] -.scrapes.- APP
  MON -.scrapes.- DB
  MON -.scrapes.- LG

8.5 Three tools: JMeter, Gatling, k6

Criterion Apache JMeter Gatling Grafana k6
Test language GUI + XML (.jmx) Java/Kotlin/Scala DSL JavaScript (ES modules)
Execution model one thread per virtual user Akka-based, non-blocking Go plus a JS engine (goja)
Resource cost per VU high low low
Reviewable in Git hard (bulky XML) excellent excellent
Protocols very broad (JDBC, JMS, LDAP, FTP…) HTTP, WebSocket, JMS, gRPC HTTP, WebSocket, gRPC, browser
Strength plugin ecosystem, non-programmer teams great reports, typed DSL for Java teams DevOps integration, thresholds as a CI gate
Weakness memory at high VU counts, hard diffs smaller community than JMeter requires JS, some features cloud-only
Senior judgment — pick the tool that matches who maintains it

All three can put an API under load. The deciding factor is who maintains the test: if the team is a Java team and the load test should live next to the code, Gatling with the Java DSL is the natural pick. If the team is DevOps-oriented and you want the load test to be an ordinary pipeline stage with an automatic pass criterion, k6 with thresholds has the least friction. If you must load JDBC, JMS or legacy enterprise protocols, JMeter is still unmatched. In an interview the answer to "which is better?" is never a tool name; it is the selection criterion.

k6 — least friction for CI

k6 is a Go binary; you write the test in JavaScript and thresholds directly determine the process exit code — so the CI gate is free.

// payments-load.js
import http from 'k6/http';
import { check, sleep } from 'k6';
import { Trend, Rate } from 'k6/metrics';

const paymentLatency = new Trend('payment_latency', true);
const businessErrors = new Rate('business_errors');

export const options = {
  scenarios: {
    // open model: arrival rate independent of how slow the system gets
    steady_traffic: {
      executor: 'ramping-arrival-rate',
      startRate: 20,
      timeUnit: '1s',
      preAllocatedVUs: 100,
      maxVUs: 500,
      stages: [
        { target: 20,  duration: '1m' },   // warm-up
        { target: 200, duration: '3m' },   // ramp to target
        { target: 200, duration: '10m' },  // steady state
        { target: 0,   duration: '1m' },   // ramp down
      ],
    },
  },
  thresholds: {
    'http_req_failed': ['rate<0.01'],
    'http_req_duration{name:createPayment}': ['p(95)<800', 'p(99)<2000'],
    'business_errors': ['rate<0.005'],
    'checks': ['rate>0.99'],
  },
};

export default function () {
  const payload = JSON.stringify({ accountId: 'ACC-1001', amount: 250000, currency: 'IRR' });
  const res = http.post('https://staging.example.internal/api/payments', payload, {
    headers: { 'Content-Type': 'application/json' },
    tags: { name: 'createPayment' },   // stable tag for metric grouping
  });

  const ok = check(res, {
    'status is 201': (r) => r.status === 201,
    'has payment id': (r) => r.json('paymentId') !== undefined,
  });

  paymentLatency.add(res.timings.duration);
  businessErrors.add(!ok);
  sleep(Math.random() * 2 + 1);   // realistic think time
}
# run locally
k6 run payments-load.js

# override options from the command line (for a quick smoke run)
k6 run --vus 5 --duration 30s payments-load.js

# emit JSON for later analysis
k6 run --out json=results.json payments-load.js

# scaffold a new script
k6 new payments-load.js

Why tags matters — If the URL contains a variable id (/api/payments/9f3a...), k6 counts each one as a separate name and your metrics shatter. With tags: { name: 'createPayment' } you group them under one stable name and can attach a threshold to it — exactly like http_req_duration{name:createPayment} above.

Gatling — when the team is a Java team

Since 3.7 Gatling has a full Java DSL; you do not need Scala. The test lives beside the project code in src/test/java and runs through the Maven plugin.

package simulations;

import io.gatling.javaapi.core.*;
import io.gatling.javaapi.http.*;

import java.time.Duration;

import static io.gatling.javaapi.core.CoreDsl.*;
import static io.gatling.javaapi.http.HttpDsl.*;

public class PaymentSimulation extends Simulation {

    HttpProtocolBuilder httpProtocol = http
        .baseUrl("https://staging.example.internal")
        .acceptHeader("application/json")
        .contentTypeHeader("application/json")
        .shareConnections();

    FeederBuilder<String> accounts = csv("accounts.csv").random();

    ScenarioBuilder createPayment = scenario("Create payment")
        .feed(accounts)
        .exec(
            http("createPayment")
                .post("/api/payments")
                .body(StringBody("""
                    {"accountId":"#{accountId}","amount":250000,"currency":"IRR"}
                    """))
                .check(status().is(201))
                .check(jsonPath("$.paymentId").saveAs("paymentId"))
        )
        .pause(Duration.ofSeconds(1), Duration.ofSeconds(3))   // think time
        .exec(
            http("getPayment")
                .get("/api/payments/#{paymentId}")
                .check(status().is(200))
        );

    {
        setUp(
            createPayment.injectOpen(          // open model: arrival rate
                nothingFor(Duration.ofSeconds(5)),
                rampUsersPerSec(10).to(200).during(Duration.ofMinutes(3)),
                constantUsersPerSec(200).during(Duration.ofMinutes(10))
            )
        )
        .protocols(httpProtocol)
        .assertions(
            global().responseTime().percentile3().lt(800),   // p95 by default
            global().failedRequests().percent().lt(1.0),
            details("createPayment").responseTime().percentile(99.0).lt(2000)
        );
    }
}
# run with the Maven plugin (io.gatling:gatling-maven-plugin)
mvn gatling:test -Dgatling.simulationClass=simulations.PaymentSimulation

# HTML report at target/gatling/<simulation>-<timestamp>/index.html

Open versus closed model in GatlinginjectOpen(...) is the open model: rampUsersPerSec/constantUsersPerSec control the arrival rate. injectClosed(...) is the closed model: constantConcurrentUsers/rampConcurrentUsers hold the number of concurrent users fixed. For public APIs open is almost always right; keep closed for systems with a bounded user population or when simulating a fixed worker pool.

JMeter — protocol breadth and non-GUI execution

JMeter's stable release is 5.6.3 and it requires Java 8+ (Java 17+ recommended). The golden rule: the GUI is for building and debugging the test only; real runs are always non-GUI.

# non-GUI run plus HTML dashboard generation
jmeter -n -t payments.jmx -l results.jtl -e -o report/

# override parameters from the command line (read in the .jmx via ${__P(threads,50)})
jmeter -n -t payments.jmx -Jthreads=200 -Jrampup=120 -Jduration=600 \
       -l results.jtl -e -o report/

# distributed run: the controller drives several server machines
jmeter -n -t payments.jmx -R 10.0.0.11,10.0.0.12 -l results.jtl

# generate the dashboard from an existing results file
jmeter -g results.jtl -o report/

Three JMeter settings everybody gets wrong — (1) Running the GUI for a real test — the GUI consumes resources itself and distorts the results. (2) Leaving graphical listeners ("View Results Tree") in the executed plan — they hold every response in memory and will OOM a long run; remove or disable them for non-GUI execution. (3) Ignoring the default heap — for high load you must raise HEAP in the launcher script, otherwise the real bottleneck becomes JMeter's own garbage collector rather than the system under test.

8.6 Reading the results

Four families of numbers, and the order in which to look at them:

  1. Throughput — the actual completion rate (req/s). If actual is below target, either the load generator ran out of steam or the system is saturated.
  2. Latency — p50 for "the usual feel", p95/p99 for "the common worst experience", max for "what went wrong". Always look at a histogram or a time series, not a single number.
  3. Error rate — and the breakdown by error type: a 503 (capacity rejection), a client timeout and a 500 (a bug) are worlds apart.
  4. Saturation — resources: CPU, memory, thread-pool queue, connection-pool occupancy, disk IO, bandwidth. Here USE (utilisation, saturation, errors) for resources and RED (rate, errors, duration) for services are the standard frameworks.

The pattern to recognise: up to a point, more load means more throughput and roughly flat latency. Past the knee of the curve, throughput flattens while latency rises exponentially — that means queueing has begun. If throughput actually falls past the knee, the system is collapsing (thrashing, continuous GC, timeout-and-retry storms).

Ten common load-testing mistakes

(1) Measuring on an environment half the size of production and extrapolating linearly. (2) Ignoring warm-up: JIT has not compiled, caches are cold, pools are empty — discard the first minutes. (3) Using one user or one fixed id so everything sits in cache and the numbers become fiction. (4) Saturating the load generator itself (CPU, ephemeral ports, bandwidth) and blaming the system. (5) No think time, producing an unrealistic concurrency pattern. (6) Happy path only, with no errors anywhere in the load. (7) No server-side monitoring — a test without server metrics tells you "it is slow", never "why". (8) A test database at one hundredth of production volume (the execution plans differ completely). (9) Real external services that hit rate limits and poison the result. (10) No pre-agreed pass criterion, which leads to hunting for a narrative that makes the result look acceptable after the fact.

How would you design a load test for a new payments API?

First the goal and pass criterion: I agree an SLO with the business, for example "p95 under 800ms and error rate under 1% at a peak load of 200 transactions per second". Without that number the test only produces charts.

Then the workload model from real data: endpoint distribution and ratios from the peak-hour logs, a six-month growth factor, and the daily pattern. I add realistic think time and use the open (arrival-rate) model so I do not fall into coordinated omission.

Then environment and data: an environment equal to production, or with a documented and fixed ratio; a database with volume close to production, because execution plans on a small table are completely different; varied data per virtual user so caching does not fake the result; and partner services replaced by contract-based stubs with realistic latency and error rates.

Then execution: first a smoke run with a few users to validate the script, then warm-up, then a stepped ramp to target, then steady state for at least 15 minutes, then a capacity test up to the breaking point, and finally an overnight soak for leaks.

And throughout, server-side monitoring: RED for the service, USE for resources, GC statistics, connection-pool occupancy and slow queries. I close the analysis with an identified bottleneck and a proposed action — not with a screenshot of a graph.

Mid load test, p99 has exploded but CPU is only at 30%. What is happening?

Low CPU with high latency almost always means queueing, not a shortage of compute. My investigation has an order.

First bounded pools: the database connection pool. If it is small, threads queue waiting to acquire a connection; the "time waiting for a connection" metric shows this instantly. Little's Law says the same thing: given the arrival rate and the time in system, the minimum pool size is determined.

Then locks and contention: a lock on a hot row, a SELECT ... FOR UPDATE on a shared counter, or a synchronized block on the hot path. A thread dump taken at peak reveals this in seconds.

Then slow dependencies: a downstream service or a slow query holding threads. Here the absence of timeouts and bulkheads lets one slow dependency freeze the entire service.

Then GC and memory: long pauses raise latency without sustained CPU load.

And finally the things everybody forgets: saturation of the load generator itself, ephemeral port exhaustion, and blocking on DNS or the TLS handshake. The senior-level observation is that low CPU is usually good news: it means the bottleneck is structural and fixable with a configuration change or a concurrency improvement, not by adding machines.


Part 9 — Security testing: the minimum expected from a backend engineer

Vulnerability detail belongs to the application-security and OWASP chapter; here we only place each tool in the lifecycle, because that is what interviews probe.

Approach What it sees When it runs Weakness
SAST (static analysis) source code without executing it every PR many false positives, blind to runtime logic
SCA (composition analysis) third-party dependencies and CVEs every build + nightly CVEs without real exploitability; needs triage
Secret scanning keys/tokens in code and history pre-commit + CI pattern-based; misses unusual secrets
DAST (dynamic analysis) the running application from outside nightly against staging needs a live environment; coverage depends on crawling
IAST / RASP inside the runtime during tests alongside functional tests requires an agent
Pen test / red team a real attack chain driven by humans periodic expensive and point-in-time

Diagram: where each security test type sits in the pipeline — نمودار: جای هر نوع تست امنیت در خط لوله.

flowchart LR
  DEV[Commit] --> SEC1[Secret scan + SAST]
  SEC1 --> BUILD[Build]
  BUILD --> SCA[Dependency / SCA scan]
  SCA --> IMG[Container image scan]
  IMG --> DEPLOY[Deploy to staging]
  DEPLOY --> DAST[DAST baseline scan]
  DAST --> PROD[Production]
  PROD -.periodic.-> PEN[Penetration test]
# SCA of Maven dependencies with OWASP Dependency-Check
mvn org.owasp:dependency-check-maven:12.1.3:check \
    -DfailBuildOnCVSS=7 -Dformats=HTML,SARIF

# image and filesystem scanning with Trivy
trivy image --severity HIGH,CRITICAL --exit-code 1 myapp:3.4.0
trivy fs --scanners vuln,secret,misconfig .

# a DAST baseline scan with OWASP ZAP (containerised)
docker run --rm -v "$(pwd)":/zap/wrk/:rw \
  ghcr.io/zaproxy/zaproxy:stable zap-baseline.py \
  -t https://staging.example.internal -r zap-report.html

Failing the build on every CVE teaches the team to bypass the gate — If every CVE above CVSS 7 turns the build red, within two weeks somebody adds a flag that skips the scan and your security posture becomes zero. A workable policy: CVEs exploitable through a real code path fail the build; everything else is filed as an issue with a time-based SLA. That is why VEX (a statement that a given CVE is not exploitable in your product) and reachability analysis have become important. And the real secret: every entry in suppressions.xml must carry an expiry date and a reason, or the file turns into a landfill.


Part 10 — Accessibility and usability

Two test types backend engineers ignore and that are legally mandated in large organisations, especially in the public sector.

  • Accessibility (a11y): can someone using a screen reader, no mouse, or with low vision complete the task? The reference is WCAG 2.2 with three conformance levels A/AA/AAA — the practical industry target is usually AA.
  • Usability: can a user finish the task without training? Measured by testing with a handful of real users, not by team opinion.
# automated a11y scan of a page with axe (open source)
npx @axe-core/cli https://staging.example.internal --tags wcag2a,wcag2aa

Automated tools only do part of the job — Industry analyses consistently show that automated tools catch only a portion — roughly a third to a half — of WCAG issues: missing alt, low contrast, missing labels. But "is this alternative text meaningful?" and "is the focus order logical?" require human judgment. The minimum every team should do: navigate the critical forms end to end with the keyboard only. That single exercise finds more defects than a month of automated scanning.


Part 11 — TDD and BDD, honestly

11.1 The red-green-refactor loop

Climbing with a rope

TDD is like clipping in before every move. You go up a little, secure the rope, then make the next move. It looks slow right up until your foot slips. Refactoring without tests is climbing without a rope: faster, until it is not.

Diagram: the TDD cycle and the rule of each step — نمودار: چرخه‌ی TDD و قانون هر مرحله.

stateDiagram-v2
  [*] --> RED
  RED --> GREEN: write the simplest code that passes
  GREEN --> REFACTOR: tests stay green
  REFACTOR --> RED: next small behaviour
  note right of RED
    Test must fail for the RIGHT reason
  end note
  note right of REFACTOR
    Change structure, never behaviour
  end note

The rules that separate real TDD from "writing tests after the code":

  1. Write the test first and confirm it fails for the right reason (not because of a compile error or a typo).
  2. Write the simplest code that makes it pass, even if it is naive.
  3. Refactor only while green — and add no new behaviour during the refactor step.
  4. Keep the steps small; if you have been red for more than a few minutes, the step was too big.
Senior judgment — when TDD helps and when it does not

It helps when: the logic is well defined and expressible (calculations, rules, parsers, state machines); the requirement is reasonably stable; or you are fixing a bug (write the reproducing test first). It does not help when: you are still exploring and the design changes hourly (spike, then throw it away); the work is essentially integration and configuration (writing a test first against an unknown external API is guesswork); or the UI is exploratory. The mature interview answer is not "always TDD"; it is "yes for domain logic; for adapters I spike first to learn the API shape, then lock it down with tests".

Tests that lock in the implementation — If your tests mock every internal collaborator and verify call ordering, you have locked in structure rather than behaviour. The symptom: any refactoring — even with unchanged behaviour — breaks ten tests. That is the exact opposite of the purpose of testing. The working rule: keep mocks for external boundaries (payments, email, queues); use real objects or fakes for internal collaborators. A test should state what happened, not how.

11.2 BDD and Gherkin

BDD (behaviour-driven development) is a conversation before it is a tool: business, development and testing build concrete examples together (the "three amigos"). The Gherkin language writes those examples in a format both humans and machines can read.

# src/test/resources/features/withdrawal.feature
# language: en
Feature: Cash withdrawal limits
  As a bank customer
  I want withdrawals above my daily limit to be reviewed
  So that fraudulent transfers can be stopped

  Background:
    Given an active account "ACC-1001" with balance 50000000 IRR
    And two-factor authentication is enabled for "ACC-1001"

  Scenario: Withdrawal within the daily limit is approved immediately
    When a withdrawal of 5000000 IRR is requested from "ACC-1001"
    Then the withdrawal is approved
    And the balance of "ACC-1001" becomes 45000000 IRR

  Scenario Outline: Withdrawals above the daily limit need manual review
    When a withdrawal of <amount> IRR is requested from "ACC-1001"
    Then the withdrawal status is "<status>"

    Examples:
      | amount   | status         |
      | 20000000 | PENDING_REVIEW |
      | 20000001 | PENDING_REVIEW |
      | 19999999 | APPROVED       |

The step definitions in Java with Cucumber-JVM:

package com.example.acceptance;

import io.cucumber.java.en.Given;
import io.cucumber.java.en.When;
import io.cucumber.java.en.Then;
import static org.assertj.core.api.Assertions.assertThat;

public class WithdrawalSteps {

    private final WithdrawalService service;   // injected via cucumber-spring
    private WithdrawalResult result;

    public WithdrawalSteps(WithdrawalService service) { this.service = service; }

    @Given("an active account {string} with balance {long} IRR")
    public void anActiveAccount(String accountId, long balance) {
        service.createAccount(accountId, balance);
    }

    @Given("two-factor authentication is enabled for {string}")
    public void mfaEnabled(String accountId) {
        service.enableMfa(accountId);
    }

    @When("a withdrawal of {long} IRR is requested from {string}")
    public void requestWithdrawal(long amount, String accountId) {
        result = service.withdraw(accountId, amount);
    }

    @Then("the withdrawal status is {string}")
    public void statusIs(String expected) {
        assertThat(result.status().name()).isEqualTo(expected);
    }
}
package com.example.acceptance;

import org.junit.platform.suite.api.ConfigurationParameter;
import org.junit.platform.suite.api.IncludeEngines;
import org.junit.platform.suite.api.SelectClasspathResource;
import org.junit.platform.suite.api.Suite;

import static io.cucumber.junit.platform.engine.Constants.GLUE_PROPERTY_NAME;
import static io.cucumber.junit.platform.engine.Constants.PLUGIN_PROPERTY_NAME;

@Suite
@IncludeEngines("cucumber")
@SelectClasspathResource("features")
@ConfigurationParameter(key = GLUE_PROPERTY_NAME, value = "com.example.acceptance")
@ConfigurationParameter(key = PLUGIN_PROPERTY_NAME,
                        value = "pretty, html:target/cucumber-report.html")
class RunAcceptanceTest { }
<dependency>
  <groupId>io.cucumber</groupId>
  <artifactId>cucumber-java</artifactId>
  <version>7.22.1</version>
  <scope>test</scope>
</dependency>
<dependency>
  <groupId>io.cucumber</groupId>
  <artifactId>cucumber-junit-platform-engine</artifactId>
  <version>7.22.1</version>
  <scope>test</scope>
</dependency>
<dependency>
  <groupId>org.junit.platform</groupId>
  <artifactId>junit-platform-suite</artifactId>
  <scope>test</scope>
</dependency>
Gherkin turned into a UI script is the worst of both worlds

The bad scenario: "Given the user clicks the login button / And types a username in the first field / And clicks the third button". That is neither readable for the business nor maintainable for engineers, and it breaks on every small UI change. Write the scenario at the level of business behaviour and hide the interaction detail in the step code. The simple test: if someone from the business reads your scenario and says "so what?", the scenario is written wrong.

Pay Cucumber's cost only when there is a buyer for it — Cucumber adds a layer of indirection: between the feature file and the code there is a regex mapping to maintain. That cost is worth paying only when somebody other than developers actually reads or writes those feature files. If the engineering team is the only reader, the same test in plain JUnit and AssertJ is more readable, faster and easier to debug. The senior-level interview answer is exactly this: "I use BDD as a conversation technique always; I use Cucumber only when there is a real non-technical audience."

Living documentation means publishing the output of those same scenario runs (the HTML report) as the official description of system behaviour. Its advantage is that it can never go stale: if behaviour changes and the document is not updated, the test goes red. It is the only kind of documentation that keeps itself honest.

What is the difference between TDD and BDD?

TDD is a development technique at the developer level: red, green, refactor; its purpose is fast feedback and better design. BDD is a collaboration approach at the team-and-business level: before writing code you agree on the meaning of the behaviour through concrete examples, and those examples become executable tests.

Technically, BDD is TDD with a vocabulary shift that focuses on behaviour rather than methods; that is also why it is called ATDD or specification by example. Neither replaces the other: on a real project I apply BDD at the acceptance level for a handful of key flows and TDD at the unit level for domain logic.

And honestly: the biggest value of BDD is not the tool, it is the three-way conversation before any code is written. Teams that install Cucumber without that conversation have merely added a regex layer to their tests.

When you say "I refactored the code", how do you know nothing broke?

First I make the definition precise: refactoring means changing internal structure without changing observable behaviour. If behaviour changes, that is not a refactor, it is a change — and the tests should change with it.

My tool is a test suite that covers behaviour from outside the unit being changed. If the tests are glued to internal details I have no safety net; I am simply writing the same code twice. So before a large refactor I first write higher-level characterisation tests that pin down current behaviour, even if that behaviour is odd.

For legacy code with no tests I use the seam technique: a point where I can make a dependency injectable without changing behaviour, so I can write a test. And to check the quality of that safety net itself I run PIT once over the package; if many mutants survive, my tests are not good enough to refactor behind.

The last layer: I keep refactoring in small commits separate from behaviour changes, so if something breaks in production the revert is quick and unambiguous.

A bug has been reported in production. What is the first thing you do?

The first priority is containment, not root cause: I measure the blast radius (how many users, how much money, since when), and if needed I stop the bleeding with a feature flag or a rollback. Root-causing comes after stabilisation.

Then reproduction: from the logs with the traceId and from real data, I build the smallest case that exhibits the bug. Until I can reproduce it, any "fix" is a guess.

Then the failing test first: I write an automated test at the lowest possible level that reproduces exactly that bug and is red. This both proves I understand the problem and becomes a permanent regression test.

Then fix and confirm: I change the code until that test goes green (the confirmation test) and run the regression suite.

And finally the systemic question: why did this bug pass every layer? Was a design technique missing (a boundary value, for instance)? Was it visible in code review? Should monitoring have alerted earlier? The output of a blameless postmortem is usually a process change, not a line of code.


Part 12 — How to talk about testing in an interview

Three common mistakes: (1) listing tool names instead of explaining decisions; (2) claiming "always TDD", which collapses under one follow-up question; (3) having no numbers at all.

The framework that works when you are asked "how would you test X?":

  1. Risk — what causes the most damage if it breaks? Start there.
  2. Level — which layer gives the cheapest answer.
  3. Technique — name them: equivalence partitioning, boundary values, decision table, state transition.
  4. Data and environment — where the data comes from, how external services are doubled.
  5. Pass criterion — when you declare it done.
  6. What you deliberately do not test — and why that risk is acceptable.

The numbers a senior brings to a quality conversation (siblings of the DORA metrics): escaped defect rate (bugs that reached production), change failure rate, time to restore (MTTR), pipeline duration, flakiness rate, and P1 requirement coverage. If you can say "the pipeline went from 40 minutes to 9, and escaped defects halved in three months", nobody needs further convincing that you understand testing.


Part 13 — A checklist for judging a test suite

On your first day in a new codebase, check these. The answers reveal the team's engineering health better than any document:

# Question Sign of health
1 Do tests run on a fresh laptop with one command? mvn verify is enough
2 How long do unit tests take? the whole unit suite under 90 seconds
3 Are tests interdependent? green under random order and parallel execution
4 Are assertions meaningful or just "not null"? assertions on specific behaviour
5 Do negative and boundary tests exist? the ratio favours failure paths
6 Where are the mocks? only at external boundaries
7 How is test data built? builders, not repeated hand-written SQL
8 Are there known flaky tests? a quarantine list with owners and deadlines
9 What does a red CI mean? it blocks the merge and nobody re-runs to dodge it
10 Do test names describe behaviour? shouldRejectWithdrawalAboveLimit, not test1
11 Is there any non-functional testing? at least one automated load test on the critical path
12 Does the last production bug have a regression test? every incident produced a new test

Command cheat sheet

Task Command
Unit tests only mvn -q test
Unit + integration mvn -q verify
A single test mvn test -Dtest=PaymentServiceTest#shouldRejectAboveLimit
Tests with a tag mvn test -Dgroups=integration
JaCoCo coverage report mvn verify then target/site/jacoco/index.html
Mutation testing mvn test-compile org.pitest:pitest-maven:mutationCoverage
Mutation on changes only mvn org.pitest:pitest-maven:scmMutationCoverage -Dinclude=ADDED,MODIFIED
Run Cucumber mvn test -Dtest=RunAcceptanceTest
Gatling mvn gatling:test -Dgatling.simulationClass=simulations.PaymentSimulation
k6 locally k6 run payments-load.js
k6 with overrides k6 run --vus 50 --duration 5m payments-load.js
JMeter non-GUI + report jmeter -n -t plan.jmx -l out.jtl -e -o report/
JMeter distributed jmeter -n -t plan.jmx -R host1,host2 -l out.jtl
Dependency scan mvn org.owasp:dependency-check-maven:check -DfailBuildOnCVSS=7
Image scan trivy image --severity HIGH,CRITICAL myapp:tag
DAST baseline zap-baseline.py -t https://staging.example.internal -r report.html
Generate a pairwise set pict model.txt > pairs.tsv
Wrapping up

Testing is not "writing @Test"; it is an engineering discipline. The level says where you test and the type says which property you are hunting. The design techniques — equivalence partitioning, boundary values, decision tables, state transition, pairwise, error guessing — are what turn a vague requirement into a complete list of test conditions. The documentation (a Test Plan with an explicit "out of scope", test cases with unarguable expected results, a traceability matrix, defect reports that separate severity from priority, and real exit criteria) is what a mature organisation asks of you. Keep the pyramid economic, fill the gap left by removed E2E tests with contract tests, put flaky tests in time-boxed quarantine, treat coverage as a signal and audit it with mutation testing. In non-functional testing, start with the SLO, use an open workload model to escape coordinated omission, then choose the tool — k6 for the CI gate, Gatling for a Java team, JMeter for enterprise protocols — and read the results through throughput, percentiles, error rate and saturation. Put SAST/SCA/DAST at the right points of the pipeline and do not leave accessibility to automated tools alone. Apply TDD to domain logic and reach for Cucumber only when a genuine non-technical audience exists. And in interviews, instead of tool names, talk about risk, decisions, pass criteria and what you deliberately chose not to test — that is what separates a senior from a mid-level engineer.