Concurrency · همزمانی سنیورSenior ~59 دقیقه مطالعه~51 min read

همگام‌سازی، قفل‌ها، AQS و مدل حافظهٔ جاواSynchronization, Locks, AQS & the Java Memory Model

یاد می‌گیری چرا حافظهٔ جاوا آن‌طور که فکر می‌کنی رفتار نمی‌کند و چطور با happens-before، ‏volatile، ‏synchronized، خانوادهٔ Lock، ‏AQS و CAS دقیقاً همان تضمین‌هایی را که نیاز داری بخری—و بدانی هرکدام چه چیزی را تضمین نمی‌کند.You will learn why Java's memory really doesn't behave the way you assume, and how to buy exactly the guarantees you need with happens-before, volatile, synchronized, the Lock family, AQS, and CAS—while knowing precisely what each one does not promise.

پیش‌نیاز:Prerequisites: نخ‌ها، Runnable/Callable و ExecutorهاThreads, Runnable/Callable & Executors


بیا رک باشیم: بیشتر باگ‌های همروندی (concurrency) به این خاطر پیش می‌آیند که برنامه‌نویس فرض می‌کند حافظه مثل یک دفترچهٔ مشترک رفتار می‌کند که همه هم‌زمان همان صفحه را می‌بینند. حافظهٔ واقعیِ یک ماشین چندهسته‌ای اصلاً این‌طور نیست. این فصل قرار است آن مدل ذهنیِ غلط را کامل بشکند و به‌جایش یک مدل درست بسازد—و بعد ابزارهایی را که با آن‌ها می‌توانی ترتیب و رؤیت‌پذیری را کنترل کنی، یکی‌یکی و از صفر بسازیم.

نقشهٔ راه این فصل

اول می‌بینیم چرا JMM اصلاً وجود دارد و سه عاملی که کدت را بازچینش می‌کنند. بعد به قلب ماجرا می‌رسیم: happens-before، تنها قاعده‌ای که تضمین می‌کند یک نخ نوشتهٔ نخ دیگر را ببیند. سپس ابزارها را از سبک به سنگین می‌سازیم: ‏volatile، ‏synchronized و بهینه‌سازی‌های قفل، خانوادهٔ Lock (‏ReentrantLock / ReadWriteLock / StampedLock)، موتور زیرین AQS، برنامه‌نویسی بدون‌قفل با CAS و مسئلهٔ ABA، ‏false sharing، و بالاخره الگوی معروف double-checked locking. در پایان یک بخش کامل پرسش‌های مصاحبه و یک جمع‌بندی داریم.

بخش ۰ — واژه‌هایی که باید از قبل بشناسی

قبل از هر چیز چند واژه را با تشبیه محکم کنیم تا بعداً سرد و ناگهانی رهایشان نکنم.

نخ، هسته و کش مثل چند آشپز در یک آشپزخانه

یک نخ (thread) را مثل یک آشپز فرض کن که مشغول اجرای دستورپخت (کد) توست. یک CPU چندهسته‌ای یعنی چند آشپز که هم‌زمان کار می‌کنند. حالا نکتهٔ مهم: هر آشپز یک میز کار کوچک کنار دستش (cache هستهٔ خودش) دارد و انبار اصلی (RAM) آن‌طرف آشپزخانه است. وقتی آشپز نمکدان را برمی‌دارد، یک کپی روی میز خودش می‌گذارد و مدتی همان کپی را استفاده می‌کند. پس اگر آشپز دیگری در انبار اصلی نمک را عوض کند، آشپز اول تا مدتی خبردار نمی‌شود—او هنوز کپیِ کهنهٔ روی میزش را می‌بیند. این «کپیِ کهنه» ریشهٔ اکثر باگ‌های رؤیت‌پذیری (visibility) در جاواست.

  • بازچینش (reordering): جابه‌جاکردن ترتیب اجرای دستورها برای سرعت بیشتر—مثل آشپزی که به‌جای ترتیب نوشته‌شدهٔ دستور، کارها را به کارآمدترین ترتیب انجام می‌دهد.
  • مانع/حصار حافظه (memory barrier / fence): یک دستور ویژه که به سخت‌افزار می‌گوید «تا اینجا را واقعاً تمام کن و همه ببینند، بعد برو جلو»—مثل قانونی که می‌گوید «تا دیس را روی میز پاس اصلی نگذاشتی، غذای بعدی را شروع نکن.»
  • اتمیک (atomic): عملیاتی که یا کامل انجام می‌شود یا اصلاً؛ هیچ نخ دیگری نمی‌تواند وسطش حالت نیمه‌کاره را ببیند.

مدل ذهنی: چرا JMM وجود دارد؟

به‌طور ساده تصور می‌کنی برنامه‌ات یک فهرست ترتیبیِ واحد از خواندن‌ها و نوشتن‌های حافظه است که دقیقاً به ترتیب سورس اجرا می‌شود و بی‌درنگ برای همهٔ نخ‌ها دیده می‌شود. این مدل، هرچقدر هم راحت باشد، دروغ است. میان سورس تو و الکترون‌های داخل تراشه، سه عامل بازچینش وجود دارد:

  1. کامپایلر (javac و JIT): می‌تواند دستورها را بازچینش، بالابری (hoist)، حذف یا ادغام کند. «بالابری» یعنی چیزی را که داخل حلقه هربار خوانده می‌شد، یک بار بیرون حلقه بخواند و در یک رجیستر نگه دارد—بعداً می‌بینی چرا این می‌تواند فاجعه بسازد.
  2. CPU: دستورها را خارج از ترتیب (out-of-order) و گاهی به‌صورت گمانه‌زنانه (speculative، یعنی حدس می‌زند کدام شاخه اجرا می‌شود و از قبل شروعش می‌کند) اجرا می‌کند.
  3. سلسله‌مراتب حافظه: بافر نوشتن (store buffer) و کش هر هسته باعث می‌شوند نوشتهٔ یک هسته با تأخیر برای هستهٔ دیگر دیده شود—همان «کپیِ کهنه روی میز آشپز».
این‌ها باگ نیستند، منبع سرعت‌اند

وسوسه می‌شوی فکر کنی این سه عامل «خرابکار»اند. برعکس: تمام کارایی مدرن از همین بازچینش‌ها می‌آید. سکو فقط یک قول به تو می‌دهد: برای کد تک‌نخی معنای as-if-serial را تضمین می‌کند—یعنی هر بازچینشی مجاز است تا وقتی نتیجه‌ای که خودت در همان نخ مشاهده می‌کنی عوض نشود. مشکل دقیقاً وقتی شروع می‌شود که نخ دیگری بخواهد وسط کار تو را تماشا کند.

اینجاست که مدل حافظهٔ جاوا (Java Memory Model / JMM) وارد می‌شود. این مدل در فصل ۱۷ از JLS تعریف و با JSR-133 (جاوا ۵) بازنویسی شد، و همان قراردادی است که دقیقاً می‌گوید کدام خواندن‌های بین‌نخی مجازند کدام نوشتن‌ها را ببینند. هرچه در بقیهٔ این فصل می‌آید—‏volatile، ‏synchronized، ‏Lock، اتمیک‌ها—در واقع راهی است برای خریدن یال‌های happens-before از این قرارداد.

Happens-before: تنها قاعده‌ای که واقعاً اهمیت دارد

اگر فقط یک چیز از این فصل با خودت ببری، همین باشد.

happens-before مثل امضای تحویل بسته

تصور کن دو دپارتمان در یک شرکت‌اند. دپارتمان A سندی را آماده می‌کند و دپارتمان B باید رویش کار کند. اگر هیچ رویه‌ای بینشان نباشد، B ممکن است نسخهٔ نیمه‌کاره یا اصلاً نسخهٔ دیروز را بردارد. اما اگر قانون این باشد که «A سند را در صندوق مشترک می‌گذارد و امضا می‌کند، و B فقط بعد از دیدن آن امضا سند را برمی‌دارد»، آنگاه هر چیزی که A پیش از امضا نوشته، تضمیناً برای B که بعد از امضا می‌خواند دیده می‌شود. آن جفتِ «امضا» (release) و «دیدن امضا» (acquire) دقیقاً همان یال happens-before است.

حالا دقیق شویم. JMM بر پایهٔ یک ترتیب جزئی (partial order) به نام happens-before (به‌اختصار HB) تعریف می‌شود. «ترتیب جزئی» یعنی لازم نیست هر دو کنش با هم مقایسه‌شدنی باشند؛ بعضی جفت‌ها مرتب‌اند و بعضی نه. قاعده این است: اگر کنش A نسبت به کنش B رابطهٔ happens-before داشته باشد، آنگاه تمام اثرهای حافظهٔ A برای B دیده می‌شوند و پیش از آن مرتب شده‌اند.

و اما نکتهٔ خطرناک: اگر دو کنش با HB مرتب نشده باشند و دست‌کم یکی نوشتن روی همان مکان باشد، یک مسابقهٔ داده (data race) داری. در این حالت JMM اجازه می‌دهد خواندن‌ها مقدارهای کهنه (stale)، پاره‌شده (torn، یعنی نیمی از مقدار قدیم و نیمی جدید) یا حتی به‌ظاهر «ناممکن» برگردانند.

یال‌های HB که واقعاً به دست می‌آوری، اینهاست—این فهرست را حفظ کن:

  • ترتیب برنامه (program order): درون یک نخ، هر کنش نسبت به هر کنش بعدیِ همان نخ HB دارد.
  • قفل مانیتور (monitor): یک unlock روی یک مانیتور نسبت به هر lock بعدی روی همان مانیتور HB دارد.
  • volatile: نوشتن روی یک فیلد volatile نسبت به هر خواندن بعدی همان فیلد HB دارد.
  • شروع نخ: Thread.start() نسبت به هر کنش در نخ آغازشده HB دارد.
  • پیوستن نخ (join): هر کنش در یک نخ نسبت به بازگشت موفق نخ دیگر از join() روی آن HB دارد.
  • وقفه (interrupt): فراخوانی interrupt() نسبت به تشخیص آن توسط نخ وقفه‌خورده HB دارد.
  • فیلدهای final: پایان سازنده (constructor) نسبت به «انجماد» (freeze) فیلدهای final HB دارد؛ این پایهٔ انتشار امن (safe publication) اشیای تغییرناپذیر است.
  • ترایایی (transitivity): اگر A HB B و B HB C، آنگاه A HB C. یعنی یال‌ها زنجیر می‌شوند.
happens-before دربارهٔ زمان نیست، دربارهٔ ترتیب است

مهم‌ترین سوءتفاهم را همین‌جا خاک کنیم. دو رویداد می‌توانند از نظر ساعت دیواری کاملاً «هم‌زمان» باشند؛ آنچه اهمیت دارد این نیست که کدام «زودتر» رخ داد، بلکه این است که آیا مدل وادار می‌کند نوشته‌های یکی توسط دیگری دیده شود. و این وادارسازی فقط میان یک کنش release و یک کنش acquire بعدی روی همان متغیر همگام‌سازی برقرار می‌شود. آزادکردن قفل A هیچ چیزی دربارهٔ نخی که قفل B را می‌گیرد به تو نمی‌گوید—امضای صندوق A به درد صندوق B نمی‌خورد.

بیا با یک مثال ملموس این زنجیر را ببینیم:

نخ ۱                              نخ ۲
-----                            -----
data = 42;        (1)
ready = true;     (2، نوشتن volatile / release)
                                 while(!ready) {}  (3، خواندن volatile / acquire)
                                 print(data);      (4)  -> تضمیناً 42 را می‌بیند

چرا کار می‌کند؟ چون (2) نوشتن volatile و (3) خواندن volatile همان فیلد است، پس (2) HB (3). طبق ترتیب برنامه (1) HB (2) و همچنین (3) HB (4). حالا ترایایی را زنجیر کن: ‏(1) HB (2) HB (3) HB (4)، پس (1) HB (4). یعنی خواندن data در خط (4) حتماً ۴۲ را می‌بیند.

اگر volatile را برداری چه می‌شود؟

اگر volatile را از ready حذف کنی، دو چیز هم‌زمان می‌شکند: هم رؤیت‌پذیری data (ممکن است نخ ۲ مقدار کهنه ببیند) و هم پایان حلقه. JIT حق دارد ready را از داخل حلقه بالا ببرد و در یک رجیستر نگه دارد—آن‌وقت while(!ready) برای همیشه همان مقدار رجیستری را می‌بیند و تا ابد حلقه می‌زند، حتی بعد از اینکه نخ ۱ مقدار را عوض کرد. این یکی از رایج‌ترین باگ‌های واقعی است.

volatile: چه می‌کند و چه نمی‌کند

حالا که happens-before را داری، ‏volatile ساده می‌شود. آن را مثل یک نمکدانِ شیشه‌ای فرض کن که همه مجبورند مستقیم از انبار اصلی بردارند، نه از کپیِ روی میز—و هر بار که کسی چیزی در آن می‌گذارد، مجبور است اول همهٔ کارهای قبلی‌اش را روی میز پاس اصلی مرتب کند.

volatile دقیقاً دو چیز به تو می‌دهد:

  1. رؤیت‌پذیری + ترتیب (release/acquire): نوشتن volatile نه‌فقط همان متغیر، بلکه یک یال HB می‌سازد و تمام نوشتن‌هایی را که در ترتیب برنامه پیش از آن بودند، برای نخی که بعداً همان volatile را می‌خواند دیدنی می‌کند. همین «سوارشدن» (piggybacking) بود که مثال data/ready بالا را کار انداخت.
  2. اتمیک‌بودن خودِ فیلد، حتی برای long/double. توجه: طبق JLS، نوشتن ۶۴بیتیِ غیر‌volatile مجاز است به دو ذخیرهٔ ۳۲بیتی شکسته شود؛ ‏volatile این پاره‌شدن را ممنوع می‌کند.

و اما آنچه volatile نمی‌دهد—اینجا اکثر آدم‌ها اشتباه می‌کنند:

  • بدون انحصار متقابل (mutual exclusion).volatile int x; x++; در واقع یک read-modify-write است (بخوان، جمع کن، بنویس) و اتمیک نیست. دو نخ می‌توانند هردو ۵ را بخوانند، هردو ۶ حساب کنند و هردو ۶ بنویسند—و یک افزایش گم می‌شود. برای این کار AtomicInteger یا قفل لازم داری.
  • بدون اتمیک‌بودن کنش مرکب.if (v == null) v = new X(); حتی اگر v volatile باشد باز هم مسابقه دارد، چون بین بررسی و انتساب فاصله هست.
پشت صحنه: چرا x86 «ارزان» است ولی ARM نه

روی معماری x86 نوشتن volatile به یک ذخیرهٔ ساده و سپس یک مانع store-load (اغلب دستوری با پیشوند lock یا mfence) کامپایل می‌شود؛ اما خواندن‌ها عملاً رایگان‌اند، چون x86 مدل حافظهٔ قوی‌ای به نام TSO (Total Store Order) دارد. روی مدل‌های حافظهٔ ضعیف‌تر مثل ARM و Power، هم خواندن و هم نوشتن باید دستورهای مانع (fence) واقعی تولید کنند. یعنی همان کد جاوا روی موبایل ARM ممکن است گران‌تر از سرور x86 تمام شود.

synchronized: مانیتور ذاتی

اگر volatile نمکدان مشترک بود، ‏synchronized کلید یک اتاق است که هر بار فقط یک نفر داخلش می‌رود.

مانیتور مثل دستشویی تک‌نفرهٔ با کلید

هر شیء جاوا یک قفل نامرئی به نام مانیتور (monitor) یا قفل ذاتی (intrinsic lock) همراه دارد—درست مثل یک دستشویی که کلیدش روی در آویزان است. برای واردشدن باید کلید را برداری (lock)، و موقع خروج آویزانش کنی (unlock). تا وقتی کلید دست توست، هیچ‌کس دیگر نمی‌تواند وارد شود. بلوک synchronized این «برداشتن و آویزان‌کردن کلید» را به‌طور خودکار انجام می‌دهد—حتی اگر وسط کار استثنا پرتاب شود، خروجش مثل یک finally عمل می‌کند و کلید را برمی‌گرداند، پس هرگز نمی‌توانی کلید را «نشت» بدهی.

تضمین‌های synchronized:

  • انحصار متقابل: تنها یک نخ در هر لحظه یک مانیتور معین را نگه می‌دارد.
  • رؤیت‌پذیری: ‏unlock نسبت به lock بعدی روی همان مانیتور HB دارد، پس تمام نوشته‌های داخل ناحیهٔ بحرانی (critical section) منتشر می‌شوند.
  • بازورودی (reentrancy): همان نخ می‌تواند مانیتوری را که پیشاپیش نگه داشته دوباره بگیرد (یک شمارندهٔ نگه‌داشت به‌ازای هر نخ نگه‌داشته می‌شود). همین است که اجازه می‌دهد یک متد synchronized، متد synchronized دیگری روی this را صدا بزند بدون اینکه خودش را در بن‌بست بیندازد.

نکتهٔ ظریف دربارهٔ اینکه روی چه چیزی قفل می‌کنی: ‏synchronized(this) و یک متد نمونهٔ synchronized، هردو روی this قفل می‌کنند؛ یک متد استاتیک synchronized روی شیء Class قفل می‌کند.

هرگز روی String یا Integer باکس‌شده قفل نکن

synchronized("lock") یا قفل روی یک Integer که از باکسینگ آمده، یک تلهٔ کلاسیک است. لیترال‌های String در استخر رشته (intern) می‌شوند و Integerهای کوچک کش می‌شوند، پس همان شیء دقیق ممکن است در کدهای کاملاً بی‌ربطِ دیگر هم استفاده شود—و ناگهان دو بخش نامرتبط برنامه روی یک قفل مشترک، بی‌آنکه بدانند، بن‌بست بسازند. همیشه یک قفل اختصاصی و خصوصی بساز: ‏private final Object lock = new Object();.

بهینه‌سازی‌های قفل — و آنچه حذف شد

از نظر تاریخی HotSpot سه شیوهٔ قفل را لایه‌بندی می‌کرد که خوب است بشناسی، به‌ویژه چون یکی‌شان حذف شده و سؤال مصاحبه است:

  • قفل مغرضانه (biased locking): فرض می‌کرد یک قفل معمولاً توسط همان نخ دوباره گرفته می‌شود، پس شیء را به آن نخ «مغرض» می‌کرد و در بازورود از CAS اتمیک به‌کلی می‌گذشت. این ویژگی به‌طور پیش‌فرض غیرفعال و منسوخ (deprecated) شد در JDK 15 (JEP 374) و پیاده‌سازی‌اش بعداً کهنه/حذف شد (JDK 18، ‏JDK-8256425). دلیل؟ هزینهٔ دفترداری و ابطالش، بارهای کاری مدرن با نخ‌های کوتاه‌عمرِ فراوان و ساختارهای دادهٔ همروندِ پرمناقشه را کند می‌کرد.
  • قفل سبک (lightweight / thin): قفل‌های بدون مناقشه با یک CAS روی «mark word» سرآیند شیء، یک رکورد قفل روی پشتهٔ نخ تخصیص می‌دهند—بدون میوتکس سطح سیستم‌عامل، یعنی خیلی ارزان.
  • قفل سنگین (heavyweight / inflated): وقتی مناقشه بالا می‌گیرد، مانیتور به یک ObjectMonitor سطح سیستم‌عامل با یک صف انتظار واقعی و پارک‌کردن OS «باد می‌کند» (inflate). این گران است ولی زیر فشار لازم.
نکتهٔ مدرن: JEP 491 و نخ‌های مجازی

یک نکتهٔ در سطح ارشد که حتماً باید بدانی: از JDK 24، ‏JEP 491 باعث می‌شود synchronized دیگر نخ‌های مجازی (virtual threads) را سنجاق (pin) نکند. پیش از ۲۴، وقتی یک نخ مجازی داخل بلوک synchronized مسدود می‌شد، نخ سکوی حاملش (carrier platform thread) را گروگان می‌گرفت—به این «pinning» می‌گویند—و مقیاس‌پذیری را خفه می‌کرد. برای همین توصیه می‌شد در کدی که پر از نخ مجازی است ReentrantLock را ترجیح بدهی. از JDK 24 آن توصیه منسوخ است: حالا synchronized در برابر java.util.concurrent.locks را صرفاً بر پایهٔ راحتی و امکانات انتخاب کن، نه مقیاس‌پذیری.

خانوادهٔ Lock: ReentrantLock، ReadWriteLock، StampedLock

مانیتور ذاتی ساده و امن است، ولی خشک است: نمی‌توانی بگویی «اگر قفل تا ۲ ثانیه آزاد نشد بی‌خیال شو». بستهٔ java.util.concurrent.locks قفل‌های صریحی می‌دهد که این انعطاف را دارند.

ReentrantLock — اسب بارکش

ReentrantLock همان انحصار متقابلِ synchronized را می‌دهد، اما با قابلیت‌های اضافه: ‏tryLock() (با مهلت اختیاری—«اگر تا فلان زمان نشد، برگرد»)، گرفتنِ قابل‌وقفه با lockInterruptibly()، ‏انصاف (fairness) اختیاری (ترتیب FIFO برای نخ‌های منتظر، به بهای کمی افت توان عملیاتی)، و چند شیء Condition به‌ازای هر قفل (در برابر تنها یک مجموعه‌انتظار برای هر مانیتور).

اما یک اصطلاح غیرقابل‌مذاکره دارد—این را در خواب هم باید بلد باشی:

private final ReentrantLock lock = new ReentrantLock();

void doWork() {
    lock.lock();
    try {
        // ناحیهٔ بحرانی
    } finally {
        lock.unlock(); // باید در finally باشد — استثنا نباید قفل را نشت دهد
    }
}
فرق کلیدی با synchronized: unlock خودکار نیست

با synchronized کلید همیشه خودبه‌خود برمی‌گردد. با Lock صریح، اگر بین lock() و unlock() استثنایی پرتاب شود و unlock() را در finally نگذاشته باشی، قفل برای همیشه نگه‌داشته می‌ماند و هر نخ دیگری که آن را بخواهد تا ابد منتظر می‌ماند. همیشه unlock() در finally.

ReentrantReadWriteLock — قفل خواندن/نوشتن جدا

تخته‌سفید کلاس

تصور کن یک تخته‌سفید داری. خواندن یعنی نگاه‌کردن به تخته—ده نفر می‌توانند هم‌زمان نگاه کنند، هیچ مشکلی نیست. نوشتن یعنی پاک‌کردن و نوشتن دوباره—موقع نوشتن هیچ‌کس دیگری نباید نه بنویسد نه بخواند، وگرنه چیز نیمه‌پاک‌شده می‌بیند. ‏ReentrantReadWriteLock دقیقاً همین است: قفل خواندن را خیلی‌ها هم‌زمان می‌گیرند، ولی قفل نوشتن انحصاری است.

این وقتی خوب است که خواندن‌ها به‌شدت بر نوشتن‌ها غلبه دارند و ناحیه‌های بحرانی بی‌اهمیت (خیلی کوتاه) نیستند. اما دو تلهٔ مهم:

  • گرسنگی نویسنده (writer starvation): اگر سیل خوانندگان بی‌وقفه بیاید، نویسنده ممکن است هرگز نوبت نگیرد. با سازندهٔ fair کاهشش بده.
  • تنزل مجاز، ارتقا ممنوع:تنزل (downgrade) یعنی از قفل نوشتن به خواندن بروی—این مجاز است، به‌شرطی که قفل خواندن را پیش از آزادکردن قفل نوشتن بگیری. اما ارتقا (upgrade) یعنی از خواندن به نوشتن—این بن‌بست می‌سازد و ممنوع است. (چرایی‌اش را در پرسش مصاحبهٔ آخر کامل باز می‌کنیم.)

StampedLock (جاوا ۸) — خواندن خوش‌بینانه

StampedLock بازورودی نیست، اما یک ویژگی مرگبار می‌افزاید: خواندن خوش‌بینانه (optimistic read).

نگاه دزدکی به ساعت دیواری

فرض کن می‌خواهی ساعت را ببینی. حالت بدبینانه این است که بروی جلوی ساعت بایستی و مانع شوی کسی عقربه‌ها را تکان دهد (قفل خواندن). حالت خوش‌بینانه این است که فقط یک نگاه سریع بیندازی، عدد را حفظ کنی، و بعد چک کنی «آیا در این فاصله کسی به ساعت دست زد؟» اگر نه، عددت معتبر است و اصلاً مزاحم کسی نشدی. اگر بله، آن‌وقت به‌ناچار می‌روی جلوی ساعت می‌ایستی. این نگاه‌سریع همان tryOptimisticRead و آن چک همان validate است.

private final StampedLock sl = new StampedLock();
private double x, y;

double distanceFromOrigin() {
    long stamp = sl.tryOptimisticRead();      // بدون CAS، بدون مسدودشدن
    double cx = x, cy = y;                     // خواندن عکس‌فوری (snapshot)
    if (!sl.validate(stamp)) {                 // ممکن است نویسنده‌ای اجرا شده باشد
        stamp = sl.readLock();                 // بازگشت به قفل خواندن بدبینانه
        try { cx = x; cy = y; }
        finally { sl.unlockRead(stamp); }
    }
    return Math.sqrt(cx * cx + cy * cy);
}

سود بزرگ: وقتی نوشتن‌ها نادرند، مسیر خواندن اصلاً به خط کش قفل دست نمی‌زند، پس مناقشهٔ cache line به‌کلی حذف می‌شود.

تله‌های StampedLock

سه چیز را یادت باشد: (۱) بازورودی نیست—اگر همان نخ دوباره قفل کند، خودش را در بن‌بست می‌اندازد. (۲) از Condition پشتیبانی نمی‌کند. (۳) مستقیماً قابل‌وقفه نیست (باید از گونه‌های Interruptibly استفاده کنی). و مهم‌ترین قانون: در خواندن خوش‌بینانه باید فیلدها را پیش از validate() درون متغیرهای محلی کپی کنی، و نباید وسط خواندن متد دیگری صدا بزنی یا یک مرجع بالقوه‌ناسازگار را واکاوی (dereference) کنی—چون ممکن است آن مرجع نیمه‌کاره باشد.

AbstractQueuedSynchronizer (AQS): موتور زیرین

حالا پرده را کنار بزنیم. ‏ReentrantLock، ‏Semaphore، ‏CountDownLatch، ‏ReentrantReadWriteLock و حتی دروازهٔ کارگرِ ThreadPoolExecutor—همه پوشش‌های نازکی روی یک موتور مشترک به نام AQS هستند. اگر این یک قطعه را بفهمی، انگار سورس نصف java.util.concurrent را بلدی.

AQS مثل سیستم نوبت‌دهی بانک

یک بانک را تصور کن با یک تابلوی «شمارهٔ فعلی» (state) و یک صف مرتب از مشتری‌ها که شماره کشیده‌اند. وقتی باجه آزاد است، نفر جلوی صف را صدا می‌زنند. وقتی مشغول است، تازه‌واردها یک شماره می‌گیرند و می‌نشینند/می‌خوابند (park) تا نوبتشان شود. وقتی مشتری فعلی کارش تمام شد، دقیقاً نفر بعدی را بیدار می‌کند (unpark). ‏AQS همین سیستم است: یک عدد وضعیت مشترک، به‌علاوهٔ یک صف منظمِ نخ‌های خوابیده.

مدل ذهنی دقیق:

  • AQS یک تک volatile int state و یک صف انتظار FIFO مبتنی بر CLH از نخ‌ها نگه می‌دارد. (‏CLH نوع خاصی از صف مبتنی بر لیست پیوندی است که هر نخ روی گرهٔ خودش می‌چرخد/می‌خوابد.)
  • یک زیرکلاس تعریف می‌کند که state چه معنایی دارد، و متدهای tryAcquire(int) / tryRelease(int) (برای حالت انحصاری) یا tryAcquireShared / tryReleaseShared (برای حالت اشتراکی) را پیاده می‌کند.
  • خودِ AQS بخش سختِ ماشین‌کاری را انجام می‌دهد: ‏CAS اتمیک روی state، صف‌بندی بازندگان به‌صورت گره‌های صف، پارک‌کردنشان با LockSupport.park()، و بیدارکردن جانشین هنگام آزادسازی.

حالا ببین چطور همان موتور، چند کلاس مختلف می‌سازد:

  • ReentrantLock:state همان شمارندهٔ نگه‌داشت است—۰ یعنی آزاد، N یعنی N بار بازوردانه نگه‌داشته‌شده. ‏tryAcquire مقدار ۰→۱ را CAS می‌کند، یا اگر نخ فعلی خودش صاحب قفل است، شمارنده را یکی بالا می‌برد.
  • Semaphore:state همان شمار مجوز (permit) است و از گرفتن اشتراکی استفاده می‌کند.
  • CountDownLatch:state همان شمارنده است، و وقتی به ۰ می‌رسد همهٔ منتظران را یک‌جا رد می‌کند—این «آزادسازی اشتراکی» است، دلیل اینکه یک latch می‌تواند چندین نخ را هم‌زمان رها کند.
چرا AQS اهمیت دارد

اگر «‏state + صف CLH + tryAcquire» را بفهمی، دیگر لازم نیست هر کلاس همگام‌سازی را جدا حفظ کنی—فقط باید بپرسی «این کلاس، state را چه معنا می‌کند و tryAcquire‌اش چه‌کار می‌کند؟» و بقیهٔ رفتار (صف، park، unpark) از موتور مشترک می‌آید.

CAS، کلاس‌های *Atomic و مسئلهٔ ABA

تا اینجا همه‌چیز دربارهٔ قفل بود. اما یک لایهٔ بالاتر—و اغلب سریع‌تر—هم هست: برنامه‌نویسی بدون‌قفل (lock-free).

CAS مثل ویرایش هم‌زمان یک سند با «آیا کسی عوضش کرد؟»

فرض کن روی یک سند مشترک کار می‌کنی. به‌جای اینکه سند را قفل کنی، این‌طور عمل می‌کنی: نسخهٔ فعلی را می‌خوانی، تغییرت را آماده می‌کنی، و موقع ذخیره می‌گویی «فقط اگر سند هنوز همان چیزی است که خواندم، تغییرم را اعمال کن؛ وگرنه بگو شکست خورد.» اگر کسی وسط کار سند را عوض کرده باشد، ذخیره‌ات رد می‌شود و از اول تلاش می‌کنی. این همان compare-and-swap (CAS) است.

‏CAS یک تک دستور سخت‌افزاری است (lock cmpxchg روی x86، ‏LL/SC روی ARM) که به‌صورت اتمیک انجام می‌دهد: «اگر این خانهٔ حافظه برابر مورد انتظار است، آن را به مقدار جدید بگذار و موفقیت را گزارش کن؛ وگرنه دست نزن و شکست را گزارش کن.» بستهٔ java.util.concurrent.atomic این را در اختیارت می‌گذارد: ‏AtomicInteger، ‏AtomicLong، ‏AtomicReference و گونه‌های آرایه و field-updater.

AtomicInteger counter = new AtomicInteger();

int incrementAndGet() {
    int prev, next;
    do {
        prev = counter.get();
        next = prev + 1;
    } while (!counter.compareAndSet(prev, next)); // تا بردن در مسابقه تلاش مجدد
    return next;
}

این حلقه خوش‌بینانه است: هیچ نخی مسدود نمی‌شود؛ مناقشه فقط باعث تلاش مجدد می‌شود.

کِی CAS برنده است و کِی بازنده

زیر مناقشهٔ کم، این حلقه قفل‌ها را له می‌کند—چون اصلاً کسی نمی‌خوابد و بیدار نمی‌شود. اما زیر مناقشهٔ سنگین، یک توفان تلاش مجدد راه می‌افتد: ده‌ها نخ مدام همدیگر را رد می‌کنند و CPU می‌سوزد بی‌آنکه کار مفیدی بشود. دقیقاً به همین دلیل جاوا ۸ ‏**LongAdder/LongAccumulator** را افزود که شمارش را روی چند سلول جدا (که padded شده‌اند تا false sharing رخ ندهد) رگه‌رگه می‌کند و فقط موقع تقاضا جمع می‌زند. برای شمارنده‌های داغ به‌مراتب بهتر از یک تک AtomicLong است.

مسئلهٔ ABA

CAS یک نقطه‌ضعف ظریف دارد: برابری مقدار را بررسی می‌کند، نه اینکه مقدار تغییر کرده و بازگشته باشد.

کلید یدکی و دزد باهوش

یک صندوق داری که با یک عدد قفل می‌شود. قانونت این است: «اگر عدد هنوز A است، آن را باز کن.» یک دزد باهوش عدد را از A به B و بعد دوباره به A برمی‌گرداند. حالا وقتی تو چک می‌کنی «هنوز A است؟»، بله هست—پس بازش می‌کنی، غافل از اینکه در این فاصله همه‌چیز عوض شده. مقدار همان است، اما دنیا عوض شده.

به زبان دقیق: اگر نخی A را بخواند، نخ دیگری A→B→A را برگرداند، بعد compareAndSet(A, ...) نخ اول موفق می‌شود—هرچند جهان زیر پایش جابه‌جا شده. برای یک شمارندهٔ int این بی‌ضرر است. اما برای ساختاری که اشاره‌گر جابه‌جا می‌کند (مثلاً یک پشتهٔ بدون‌قفل که گرهی را pop می‌کند که در این فاصله آزاد و بازتخصیص شده)، این وضعیت را خراب می‌کند. راه‌حل، جفت‌کردن مقدار با یک مُهر/نسخهٔ یکنوا (monotonic stamp) است: ‏AtomicStampedReference (مقدار + مُهر int) یا AtomicMarkableReference (مقدار + بولین).

اشتراک کاذب (false sharing): مالیات نامرئی

این یکی از آن باگ‌های عملکردی است که در کد هیچ ردی از خودش نمی‌گذارد و فقط پروفایلر پیدایش می‌کند.

یک سینی مشترک برای دو نفر

کش‌ها حافظه را نه متغیر به متغیر، بلکه در بسته‌های ۶۴بایتی به نام خط کش (cache line) جابه‌جا می‌کنند. حالا تصور کن دو نفر روی یک سینی مشترک هرکدام یک لیوان جدا دارند. لیوان‌ها مستقل‌اند، ولی چون روی یک سینی‌اند، هر بار که یکی لیوانش را جابه‌جا می‌کند، مجبور است کل سینی را بردارد و دیگری باید صبر کند تا سینی برگردد. دو متغیرِ مستقل که تصادفاً در یک خط کش نشسته‌اند دقیقاً همین‌اند: هر نوشتن روی یکی، نسخهٔ هستهٔ دیگر را باطل می‌کند و خط میان هسته‌ها «پینگ‌پنگ» می‌شود، هرچند هیچ اشتراک منطقی‌ای نیست.

این اشتراک کاذب است و می‌تواند بی‌صدا یک مرتبهٔ بزرگی (۱۰ برابر) کندت کند.

راه‌حل: فیلدهای داغ را روی خط کش خودشان بالشتک‌گذاری (pad) کن تا هرکدام سینیِ جدا داشته باشند. جاوا ۸ به بعد @jdk.internal.vm.annotation.Contended را می‌دهد (برای کد اپلیکیشن به فلگ -XX:-RestrictContended نیاز دارد)؛ ‏Cell داخلیِ LongAdder دقیقاً همین‌طور annotate شده—برای همین LongAdder این‌قدر خوب مقیاس می‌گیرد. بالشتک‌گذاری دستی (long p1..p7) هم کار می‌کند، اما JIT ممکن است فیلدهای بی‌استفاده را حذف کند، پس هرجا در دسترس است @Contended ارجح است.

قفل‌گذاری دوبار-بررسی‌شده (double-checked locking)، درست‌شده

این معروف‌ترین الگوی «به‌ظاهر درست ولی خراب» در همروندی جاواست. هدف: یک شیء گران را فقط یک بار و تنبل (lazy) بساز، و بعد بدون قفل سریع برش گردان.

اصطلاح کلاسیکِ خراب:

// پیش از جاوا ۵ خراب بود و امروز هم بدون volatile خراب است
private Helper helper;
Helper getHelper() {
    if (helper == null) {                 // بررسی اول (بدون قفل)
        synchronized (this) {
            if (helper == null)           // بررسی دوم (با قفل)
                helper = new Helper();    // انتشار
        }
    }
    return helper;
}

چرا بدون volatile خراب است؟ چون helper = new Helper() یک عملیات اتمیک نیست، بلکه سه گام است: (۱) حافظه تخصیص بده، (۲) سازنده را اجرا کن و فیلدها را مقداردهی کن، (۳) مرجع را به helper انتساب بده. و JMM اجازه می‌دهد که گام ۳ (انتساب مرجع) پیش از گام ۲ (نوشته‌های سازنده) دیده شود.

شیء نیمه‌ساخته‌شده: باگ ترسناک

تصور کن نخ دوم روی مسیر سریعِ «بررسی اول» است. به‌خاطر آن بازچینش، این نخ می‌تواند یک helper غیرتهی ببیند که هنوز به یک شیء نیمه‌ساخته‌شده اشاره می‌کند—مرجع رسیده، ولی فیلدهای داخلی هنوز مقدار پیش‌فرض (۰/null) دارند. نخ دوم آن را برمی‌گرداند و ازش استفاده می‌کند، و تو یک باگ کاملاً غیرقابل‌بازتولید داری.

راه‌حل، volatile کردن فیلد است. این کار مانع release/acquire را درج می‌کند که تضمین می‌کند نوشته‌های سازنده نسبت به هر خواندن مرجع HB داشته باشند:

private volatile Helper helper;           // volatile اجباری است
Helper getHelper() {
    Helper result = helper;               // یک بار volatile را در محلی بخوان
    if (result == null) {
        synchronized (this) {
            result = helper;
            if (result == null)
                helper = result = new Helper();
        }
    }
    return result;
}

خواندن در متغیر محلی result یک بهینه‌سازی واقعی است: در مسیر داغ، دو خواندن volatile را به یکی فرومی‌کاهد.

بهترین راه برای singleton: اصلاً DCL ننویس

اگر یک singleton استاتیک می‌خواهی، کل این پیچیدگی را دور بزن و از اصطلاح holder با راه‌اندازی به‌تقاضا (initialization-on-demand holder) استفاده کن. این الگو به تضمین‌های راه‌اندازی تنبلِ کلاس در خود JLS تکیه می‌کند: کلاس داخلی تا اولین دسترسی به Holder.INSTANCE اصلاً بارگذاری نمی‌شود، و خودِ JVM امن‌بودن و تنبل‌بودن این راه‌اندازی را تضمین می‌کند—بدون هیچ volatile یا synchronized.

class Singleton {
    private Singleton() {}
    private static class Holder { static final Singleton INSTANCE = new Singleton(); }
    static Singleton getInstance() { return Holder.INSTANCE; } // JVM راه‌اندازی امن و تنبل را تضمین می‌کند
}

دام‌ها و نکات ظریف رایج

اینها را مثل یک چک‌لیست موقع بازبینی کد همروند نگه دار:

  • تکیه به زمان به‌جای HB. «یک Thread.sleep هست، حتماً نخ دیگر تا حالا تمام کرده.» ‏sleep هیچ یال HB نمی‌سازد؛ خواندن کهنه همچنان کاملاً مجاز است.
  • volatile روی مرجع شیء تغییرپذیر. مرجع را امن منتشر می‌کند، اما تغییرهای بعدیِ فیلدهای همان شیء پوشش داده نمی‌شوند. به‌جایش یک عکس‌فوری تغییرناپذیر (immutable snapshot) منتشر کن.
  • فیلدهای غیر‌final در اشیای «تغییرناپذیر». فقط فیلدهای final تضمین انجماد سازنده را می‌گیرند؛ یک فیلد غیر‌final اگر شیء از راه مسابقهٔ داده منتشر شود، می‌تواند با مقدار پیش‌فرضش دیده شود.
  • قفل روی this در یک کتابخانه. فراخوانندگان می‌توانند روی شیء تو هم قفل بگذارند و مناقشه یا بن‌بستِ غافلگیرکننده بسازند. از یک قفل خصوصی استفاده کن.
  • check-then-act روی کالکشن‌های همروند.if (!map.containsKey(k)) map.put(k, v); مسابقه دارد؛ به‌جایش از putIfAbsent/computeIfAbsent استفاده کن.
  • فراموشیِ اینکه unlock() خودکار نیست. قفل‌های صریح Lock نیاز به try/finally دارند؛ بلوک synchronized نمی‌تواند نشت کند.
  • فرض اینکه size()/isEmpty() روی کالکشن‌های همروند دقیق‌اند. این‌ها عکس‌فوری‌های سازگارِ ضعیف (weakly consistent) هستند، نه اعداد لحظه‌ای دقیق.

بهترین شیوه‌ها

  • تغییرناپذیری (immutability) و محصورسازی (confinement، یعنی داده را در یک نخ نگه‌داشتن) را بر قفل‌گذاری ترجیح بده؛ ارزان‌ترین قفل، قفلی است که اصلاً نمی‌گیری.
  • پیش از دست‌ساز نوشتنِ قفل، از ابزارهای سطح‌بالا استفاده کن: کالکشن‌های java.util.concurrent، ‏CompletableFuture، ‏ExecutorService.
  • ناحیه‌های بحرانی را کوتاه و بدون I/O نگه دار. هرگز هنگام نگه‌داشتن قفل، کد بیگانه یا callback صدا نزن—این دعوت مستقیم به بن‌بست است.
  • یک ترتیب سراسری قفل (global lock ordering) برقرار و مستند کن و همه‌جا قفل‌ها را با همان ترتیب بگیر؛ این جلوی بن‌بست را می‌گیرد.
  • برای شمارنده‌ها و انباشتگرها زیر مناقشه، به‌جای AtomicLong سراغ LongAdder برو.
  • هر فیلدی که میان نخ‌ها می‌رود را عمداً منتشر کن—با final، ‏volatile، یک قفل، یا یک کالکشن همروند. اگر نمی‌توانی یال HB مربوطه را نام ببری، یعنی باگ داری.

پرسش‌های مصاحبه

پ۱: volatile چه تضمین‌هایی می‌دهد و عمداً چه نمی‌دهد؟

رؤیت‌پذیری و ترتیب از راه یک یال release (نوشتن) / acquire (خواندن): تمام نوشته‌های پیش از نوشتن volatile، پس از خواندن بعدیِ همان فیلد دیده می‌شوند—به‌علاوهٔ اتمیک‌بودن خودِ فیلد، شامل long/double ۶۴بیتی. اما انحصار متقابل نمی‌دهد و عملیات مرکبی مانند x++ را اتمیک نمی‌کند.

پ۲: happens-before را بدون گفتن «پیش‌تر در زمان» توضیح بده

یک ترتیب جزئی روی کنش‌هاست. اگر A HB B، اثرهای حافظهٔ A تضمیناً برای B دیده و پیش از آن مرتب‌اند. این رابطه میان یک release روی یک متغیر همگام‌سازی و یک acquire بعدی روی همان متغیر برقرار می‌شود، به‌علاوهٔ ترتیب برنامه، شروع/پیوستن نخ، و ترایایی. دو کنشی که با HB مرتب نیستند و نوشتن متعارض دارند، یک مسابقهٔ داده می‌سازند.

پ۳ (دام): حلقهٔ پرچم بدون volatile
boolean running = true;      // volatile نیست
void stop() { running = false; }
void run() { while (running) { /* کار */ } }

چه می‌تواند رخ دهد؟ JIT مجاز است running را از داخل حلقه در یک رجیستر بالا ببرد (بالابری حلقه / loop hoisting)، پس run() می‌تواند تا ابد حلقه بزند، حتی بعد از اینکه stop() برگشته—چون هیچ یال HB نوشتن را وادار به مشاهده نمی‌کند. ‏volatile کردن running مشکل را حل می‌کند.

پ۴ (دام): این چه چاپ می‌کند؟
static int a = 0, b = 0;
// نخ ۱: a = 1; int r1 = b;
// نخ ۲: b = 1; int r2 = a;

آیا r1 == 0 && r2 == 0 ممکن است؟ بله. بدون همگام‌سازی، بارگذاری‌ها و ذخیره‌ها می‌توانند بازچینش شوند (به‌خاطر store buffering)، پس هردو نخ می‌توانند مقدارِ پیش‌از‌نوشتنِ دیگری را بخوانند. سازگاری ترتیبی (sequential consistency) برای کد مسابقه‌دار تضمین نیست.

پ۵: دقیقاً چرا double-checked locking بدون volatile خراب است؟

helper = new Helper() یعنی «تخصیص + ساخت + انتساب»، و JMM اجازه می‌دهد انتساب مرجع پیش از نوشته‌های فیلدِ سازنده بازچینش (دیده) شود. یک خوانندهٔ مسابقه‌دار روی مسیر بدون‌قفل می‌تواند مرجع غیرتهی به یک شیء نیمه‌راه‌اندازی‌شده ببیند. ‏volatile مانع release/acquire را درج می‌کند که نوشته‌های سازنده را نسبت به خواندن مرجع HB می‌کند.

پ۶: چرا biased locking حذف شد و منطقش چه بود؟

حالت «همان نخ مکرراً دوباره قفل می‌کند» را بهینه می‌کرد، با مغرض‌کردن شیء به آن نخ و گذشتن کامل از CAS. اما هزینهٔ ابطال (revocation) و دفترداری‌اش با بارهای کاری مدرنِ چندنخی و ساختارهای دادهٔ همروند گران شد. ‏JEP 374 آن را در JDK 15 به‌طور پیش‌فرض غیرفعال کرد؛ بعداً حذف شد (JDK 18). قفل سبک (lightweight) و سنگین (inflated) باقی مانده‌اند.

پ۷: synchronized در برابر ReentrantLock — هرکدام کِی، و آیا نخ‌های مجازی پاسخ را عوض کردند؟

ReentrantLock امکانات tryLock، مهلت، قابلیت‌وقفه، انصاف و چند Condition را می‌افزاید؛ ‏synchronized ساده‌تر است و نمی‌تواند قفل را نشت دهد. از نظر تاریخی در کدِ پر از نخ مجازی ReentrantLock را ترجیح می‌دادی، چون synchronized نخ حامل را سنجاق (pin) می‌کرد—اما JEP 491 (JDK 24) این سنجاق را حذف کرد، پس اکنون انتخاب را بر پایهٔ امکانات و راحتی بگذار، نه مقیاس‌پذیری.

پ۸: AQS را در یک نفس توصیف کن

یک volatile int state به‌علاوهٔ یک صف انتظار FIFO مبتنی بر CLH. زیرکلاس‌ها تعریف می‌کنند state چه معنایی دارد و tryAcquire/tryRelease (یا گونه‌های اشتراکی) را پیاده می‌کنند؛ ‏AQS کارِ CAS، صف‌بندی و park/unpark را انجام می‌دهد. ‏ReentrantLock (‏state = شمار نگه‌داشت)، ‏Semaphore (‏permits) و CountDownLatch (شمارنده، با آزادسازی اشتراکی) همه بر آن ساخته شده‌اند.

پ۹ (سخت): مسئلهٔ ABA چیست و کِی واقعاً گاز می‌گیرد؟

‏CAS مقدارها را مقایسه می‌کند نه تاریخچه را. اگر مقدار A→B→A برود، CASی که انتظار A دارد موفق می‌شود، هرچند وضعیت میانی تغییر کرده. برای یک شمارندهٔ عددی بی‌ضرر است؛ اما برای بازاستفادهٔ اشاره‌گر/گره در ساختارهای بدون‌قفل خطرناک است—جایی که یک آدرسِ بازاستفاده‌شده از بررسیِ CAS می‌گذرد اما به پیوند کهنه ارجاع می‌دهد. با AtomicStampedReference (مقدار + نسخه) رفعش کن.

پ۱۰: false sharing چیست و چطور تشخیص/رفعش می‌کنی؟

دو متغیر مستقل که در همان خط کش ۶۴بایتی نشسته‌اند، در هر نوشتن باعث ابطال میان‌هسته‌ای می‌شوند و خط را میان هسته‌ها پینگ‌پنگ می‌کنند. با شمارنده‌های perf (مناقشهٔ خط کش / رویدادهای HITM) یا با پروفایلینگ تشخیصش بده؛ با بالشتک‌گذاری فیلدهای داغ روی خطوط جدا رفعش کن، مثلاً با @Contended (همان کاری که Cell در LongAdder می‌کند).

پ۱۱: کِی StampedLock از ReadWriteLock بهتر است و تله‌اش چیست؟

وقتی خواندن‌ها غالب‌اند و می‌توانی با خواندن خوش‌بینانه در مسیر خواندن اصلاً به خط کش قفل دست نزنی. تله‌ها: بازورودی نیست، ‏Condition ندارد، و خواندن خوش‌بینانه باید فیلدها را در متغیر محلی کپی و سپس validate() کند—خواندن از یک مرجعِ بالقوه‌کهنه در وسط خواندن، یک باگ است.

پ۱۲: آیا AtomicLong همیشه شمارندهٔ درست است؟

نه. زیر مناقشهٔ سنگین، حلقهٔ CAS-retry آن می‌کوبد و CPU هدر می‌رود. ‏LongAdder/LongAccumulator افزایش‌ها را روی سلول‌های padded رگه‌رگه می‌کنند و تنبل جمع می‌زنند، و برای شمارنده‌های write-hot و read-rare خیلی بهتر مقیاس می‌گیرند—به بهای sum() کمی گران‌تر و نبودِ یک read-modify-write اتمیک روی کل مقدار.

پ۱۳ (دام): آیا `long x; x = someLong;` میان نخ‌ها اتمیک است؟

برای یک long/double سادهٔ غیر‌volatile تضمین نیست—JLS اجازه می‌دهد نوشتن ۶۴بیتی به دو ذخیرهٔ ۳۲بیتی شکسته شود، پس یک خوانندهٔ مسابقه‌دار می‌تواند مقدار پاره‌شده (torn) ببیند—نیمی قدیم، نیمی جدید. ‏volatile، ‏AtomicLong یا یک قفل این پاره‌شدن را حذف می‌کند.

پ۱۴ (سخت): انتشار شیء تغییرناپذیر با فیلد final در برابر فیلد ساده، درون مسابقهٔ داده—فرق چیست؟

فیلدهای final یک انجماد ویژه در پایان سازنده می‌گیرند: هر نخی که مرجع شیء را بخواند، تضمیناً فیلدهای finalِ درست‌راه‌اندازی‌شده را می‌بیند، حتی زیر انتشار مسابقه‌دار. اما فیلد سادهٔ (غیر‌final) چنین تضمینی ندارد—نخ دیگر می‌تواند مقدار پیش‌فرضش (۰/null) را ببیند. برای همین «تغییرناپذیر = همهٔ فیلدها final» یک قاعدهٔ درستی است، نه سلیقه.

پ۱۵: چرا تنزل قفل کار می‌کند اما ارتقا در ReentrantReadWriteLock بن‌بست می‌شود؟

تنزل (گرفتن قفل خواندن هنگام نگه‌داشتنِ نوشتن، سپس آزادکردن نوشتن) امن است، چون هرگز لازم نیست منتظر رسیدنِ خوانندگان/نویسندگان دیگر به وضعیت قوی‌تر بمانی. اما ارتقا (نگه‌داشتنِ خواندن، درخواست نوشتن) نیازمند آزادکردن همهٔ خوانندگانِ دیگر است—و اگر دو خواننده هردو بخواهند ارتقا دهند، هرکدام منتظر می‌ماند دیگری قفل خواندنش را رها کند: بن‌بست. برای همین API آن را ممنوع می‌کند.

نکاتِ سنیور و موارد پیشرفته

تا اینجا مدل ذهنیِ درست را ساختی و ابزارها را می‌شناسی. حالا می‌رویم سراغ لایه‌ای که مرزِ یک سنیورِ واقعی از یک برنامه‌نویسِ خوب را مشخص می‌کند: تضمین‌های عمیق‌ترِ خودِ JMM، انتشار امن، نردبان access-modeهای مدرن با VarHandle، معناشناسیِ درستِ wait/notify، رده‌بندی تضمین‌های پیشرفت (progress)، و باگ‌هایی که فقط زیر بار تولید و با یک profiler خودشان را نشان می‌دهند.

نقشهٔ راهِ این بخش

۱) SC-DRF و مقادیرِ out-of-thin-air. ۲) تفاوتِ coherence و consistency. ۳) انتشار امن و فرارِ this. ۴) VarHandle و نردبانِ plain/opaque/acquire-release/volatile و fenceها. ۵) معناشناسیِ wait/notify/Condition و بیداریِ کاذب. ۶) رده‌بندیِ wait-free / lock-free / obstruction-free. ۷) deadlock/livelock/starvation و بهینه‌سازی‌های JITِ قفل. ۸) گاچاهای نامرئی و ابزارِ تست. در آخر ۹ پرسشِ سختِ مصاحبه.

۱) قضیهٔ SC-DRF: دقیقاً JMM چه قول می‌دهد؟

فصل نشان داد که کدِ ریسی می‌تواند مقادیرِ کهنه یا پاره ببیند. اما نکتهٔ آرامش‌بخش این است: JMM یک قضیهٔ محوری دارد به نامِ SC-DRF (Sequential Consistency for Data-Race-Free programs).

قولِ اصلیِ JMM را حفظ کن

اگر برنامه‌ات کاملاً همگام‌سازی‌شده باشد—هیچ جفت‌دسترسیِ متضاد (حداقل یکی نوشتن) بدونِ یالِ happens-before نداشته باشد، یعنی بدونِ مسابقهٔ داده (DRF)—آنگاه JMM تضمین می‌کند دقیقاً مثلِ یک برنامهٔ پیوستهٔ ترتیبی (sequentially consistent) رفتار می‌کند: یک ترتیبِ سراسریِ واحد که همهٔ نخ‌ها رویش توافق دارند. تا وقتی DRF باشی، تمام آن بازچینش‌های ترسناک نامرئی‌اند.

یعنی هدفت ساده می‌شود: لازم نیست دربارهٔ تک‌تکِ بازچینش‌ها فکر کنی، فقط اثبات کن DRF هستی—و آن جملهٔ طلاییِ فصل معنا می‌گیرد: «اگر نتوانی یالِ HB را نام ببری، باگ داری.» اما آن‌طرفِ سکه: برای کدِ racy، JMM هنوز یک مشکلِ حل‌نشدهٔ نظری دارد.

مقادیرِ Out-of-Thin-Air (از هیچ)

مدلِ رسمیِ JMM برای جلوگیری از مقادیرِ «از هیچ» (out-of-thin-air) قواعدِ علّیت دارد—نباید مقداری ظاهر شود که هیچ نوشته‌ای تولیدش نکرده و از یک حدسِ خودتوجیه‌گر بیرون آمده. اما ثابت شده فرمول‌بندیِ فعلی هم بعضی اجراهای مطلوب را ممنوع می‌کند و هم OoTA را کامل نمی‌بندد؛ این «سوراخِ» شناخته‌شدهٔ مدل و انگیزهٔ کار روی JMM جدید است. نتیجهٔ عملی: هرگز روی رفتارِ کدِ racy حساب نکن؛ حتی مقادیرِ «غیرممکن» هم مجازند.

۲) Coherence در برابر Consistency — یک اصلاحِ ظریف

تشبیهِ «کپیِ کهنه روی میزِ آشپز» برای شهود عالی است، ولی سنیور باید دقیق‌تر بداند. سخت‌افزارِ مدرن cache coherence را با پروتکل‌هایی مثل MESI تضمین می‌کند: برای یک آدرسِ واحد، همهٔ هسته‌ها روی یک ترتیبِ سراسریِ واحد از نوشتن‌ها توافق دارند و یک نوشتن، کپیِ هسته‌های دیگر را باطل (invalidate) می‌کند. پس کش‌ها واقعاً «برای همیشه» کهنه نمی‌مانند.

منبعِ واقعیِ کهنگی: store buffer و رجیستر، نه ناهماهنگیِ کش

پس چرا هنوز مقدارِ کهنه می‌بینی؟ دو دلیل: (۱) store buffer — نوشتنِ یک هسته لحظه‌ای در بافرِ محلی می‌نشیند و با تأخیر به کشِ منسجم می‌رسد؛ در این فاصله هستهٔ خودت مقدارِ جدید را می‌بیند ولی بقیه نه (همان «store buffering»ِ پشتِ سؤالِ کلاسیکِ r1==0 && r2==0). (۲) کامپایلر/JIT که مقدار را در رجیستر cache می‌کند (hoisting). پس مشکل coherence نیست بلکه consistency (ترتیبِ بینِ آدرس‌ها) است. volatile و قفل fence صادر می‌کنند که store buffer را تخلیه و بازچینش را مهار می‌کند—«کش را تازه» نمی‌کنند.

۳) انتشارِ امن (Safe Publication) و فرارِ this

فصل «انتشارِ عمدی» را به‌عنوان قاعدهٔ طلایی گفت. حالا دقیق‌تر: چهار راهِ متعارفِ انتشارِ امنِ یک شیء وجود دارد:

  1. مقداردهیِ اولیهٔ آن از یک بلاکِ static initializer (تضمینِ init کلاس).
  2. ذخیرهٔ ارجاعش در یک فیلدِ volatile (یا AtomicReference).
  3. ذخیرهٔ ارجاعش در یک فیلدِ final که در سازنده مقدار می‌گیرد.
  4. ذخیرهٔ ارجاعش با محافظتِ یک قفل (یا گذاشتنش در یک کالکشنِ concurrent).
فرارِ `this` در سازنده: باگی که پیش از پایانِ سازنده رخ می‌دهد

تضمینِ freezeِ فیلدهای final فقط وقتی معتبر است که ارجاعِ شیء قبل از پایانِ سازنده فرار نکند. اگر داخلِ سازنده this بیرون درز کند، این تضمین می‌شکند:

public class Listener {
    private final int id;
    public Listener(EventBus bus) {
        bus.register(this);   // ❌ this فرار کرد—هنوز سازنده تمام نشده
        this.id = computeId(); // نخِ دیگر ممکن است id را 0 ببیند
    }
}

الگوهای رایجِ فرارِ this: ثبتِ listener/callback، شروعِ نخ داخلِ سازنده، یا انتشارِ this در متغیرِ static. راهِ درست: سازنده را کامل کن، بعد در متدِ کارخانه‌ایِ static create(...) شیء را بساز و سپس register کن.

۴) VarHandle و نردبانِ access-modeها — مدلِ مدرنِ حافظه

فصل CAS را با AtomicInteger نشان داد. اما از Java 9 (JEP 193) ابزارِ سطحِ پایین و رسمیِ همهٔ اینها VarHandle است—جانشینِ امنِ sun.misc.Unsafe و Atomic*FieldUpdater (کلاس‌های java.util.concurrent خودشان از JDK 9 رویش مهاجرت کرده‌اند).

هدیهٔ اصلی‌اش: به‌جای دوگانهٔ «plain یا volatile»، یک نردبانِ چهارپله‌ای از قوّتِ ترتیب می‌دهد، از ضعیف به قوی.

نردبانِ قوّتِ access-mode در VarHandle (از ضعیف به قوی):

flowchart LR
  Plain["Plain<br/>get/set"] --> Opaque["Opaque<br/>getOpaque/setOpaque"]
  Opaque --> RelAcq["Release/Acquire<br/>setRelease/getAcquire"]
  RelAcq --> Volatile["Volatile (SC)<br/>getVolatile/setVolatile"]
  • Plain: بی‌ترتیب و بی‌رؤیت‌پذیریِ بینِ نخی—فقط اتمیکِ خودِ دسترسی (به‌جز long/double). مثلِ فیلدِ معمولی.
  • Opaque: دسترسی‌ها پاک نمی‌شوند و برای یک آدرس coherent می‌مانند، ولی نسبت به آدرس‌های دیگر بازچینش‌پذیر است—برای فلگِ توقف یا شمارندهٔ آماری که فقط باید «بالاخره» دیده شود، از volatile ارزان‌تر.
  • Release/Acquire: همان نیمه‌حصارِ release/acquireِ happens-before، ولی بدونِ حصارِ سنگینِ StoreLoad—«امضا/دیدنِ امضا» بی‌پرداختِ هزینهٔ کاملِ SC.
  • Volatile: قوی‌ترین—معادلِ کلیدواژهٔ volatile، با ترتیبِ سراسریِ SC.
class Node {
    Object item;
    volatile Node next;      // فیلدِ عادی، دستکاری با VarHandle
    private static final VarHandle NEXT;
    static {
        try {
            NEXT = MethodHandles.lookup()
                .findVarHandle(Node.class, "next", Node.class);
        } catch (ReflectiveOperationException e) { throw new ExceptionInInitializerError(e); }
    }
    boolean casNext(Node expect, Node update) {
        return NEXT.compareAndSet(this, expect, update);
    }
    void publish(Node n) { NEXT.setRelease(this, n); } // release، ارزان‌تر از setVolatile
}
compareAndExchange و CASِ ضعیف

دو نکتهٔ سنیور: (۱) compareAndExchange مثل compareAndSet است ولی به‌جای boolean مقدارِ شاهد (witness) را برمی‌گرداند—در حلقهٔ retry یک get()ِ اضافه را حذف می‌کند. (۲) weakCompareAndSet مجاز است کاذب (spuriously) شکست بخورد حتی وقتی مقدار برابرِ expected است؛ در عوض روی معماری‌های LL/SC مثلِ ARM ارزان‌تر است. برای همین فقط داخلِ حلقهٔ do/whileِ خودت درست است، نه به‌تنهایی.

VarHandle چهار متدِ استاتیکِ fence هم دارد (fullFence، acquireFence، releaseFence، loadLoadFence/storeStoreFence) که حصارِ حافظه صادر می‌کنند بدونِ گره‌خوردن به فیلدِ خاص—معادلِ مدرن و امنِ Unsafe.fullFence برای الگوهای انتشارِ دستیِ خیلی خاص.

۵) wait / notify / Condition — معناشناسیِ درست

فصل قفل و AQS را گفت اما به هماهنگیِ نخ‌ها با wait/notify نپرداخت—و این یکی از پرباگ‌ترین نقاطِ کدِ واقعی است.

اتاقِ انتظارِ مطب

wait() یعنی «کلید را پس بده و در اتاقِ انتظار بخواب» و notify() یعنی «یکی از خواب‌ها را صدا بزن». نکتهٔ حیاتی: نخِ بیدارشده بلافاصله اجرا نمی‌شود—باید دوباره برای قفل رقابت کند، پس بینِ «بیدار شدن» و «اجرا شدن» دنیا ممکن است عوض شده باشد.

سه قانونِ آهنین:

  1. wait/notify/notifyAll را فقط وقتی صدا بزن که monitorِ همان شیء را در دست داری، وگرنه IllegalMonitorStateException می‌خوری.
  2. همیشه wait را داخلِ یک حلقه بگذار که شرط را دوباره چک می‌کند، نه داخلِ if. دو دلیل: بیداریِ کاذب (spurious wakeup) که JLS صراحتاً مجازش می‌داند، و اینکه ممکن است تا وقتی نوبتت شود شرط دوباره باطل شده باشد.
synchronized (queue) {
    while (queue.isEmpty()) {   // ✅ while نه if
        queue.wait();
    }
    return queue.remove();
}
notify در برابر notifyAll و «سیگنالِ گم‌شده»

notify() فقط یک نخ را بیدار می‌کند و کدام را هم تضمین نمی‌کند. اگر چند نخ روی شرط‌های متفاوتِ همان monitor منتظر باشند، ممکن است نخِ «اشتباه» را بیدار کند که شرطش برقرار نیست؛ او دوباره می‌خوابد و سیگنال گم می‌شود—نخ‌های واجدِ شرایط هرگز بیدار نمی‌شوند (deadlock نرم). قاعده: مگر اینکه همهٔ waiterها هم‌ارز باشند، notifyAll؛ بهتر: با Condition (از ReentrantLock) صف‌های مجزا (notFull/notEmpty) بساز و فقط نخِ مرتبط را بیدار کن—هم امن‌تر، هم کارآمدتر.

park/unpark در برابر wait/notify

زیرِ AQS، مکانیزمِ خواباندن LockSupport.park()/unpark() است، نه wait/notify. تفاوتِ ظریف: unpark یک permit می‌دهد که انباشته نمی‌شود (حداکثر یکی) ولی می‌تواند قبل از park برسد—پس «سیگنالِ گم‌شده»ی wait/notify را ندارد. park هم می‌تواند به‌صورتِ کاذب برگردد، پس آن هم باید داخلِ حلقهٔ شرط باشد.

۶) رده‌بندیِ تضمین‌های پیشرفت: wait-free / lock-free / obstruction-free

فصل CAS را «lock-free» نامید. سنیور باید این اصطلاح را دقیق بداند چون در مصاحبه تله می‌گذارند.

سه سطحِ ضمانتِ پیشرفت
  • Obstruction-free: یک نخِ بدونِ مزاحمت در گام‌های محدود تمام می‌کند. (ضعیف‌ترین.)
  • Lock-free: همیشه حداقل یک نخ پیشرفت می‌کند—سیستم گیر نمی‌کند، هرچند یک نخِ بدشانس ممکن است تا ابد retry کند (starvation).
  • Wait-free: هر نخ در گام‌های محدود تمام می‌کند—بی‌starvation. (قوی‌ترین.)
حلقهٔ CAS، lock-free است نه wait-free

آن حلقهٔ do { ... } while(!cas())ِ فصل lock-free است نه wait-free: یکی همیشه برنده می‌شود، اما یک نخِ خاص می‌تواند بارها ببازد و گرسنه بماند. برای همین زیرِ contentionِ سنگین LongAdder بهتر است. و برخلاف تصورِ رایج «lock-free» یعنی «سریع‌تر» نیست—زیرِ بارِ سنگین یک قفلِ منصفانه ممکن است throughputِ بهتری بدهد. Thread.onSpinWait() (Java 9، JEP 285) هم در حلقه‌های spin کمک می‌کند: هینتِ PAUSE به CPU می‌دهد تا مصرفِ توان و ترافیکِ coherence کم شود.

۷) Deadlock در برابر Livelock در برابر Starvation و بهینه‌سازی‌های JIT

فصل «ترتیبِ سراسریِ قفل» را برای جلوگیری از deadlock گفت. سه شکستِ همروندی را از هم جدا کن:

  • Deadlock: نخ‌ها برای همیشه منتظرِ هم‌اند (چرخهٔ انتظار). چهار شرطِ Coffman لازم‌اند (mutual exclusion، hold-and-wait، no-preemption، circular-wait)؛ شکستنِ یکی کافی است—معمولاً circular-wait را با ترتیبِ سراسریِ قفل.
  • Livelock: نخ‌ها بلاک نیستند و مدام کار می‌کنند ولی پیشرفتی نمی‌کنند—مثلِ حلقهٔ tryLockای که هر بار شکست، قفل را رها و فوراً دوباره تلاش می‌کند؛ راهِ حل backoffِ تصادفی.
  • Starvation: بعضی نخ‌ها پیشرفت می‌کنند اما یکی هرگز نوبت نمی‌گیرد (writer starvation در ReadWriteLock).
بهینه‌سازی‌های JITِ قفل که باید بشناسی

سه بهینه‌سازیِ HotSpot که فصل نگفت: (۱) Lock elision — اگر escape analysis ثابت کند شیء از یک نخ فرار نمی‌کند، JIT قفلش را کاملاً حذف می‌کند (مثلِ StringBufferِ محلی). (۲) Lock coarsening — چند synchronizedِ پشتِ‌سرِ همِ روی یک شیء را در یک ناحیه ادغام می‌کند تا lock/unlockِ مکرر حذف شود. (۳) Adaptive spinning — قبل از park کمی spin می‌کند به امیدِ آزادیِ زودِ قفل و مدتش را از تاریخچهٔ همان قفل تنظیم می‌کند. یعنی micro-benchmarkِ ساده اغلب هزینهٔ واقعیِ قفل را نشان نمی‌دهد.

۸) گاچاهای نامرئی که در تولید می‌زنند

`volatile` روی آرایه فقط ارجاع را پوشش می‌دهد، نه عناصر را
volatile int[] data;   // فقط خودِ ارجاعِ آرایه volatile است
data[5] = 42;          // ❌ این نوشتن هیچ معناشناسیِ volatile ندارد!

یک اشتباهِ فاجعه‌بار: نوشتن در data[i] یک نوشتنِ عادی است و هیچ یالِ HB نمی‌سازد. اگر عناصر باید ترتیب/رؤیت‌پذیریِ volatile داشته باشند، از AtomicIntegerArray یا VarHandle رویِ آرایه (MethodHandles.arrayElementVarHandle) استفاده کن.

نشتِ ThreadLocal در thread pool

ThreadLocal مقدار را به نخ گره می‌زند؛ در thread pool نخ‌ها بازاستفاده می‌شوند و نمی‌میرند. اگر بعد از هر task مقدار را remove() نکنی: (۱) داده‌ی یک درخواست به درخواستِ بعدیِ همان نخ نشت می‌کند (باگِ امنیتی/صحت)، و (۲) شیءِ بزرگ تا ابد زنده می‌ماند (memory leak). همیشه در finally یا فیلترِ خروجی remove() کن.

کالکشن‌های concurrent خودشان یالِ HB می‌سازند

یک اصلِ مستند اما کم‌دانسته: کلاس‌های java.util.concurrent تضمینِ «memory consistency» می‌دهند—گذاشتنِ یک شیء در BlockingQueue happens-before برداشتنش با take()، و نوشتن در ConcurrentHashMap قبل از خواندنِ بعدیِ همان کلید HB است. یعنی وقتی شیءِ mutable را از میانِ صفِ concurrent رد می‌کنی، به volatileِ جدا نیاز نداری—خودِ صف انتشارِ امن را انجام می‌دهد.

۹) چطور همروندی را واقعاً تست کنیم — jcstress و JMH

نوشتنِ کدِ همروند بدونِ تستِ درست خودفریبی است—باگ ممکن است روی x86ات دیده نشود ولی روی ARM یا زیرِ بار بترکد. jcstress (رسمیِ OpenJDK) litmus test می‌سازد و میلیون‌ها بار اجرا می‌کند تا بازچینش‌های نادر را شکار کند—برای اثباتِ DRF بودنِ یک الگو بی‌بدیل. و JMH برای بنچمارکِ درست که lock elision/coarsening و warmupِ JIT را لحاظ می‌کند (یک System.nanoTime()ِ دستی تقریباً همیشه دروغ می‌گوید).

پرسش ۱ (سخت): JMM دقیقاً برای یک برنامهٔ درست‌همگام چه تضمینی می‌دهد؟ SC-DRF را توضیح بده.

قضیهٔ SC-DRF: اگر برنامه data-race-free باشد (هر جفت‌دسترسیِ متضاد با یالِ happens-before مرتب)، JMM تضمین می‌کند اجرا پیوستهٔ ترتیبی به‌نظر می‌رسد—یک ترتیبِ سراسریِ واحد و همهٔ بازچینش‌ها نامرئی. پس به‌جای استدلال دربارهٔ تک‌تکِ بازچینش‌ها، فقط اثبات کن DRF هستی. آن‌طرفِ سکه: برای کدِ racy مدل حتی «out-of-thin-air» را کامل نمی‌بندد و سوراخِ نظری دارد؛ هرگز روی رفتارِ racy حساب نکن.

پرسش ۲ (سخت): آیا دیدنِ مقدارِ کهنه یعنی cache ناهماهنگ است؟

نه. سخت‌افزار cache coherence را (با MESI) تضمین می‌کند: برای یک آدرس یک ترتیبِ سراسریِ نوشتن هست و یک نوشتن کپیِ بقیه را باطل می‌کند. کهنگی از store buffer (نوشتنِ محلیِ هنوز-flush-نشده) و cacheکردنِ JIT در رجیستر می‌آید. پس مشکلْ coherence نیست بلکه memory consistency (ترتیبِ بینِ آدرس‌ها) است. volatile/قفل با fence، store buffer را تخلیه و بازچینش را مهار می‌کنند؛ «کش را رفرش» نمی‌کنند.

پرسش ۳ (سخت): نردبانِ access-modeهای VarHandle را توضیح بده؛ کِی opaque، کِی release/acquire، کِی volatile؟

چهار پله از ضعیف به قوی: plain (فقط اتمیک، بی‌ترتیب)، opaque (coherence و progress برای همان متغیر، ولی نسبت به بقیه بازچینش‌پذیر—برای فلگِ توقف یا شمارندهٔ آماری ارزان‌تر از volatile)، release/acquire (نیمه‌حصار؛ همان HBِ امضا/دیدنِ امضا بدونِ حصارِ سنگینِ StoreLoad)، و volatile (قوی‌ترین، SC سراسری). قاعده: ضعیف‌ترین را بگیر که درستیِ الگو را تأمین کند؛ در ساختارهای lock-free معمولاً release برای publish و acquire برای خواندن کافی است.

پرسش ۴ (گاچا): چرا باید `wait()` همیشه داخلِ حلقه باشد نه `if`؟

دو دلیل: (۱) بیداریِ کاذب (spurious wakeup) که JLS صراحتاً مجاز می‌داند—نخ می‌تواند بدونِ هیچ notifyای بیدار شود. (۲) بینِ لحظهٔ بیدار شدن و لحظهٔ واقعیِ گرفتنِ قفل، نخِ دیگری ممکن است شرط را دوباره باطل کرده باشد (مثلاً صف که خالی نبود، دوباره خالی شد). پس باید بعد از بیداری شرط را دوباره چک کنی؛ فقط while (!condition) wait(); امن است. همین دلیل برای Condition.await() و LockSupport.park() هم صادق است.

پرسش ۵ (سخت): تفاوتِ `notify` و `notifyAll` و «سیگنالِ گم‌شده» چیست؟

notify() یک نخِ نامشخص را بیدار می‌کند. اگر waiterها روی شرط‌های متفاوت منتظر باشند، ممکن است نخی بیدار شود که شرطش برقرار نیست، دوباره بخوابد و سیگنال گم شود—نخی که باید بیدار می‌شد هرگز بیدار نمی‌شود (deadlockِ نرم/lost-wakeup). پیش‌فرض notifyAll مگر اینکه همهٔ waiterها هم‌ارز باشند؛ بهتر: Conditionهای مجزا (notFull/notEmpty) و signal فقط نخِ مرتبط.

پرسش ۶ (سخت): wait-free، lock-free و obstruction-free را تعریف کن؛ حلقهٔ CAS کدام است؟

obstruction-free: یک نخِ بدونِ مزاحمت در گام‌های محدود تمام می‌شود. lock-free: همیشه حداقل یک نخ پیشرفت می‌کند (سیستم گیر نمی‌کند) اما یک نخِ خاص می‌تواند گرسنه بماند. wait-free: هر نخ در گام‌های محدود تمام می‌شود (بی‌starvation). حلقهٔ do{}while(!cas()) lock-free است نه wait-free—یکی همیشه برنده می‌شود ولی یک بدشانس بارها می‌بازد. پس زیرِ contentionِ سنگین LongAdder؛ و «lock-free» لزوماً «سریع‌تر» نیست.

پرسش ۷ (سخت): فرارِ `this` در سازنده چیست و چرا حتی بدونِ نخِ آشکار خطرناک است؟

اگر ارجاعِ شیء قبل از پایانِ سازنده درز کند (ثبتِ listener، start کردنِ نخ، انتشار در فیلدِ static)، تضمینِ freezeِ final می‌شکند: نخِ دیگری می‌تواند فیلدهای مقداردهی‌نشده (0/null) ببیند، چون نوشتن‌های سازنده کامل نشده‌اند. حتی تک‌نخی هم خطرناک است چون کدِ کال‌بک ممکن است روی شیءِ نیمه‌ساخته کار کند. راهِ درست: سازنده را کامل کن و انتشار را به متدِ کارخانه‌ایِ static بعد از ساخت منتقل کن.

پرسش ۸ (گاچا): `volatile int[] a; a[i] = x;` چه معناشناسی‌ای دارد؟

فقط ارجاعِ آرایه volatile است نه عناصرش. نوشتنِ a[i] یک نوشتنِ عادی است و هیچ یالِ happens-before نمی‌سازد—نخِ دیگر ممکن است مقدارِ کهنه ببیند یا اصلاً نبیند. برای عناصر از AtomicIntegerArray یا MethodHandles.arrayElementVarHandle استفاده کن. همین تله برای AtomicReference<mutableObject> هم هست: ارجاع امن منتشر می‌شود ولی دستکاریِ بعدیِ فیلدهای شیء پوشش نمی‌یابد.

پرسش ۹ (سخت): تفاوتِ livelock و deadlock چیست و چطور رفعشان می‌کنی؟

deadlock: نخ‌ها بلاک و در چرخهٔ انتظارِ متقابل‌اند؛ چهار شرطِ Coffman لازم و شکستنِ یکی (معمولاً circular-wait با ترتیبِ سراسریِ قفل) کافی است. livelock: نخ‌ها بلاک نیستند و مدام کار می‌کنند ولی پیشرفتی نمی‌کنند—مثلِ حلقهٔ tryLockای که هر بار شکست، رها و فوراً دوباره تلاش می‌کند. رفع: backoffِ تصادفی. هر دو با thread dump و detectِ چرخه عیب‌یابی می‌شوند، ولی جلوگیری بهتر از درمان است.

در یک نگاه (سنیور)
  • SC-DRF: اگر DRF باشی دنیا پیوستهٔ ترتیبی به‌نظر می‌رسد؛ برای کدِ racy حتی OoTA بسته نیست. کهنگی از store buffer و رجیستر می‌آید نه ناهماهنگیِ کش (consistency نه coherence).
  • انتشارِ امن چهار راه دارد و با فرارِ this در سازنده می‌شکند. VarHandle نردبانِ plain/opaque/release-acquire/volatile را می‌دهد—ضعیف‌ترینِ درست را بگیر.
  • wait را همیشه در حلقه بگذار؛ پیش‌فرض notifyAll یا Conditionهای مجزا. حلقهٔ CAS lock-free است نه wait-free؛ زیرِ بار LongAdder.
  • deadlock/livelock/starvation را جدا کن؛ JIT قفل را elide/coarsen می‌کند. گاچاها: volatile آرایه فقط ارجاع، نشتِ ThreadLocal در pool، و HB رایگانِ کالکشن‌های concurrent. با jcstress/JMH تست کن.
جمع‌بندی
  • حافظه در سطح چندهسته‌ای مثل دفترچهٔ مشترک نیست؛ کامپایلر، CPU و کش، کدت را بازچینش می‌کنند. ‏JMM (فصل ۱۷ JLS، بازنویسی‌شده با JSR-133 در جاوا ۵) قرارداد رسمیِ این دنیاست.
  • happens-before تنها قاعده‌ای است که اهمیت دارد: دربارهٔ ترتیب است نه زمان، و فقط میان یک release و یک acquire بعدی روی همان متغیر همگام‌سازی برقرار می‌شود.
  • volatile رؤیت‌پذیری/ترتیب و اتمیک‌بودنِ فیلد (حتی ۶۴بیتی) می‌دهد، اما انحصار متقابل نمی‌دهد و x++ را اتمیک نمی‌کند.
  • synchronized انحصار + رؤیت‌پذیری + بازورودی می‌دهد و کلید را نشت نمی‌دهد؛ ‏biased locking حذف شده (JEP 374 در JDK 15، حذف در JDK 18)، و از JDK 24 با JEP 491 دیگر نخ مجازی را سنجاق نمی‌کند.
  • خانوادهٔ Lock انعطاف می‌افزاید: ‏ReentrantLock (‏tryLock/مهلت/انصاف/Condition)، ‏ReadWriteLock (تنزل مجاز، ارتقا ممنوع)، ‏StampedLock (خواندن خوش‌بینانه، اما بدون بازورودی).
  • AQS موتور مشترک زیر همهٔ این‌هاست: ‏volatile int state + صف CLH + ‏tryAcquire.
  • CAS برنامه‌نویسی بدون‌قفل می‌سازد اما مسئلهٔ ABA دارد (با AtomicStampedReference رفع کن)؛ زیر مناقشهٔ سنگین سراغ LongAdder برو.
  • مراقب false sharing (با @Contended pad کن) و DCL بدون volatile باش؛ برای singleton از holder idiom استفاده کن.
  • قانون طلایی: هر فیلد بین‌نخی را عمداً منتشر کن. اگر نمی‌توانی یال HB را نام ببری، باگ داری.

Let's be honest: most concurrency bugs happen because a programmer assumes memory behaves like a shared notebook that every thread reads from the same page at the same instant. On a real multi-core machine, memory does not work like that at all. This chapter is going to fully break that wrong mental model and build a correct one in its place — and then construct, one at a time and from scratch, the tools you use to control ordering and visibility.

Roadmap for this chapter

First we see why the JMM exists at all and the three agents that reorder your code. Then we reach the heart of everything: happens-before, the only rule that guarantees one thread sees another's writes. Next we build the tools from light to heavy: volatile, synchronized and its lock optimizations, the Lock family (ReentrantLock / ReadWriteLock / StampedLock), the underlying engine AQS, lock-free programming with CAS and the ABA problem, false sharing, and finally the famous double-checked locking idiom. We close with a full interview questions section and a nutshell summary.

Part 0 — words you must know first

Before anything else, let's anchor a few terms with analogies so I never drop them on you cold later.

Thread, core, and cache — like several chefs in one kitchen

Think of a thread as a chef working through your recipe (your code). A multi-core CPU means several chefs cooking at once. Here's the key detail: each chef has a little countertop right beside them (their core's cache), while the main pantry (RAM) is across the kitchen. When a chef grabs the salt, they put a copy on their own counter and use that copy for a while. So if another chef changes the salt in the main pantry, the first chef won't notice for a while — they still see the stale copy on their counter. That stale copy is the root of most visibility bugs in Java.

  • Reordering: shuffling the execution order of instructions for speed — like a chef doing tasks in the most efficient order rather than the written order.
  • Memory barrier / fence: a special instruction telling the hardware "actually finish everything up to here and make it visible before moving on" — like a rule that says "don't start the next dish until the current one is on the pass."
  • Atomic: an operation that either happens completely or not at all; no other thread can observe a half-done intermediate state.

Mental model: why the JMM exists

Naively you imagine your program as a single, sequential list of memory reads and writes executed exactly in source order, instantly visible to every thread. Convenient as it is, that model is a lie. Between your source code and the electrons in the chip, there are three reordering agents:

  1. The compiler (javac + JIT): may reorder, hoist, eliminate, or fold instructions. "Hoist" means taking something read every loop iteration and reading it once outside the loop into a register — you'll soon see why that can be catastrophic.
  2. The CPU: executes out-of-order and sometimes speculatively (it guesses which branch runs and starts it early).
  3. The memory hierarchy: store buffers and per-core caches delay when one core's write becomes visible to another — that "stale copy on the chef's counter."
These aren't bugs — they're the source of speed

You're tempted to think these three agents are "saboteurs." The opposite: all modern performance comes from exactly these reorderings. The platform makes you only one promise: for single-threaded code it guarantees as-if-serial semantics — any reordering is allowed as long as it never changes the result you would observe in that same thread. The trouble starts precisely when another thread wants to watch you mid-work.

This is where the Java Memory Model (JMM) enters. It's specified in JLS Chapter 17 and was reworked by JSR-133 (Java 5), and it's the contract that tells you exactly which cross-thread reads are allowed to see which writes. Everything else in this chapter — volatile, synchronized, Lock, atomics — is really a way to buy happens-before edges from that contract.

Happens-before: the only rule that really matters

If you take just one thing from this chapter, make it this.

happens-before is like signing for a delivery

Imagine two departments in a company. Department A prepares a document and Department B must work on it. With no procedure between them, B might grab a half-finished draft or even yesterday's version. But if the rule is "A puts the document in the shared inbox and signs it, and B only picks it up after seeing that signature," then everything A wrote before signing is guaranteed visible to B who reads after the signature. That pairing — "sign" (release) and "see the signature" (acquire) — is exactly a happens-before edge.

Now let's be precise. The JMM is defined in terms of a partial order called happens-before (HB). "Partial" means not every two actions are comparable; some pairs are ordered and some aren't. The rule: if action A happens-before action B, then all of A's memory effects are visible to and ordered before B.

And here's the dangerous part: if two actions are not ordered by HB and at least one is a write to the same location, you have a data race. In that case the JMM permits the reads to return stale, torn (half the old value, half the new), or even seemingly "impossible" values.

The HB edges you actually get — memorize this list:

  • Program order: within a single thread, each action HB every later action in that thread.
  • Monitor lock: an unlock on a monitor HB every subsequent lock on that same monitor.
  • Volatile: a write to a volatile field HB every subsequent read of that same field.
  • Thread start: Thread.start() HB every action in the started thread.
  • Thread join: every action in a thread HB another thread's successful return from join() on it.
  • Interrupt: a call to interrupt() HB the interrupted thread detecting it.
  • Final fields: the end of a constructor HB the freeze of final fields (the basis for safe publication of immutable objects).
  • Transitivity: if A HB B and B HB C, then A HB C. Edges chain together.
happens-before is not about time, it's about ordering

Let's bury the biggest misconception right here. Two events can be perfectly "simultaneous" in wall-clock terms; what matters is not which happened "earlier" but whether the model forces one's writes to be seen by the other. And that forcing is established only between a release action and a subsequent acquire action on the same synchronization variable. Releasing lock A tells you nothing about a thread that acquires lock B — inbox A's signature is useless for inbox B.

Let's watch the chain with a concrete example:

Thread 1                         Thread 2
--------                         --------
data = 42;        (1)
ready = true;     (2, volatile write / release)
                                 while(!ready) {}  (3, volatile read / acquire)
                                 print(data);      (4)  -> guaranteed to see 42

Why does it work? Because (2) is a volatile write and (3) a volatile read of the same field, (2) HB (3). By program order (1) HB (2), and also (3) HB (4). Now chain via transitivity: (1) HB (2) HB (3) HB (4), so (1) HB (4). That means the read of data at line (4) must observe 42.

What happens if you remove volatile?

Remove volatile from ready and two things break at once: the visibility of data (Thread 2 may see a stale value) and the termination of the loop. The JIT is allowed to hoist ready out of the loop into a register — then while(!ready) keeps seeing the same register value and loops forever, even after Thread 1 changed the value. This is one of the most common real-world bugs.

volatile: what it does and what it doesn't

Now that you have happens-before, volatile becomes simple. Picture it as a glass jar everyone must read directly from the main pantry, never from their counter copy — and every time someone puts something in it, they must first commit all their prior work to the pass.

volatile gives you exactly two things:

  1. Visibility + ordering (release/acquire): a volatile write establishes an HB edge, making not just that variable but all writes that preceded it in program order visible to a thread that subsequently reads the same volatile. This "piggybacking" is what made the data/ready example work.
  2. Atomicity of the field itself, even for long/double. Note: under the JLS, a non-volatile 64-bit write may be split into two 32-bit stores; volatile forbids that tearing.

And what volatile does not give you — this is where most people go wrong:

  • No mutual exclusion. volatile int x; x++; is really a read-modify-write (read, add, write) and is not atomic. Two threads can both read 5, both compute 6, both write 6 — and one increment is lost. Use an AtomicInteger or a lock.
  • No compound-action atomicity. if (v == null) v = new X(); still races even if v is volatile, because there's a gap between the check and the assignment.
Under the hood: why x86 is "cheap" but ARM isn't

On x86, a volatile write compiles to a plain store followed by a store-load barrier (often a lock-prefixed instruction or mfence); but reads are essentially free, because x86 has a strong memory model called TSO (Total Store Order). On weaker memory models like ARM and Power, both reads and writes must emit real fence instructions. So the same Java code may cost more on an ARM phone than on an x86 server.

synchronized: the intrinsic monitor

If volatile was the shared jar, synchronized is the key to a room only one person enters at a time.

A monitor is like a single-occupancy restroom with a key

Every Java object carries an invisible lock called its monitor (intrinsic lock) — just like a restroom with the key hanging on the door. To enter you take the key (lock), and on exit you hang it back (unlock). While the key is in your hand, nobody else can enter. A synchronized block does this "take and hang the key" automatically — even if an exception is thrown mid-work, the exit behaves like a finally and returns the key, so you can never "leak" it.

The guarantees of synchronized:

  • Mutual exclusion: only one thread holds a given monitor at a time.
  • Visibility: unlock HB the subsequent lock on the same monitor, so the whole critical section's writes publish.
  • Reentrancy: the same thread may re-acquire a monitor it already holds (a per-thread hold count is kept). That's what lets a synchronized method call another synchronized method on this without deadlocking itself.

A subtle point about what you lock on: synchronized(this) and a synchronized instance method both lock on this; a synchronized static method locks on the Class object.

Never lock on a String or a boxed Integer

synchronized("lock") or locking on an Integer that came from boxing is a classic trap. String literals are interned into a shared pool, and small Integers are cached, so that exact object may also be used by completely unrelated code — and suddenly two unrelated parts of your program deadlock on a shared lock without knowing it. Always create a dedicated private lock: private final Object lock = new Object();.

Lock optimizations — and what got removed

Historically HotSpot layered three locking schemes worth knowing, especially since one was removed and is a favorite interview question:

  • Biased locking: assumed a lock is usually re-acquired by the same thread, so it "biased" the object to that thread and skipped atomic CAS entirely on re-entry. It was disabled by default and deprecated in JDK 15 (JEP 374) and the implementation was subsequently obsoleted/removed (JDK 18, JDK-8256425). Why? Its revocation and bookkeeping costs hurt modern workloads with lots of short-lived threads and contended concurrent data structures.
  • Lightweight (thin) locking: uncontended locks use a CAS on the object header's mark word to stack-allocate a lock record on the thread's stack — no OS mutex, so very cheap.
  • Heavyweight (inflated) locking: under contention the monitor "inflates" to an OS-level ObjectMonitor with a real wait queue and OS parking. Expensive, but needed under pressure.
The modern note: JEP 491 and virtual threads

A senior-level point you must know: as of JDK 24, JEP 491 makes synchronized no longer pin virtual threads. Before 24, when a virtual thread blocked inside a synchronized block, it held its carrier platform thread hostage — this is called pinning — throttling scalability. That's why you were advised to prefer ReentrantLock in virtual-thread-heavy code. From JDK 24 on that advice is obsolete: now choose synchronized vs. java.util.concurrent.locks purely on ergonomics and features, not scalability.

The Lock family: ReentrantLock, ReadWriteLock, StampedLock

The intrinsic monitor is simple and safe, but rigid: you can't say "give up if the lock isn't free within 2 seconds." The java.util.concurrent.locks package gives explicit locks with that flexibility.

ReentrantLock — the workhorse

ReentrantLock provides the same mutual exclusion as synchronized, but with extra capabilities: tryLock() (with an optional timeout — "if it isn't free by then, come back"), interruptible acquisition via lockInterruptibly(), optional fairness (FIFO ordering for waiting threads, at a throughput cost), and multiple Condition objects per lock (vs. one wait-set per monitor).

But it has one non-negotiable idiom — know it in your sleep:

private final ReentrantLock lock = new ReentrantLock();

void doWork() {
    lock.lock();
    try {
        // critical section
    } finally {
        lock.unlock(); // MUST be in finally — an exception must not leak the lock
    }
}
Key difference from synchronized: unlock is not automatic

With synchronized the key always returns on its own. With an explicit Lock, if an exception is thrown between lock() and unlock() and you didn't put unlock() in a finally, the lock stays held forever and any other thread that wants it waits forever. Always unlock() in finally.

ReentrantReadWriteLock — separate read/write locks

A classroom whiteboard

Picture a whiteboard. Reading means looking at it — ten people can look at once, no problem. Writing means erasing and rewriting — while writing, nobody else should write or even read, or they'd see a half-erased mess. ReentrantReadWriteLock is exactly this: many threads take the read lock at once, but the write lock is exclusive.

This is good when reads massively dominate writes and critical sections are non-trivial (not tiny). But two important traps:

  • Writer starvation: if readers pour in without pause, a writer may never get a turn. Mitigate with the fair constructor.
  • Downgrading allowed, upgrading forbidden: downgrading (write → read) is legal — provided you acquire the read lock before releasing the write lock. But upgrading (read → write) deadlocks and is forbidden. (We fully unpack why in the last interview question.)

StampedLock (Java 8) — optimistic reads

StampedLock is not reentrant, but adds a killer feature: optimistic reads.

A quick glance at the wall clock

Suppose you want the time. The pessimistic way is to stand in front of the clock and stop anyone from moving the hands (a read lock). The optimistic way is to just take a quick glance, memorize the number, then check "did anyone touch the clock in the meantime?" If not, your number is valid and you never disturbed anyone. If yes, then you reluctantly go stand in front of the clock. That quick glance is tryOptimisticRead and that check is validate.

private final StampedLock sl = new StampedLock();
private double x, y;

double distanceFromOrigin() {
    long stamp = sl.tryOptimisticRead();      // no CAS, no blocking
    double cx = x, cy = y;                     // read snapshot
    if (!sl.validate(stamp)) {                 // a writer may have run
        stamp = sl.readLock();                 // fall back to a pessimistic read lock
        try { cx = x; cy = y; }
        finally { sl.unlockRead(stamp); }
    }
    return Math.sqrt(cx * cx + cy * cy);
}

The big win: when writes are rare, the read path never even touches the lock's cache line, so cache-line contention is eliminated entirely.

StampedLock traps

Remember three things: (1) it is not reentrant — if the same thread re-locks, it self-deadlocks. (2) It does not support Condition. (3) It is not directly interruptible (use the Interruptibly variants). And the most important rule: in an optimistic read you must copy fields into locals before validate(), and you must not call other methods or dereference a potentially-inconsistent reference mid-read — that reference could be half-built.

AbstractQueuedSynchronizer (AQS): the engine underneath

Now let's pull back the curtain. ReentrantLock, Semaphore, CountDownLatch, ReentrantReadWriteLock, and even ThreadPoolExecutor's worker gate — all are thin wrappers over one shared engine called AQS. Understand this one piece and you practically know the source of half of java.util.concurrent.

AQS is like a bank's take-a-number system

Picture a bank with a "now serving" display (state) and an orderly queue of customers who took a number. When the window is free, they call the front of the line. When busy, newcomers take a number and sit down / go to sleep (park) until their turn. When the current customer finishes, they wake exactly the next person (unpark). AQS is that system: one shared status number, plus an orderly queue of sleeping threads.

The precise mental model:

  • AQS holds a single volatile int state and a CLH-based FIFO wait queue of threads. (CLH is a particular linked-list-based queue where each thread spins/sleeps on its own node.)
  • A subclass defines what state means and implements tryAcquire(int) / tryRelease(int) (exclusive) or tryAcquireShared / tryReleaseShared (shared).
  • AQS itself does the hard machining: atomically CAS-ing state, enqueuing losers as queue nodes, parking them via LockSupport.park(), and unparking the successor on release.

Now see how the same engine builds several different classes:

  • ReentrantLock: state is the hold count — 0 = free, N = held N times reentrantly. tryAcquire CAS-es 0→1, or, if the current thread already owns it, bumps the count.
  • Semaphore: state is the permit count, using shared acquisition.
  • CountDownLatch: state is the count, and when it hits 0 it lets all waiters through at once — this is the "shared release," which is why one latch can release many threads simultaneously.
Why AQS matters

Once you understand "state + CLH queue + tryAcquire," you no longer memorize each synchronizer class separately — you just ask "what does this class make state mean, and what does its tryAcquire do?" and the rest of the behavior (queue, park, unpark) comes from the shared engine.

CAS, the Atomic* classes, and the ABA problem

So far everything was about locks. But there's a layer above them — and often faster: lock-free programming.

CAS is like editing a shared doc with "did anyone change it?"

Suppose you're editing a shared document. Instead of locking it, you do this: read the current version, prepare your change, and on save say "only apply my change if the doc is still exactly what I read; otherwise report failure." If someone changed the doc in the meantime, your save is rejected and you retry from scratch. That's compare-and-swap (CAS).

CAS is a single hardware instruction (lock cmpxchg on x86, LL/SC on ARM) that atomically does: "if this memory equals expected, set it to new and report success; otherwise leave it and report failure." The java.util.concurrent.atomic package exposes it: AtomicInteger, AtomicLong, AtomicReference, and array/field-updater variants.

AtomicInteger counter = new AtomicInteger();

int incrementAndGet() {
    int prev, next;
    do {
        prev = counter.get();
        next = prev + 1;
    } while (!counter.compareAndSet(prev, next)); // retry until we win the race
    return next;
}

This loop is optimistic: no thread blocks; contention just causes retries.

When CAS wins and when it loses

Under low contention, this loop crushes locks — because nobody ever sleeps and wakes. But under heavy contention a retry storm kicks in: dozens of threads keep beating each other and the CPU burns without doing useful work. That is exactly why Java 8 added LongAdder/LongAccumulator, which stripe the count across multiple separate cells (padded to avoid false sharing) and sum only on demand. For hot counters it's vastly better than a single AtomicLong.

The ABA problem

CAS has a subtle weakness: it checks value equality, not whether the value changed and changed back.

The spare key and the clever thief

You have a safe locked by a number. Your rule is "if the number is still A, open it." A clever thief changes the number from A to B and back to A. Now when you check "is it still A?", yes it is — so you open it, unaware that everything changed in between. The value is the same, but the world moved.

Precisely: if a thread reads A, another thread flips A→B→A, then the first thread's compareAndSet(A, ...) succeeds — even though the world moved underneath it. For an int counter this is harmless. But for a pointer-swapping structure (say a lock-free stack popping a node that was freed and reallocated in the meantime) it corrupts state. The fix is to pair the value with a monotonic stamp/version: AtomicStampedReference (value + int stamp) or AtomicMarkableReference (value + boolean).

False sharing: the invisible tax

This is one of those performance bugs that leaves no trace in the code and only a profiler finds it.

One shared tray for two people

Caches move memory not variable-by-variable but in 64-byte packages called cache lines. Now imagine two people each have their own cup on a single shared tray. The cups are independent, but because they're on one tray, every time one person moves their cup they must pick up the whole tray, and the other must wait for it to come back. Two independent variables that happen to sit on the same cache line are exactly this: every write to one invalidates the other core's copy, and the line "ping-pongs" between cores even though there's no logical sharing.

This is false sharing, and it can silently cost you an order of magnitude (10x).

The fix: pad hot fields onto their own cache line so each gets its own tray. Java 8+ provides @jdk.internal.vm.annotation.Contended (application code needs the -XX:-RestrictContended flag); LongAdder's internal Cell is annotated exactly this way — which is why LongAdder scales so well. Manual padding (long p1..p7) also works, but the JIT may eliminate unused fields, so @Contended is preferred where available.

Double-checked locking, done right

This is the most famous "looks correct but is broken" pattern in Java concurrency. The goal: build an expensive object only once and lazily, then return it quickly without locking.

The classic broken idiom:

// BROKEN before Java 5 and still broken today without volatile
private Helper helper;
Helper getHelper() {
    if (helper == null) {                 // 1st check (no lock)
        synchronized (this) {
            if (helper == null)           // 2nd check (locked)
                helper = new Helper();    // publish
        }
    }
    return helper;
}

Why is it broken without volatile? Because helper = new Helper() is not one atomic operation but three steps: (1) allocate memory, (2) run the constructor and initialize fields, (3) assign the reference to helper. And the JMM permits step 3 (the reference assignment) to become visible before step 2 (the constructor's writes).

A partially constructed object: the scary bug

Imagine a second thread on the fast "1st check" path. Because of that reordering, this thread can see a non-null helper that still points at a partially constructed object — the reference has landed, but the internal fields still hold default values (0/null). The second thread returns it and uses it, and you have a completely unreproducible bug.

The fix is to make the field volatile. That inserts the release/acquire barrier which guarantees the constructor's writes HB any read of the reference:

private volatile Helper helper;           // volatile is mandatory
Helper getHelper() {
    Helper result = helper;               // read volatile once into a local
    if (result == null) {
        synchronized (this) {
            result = helper;
            if (result == null)
                helper = result = new Helper();
        }
    }
    return result;
}

Reading into the local result is a real optimization: on the hot path it collapses two volatile reads into one.

The best way for a singleton: don't write DCL at all

If you want a static singleton, skip all this complexity and use the initialization-on-demand holder idiom. It leans on the JLS's own lazy class-initialization guarantees: the inner class isn't loaded until the first access to Holder.INSTANCE, and the JVM itself guarantees this init is safe and lazy — with no volatile or synchronized at all.

class Singleton {
    private Singleton() {}
    private static class Holder { static final Singleton INSTANCE = new Singleton(); }
    static Singleton getInstance() { return Holder.INSTANCE; } // JVM guarantees safe, lazy init
}

Common pitfalls & gotchas

Keep these as a checklist when reviewing concurrent code:

  • Relying on time instead of HB. "There's a Thread.sleep, surely the other thread finished by now." sleep creates no HB edge; stale reads remain perfectly legal.
  • volatile on a mutable object reference. It publishes the reference safely, but subsequent mutations of that object's fields are not covered. Publish an immutable snapshot instead.
  • Non-final fields in "immutable" objects. Only final fields get the constructor freeze guarantee; a non-final field can be seen with its default value if the object is published via a data race.
  • Locking on this in a library. Callers can lock on your object too, causing surprise contention or deadlock. Use a private lock.
  • check-then-act on concurrent collections. if (!map.containsKey(k)) map.put(k, v); races; use putIfAbsent/computeIfAbsent instead.
  • Forgetting unlock() isn't automatic. Explicit Locks need try/finally; a synchronized block cannot leak.
  • Assuming size()/isEmpty() on concurrent collections are exact. They're weakly consistent snapshots, not precise instantaneous counts.

Best practices

  • Prefer immutability and confinement (keeping data within one thread) over locking; the cheapest lock is the one you never take.
  • Reach for higher-level tools before hand-rolling locks: java.util.concurrent collections, CompletableFuture, ExecutorService.
  • Keep critical sections short and free of I/O. Never call foreign/callback code while holding a lock — that's a direct invitation to deadlock.
  • Establish and document a global lock ordering and acquire locks in that consistent order everywhere; this prevents deadlock.
  • For counters/accumulators under contention, reach for LongAdder, not AtomicLong.
  • Deliberately publish every field that crosses threads — via final, volatile, a lock, or a concurrent collection. If you can't name the HB edge, it's a bug.

Interview Questions

Q1: What guarantees does volatile provide, and what does it deliberately not?

Visibility and ordering via a release (write) / acquire (read) edge: all writes before a volatile write are visible after a subsequent read of that field — plus atomicity of the field itself, including 64-bit long/double. It does not provide mutual exclusion, and it does not make compound operations like x++ atomic.

Q2: Explain happens-before without saying "before in time"

It's a partial order over actions. If A HB B, the memory effects of A are guaranteed visible to and ordered before B. It's established between a release on a sync variable and a subsequent acquire on the same variable, plus program order, thread start/join, and transitivity. Two actions unordered by HB with a conflicting write form a data race.

Q3 (gotcha): A flag loop without volatile
boolean running = true;      // not volatile
void stop() { running = false; }
void run() { while (running) { /* work */ } }

What can happen? The JIT may hoist running out of the loop into a register (loop hoisting), so run() can loop forever even after stop() returns — there's no HB edge forcing the write to be observed. Making running volatile fixes it.

Q4 (gotcha): What does this print?
static int a = 0, b = 0;
// Thread 1: a = 1; int r1 = b;
// Thread 2: b = 1; int r2 = a;

Can r1 == 0 && r2 == 0? Yes. With no synchronization, the loads and stores can be reordered (store buffering), so both threads can read the other's pre-write value. Sequential consistency is not guaranteed for racy code.

Q5: Why is double-checked locking broken without volatile, exactly?

helper = new Helper() is allocate + construct + assign, and the JMM lets the reference assignment be reordered before (seen ahead of) the constructor's field writes. A racing reader on the lock-free path can see a non-null reference to a partially initialized object. volatile inserts the release/acquire barrier that makes the constructor's writes HB the reference read.

Q6: Why was biased locking removed, and what was the reasoning?

It optimized the "same thread repeatedly re-locks" case by biasing an object to that thread and skipping CAS entirely. But its revocation and bookkeeping costs grew expensive with modern many-thread workloads and concurrent data structures. JEP 374 disabled it by default in JDK 15; it was later removed (JDK 18). Lightweight and inflated locking remain.

Q7: synchronized vs ReentrantLock — when each, and did virtual threads change the answer?

ReentrantLock adds tryLock, timeouts, interruptibility, fairness, and multiple Conditions; synchronized is simpler and can't leak the lock. Historically you preferred ReentrantLock in virtual-thread code because synchronized pinned the carrier — but JEP 491 (JDK 24) removed that pinning, so now choose on features and ergonomics, not scalability.

Q8: Describe AQS in one breath

A volatile int state plus a CLH-based FIFO wait queue. Subclasses define what state means and implement tryAcquire/tryRelease (or the shared variants); AQS does the CAS, enqueue, and park/unpark. ReentrantLock (state = hold count), Semaphore (permits), and CountDownLatch (count, shared release) are all built on it.

Q9 (hard): What is the ABA problem and when does it actually bite?

CAS compares values, not history. If a value goes A→B→A, a CAS expecting A succeeds even though intermediate state changed. Harmless for a numeric counter; dangerous for pointer/node reuse in lock-free structures where a reused address passes the CAS check but references stale linkage. Fix with AtomicStampedReference (value + version).

Q10: What is false sharing and how do you detect/fix it?

Two independent variables sitting on the same 64-byte cache line cause cross-core invalidation on every write, ping-ponging the line between cores. Detect via perf counters (cache-line contention / HITM events) or by profiling; fix by padding hot fields onto separate lines, e.g. @Contended (as LongAdder's Cell does).

Q11: When does StampedLock beat ReadWriteLock, and what's the trap?

When reads dominate and you can use optimistic reads to avoid touching the lock's cache line at all on the read path. Traps: it's not reentrant, has no Condition, and optimistic reads must copy fields to locals and then validate() before using them — reading through a possibly-stale reference mid-read is a bug.

Q12: Is AtomicLong always the right counter?

No. Under heavy contention its CAS-retry loop thrashes and burns CPU. LongAdder/LongAccumulator stripe increments across padded cells and sum lazily, scaling far better for write-hot, read-rare counters — at the cost of a slightly more expensive sum() and no atomic read-modify-write across the whole value.

Q13 (gotcha): Is `long x; x = someLong;` atomic across threads?

Not guaranteed for a plain non-volatile long/double — the JLS permits a 64-bit write to be split into two 32-bit stores, so a racing reader can see a torn value (half old, half new). volatile, AtomicLong, or a lock removes the tearing.

Q14 (hard): Publishing an immutable object via a final field vs a plain field into a data race — the difference?

final fields get a special freeze at constructor end: any thread that reads the object reference is guaranteed to see the correctly-initialized final fields, even under racy publication. A plain (non-final) field carries no such guarantee — another thread can observe its default (0/null) value. This is why "immutable = all fields final" is a correctness rule, not style.

Q15: Why can lock downgrading work but upgrading deadlock in ReentrantReadWriteLock?

Downgrading (acquire the read lock while holding write, then release write) is safe because you never need to wait for other readers/writers to reach a stronger state. Upgrading (holding read, requesting write) requires all other readers to release — but if two readers both try to upgrade, each waits for the other to drop its read lock: deadlock. So the API forbids it.

Senior notes & advanced edge cases

You now hold the correct mental model and the tools. This section covers the layer that separates a genuine senior from a good programmer: the deeper JMM guarantees, safe publication, the modern access-mode ladder with VarHandle, the real semantics of wait/notify, the taxonomy of progress guarantees, and bugs that only surface under production load and a profiler.

Roadmap for this section
  1. SC-DRF and out-of-thin-air values. 2) coherence vs consistency. 3) Safe publication and this-escape. 4) VarHandle and the plain/opaque/acquire-release/volatile ladder plus fences. 5) wait/notify/Condition and spurious wakeups. 6) The wait-free / lock-free / obstruction-free taxonomy. 7) deadlock/livelock/starvation and JIT lock optimizations. 8) invisible gotchas and testing tools. Then 9 hard interview questions.

1) The SC-DRF theorem: what the JMM actually promises

The chapter showed racy code can see stale or torn values. The reassuring part: the JMM has one central theorem, SC-DRF (Sequential Consistency for Data-Race-Free programs).

Memorize the JMM's real promise

If your program is correctly synchronized — no pair of conflicting accesses (at least one a write) unordered by happens-before, i.e. data-race-free — then the JMM guarantees it behaves exactly as if sequentially consistent: a single global order of operations all threads agree on. All those frightening reorderings become invisible as long as you stay DRF.

This simplifies your real job: never reason about individual reorderings, just prove you're DRF — which is exactly what makes the chapter's golden line land: "if you can't name the HB edge, you have a bug." The flip side: for racy code, the JMM still has an unsolved theoretical problem.

Out-of-thin-air (OoTA) values

The formal JMM has causality rules meant to forbid "out-of-thin-air" values — a value no write produced, conjured from a self-justifying speculation. But the current formulation provably both forbids some desirable executions and fails to fully close OoTA; this is the model's well-known "hole" and the reason for ongoing work on a new JMM. Practical takeaway: never rely on racy code — even values that look "impossible" are permitted by the model.

2) Coherence vs consistency — a subtle correction

The "stale copy on the chef's counter" analogy is great for intuition, but a senior must be more precise. Modern hardware guarantees cache coherence via protocols like MESI: for a single address all cores agree on one global write order, and a write invalidates other cores' copies. So caches don't actually stay stale "forever."

The real source of staleness: store buffers and registers, not incoherent caches

So why do you still see stale values? Two reasons: (1) the store buffer — a core's write first sits in a local buffer and reaches the coherent cache with a delay; in that window your core sees the new value but others don't (the "store buffering" behind the classic r1==0 && r2==0 puzzle). (2) The compiler/JIT caching a value in a register (hoisting). So the problem isn't coherence, it's consistency. volatile and locks emit fences that drain the store buffer and constrain reordering — they don't "refresh the cache."

3) Safe publication and this-escape

The chapter stated "deliberate publication" as the golden rule. More precisely, there are four canonical ways to safely publish an object:

  1. Initialize it from a static initializer (class-init guarantee).
  2. Store its reference in a volatile field (or an AtomicReference).
  3. Store its reference in a final field assigned in a constructor.
  4. Store its reference guarded by a lock (or place it in a concurrent collection).
`this`-escape in a constructor: a bug that happens before the constructor finishes

The final-field freeze guarantee holds only if the reference does not escape before the constructor completes. If you let this leak from inside the constructor, it breaks:

public class Listener {
    private final int id;
    public Listener(EventBus bus) {
        bus.register(this);    // ❌ this escaped — constructor isn't done yet
        this.id = computeId(); // another thread may observe id == 0
    }
}

Common this-escape patterns: registering a listener/callback, starting a thread inside a constructor, or publishing this into a static variable. The fix: finish the constructor, then in a static create(...) factory build the object and only then register it.

4) VarHandle and the access-mode ladder — the modern memory model

The chapter demonstrated CAS with AtomicInteger. But since Java 9 (JEP 193) the low-level, official machinery underneath it all is VarHandle — the safe, supported successor to sun.misc.Unsafe and the Atomic*FieldUpdater classes (the java.util.concurrent classes themselves migrated onto it in JDK 9).

Its key gift: instead of the "plain or volatile" binary, it offers a four-rung ladder of ordering strength, weakest to strongest.

The VarHandle access-mode strength ladder (weak to strong):

flowchart LR
  Plain["Plain<br/>get/set"] --> Opaque["Opaque<br/>getOpaque/setOpaque"]
  Opaque --> RelAcq["Release/Acquire<br/>setRelease/getAcquire"]
  RelAcq --> Volatile["Volatile (SC)<br/>getVolatile/setVolatile"]
  • Plain: no cross-thread ordering or visibility — just atomicity of the access itself (except plain long/double). Like a normal field read.
  • Opaque: accesses aren't elided and stay coherent for a single address, but are reorderable relative to other addresses — for a stop flag or a stats counter that only needs to be seen "eventually," cheaper than volatile.
  • Release/Acquire: the same release/acquire half-fence from happens-before, but without the heavy StoreLoad fence — the "sign / see the signature" edge without full SC cost.
  • Volatile: the strongest — equivalent to the volatile keyword, with a global SC order.
class Node {
    Object item;
    volatile Node next;      // ordinary field, manipulated via VarHandle
    private static final VarHandle NEXT;
    static {
        try {
            NEXT = MethodHandles.lookup()
                .findVarHandle(Node.class, "next", Node.class);
        } catch (ReflectiveOperationException e) { throw new ExceptionInInitializerError(e); }
    }
    boolean casNext(Node expect, Node update) {
        return NEXT.compareAndSet(this, expect, update);
    }
    void publish(Node n) { NEXT.setRelease(this, n); } // release, cheaper than setVolatile
}
compareAndExchange and weak CAS

Two senior points: (1) compareAndExchange is like compareAndSet but returns the witness value instead of a boolean — in a retry loop it eliminates a redundant get(). (2) weakCompareAndSet may fail spuriously even when the value equals expected; in exchange it's cheaper on LL/SC architectures like ARM. That's why it's only correct inside a do/while loop you write yourself, never standalone.

Standalone fences: when you have no volatile variable at all

VarHandle also has four static fence methods (fullFence, acquireFence, releaseFence, loadLoadFence/storeStoreFence) that emit a memory barrier without being tied to a specific field — the modern, safe equivalent of Unsafe.fullFence for very specific hand-rolled publication patterns.

5) wait / notify / Condition — correct semantics

The chapter covered locks and AQS but never touched thread coordination via wait/notify — one of the buggiest spots in real code.

A doctor's waiting room

wait() means "hand back the key and sleep in the waiting room," notify() means "call one of the sleepers." The crucial detail: the awoken thread does not run immediately — it must re-contend for the lock, so between "waking" and "running" the world may have changed.

Three iron rules:

  1. Call wait/notify/notifyAll only while you hold that object's monitor, or you get an IllegalMonitorStateException.
  2. Always put wait inside a loop that re-checks the condition, not an if. Two reasons: spurious wakeups, which the JLS explicitly permits, and that by the time it's your turn the condition may be false again.
synchronized (queue) {
    while (queue.isEmpty()) {   // ✅ while, not if
        queue.wait();
    }
    return queue.remove();
}
notify vs notifyAll and the "lost signal"

notify() wakes only one thread and doesn't guarantee which. If several threads wait on different conditions on the same monitor, it may wake the "wrong" one whose condition still isn't satisfied; it sleeps again and the signal is lost — the eligible threads never wake (a soft deadlock). Rule: unless all waiters are provably equivalent, choose notifyAll. With Condition (from ReentrantLock) you build separate wait queues (notFull/notEmpty) and needn't wake everyone — safer and more efficient.

park/unpark vs wait/notify

Underneath AQS, the parking mechanism is LockSupport.park()/unpark(), not wait/notify. The subtle difference: unpark grants a permit that doesn't accumulate (at most one) but can arrive before park — so it avoids wait/notify's "lost signal." park can also return spuriously, so it too sits inside a condition loop.

6) Progress-guarantee taxonomy: wait-free / lock-free / obstruction-free

The chapter called CAS "lock-free." A senior must know this term precisely because interviewers set traps around it.

Three levels of progress guarantee
  • Obstruction-free: a thread running alone finishes in bounded steps. (Weakest.)
  • Lock-free: at any moment at least one thread makes progress — the system never stalls, though an unlucky thread may retry forever (starvation).
  • Wait-free: every thread finishes in bounded steps — no starvation. (Strongest.)
A CAS loop is lock-free, not wait-free

The do { ... } while(!cas()) loop from the chapter is lock-free, not wait-free: someone always wins, but a particular thread can lose repeatedly and starve. That's why LongAdder is better under heavy contention. And contrary to popular belief, "lock-free" doesn't mean "faster" — under very heavy load a fair lock may give better throughput. Thread.onSpinWait() (Java 9, JEP 285) helps in spin loops: it gives the CPU a PAUSE hint to cut power draw and coherence traffic.

7) Deadlock vs livelock vs starvation, and the JIT's lock optimizations

The chapter mentioned "global lock ordering" to prevent deadlock. Distinguish the three concurrency failures:

  • Deadlock: threads wait for each other forever (a wait cycle). The four Coffman conditions must all hold (mutual exclusion, hold-and-wait, no-preemption, circular-wait); breaking one suffices — usually circular-wait, via a global lock order.
  • Livelock: threads aren't blocked and keep working but make no progress — like a tryLock loop that releases and immediately retries on every failure; the fix is randomized backoff.
  • Starvation: some threads progress but one never gets a turn (writer starvation in ReadWriteLock).
JIT lock optimizations you should know

Three HotSpot optimizations the chapter didn't mention: (1) Lock elision — if escape analysis proves an object never escapes one thread, the JIT removes its locking entirely (e.g. a local StringBuffer). (2) Lock coarsening — it merges back-to-back synchronized regions on the same object into one to avoid repeated lock/unlock cost. (3) Adaptive spinning — before parking it spins briefly hoping the lock frees soon, tuning the spin from that lock's history. So a naive micro-benchmark rarely reveals a lock's true cost.

8) Invisible gotchas that bite in production

`volatile` on an array covers only the reference, not the elements
volatile int[] data;   // only the array reference itself is volatile
data[5] = 42;          // ❌ this write has no volatile semantics!

A catastrophic mistake: writing to data[i] is a plain write and creates no HB edge. If the elements need volatile ordering/visibility, use AtomicIntegerArray or a VarHandle over the array (MethodHandles.arrayElementVarHandle).

ThreadLocal leaks in a thread pool

ThreadLocal binds a value to a thread; in a pool, threads are reused and never die. If you don't remove() after each task: (1) one request's data leaks into the next request on the same thread (a security/correctness bug), and (2) a large object stays alive forever (a memory leak). Always remove() in a finally or exit filter.

Concurrent collections establish HB edges for you

A documented but under-known principle: the java.util.concurrent classes give "memory consistency" guarantees — e.g. placing an object into a BlockingQueue happens-before its retrieval via take(), and a write into ConcurrentHashMap happens-before a later read of the same key. So passing a mutable object through a concurrent queue needs no separate volatile — the queue itself performs safe publication.

9) How to actually test concurrency — jcstress and JMH

Writing concurrent code without proper testing is self-deception — a bug may never appear on your x86 laptop yet blow up on an ARM server or under load. jcstress (official OpenJDK) builds litmus tests and runs them millions of times to hunt rare reorderings — unmatched for proving a pattern is DRF. And JMH for correct benchmarking that accounts for lock elision/coarsening and JIT warmup (a hand-rolled System.nanoTime() almost always lies).

Q1 (hard): What exactly does the JMM guarantee for a correctly-synchronized program? Explain SC-DRF.

The SC-DRF theorem: if a program is data-race-free (every conflicting-access pair ordered by happens-before), the JMM guarantees the execution appears sequentially consistent — a single global order all threads agree on, reorderings invisible. So instead of reasoning about individual reorderings, just prove you're DRF. Flip side: for racy code the model doesn't fully close "out-of-thin-air" values and has a known theoretical hole — never rely on racy behavior.

Q2 (hard): Does seeing a stale value mean the cache is incoherent?

No. Hardware guarantees cache coherence (via MESI): for a single address there's a global write order and a write invalidates other copies. Staleness comes from the store buffer (a not-yet-flushed local write) and the JIT caching a value in a register. So the problem isn't coherence, it's memory consistency (ordering across addresses). volatile/locks emit fences that drain the store buffer and constrain reordering; they don't "refresh the cache."

Q3 (hard): Explain the VarHandle access-mode ladder; when opaque, when release/acquire, when volatile?

Four rungs weak to strong: plain (atomic only, no ordering), opaque (coherence and progress for one variable, reorderable relative to others — cheaper than volatile for a stop flag or stats counter), release/acquire (half-fence; the same sign/see-the-signature HB without the heavy StoreLoad), and volatile (strongest, global SC). Rule: pick the weakest mode that keeps your pattern correct; in lock-free structures release-for-publish plus acquire-for-read usually suffices and shaves the full volatile cost.

Q4 (gotcha): Why must `wait()` always be in a loop, not an `if`?

Two reasons: (1) spurious wakeups, which the JLS explicitly permits — a thread can wake with no notify at all. (2) Between waking and reacquiring the lock, another thread may have invalidated the condition again. So re-check after waking; only while (!condition) wait(); is safe. The same applies to Condition.await() and LockSupport.park().

Q5 (hard): What's the difference between `notify` and `notifyAll`, and the "lost signal"?

notify() wakes one unspecified thread. If waiters block on different conditions, it may wake one whose condition doesn't hold; it sleeps again and the signal is lost — so the thread that should have woken never does (a soft deadlock / lost-wakeup). Default to notifyAll unless all waiters are provably equivalent; better still, use separate Conditions (notFull/notEmpty) and signal only the relevant thread.

Q6 (hard): Define wait-free, lock-free, and obstruction-free; which is a CAS loop?

obstruction-free: a thread with no interference finishes in bounded steps. lock-free: at least one thread always makes progress (the system never stalls) but a particular thread can starve. wait-free: every thread finishes in bounded steps (no starvation). A do{}while(!cas()) loop is lock-free, not wait-free — someone always wins but an unlucky thread can lose repeatedly. Hence LongAdder under heavy contention; and "lock-free" doesn't necessarily mean "faster."

Q7 (hard): What is `this`-escape in a constructor and why is it dangerous even without an obvious thread?

If the reference leaks before the constructor completes (registering a listener, starting a thread, publishing into a static field), the final-field freeze guarantee breaks: another thread can observe uninitialized fields (0/null) because the constructor's writes aren't done. Even single-threaded it's dangerous because callback code may operate on a half-built object. The fix: finish the constructor and move publication into a static factory after construction.

Q8 (gotcha): What are the semantics of `volatile int[] a; a[i] = x;`?

Only the array reference is volatile, not its elements. Writing a[i] is a plain write with no happens-before edge — another thread may see a stale value or none. For volatile semantics on elements use AtomicIntegerArray or MethodHandles.arrayElementVarHandle. Same trap with AtomicReference<mutableObject>: the reference is safely published but later mutations of that object's fields aren't covered.

Q9 (hard): What's the difference between livelock and deadlock, and how do you fix them?

deadlock: threads blocked in a mutual wait cycle; all four Coffman conditions must hold and breaking one (usually circular-wait, via a global lock order) suffices. livelock: threads aren't blocked and keep working but make no progress — like a tryLock loop that releases and immediately retries on each failure, everyone in lockstep. Fix: randomized backoff. Both are diagnosable via a thread dump and cycle detection, but prevention beats cure.

In a nutshell (senior)
  • SC-DRF: if you're DRF the world looks sequentially consistent; for racy code even OoTA isn't closed. Staleness comes from the store buffer and registers, not incoherent caches (consistency, not coherence).
  • Safe publication has four canonical forms and breaks on this-escape. VarHandle gives the plain/opaque/release-acquire/volatile ladder — pick the weakest that's correct.
  • Always loop wait; default to notifyAll or separate Conditions. A CAS loop is lock-free, not wait-free; under load reach for LongAdder.
  • Separate deadlock/livelock/starvation; the JIT elides/coarsens locks. Gotchas: volatile array covers only the reference, ThreadLocal leaks in pools, concurrent collections establish HB for free. Test with jcstress/JMH.
In a nutshell
  • Memory at the multi-core level is not a shared notebook; the compiler, CPU, and caches reorder your code. The JMM (JLS Ch. 17, reworked by JSR-133 in Java 5) is the formal contract for this world.
  • happens-before is the only rule that matters: it's about ordering, not time, and holds only between a release and a subsequent acquire on the same synchronization variable.
  • volatile gives visibility/ordering and field atomicity (even 64-bit), but not mutual exclusion, and does not make x++ atomic.
  • synchronized gives exclusion + visibility + reentrancy and can't leak the key; biased locking is gone (JEP 374 in JDK 15, removed in JDK 18), and as of JDK 24 with JEP 491 it no longer pins virtual threads.
  • The Lock family adds flexibility: ReentrantLock (tryLock/timeout/fairness/Condition), ReadWriteLock (downgrade allowed, upgrade forbidden), StampedLock (optimistic reads, but not reentrant).
  • AQS is the shared engine under all of them: volatile int state + CLH queue + tryAcquire.
  • CAS enables lock-free programming but has the ABA problem (fix with AtomicStampedReference); under heavy contention reach for LongAdder.
  • Watch out for false sharing (pad with @Contended) and DCL without volatile; for singletons use the holder idiom.
  • The golden rule: deliberately publish every cross-thread field. If you can't name the HB edge, you have a bug.