Concurrency · همزمانی سنیورSenior ~59 دقیقه مطالعه~51 min read
همگامسازی، قفلها، AQS و مدل حافظهٔ جاواSynchronization, Locks, AQS & the Java Memory Model
یاد میگیری چرا حافظهٔ جاوا آنطور که فکر میکنی رفتار نمیکند و چطور با happens-before، volatile، synchronized، خانوادهٔ Lock، AQS و CAS دقیقاً همان تضمینهایی را که نیاز داری بخری—و بدانی هرکدام چه چیزی را تضمین نمیکند.You will learn why Java's memory really doesn't behave the way you assume, and how to buy exactly the guarantees you need with happens-before, volatile, synchronized, the Lock family, AQS, and CAS—while knowing precisely what each one does not promise.
پیشنیاز:Prerequisites: نخها، Runnable/Callable و ExecutorهاThreads, Runnable/Callable & Executors
بیا رک باشیم: بیشتر باگهای همروندی (concurrency) به این خاطر پیش میآیند که برنامهنویس فرض میکند حافظه مثل یک دفترچهٔ مشترک رفتار میکند که همه همزمان همان صفحه را میبینند. حافظهٔ واقعیِ یک ماشین چندهستهای اصلاً اینطور نیست. این فصل قرار است آن مدل ذهنیِ غلط را کامل بشکند و بهجایش یک مدل درست بسازد—و بعد ابزارهایی را که با آنها میتوانی ترتیب و رؤیتپذیری را کنترل کنی، یکییکی و از صفر بسازیم.
اول میبینیم چرا JMM اصلاً وجود دارد و سه عاملی که کدت را بازچینش میکنند. بعد به قلب ماجرا میرسیم: happens-before، تنها قاعدهای که تضمین میکند یک نخ نوشتهٔ نخ دیگر را ببیند. سپس ابزارها را از سبک به سنگین میسازیم: volatile، synchronized و بهینهسازیهای قفل، خانوادهٔ Lock (ReentrantLock / ReadWriteLock / StampedLock)، موتور زیرین AQS، برنامهنویسی بدونقفل با CAS و مسئلهٔ ABA، false sharing، و بالاخره الگوی معروف double-checked locking. در پایان یک بخش کامل پرسشهای مصاحبه و یک جمعبندی داریم.
بخش ۰ — واژههایی که باید از قبل بشناسی
قبل از هر چیز چند واژه را با تشبیه محکم کنیم تا بعداً سرد و ناگهانی رهایشان نکنم.
یک نخ (thread) را مثل یک آشپز فرض کن که مشغول اجرای دستورپخت (کد) توست. یک CPU چندهستهای یعنی چند آشپز که همزمان کار میکنند. حالا نکتهٔ مهم: هر آشپز یک میز کار کوچک کنار دستش (cache هستهٔ خودش) دارد و انبار اصلی (RAM) آنطرف آشپزخانه است. وقتی آشپز نمکدان را برمیدارد، یک کپی روی میز خودش میگذارد و مدتی همان کپی را استفاده میکند. پس اگر آشپز دیگری در انبار اصلی نمک را عوض کند، آشپز اول تا مدتی خبردار نمیشود—او هنوز کپیِ کهنهٔ روی میزش را میبیند. این «کپیِ کهنه» ریشهٔ اکثر باگهای رؤیتپذیری (visibility) در جاواست.
- بازچینش (reordering): جابهجاکردن ترتیب اجرای دستورها برای سرعت بیشتر—مثل آشپزی که بهجای ترتیب نوشتهشدهٔ دستور، کارها را به کارآمدترین ترتیب انجام میدهد.
- مانع/حصار حافظه (memory barrier / fence): یک دستور ویژه که به سختافزار میگوید «تا اینجا را واقعاً تمام کن و همه ببینند، بعد برو جلو»—مثل قانونی که میگوید «تا دیس را روی میز پاس اصلی نگذاشتی، غذای بعدی را شروع نکن.»
- اتمیک (atomic): عملیاتی که یا کامل انجام میشود یا اصلاً؛ هیچ نخ دیگری نمیتواند وسطش حالت نیمهکاره را ببیند.
مدل ذهنی: چرا JMM وجود دارد؟
بهطور ساده تصور میکنی برنامهات یک فهرست ترتیبیِ واحد از خواندنها و نوشتنهای حافظه است که دقیقاً به ترتیب سورس اجرا میشود و بیدرنگ برای همهٔ نخها دیده میشود. این مدل، هرچقدر هم راحت باشد، دروغ است. میان سورس تو و الکترونهای داخل تراشه، سه عامل بازچینش وجود دارد:
- کامپایلر (javac و JIT): میتواند دستورها را بازچینش، بالابری (hoist)، حذف یا ادغام کند. «بالابری» یعنی چیزی را که داخل حلقه هربار خوانده میشد، یک بار بیرون حلقه بخواند و در یک رجیستر نگه دارد—بعداً میبینی چرا این میتواند فاجعه بسازد.
- CPU: دستورها را خارج از ترتیب (out-of-order) و گاهی بهصورت گمانهزنانه (speculative، یعنی حدس میزند کدام شاخه اجرا میشود و از قبل شروعش میکند) اجرا میکند.
- سلسلهمراتب حافظه: بافر نوشتن (store buffer) و کش هر هسته باعث میشوند نوشتهٔ یک هسته با تأخیر برای هستهٔ دیگر دیده شود—همان «کپیِ کهنه روی میز آشپز».
وسوسه میشوی فکر کنی این سه عامل «خرابکار»اند. برعکس: تمام کارایی مدرن از همین بازچینشها میآید. سکو فقط یک قول به تو میدهد: برای کد تکنخی معنای as-if-serial را تضمین میکند—یعنی هر بازچینشی مجاز است تا وقتی نتیجهای که خودت در همان نخ مشاهده میکنی عوض نشود. مشکل دقیقاً وقتی شروع میشود که نخ دیگری بخواهد وسط کار تو را تماشا کند.
اینجاست که مدل حافظهٔ جاوا (Java Memory Model / JMM) وارد میشود. این مدل در فصل ۱۷ از JLS تعریف و با JSR-133 (جاوا ۵) بازنویسی شد، و همان قراردادی است که دقیقاً میگوید کدام خواندنهای بیننخی مجازند کدام نوشتنها را ببینند. هرچه در بقیهٔ این فصل میآید—volatile، synchronized، Lock، اتمیکها—در واقع راهی است برای خریدن یالهای happens-before از این قرارداد.
Happens-before: تنها قاعدهای که واقعاً اهمیت دارد
اگر فقط یک چیز از این فصل با خودت ببری، همین باشد.
تصور کن دو دپارتمان در یک شرکتاند. دپارتمان A سندی را آماده میکند و دپارتمان B باید رویش کار کند. اگر هیچ رویهای بینشان نباشد، B ممکن است نسخهٔ نیمهکاره یا اصلاً نسخهٔ دیروز را بردارد. اما اگر قانون این باشد که «A سند را در صندوق مشترک میگذارد و امضا میکند، و B فقط بعد از دیدن آن امضا سند را برمیدارد»، آنگاه هر چیزی که A پیش از امضا نوشته، تضمیناً برای B که بعد از امضا میخواند دیده میشود. آن جفتِ «امضا» (release) و «دیدن امضا» (acquire) دقیقاً همان یال happens-before است.
حالا دقیق شویم. JMM بر پایهٔ یک ترتیب جزئی (partial order) به نام happens-before (بهاختصار HB) تعریف میشود. «ترتیب جزئی» یعنی لازم نیست هر دو کنش با هم مقایسهشدنی باشند؛ بعضی جفتها مرتباند و بعضی نه. قاعده این است: اگر کنش A نسبت به کنش B رابطهٔ happens-before داشته باشد، آنگاه تمام اثرهای حافظهٔ A برای B دیده میشوند و پیش از آن مرتب شدهاند.
و اما نکتهٔ خطرناک: اگر دو کنش با HB مرتب نشده باشند و دستکم یکی نوشتن روی همان مکان باشد، یک مسابقهٔ داده (data race) داری. در این حالت JMM اجازه میدهد خواندنها مقدارهای کهنه (stale)، پارهشده (torn، یعنی نیمی از مقدار قدیم و نیمی جدید) یا حتی بهظاهر «ناممکن» برگردانند.
یالهای HB که واقعاً به دست میآوری، اینهاست—این فهرست را حفظ کن:
- ترتیب برنامه (program order): درون یک نخ، هر کنش نسبت به هر کنش بعدیِ همان نخ HB دارد.
- قفل مانیتور (monitor): یک
unlockروی یک مانیتور نسبت به هرlockبعدی روی همان مانیتور HB دارد. - volatile: نوشتن روی یک فیلد
volatileنسبت به هر خواندن بعدی همان فیلد HB دارد. - شروع نخ:
Thread.start()نسبت به هر کنش در نخ آغازشده HB دارد. - پیوستن نخ (join): هر کنش در یک نخ نسبت به بازگشت موفق نخ دیگر از
join()روی آن HB دارد. - وقفه (interrupt): فراخوانی
interrupt()نسبت به تشخیص آن توسط نخ وقفهخورده HB دارد. - فیلدهای final: پایان سازنده (constructor) نسبت به «انجماد» (freeze) فیلدهای
finalHB دارد؛ این پایهٔ انتشار امن (safe publication) اشیای تغییرناپذیر است. - ترایایی (transitivity): اگر A HB B و B HB C، آنگاه A HB C. یعنی یالها زنجیر میشوند.
مهمترین سوءتفاهم را همینجا خاک کنیم. دو رویداد میتوانند از نظر ساعت دیواری کاملاً «همزمان» باشند؛ آنچه اهمیت دارد این نیست که کدام «زودتر» رخ داد، بلکه این است که آیا مدل وادار میکند نوشتههای یکی توسط دیگری دیده شود. و این وادارسازی فقط میان یک کنش release و یک کنش acquire بعدی روی همان متغیر همگامسازی برقرار میشود. آزادکردن قفل A هیچ چیزی دربارهٔ نخی که قفل B را میگیرد به تو نمیگوید—امضای صندوق A به درد صندوق B نمیخورد.
بیا با یک مثال ملموس این زنجیر را ببینیم:
نخ ۱ نخ ۲
----- -----
data = 42; (1)
ready = true; (2، نوشتن volatile / release)
while(!ready) {} (3، خواندن volatile / acquire)
print(data); (4) -> تضمیناً 42 را میبیند
چرا کار میکند؟ چون (2) نوشتن volatile و (3) خواندن volatile همان فیلد است، پس (2) HB (3). طبق ترتیب برنامه (1) HB (2) و همچنین (3) HB (4). حالا ترایایی را زنجیر کن: (1) HB (2) HB (3) HB (4)، پس (1) HB (4). یعنی خواندن data در خط (4) حتماً ۴۲ را میبیند.
اگر volatile را از ready حذف کنی، دو چیز همزمان میشکند: هم رؤیتپذیری data (ممکن است نخ ۲ مقدار کهنه ببیند) و هم پایان حلقه. JIT حق دارد ready را از داخل حلقه بالا ببرد و در یک رجیستر نگه دارد—آنوقت while(!ready) برای همیشه همان مقدار رجیستری را میبیند و تا ابد حلقه میزند، حتی بعد از اینکه نخ ۱ مقدار را عوض کرد. این یکی از رایجترین باگهای واقعی است.
volatile: چه میکند و چه نمیکند
حالا که happens-before را داری، volatile ساده میشود. آن را مثل یک نمکدانِ شیشهای فرض کن که همه مجبورند مستقیم از انبار اصلی بردارند، نه از کپیِ روی میز—و هر بار که کسی چیزی در آن میگذارد، مجبور است اول همهٔ کارهای قبلیاش را روی میز پاس اصلی مرتب کند.
volatile دقیقاً دو چیز به تو میدهد:
- رؤیتپذیری + ترتیب (release/acquire): نوشتن volatile نهفقط همان متغیر، بلکه یک یال HB میسازد و تمام نوشتنهایی را که در ترتیب برنامه پیش از آن بودند، برای نخی که بعداً همان volatile را میخواند دیدنی میکند. همین «سوارشدن» (piggybacking) بود که مثال
data/readyبالا را کار انداخت. - اتمیکبودن خودِ فیلد، حتی برای
long/double. توجه: طبق JLS، نوشتن ۶۴بیتیِ غیرvolatile مجاز است به دو ذخیرهٔ ۳۲بیتی شکسته شود؛ volatileاین پارهشدن را ممنوع میکند.
و اما آنچه volatile نمیدهد—اینجا اکثر آدمها اشتباه میکنند:
- بدون انحصار متقابل (mutual exclusion).
volatile int x; x++;در واقع یک read-modify-write است (بخوان، جمع کن، بنویس) و اتمیک نیست. دو نخ میتوانند هردو ۵ را بخوانند، هردو ۶ حساب کنند و هردو ۶ بنویسند—و یک افزایش گم میشود. برای این کارAtomicIntegerیا قفل لازم داری. - بدون اتمیکبودن کنش مرکب.
if (v == null) v = new X();حتی اگرvvolatile باشد باز هم مسابقه دارد، چون بین بررسی و انتساب فاصله هست.
روی معماری x86 نوشتن volatile به یک ذخیرهٔ ساده و سپس یک مانع store-load (اغلب دستوری با پیشوند lock یا mfence) کامپایل میشود؛ اما خواندنها عملاً رایگاناند، چون x86 مدل حافظهٔ قویای به نام TSO (Total Store Order) دارد. روی مدلهای حافظهٔ ضعیفتر مثل ARM و Power، هم خواندن و هم نوشتن باید دستورهای مانع (fence) واقعی تولید کنند. یعنی همان کد جاوا روی موبایل ARM ممکن است گرانتر از سرور x86 تمام شود.
synchronized: مانیتور ذاتی
اگر volatile نمکدان مشترک بود، synchronized کلید یک اتاق است که هر بار فقط یک نفر داخلش میرود.
هر شیء جاوا یک قفل نامرئی به نام مانیتور (monitor) یا قفل ذاتی (intrinsic lock) همراه دارد—درست مثل یک دستشویی که کلیدش روی در آویزان است. برای واردشدن باید کلید را برداری (lock)، و موقع خروج آویزانش کنی (unlock). تا وقتی کلید دست توست، هیچکس دیگر نمیتواند وارد شود. بلوک synchronized این «برداشتن و آویزانکردن کلید» را بهطور خودکار انجام میدهد—حتی اگر وسط کار استثنا پرتاب شود، خروجش مثل یک finally عمل میکند و کلید را برمیگرداند، پس هرگز نمیتوانی کلید را «نشت» بدهی.
تضمینهای synchronized:
- انحصار متقابل: تنها یک نخ در هر لحظه یک مانیتور معین را نگه میدارد.
- رؤیتپذیری: unlock نسبت به lock بعدی روی همان مانیتور HB دارد، پس تمام نوشتههای داخل ناحیهٔ بحرانی (critical section) منتشر میشوند.
- بازورودی (reentrancy): همان نخ میتواند مانیتوری را که پیشاپیش نگه داشته دوباره بگیرد (یک شمارندهٔ نگهداشت بهازای هر نخ نگهداشته میشود). همین است که اجازه میدهد یک متد
synchronized، متدsynchronizedدیگری رویthisرا صدا بزند بدون اینکه خودش را در بنبست بیندازد.
نکتهٔ ظریف دربارهٔ اینکه روی چه چیزی قفل میکنی: synchronized(this) و یک متد نمونهٔ synchronized، هردو روی this قفل میکنند؛ یک متد استاتیک synchronized روی شیء Class قفل میکند.
synchronized("lock") یا قفل روی یک Integer که از باکسینگ آمده، یک تلهٔ کلاسیک است. لیترالهای String در استخر رشته (intern) میشوند و Integerهای کوچک کش میشوند، پس همان شیء دقیق ممکن است در کدهای کاملاً بیربطِ دیگر هم استفاده شود—و ناگهان دو بخش نامرتبط برنامه روی یک قفل مشترک، بیآنکه بدانند، بنبست بسازند. همیشه یک قفل اختصاصی و خصوصی بساز: private final Object lock = new Object();.
بهینهسازیهای قفل — و آنچه حذف شد
از نظر تاریخی HotSpot سه شیوهٔ قفل را لایهبندی میکرد که خوب است بشناسی، بهویژه چون یکیشان حذف شده و سؤال مصاحبه است:
- قفل مغرضانه (biased locking): فرض میکرد یک قفل معمولاً توسط همان نخ دوباره گرفته میشود، پس شیء را به آن نخ «مغرض» میکرد و در بازورود از CAS اتمیک بهکلی میگذشت. این ویژگی بهطور پیشفرض غیرفعال و منسوخ (deprecated) شد در JDK 15 (JEP 374) و پیادهسازیاش بعداً کهنه/حذف شد (JDK 18، JDK-8256425). دلیل؟ هزینهٔ دفترداری و ابطالش، بارهای کاری مدرن با نخهای کوتاهعمرِ فراوان و ساختارهای دادهٔ همروندِ پرمناقشه را کند میکرد.
- قفل سبک (lightweight / thin): قفلهای بدون مناقشه با یک CAS روی «mark word» سرآیند شیء، یک رکورد قفل روی پشتهٔ نخ تخصیص میدهند—بدون میوتکس سطح سیستمعامل، یعنی خیلی ارزان.
- قفل سنگین (heavyweight / inflated): وقتی مناقشه بالا میگیرد، مانیتور به یک
ObjectMonitorسطح سیستمعامل با یک صف انتظار واقعی و پارککردن OS «باد میکند» (inflate). این گران است ولی زیر فشار لازم.
یک نکتهٔ در سطح ارشد که حتماً باید بدانی: از JDK 24، JEP 491 باعث میشود synchronized دیگر نخهای مجازی (virtual threads) را سنجاق (pin) نکند. پیش از ۲۴، وقتی یک نخ مجازی داخل بلوک synchronized مسدود میشد، نخ سکوی حاملش (carrier platform thread) را گروگان میگرفت—به این «pinning» میگویند—و مقیاسپذیری را خفه میکرد. برای همین توصیه میشد در کدی که پر از نخ مجازی است ReentrantLock را ترجیح بدهی. از JDK 24 آن توصیه منسوخ است: حالا synchronized در برابر java.util.concurrent.locks را صرفاً بر پایهٔ راحتی و امکانات انتخاب کن، نه مقیاسپذیری.
خانوادهٔ Lock: ReentrantLock، ReadWriteLock، StampedLock
مانیتور ذاتی ساده و امن است، ولی خشک است: نمیتوانی بگویی «اگر قفل تا ۲ ثانیه آزاد نشد بیخیال شو». بستهٔ java.util.concurrent.locks قفلهای صریحی میدهد که این انعطاف را دارند.
ReentrantLock — اسب بارکش
ReentrantLock همان انحصار متقابلِ synchronized را میدهد، اما با قابلیتهای اضافه: tryLock() (با مهلت اختیاری—«اگر تا فلان زمان نشد، برگرد»)، گرفتنِ قابلوقفه با lockInterruptibly()، انصاف (fairness) اختیاری (ترتیب FIFO برای نخهای منتظر، به بهای کمی افت توان عملیاتی)، و چند شیء Condition بهازای هر قفل (در برابر تنها یک مجموعهانتظار برای هر مانیتور).
اما یک اصطلاح غیرقابلمذاکره دارد—این را در خواب هم باید بلد باشی:
private final ReentrantLock lock = new ReentrantLock();
void doWork() {
lock.lock();
try {
// ناحیهٔ بحرانی
} finally {
lock.unlock(); // باید در finally باشد — استثنا نباید قفل را نشت دهد
}
}
با synchronized کلید همیشه خودبهخود برمیگردد. با Lock صریح، اگر بین lock() و unlock() استثنایی پرتاب شود و unlock() را در finally نگذاشته باشی، قفل برای همیشه نگهداشته میماند و هر نخ دیگری که آن را بخواهد تا ابد منتظر میماند. همیشه unlock() در finally.
ReentrantReadWriteLock — قفل خواندن/نوشتن جدا
تصور کن یک تختهسفید داری. خواندن یعنی نگاهکردن به تخته—ده نفر میتوانند همزمان نگاه کنند، هیچ مشکلی نیست. نوشتن یعنی پاککردن و نوشتن دوباره—موقع نوشتن هیچکس دیگری نباید نه بنویسد نه بخواند، وگرنه چیز نیمهپاکشده میبیند. ReentrantReadWriteLock دقیقاً همین است: قفل خواندن را خیلیها همزمان میگیرند، ولی قفل نوشتن انحصاری است.
این وقتی خوب است که خواندنها بهشدت بر نوشتنها غلبه دارند و ناحیههای بحرانی بیاهمیت (خیلی کوتاه) نیستند. اما دو تلهٔ مهم:
- گرسنگی نویسنده (writer starvation): اگر سیل خوانندگان بیوقفه بیاید، نویسنده ممکن است هرگز نوبت نگیرد. با سازندهٔ
fairکاهشش بده. - تنزل مجاز، ارتقا ممنوع: تنزل (downgrade) یعنی از قفل نوشتن به خواندن بروی—این مجاز است، بهشرطی که قفل خواندن را پیش از آزادکردن قفل نوشتن بگیری. اما ارتقا (upgrade) یعنی از خواندن به نوشتن—این بنبست میسازد و ممنوع است. (چراییاش را در پرسش مصاحبهٔ آخر کامل باز میکنیم.)
StampedLock (جاوا ۸) — خواندن خوشبینانه
StampedLock بازورودی نیست، اما یک ویژگی مرگبار میافزاید: خواندن خوشبینانه (optimistic read).
فرض کن میخواهی ساعت را ببینی. حالت بدبینانه این است که بروی جلوی ساعت بایستی و مانع شوی کسی عقربهها را تکان دهد (قفل خواندن). حالت خوشبینانه این است که فقط یک نگاه سریع بیندازی، عدد را حفظ کنی، و بعد چک کنی «آیا در این فاصله کسی به ساعت دست زد؟» اگر نه، عددت معتبر است و اصلاً مزاحم کسی نشدی. اگر بله، آنوقت بهناچار میروی جلوی ساعت میایستی. این نگاهسریع همان tryOptimisticRead و آن چک همان validate است.
private final StampedLock sl = new StampedLock();
private double x, y;
double distanceFromOrigin() {
long stamp = sl.tryOptimisticRead(); // بدون CAS، بدون مسدودشدن
double cx = x, cy = y; // خواندن عکسفوری (snapshot)
if (!sl.validate(stamp)) { // ممکن است نویسندهای اجرا شده باشد
stamp = sl.readLock(); // بازگشت به قفل خواندن بدبینانه
try { cx = x; cy = y; }
finally { sl.unlockRead(stamp); }
}
return Math.sqrt(cx * cx + cy * cy);
}
سود بزرگ: وقتی نوشتنها نادرند، مسیر خواندن اصلاً به خط کش قفل دست نمیزند، پس مناقشهٔ cache line بهکلی حذف میشود.
سه چیز را یادت باشد: (۱) بازورودی نیست—اگر همان نخ دوباره قفل کند، خودش را در بنبست میاندازد. (۲) از Condition پشتیبانی نمیکند. (۳) مستقیماً قابلوقفه نیست (باید از گونههای Interruptibly استفاده کنی). و مهمترین قانون: در خواندن خوشبینانه باید فیلدها را پیش از validate() درون متغیرهای محلی کپی کنی، و نباید وسط خواندن متد دیگری صدا بزنی یا یک مرجع بالقوهناسازگار را واکاوی (dereference) کنی—چون ممکن است آن مرجع نیمهکاره باشد.
AbstractQueuedSynchronizer (AQS): موتور زیرین
حالا پرده را کنار بزنیم. ReentrantLock، Semaphore، CountDownLatch، ReentrantReadWriteLock و حتی دروازهٔ کارگرِ ThreadPoolExecutor—همه پوششهای نازکی روی یک موتور مشترک به نام AQS هستند. اگر این یک قطعه را بفهمی، انگار سورس نصف java.util.concurrent را بلدی.
یک بانک را تصور کن با یک تابلوی «شمارهٔ فعلی» (state) و یک صف مرتب از مشتریها که شماره کشیدهاند. وقتی باجه آزاد است، نفر جلوی صف را صدا میزنند. وقتی مشغول است، تازهواردها یک شماره میگیرند و مینشینند/میخوابند (park) تا نوبتشان شود. وقتی مشتری فعلی کارش تمام شد، دقیقاً نفر بعدی را بیدار میکند (unpark). AQS همین سیستم است: یک عدد وضعیت مشترک، بهعلاوهٔ یک صف منظمِ نخهای خوابیده.
مدل ذهنی دقیق:
- AQS یک تک
volatile int stateو یک صف انتظار FIFO مبتنی بر CLH از نخها نگه میدارد. (CLH نوع خاصی از صف مبتنی بر لیست پیوندی است که هر نخ روی گرهٔ خودش میچرخد/میخوابد.) - یک زیرکلاس تعریف میکند که
stateچه معنایی دارد، و متدهایtryAcquire(int)/tryRelease(int)(برای حالت انحصاری) یاtryAcquireShared/tryReleaseShared(برای حالت اشتراکی) را پیاده میکند. - خودِ AQS بخش سختِ ماشینکاری را انجام میدهد: CAS اتمیک روی
state، صفبندی بازندگان بهصورت گرههای صف، پارککردنشان باLockSupport.park()، و بیدارکردن جانشین هنگام آزادسازی.
حالا ببین چطور همان موتور، چند کلاس مختلف میسازد:
ReentrantLock: stateهمان شمارندهٔ نگهداشت است—۰ یعنی آزاد، N یعنی N بار بازوردانه نگهداشتهشده. tryAcquireمقدار ۰→۱ را CAS میکند، یا اگر نخ فعلی خودش صاحب قفل است، شمارنده را یکی بالا میبرد.Semaphore: stateهمان شمار مجوز (permit) است و از گرفتن اشتراکی استفاده میکند.CountDownLatch: stateهمان شمارنده است، و وقتی به ۰ میرسد همهٔ منتظران را یکجا رد میکند—این «آزادسازی اشتراکی» است، دلیل اینکه یک latch میتواند چندین نخ را همزمان رها کند.
اگر «state + صف CLH + tryAcquire» را بفهمی، دیگر لازم نیست هر کلاس همگامسازی را جدا حفظ کنی—فقط باید بپرسی «این کلاس، state را چه معنا میکند و tryAcquireاش چهکار میکند؟» و بقیهٔ رفتار (صف، park، unpark) از موتور مشترک میآید.
CAS، کلاسهای *Atomic و مسئلهٔ ABA
تا اینجا همهچیز دربارهٔ قفل بود. اما یک لایهٔ بالاتر—و اغلب سریعتر—هم هست: برنامهنویسی بدونقفل (lock-free).
فرض کن روی یک سند مشترک کار میکنی. بهجای اینکه سند را قفل کنی، اینطور عمل میکنی: نسخهٔ فعلی را میخوانی، تغییرت را آماده میکنی، و موقع ذخیره میگویی «فقط اگر سند هنوز همان چیزی است که خواندم، تغییرم را اعمال کن؛ وگرنه بگو شکست خورد.» اگر کسی وسط کار سند را عوض کرده باشد، ذخیرهات رد میشود و از اول تلاش میکنی. این همان compare-and-swap (CAS) است.
CAS یک تک دستور سختافزاری است (lock cmpxchg روی x86، LL/SC روی ARM) که بهصورت اتمیک انجام میدهد: «اگر این خانهٔ حافظه برابر مورد انتظار است، آن را به مقدار جدید بگذار و موفقیت را گزارش کن؛ وگرنه دست نزن و شکست را گزارش کن.» بستهٔ java.util.concurrent.atomic این را در اختیارت میگذارد: AtomicInteger، AtomicLong، AtomicReference و گونههای آرایه و field-updater.
AtomicInteger counter = new AtomicInteger();
int incrementAndGet() {
int prev, next;
do {
prev = counter.get();
next = prev + 1;
} while (!counter.compareAndSet(prev, next)); // تا بردن در مسابقه تلاش مجدد
return next;
}
این حلقه خوشبینانه است: هیچ نخی مسدود نمیشود؛ مناقشه فقط باعث تلاش مجدد میشود.
زیر مناقشهٔ کم، این حلقه قفلها را له میکند—چون اصلاً کسی نمیخوابد و بیدار نمیشود. اما زیر مناقشهٔ سنگین، یک توفان تلاش مجدد راه میافتد: دهها نخ مدام همدیگر را رد میکنند و CPU میسوزد بیآنکه کار مفیدی بشود. دقیقاً به همین دلیل جاوا ۸ **LongAdder/LongAccumulator** را افزود که شمارش را روی چند سلول جدا (که padded شدهاند تا false sharing رخ ندهد) رگهرگه میکند و فقط موقع تقاضا جمع میزند. برای شمارندههای داغ بهمراتب بهتر از یک تک AtomicLong است.
مسئلهٔ ABA
CAS یک نقطهضعف ظریف دارد: برابری مقدار را بررسی میکند، نه اینکه مقدار تغییر کرده و بازگشته باشد.
یک صندوق داری که با یک عدد قفل میشود. قانونت این است: «اگر عدد هنوز A است، آن را باز کن.» یک دزد باهوش عدد را از A به B و بعد دوباره به A برمیگرداند. حالا وقتی تو چک میکنی «هنوز A است؟»، بله هست—پس بازش میکنی، غافل از اینکه در این فاصله همهچیز عوض شده. مقدار همان است، اما دنیا عوض شده.
به زبان دقیق: اگر نخی A را بخواند، نخ دیگری A→B→A را برگرداند، بعد compareAndSet(A, ...) نخ اول موفق میشود—هرچند جهان زیر پایش جابهجا شده. برای یک شمارندهٔ int این بیضرر است. اما برای ساختاری که اشارهگر جابهجا میکند (مثلاً یک پشتهٔ بدونقفل که گرهی را pop میکند که در این فاصله آزاد و بازتخصیص شده)، این وضعیت را خراب میکند. راهحل، جفتکردن مقدار با یک مُهر/نسخهٔ یکنوا (monotonic stamp) است: AtomicStampedReference (مقدار + مُهر int) یا AtomicMarkableReference (مقدار + بولین).
اشتراک کاذب (false sharing): مالیات نامرئی
این یکی از آن باگهای عملکردی است که در کد هیچ ردی از خودش نمیگذارد و فقط پروفایلر پیدایش میکند.
کشها حافظه را نه متغیر به متغیر، بلکه در بستههای ۶۴بایتی به نام خط کش (cache line) جابهجا میکنند. حالا تصور کن دو نفر روی یک سینی مشترک هرکدام یک لیوان جدا دارند. لیوانها مستقلاند، ولی چون روی یک سینیاند، هر بار که یکی لیوانش را جابهجا میکند، مجبور است کل سینی را بردارد و دیگری باید صبر کند تا سینی برگردد. دو متغیرِ مستقل که تصادفاً در یک خط کش نشستهاند دقیقاً همیناند: هر نوشتن روی یکی، نسخهٔ هستهٔ دیگر را باطل میکند و خط میان هستهها «پینگپنگ» میشود، هرچند هیچ اشتراک منطقیای نیست.
این اشتراک کاذب است و میتواند بیصدا یک مرتبهٔ بزرگی (۱۰ برابر) کندت کند.
راهحل: فیلدهای داغ را روی خط کش خودشان بالشتکگذاری (pad) کن تا هرکدام سینیِ جدا داشته باشند. جاوا ۸ به بعد @jdk.internal.vm.annotation.Contended را میدهد (برای کد اپلیکیشن به فلگ -XX:-RestrictContended نیاز دارد)؛ Cell داخلیِ LongAdder دقیقاً همینطور annotate شده—برای همین LongAdder اینقدر خوب مقیاس میگیرد. بالشتکگذاری دستی (long p1..p7) هم کار میکند، اما JIT ممکن است فیلدهای بیاستفاده را حذف کند، پس هرجا در دسترس است @Contended ارجح است.
قفلگذاری دوبار-بررسیشده (double-checked locking)، درستشده
این معروفترین الگوی «بهظاهر درست ولی خراب» در همروندی جاواست. هدف: یک شیء گران را فقط یک بار و تنبل (lazy) بساز، و بعد بدون قفل سریع برش گردان.
اصطلاح کلاسیکِ خراب:
// پیش از جاوا ۵ خراب بود و امروز هم بدون volatile خراب است
private Helper helper;
Helper getHelper() {
if (helper == null) { // بررسی اول (بدون قفل)
synchronized (this) {
if (helper == null) // بررسی دوم (با قفل)
helper = new Helper(); // انتشار
}
}
return helper;
}
چرا بدون volatile خراب است؟ چون helper = new Helper() یک عملیات اتمیک نیست، بلکه سه گام است: (۱) حافظه تخصیص بده، (۲) سازنده را اجرا کن و فیلدها را مقداردهی کن، (۳) مرجع را به helper انتساب بده. و JMM اجازه میدهد که گام ۳ (انتساب مرجع) پیش از گام ۲ (نوشتههای سازنده) دیده شود.
تصور کن نخ دوم روی مسیر سریعِ «بررسی اول» است. بهخاطر آن بازچینش، این نخ میتواند یک helper غیرتهی ببیند که هنوز به یک شیء نیمهساختهشده اشاره میکند—مرجع رسیده، ولی فیلدهای داخلی هنوز مقدار پیشفرض (۰/null) دارند. نخ دوم آن را برمیگرداند و ازش استفاده میکند، و تو یک باگ کاملاً غیرقابلبازتولید داری.
راهحل، volatile کردن فیلد است. این کار مانع release/acquire را درج میکند که تضمین میکند نوشتههای سازنده نسبت به هر خواندن مرجع HB داشته باشند:
private volatile Helper helper; // volatile اجباری است
Helper getHelper() {
Helper result = helper; // یک بار volatile را در محلی بخوان
if (result == null) {
synchronized (this) {
result = helper;
if (result == null)
helper = result = new Helper();
}
}
return result;
}
خواندن در متغیر محلی result یک بهینهسازی واقعی است: در مسیر داغ، دو خواندن volatile را به یکی فرومیکاهد.
اگر یک singleton استاتیک میخواهی، کل این پیچیدگی را دور بزن و از اصطلاح holder با راهاندازی بهتقاضا (initialization-on-demand holder) استفاده کن. این الگو به تضمینهای راهاندازی تنبلِ کلاس در خود JLS تکیه میکند: کلاس داخلی تا اولین دسترسی به Holder.INSTANCE اصلاً بارگذاری نمیشود، و خودِ JVM امنبودن و تنبلبودن این راهاندازی را تضمین میکند—بدون هیچ volatile یا synchronized.
class Singleton {
private Singleton() {}
private static class Holder { static final Singleton INSTANCE = new Singleton(); }
static Singleton getInstance() { return Holder.INSTANCE; } // JVM راهاندازی امن و تنبل را تضمین میکند
}
دامها و نکات ظریف رایج
اینها را مثل یک چکلیست موقع بازبینی کد همروند نگه دار:
- تکیه به زمان بهجای HB. «یک
Thread.sleepهست، حتماً نخ دیگر تا حالا تمام کرده.» sleepهیچ یال HB نمیسازد؛ خواندن کهنه همچنان کاملاً مجاز است. - volatile روی مرجع شیء تغییرپذیر. مرجع را امن منتشر میکند، اما تغییرهای بعدیِ فیلدهای همان شیء پوشش داده نمیشوند. بهجایش یک عکسفوری تغییرناپذیر (immutable snapshot) منتشر کن.
- فیلدهای غیرfinal در اشیای «تغییرناپذیر». فقط فیلدهای
finalتضمین انجماد سازنده را میگیرند؛ یک فیلد غیرfinal اگر شیء از راه مسابقهٔ داده منتشر شود، میتواند با مقدار پیشفرضش دیده شود. - قفل روی
thisدر یک کتابخانه. فراخوانندگان میتوانند روی شیء تو هم قفل بگذارند و مناقشه یا بنبستِ غافلگیرکننده بسازند. از یک قفل خصوصی استفاده کن. - check-then-act روی کالکشنهای همروند.
if (!map.containsKey(k)) map.put(k, v);مسابقه دارد؛ بهجایش ازputIfAbsent/computeIfAbsentاستفاده کن. - فراموشیِ اینکه
unlock()خودکار نیست. قفلهای صریحLockنیاز بهtry/finallyدارند؛ بلوکsynchronizedنمیتواند نشت کند. - فرض اینکه
size()/isEmpty()روی کالکشنهای همروند دقیقاند. اینها عکسفوریهای سازگارِ ضعیف (weakly consistent) هستند، نه اعداد لحظهای دقیق.
بهترین شیوهها
- تغییرناپذیری (immutability) و محصورسازی (confinement، یعنی داده را در یک نخ نگهداشتن) را بر قفلگذاری ترجیح بده؛ ارزانترین قفل، قفلی است که اصلاً نمیگیری.
- پیش از دستساز نوشتنِ قفل، از ابزارهای سطحبالا استفاده کن: کالکشنهای
java.util.concurrent، CompletableFuture، ExecutorService. - ناحیههای بحرانی را کوتاه و بدون I/O نگه دار. هرگز هنگام نگهداشتن قفل، کد بیگانه یا callback صدا نزن—این دعوت مستقیم به بنبست است.
- یک ترتیب سراسری قفل (global lock ordering) برقرار و مستند کن و همهجا قفلها را با همان ترتیب بگیر؛ این جلوی بنبست را میگیرد.
- برای شمارندهها و انباشتگرها زیر مناقشه، بهجای
AtomicLongسراغLongAdderبرو. - هر فیلدی که میان نخها میرود را عمداً منتشر کن—با
final، volatile، یک قفل، یا یک کالکشن همروند. اگر نمیتوانی یال HB مربوطه را نام ببری، یعنی باگ داری.
پرسشهای مصاحبه
رؤیتپذیری و ترتیب از راه یک یال release (نوشتن) / acquire (خواندن): تمام نوشتههای پیش از نوشتن volatile، پس از خواندن بعدیِ همان فیلد دیده میشوند—بهعلاوهٔ اتمیکبودن خودِ فیلد، شامل long/double ۶۴بیتی. اما انحصار متقابل نمیدهد و عملیات مرکبی مانند x++ را اتمیک نمیکند.
یک ترتیب جزئی روی کنشهاست. اگر A HB B، اثرهای حافظهٔ A تضمیناً برای B دیده و پیش از آن مرتباند. این رابطه میان یک release روی یک متغیر همگامسازی و یک acquire بعدی روی همان متغیر برقرار میشود، بهعلاوهٔ ترتیب برنامه، شروع/پیوستن نخ، و ترایایی. دو کنشی که با HB مرتب نیستند و نوشتن متعارض دارند، یک مسابقهٔ داده میسازند.
boolean running = true; // volatile نیست
void stop() { running = false; }
void run() { while (running) { /* کار */ } }
چه میتواند رخ دهد؟ JIT مجاز است running را از داخل حلقه در یک رجیستر بالا ببرد (بالابری حلقه / loop hoisting)، پس run() میتواند تا ابد حلقه بزند، حتی بعد از اینکه stop() برگشته—چون هیچ یال HB نوشتن را وادار به مشاهده نمیکند. volatile کردن running مشکل را حل میکند.
static int a = 0, b = 0;
// نخ ۱: a = 1; int r1 = b;
// نخ ۲: b = 1; int r2 = a;
آیا r1 == 0 && r2 == 0 ممکن است؟ بله. بدون همگامسازی، بارگذاریها و ذخیرهها میتوانند بازچینش شوند (بهخاطر store buffering)، پس هردو نخ میتوانند مقدارِ پیشازنوشتنِ دیگری را بخوانند. سازگاری ترتیبی (sequential consistency) برای کد مسابقهدار تضمین نیست.
helper = new Helper() یعنی «تخصیص + ساخت + انتساب»، و JMM اجازه میدهد انتساب مرجع پیش از نوشتههای فیلدِ سازنده بازچینش (دیده) شود. یک خوانندهٔ مسابقهدار روی مسیر بدونقفل میتواند مرجع غیرتهی به یک شیء نیمهراهاندازیشده ببیند. volatile مانع release/acquire را درج میکند که نوشتههای سازنده را نسبت به خواندن مرجع HB میکند.
حالت «همان نخ مکرراً دوباره قفل میکند» را بهینه میکرد، با مغرضکردن شیء به آن نخ و گذشتن کامل از CAS. اما هزینهٔ ابطال (revocation) و دفترداریاش با بارهای کاری مدرنِ چندنخی و ساختارهای دادهٔ همروند گران شد. JEP 374 آن را در JDK 15 بهطور پیشفرض غیرفعال کرد؛ بعداً حذف شد (JDK 18). قفل سبک (lightweight) و سنگین (inflated) باقی ماندهاند.
ReentrantLock امکانات tryLock، مهلت، قابلیتوقفه، انصاف و چند Condition را میافزاید؛ synchronized سادهتر است و نمیتواند قفل را نشت دهد. از نظر تاریخی در کدِ پر از نخ مجازی ReentrantLock را ترجیح میدادی، چون synchronized نخ حامل را سنجاق (pin) میکرد—اما JEP 491 (JDK 24) این سنجاق را حذف کرد، پس اکنون انتخاب را بر پایهٔ امکانات و راحتی بگذار، نه مقیاسپذیری.
یک volatile int state بهعلاوهٔ یک صف انتظار FIFO مبتنی بر CLH. زیرکلاسها تعریف میکنند state چه معنایی دارد و tryAcquire/tryRelease (یا گونههای اشتراکی) را پیاده میکنند؛ AQS کارِ CAS، صفبندی و park/unpark را انجام میدهد. ReentrantLock (state = شمار نگهداشت)، Semaphore (permits) و CountDownLatch (شمارنده، با آزادسازی اشتراکی) همه بر آن ساخته شدهاند.
CAS مقدارها را مقایسه میکند نه تاریخچه را. اگر مقدار A→B→A برود، CASی که انتظار A دارد موفق میشود، هرچند وضعیت میانی تغییر کرده. برای یک شمارندهٔ عددی بیضرر است؛ اما برای بازاستفادهٔ اشارهگر/گره در ساختارهای بدونقفل خطرناک است—جایی که یک آدرسِ بازاستفادهشده از بررسیِ CAS میگذرد اما به پیوند کهنه ارجاع میدهد. با AtomicStampedReference (مقدار + نسخه) رفعش کن.
دو متغیر مستقل که در همان خط کش ۶۴بایتی نشستهاند، در هر نوشتن باعث ابطال میانهستهای میشوند و خط را میان هستهها پینگپنگ میکنند. با شمارندههای perf (مناقشهٔ خط کش / رویدادهای HITM) یا با پروفایلینگ تشخیصش بده؛ با بالشتکگذاری فیلدهای داغ روی خطوط جدا رفعش کن، مثلاً با @Contended (همان کاری که Cell در LongAdder میکند).
وقتی خواندنها غالباند و میتوانی با خواندن خوشبینانه در مسیر خواندن اصلاً به خط کش قفل دست نزنی. تلهها: بازورودی نیست، Condition ندارد، و خواندن خوشبینانه باید فیلدها را در متغیر محلی کپی و سپس validate() کند—خواندن از یک مرجعِ بالقوهکهنه در وسط خواندن، یک باگ است.
نه. زیر مناقشهٔ سنگین، حلقهٔ CAS-retry آن میکوبد و CPU هدر میرود. LongAdder/LongAccumulator افزایشها را روی سلولهای padded رگهرگه میکنند و تنبل جمع میزنند، و برای شمارندههای write-hot و read-rare خیلی بهتر مقیاس میگیرند—به بهای sum() کمی گرانتر و نبودِ یک read-modify-write اتمیک روی کل مقدار.
برای یک long/double سادهٔ غیرvolatile تضمین نیست—JLS اجازه میدهد نوشتن ۶۴بیتی به دو ذخیرهٔ ۳۲بیتی شکسته شود، پس یک خوانندهٔ مسابقهدار میتواند مقدار پارهشده (torn) ببیند—نیمی قدیم، نیمی جدید. volatile، AtomicLong یا یک قفل این پارهشدن را حذف میکند.
فیلدهای final یک انجماد ویژه در پایان سازنده میگیرند: هر نخی که مرجع شیء را بخواند، تضمیناً فیلدهای finalِ درستراهاندازیشده را میبیند، حتی زیر انتشار مسابقهدار. اما فیلد سادهٔ (غیرfinal) چنین تضمینی ندارد—نخ دیگر میتواند مقدار پیشفرضش (۰/null) را ببیند. برای همین «تغییرناپذیر = همهٔ فیلدها final» یک قاعدهٔ درستی است، نه سلیقه.
تنزل (گرفتن قفل خواندن هنگام نگهداشتنِ نوشتن، سپس آزادکردن نوشتن) امن است، چون هرگز لازم نیست منتظر رسیدنِ خوانندگان/نویسندگان دیگر به وضعیت قویتر بمانی. اما ارتقا (نگهداشتنِ خواندن، درخواست نوشتن) نیازمند آزادکردن همهٔ خوانندگانِ دیگر است—و اگر دو خواننده هردو بخواهند ارتقا دهند، هرکدام منتظر میماند دیگری قفل خواندنش را رها کند: بنبست. برای همین API آن را ممنوع میکند.
نکاتِ سنیور و موارد پیشرفته
تا اینجا مدل ذهنیِ درست را ساختی و ابزارها را میشناسی. حالا میرویم سراغ لایهای که مرزِ یک سنیورِ واقعی از یک برنامهنویسِ خوب را مشخص میکند: تضمینهای عمیقترِ خودِ JMM، انتشار امن، نردبان access-modeهای مدرن با VarHandle، معناشناسیِ درستِ wait/notify، ردهبندی تضمینهای پیشرفت (progress)، و باگهایی که فقط زیر بار تولید و با یک profiler خودشان را نشان میدهند.
۱) SC-DRF و مقادیرِ out-of-thin-air. ۲) تفاوتِ coherence و consistency. ۳) انتشار امن و فرارِ this. ۴) VarHandle و نردبانِ plain/opaque/acquire-release/volatile و fenceها. ۵) معناشناسیِ wait/notify/Condition و بیداریِ کاذب. ۶) ردهبندیِ wait-free / lock-free / obstruction-free. ۷) deadlock/livelock/starvation و بهینهسازیهای JITِ قفل. ۸) گاچاهای نامرئی و ابزارِ تست. در آخر ۹ پرسشِ سختِ مصاحبه.
۱) قضیهٔ SC-DRF: دقیقاً JMM چه قول میدهد؟
فصل نشان داد که کدِ ریسی میتواند مقادیرِ کهنه یا پاره ببیند. اما نکتهٔ آرامشبخش این است: JMM یک قضیهٔ محوری دارد به نامِ SC-DRF (Sequential Consistency for Data-Race-Free programs).
اگر برنامهات کاملاً همگامسازیشده باشد—هیچ جفتدسترسیِ متضاد (حداقل یکی نوشتن) بدونِ یالِ happens-before نداشته باشد، یعنی بدونِ مسابقهٔ داده (DRF)—آنگاه JMM تضمین میکند دقیقاً مثلِ یک برنامهٔ پیوستهٔ ترتیبی (sequentially consistent) رفتار میکند: یک ترتیبِ سراسریِ واحد که همهٔ نخها رویش توافق دارند. تا وقتی DRF باشی، تمام آن بازچینشهای ترسناک نامرئیاند.
یعنی هدفت ساده میشود: لازم نیست دربارهٔ تکتکِ بازچینشها فکر کنی، فقط اثبات کن DRF هستی—و آن جملهٔ طلاییِ فصل معنا میگیرد: «اگر نتوانی یالِ HB را نام ببری، باگ داری.» اما آنطرفِ سکه: برای کدِ racy، JMM هنوز یک مشکلِ حلنشدهٔ نظری دارد.
مدلِ رسمیِ JMM برای جلوگیری از مقادیرِ «از هیچ» (out-of-thin-air) قواعدِ علّیت دارد—نباید مقداری ظاهر شود که هیچ نوشتهای تولیدش نکرده و از یک حدسِ خودتوجیهگر بیرون آمده. اما ثابت شده فرمولبندیِ فعلی هم بعضی اجراهای مطلوب را ممنوع میکند و هم OoTA را کامل نمیبندد؛ این «سوراخِ» شناختهشدهٔ مدل و انگیزهٔ کار روی JMM جدید است. نتیجهٔ عملی: هرگز روی رفتارِ کدِ racy حساب نکن؛ حتی مقادیرِ «غیرممکن» هم مجازند.
۲) Coherence در برابر Consistency — یک اصلاحِ ظریف
تشبیهِ «کپیِ کهنه روی میزِ آشپز» برای شهود عالی است، ولی سنیور باید دقیقتر بداند. سختافزارِ مدرن cache coherence را با پروتکلهایی مثل MESI تضمین میکند: برای یک آدرسِ واحد، همهٔ هستهها روی یک ترتیبِ سراسریِ واحد از نوشتنها توافق دارند و یک نوشتن، کپیِ هستههای دیگر را باطل (invalidate) میکند. پس کشها واقعاً «برای همیشه» کهنه نمیمانند.
پس چرا هنوز مقدارِ کهنه میبینی؟ دو دلیل: (۱) store buffer — نوشتنِ یک هسته لحظهای در بافرِ محلی مینشیند و با تأخیر به کشِ منسجم میرسد؛ در این فاصله هستهٔ خودت مقدارِ جدید را میبیند ولی بقیه نه (همان «store buffering»ِ پشتِ سؤالِ کلاسیکِ r1==0 && r2==0). (۲) کامپایلر/JIT که مقدار را در رجیستر cache میکند (hoisting). پس مشکل coherence نیست بلکه consistency (ترتیبِ بینِ آدرسها) است. volatile و قفل fence صادر میکنند که store buffer را تخلیه و بازچینش را مهار میکند—«کش را تازه» نمیکنند.
۳) انتشارِ امن (Safe Publication) و فرارِ this
فصل «انتشارِ عمدی» را بهعنوان قاعدهٔ طلایی گفت. حالا دقیقتر: چهار راهِ متعارفِ انتشارِ امنِ یک شیء وجود دارد:
- مقداردهیِ اولیهٔ آن از یک بلاکِ static initializer (تضمینِ init کلاس).
- ذخیرهٔ ارجاعش در یک فیلدِ
volatile(یاAtomicReference). - ذخیرهٔ ارجاعش در یک فیلدِ
finalکه در سازنده مقدار میگیرد. - ذخیرهٔ ارجاعش با محافظتِ یک قفل (یا گذاشتنش در یک کالکشنِ concurrent).
تضمینِ freezeِ فیلدهای final فقط وقتی معتبر است که ارجاعِ شیء قبل از پایانِ سازنده فرار نکند. اگر داخلِ سازنده this بیرون درز کند، این تضمین میشکند:
public class Listener {
private final int id;
public Listener(EventBus bus) {
bus.register(this); // ❌ this فرار کرد—هنوز سازنده تمام نشده
this.id = computeId(); // نخِ دیگر ممکن است id را 0 ببیند
}
}
الگوهای رایجِ فرارِ this: ثبتِ listener/callback، شروعِ نخ داخلِ سازنده، یا انتشارِ this در متغیرِ static. راهِ درست: سازنده را کامل کن، بعد در متدِ کارخانهایِ static create(...) شیء را بساز و سپس register کن.
۴) VarHandle و نردبانِ access-modeها — مدلِ مدرنِ حافظه
فصل CAS را با AtomicInteger نشان داد. اما از Java 9 (JEP 193) ابزارِ سطحِ پایین و رسمیِ همهٔ اینها VarHandle است—جانشینِ امنِ sun.misc.Unsafe و Atomic*FieldUpdater (کلاسهای java.util.concurrent خودشان از JDK 9 رویش مهاجرت کردهاند).
هدیهٔ اصلیاش: بهجای دوگانهٔ «plain یا volatile»، یک نردبانِ چهارپلهای از قوّتِ ترتیب میدهد، از ضعیف به قوی.
نردبانِ قوّتِ access-mode در VarHandle (از ضعیف به قوی):
flowchart LR
Plain["Plain<br/>get/set"] --> Opaque["Opaque<br/>getOpaque/setOpaque"]
Opaque --> RelAcq["Release/Acquire<br/>setRelease/getAcquire"]
RelAcq --> Volatile["Volatile (SC)<br/>getVolatile/setVolatile"]
- Plain: بیترتیب و بیرؤیتپذیریِ بینِ نخی—فقط اتمیکِ خودِ دسترسی (بهجز long/double). مثلِ فیلدِ معمولی.
- Opaque: دسترسیها پاک نمیشوند و برای یک آدرس coherent میمانند، ولی نسبت به آدرسهای دیگر بازچینشپذیر است—برای فلگِ توقف یا شمارندهٔ آماری که فقط باید «بالاخره» دیده شود، از volatile ارزانتر.
- Release/Acquire: همان نیمهحصارِ release/acquireِ happens-before، ولی بدونِ حصارِ سنگینِ StoreLoad—«امضا/دیدنِ امضا» بیپرداختِ هزینهٔ کاملِ SC.
- Volatile: قویترین—معادلِ کلیدواژهٔ
volatile، با ترتیبِ سراسریِ SC.
class Node {
Object item;
volatile Node next; // فیلدِ عادی، دستکاری با VarHandle
private static final VarHandle NEXT;
static {
try {
NEXT = MethodHandles.lookup()
.findVarHandle(Node.class, "next", Node.class);
} catch (ReflectiveOperationException e) { throw new ExceptionInInitializerError(e); }
}
boolean casNext(Node expect, Node update) {
return NEXT.compareAndSet(this, expect, update);
}
void publish(Node n) { NEXT.setRelease(this, n); } // release، ارزانتر از setVolatile
}
دو نکتهٔ سنیور: (۱) compareAndExchange مثل compareAndSet است ولی بهجای boolean مقدارِ شاهد (witness) را برمیگرداند—در حلقهٔ retry یک get()ِ اضافه را حذف میکند. (۲) weakCompareAndSet مجاز است کاذب (spuriously) شکست بخورد حتی وقتی مقدار برابرِ expected است؛ در عوض روی معماریهای LL/SC مثلِ ARM ارزانتر است. برای همین فقط داخلِ حلقهٔ do/whileِ خودت درست است، نه بهتنهایی.
VarHandle چهار متدِ استاتیکِ fence هم دارد (fullFence، acquireFence، releaseFence، loadLoadFence/storeStoreFence) که حصارِ حافظه صادر میکنند بدونِ گرهخوردن به فیلدِ خاص—معادلِ مدرن و امنِ Unsafe.fullFence برای الگوهای انتشارِ دستیِ خیلی خاص.
۵) wait / notify / Condition — معناشناسیِ درست
فصل قفل و AQS را گفت اما به هماهنگیِ نخها با wait/notify نپرداخت—و این یکی از پرباگترین نقاطِ کدِ واقعی است.
wait() یعنی «کلید را پس بده و در اتاقِ انتظار بخواب» و notify() یعنی «یکی از خوابها را صدا بزن». نکتهٔ حیاتی: نخِ بیدارشده بلافاصله اجرا نمیشود—باید دوباره برای قفل رقابت کند، پس بینِ «بیدار شدن» و «اجرا شدن» دنیا ممکن است عوض شده باشد.
سه قانونِ آهنین:
wait/notify/notifyAllرا فقط وقتی صدا بزن که monitorِ همان شیء را در دست داری، وگرنهIllegalMonitorStateExceptionمیخوری.- همیشه
waitرا داخلِ یک حلقه بگذار که شرط را دوباره چک میکند، نه داخلِif. دو دلیل: بیداریِ کاذب (spurious wakeup) که JLS صراحتاً مجازش میداند، و اینکه ممکن است تا وقتی نوبتت شود شرط دوباره باطل شده باشد.
synchronized (queue) {
while (queue.isEmpty()) { // ✅ while نه if
queue.wait();
}
return queue.remove();
}
notify() فقط یک نخ را بیدار میکند و کدام را هم تضمین نمیکند. اگر چند نخ روی شرطهای متفاوتِ همان monitor منتظر باشند، ممکن است نخِ «اشتباه» را بیدار کند که شرطش برقرار نیست؛ او دوباره میخوابد و سیگنال گم میشود—نخهای واجدِ شرایط هرگز بیدار نمیشوند (deadlock نرم). قاعده: مگر اینکه همهٔ waiterها همارز باشند، notifyAll؛ بهتر: با Condition (از ReentrantLock) صفهای مجزا (notFull/notEmpty) بساز و فقط نخِ مرتبط را بیدار کن—هم امنتر، هم کارآمدتر.
زیرِ AQS، مکانیزمِ خواباندن LockSupport.park()/unpark() است، نه wait/notify. تفاوتِ ظریف: unpark یک permit میدهد که انباشته نمیشود (حداکثر یکی) ولی میتواند قبل از park برسد—پس «سیگنالِ گمشده»ی wait/notify را ندارد. park هم میتواند بهصورتِ کاذب برگردد، پس آن هم باید داخلِ حلقهٔ شرط باشد.
۶) ردهبندیِ تضمینهای پیشرفت: wait-free / lock-free / obstruction-free
فصل CAS را «lock-free» نامید. سنیور باید این اصطلاح را دقیق بداند چون در مصاحبه تله میگذارند.
- Obstruction-free: یک نخِ بدونِ مزاحمت در گامهای محدود تمام میکند. (ضعیفترین.)
- Lock-free: همیشه حداقل یک نخ پیشرفت میکند—سیستم گیر نمیکند، هرچند یک نخِ بدشانس ممکن است تا ابد retry کند (starvation).
- Wait-free: هر نخ در گامهای محدود تمام میکند—بیstarvation. (قویترین.)
آن حلقهٔ do { ... } while(!cas())ِ فصل lock-free است نه wait-free: یکی همیشه برنده میشود، اما یک نخِ خاص میتواند بارها ببازد و گرسنه بماند. برای همین زیرِ contentionِ سنگین LongAdder بهتر است. و برخلاف تصورِ رایج «lock-free» یعنی «سریعتر» نیست—زیرِ بارِ سنگین یک قفلِ منصفانه ممکن است throughputِ بهتری بدهد. Thread.onSpinWait() (Java 9، JEP 285) هم در حلقههای spin کمک میکند: هینتِ PAUSE به CPU میدهد تا مصرفِ توان و ترافیکِ coherence کم شود.
۷) Deadlock در برابر Livelock در برابر Starvation و بهینهسازیهای JIT
فصل «ترتیبِ سراسریِ قفل» را برای جلوگیری از deadlock گفت. سه شکستِ همروندی را از هم جدا کن:
- Deadlock: نخها برای همیشه منتظرِ هماند (چرخهٔ انتظار). چهار شرطِ Coffman لازماند (mutual exclusion، hold-and-wait، no-preemption، circular-wait)؛ شکستنِ یکی کافی است—معمولاً circular-wait را با ترتیبِ سراسریِ قفل.
- Livelock: نخها بلاک نیستند و مدام کار میکنند ولی پیشرفتی نمیکنند—مثلِ حلقهٔ
tryLockای که هر بار شکست، قفل را رها و فوراً دوباره تلاش میکند؛ راهِ حل backoffِ تصادفی. - Starvation: بعضی نخها پیشرفت میکنند اما یکی هرگز نوبت نمیگیرد (writer starvation در ReadWriteLock).
سه بهینهسازیِ HotSpot که فصل نگفت: (۱) Lock elision — اگر escape analysis ثابت کند شیء از یک نخ فرار نمیکند، JIT قفلش را کاملاً حذف میکند (مثلِ StringBufferِ محلی). (۲) Lock coarsening — چند synchronizedِ پشتِسرِ همِ روی یک شیء را در یک ناحیه ادغام میکند تا lock/unlockِ مکرر حذف شود. (۳) Adaptive spinning — قبل از park کمی spin میکند به امیدِ آزادیِ زودِ قفل و مدتش را از تاریخچهٔ همان قفل تنظیم میکند. یعنی micro-benchmarkِ ساده اغلب هزینهٔ واقعیِ قفل را نشان نمیدهد.
۸) گاچاهای نامرئی که در تولید میزنند
volatile int[] data; // فقط خودِ ارجاعِ آرایه volatile است
data[5] = 42; // ❌ این نوشتن هیچ معناشناسیِ volatile ندارد!
یک اشتباهِ فاجعهبار: نوشتن در data[i] یک نوشتنِ عادی است و هیچ یالِ HB نمیسازد. اگر عناصر باید ترتیب/رؤیتپذیریِ volatile داشته باشند، از AtomicIntegerArray یا VarHandle رویِ آرایه (MethodHandles.arrayElementVarHandle) استفاده کن.
ThreadLocal مقدار را به نخ گره میزند؛ در thread pool نخها بازاستفاده میشوند و نمیمیرند. اگر بعد از هر task مقدار را remove() نکنی: (۱) دادهی یک درخواست به درخواستِ بعدیِ همان نخ نشت میکند (باگِ امنیتی/صحت)، و (۲) شیءِ بزرگ تا ابد زنده میماند (memory leak). همیشه در finally یا فیلترِ خروجی remove() کن.
یک اصلِ مستند اما کمدانسته: کلاسهای java.util.concurrent تضمینِ «memory consistency» میدهند—گذاشتنِ یک شیء در BlockingQueue happens-before برداشتنش با take()، و نوشتن در ConcurrentHashMap قبل از خواندنِ بعدیِ همان کلید HB است. یعنی وقتی شیءِ mutable را از میانِ صفِ concurrent رد میکنی، به volatileِ جدا نیاز نداری—خودِ صف انتشارِ امن را انجام میدهد.
نوشتنِ کدِ همروند بدونِ تستِ درست خودفریبی است—باگ ممکن است روی x86ات دیده نشود ولی روی ARM یا زیرِ بار بترکد. jcstress (رسمیِ OpenJDK) litmus test میسازد و میلیونها بار اجرا میکند تا بازچینشهای نادر را شکار کند—برای اثباتِ DRF بودنِ یک الگو بیبدیل. و JMH برای بنچمارکِ درست که lock elision/coarsening و warmupِ JIT را لحاظ میکند (یک System.nanoTime()ِ دستی تقریباً همیشه دروغ میگوید).
قضیهٔ SC-DRF: اگر برنامه data-race-free باشد (هر جفتدسترسیِ متضاد با یالِ happens-before مرتب)، JMM تضمین میکند اجرا پیوستهٔ ترتیبی بهنظر میرسد—یک ترتیبِ سراسریِ واحد و همهٔ بازچینشها نامرئی. پس بهجای استدلال دربارهٔ تکتکِ بازچینشها، فقط اثبات کن DRF هستی. آنطرفِ سکه: برای کدِ racy مدل حتی «out-of-thin-air» را کامل نمیبندد و سوراخِ نظری دارد؛ هرگز روی رفتارِ racy حساب نکن.
نه. سختافزار cache coherence را (با MESI) تضمین میکند: برای یک آدرس یک ترتیبِ سراسریِ نوشتن هست و یک نوشتن کپیِ بقیه را باطل میکند. کهنگی از store buffer (نوشتنِ محلیِ هنوز-flush-نشده) و cacheکردنِ JIT در رجیستر میآید. پس مشکلْ coherence نیست بلکه memory consistency (ترتیبِ بینِ آدرسها) است. volatile/قفل با fence، store buffer را تخلیه و بازچینش را مهار میکنند؛ «کش را رفرش» نمیکنند.
چهار پله از ضعیف به قوی: plain (فقط اتمیک، بیترتیب)، opaque (coherence و progress برای همان متغیر، ولی نسبت به بقیه بازچینشپذیر—برای فلگِ توقف یا شمارندهٔ آماری ارزانتر از volatile)، release/acquire (نیمهحصار؛ همان HBِ امضا/دیدنِ امضا بدونِ حصارِ سنگینِ StoreLoad)، و volatile (قویترین، SC سراسری). قاعده: ضعیفترین را بگیر که درستیِ الگو را تأمین کند؛ در ساختارهای lock-free معمولاً release برای publish و acquire برای خواندن کافی است.
دو دلیل: (۱) بیداریِ کاذب (spurious wakeup) که JLS صراحتاً مجاز میداند—نخ میتواند بدونِ هیچ notifyای بیدار شود. (۲) بینِ لحظهٔ بیدار شدن و لحظهٔ واقعیِ گرفتنِ قفل، نخِ دیگری ممکن است شرط را دوباره باطل کرده باشد (مثلاً صف که خالی نبود، دوباره خالی شد). پس باید بعد از بیداری شرط را دوباره چک کنی؛ فقط while (!condition) wait(); امن است. همین دلیل برای Condition.await() و LockSupport.park() هم صادق است.
notify() یک نخِ نامشخص را بیدار میکند. اگر waiterها روی شرطهای متفاوت منتظر باشند، ممکن است نخی بیدار شود که شرطش برقرار نیست، دوباره بخوابد و سیگنال گم شود—نخی که باید بیدار میشد هرگز بیدار نمیشود (deadlockِ نرم/lost-wakeup). پیشفرض notifyAll مگر اینکه همهٔ waiterها همارز باشند؛ بهتر: Conditionهای مجزا (notFull/notEmpty) و signal فقط نخِ مرتبط.
obstruction-free: یک نخِ بدونِ مزاحمت در گامهای محدود تمام میشود. lock-free: همیشه حداقل یک نخ پیشرفت میکند (سیستم گیر نمیکند) اما یک نخِ خاص میتواند گرسنه بماند. wait-free: هر نخ در گامهای محدود تمام میشود (بیstarvation). حلقهٔ do{}while(!cas()) lock-free است نه wait-free—یکی همیشه برنده میشود ولی یک بدشانس بارها میبازد. پس زیرِ contentionِ سنگین LongAdder؛ و «lock-free» لزوماً «سریعتر» نیست.
اگر ارجاعِ شیء قبل از پایانِ سازنده درز کند (ثبتِ listener، start کردنِ نخ، انتشار در فیلدِ static)، تضمینِ freezeِ final میشکند: نخِ دیگری میتواند فیلدهای مقداردهینشده (0/null) ببیند، چون نوشتنهای سازنده کامل نشدهاند. حتی تکنخی هم خطرناک است چون کدِ کالبک ممکن است روی شیءِ نیمهساخته کار کند. راهِ درست: سازنده را کامل کن و انتشار را به متدِ کارخانهایِ static بعد از ساخت منتقل کن.
فقط ارجاعِ آرایه volatile است نه عناصرش. نوشتنِ a[i] یک نوشتنِ عادی است و هیچ یالِ happens-before نمیسازد—نخِ دیگر ممکن است مقدارِ کهنه ببیند یا اصلاً نبیند. برای عناصر از AtomicIntegerArray یا MethodHandles.arrayElementVarHandle استفاده کن. همین تله برای AtomicReference<mutableObject> هم هست: ارجاع امن منتشر میشود ولی دستکاریِ بعدیِ فیلدهای شیء پوشش نمییابد.
deadlock: نخها بلاک و در چرخهٔ انتظارِ متقابلاند؛ چهار شرطِ Coffman لازم و شکستنِ یکی (معمولاً circular-wait با ترتیبِ سراسریِ قفل) کافی است. livelock: نخها بلاک نیستند و مدام کار میکنند ولی پیشرفتی نمیکنند—مثلِ حلقهٔ tryLockای که هر بار شکست، رها و فوراً دوباره تلاش میکند. رفع: backoffِ تصادفی. هر دو با thread dump و detectِ چرخه عیبیابی میشوند، ولی جلوگیری بهتر از درمان است.
- SC-DRF: اگر DRF باشی دنیا پیوستهٔ ترتیبی بهنظر میرسد؛ برای کدِ racy حتی OoTA بسته نیست. کهنگی از store buffer و رجیستر میآید نه ناهماهنگیِ کش (consistency نه coherence).
- انتشارِ امن چهار راه دارد و با فرارِ
thisدر سازنده میشکند. VarHandle نردبانِ plain/opaque/release-acquire/volatile را میدهد—ضعیفترینِ درست را بگیر. waitرا همیشه در حلقه بگذار؛ پیشفرضnotifyAllیاConditionهای مجزا. حلقهٔ CAS lock-free است نه wait-free؛ زیرِ بارLongAdder.- deadlock/livelock/starvation را جدا کن؛ JIT قفل را elide/coarsen میکند. گاچاها:
volatileآرایه فقط ارجاع، نشتِThreadLocalدر pool، و HB رایگانِ کالکشنهای concurrent. با jcstress/JMH تست کن.
- حافظه در سطح چندهستهای مثل دفترچهٔ مشترک نیست؛ کامپایلر، CPU و کش، کدت را بازچینش میکنند. JMM (فصل ۱۷ JLS، بازنویسیشده با JSR-133 در جاوا ۵) قرارداد رسمیِ این دنیاست.
- happens-before تنها قاعدهای است که اهمیت دارد: دربارهٔ ترتیب است نه زمان، و فقط میان یک release و یک acquire بعدی روی همان متغیر همگامسازی برقرار میشود.
volatileرؤیتپذیری/ترتیب و اتمیکبودنِ فیلد (حتی ۶۴بیتی) میدهد، اما انحصار متقابل نمیدهد وx++را اتمیک نمیکند.synchronizedانحصار + رؤیتپذیری + بازورودی میدهد و کلید را نشت نمیدهد؛ biased locking حذف شده (JEP 374 در JDK 15، حذف در JDK 18)، و از JDK 24 با JEP 491 دیگر نخ مجازی را سنجاق نمیکند.- خانوادهٔ
Lockانعطاف میافزاید: ReentrantLock(tryLock/مهلت/انصاف/Condition)، ReadWriteLock(تنزل مجاز، ارتقا ممنوع)، StampedLock(خواندن خوشبینانه، اما بدون بازورودی). - AQS موتور مشترک زیر همهٔ اینهاست:
volatile int state+ صف CLH + tryAcquire. - CAS برنامهنویسی بدونقفل میسازد اما مسئلهٔ ABA دارد (با
AtomicStampedReferenceرفع کن)؛ زیر مناقشهٔ سنگین سراغLongAdderبرو. - مراقب false sharing (با
@Contendedpad کن) و DCL بدون volatile باش؛ برای singleton از holder idiom استفاده کن. - قانون طلایی: هر فیلد بیننخی را عمداً منتشر کن. اگر نمیتوانی یال HB را نام ببری، باگ داری.
Let's be honest: most concurrency bugs happen because a programmer assumes memory behaves like a shared notebook that every thread reads from the same page at the same instant. On a real multi-core machine, memory does not work like that at all. This chapter is going to fully break that wrong mental model and build a correct one in its place — and then construct, one at a time and from scratch, the tools you use to control ordering and visibility.
First we see why the JMM exists at all and the three agents that reorder your code. Then we reach the heart of everything: happens-before, the only rule that guarantees one thread sees another's writes. Next we build the tools from light to heavy: volatile, synchronized and its lock optimizations, the Lock family (ReentrantLock / ReadWriteLock / StampedLock), the underlying engine AQS, lock-free programming with CAS and the ABA problem, false sharing, and finally the famous double-checked locking idiom. We close with a full interview questions section and a nutshell summary.
Part 0 — words you must know first
Before anything else, let's anchor a few terms with analogies so I never drop them on you cold later.
Think of a thread as a chef working through your recipe (your code). A multi-core CPU means several chefs cooking at once. Here's the key detail: each chef has a little countertop right beside them (their core's cache), while the main pantry (RAM) is across the kitchen. When a chef grabs the salt, they put a copy on their own counter and use that copy for a while. So if another chef changes the salt in the main pantry, the first chef won't notice for a while — they still see the stale copy on their counter. That stale copy is the root of most visibility bugs in Java.
- Reordering: shuffling the execution order of instructions for speed — like a chef doing tasks in the most efficient order rather than the written order.
- Memory barrier / fence: a special instruction telling the hardware "actually finish everything up to here and make it visible before moving on" — like a rule that says "don't start the next dish until the current one is on the pass."
- Atomic: an operation that either happens completely or not at all; no other thread can observe a half-done intermediate state.
Mental model: why the JMM exists
Naively you imagine your program as a single, sequential list of memory reads and writes executed exactly in source order, instantly visible to every thread. Convenient as it is, that model is a lie. Between your source code and the electrons in the chip, there are three reordering agents:
- The compiler (javac + JIT): may reorder, hoist, eliminate, or fold instructions. "Hoist" means taking something read every loop iteration and reading it once outside the loop into a register — you'll soon see why that can be catastrophic.
- The CPU: executes out-of-order and sometimes speculatively (it guesses which branch runs and starts it early).
- The memory hierarchy: store buffers and per-core caches delay when one core's write becomes visible to another — that "stale copy on the chef's counter."
You're tempted to think these three agents are "saboteurs." The opposite: all modern performance comes from exactly these reorderings. The platform makes you only one promise: for single-threaded code it guarantees as-if-serial semantics — any reordering is allowed as long as it never changes the result you would observe in that same thread. The trouble starts precisely when another thread wants to watch you mid-work.
This is where the Java Memory Model (JMM) enters. It's specified in JLS Chapter 17 and was reworked by JSR-133 (Java 5), and it's the contract that tells you exactly which cross-thread reads are allowed to see which writes. Everything else in this chapter — volatile, synchronized, Lock, atomics — is really a way to buy happens-before edges from that contract.
Happens-before: the only rule that really matters
If you take just one thing from this chapter, make it this.
Imagine two departments in a company. Department A prepares a document and Department B must work on it. With no procedure between them, B might grab a half-finished draft or even yesterday's version. But if the rule is "A puts the document in the shared inbox and signs it, and B only picks it up after seeing that signature," then everything A wrote before signing is guaranteed visible to B who reads after the signature. That pairing — "sign" (release) and "see the signature" (acquire) — is exactly a happens-before edge.
Now let's be precise. The JMM is defined in terms of a partial order called happens-before (HB). "Partial" means not every two actions are comparable; some pairs are ordered and some aren't. The rule: if action A happens-before action B, then all of A's memory effects are visible to and ordered before B.
And here's the dangerous part: if two actions are not ordered by HB and at least one is a write to the same location, you have a data race. In that case the JMM permits the reads to return stale, torn (half the old value, half the new), or even seemingly "impossible" values.
The HB edges you actually get — memorize this list:
- Program order: within a single thread, each action HB every later action in that thread.
- Monitor lock: an
unlockon a monitor HB every subsequentlockon that same monitor. - Volatile: a write to a
volatilefield HB every subsequent read of that same field. - Thread start:
Thread.start()HB every action in the started thread. - Thread join: every action in a thread HB another thread's successful return from
join()on it. - Interrupt: a call to
interrupt()HB the interrupted thread detecting it. - Final fields: the end of a constructor HB the freeze of
finalfields (the basis for safe publication of immutable objects). - Transitivity: if A HB B and B HB C, then A HB C. Edges chain together.
Let's bury the biggest misconception right here. Two events can be perfectly "simultaneous" in wall-clock terms; what matters is not which happened "earlier" but whether the model forces one's writes to be seen by the other. And that forcing is established only between a release action and a subsequent acquire action on the same synchronization variable. Releasing lock A tells you nothing about a thread that acquires lock B — inbox A's signature is useless for inbox B.
Let's watch the chain with a concrete example:
Thread 1 Thread 2
-------- --------
data = 42; (1)
ready = true; (2, volatile write / release)
while(!ready) {} (3, volatile read / acquire)
print(data); (4) -> guaranteed to see 42
Why does it work? Because (2) is a volatile write and (3) a volatile read of the same field, (2) HB (3). By program order (1) HB (2), and also (3) HB (4). Now chain via transitivity: (1) HB (2) HB (3) HB (4), so (1) HB (4). That means the read of data at line (4) must observe 42.
Remove volatile from ready and two things break at once: the visibility of data (Thread 2 may see a stale value) and the termination of the loop. The JIT is allowed to hoist ready out of the loop into a register — then while(!ready) keeps seeing the same register value and loops forever, even after Thread 1 changed the value. This is one of the most common real-world bugs.
volatile: what it does and what it doesn't
Now that you have happens-before, volatile becomes simple. Picture it as a glass jar everyone must read directly from the main pantry, never from their counter copy — and every time someone puts something in it, they must first commit all their prior work to the pass.
volatile gives you exactly two things:
- Visibility + ordering (release/acquire): a volatile write establishes an HB edge, making not just that variable but all writes that preceded it in program order visible to a thread that subsequently reads the same volatile. This "piggybacking" is what made the
data/readyexample work. - Atomicity of the field itself, even for
long/double. Note: under the JLS, a non-volatile 64-bit write may be split into two 32-bit stores;volatileforbids that tearing.
And what volatile does not give you — this is where most people go wrong:
- No mutual exclusion.
volatile int x; x++;is really a read-modify-write (read, add, write) and is not atomic. Two threads can both read 5, both compute 6, both write 6 — and one increment is lost. Use anAtomicIntegeror a lock. - No compound-action atomicity.
if (v == null) v = new X();still races even ifvis volatile, because there's a gap between the check and the assignment.
On x86, a volatile write compiles to a plain store followed by a store-load barrier (often a lock-prefixed instruction or mfence); but reads are essentially free, because x86 has a strong memory model called TSO (Total Store Order). On weaker memory models like ARM and Power, both reads and writes must emit real fence instructions. So the same Java code may cost more on an ARM phone than on an x86 server.
synchronized: the intrinsic monitor
If volatile was the shared jar, synchronized is the key to a room only one person enters at a time.
Every Java object carries an invisible lock called its monitor (intrinsic lock) — just like a restroom with the key hanging on the door. To enter you take the key (lock), and on exit you hang it back (unlock). While the key is in your hand, nobody else can enter. A synchronized block does this "take and hang the key" automatically — even if an exception is thrown mid-work, the exit behaves like a finally and returns the key, so you can never "leak" it.
The guarantees of synchronized:
- Mutual exclusion: only one thread holds a given monitor at a time.
- Visibility: unlock HB the subsequent lock on the same monitor, so the whole critical section's writes publish.
- Reentrancy: the same thread may re-acquire a monitor it already holds (a per-thread hold count is kept). That's what lets a
synchronizedmethod call anothersynchronizedmethod onthiswithout deadlocking itself.
A subtle point about what you lock on: synchronized(this) and a synchronized instance method both lock on this; a synchronized static method locks on the Class object.
synchronized("lock") or locking on an Integer that came from boxing is a classic trap. String literals are interned into a shared pool, and small Integers are cached, so that exact object may also be used by completely unrelated code — and suddenly two unrelated parts of your program deadlock on a shared lock without knowing it. Always create a dedicated private lock: private final Object lock = new Object();.
Lock optimizations — and what got removed
Historically HotSpot layered three locking schemes worth knowing, especially since one was removed and is a favorite interview question:
- Biased locking: assumed a lock is usually re-acquired by the same thread, so it "biased" the object to that thread and skipped atomic CAS entirely on re-entry. It was disabled by default and deprecated in JDK 15 (JEP 374) and the implementation was subsequently obsoleted/removed (JDK 18, JDK-8256425). Why? Its revocation and bookkeeping costs hurt modern workloads with lots of short-lived threads and contended concurrent data structures.
- Lightweight (thin) locking: uncontended locks use a CAS on the object header's mark word to stack-allocate a lock record on the thread's stack — no OS mutex, so very cheap.
- Heavyweight (inflated) locking: under contention the monitor "inflates" to an OS-level
ObjectMonitorwith a real wait queue and OS parking. Expensive, but needed under pressure.
A senior-level point you must know: as of JDK 24, JEP 491 makes synchronized no longer pin virtual threads. Before 24, when a virtual thread blocked inside a synchronized block, it held its carrier platform thread hostage — this is called pinning — throttling scalability. That's why you were advised to prefer ReentrantLock in virtual-thread-heavy code. From JDK 24 on that advice is obsolete: now choose synchronized vs. java.util.concurrent.locks purely on ergonomics and features, not scalability.
The Lock family: ReentrantLock, ReadWriteLock, StampedLock
The intrinsic monitor is simple and safe, but rigid: you can't say "give up if the lock isn't free within 2 seconds." The java.util.concurrent.locks package gives explicit locks with that flexibility.
ReentrantLock — the workhorse
ReentrantLock provides the same mutual exclusion as synchronized, but with extra capabilities: tryLock() (with an optional timeout — "if it isn't free by then, come back"), interruptible acquisition via lockInterruptibly(), optional fairness (FIFO ordering for waiting threads, at a throughput cost), and multiple Condition objects per lock (vs. one wait-set per monitor).
But it has one non-negotiable idiom — know it in your sleep:
private final ReentrantLock lock = new ReentrantLock();
void doWork() {
lock.lock();
try {
// critical section
} finally {
lock.unlock(); // MUST be in finally — an exception must not leak the lock
}
}
With synchronized the key always returns on its own. With an explicit Lock, if an exception is thrown between lock() and unlock() and you didn't put unlock() in a finally, the lock stays held forever and any other thread that wants it waits forever. Always unlock() in finally.
ReentrantReadWriteLock — separate read/write locks
Picture a whiteboard. Reading means looking at it — ten people can look at once, no problem. Writing means erasing and rewriting — while writing, nobody else should write or even read, or they'd see a half-erased mess. ReentrantReadWriteLock is exactly this: many threads take the read lock at once, but the write lock is exclusive.
This is good when reads massively dominate writes and critical sections are non-trivial (not tiny). But two important traps:
- Writer starvation: if readers pour in without pause, a writer may never get a turn. Mitigate with the
fairconstructor. - Downgrading allowed, upgrading forbidden: downgrading (write → read) is legal — provided you acquire the read lock before releasing the write lock. But upgrading (read → write) deadlocks and is forbidden. (We fully unpack why in the last interview question.)
StampedLock (Java 8) — optimistic reads
StampedLock is not reentrant, but adds a killer feature: optimistic reads.
Suppose you want the time. The pessimistic way is to stand in front of the clock and stop anyone from moving the hands (a read lock). The optimistic way is to just take a quick glance, memorize the number, then check "did anyone touch the clock in the meantime?" If not, your number is valid and you never disturbed anyone. If yes, then you reluctantly go stand in front of the clock. That quick glance is tryOptimisticRead and that check is validate.
private final StampedLock sl = new StampedLock();
private double x, y;
double distanceFromOrigin() {
long stamp = sl.tryOptimisticRead(); // no CAS, no blocking
double cx = x, cy = y; // read snapshot
if (!sl.validate(stamp)) { // a writer may have run
stamp = sl.readLock(); // fall back to a pessimistic read lock
try { cx = x; cy = y; }
finally { sl.unlockRead(stamp); }
}
return Math.sqrt(cx * cx + cy * cy);
}
The big win: when writes are rare, the read path never even touches the lock's cache line, so cache-line contention is eliminated entirely.
Remember three things: (1) it is not reentrant — if the same thread re-locks, it self-deadlocks. (2) It does not support Condition. (3) It is not directly interruptible (use the Interruptibly variants). And the most important rule: in an optimistic read you must copy fields into locals before validate(), and you must not call other methods or dereference a potentially-inconsistent reference mid-read — that reference could be half-built.
AbstractQueuedSynchronizer (AQS): the engine underneath
Now let's pull back the curtain. ReentrantLock, Semaphore, CountDownLatch, ReentrantReadWriteLock, and even ThreadPoolExecutor's worker gate — all are thin wrappers over one shared engine called AQS. Understand this one piece and you practically know the source of half of java.util.concurrent.
Picture a bank with a "now serving" display (state) and an orderly queue of customers who took a number. When the window is free, they call the front of the line. When busy, newcomers take a number and sit down / go to sleep (park) until their turn. When the current customer finishes, they wake exactly the next person (unpark). AQS is that system: one shared status number, plus an orderly queue of sleeping threads.
The precise mental model:
- AQS holds a single
volatile int stateand a CLH-based FIFO wait queue of threads. (CLH is a particular linked-list-based queue where each thread spins/sleeps on its own node.) - A subclass defines what
statemeans and implementstryAcquire(int)/tryRelease(int)(exclusive) ortryAcquireShared/tryReleaseShared(shared). - AQS itself does the hard machining: atomically CAS-ing
state, enqueuing losers as queue nodes, parking them viaLockSupport.park(), and unparking the successor on release.
Now see how the same engine builds several different classes:
ReentrantLock:stateis the hold count — 0 = free, N = held N times reentrantly.tryAcquireCAS-es 0→1, or, if the current thread already owns it, bumps the count.Semaphore:stateis the permit count, using shared acquisition.CountDownLatch:stateis the count, and when it hits 0 it lets all waiters through at once — this is the "shared release," which is why one latch can release many threads simultaneously.
Once you understand "state + CLH queue + tryAcquire," you no longer memorize each synchronizer class separately — you just ask "what does this class make state mean, and what does its tryAcquire do?" and the rest of the behavior (queue, park, unpark) comes from the shared engine.
CAS, the Atomic* classes, and the ABA problem
So far everything was about locks. But there's a layer above them — and often faster: lock-free programming.
Suppose you're editing a shared document. Instead of locking it, you do this: read the current version, prepare your change, and on save say "only apply my change if the doc is still exactly what I read; otherwise report failure." If someone changed the doc in the meantime, your save is rejected and you retry from scratch. That's compare-and-swap (CAS).
CAS is a single hardware instruction (lock cmpxchg on x86, LL/SC on ARM) that atomically does: "if this memory equals expected, set it to new and report success; otherwise leave it and report failure." The java.util.concurrent.atomic package exposes it: AtomicInteger, AtomicLong, AtomicReference, and array/field-updater variants.
AtomicInteger counter = new AtomicInteger();
int incrementAndGet() {
int prev, next;
do {
prev = counter.get();
next = prev + 1;
} while (!counter.compareAndSet(prev, next)); // retry until we win the race
return next;
}
This loop is optimistic: no thread blocks; contention just causes retries.
Under low contention, this loop crushes locks — because nobody ever sleeps and wakes. But under heavy contention a retry storm kicks in: dozens of threads keep beating each other and the CPU burns without doing useful work. That is exactly why Java 8 added LongAdder/LongAccumulator, which stripe the count across multiple separate cells (padded to avoid false sharing) and sum only on demand. For hot counters it's vastly better than a single AtomicLong.
The ABA problem
CAS has a subtle weakness: it checks value equality, not whether the value changed and changed back.
You have a safe locked by a number. Your rule is "if the number is still A, open it." A clever thief changes the number from A to B and back to A. Now when you check "is it still A?", yes it is — so you open it, unaware that everything changed in between. The value is the same, but the world moved.
Precisely: if a thread reads A, another thread flips A→B→A, then the first thread's compareAndSet(A, ...) succeeds — even though the world moved underneath it. For an int counter this is harmless. But for a pointer-swapping structure (say a lock-free stack popping a node that was freed and reallocated in the meantime) it corrupts state. The fix is to pair the value with a monotonic stamp/version: AtomicStampedReference (value + int stamp) or AtomicMarkableReference (value + boolean).
False sharing: the invisible tax
This is one of those performance bugs that leaves no trace in the code and only a profiler finds it.
Caches move memory not variable-by-variable but in 64-byte packages called cache lines. Now imagine two people each have their own cup on a single shared tray. The cups are independent, but because they're on one tray, every time one person moves their cup they must pick up the whole tray, and the other must wait for it to come back. Two independent variables that happen to sit on the same cache line are exactly this: every write to one invalidates the other core's copy, and the line "ping-pongs" between cores even though there's no logical sharing.
This is false sharing, and it can silently cost you an order of magnitude (10x).
The fix: pad hot fields onto their own cache line so each gets its own tray. Java 8+ provides @jdk.internal.vm.annotation.Contended (application code needs the -XX:-RestrictContended flag); LongAdder's internal Cell is annotated exactly this way — which is why LongAdder scales so well. Manual padding (long p1..p7) also works, but the JIT may eliminate unused fields, so @Contended is preferred where available.
Double-checked locking, done right
This is the most famous "looks correct but is broken" pattern in Java concurrency. The goal: build an expensive object only once and lazily, then return it quickly without locking.
The classic broken idiom:
// BROKEN before Java 5 and still broken today without volatile
private Helper helper;
Helper getHelper() {
if (helper == null) { // 1st check (no lock)
synchronized (this) {
if (helper == null) // 2nd check (locked)
helper = new Helper(); // publish
}
}
return helper;
}
Why is it broken without volatile? Because helper = new Helper() is not one atomic operation but three steps: (1) allocate memory, (2) run the constructor and initialize fields, (3) assign the reference to helper. And the JMM permits step 3 (the reference assignment) to become visible before step 2 (the constructor's writes).
Imagine a second thread on the fast "1st check" path. Because of that reordering, this thread can see a non-null helper that still points at a partially constructed object — the reference has landed, but the internal fields still hold default values (0/null). The second thread returns it and uses it, and you have a completely unreproducible bug.
The fix is to make the field volatile. That inserts the release/acquire barrier which guarantees the constructor's writes HB any read of the reference:
private volatile Helper helper; // volatile is mandatory
Helper getHelper() {
Helper result = helper; // read volatile once into a local
if (result == null) {
synchronized (this) {
result = helper;
if (result == null)
helper = result = new Helper();
}
}
return result;
}
Reading into the local result is a real optimization: on the hot path it collapses two volatile reads into one.
If you want a static singleton, skip all this complexity and use the initialization-on-demand holder idiom. It leans on the JLS's own lazy class-initialization guarantees: the inner class isn't loaded until the first access to Holder.INSTANCE, and the JVM itself guarantees this init is safe and lazy — with no volatile or synchronized at all.
class Singleton {
private Singleton() {}
private static class Holder { static final Singleton INSTANCE = new Singleton(); }
static Singleton getInstance() { return Holder.INSTANCE; } // JVM guarantees safe, lazy init
}
Common pitfalls & gotchas
Keep these as a checklist when reviewing concurrent code:
- Relying on time instead of HB. "There's a
Thread.sleep, surely the other thread finished by now."sleepcreates no HB edge; stale reads remain perfectly legal. volatileon a mutable object reference. It publishes the reference safely, but subsequent mutations of that object's fields are not covered. Publish an immutable snapshot instead.- Non-
finalfields in "immutable" objects. Onlyfinalfields get the constructor freeze guarantee; a non-final field can be seen with its default value if the object is published via a data race. - Locking on
thisin a library. Callers can lock on your object too, causing surprise contention or deadlock. Use a private lock. check-then-acton concurrent collections.if (!map.containsKey(k)) map.put(k, v);races; useputIfAbsent/computeIfAbsentinstead.- Forgetting
unlock()isn't automatic. ExplicitLocks needtry/finally; asynchronizedblock cannot leak. - Assuming
size()/isEmpty()on concurrent collections are exact. They're weakly consistent snapshots, not precise instantaneous counts.
Best practices
- Prefer immutability and confinement (keeping data within one thread) over locking; the cheapest lock is the one you never take.
- Reach for higher-level tools before hand-rolling locks:
java.util.concurrentcollections,CompletableFuture,ExecutorService. - Keep critical sections short and free of I/O. Never call foreign/callback code while holding a lock — that's a direct invitation to deadlock.
- Establish and document a global lock ordering and acquire locks in that consistent order everywhere; this prevents deadlock.
- For counters/accumulators under contention, reach for
LongAdder, notAtomicLong. - Deliberately publish every field that crosses threads — via
final,volatile, a lock, or a concurrent collection. If you can't name the HB edge, it's a bug.
Interview Questions
Visibility and ordering via a release (write) / acquire (read) edge: all writes before a volatile write are visible after a subsequent read of that field — plus atomicity of the field itself, including 64-bit long/double. It does not provide mutual exclusion, and it does not make compound operations like x++ atomic.
It's a partial order over actions. If A HB B, the memory effects of A are guaranteed visible to and ordered before B. It's established between a release on a sync variable and a subsequent acquire on the same variable, plus program order, thread start/join, and transitivity. Two actions unordered by HB with a conflicting write form a data race.
boolean running = true; // not volatile
void stop() { running = false; }
void run() { while (running) { /* work */ } }
What can happen? The JIT may hoist running out of the loop into a register (loop hoisting), so run() can loop forever even after stop() returns — there's no HB edge forcing the write to be observed. Making running volatile fixes it.
static int a = 0, b = 0;
// Thread 1: a = 1; int r1 = b;
// Thread 2: b = 1; int r2 = a;
Can r1 == 0 && r2 == 0? Yes. With no synchronization, the loads and stores can be reordered (store buffering), so both threads can read the other's pre-write value. Sequential consistency is not guaranteed for racy code.
helper = new Helper() is allocate + construct + assign, and the JMM lets the reference assignment be reordered before (seen ahead of) the constructor's field writes. A racing reader on the lock-free path can see a non-null reference to a partially initialized object. volatile inserts the release/acquire barrier that makes the constructor's writes HB the reference read.
It optimized the "same thread repeatedly re-locks" case by biasing an object to that thread and skipping CAS entirely. But its revocation and bookkeeping costs grew expensive with modern many-thread workloads and concurrent data structures. JEP 374 disabled it by default in JDK 15; it was later removed (JDK 18). Lightweight and inflated locking remain.
ReentrantLock adds tryLock, timeouts, interruptibility, fairness, and multiple Conditions; synchronized is simpler and can't leak the lock. Historically you preferred ReentrantLock in virtual-thread code because synchronized pinned the carrier — but JEP 491 (JDK 24) removed that pinning, so now choose on features and ergonomics, not scalability.
A volatile int state plus a CLH-based FIFO wait queue. Subclasses define what state means and implement tryAcquire/tryRelease (or the shared variants); AQS does the CAS, enqueue, and park/unpark. ReentrantLock (state = hold count), Semaphore (permits), and CountDownLatch (count, shared release) are all built on it.
CAS compares values, not history. If a value goes A→B→A, a CAS expecting A succeeds even though intermediate state changed. Harmless for a numeric counter; dangerous for pointer/node reuse in lock-free structures where a reused address passes the CAS check but references stale linkage. Fix with AtomicStampedReference (value + version).
Two independent variables sitting on the same 64-byte cache line cause cross-core invalidation on every write, ping-ponging the line between cores. Detect via perf counters (cache-line contention / HITM events) or by profiling; fix by padding hot fields onto separate lines, e.g. @Contended (as LongAdder's Cell does).
When reads dominate and you can use optimistic reads to avoid touching the lock's cache line at all on the read path. Traps: it's not reentrant, has no Condition, and optimistic reads must copy fields to locals and then validate() before using them — reading through a possibly-stale reference mid-read is a bug.
No. Under heavy contention its CAS-retry loop thrashes and burns CPU. LongAdder/LongAccumulator stripe increments across padded cells and sum lazily, scaling far better for write-hot, read-rare counters — at the cost of a slightly more expensive sum() and no atomic read-modify-write across the whole value.
Not guaranteed for a plain non-volatile long/double — the JLS permits a 64-bit write to be split into two 32-bit stores, so a racing reader can see a torn value (half old, half new). volatile, AtomicLong, or a lock removes the tearing.
final fields get a special freeze at constructor end: any thread that reads the object reference is guaranteed to see the correctly-initialized final fields, even under racy publication. A plain (non-final) field carries no such guarantee — another thread can observe its default (0/null) value. This is why "immutable = all fields final" is a correctness rule, not style.
Downgrading (acquire the read lock while holding write, then release write) is safe because you never need to wait for other readers/writers to reach a stronger state. Upgrading (holding read, requesting write) requires all other readers to release — but if two readers both try to upgrade, each waits for the other to drop its read lock: deadlock. So the API forbids it.
Senior notes & advanced edge cases
You now hold the correct mental model and the tools. This section covers the layer that separates a genuine senior from a good programmer: the deeper JMM guarantees, safe publication, the modern access-mode ladder with VarHandle, the real semantics of wait/notify, the taxonomy of progress guarantees, and bugs that only surface under production load and a profiler.
- SC-DRF and out-of-thin-air values. 2) coherence vs consistency. 3) Safe publication and
this-escape. 4) VarHandle and the plain/opaque/acquire-release/volatile ladder plus fences. 5)wait/notify/Conditionand spurious wakeups. 6) The wait-free / lock-free / obstruction-free taxonomy. 7) deadlock/livelock/starvation and JIT lock optimizations. 8) invisible gotchas and testing tools. Then 9 hard interview questions.
1) The SC-DRF theorem: what the JMM actually promises
The chapter showed racy code can see stale or torn values. The reassuring part: the JMM has one central theorem, SC-DRF (Sequential Consistency for Data-Race-Free programs).
If your program is correctly synchronized — no pair of conflicting accesses (at least one a write) unordered by happens-before, i.e. data-race-free — then the JMM guarantees it behaves exactly as if sequentially consistent: a single global order of operations all threads agree on. All those frightening reorderings become invisible as long as you stay DRF.
This simplifies your real job: never reason about individual reorderings, just prove you're DRF — which is exactly what makes the chapter's golden line land: "if you can't name the HB edge, you have a bug." The flip side: for racy code, the JMM still has an unsolved theoretical problem.
The formal JMM has causality rules meant to forbid "out-of-thin-air" values — a value no write produced, conjured from a self-justifying speculation. But the current formulation provably both forbids some desirable executions and fails to fully close OoTA; this is the model's well-known "hole" and the reason for ongoing work on a new JMM. Practical takeaway: never rely on racy code — even values that look "impossible" are permitted by the model.
2) Coherence vs consistency — a subtle correction
The "stale copy on the chef's counter" analogy is great for intuition, but a senior must be more precise. Modern hardware guarantees cache coherence via protocols like MESI: for a single address all cores agree on one global write order, and a write invalidates other cores' copies. So caches don't actually stay stale "forever."
So why do you still see stale values? Two reasons: (1) the store buffer — a core's write first sits in a local buffer and reaches the coherent cache with a delay; in that window your core sees the new value but others don't (the "store buffering" behind the classic r1==0 && r2==0 puzzle). (2) The compiler/JIT caching a value in a register (hoisting). So the problem isn't coherence, it's consistency. volatile and locks emit fences that drain the store buffer and constrain reordering — they don't "refresh the cache."
3) Safe publication and this-escape
The chapter stated "deliberate publication" as the golden rule. More precisely, there are four canonical ways to safely publish an object:
- Initialize it from a static initializer (class-init guarantee).
- Store its reference in a
volatilefield (or anAtomicReference). - Store its reference in a
finalfield assigned in a constructor. - Store its reference guarded by a lock (or place it in a concurrent collection).
The final-field freeze guarantee holds only if the reference does not escape before the constructor completes. If you let this leak from inside the constructor, it breaks:
public class Listener {
private final int id;
public Listener(EventBus bus) {
bus.register(this); // ❌ this escaped — constructor isn't done yet
this.id = computeId(); // another thread may observe id == 0
}
}
Common this-escape patterns: registering a listener/callback, starting a thread inside a constructor, or publishing this into a static variable. The fix: finish the constructor, then in a static create(...) factory build the object and only then register it.
4) VarHandle and the access-mode ladder — the modern memory model
The chapter demonstrated CAS with AtomicInteger. But since Java 9 (JEP 193) the low-level, official machinery underneath it all is VarHandle — the safe, supported successor to sun.misc.Unsafe and the Atomic*FieldUpdater classes (the java.util.concurrent classes themselves migrated onto it in JDK 9).
Its key gift: instead of the "plain or volatile" binary, it offers a four-rung ladder of ordering strength, weakest to strongest.
The VarHandle access-mode strength ladder (weak to strong):
flowchart LR
Plain["Plain<br/>get/set"] --> Opaque["Opaque<br/>getOpaque/setOpaque"]
Opaque --> RelAcq["Release/Acquire<br/>setRelease/getAcquire"]
RelAcq --> Volatile["Volatile (SC)<br/>getVolatile/setVolatile"]
- Plain: no cross-thread ordering or visibility — just atomicity of the access itself (except plain long/double). Like a normal field read.
- Opaque: accesses aren't elided and stay coherent for a single address, but are reorderable relative to other addresses — for a stop flag or a stats counter that only needs to be seen "eventually," cheaper than volatile.
- Release/Acquire: the same release/acquire half-fence from happens-before, but without the heavy StoreLoad fence — the "sign / see the signature" edge without full SC cost.
- Volatile: the strongest — equivalent to the
volatilekeyword, with a global SC order.
class Node {
Object item;
volatile Node next; // ordinary field, manipulated via VarHandle
private static final VarHandle NEXT;
static {
try {
NEXT = MethodHandles.lookup()
.findVarHandle(Node.class, "next", Node.class);
} catch (ReflectiveOperationException e) { throw new ExceptionInInitializerError(e); }
}
boolean casNext(Node expect, Node update) {
return NEXT.compareAndSet(this, expect, update);
}
void publish(Node n) { NEXT.setRelease(this, n); } // release, cheaper than setVolatile
}
Two senior points: (1) compareAndExchange is like compareAndSet but returns the witness value instead of a boolean — in a retry loop it eliminates a redundant get(). (2) weakCompareAndSet may fail spuriously even when the value equals expected; in exchange it's cheaper on LL/SC architectures like ARM. That's why it's only correct inside a do/while loop you write yourself, never standalone.
VarHandle also has four static fence methods (fullFence, acquireFence, releaseFence, loadLoadFence/storeStoreFence) that emit a memory barrier without being tied to a specific field — the modern, safe equivalent of Unsafe.fullFence for very specific hand-rolled publication patterns.
5) wait / notify / Condition — correct semantics
The chapter covered locks and AQS but never touched thread coordination via wait/notify — one of the buggiest spots in real code.
wait() means "hand back the key and sleep in the waiting room," notify() means "call one of the sleepers." The crucial detail: the awoken thread does not run immediately — it must re-contend for the lock, so between "waking" and "running" the world may have changed.
Three iron rules:
- Call
wait/notify/notifyAllonly while you hold that object's monitor, or you get anIllegalMonitorStateException. - Always put
waitinside a loop that re-checks the condition, not anif. Two reasons: spurious wakeups, which the JLS explicitly permits, and that by the time it's your turn the condition may be false again.
synchronized (queue) {
while (queue.isEmpty()) { // ✅ while, not if
queue.wait();
}
return queue.remove();
}
notify() wakes only one thread and doesn't guarantee which. If several threads wait on different conditions on the same monitor, it may wake the "wrong" one whose condition still isn't satisfied; it sleeps again and the signal is lost — the eligible threads never wake (a soft deadlock). Rule: unless all waiters are provably equivalent, choose notifyAll. With Condition (from ReentrantLock) you build separate wait queues (notFull/notEmpty) and needn't wake everyone — safer and more efficient.
Underneath AQS, the parking mechanism is LockSupport.park()/unpark(), not wait/notify. The subtle difference: unpark grants a permit that doesn't accumulate (at most one) but can arrive before park — so it avoids wait/notify's "lost signal." park can also return spuriously, so it too sits inside a condition loop.
6) Progress-guarantee taxonomy: wait-free / lock-free / obstruction-free
The chapter called CAS "lock-free." A senior must know this term precisely because interviewers set traps around it.
- Obstruction-free: a thread running alone finishes in bounded steps. (Weakest.)
- Lock-free: at any moment at least one thread makes progress — the system never stalls, though an unlucky thread may retry forever (starvation).
- Wait-free: every thread finishes in bounded steps — no starvation. (Strongest.)
The do { ... } while(!cas()) loop from the chapter is lock-free, not wait-free: someone always wins, but a particular thread can lose repeatedly and starve. That's why LongAdder is better under heavy contention. And contrary to popular belief, "lock-free" doesn't mean "faster" — under very heavy load a fair lock may give better throughput. Thread.onSpinWait() (Java 9, JEP 285) helps in spin loops: it gives the CPU a PAUSE hint to cut power draw and coherence traffic.
7) Deadlock vs livelock vs starvation, and the JIT's lock optimizations
The chapter mentioned "global lock ordering" to prevent deadlock. Distinguish the three concurrency failures:
- Deadlock: threads wait for each other forever (a wait cycle). The four Coffman conditions must all hold (mutual exclusion, hold-and-wait, no-preemption, circular-wait); breaking one suffices — usually circular-wait, via a global lock order.
- Livelock: threads aren't blocked and keep working but make no progress — like a
tryLockloop that releases and immediately retries on every failure; the fix is randomized backoff. - Starvation: some threads progress but one never gets a turn (writer starvation in ReadWriteLock).
Three HotSpot optimizations the chapter didn't mention: (1) Lock elision — if escape analysis proves an object never escapes one thread, the JIT removes its locking entirely (e.g. a local StringBuffer). (2) Lock coarsening — it merges back-to-back synchronized regions on the same object into one to avoid repeated lock/unlock cost. (3) Adaptive spinning — before parking it spins briefly hoping the lock frees soon, tuning the spin from that lock's history. So a naive micro-benchmark rarely reveals a lock's true cost.
8) Invisible gotchas that bite in production
volatile int[] data; // only the array reference itself is volatile
data[5] = 42; // ❌ this write has no volatile semantics!
A catastrophic mistake: writing to data[i] is a plain write and creates no HB edge. If the elements need volatile ordering/visibility, use AtomicIntegerArray or a VarHandle over the array (MethodHandles.arrayElementVarHandle).
ThreadLocal binds a value to a thread; in a pool, threads are reused and never die. If you don't remove() after each task: (1) one request's data leaks into the next request on the same thread (a security/correctness bug), and (2) a large object stays alive forever (a memory leak). Always remove() in a finally or exit filter.
A documented but under-known principle: the java.util.concurrent classes give "memory consistency" guarantees — e.g. placing an object into a BlockingQueue happens-before its retrieval via take(), and a write into ConcurrentHashMap happens-before a later read of the same key. So passing a mutable object through a concurrent queue needs no separate volatile — the queue itself performs safe publication.
Writing concurrent code without proper testing is self-deception — a bug may never appear on your x86 laptop yet blow up on an ARM server or under load. jcstress (official OpenJDK) builds litmus tests and runs them millions of times to hunt rare reorderings — unmatched for proving a pattern is DRF. And JMH for correct benchmarking that accounts for lock elision/coarsening and JIT warmup (a hand-rolled System.nanoTime() almost always lies).
The SC-DRF theorem: if a program is data-race-free (every conflicting-access pair ordered by happens-before), the JMM guarantees the execution appears sequentially consistent — a single global order all threads agree on, reorderings invisible. So instead of reasoning about individual reorderings, just prove you're DRF. Flip side: for racy code the model doesn't fully close "out-of-thin-air" values and has a known theoretical hole — never rely on racy behavior.
No. Hardware guarantees cache coherence (via MESI): for a single address there's a global write order and a write invalidates other copies. Staleness comes from the store buffer (a not-yet-flushed local write) and the JIT caching a value in a register. So the problem isn't coherence, it's memory consistency (ordering across addresses). volatile/locks emit fences that drain the store buffer and constrain reordering; they don't "refresh the cache."
Four rungs weak to strong: plain (atomic only, no ordering), opaque (coherence and progress for one variable, reorderable relative to others — cheaper than volatile for a stop flag or stats counter), release/acquire (half-fence; the same sign/see-the-signature HB without the heavy StoreLoad), and volatile (strongest, global SC). Rule: pick the weakest mode that keeps your pattern correct; in lock-free structures release-for-publish plus acquire-for-read usually suffices and shaves the full volatile cost.
Two reasons: (1) spurious wakeups, which the JLS explicitly permits — a thread can wake with no notify at all. (2) Between waking and reacquiring the lock, another thread may have invalidated the condition again. So re-check after waking; only while (!condition) wait(); is safe. The same applies to Condition.await() and LockSupport.park().
notify() wakes one unspecified thread. If waiters block on different conditions, it may wake one whose condition doesn't hold; it sleeps again and the signal is lost — so the thread that should have woken never does (a soft deadlock / lost-wakeup). Default to notifyAll unless all waiters are provably equivalent; better still, use separate Conditions (notFull/notEmpty) and signal only the relevant thread.
obstruction-free: a thread with no interference finishes in bounded steps. lock-free: at least one thread always makes progress (the system never stalls) but a particular thread can starve. wait-free: every thread finishes in bounded steps (no starvation). A do{}while(!cas()) loop is lock-free, not wait-free — someone always wins but an unlucky thread can lose repeatedly. Hence LongAdder under heavy contention; and "lock-free" doesn't necessarily mean "faster."
If the reference leaks before the constructor completes (registering a listener, starting a thread, publishing into a static field), the final-field freeze guarantee breaks: another thread can observe uninitialized fields (0/null) because the constructor's writes aren't done. Even single-threaded it's dangerous because callback code may operate on a half-built object. The fix: finish the constructor and move publication into a static factory after construction.
Only the array reference is volatile, not its elements. Writing a[i] is a plain write with no happens-before edge — another thread may see a stale value or none. For volatile semantics on elements use AtomicIntegerArray or MethodHandles.arrayElementVarHandle. Same trap with AtomicReference<mutableObject>: the reference is safely published but later mutations of that object's fields aren't covered.
deadlock: threads blocked in a mutual wait cycle; all four Coffman conditions must hold and breaking one (usually circular-wait, via a global lock order) suffices. livelock: threads aren't blocked and keep working but make no progress — like a tryLock loop that releases and immediately retries on each failure, everyone in lockstep. Fix: randomized backoff. Both are diagnosable via a thread dump and cycle detection, but prevention beats cure.
- SC-DRF: if you're DRF the world looks sequentially consistent; for racy code even OoTA isn't closed. Staleness comes from the store buffer and registers, not incoherent caches (consistency, not coherence).
- Safe publication has four canonical forms and breaks on
this-escape. VarHandle gives the plain/opaque/release-acquire/volatile ladder — pick the weakest that's correct. - Always loop
wait; default tonotifyAllor separateConditions. A CAS loop is lock-free, not wait-free; under load reach forLongAdder. - Separate deadlock/livelock/starvation; the JIT elides/coarsens locks. Gotchas:
volatilearray covers only the reference,ThreadLocalleaks in pools, concurrent collections establish HB for free. Test with jcstress/JMH.
- Memory at the multi-core level is not a shared notebook; the compiler, CPU, and caches reorder your code. The JMM (JLS Ch. 17, reworked by JSR-133 in Java 5) is the formal contract for this world.
- happens-before is the only rule that matters: it's about ordering, not time, and holds only between a release and a subsequent acquire on the same synchronization variable.
volatilegives visibility/ordering and field atomicity (even 64-bit), but not mutual exclusion, and does not makex++atomic.synchronizedgives exclusion + visibility + reentrancy and can't leak the key; biased locking is gone (JEP 374 in JDK 15, removed in JDK 18), and as of JDK 24 with JEP 491 it no longer pins virtual threads.- The
Lockfamily adds flexibility:ReentrantLock(tryLock/timeout/fairness/Condition),ReadWriteLock(downgrade allowed, upgrade forbidden),StampedLock(optimistic reads, but not reentrant). - AQS is the shared engine under all of them:
volatile int state+ CLH queue +tryAcquire. - CAS enables lock-free programming but has the ABA problem (fix with
AtomicStampedReference); under heavy contention reach forLongAdder. - Watch out for false sharing (pad with
@Contended) and DCL without volatile; for singletons use the holder idiom. - The golden rule: deliberately publish every cross-thread field. If you can't name the HB edge, you have a bug.