Java Core · جاوا پایه متوسطIntermediate ~62 دقیقه مطالعه~52 min read
لامبدا، استریم و OptionalLambdas, Streams & Optional
از صفر تا استادی در جعبهابزار تابعی جاوا: یاد میگیری لامبدا واقعاً پشت پرده چه میشود، خطلولهٔ تنبل استریم چطور کار میکند، کالکتورها و استریم موازی کجا میدرخشند و کجا گاز میگیرند، و چطور Optional را درست بهکار ببری — با هر کد، جدول و سؤال مصاحبهٔ اصلی.A zero-to-hero walk through Java's functional toolkit: what a lambda really becomes under the hood, how the lazy Stream pipeline flows, where Collectors and parallel streams shine or bite, and how to use Optional correctly — with every original code sample, table, and interview question taught in full.
جاوا ۸ یکی از بزرگترین تحولهای تاریخ زبان بود: یک لایهٔ کاملاً تابعی (functional) روی زبانی که تا آن روز صددرصد شیءگرا بود سوار شد. اما نکتهٔ ظریف اینجاست که JVM هیچ عوض نشد — هر لامبدا هنوز یک شیء معمولی است و هر استریم با فراخوانی متدهای عادی پیش میرود. کسی که این را میفهمد نهتنها میتواند list.stream().map(...) بنویسد، بلکه میتواند وقتی زیر بارِ تولید کند شد آن را دیباگ کند. در این فصل قدمبهقدم و از پایه همهچیز را میسازیم.
سه ستون داریم که کل این فصل رویشان بنا میشود:
۱. رابطهای تابعی (functional interfaces) — تایپهایی با تنها یک متد انتزاعی که «هدف» لامبداها هستند.
۲. استریمها (streams) — خطلولهٔ تنبل و یکبارمصرفِ عملیات روی یک منبع داده.
۳. Optional — ظرفی که «شاید مقدار نباشد» را در خودِ تایپ صریح میکند و جای null را در مرزهای API میگیرد.
در راه، به کالکتورها، استریمهای موازی، تلههای رایج و در پایان ۱۵ سؤال مصاحبهٔ واقعی میرسیم.
بخش ۰ — چند واژه که باید از همین اول بدانی
قبل از اینکه جلو برویم، چند واژه را که در متن مدام برمیگردند از پایه میسازم تا هیچجا سرگردان نشوی.
یک رستوران را تصور کن که فقط یک غذا سرو میکند: قورمهسبزی. منوی رستوران میتواند توضیحات، ساعت کار و آدرس داشته باشد، اما در نهایت فقط یک «کار اصلی» انجام میدهد: قورمه میپزد. رابط تابعی هم دقیقاً همین است — فقط یک متد انتزاعی دارد که کار اصلیاش است. به این متد میگویند SAM یعنی Single Abstract Method (تنها متد انتزاعی). لامبدا در واقع نسخهٔ کوتاهشدهٔ «بگو آن یک کار را چطور انجام بده».
- desugaring (شکرزدایی): کامپایلر خیلی از نحوهای شیرین و کوتاه را پشت پرده به شکل مفصلتر و ابتداییترشان تبدیل میکند. لامبدا «شکر نحوی» است؛ desugaring یعنی دیدن آن چیزِ مفصلی که واقعاً تولید میشود.
- boxing (باکسینگ): جاوا دو دنیای عدد دارد؛ نوع اولیه مثل
int(سبک، روی پشته) و نوع شیء مثلInteger(سنگین، روی heap). هر بار کهintرا درIntegerمیپیچی، یک شیء تازه ساخته میشود؛ به این پیچیدن میگویند boxing و در حلقههای داغ گران است. - hot path (مسیر داغ): بخشی از کد که میلیونها بار در ثانیه اجرا میشود. هر تخصیص حافظهٔ اضافی اینجا ضرب در میلیون میشود.
- lazy / eager (تنبل / حریص): تنبل یعنی کار را تا آخرین لحظهٔ لازم عقب میاندازد؛ حریص یعنی همینالان انجامش میدهد.
بخش ۱ — مدل ذهنی (mental model)
بیایید کل ماجرا را در یک جمله بگیریم: جاوا ۸ یک لایهٔ تابعی را روی زبانی شیءگرا پیوند زد، بدون آنکه سیستم تایپ JVM را عوض کند. هر لامبدا هنوز یک شیء است که یک رابط را پیادهسازی میکند؛ هر استریم هنوز با متدهای معمولی کار میکند. همین که بدانی desugaring چیست — یعنی آنچه کامپایلر و رانتایم واقعاً انجام میدهند — تفاوتِ بینِ یک برنامهنویسِ معمولی و یک مهندسِ ارشد است.
هر جا لامبدا دیدی، در ذهنت بگو: «این یک شیءِ کوچک است که یک متد دارد.» هر جا استریم دیدی بگو: «این یک دستورِ آشپزی است، نه خودِ غذا — تا وقتی کسی نگوید بپز (عملیات پایانی)، هیچچیز پخته نمیشود.»
بخش ۲ — رابطهای تابعی (functional interfaces)
تشبیه، مفهوم، و بازگشت به جاوا
یک رابط تابعی دقیقاً یک متد انتزاعی دارد (همان SAM که در بخش ۰ ساختیم). حاشیهنویسی @FunctionalInterface اختیاری است اما دو کار میکند: نیت تو را مستند میکند، و اگر اشتباهاً متد انتزاعیِ دومی اضافه کنی کامپایلر داد میزند. توجه کن که متدهای default و static در شمارشِ SAM حساب نمیشوند — چون آنها بدنه دارند و انتزاعی نیستند.
متد انتزاعی یعنی متدی بدون بدنه که کسی باید بعداً پرش کند. default و static خودشان بدنه دارند، پس چیزی برای پر کردن باقی نمیگذارند. رستورانِ قورمه هنوز یک غذای اصلی دارد، حتی اگر منویش پر از توضیحات آماده باشد.
بستهٔ java.util.function مجموعهٔ پایه را آماده به تو میدهد. این جدول را حفظ کن؛ در مصاحبه بارها لازمش داری:
| رابط | متد انتزاعی | شکل |
|---|---|---|
Supplier<T> |
T get() |
() -> T |
Consumer<T> |
void accept(T) |
T -> void |
Function<T,R> |
R apply(T) |
T -> R |
Predicate<T> |
boolean test(T) |
T -> boolean |
UnaryOperator<T> |
T apply(T) |
T -> T |
BiFunction<T,U,R> |
R apply(T,U) |
(T,U) -> R |
BinaryOperator<T> |
T apply(T,T) |
(T,T) -> T |
BiConsumer<T,U> |
void accept(T,U) |
(T,U) -> void |
BiPredicate<T,U> |
boolean test(T,U) |
(T,U) -> boolean |
تصور کن یک شرکت داری: Supplier انباردار است که چیزی میآورد بدون آنکه چیزی از تو بگیرد (() -> T). Consumer سطلِ زباله است که چیزی میگیرد و هیچ برنمیگرداند (T -> void). Function کارگرِ خط تولید است؛ ورودی میگیرد و خروجیِ متفاوت پس میدهد (T -> R). Predicate نگهبانِ در است؛ نگاه میکند و فقط «بله/خیر» میگوید (T -> boolean). پیشوندِ Bi یعنی همان نقش، اما با دو ورودی بهجای یک.
نسخههای تخصصیشده برای انواع اولیه هم داریم: IntFunction, ToIntFunction, IntPredicate, IntUnaryOperator, ObjIntConsumer و غیره. چرا وجود دارند؟ تا از باکسینگ فرار کنند. اگر با Function<Integer,Integer> کار کنی، هر عدد باید در Integer پیچیده شود؛ اما IntUnaryOperator مستقیم با int کار میکند و هیچ شیءِ اضافهای نمیسازد. در مسیرهای داغ حتماً اینها را ترجیح بده.
ترکیب کردن (composition)
زیبایی رابطهای تابعی این است که خودشان متدهای default دارند که به تو اجازه میدهند دو تابعِ کوچک را به یک تابعِ بزرگتر بچسبانی — دقیقاً مثل وصل کردنِ چند لولهٔ آب به هم.
Predicate<String> nonEmpty = s -> !s.isEmpty();
Predicate<String> shortish = s -> s.length() < 10;
Predicate<String> ok = nonEmpty.and(shortish).negate(); // نیازی به دِمورگان دستی نیست
Function<Integer,Integer> plus1 = x -> x + 1;
Function<Integer,Integer> times2 = x -> x * 2;
plus1.andThen(times2).apply(3); // (3+1)*2 = 8
plus1.compose(times2).apply(3); // (3*2)+1 = 7
andThen یعنی «اول من، بعد اون یکی»؛ پس plus1.andThen(times2) اول ۱ اضافه میکند بعد ضرب در ۲ میکند → (3+1)*2 = 8. اما compose برعکس است: «اول اون یکی، بعد من»؛ پس plus1.compose(times2) اول ضرب در ۲ میکند بعد ۱ اضافه میکند → (3*2)+1 = 7. اگر یادت باشد که compose مثل ترکیبِ ریاضیِ f∘g است (اول g بعد f)، هیچوقت اشتباه نمیکنی.
Comparator هم یک رابط تابعی است و یک API روان و غنی دارد:
Comparator<Person> byAgeThenName =
Comparator.comparingInt(Person::age) // تخصصیشده برای int، بدون باکسینگ
.thenComparing(Person::name)
.reversed();
اینجا comparingInt را دیدی؟ همان منطقِ «از باکسینگ فرار کن» است: چون سن یک int است، از نسخهٔ comparingInt استفاده میکنیم تا هر مقایسه یک Integer اضافه نسازد. thenComparing میگوید «اگر سنها برابر بود، با نام تصمیم بگیر» و reversed کل ترتیب را برعکس میکند.
لامبدا در برابر کلاس ناشناس (anonymous class)
ظاهرشان شبیه است اما در سه چیز فرق دارند که دقیقاً همانها در مصاحبهها پرسیده میشوند:
- اتصال
this. در یک کلاس ناشناس،thisبه همان نمونهٔ ناشناس اشاره میکند. در لامبدا،thisبه نمونهٔ دربرگیرنده (enclosing، یعنی کلاسی که لامبدا داخلش نوشته شده) اشاره دارد. لامبداthisجداگانهٔ خودش را ندارد. - دامنه و سایهاندازی (shadowing). لامبدا همان دامنهٔ کدِ اطرافش را به اشتراک میگذارد؛ نمیتوانی داخلش متغیری همنامِ یک متغیرِ محلیِ بیرون اعلام کنی. کلاس ناشناس یک دامنهٔ تازه باز میکند و میتواند نامِ بیرونی را «سایه» بیندازد (یعنی نامِ تازهای با همان اسم بسازد که نامِ بیرونی را میپوشاند).
- کامپایل. کلاس ناشناس یک فایل
.classواقعی تولید میکند (مثلOuter$1.class) و باnewنمونه میشود. اما لامبدا به یک متدِ سنتتیک (synthetic، یعنی متدی که خودِ کامپایلر پنهانی میسازد) خصوصی بهعلاوهٔ یک بوتاسترپِinvokedynamicاز طریقLambdaMetafactoryکامپایل میشود.
کلاس ناشناس مثل این است که از همین حالا برای هر کاری یک کارمندِ دائمی استخدام کنی (فایل .class از پیش ساخته میشود). اما لامبدا مثل قراردادِ «هنگام نیاز» است: JVM تا اولین باری که واقعاً به آن لامبدا نیاز شود، کلاسِ پیادهسازش را نمیسازد. و اگر لامبدا هیچ متغیری از بیرون نگیرد (بدونحالت / non-capturing)، JVM فقط یک نمونه میسازد و بهعنوان singleton کش میکند — چون همهشان یکساناند، چرا چند تا؟
Runnable a = new Runnable() {
public void run() { System.out.println(this.getClass()); } // کلاس ناشناس
};
Runnable l = () -> System.out.println(this.getClass()); // 'this' دربرگیرنده
در خطِ اول this.getClass() نامِ کلاسِ ناشناس (چیزی مثل Outer$1) را چاپ میکند، چون this همان نمونهٔ ناشناس است. اما در لامبدا this.getClass() نامِ کلاسِ دربرگیرنده را چاپ میکند، چون لامبدا this مستقلی ندارد.
گرفتنِ متغیر با قاعدهٔ effectively-final
اینجا یک قانونِ بهظاهر عجیب است که خیلیها فقط حفظش میکنند بدون آنکه بفهمند: لامبدا فقط میتواند متغیرهای محلیای را «بگیرد» (capture) که final یا effectively final باشند. effectively final یعنی متغیری که یکبار مقداردهی شده و دیگر هرگز تخصیصِ مجدد نمیشود، حتی اگر کلمهٔ final را ننوشته باشی.
تصور کن یک محلی روی پشته (stack) مثل یادداشتی است که روی میزِ آشپزخانه گذاشتهای. وقتی متد تمام میشود، میز جمع میشود و یادداشت دور ریخته میشود. اما لامبدا ممکن است بعد از پایانِ متد هم زنده بماند (مثلاً به یک نخِ دیگر داده شده باشد). پس جاوا مقدارِ آن یادداشت را کپی میکند و درونِ خودِ لامبدا نگه میدارد. حالا اگر اجازه میداد متغیر را عوض کنی، دو نسخهٔ متفاوت پیدا میشد: یکی روی میز، یکی درونِ لامبدا — و کدام درست است؟ برای پرهیز از این ابهام، جاوا اصلاً تغییر را ممنوع میکند.
پس این یک قانونِ سبک نیست، بلکه یک تضمینِ مدل حافظه (memory model) است. محلیها روی پشته زندگی میکنند؛ لامبدا ممکن است بیش از متد عمر کند، پس مقدار درونِ نمونهٔ سنتتیک کپی میشود. در مقابل، فیلدها (fields، یعنی متغیرهای عضوِ کلاس) از طریقِ thisِ دربرگیرنده با ارجاع گرفته میشوند، پس میتوانند تغییر کنند — و همین راهی رایج است که مردم حالتِ تغییرپذیر را به یک خطلولهٔ ظاهراً «تابعی» قاچاق میکنند.
int base = 10; // effectively final -> مجاز
IntUnaryOperator f = x -> x + base;
// base = 11; // capture را میشکند: خطای کامپایل
int[] counter = {0}; // درِ فرارِ کلاسیک: ارجاعِ آرایه final است،
list.forEach(x -> counter[0]++); // اما محتوایش تغییر میکند. قانونی، ولی بوی بد.
آن int[] counter = {0} یک حقهٔ معروف است: خودِ ارجاعِ آرایه هرگز عوض نمیشود (پس effectively final است)، اما محتوایش را میتوانی تغییر بدهی. این کد کامپایل میشود، ولی بوی بدی میدهد چون داری حالتِ تغییرپذیر را وارد کدِ تابعی میکنی. لحظهای که کسی .parallel() اضافه کند، همین کد به یک باگِ رِیسِ داده تبدیل میشود.
ارجاع به متد (method references)
وقتی لامبدایت فقط یک متدِ موجود را صدا میزند، جاوا یک راهِ کوتاهتر میدهد: ارجاع به متد. چهار شکل دارد و هرکدام صرفاً «شکر نحوی» برای یک لامبدا هستند:
Function<String,Integer> len = String::length; // ۱. متد نمونه از شیء دلخواه: s -> s.length()
Supplier<List<String>> mk = ArrayList::new; // ۲. سازنده: () -> new ArrayList<>()
Consumer<String> pr = System.out::println; // ۳. متد نمونه از یک شیء مشخص
Function<String,Integer> parse = Integer::parseInt; // ۴. متد استاتیک: s -> Integer.parseInt(s)
نکتهٔ پیچیده شکلِ ۱ است: String::length یک ارجاعِ نامقید (unbound) است. «نامقید» یعنی گیرنده (receiver، یعنی آن شیءی که متد رویش صدا زده میشود) از پیش مشخص نیست. پس اولین پارامترِ تابع تبدیل به گیرنده میشود: s -> s.length(). مقایسهاش کن با myList::add که در آن گیرنده (myList) از پیش ثابت است — به این میگویند مقید (bound).
اگر جای گیرنده یک نامِ تایپ بگذاری (String::length)، نامقید است و گیرنده اولین آرگومان میشود. اگر جای گیرنده یک شیءِ واقعی بگذاری ("hi"::length یا myList::add)، مقید است و گیرنده همان شیءِ ثابت میماند.
بخش ۳ — خطلولهٔ استریم (Stream pipeline)
استریم یک ساختار داده نیست. یک توصیفِ محاسبه روی یک منبع است — مثل دستور آشپزی، نه خودِ غذا. تا وقتی کسی «بپز» را نگوید، هیچچیز اجرا نمیشود.
خطلوله سه بخش دارد:
منبع ──► عملیات میانی (0..n) ──► عملیات پایانی (دقیقاً 1)
List filter/map/sorted/... collect/forEach/reduce/count/...
(تنبل، Stream برمیگرداند) (حریص، اجرا را میآغازد)
سه ویژگیِ کلیدی را باید در خون داشته باشی:
- تنبلی (laziness). عملیاتِ میانی تا وقتی عملیاتِ پایانی اجرا نشود هیچ کاری نمیکنند. آنگاه عناصر یکییکی کشیده میشوند و عمودی از کلِ زنجیره عبور میکنند — نه اینکه هر عملگر یکبار کلِ داده را بپیماید. این همان چیزی است که کوتاهبستن (short-circuit) و ادغام (fusion) را ممکن میکند.
- یکبارمصرف. یک استریم فقط یکبار مصرف میشود. اگر دوباره استفادهاش کنی،
IllegalStateException: stream has already been operated upon or closedپرتاب میشود. - بدون تغییرِ منبع. یک خطلولهٔ خوشرفتار هرگز منبعِ خودش را تغییر نمیدهد.
تصور کن چهار نفر در صف کارخانهاند: فیلترچی، رنگکار، بستهبند. پردازشِ افقی یعنی اول همهٔ محصولات از فیلترچی رد شوند، بعد همه بروند پیش رنگکار. پردازشِ عمودی (کاری که استریم میکند) یعنی یک محصول تا آخرِ خط میرود — فیلتر، رنگ، بستهبندی — بعد نوبتِ محصولِ بعدی میشود. مزیتش این است که اگر فقط یک محصولِ سالم بخواهی (findFirst)، بهمحضِ رسیدن به آن، بقیهٔ خط میایستد و انرژی هدر نمیرود.
نمایشِ تنبلی و کوتاهبستن
List<String> names = List.of("alpha", "beta", "gamma", "delta");
Optional<String> first = names.stream()
.peek(s -> System.out.println("filter " + s))
.filter(s -> s.length() == 5)
.peek(s -> System.out.println("map " + s))
.map(String::toUpperCase)
.findFirst(); // بعد از اولین تطبیق کوتاه میبندد
این کد filter alpha سپس map alpha را چاپ میکند و بعد متوقف میشود — هرگز به beta/gamma/delta دست نمیزند. چرا؟ چون findFirst کوتاهبند است (بهمحضِ اولین نتیجه دست از کار میکشد) و پردازش هم عنصربهعنصر است. alpha طولش ۵ است پس از فیلتر رد میشود، وارد map میشود، و چون اولین نتیجه است کلِ خطلوله همانجا میایستد.
خیلیها فکر میکنند filter اول کلِ لیست را فیلتر میکند و بعد map روی نتیجه اجرا میشود. نه! هر عنصر تکبهتک کلِ زنجیره را طی میکند. همین کوتاهبستن را ممکن میکند؛ اگر افقی بود، findFirst مجبور بود منتظر تمامشدنِ فیلترِ کلِ لیست بماند.
map / filter / reduce / flatMap
// map: تبدیل یکبهیک
List<Integer> lengths = names.stream().map(String::length).toList();
// filter: نگهداشتنِ منطبقها
List<String> fives = names.stream().filter(s -> s.length() == 5).toList();
// reduce: تاکردن به یک مقدار. شکل ۳ آرگومانی: identity, accumulator, combiner
int total = names.stream().reduce(0, (acc, s) -> acc + s.length(), Integer::sum);
// identity ^ accumulator ^^^^^^^^^^^^^^^^^^^^ combiner ^^^^^^^^^^
// combiner نتایج جزئی را در حالت موازی ادغام میکند؛ باید شرکتپذیر (associative)
// و سازگار با accumulator باشد، وگرنه نتایج موازی بیصدا فرق میکنند.
// flatMap: یکبهچند، سپس یک سطح تختکردن
List<List<Integer>> matrix = List.of(List.of(1,2), List.of(3,4));
List<Integer> flat = matrix.stream()
.flatMap(List::stream) // Stream<List<Integer>> -> Stream<Integer>
.toList(); // [1, 2, 3, 4]
بیایید هرکدام را زنده کنیم. map مثل خطِ رنگکاری است: هر عنصر میرود تو، همان تعداد عنصرِ تغییریافته میآید بیرون (یکبهیک). filter مثل نگهبانِ در است: بعضی رد میشوند، بعضی نه. reduce مثل تاکردنِ یک ورقِ کاغذِ بلند است: بارها تا میزنی تا به یک چیزِ کوچک برسی — مجموع، بیشینه، هرچه.
تصور کن چند جعبه داری و درونِ هر جعبه چند توپ است. اگر map بزنی، باز هم چند جعبه داری. اما flatMap درِ هر جعبه را باز میکند و همهٔ توپها را در یک سبدِ واحد میریزد — یک سطح تودرتویی را «تخت» میکند. برای همین Stream<List<Integer>> به Stream<Integer> تبدیل میشود.
شکلِ سهآرگومانیِ reduce جایی است که خیلیها میلغزند. سه جزء دارد: identity (مقدارِ شروع)، accumulator (چطور یک عنصر را در نتیجهٔ جاری تا کنیم)، و combiner (چطور دو نتیجهٔ جزئی را به هم بچسبانیم). identity باید identityِ واقعیِ combiner باشد؛ یعنی combiner.apply(identity, x) == x. اگر این شرط نقض شود، اجرای موازی و ترتیبی با هم اختلاف پیدا میکنند. نکتهٔ ظریفِ دیگر: نوعِ نتیجهٔ accumulator (R) میتواند با نوعِ عنصر (T) فرق کند، و دقیقاً به همین دلیل به combinerِ جداگانه نیاز داریم.
تصور کن میخواهی طولِ کلِ یک کتابِ ۱۰۰۰ صفحهای را بشماری و ۴ دوست داری. کارِ عاقلانه: کتاب را به ۴ بخش تقسیم کن، هر کس بخشِ خودش را بشمارد (این کارِ accumulator است)، و بعد ۴ عدد را با هم جمع کنید (این کارِ combiner است). combiner همان مرحلهٔ «چهار عددِ جداگانه را چطور یکی کنیم» است. بدونِ آن، موازیسازی ممکن نبود. و اگر شمارشِ تو با جمعکردنِ نهایی ناسازگار باشد، جوابِ نهایی غلط درمیآید.
از جاوا ۱۶، متدِ mapMulti یک جایگزینِ ارزانتر برای flatMap است. بهجای اینکه برای هر عنصر یک Stream جدید بسازد (که تخصیص حافظه دارد)، نتایج را مستقیم درونِ یک Consumer (که آن را sink مینامند) میریزد. وقتی هر عنصر فقط به چند عنصر باز میشود (fan-out کوچک)، این ارزانتر است:
Stream.of(1,2,3).<Integer>mapMulti((n, sink) -> { sink.accept(n); sink.accept(-n); });
استریمهای اولیه (primitive streams)
IntStream، LongStream، DoubleStream سه استریمِ ویژهاند که مستقیم با انواعِ اولیه کار میکنند تا از باکسینگ فرار کنند، و در ضمن عملیاتِ پایانیِ عددیِ مفیدی (sum, average, max) اضافه میکنند. با mapToInt / boxed / asLongStream میتوانی بینشان پل بزنی:
int sum = names.stream().mapToInt(String::length).sum();
IntSummaryStatistics stats = IntStream.rangeClosed(1, 100).summaryStatistics();
stats.getAverage(); stats.getMax(); stats.getCount();
double avg = names.stream().mapToInt(String::length).average().orElse(0);
average()، max()، min() مقدارِ OptionalDouble/OptionalInt برمیگردانند، نه یک عددِ خام. دلیلش ساده است: میانگینِ یک استریمِ خالی چیست؟ هیچ! پس بهجای اینکه صفرِ گمراهکننده یا خطا بدهند، یک ظرفِ «شاید خالی» برمیگردانند و تصمیم را به تو میسپارند (.orElse(0)).
بخش ۴ — کالکتورها (Collectors)
collect عملیاتِ پایانیِ عاممنظورهٔ کاهشِ تغییرپذیر (mutable reduction) است. «کاهش» یعنی از خیلی عنصر به یک نتیجه میرسیم؛ «تغییرپذیر» یعنی این کار را با پُر کردنِ یک ظرفِ قابلتغییر (مثل یک List یا Map که کمکم بزرگ میشود) انجام میدهیم. کارخانهٔ Collectors تقریباً همهٔ نیازهایت را آماده دارد:
// گروهبندی: Map<K, List<V>>
Map<Integer, List<String>> byLen =
names.stream().collect(Collectors.groupingBy(String::length));
// گروهبندی با کالکتورِ پاییندستی (downstream): Map<K, aggregate>
Map<Integer, Long> countByLen =
names.stream().collect(Collectors.groupingBy(String::length, Collectors.counting()));
Map<Integer, String> joinedByLen =
names.stream().collect(Collectors.groupingBy(
String::length, Collectors.joining(", ", "[", "]")));
// toMap: مراقبِ کلیدهای تکراری باشید -> بدون تابعِ ادغام، IllegalStateException
Map<Integer, String> byLenFirstWins = names.stream()
.collect(Collectors.toMap(String::length, s -> s, (a, b) -> a)); // ادغام = اولی بماند
// partitioningBy: همیشه Map<Boolean, List<V>> با حضورِ هر دو کلید true و false
Map<Boolean, List<String>> parts =
names.stream().collect(Collectors.partitioningBy(s -> s.length() > 4));
// teeing (جاوا ۱۲): دو کالکتور را اجرا کن، نتایج را ادغام کن — با یک پیمایش
record MinMax(int min, int max) {}
MinMax mm = IntStream.rangeClosed(1, 10).boxed().collect(Collectors.teeing(
Collectors.minBy(Integer::compareTo),
Collectors.maxBy(Integer::compareTo),
(lo, hi) -> new MinMax(lo.orElseThrow(), hi.orElseThrow())));
groupingBy مثل این است که کارگرها را بر اساسِ قد در چند صف بچینی. اما بعدش میخواهی با هر صف چه کنی؟ آن «کالکتورِ پاییندستی» (downstream collector) همان کاری است که درونِ هر صف انجام میشود: بشمار (counting())، به هم بچسبان (joining(...))، یا فقط در لیست بریز (پیشفرض). پس groupingBy(len, counting()) یعنی «بر اساسِ طول گروه کن، و در هر گروه فقط تعداد را نگه دار».
teeing (جاوا ۱۲) هم زیباست: مثل حرفِ T که دو شاخه دارد، استریم را همزمان به دو کالکتور میدهد و بعد نتایجشان را با یک تابعِ سومی به هم میآمیزد — همهٔ اینها با فقط یک بار پیمایشِ استریم. در مثالِ بالا همزمان کمینه و بیشینه را میگیرد و در یک record میریزد.
حالا دو باگِ پرتکرار که در مصاحبه عاشقِ پرسیدنشاناند:
toMap اگر دو عنصر کلیدِ یکسان بسازند، در زمانِ اجرا IllegalStateException پرتاب میکند — مگر یک تابعِ ادغام (merge function) بدهی که بگوید «وقتی دو مقدار سرِ یک کلید دعوا کردند، کدام بماند». مثلاً (a, b) -> a یعنی «اولی بماند». اما groupingBy هرگز این مشکل را ندارد، چون بهصورتِ پیشفرض مقادیرِ همکلید را در یک لیست کنارِ هم میگذارد.
Collectors.toList() هیچ تضمینی دربارهٔ نوع یا تغییرپذیریِ لیستِ بازگشتی نمیدهد. اگر نتیجهٔ تغییرناپذیر میخواهی، از Stream.toList() (جاوا ۱۶+، تغییرناپذیر) یا Collectors.toUnmodifiableList() استفاده کن. اگر نوعِ تغییرپذیرِ مشخصی میخواهی، از Collectors.toCollection(ArrayList::new) استفاده کن.
نکتهٔ مهمی که سوالِ ارشد است: Stream.toList() در برابر Collectors.toList(). اولی یک لیستِ تغییرناپذیر برمیگرداند و عناصرِ null را هم میپذیرد؛ دومی بهطورِ تاریخی ArrayList میداد اما این در قرارداد مشخص نشده. مهاجرتِ کورکورانه میتواند کدی را که بعداً نتیجه را تغییر میدهد بشکند.
بخش ۵ — استریمهای موازی: قدرت و خطر
تصور کن یک ساختمانِ اداری فقط یک آشپزخانهٔ مشترک دارد. اگر یک نفر آنجا برود و ساعتها منتظرِ جوشآمدنِ آب بماند (کارِ مسدودکننده)، بقیهٔ ساختمان گرسنه میمانند. ForkJoinPool.commonPool() دقیقاً همان آشپزخانهٔ مشترکِ کلِ JVM است.
stream.parallel() (یا Collection.parallelStream()) منبع را از طریقِ یک Spliterator (شکنندهٔ استریم به تکهها) میشکند و کار را به ForkJoinPool مشترک میسپارد — همان ForkJoinPool.commonPool() که بهطورِ پیشفرض #cores - 1 نخ دارد. این پرسوءاستفادهترین قابلیتِ کلِ API است.
جایی که موازیسازی کمک میکند:
- N بزرگ (دهها هزار به بالا) و کارِ CPU-محور بهازای هر عنصر.
- منبعی که ارزان و یکنواخت میشکند: آرایه،
ArrayList،IntStream.range. در مقابلLinkedListو بیشترِ منابعِ مبتنیبرIteratorبد میشکنند (چون برای رسیدن به وسطشان باید از اول بپیمایی). - بدونِ قیدِ ترتیب، یا اینکه بتوانی
unordered()را تحمل کنی.
جایی که آسیب میزند یا کاملاً غلط است:
- پولِ مشترک. همهٔ استریمهای موازیِ JVM یک پول را به اشتراک میگذارند. یک کارِ مسدودکننده (I/O، JDBC،
sleep) درونِ استریمِ موازی همهٔ کاربرانِ دیگر — حتی موارد داخلیِ JDK — را قحطی میدهد. هرگز I/O مسدودکننده در استریمِ موازی نکن؛ اگر ناچاری، درForkJoinPoolِ خودت بپیچ و خطلوله را بهعنوان یک task ثبت کن. reduceِ غیرشرکتپذیر نتایجِ نامعین میدهد.- لامبدای دارایحالت / حالتِ تغییرپذیرِ مشترک رِیسِ داده (data race، یعنی دو نخ همزمان روی یک چیز مینویسند) میسازد. این کد خراب است:
List<Integer> out = new ArrayList<>(); // thread-safe نیست
IntStream.range(0, 10_000).parallel()
.forEach(out::add); // رِیس: بهروزرسانیهای گمشده یا استثنا
// اصلاح: .collect(Collectors.toList()) یا .boxed().collect(...) که ذاتاً بدون رِیس است.
forEachدر حالتِ موازی ترتیب را حفظ نمیکند؛ اگر به ترتیبِ برخورد (encounter order، ترتیبی که عناصر واقعاً در منبع بودند) نیاز داری ازforEachOrderedاستفاده کن (با هزینهٔ کارایی).- N کوچک یا کارِ ارزان: سربارِ راهاندازیِ fork/join بر منفعت غلبه میکند؛ حالتِ ترتیبی سریعتر است.
پیشفرض را همیشه ترتیبی بگذار. فقط وقتی سراغِ .parallel() برو که هر سه شرط برقرار باشد: یک بنچمارک (مثل JMH) سرعتش را ثابت کند، منبع شکستپذیر باشد (آرایه/ArrayList)، و صفر حالتِ تغییرپذیرِ مشترک داشته باشی.
ترتیب، حالتداری و اثرات جانبی
- ترتیبِ برخورد ویژگیِ منبع است:
Listآن را دارد (چون ترتیبدار است)،HashSetندارد. عملیاتِsorted/distinct/limitعملیاتِ میانیِ دارایحالت (stateful) هستند — یعنی برای کارشان باید عناصرِ قبلی را بهخاطر بسپارند. اینها ممکن است کلِ استریم را در حافظه بافر کنند، که تنبلی را نقض میکند و روی استریمهای بینهایت میتواند OOM (کمبودِ حافظه) بدهد. - اثراتِ جانبی در
map/filterبوی بد کد است و در حالتِ موازی ناامن.peekفقط برای دیباگ در نظر گرفته شده؛ JDK صریحاً هشدار میدهد که وقتی یک عملیاتِ پاییندستی (مثلcount) بدونِ پیمایش قابلِ محاسبه است، ممکن استpeekبرای هر عنصر اجرا نشود.
long n = Stream.of("a","b","c").peek(System.out::println).count();
// ممکن است هیچ چاپ نکند: از جاوا ۹، count() چون اندازهٔ استریم را مستقیم میداند کوتاه میبندد.
از جاوا ۹، اگر هیچ عملیاتِ اندازهعوضکن (filter/flatMap) پیش از count() نباشد، جاوا میتواند تعداد را بدونِ حتی نگاهکردن به عناصر بگوید. پس peek(System.out::println) هیچچیز چاپ نمیکند. درسِ بزرگتر: peek را فقط برای دیباگِ موقت بهکار ببر، هرگز برای منطقِ برنامه.
بخش ۶ — Optional: استفادهٔ درست و ضدالگوها
تصور کن یک پاکت به تو میدهند و میگویند «شاید داخلش نامه باشد، شاید خالی باشد». همین که پاکت را میبینی، میدانی که باید احتمالِ خالیبودن را در نظر بگیری — دیگر غافلگیر نمیشوی. Optional<T> همان پاکت است: بهجای اینکه بیخبر یک null به تو بدهند و در زمانِ اجرا با NullPointerException غافلگیر شوی، تایپِ متد صریحاً میگوید «مواظب باش، شاید مقدار نباشد».
Optional<T> «مقداری که ممکن است غایب باشد» را بهعنوانِ نوعِ بازگشتی منتقل میکند. این نکته را برجسته کن: برای مقادیرِ بازگشتی طراحی شده، نه برای فیلدها و نه پارامترها.
Optional<User> found = repo.findById(id);
// خوب: fallback / انشعاب را روان بیان کن
String name = found.map(User::name).orElse("anonymous");
found.ifPresentOrElse(u -> log.info("hit {}", u), () -> log.warn("miss")); // جاوا ۹
User u = found.orElseThrow(() -> new NotFoundException(id)); // پرتاب با زمینه
Optional<User> chained = found.or(() -> repo.findInCache(id)); // جاوا ۹، fallback تنبل
ببین چقدر روان است: map(User::name) میگوید «اگر کاربری بود، نامش را بگیر»، و orElse("anonymous") میگوید «وگرنه anonymous». هیچ if و nullی در کار نیست. ifPresentOrElse (جاوا ۹) دو مسیر میدهد؛ orElseThrow با یک پیامِ بامعنا پرتاب میکند؛ و or (جاوا ۹) یک fallbackِ تنبل میدهد که فقط وقتی خالی بود اجرا میشود.
ضدالگوها
// ۱. isPresent()/get() — دوباره null-checking را میسازد، هدف را نقض میکند
if (found.isPresent()) return found.get(); // پرهیز: از map/orElse/orElseThrow استفاده کن
// ۲. orElse با آرگومانِ گران/اثردار — همیشه ارزیابی میشود، حتی وقتی مقدار حاضر است
User u = found.orElse(createExpensiveDefault()); // باگ: default هر بار ساخته میشود
User u2 = found.orElseGet(() -> createExpensiveDefault()); // اصلاح: supplier تنبل
// ۳. فیلد/پارامترِ Optional — یک لفاف اضافه، سریالسازی را میشکند، بیفایده
class Order { private Optional<Coupon> coupon; } // پرهیز
void apply(Optional<Coupon> c) { } // پرهیز: overload یا پذیرشِ null
// ۴. Optional.get() بدون بررسی — NoSuchElementException پرتاب میکند
found.get(); // پرهیز مگر تازه isPresent را چک کرده باشی
// ۵. پیچیدن و باز کردنِ بیمورد
return Optional.ofNullable(x).orElse(y); // فقط: return x != null ? x : y
به ضدالگوی شمارهٔ ۲ خوب دقت کن، چون در مصاحبه کلاسیک است. orElse(v) یک مقدارِ ازپیشمحاسبهشده میگیرد — پس آرگومانش همیشه ساخته میشود، حتی وقتی پاکت پُر است و اصلاً به default نیازی نیست! اگر createExpensiveDefault() گران است یا اثرِ جانبی دارد (مثلاً در دیتابیس مینویسد)، این یک باگِ واقعی است. orElseGet(() -> ...) تنبل است و supplier را فقط وقتی پاکت خالی است اجرا میکند.
نکات باقیمانده را هم روشن کنیم. Optional.of(x) اگر x نال باشد فوراً NPE میدهد — از آن بهعنوانِ یک اظهار (assertion) استفاده کن که میگوید «مطمئنم این نال نیست». Optional.ofNullable(x) نرمتر است و نال را تحمل میکند (اگر نال بود، empty میشود). و Optional.stream() (جاوا ۹) یک Optional را به استریمِ صفر-یا-یک عنصری تبدیل میکند — ابزاری عالی برای flat-map کردن و دور ریختنِ خالیها:
List<User> users = ids.stream()
.map(repo::findById) // Stream<Optional<User>>
.flatMap(Optional::stream) // خالیها را میاندازد، حاضرها را باز میکند
.toList();
اگر متدت قرار است List یا Map برگرداند، هرگز Optional<List> نده — بهجایش یک List/Mapِ خالی برگردان. یک مجموعهٔ خالی خودش دقیقاً یعنی «هیچچیز»، پس لایهٔ Optional زائد است و فقط کارِ فراخوان را سختتر میکند.
بخش ۷ — تلهها و بهترینروشها
اینها را مثل چکلیستِ نهایی نگه دار:
- برای نتایجِ فقطخواندنی
Stream.toList()(جاوا ۱۶+) را برcollect(toList())ترجیح بده؛ اما اول تفاوتِ تغییرپذیری را بدان. - لامبداها را کوتاه و خالص (pure، یعنی بدونِ اثرِ جانبی) نگه دار؛ وقتی منطق بزرگ شد یا در stack trace به نام نیاز داشت، آن را به یک متدِ نامدار استخراج کن (و ارجاع به متد بگذار) — فریمهای لامبدا در stack trace بهشکلِ زشتِ
lambda$method$0نمایش داده میشوند. - هرگز حالتِ مشترک را از درونِ استریم تغییر نده، حتی بهصورتِ ترتیبی — چون لحظهای که کسی
.parallel()اضافه کند میشکند. - در مسیرهای عددیِ داغ از استریمهای اولیه استفاده کن تا باکسینگ را حذف کنی.
groupingBy+ کالکتورِ پاییندستی همیشه بهتر از «جمع در لیست و بعد دوباره استریمکردن» است.- مراقبِ استریمهای بینهایت (
Stream.iterate،generate) با عملیاتِ دارایحالتِ مثلsorted/distinctباش — هرگز پایان نمییابند. اولlimit/takeWhileبگذار. takeWhile/dropWhile(جاوا ۹) روی استریمِ مرتبگونه کوتاه میبندند؛filterنه. (takeWhileتا اولین شکستِ شرط عناصر را برمیدارد و بعد میایستد؛filterکلِ استریم را میپیماید.)
بخش ۸ — سؤالات مصاحبه
هر کدام را جدی بخوان؛ اینها همان جاهاییاند که مصاحبهکننده سطحِ ارشد را از میانی جدا میکند.
محلیها روی پشته زندگی میکنند و بهصورتِ مقداری درونِ نمونهٔ سنتتیکِ لامبدا کپی میشوند؛ لامبدا میتواند بیش از فریمِ متد عمر کند، پس اجازهٔ تخصیصِ مجدد دو نسخهٔ ناسازگار میساخت (یکی روی پشته، یکی درونِ لامبدا). فیلدها اما از طریقِ thisِ دربرگیرندهٔ گرفتهشده در دسترساند، پس تغییرات از طریقِ همان ارجاعِ مشترک دیده میشوند — کامپایلر فقط به پایداریِ خودِ ارجاع نیاز دارد، نه به ثابتبودنِ محتوای فیلد.
orElse(v) یک مقدارِ ازپیشمحاسبهشده میگیرد — آرگومانش همیشه ارزیابی میشود، حتی وقتی Optional حاضر است. orElseGet(supplier) تنبل است: supplier فقط وقتی خالی است اجرا میشود. پاسدادنِ defaultِ گران یا اثردار به orElse یک باگِ واقعیِ کارایی/صحت است.
نه. استریمها عنصربهعنصر و عمودی از کلِ خطلوله پردازش میکنند. هر عنصر از filter سپس map عبور میکند پیش از آنکه عنصرِ بعدی شروع شود. همین چیز است که کوتاهبستن (findFirst، anyMatch، limit) را بدونِ پردازشِ بقیه ممکن میکند.
long n = Stream.of("a","b","c").peek(System.out::println).count();
System.out.println(n);
احتمالاً فقط 3. از جاوا ۹، count() میتواند اندازه را بدونِ پیمایش تعیین کند وقتی هیچ عملیاتِ اندازهعوضکن (filter/flatMap) پیش از آن نباشد، پس peek ممکن است هرگز شلیک نکند. تکیه بر peek برای هر چیزی جز دیباگ ناامن است.
Map<String,Integer> m = words.stream()
.collect(Collectors.toMap(w -> w.substring(0,1), String::length));
toMap روی کلیدهای تکراری (دو واژه با حرفِ اولِ یکسان) IllegalStateException پرتاب میکند. یک تابعِ ادغام اضافه کن: Collectors.toMap(k, v, (a,b) -> a) یا از groupingBy استفاده کن.
کندتر: N کوچک، کارِ ارزان بهازای هر عنصر، منابعِ بدشکن (LinkedList)، یا I/O مسدودکننده (که پولِ مشترکِ ForkJoinPool را قحطی میدهد). غلط: حالتِ تغییرپذیرِ مشترک (رِیسِ داده)، reduceِ غیرشرکتپذیر، یا تکیه بر ترتیبِ برخورد با forEachِ ساده.
Stream.toList() (جاوا ۱۶+) یک لیستِ تغییرناپذیر با قراردادِ مشخص برمیگرداند و نال را میپذیرد. Collectors.toList() یک ArrayListِ نامشخص و معمولاً تغییرپذیر برمیگرداند — نباید بر نوع یا تغییرپذیریِ آن تکیه کنی. جایگزینیِ یکی با دیگری میتواند کدی را که نتیجه را تغییر میدهد یا برعکس انتظارِ تغییرناپذیری دارد بشکند.
reduce(identity, accumulator, combiner): accumulator: (R,T)->R یک عنصر را در نتیجهٔ جزئیِ نوعِ متفاوتِ R تا میکند؛ combiner: (R,R)->R دو نتیجهٔ جزئی را ادغام میکند. combiner برای این هست که اجرای موازی بتواند زیربازهها را مستقل تا کند و بعد ادغام کند. identity باید combiner(identity, x) == x را برآورده کند و accumulator باید شرکتپذیر/سازگار با combiner باشد، وگرنه نتایجِ موازی و ترتیبی واگرا میشوند.
یک تخصیص و یک لایهٔ لفاف بدونِ هیچ سودِ خوانایی اضافه میکند، فریمورکهای رایجِ سریالسازی را میشکند (Optional سریالپذیر / Serializable نیست)، و فراخوانها را وادار به ساختِ لفاف میکند. برای پارامترها overload یا آرگومانِ nullable را ترجیح بده؛ برای فیلدها مقدارِ خام (احتمالاً نال) را ذخیره کن و از getter مقدارِ Optional برگردان.
String::length نامقید است: گیرندهٔ ثابتی ندارد، پس اولین پارامترِ نوعِ تابعیِ هدف به گیرنده تبدیل میشود — به s -> s.length() باز میشود که با Function<String,Integer> میخوانَد. در مقابلِ "hi"::length (مقید: گیرندهٔ ثابت، Supplier<Integer>).
IllegalStateException: stream has already been operated upon or closed. استریمها یکبارمصرفاند؛ منبع را به یک متغیر تخصیص بده و هر بار استریمِ تازه بساز، یا کد را به یک خطلولهٔ واحد بازساختار بده.
آنها خطلوله را به جزئیاتِ اجرایی گره میزنند که رانتایم آزاد است بهینهشان کند و حذفشان کند (مثلِ رد کردنِ پیمایش توسطِ count) یا بازچینی/موازیشان کند. اثراتِ جانبی همچنین کد را در لحظهای که کسی .parallel() اضافه میکند ناایمن میکنند و یک رِیسِ خاموش را به ازدسترفتنِ دادهٔ تولید تبدیل میکنند.
partitioningBy همیشه مپی با هر دو کلیدِ true و false برمیگرداند، حتی وقتی یک بخش خالی است. groupingBy روی کلیدِ بولی، کلیدهایی را که عنصر ندارند حذف میکند. کدِ پاییندستیای که فرض میکند هر دو کلید وجود دارند، با groupingBy دچارِ NPE میشود.
عملیاتِ پایانی را بهعنوانِ یک task به ForkJoinPoolِ خودت ثبت کن: myPool.submit(() -> stream.parallel().reduce(...)).get(). استریمِ موازی پولِ نخِ ثبتکننده را به ارث میبرد. این کارِ مسدودکننده یا طولانی را از پولِ مشترک جدا میکند و بقیهٔ JVM را قحطی نمیدهد.
Optional.of(null) بلافاصله NullPointerException پرتاب میکند — از آن برای اظهارِ نانالبودن استفاده کن. Optional.ofNullable(null) مقدارِ Optional.empty() برمیگرداند. انتخابِ اشتباه یا باگی را پنهان میکند یا بهطورِ غیرمنتظره پرتاب میکند.
نکاتِ سنیور و موارد پیشرفته
تا اینجا کل جعبهابزار تابعی جاوا را ساختیم. اما چیزی که یک سنیور واقعی را از یک برنامهنویسِ خوب جدا میکند، دانستنِ همان لبههای تیزی است که در دمو هیچوقت دیده نمیشوند و فقط ساعت سه صبح، وقتی صفحهٔ pager روشن میشود، خودشان را نشان میدهند. این بخش دقیقاً همان لبههاست: exception در لامبدا، ساختِ Collector سفارشی، مشخصههای Spliterator، تلههای استریمِ موازی در دنیای virtual thread، و بهروزرسانیهای جاوای مدرن (Gatherers).
۱. checked exception داخل لامبدا — دردِ روزمرهٔ هر تیمی که واقعاً کد مینویسد.
۲. آناتومی خودِ Collector — چهار تابع و مشخصهها، و کالکتورهای concurrent.
۳. مشخصههای Spliterator — چرا بهینهساز گاهی کل کار را رد میکند.
۴. جدولِ درستیِ short-circuit — دامِ «صدقِ پوچ» و findAny در برابر findFirst.
۵. Gatherers (جاوا ۲۴) — بالاخره میشود عملیاتِ میانیِ سفارشی نوشت.
۶. نشت حافظه از لامبدا و لامبدای Serializable.
۷. استریمِ موازی در برابر virtual thread برای کارِ I/O.
۸. سؤالات مصاحبهٔ سختِ سطحِ سنیور.
۱) checked exception داخل لامبدا — واقعیتِ تلخِ روزمره
این را همین اول بگویم چون بیشترین وقتِ تیمها را میخورد: هیچکدام از رابطهای java.util.function هیچ checked exceptionی declare نمیکنند. یعنی لحظهای که داخل map بخواهی متدی صدا بزنی که throws IOException دارد، کد کامپایل نمیشود.
List<String> paths = List.of("a.txt", "b.txt");
paths.stream()
.map(p -> Files.readString(Path.of(p))) // خطای کامپایل: unhandled IOException
.toList();
Function<T,R> یک قراردادِ عمومی است که قرار است در هزاران جای مختلف — از جمله استریمِ موازی روی چند thread — کار کند. اگر اجازه میدادند هر لامبدا هر checked exceptionی پرتاب کند، آنوقت موتورِ استریم باید میدانست آن exception را کجا و روی کدام thread تحویل بدهد. طراحها ترجیح دادند این پیچیدگی اصلاً وارد نشود؛ به همین خاطر امضای این رابطها «تمیز» است و فقط unchecked exception عبور میکند.
سه راهِ عملی داری و هر سه در کدِ واقعی دیده میشوند:
// راهِ ۱: try/catch داخلِ لامبدا و تبدیل به unchecked (خواناترین برای منطقِ ساده)
.map(p -> {
try { return Files.readString(Path.of(p)); }
catch (IOException e) { throw new UncheckedIOException(e); } // نوعِ آمادهٔ JDK
})
// راهِ ۲: یک رابطِ تابعیِ throwing و یک آداپتور که آن را wrap میکند
@FunctionalInterface interface ThrowingFn<T,R> { R apply(T t) throws Exception; }
static <T,R> Function<T,R> unchecked(ThrowingFn<T,R> f) {
return t -> { try { return f.apply(t); }
catch (Exception e) { throw new RuntimeException(e); } };
}
// استفاده: .map(unchecked(p -> Files.readString(Path.of(p))))
یک ترفندِ معروف با generic این است که checked exception را «قاچاقی» و بدون wrap کردن پرتاب کنی (به آن sneaky-throw میگویند). کامپایل میشود و امضاها تمیز میمانند، اما یک هیولا میسازی: caller حالا یک IOException میگیرد که در امضای هیچ متدی ننوشته، پس نمیتواند درست catchش کند و ابزارها هم دربارهاش هشدار نمیدهند. در کدِ کتابخانهای هرگز این کار را نکن؛ همیشه با UncheckedIOException/CompletionException تمیز wrap کن تا stack و نوعِ خطا صادق بماند.
اگر لامبدا آنقدر بزرگ شده که سهخط try/catch لازم دارد، همانجا علامتِ این است که باید به یک متدِ نامدار استخراجش کنی و با method reference صدایش بزنی. هم stack trace تمیزتر میشود (بهجای lambda$process$3 نامِ واقعیِ متد را میبینی)، هم تستپذیر میشود.
۲) آناتومیِ خودِ Collector — پشتِ groupingBy چه میگذرد
فصل نشان داد چطور از Collectors آماده استفاده کنی. سنیور باید بتواند خودش یکی بسازد، چون هر Collector در واقع چهار قطعه است:
public interface Collector<T, A, R> {
Supplier<A> supplier(); // ظرفِ خالیِ تازه بساز (accumulator container)
BiConsumer<A, T> accumulator(); // یک عنصر را داخلِ ظرف بریز
BinaryOperator<A> combiner(); // دو ظرفِ نیمهپُر را (در حالت موازی) ادغام کن
Function<A, R> finisher(); // ظرفِ داخلی A را به نتیجهٔ نهاییِ R تبدیل کن
Set<Characteristics> characteristics();
}
سه حرفِ generic معنی دارند: T نوعِ عنصرِ ورودی، A نوعِ ظرفِ میانیِ قابلتغییر (accumulation)، و R نوعِ نتیجهٔ نهایی. مثلاً در joining، A یک StringBuilder است اما R یک String؛ اینجا finisher همان .toString() است.
سه برچسب داری. IDENTITY_FINISH یعنی «finisher کاری نمیکند، همان ظرفِ A خودش نتیجه است» — پس موتور میتواند این مرحله را کاملاً حذف کند. UNORDERED یعنی «ترتیبِ عناصر برایم مهم نیست» — به موتور اجازه میدهد در حالتِ موازی سربارِ حفظِ ترتیب را کنار بگذارد. CONCURRENT یعنی «accumulator من thread-safe است، پس چند thread میتوانند در یک ظرفِ مشترک بریزند و اصلاً به combiner نیازی نیست».
اینجا همانجایی است که سنیورها در بحثِ کارایی جدا میشوند:
Collectors.groupingBy(...) مشخصهٔ CONCURRENT ندارد. یعنی در استریمِ موازی، هر thread یک HashMap جداگانه میسازد و بعد همه با combiner ادغام میشوند — ادغامِ مپها گران است. اگر واقعاً موازیسازی میخواهی، از groupingByConcurrent (یا toConcurrentMap) استفاده کن که یک ConcurrentHashMap مشترک را مستقیم پُر میکند و مرحلهٔ ادغام حذف میشود — اما در عوض ترتیبِ encounter را از دست میدهی. این trade-off را باید صریح انتخاب کنی، نه تصادفی.
کالکتورهای ترکیبیِ کمتر شناختهشده که سنیور باید بلد باشد:
// mapping: قبل از downstream، هر عنصر را نگاشت کن
Map<Integer,List<Character>> firstChars = words.stream().collect(
groupingBy(String::length, mapping(w -> w.charAt(0), toList())));
// filtering (جاوا ۹): بعد از گروهبندی فیلتر کن — کلیدِ خالی حفظ میشود (فرقِ مهم با filter قبل از groupingBy)
Map<Dept,List<Emp>> seniorsByDept = emps.stream().collect(
groupingBy(Emp::dept, filtering(e -> e.level() > 5, toList())));
// collectingAndThen: نتیجهٔ نهایی را یک مرحله بیشتر ببر (مثلاً immutable کن)
List<String> frozen = names.stream().collect(
collectingAndThen(toList(), List::copyOf));
// reducing/summingInt بهعنوان downstream برای aggregate درونِ هر گروه
Map<Dept,Integer> payroll = emps.stream().collect(
groupingBy(Emp::dept, summingInt(Emp::salary)));
اگر قبل از groupingBy بنویسی .filter(...)، گروههایی که هیچ عضوی از فیلتر رد نکردند اصلاً در مپ ظاهر نمیشوند. اما اگر از filtering(...) بهعنوان downstream استفاده کنی، آن کلیدها با یک لیستِ خالی باقی میمانند. اگر کدِ بعدی انتظار دارد همهٔ departmentها بهعنوان کلید موجود باشند، این فرق یعنی NPE یا نبودِ خطا.
نمودارِ بیلینگ: چهار قطعهٔ یک Collector و مسیرِ داده — Four moving parts of a Collector:
flowchart LR
Src[Stream elements T] --> Acc
Sup[supplier: new empty A] --> Acc[accumulator: A x T -> A]
Acc --> Comb[combiner: A x A -> A parallel merge]
Comb --> Fin[finisher: A -> R]
Fin --> Out[Result R]
۳) مشخصههای Spliterator — چرا موتور گاهی کارِ تو را رد میکند
فصل گفت که count() گاهی بدون پیمایش جواب میدهد و peek اجرا نمیشود. حالا چرای آن: منبعِ هر استریم یک Spliterator دارد که چند برچسبِ فراداده (metadata) با خودش حمل میکند: SIZED (اندازهام را دقیق میدانم)، SUBSIZED، ORDERED، SORTED، DISTINCT، NONNULL، IMMUTABLE.
این برچسبها به بهینهساز اجازهٔ میانبر میدهند:
count()روی منبعِSIZEDکه هیچ عملیاتِ اندازهعوضکن (filter/flatMap) نداشته باشد، فقط اندازه را میخواند — به همین خاطرpeekهرگز اجرا نمیشود.distinct()روی استریمی که از قبلDISTINCTاست (مثلاً ازTreeSetآمده) عملاً no-op است.sorted()روی منبعِSORTEDبا همان comparator رد میشود.
اگر ترتیبِ خروجی برایت مهم نیست، صریح .unordered() بزن. روی distinct/limit/skip در حالتِ موازی، حفظِ encounter order گران است چون threadها باید هماهنگ بمانند. با برداشتنِ آن قید، بهینهساز میتواند از هر thread هر عنصری را اول بردارد. بسیاری تیمها ماهها با یک distinct().limit(n)ِ کُند سر میکنند بیآنکه بدانند فقط یک .unordered() قبلش لازم بود.
۴) جدولِ درستیِ short-circuit — صدقِ پوچ و findAny
سه متدِ تطبیق (anyMatch/allMatch/noneMatch) روی استریمِ خالی رفتاری دارند که در مصاحبه دام است:
| روی استریمِ خالی | نتیجه |
|---|---|
anyMatch(p) |
false |
allMatch(p) |
true ← صدقِ پوچ (vacuous truth) |
noneMatch(p) |
true |
این ریاضی است نه باگِ جاوا: «همهٔ عناصر شرط را دارند» وقتی هیچ عنصری نیست، بهطور پوچ صادق است. کدِ واقعی که این را فراموش میکند: orders.stream().allMatch(Order::isPaid) روی لیستِ خالیِ سفارشها true برمیگرداند و ممکن است سهواً یک flowِ «همه پرداخت شده» را باز کند. همیشه اول isEmpty را جدا هندل کن.
findFirst همیشه اولین عنصر به ترتیبِ encounter را میخواهد؛ در حالتِ موازی این یعنی هماهنگی و سربار. findAny میگوید «هر عنصری که زودتر پیدا شد کافی است» — در موازی خیلی ارزانتر است. اگر منطقاً فرقی نمیکند کدام عنصر برگردد (مثلاً فقط وجودِ یکی مهم است)، findAny را انتخاب کن.
۵) Gatherers — عملیاتِ میانیِ سفارشی (جاوا ۲۴)
بزرگترین بهروزرسانیِ استریم از جاوا ۸ تا امروز: تا قبل از این، تو فقط میتوانستی از عملیاتِ میانیِ آماده (map/filter/...) استفاده کنی و نمیتوانستی عملیاتِ میانیِ حالتدارِ (stateful) خودت را بنویسی. مثلاً «هر عنصر را با مجموعِ در حال اجرا نگاشت کن» یا «عناصر را در پنجرههای سهتایی گروه کن» با API قدیمی یا ناممکن بود یا زشت. Stream::gather این را حل میکند: دقیقاً قرینهٔ collect است اما بهجای نتیجهٔ نهایی، دوباره یک Stream میدهد.
Stream Gatherers در JEP 461 (جاوا ۲۲) و JEP 473 (جاوا ۲۳) بهصورت preview آمد و در JEP 485 در جاوا ۲۴ نهایی (final) شد. در جاوا ۲۵ (LTS) بهصورت پایدار موجود است. پنج gathererِ آماده در java.util.stream.Gatherers هست: fold, scan, windowFixed, windowSliding, mapConcurrent.
// running total (اسکن) — با API قدیمی عملاً ناممکن بود بیآنکه به آرایهٔ بیرونی دست بزنی
List<Integer> runningSums = Stream.of(1, 2, 3, 4)
.gather(Gatherers.scan(() -> 0, Integer::sum))
.toList(); // [1, 3, 6, 10]
// پنجرهٔ لغزان سهتایی
List<List<Integer>> windows = Stream.of(1, 2, 3, 4, 5)
.gather(Gatherers.windowSliding(3))
.toList(); // [[1,2,3],[2,3,4],[3,4,5]]
// mapConcurrent: هر عنصر را با یک virtual thread و سقفِ همزمانیِ مشخص نگاشت کن
List<String> bodies = urls.stream()
.gather(Gatherers.mapConcurrent(10, this::httpGet)) // حداکثر ۱۰ درخواستِ همزمان
.toList();
تا قبل از این، هر منطقِ حالتداری که در یک عملیاتِ استریم لازم داشتی (running max، dedupe با یک کلید، batch کردن) تو را مجبور میکرد یا از استریم بیرون بیایی یا با int[]/متغیرِ بیرونی تقلب کنی — همان بویی که در بخشِ effectively-final دیدیم. gather این منطق را به یک واحدِ تمیز، composable و امنِ موازی تبدیل میکند.
۶) نشتِ حافظه از لامبدا و لامبدای Serializable
وقتی لامبدا یک فیلد یا متدِ نمونه را capture میکند، در واقع کلِ this را capture میکند (فیلدها through the enclosing this گرفته میشوند — همان نکتهٔ فصل). حالا اگر آن لامبدا در یک ساختارِ عمرِطولانی ذخیره شود — یک listener، یک کشِ static، یک زنجیرهٔ CompletableFuture که تمام نمیشود — آنوقت کلِ شیءِ میزبان (شاید یک کنترلرِ سنگین با کلی state) هرگز garbage نمیشود. این یک منبعِ کلاسیکِ نشتِ حافظه در UI و در event busهاست.
class HeavyController {
private final byte[] cache = new byte[50_000_000];
void register(EventBus bus) {
bus.subscribe(e -> handle(e)); // handle متدِ نمونه است → this و در نتیجه cache پین میشود
}
}
راهِ حل: لامبدا را طوری بنویس که فقط چیزهای لازم را capture کند (یک متغیرِ محلی از فیلد بردار)، یا صریح unsubscribe کن.
یک لامبدا بهطور پیشفرض Serializable نیست. اگر با intersection cast مجبورش کنی ((Runnable & Serializable) () -> ...)، جاوا مکانیزمِ سنگینِ SerializedLambda و متدِ $deserializeLambda$ را فعال میکند و سریالسازیاش شکننده است (به نامِ متدِ synthetic وابسته میشود که بینِ کامپایلها میتواند فرق کند). در سیستمهای توزیعشده مثل Spark این را میبینی؛ بدانِ ضرورت هرگز لامبدا را Serializable نکن.
۷) استریمِ موازی در برابر virtual thread — تلهٔ مدرن
فصل گفت داخلِ استریمِ موازی I/O بلاککننده نگذار چون commonPoolِ مشترک را گرسنه میکند. حالا در جاوای مدرن راهِ درست چیست؟
parallelStream() برای کارِ CPU-bound روی دادههای در حافظه است و همچنان روی ForkJoinPool.commonPool() میچرخد — که virtual thread نیست و برای بلاکشدن ساخته نشده. برای کارِ I/O-bound (صدا زدنِ ۵۰۰ سرویسِ HTTP)، ابزارِ درست virtual thread (JDK 21) است: یا Gatherers.mapConcurrent(n, ...) که مستقیماً روی virtual thread اجرا میکند، یا structured concurrency (استاندارد در JDK 25). اینها هزاران کارِ بلاکشونده را ارزان اداره میکنند بیآنکه هیچ pool مشترکی را گرسنه کنند.
// اشتباهِ رایج: I/O روی commonPool مشترک → کلِ JVM را کند میکند
urls.parallelStream().map(this::httpGet).toList();
// درستِ مدرن (JDK 21+): virtual thread، هزار درخواستِ بلاکشونده بدونِ گرسنگیِ pool
urls.stream().gather(Gatherers.mapConcurrent(50, this::httpGet)).toList();
۸) دو گازگرفتنِ کوچکِ دیگر که خون به پا میکنند
Collectors.groupingBy اگر تابعِ classifier مقدارِ null برگرداند NullPointerException میدهد (چون HashMap اجازهٔ کلیدِ null دارد ولی خودِ groupingBy در Objects.requireNonNull میگیردش). این در دادهٔ واقعیِ کثیف مدام پیش میآید: groupingBy(User::country) وقتی بعضی کاربرها country ندارند. اول null را به یک مقدارِ سنتینل (مثلِ "UNKNOWN") نگاشت کن.
opt.map(f) اگر f مقدارِ null تولید کند، بهجای ترکیدن، بیسروصدا Optional.empty() میدهد. این خوب بهنظر میرسد اما یک باگِ خاموش است: تو فکر میکنی مقدار «نبود»، در حالی که واقعاً «بود ولی نگاشتش null شد». اگر f خودش یک Optional برمیگرداند، map به تو Optional<Optional<X>> میدهد؛ آنجا باید از flatMap استفاده کنی تا یک لایه صاف شود.
چون Function<T,R> هیچ checked exceptionی declare نمیکند و لامبدا نمیتواند بیش از آنچه SAMِ هدف اجازه میدهد پرتاب کند. علتِ طراحی این است که رابطهای تابعی عمومیاند و باید در محیطهای موازی/ناهمگام هم کار کنند، جایی که مسیرِ پرتابِ checked exception مبهم میشود. در عمل: یا داخلِ لامبدا try/catch و wrap به unchecked (UncheckedIOException)، یا یک آداپتورِ unchecked(ThrowingFn)، یا استخراج به متدِ نامدار. هرگز sneaky-throw در کدِ کتابخانهای، چون نوعِ خطا را از caller پنهان میکند.
supplier (ظرفِ خالی)، accumulator (عنصر → ظرف)، combiner (ادغامِ دو ظرف در موازی)، finisher (ظرف → نتیجهٔ نهایی). مشخصهها به بهینهساز میانبر میدهند: IDENTITY_FINISH یعنی finisher حذفشدنی است چون A همان R است؛ UNORDERED یعنی میشود سربارِ حفظِ ترتیب را کنار گذاشت؛ CONCURRENT یعنی accumulator امنِthread است پس چند thread در یک ظرفِ مشترک میریزند و combiner لازم نیست. groupingByConcurrent هر سه را دارد و برای موازی از groupingBy تندتر است، اما encounter order را فدا میکند.
true. این «صدقِ پوچ» (vacuous truth) است: گزارهٔ «همهٔ عناصر شرط را دارند» روی مجموعهٔ خالی بهطور منطقی صادق است چون هیچ نمونهٔ نقضی وجود ندارد. anyMatch روی خالی false و noneMatch روی خالی true میدهد. باگِ واقعی: چکِ اعتبارسنجی روی لیستِ خالی سهواً pass میشود؛ همیشه حالتِ خالی را جدا هندل کن.
parallelStream روی commonPool مشترکِ کلِ JVM اجرا میشود که برای کارِ CPU-bound و به تعدادِ هستهها سایز شده، نه برای بلاکشدن. ۵۰۰ فراخوانیِ HTTPِ بلاکشونده این pool را گرسنه میکند و باقیِ استریمهای موازیِ JVM (و کارهای داخلیِ JDK) را کُند میکند، در حالی که موازیسازیِ واقعی هم محدود به تعدادِ هستههاست. جایگزینِ درست virtual thread است: urls.stream().gather(Gatherers.mapConcurrent(50, this::fetch)) یا structured concurrency — هزاران کارِ بلاکشونده را ارزان اداره میکنند.
Gatherer یک عملیاتِ میانیِ سفارشی و احتمالاً حالتدار است که با Stream::gather صدا زده میشود؛ قرینهٔ collect است اما دوباره Stream میدهد (نهاییشده در JEP 485، جاوا ۲۴). قبل از آن، منطقِ حالتدار مثلِ running total (scan)، پنجرهٔ لغزان (windowSliding)، یا نگاشتِ همزمانِ کنترلشده (mapConcurrent) یا ناممکن بود یا تو را مجبور به دستزدن به state بیرونی (int[]) میکرد که در موازی میترکید. Gatherer این را به واحدی تمیز و composable تبدیل میکند.
وقتی لامبدا یک فیلد یا متدِ نمونه را capture میکند، ارجاعِ enclosing this را نگه میدارد، پس کلِ شیءِ میزبان زنده میماند. اگر آن لامبدا در جایی با عمرِ طولانی ثبت شود (event bus، کشِ static، CompletableFuture معلق)، شیءِ میزبان — شاید با مگابایتها state — هرگز garbage نمیشود. راهِ حل: فقط متغیرهای لازم را در یک local کپی و capture کن (نه فیلد را مستقیم)، یا صریح unregister کن. لامبدای non-capturing این مشکل را ندارد چون singletonِ بیstate است.
reduce برای reductionِ immutable طراحی شده و فرض میکند accumulator یک مقدارِ تازه برمیگرداند، نه اینکه یک ظرفِ مشترک را mutate کند. اینجا همه یک StringBuilder مشترک را دستکاری میکنند؛ در موازی این یک data race و نتیجهٔ خراب است، و identity هم فقط یک نمونه است که بینِ sub-taskها به اشتراک گذاشته میشود. راهِ درست، mutable reduction با collect(Collectors.joining()) یا یک Collectorِ StringBuilder-محور است که برای همین ساخته شده و combiner-اش ظرفها را درست ادغام میکند.
.filter(p) قبل از groupingBy عناصر را قبلِ گروهبندی حذف میکند، پس کلیدهایی که هیچ عضوی از فیلتر رد نکردند اصلاً در مپ ظاهر نمیشوند. filtering(p, downstream) (جاوا ۹) بهعنوان downstream فیلتر میکند، پس همهٔ کلیدها میمانند اما بعضی لیستشان خالی میشود. اگر کدِ پاییندست فرض میکند همهٔ کلیدها موجودند، انتخابِ اشتباه یعنی NPE یا نبودِ داده.
- رابطهای تابعی checked exception declare نمیکنند؛ wrap به unchecked یا استخراج به متد — هرگز sneaky-throw.
- هر
Collectorچهار تابع + مشخصههاست؛CONCURRENT/UNORDERED/IDENTITY_FINISHکارایی را عوض میکنند وgroupingByConcurrentبرای موازیِ واقعی لازم است. - مشخصههای
Spliterator(SIZED/DISTINCT/SORTED) به موتور اجازهٔ میانبر میدهند؛unordered()موازی را تند میکند. allMatchروی خالیtrue(صدقِ پوچ)؛findAnyدر موازی ازfindFirstارزانتر.- Gatherers (جاوا ۲۴) عملیاتِ میانیِ سفارشیِ حالتدار میدهد؛
mapConcurrent+virtual thread جایگزینِ درستِ parallelStream برای I/O است. - لامبدا با capture کردنِ فیلد کلِ
thisرا پین میکند → نشتِ حافظه.
- لامبدا در نهایت یک شیء با یک متد است؛ رابطِ تابعی (SAM) هدفِ آن است.
@FunctionalInterface، متدهایdefault/static، و نسخههای اولیه (IntFunctionو ...) برای فرار از باکسینگ را بشناس. - لامبدا با کلاسِ ناشناس در سه چیز فرق دارد: اتصالِ
this، سایهاندازی، و کامپایل (invokedynamic+LambdaMetafactory، کشِ singleton برای بدونحالتها). محلیهای گرفتهشده باید effectively final باشند؛ فیلدها نه. - استریم یک دستورِ آشپزی است، نه غذا: تنبل، یکبارمصرف، عنصربهعنصرِ عمودی.
map/filter/reduce/flatMap/mapMultiرا بلد باش؛reduceِ سهآرگومانی به combinerِ شرکتپذیر نیاز دارد. - کالکتورها:
groupingByبا کالکتورِ پاییندستی، دامِ کلیدِ تکراریِtoMap، تفاوتِStream.toList()(تغییرناپذیر) باCollectors.toList()(نامشخص)، وpartitioningByکه همیشه هر دو کلید را دارد. - استریمِ موازی قدرتمند ولی خطرناک است: پولِ مشترک، نبودِ I/O مسدودکننده، صفر حالتِ مشترک، و همیشه با بنچمارک. پیشفرض را ترتیبی بگذار.
Optionalرا فقط بهعنوانِ نوعِ بازگشتی بهکار ببر؛orElseدر برابرorElseGet،ofدر برابرofNullable، و هرگز برای فیلد/پارامتر/مجموعه.
Java 8 was one of the biggest shifts in the language's history: a fully functional layer bolted onto a language that, until that day, was one hundred percent object-oriented. The subtle part is that the JVM itself never changed — every lambda is still an ordinary object, and every stream still runs on plain method calls. Someone who understands this can not only write list.stream().map(...), but can also debug it when it misbehaves under production load. In this chapter we build everything up step by step, from the ground floor.
Three pillars hold up this whole chapter:
- Functional interfaces — types with a single abstract method that lambdas "target".
- Streams — a lazy, single-use pipeline of operations over a data source.
- Optional — a container that makes "maybe absent" explicit in the type, replacing
nullat API boundaries. Along the way we hit Collectors, parallel streams, common pitfalls, and finally 15 real interview questions.
Part 0 — words you must know first
Before we move on, let me build a few words from scratch that keep coming back, so you're never lost.
Picture a restaurant that serves exactly one dish: a beef stew. Its menu can have descriptions, opening hours, an address — but ultimately it does one "main job": it cooks stew. A functional interface is exactly this — it has just one abstract method that is its main job. That method is called the SAM, the Single Abstract Method. A lambda is really the shorthand for "here's how to do that one job".
- desugaring: the compiler quietly rewrites many sweet, short syntaxes into their fuller, more primitive form. A lambda is "syntactic sugar"; desugaring means seeing the verbose thing that's actually generated.
- boxing: Java has two number worlds — primitives like
int(light, on the stack) and object types likeInteger(heavy, on the heap). Every time you wrap anintinto anIntegera fresh object is allocated; that wrapping is boxing, and it's expensive in tight loops. - hot path: the slice of code that runs millions of times a second. Every extra allocation here gets multiplied by a million.
- lazy / eager: lazy means postponing work until the last needed moment; eager means doing it right now.
Part 1 — the mental model
Let's capture the whole thing in one sentence: Java 8 grafted a functional layer onto an object-oriented language without changing the JVM's type system. Every lambda is still an object implementing an interface; every stream is still driven by ordinary method calls. Knowing what desugaring is — what the compiler and runtime actually do — is the difference between an ordinary developer and a senior engineer.
Whenever you see a lambda, tell yourself: "this is a tiny object with one method." Whenever you see a stream, tell yourself: "this is a recipe, not the meal — until someone says cook (a terminal op), nothing gets cooked."
Part 2 — Functional interfaces
Analogy, concept, back to Java
A functional interface has exactly one abstract method (the SAM we built in Part 0). The @FunctionalInterface annotation is optional, but it does two things: it documents your intent, and it makes the compiler shout if you accidentally add a second abstract method. Note that default and static methods don't count toward the SAM — because they have bodies, they aren't abstract.
An abstract method is one with no body that someone must fill in later. default and static methods already have bodies, so there's nothing left to fill. The stew restaurant still has one main dish, even if its menu is full of pre-written notes.
The java.util.function package hands you the core set ready-made. Memorize this table; you'll need it repeatedly in interviews:
| Interface | Abstract method | Shape |
|---|---|---|
Supplier<T> |
T get() |
() -> T |
Consumer<T> |
void accept(T) |
T -> void |
Function<T,R> |
R apply(T) |
T -> R |
Predicate<T> |
boolean test(T) |
T -> boolean |
UnaryOperator<T> |
T apply(T) |
T -> T |
BiFunction<T,U,R> |
R apply(T,U) |
(T,U) -> R |
BinaryOperator<T> |
T apply(T,T) |
(T,T) -> T |
BiConsumer<T,U> |
void accept(T,U) |
(T,U) -> void |
BiPredicate<T,U> |
boolean test(T,U) |
(T,U) -> boolean |
Picture running a company: Supplier is the stockroom clerk who brings you something without taking anything (() -> T). Consumer is the trash bin that takes something and returns nothing (T -> void). Function is the assembly-line worker; takes an input, hands back a different output (T -> R). Predicate is the door guard; looks and answers only "yes/no" (T -> boolean). The Bi prefix means the same role but with two inputs instead of one.
There are also primitive-specialized variants: IntFunction, ToIntFunction, IntPredicate, IntUnaryOperator, ObjIntConsumer, and so on. Why do they exist? To escape boxing. If you work with Function<Integer,Integer>, every number must be wrapped into an Integer; but IntUnaryOperator works directly with int and allocates no extra object. In hot paths, prefer these.
Composition
The beauty of functional interfaces is that they carry their own default methods that let you glue two small functions into one bigger function — exactly like connecting several water pipes together.
Predicate<String> nonEmpty = s -> !s.isEmpty();
Predicate<String> shortish = s -> s.length() < 10;
Predicate<String> ok = nonEmpty.and(shortish).negate(); // De Morgan by hand not needed
Function<Integer,Integer> plus1 = x -> x + 1;
Function<Integer,Integer> times2 = x -> x * 2;
plus1.andThen(times2).apply(3); // (3+1)*2 = 8
plus1.compose(times2).apply(3); // (3*2)+1 = 7
andThen means "me first, then the other one": plus1.andThen(times2) first adds 1, then multiplies by 2 → (3+1)*2 = 8. But compose is the reverse: "the other one first, then me": plus1.compose(times2) first multiplies by 2, then adds 1 → (3*2)+1 = 7. If you remember that compose mirrors the math composition f∘g (g first, then f), you'll never get it wrong.
Comparator is also a functional interface, with a rich fluent API:
Comparator<Person> byAgeThenName =
Comparator.comparingInt(Person::age) // primitive-specialized, no boxing
.thenComparing(Person::name)
.reversed();
See comparingInt? It's the same "escape boxing" logic: since age is an int, we use the comparingInt variant so each comparison doesn't build an extra Integer. thenComparing says "if ages tie, decide by name," and reversed flips the whole order.
Lambdas vs anonymous classes
They look alike but differ in three ways — and those three are exactly what interviews ask about:
thisbinding. In an anonymous class,thisrefers to the anonymous instance. In a lambda,thisrefers to the enclosing instance (the class the lambda is written inside). A lambda has no separatethisof its own.- Scope / shadowing. A lambda shares the surrounding scope; you cannot declare a variable with the same name as an outer local. An anonymous class opens a new scope and can shadow an outer name (create a new one with the same name that hides the outer).
- Compilation. An anonymous class generates a real
.classfile (likeOuter$1.class) and is instantiated withnew. A lambda, by contrast, compiles to a private synthetic method (a method the compiler secretly creates) plus aninvokedynamicbootstrap throughLambdaMetafactory.
An anonymous class is like permanently hiring a full-time employee for every task right now (the .class file is pre-built). A lambda is more like an "as-needed" contract: the JVM doesn't spin up the implementing class until the first time that lambda is actually used. And if the lambda captures nothing from outside (stateless / non-capturing), the JVM builds just one instance and caches it as a singleton — since they're all identical, why make more?
Runnable a = new Runnable() {
public void run() { System.out.println(this.getClass()); } // anon class
};
Runnable l = () -> System.out.println(this.getClass()); // enclosing 'this'
In the first line this.getClass() prints the anonymous class name (something like Outer$1), because this is the anonymous instance. In the lambda, this.getClass() prints the enclosing class name, because the lambda has no this of its own.
Effectively-final capture
Here's a seemingly odd rule many people just memorize without understanding: a lambda may only capture local variables that are final or effectively final. Effectively final means a variable assigned once and never reassigned, even if you didn't write the word final.
Think of a stack local as a sticky note you left on the kitchen counter. When the method ends, the counter is cleared and the note is thrown away. But a lambda might live on past the method (say it was handed to another thread). So Java copies the note's value and stores it inside the lambda itself. Now, if it let you change the variable, you'd have two divergent versions: one on the counter, one inside the lambda — and which is correct? To avoid the ambiguity, Java forbids the change outright.
So this isn't a style rule — it's a memory-model guarantee. Locals live on the stack; the lambda may outlive the method, so the value is copied into the synthetic instance. Fields, by contrast (member variables of a class), are captured by reference through the enclosing this, so they can mutate — which is a common way people smuggle mutable state into a supposedly "functional" pipeline.
int base = 10; // effectively final -> OK
IntUnaryOperator f = x -> x + base;
// base = 11; // would break the capture: compile error
int[] counter = {0}; // classic escape hatch: the array ref is final,
list.forEach(x -> counter[0]++); // but its contents mutate. Legal, but a smell.
That int[] counter = {0} is a well-known hack: the array reference itself never changes (so it's effectively final), but you can mutate its contents. This compiles, but it smells bad because you're smuggling mutable state into functional code. The moment someone adds .parallel(), this exact code becomes a data-race bug.
Method references
When your lambda merely calls an existing method, Java offers a shorter path: the method reference. It comes in four forms, each just "syntactic sugar" for a lambda:
Function<String,Integer> len = String::length; // 1. instance method of arbitrary object: s -> s.length()
Supplier<List<String>> mk = ArrayList::new; // 2. constructor: () -> new ArrayList<>()
Consumer<String> pr = System.out::println; // 3. instance method of a specific object
Function<String,Integer> parse = Integer::parseInt; // 4. static method: s -> Integer.parseInt(s)
The tricky one is form 1: String::length is an unbound reference. "Unbound" means the receiver (the object the method is called on) isn't fixed up front. So the first parameter of the function becomes the receiver: s -> s.length(). Contrast it with myList::add, where the receiver (myList) is fixed in advance — that's called bound.
If you put a type name in the receiver spot (String::length), it's unbound and the receiver becomes the first argument. If you put a real object in the receiver spot ("hi"::length or myList::add), it's bound and the receiver stays that fixed object.
Part 3 — The Stream pipeline
A stream is not a data structure. It is a description of a computation over a source — like a recipe, not the meal itself. Until someone says "cook," nothing runs.
A pipeline has three parts:
source ──► intermediate ops (0..n) ──► terminal op (exactly 1)
List filter/map/sorted/... collect/forEach/reduce/count/...
(lazy, return Stream) (eager, triggers execution)
Three key properties you must have in your blood:
- Laziness. Intermediate operations do nothing until a terminal op runs. Then elements are pulled one at a time and pushed vertically through the whole chain — not one full pass per operator. This is exactly what enables short-circuiting and fusion.
- Single use. A stream is consumed once. Reusing it throws
IllegalStateException: stream has already been operated upon or closed. - No source mutation. A well-behaved pipeline never modifies its source.
Picture four people at a factory line: an inspector, a painter, a packer. Horizontal processing means all products first pass the inspector, then all go to the painter. Vertical processing (what streams do) means one product goes all the way down the line — inspect, paint, pack — then the next product's turn. The payoff: if you only want one good product (findFirst), the moment you reach it the whole line stops and no energy is wasted.
Laziness and short-circuiting demonstrated
List<String> names = List.of("alpha", "beta", "gamma", "delta");
Optional<String> first = names.stream()
.peek(s -> System.out.println("filter " + s))
.filter(s -> s.length() == 5)
.peek(s -> System.out.println("map " + s))
.map(String::toUpperCase)
.findFirst(); // short-circuits after first match
This prints filter alpha, then map alpha, then stops — it never touches beta/gamma/delta. Why? Because findFirst short-circuits (quits at the first result) and processing is element-at-a-time. alpha has length 5 so it passes the filter, enters map, and being the first result the whole pipeline halts right there.
Many people think filter first filters the entire list, then map runs over the result. No! Each element flows through the whole chain one at a time. That's what makes short-circuiting possible; if it were horizontal, findFirst would have to wait for the whole list to be filtered first.
map / filter / reduce / flatMap
// map: 1-to-1 transform
List<Integer> lengths = names.stream().map(String::length).toList();
// filter: keep matching
List<String> fives = names.stream().filter(s -> s.length() == 5).toList();
// reduce: fold to a single value. 3-arg form: identity, accumulator, combiner
int total = names.stream().reduce(0, (acc, s) -> acc + s.length(), Integer::sum);
// identity ^ accumulator ^^^^^^^^^^^^^^^^^^^^ combiner ^^^^^^^^^^
// The combiner merges partial results in parallel; must be associative and
// consistent with the accumulator, or parallel results silently differ.
// flatMap: 1-to-many, then flatten one level
List<List<Integer>> matrix = List.of(List.of(1,2), List.of(3,4));
List<Integer> flat = matrix.stream()
.flatMap(List::stream) // Stream<List<Integer>> -> Stream<Integer>
.toList(); // [1, 2, 3, 4]
Let's bring each to life. map is like a painting station: each element goes in, the same number of transformed elements come out (one-to-one). filter is a door guard: some pass, some don't. reduce is like folding a long sheet of paper: you fold repeatedly until you get one small thing — a sum, a max, whatever.
Imagine you have several boxes, and inside each box are several balls. If you map, you still have several boxes. But flatMap opens each box and pours all the balls into one single basket — it "flattens" one level of nesting. That's why Stream<List<Integer>> becomes Stream<Integer>.
The three-argument form of reduce is where many people slip. It has three parts: identity (the starting value), accumulator (how to fold one element into the running result), and combiner (how to merge two partial results). The identity must be a true identity for the combiner; that is, combiner.apply(identity, x) == x. If that's violated, parallel and sequential runs disagree. Another subtle point: the accumulator's result type (R) may differ from the element type (T), and that's exactly why a separate combiner is needed.
Suppose you want to count the total length of a 1000-page book, and you have 4 friends. The smart move: split the book into 4 parts, each person counts their part (that's the accumulator's job), then add up the 4 numbers (that's the combiner's job). The combiner is the "how do we merge four separate numbers into one" step. Without it, parallelism would be impossible. And if your counting is inconsistent with the final adding-up, the final answer comes out wrong.
Since Java 16, mapMulti is a cheaper alternative to flatMap. Instead of allocating a new Stream per element (which costs an allocation), it pushes results directly into a Consumer (called the sink). When each element fans out to just a few (small fan-out), this is cheaper:
Stream.of(1,2,3).<Integer>mapMulti((n, sink) -> { sink.accept(n); sink.accept(-n); });
Primitive streams
IntStream, LongStream, DoubleStream are three special streams that work directly with primitives to escape boxing, and they add handy numeric terminal ops (sum, average, max). Bridge between them with mapToInt / boxed / asLongStream:
int sum = names.stream().mapToInt(String::length).sum();
IntSummaryStatistics stats = IntStream.rangeClosed(1, 100).summaryStatistics();
stats.getAverage(); stats.getMax(); stats.getCount();
double avg = names.stream().mapToInt(String::length).average().orElse(0);
average(), max(), min() return an OptionalDouble/OptionalInt, not a raw number. The reason is simple: what's the average of an empty stream? Nothing! So instead of returning a misleading zero or throwing, they return a "maybe empty" container and hand the decision to you (.orElse(0)).
Part 4 — Collectors
collect is the general-purpose mutable reduction terminal op. "Reduction" means going from many elements to one result; "mutable" means we do it by filling a mutable container (like a List or Map that grows step by step). The Collectors factory covers almost everything you'll need:
// grouping: Map<K, List<V>>
Map<Integer, List<String>> byLen =
names.stream().collect(Collectors.groupingBy(String::length));
// grouping with a downstream collector: Map<K, aggregate>
Map<Integer, Long> countByLen =
names.stream().collect(Collectors.groupingBy(String::length, Collectors.counting()));
Map<Integer, String> joinedByLen =
names.stream().collect(Collectors.groupingBy(
String::length, Collectors.joining(", ", "[", "]")));
// toMap: beware duplicate keys -> IllegalStateException without a merge fn
Map<Integer, String> byLenFirstWins = names.stream()
.collect(Collectors.toMap(String::length, s -> s, (a, b) -> a)); // merge = keep first
// partitioning: always Map<Boolean, List<V>> with BOTH true and false keys present
Map<Boolean, List<String>> parts =
names.stream().collect(Collectors.partitioningBy(s -> s.length() > 4));
// teeing (Java 12): run two collectors, merge results — one pass
record MinMax(int min, int max) {}
MinMax mm = IntStream.rangeClosed(1, 10).boxed().collect(Collectors.teeing(
Collectors.minBy(Integer::compareTo),
Collectors.maxBy(Integer::compareTo),
(lo, hi) -> new MinMax(lo.orElseThrow(), hi.orElseThrow())));
groupingBy is like sorting workers into several lines by height. But then what do you do with each line? That "downstream collector" is what happens inside each line: count them (counting()), join them (joining(...)), or just dump them into a list (the default). So groupingBy(len, counting()) means "group by length, and in each group keep only the count".
teeing (Java 12) is elegant too: like the letter T with two branches, it feeds the stream to two collectors at once and then blends their results with a third function — all in a single pass over the stream. In the example above it grabs the min and the max simultaneously and packs them into a record.
Now two frequent bugs that interviewers love to ask:
toMap throws IllegalStateException at runtime if two elements produce the same key — unless you supply a merge function telling it "when two values fight over one key, which stays". For example, (a, b) -> a means "keep the first". But groupingBy never has this problem, because by default it places same-key values side by side in a list.
Collectors.toList() gives no guarantee about the returned list's type or mutability. If you want an immutable result, use Stream.toList() (Java 16+, unmodifiable) or Collectors.toUnmodifiableList(). If you want a specific mutable type, use Collectors.toCollection(ArrayList::new).
A key senior-level point: Stream.toList() vs Collectors.toList(). The former returns an unmodifiable list and even allows null elements; the latter historically returned an ArrayList, but that is unspecified by contract. Migrating blindly can break code that later mutates the result.
Part 5 — Parallel streams: power and peril
Imagine an office building with only one shared kitchen. If one person goes there and waits hours for water to boil (a blocking task), the rest of the building goes hungry. ForkJoinPool.commonPool() is exactly that shared kitchen for the whole JVM.
stream.parallel() (or Collection.parallelStream()) splits the source via a Spliterator (a stream-into-chunks splitter) and farms the work to the common ForkJoinPool — that's ForkJoinPool.commonPool(), sized to #cores - 1 by default. This is the single most misused feature in the entire API.
When parallelism helps:
- Large N (tens of thousands+), CPU-bound per-element work.
- A source that splits cheaply and evenly: arrays,
ArrayList,IntStream.range. By contrastLinkedListand mostIterator-backed sources split poorly (to reach the middle you must walk from the start). - No ordering constraint, or you can tolerate
unordered().
When it hurts or is outright wrong:
- Shared common pool. All parallel streams in the JVM share one pool. A blocking task (I/O, JDBC,
sleep) inside a parallel stream starves every other user, including internal JDK ones. Never do blocking I/O in a parallel stream; if you must, wrap it in your ownForkJoinPooland submit the pipeline as a task. - Non-associative reduce gives nondeterministic results.
- Stateful lambdas / shared mutable state cause data races (two threads writing the same thing at once). This code is broken:
List<Integer> out = new ArrayList<>(); // not thread-safe
IntStream.range(0, 10_000).parallel()
.forEach(out::add); // RACE: lost updates or exceptions
// Fix: .collect(Collectors.toList()) or .boxed().collect(...), which is race-free by design.
forEachdoes not preserve order in parallel; if you need encounter order (the order the elements actually were in the source), useforEachOrdered(at a performance cost).- Small N or cheap work: fork/join startup overhead dwarfs the benefit; sequential is faster.
Always default to sequential. Only reach for .parallel() when all three hold: a benchmark (like JMH) proves it's faster, the source is splittable (array/ArrayList), and you have zero shared mutable state.
Ordering, statefulness, side-effects
- Encounter order is a property of the source: a
Listhas it (it's ordered), aHashSetdoesn't. The opssorted/distinct/limitare stateful intermediate ops — meaning they must remember previous elements to do their job. They may buffer the whole stream in memory, which defeats laziness and can OOM (run out of memory) on infinite streams. - Side effects in
map/filterare a code smell and unsafe in parallel.peekis intended only for debugging; the JDK explicitly warns that when a downstream op (likecount) can be computed without traversal,peekmay not run for every element.
long n = Stream.of("a","b","c").peek(System.out::println).count();
// May print nothing: count() can short-circuit since Java 9 sizes the stream directly.
Since Java 9, if no size-changing op (filter/flatMap) precedes count(), Java can report the count without even looking at the elements. So peek(System.out::println) prints nothing. The bigger lesson: use peek only for temporary debugging, never for program logic.
Part 6 — Optional: correct use and anti-patterns
Imagine someone hands you an envelope and says "maybe there's a letter inside, maybe it's empty." The moment you see the envelope, you know to account for the empty case — no surprises. Optional<T> is that envelope: instead of silently handing you a null and surprising you with a NullPointerException at runtime, the method's type says out loud "watch out, there may be no value."
Optional<T> communicates "a value that may be absent" as a return type. Highlight this: it was designed for return values, not for fields and not for parameters.
Optional<User> found = repo.findById(id);
// GOOD: express the fallback / branch fluently
String name = found.map(User::name).orElse("anonymous");
found.ifPresentOrElse(u -> log.info("hit {}", u), () -> log.warn("miss")); // Java 9
User u = found.orElseThrow(() -> new NotFoundException(id)); // throw with context
Optional<User> chained = found.or(() -> repo.findInCache(id)); // Java 9, lazy fallback
See how fluent that is: map(User::name) says "if there's a user, take its name," and orElse("anonymous") says "otherwise anonymous." No if, no null anywhere. ifPresentOrElse (Java 9) gives two branches; orElseThrow throws with a meaningful message; and or (Java 9) provides a lazy fallback that runs only when empty.
Anti-patterns
// 1. isPresent()/get() — reimplements null-checking, defeats the purpose
if (found.isPresent()) return found.get(); // AVOID: use map/orElse/orElseThrow
// 2. orElse with an expensive/side-effecting arg — ALWAYS evaluated, even when present
User u = found.orElse(createExpensiveDefault()); // BUG: default built every call
User u2 = found.orElseGet(() -> createExpensiveDefault()); // FIX: lazy supplier
// 3. Optional fields / parameters — adds a wrapper, breaks serialization, no benefit
class Order { private Optional<Coupon> coupon; } // AVOID
void apply(Optional<Coupon> c) { } // AVOID: overload or accept null instead
// 4. Optional.get() without checking — throws NoSuchElementException
found.get(); // AVOID unless you just checked isPresent
// 5. Wrapping then unwrapping needlessly
return Optional.ofNullable(x).orElse(y); // just: return x != null ? x : y
Look hard at anti-pattern #2, because it's an interview classic. orElse(v) takes an already-computed value — so its argument is always built, even when the envelope is full and no default is needed at all! If createExpensiveDefault() is expensive or has a side effect (say it writes to a database), this is a real bug. orElseGet(() -> ...) is lazy and runs the supplier only when the envelope is empty.
Let's clear up the remaining points. Optional.of(x) throws NPE immediately if x is null — use it as an assertion saying "I'm sure this isn't null." Optional.ofNullable(x) is gentler and tolerates null (if null, it becomes empty). And Optional.stream() (Java 9) turns an Optional into a 0-or-1-element stream — an excellent tool for flat-mapping away empties:
List<User> users = ids.stream()
.map(repo::findById) // Stream<Optional<User>>
.flatMap(Optional::stream) // drops empties, unwraps present ones
.toList();
If your method returns a List or Map, never return Optional<List> — return an empty List/Map instead. An empty collection already means exactly "nothing," so the Optional layer is redundant and only makes the caller's job harder.
Part 7 — Pitfalls and best practices
Keep these as a final checklist:
- Prefer
Stream.toList()(Java 16+) overcollect(toList())for read-only results; but first know the mutability difference. - Keep lambdas short and pure (no side effects); when logic grows or needs a name in stack traces, extract it to a named method (and use a method reference) — lambda frames show up in stack traces as the ugly
lambda$method$0. - Never mutate shared state from a stream, even sequentially — because it breaks the moment someone adds
.parallel(). - Use primitive streams in numeric hot paths to kill boxing.
groupingBy+ a downstream collector always beats "collect to lists then re-stream".- Beware infinite streams (
Stream.iterate,generate) with stateful ops likesorted/distinct— they never terminate. Putlimit/takeWhilefirst. takeWhile/dropWhile(Java 9) short-circuit on a sorted-ish stream;filterdoes not. (takeWhiletakes elements up to the first failure of the condition, then stops;filterwalks the entire stream.)
Part 8 — Interview Questions
Read each one seriously; these are exactly where an interviewer separates senior from mid-level.
Locals live on the stack and are copied by value into the lambda's synthetic instance; the lambda can outlive the method frame, so allowing reassignment would create two inconsistent copies (one on the stack, one inside the lambda). Fields, however, are reached through the captured enclosing this reference, so mutations are visible through that shared reference — the compiler only needs the reference itself to be stable, not the field's contents.
orElse(v) takes an already-computed value — its argument is always evaluated, even when the Optional is present. orElseGet(supplier) is lazy: the supplier runs only when empty. Passing an expensive or side-effecting default to orElse is a real performance/correctness bug.
No. Streams process element-by-element, vertically through the whole pipeline. Each element flows through filter then map before the next element starts. This is what makes short-circuiting (findFirst, anyMatch, limit) possible without processing the rest.
long n = Stream.of("a","b","c").peek(System.out::println).count();
System.out.println(n);
Likely just 3. Since Java 9, count() can determine size without traversal when no size-changing ops (filter/flatMap) precede it, so peek may never fire. Relying on peek for anything but debugging is unsafe.
Map<String,Integer> m = words.stream()
.collect(Collectors.toMap(w -> w.substring(0,1), String::length));
toMap throws IllegalStateException on duplicate keys (two words with the same first letter). Add a merge function: Collectors.toMap(k, v, (a,b) -> a) or use groupingBy.
Slower: small N, cheap per-element work, poorly-splittable sources (LinkedList), or blocking I/O (which starves the shared common ForkJoinPool). Wrong: shared mutable state (data races), non-associative reduce, or reliance on encounter order with plain forEach.
Stream.toList() (Java 16+) returns an unmodifiable list with a specified contract and allows nulls. Collectors.toList() returns an unspecified, usually-mutable ArrayList — you must not rely on either its type or its mutability. Swapping one for the other can break code that mutates the result or, conversely, expects immutability.
reduce(identity, accumulator, combiner): accumulator: (R,T)->R folds an element into a partial result of a different type R; combiner: (R,R)->R merges two partial results. The combiner exists so parallel execution can fold sub-ranges independently then merge them. The identity must satisfy combiner(identity, x) == x and the accumulator must be associative/consistent with the combiner, or parallel and sequential results diverge.
It adds an allocation and a wrapper layer with no readability gain, breaks common serialization frameworks (Optional isn't Serializable), and forces callers to construct wrappers. For parameters, prefer overloads or nullable arguments; for fields, store the raw value (possibly null) and return Optional from the getter.
String::length is unbound: it has no fixed receiver, so the first parameter of the target functional type becomes the receiver — it desugars to s -> s.length(), matching Function<String,Integer>. Contrast "hi"::length (bound: fixed receiver, Supplier<Integer>).
IllegalStateException: stream has already been operated upon or closed. Streams are single-use; assign the source to a variable and create a fresh stream each time, or restructure to one pipeline.
They couple the pipeline to execution details that the runtime is free to optimize away (like count skipping traversal) or reorder/parallelize. Side effects also make the code non-thread-safe the instant someone adds .parallel(), turning a silent race into production data loss.
partitioningBy always returns a map with both true and false keys, even when one partition is empty. groupingBy on a boolean key omits keys that have no elements. Downstream code that assumes both keys exist will NPE with groupingBy.
Submit the terminal operation as a task to your own ForkJoinPool: myPool.submit(() -> stream.parallel().reduce(...)).get(). The parallel stream inherits the pool of the submitting thread. This isolates blocking or long work from the shared common pool and avoids starving the rest of the JVM.
Optional.of(null) throws NullPointerException immediately — use it to assert non-null. Optional.ofNullable(null) returns Optional.empty(). Choosing the wrong one either masks a bug or throws unexpectedly.
Senior notes & advanced edge cases
We've now built the whole functional toolkit. But what separates a real senior from a solid developer is knowing the sharp edges that never show up in the demo and only reveal themselves at 3 a.m. when the pager lights up. This section is exactly those edges: exceptions inside lambdas, building a custom Collector, Spliterator characteristics, the parallel-stream trap in a virtual-thread world, and the modern-Java update (Gatherers).
- Checked exceptions inside lambdas — the daily pain of every team that actually ships code.
- The anatomy of
Collectoritself — four functions and characteristics, plus concurrent collectors. Spliteratorcharacteristics — why the optimizer sometimes skips your work entirely.- The short-circuit truth table — the "vacuous truth" trap and
findAnyvsfindFirst. - Gatherers (Java 24) — you can finally write custom intermediate operations.
- Lambda memory leaks and Serializable lambdas.
- Parallel streams vs virtual threads for I/O work.
- Hard senior-level interview questions.
1) Checked exceptions inside lambdas — the bitter daily truth
Let me lead with this because it eats more team-hours than anything else: none of the java.util.function interfaces declare any checked exception. The moment you call a method that throws IOException inside map, the code won't compile.
List<String> paths = List.of("a.txt", "b.txt");
paths.stream()
.map(p -> Files.readString(Path.of(p))) // compile error: unhandled IOException
.toList();
Function<T,R> is a general contract meant to run in thousands of different places — including a parallel stream across several threads. If every lambda could throw any checked exception, the stream engine would have to know where and on which thread to deliver that exception. The designers chose to keep this complexity out entirely, so these interfaces have "clean" signatures and only unchecked exceptions pass through.
You have three practical routes, and all three show up in real code:
// Route 1: try/catch inside the lambda and rethrow as unchecked (clearest for simple logic)
.map(p -> {
try { return Files.readString(Path.of(p)); }
catch (IOException e) { throw new UncheckedIOException(e); } // JDK's ready-made type
})
// Route 2: a throwing functional interface + an adapter that wraps it
@FunctionalInterface interface ThrowingFn<T,R> { R apply(T t) throws Exception; }
static <T,R> Function<T,R> unchecked(ThrowingFn<T,R> f) {
return t -> { try { return f.apply(t); }
catch (Exception e) { throw new RuntimeException(e); } };
}
// usage: .map(unchecked(p -> Files.readString(Path.of(p))))
A well-known generics trick lets you smuggle a checked exception out without wrapping it (called sneaky-throw). It compiles and the signatures stay clean — but you've built a monster: the caller now receives an IOException that appears in no method signature, so it can't catch it properly and tools can't warn about it. Never do this in library code; always wrap cleanly in UncheckedIOException/CompletionException so the stack and the error type stay honest.
If a lambda has grown big enough to need a three-line try/catch, that's the signal to extract it into a named method and call it via a method reference. The stack trace gets cleaner (you see the real method name instead of lambda$process$3), and it becomes testable.
2) The anatomy of Collector itself — what happens behind groupingBy
The chapter showed how to use the ready-made Collectors. A senior must be able to build one, because every Collector is really four pieces:
public interface Collector<T, A, R> {
Supplier<A> supplier(); // make a fresh empty accumulator container
BiConsumer<A, T> accumulator(); // fold one element into the container
BinaryOperator<A> combiner(); // merge two half-full containers (parallel)
Function<A, R> finisher(); // turn internal container A into final result R
Set<Characteristics> characteristics();
}
The three generic letters matter: T is the input element type, A is the mutable intermediate (accumulation) type, and R is the final result type. In joining, for example, A is a StringBuilder but R is a String; there the finisher is just .toString().
You have three labels. IDENTITY_FINISH means "the finisher does nothing, the container A is the result" — so the engine can drop that step entirely. UNORDERED means "I don't care about element order" — letting the engine skip the cost of preserving order in parallel. CONCURRENT means "my accumulator is thread-safe, so multiple threads can pour into one shared container and no combiner is even needed".
This is exactly where seniors part ways on performance:
Collectors.groupingBy(...) does not carry the CONCURRENT characteristic. So in a parallel stream, each thread builds its own separate HashMap and then they all merge via the combiner — and merging maps is expensive. If you truly want parallelism, use groupingByConcurrent (or toConcurrentMap), which fills one shared ConcurrentHashMap directly and eliminates the merge step — at the cost of losing encounter order. Choose that trade-off explicitly, not by accident.
Lesser-known composite collectors a senior should know:
// mapping: map each element before the downstream collector
Map<Integer,List<Character>> firstChars = words.stream().collect(
groupingBy(String::length, mapping(w -> w.charAt(0), toList())));
// filtering (Java 9): filter AFTER grouping — empty keys are kept (key difference from filter-before-groupingBy)
Map<Dept,List<Emp>> seniorsByDept = emps.stream().collect(
groupingBy(Emp::dept, filtering(e -> e.level() > 5, toList())));
// collectingAndThen: push the final result one more step (e.g. make it immutable)
List<String> frozen = names.stream().collect(
collectingAndThen(toList(), List::copyOf));
// reducing/summingInt as a downstream to aggregate within each group
Map<Dept,Integer> payroll = emps.stream().collect(
groupingBy(Emp::dept, summingInt(Emp::salary)));
If you write .filter(...) before groupingBy, the groups whose members all failed the filter don't appear in the map at all. But if you use filtering(...) as a downstream, those keys remain with an empty list. If downstream code expects every department present as a key, this difference means either an NPE or its absence.
Anatomy diagram — the four moving parts of a Collector / چهار قطعهٔ یک Collector:
flowchart LR
Src[Stream elements T] --> Acc
Sup[supplier: new empty A] --> Acc[accumulator: A x T -> A]
Acc --> Comb[combiner: A x A -> A parallel merge]
Comb --> Fin[finisher: A -> R]
Fin --> Out[Result R]
3) Spliterator characteristics — why the engine sometimes skips your work
The chapter noted that count() sometimes answers without traversal and peek doesn't run. Here's the why: every stream's source has a Spliterator that carries metadata flags: SIZED (I know my exact size), SUBSIZED, ORDERED, SORTED, DISTINCT, NONNULL, IMMUTABLE.
These flags let the optimizer take shortcuts:
count()on aSIZEDsource with no size-changing op (filter/flatMap) just reads the size — which is exactly whypeeknever fires.distinct()on a stream that is alreadyDISTINCT(e.g. from aTreeSet) is effectively a no-op.sorted()on aSORTEDsource with the same comparator is skipped.
If output order doesn't matter, call .unordered() explicitly. For distinct/limit/skip in parallel, preserving encounter order is expensive because threads must stay coordinated. Drop that constraint and the optimizer can take any element from any thread first. Many teams live for months with a slow distinct().limit(n) never realizing all it needed was a .unordered() in front of it.
4) The short-circuit truth table — vacuous truth and findAny
The three matching methods (anyMatch/allMatch/noneMatch) behave on an empty stream in a way that traps people in interviews:
| On an empty stream | Result |
|---|---|
anyMatch(p) |
false |
allMatch(p) |
true ← vacuous truth |
noneMatch(p) |
true |
This is math, not a Java bug: "all elements satisfy the predicate" is vacuously true when there are no elements. Real code that forgets this: orders.stream().allMatch(Order::isPaid) returns true on an empty order list and may accidentally open an "everything is paid" flow. Always handle isEmpty separately first.
findFirst always wants the first element in encounter order; in parallel that means coordination and overhead. findAny says "whichever element is found soonest is fine" — much cheaper in parallel. If it logically doesn't matter which element you get back (e.g. only existence matters), choose findAny.
5) Gatherers — custom intermediate operations (Java 24)
The biggest stream update from Java 8 to today: until now you could only use the built-in intermediate operations (map/filter/...) and could not write your own stateful intermediate operation. Something like "map each element with a running total" or "group elements into windows of three" was either impossible or ugly with the old API. Stream::gather fixes this: it is the exact mirror of collect, but instead of a final result it yields another Stream.
Stream Gatherers arrived as preview in JEP 461 (Java 22) and JEP 473 (Java 23), and were finalized in JEP 485 in Java 24. In Java 25 (LTS) they're stable. Five ready-made gatherers live in java.util.stream.Gatherers: fold, scan, windowFixed, windowSliding, mapConcurrent.
// running total (scan) — practically impossible with the old API without an external array
List<Integer> runningSums = Stream.of(1, 2, 3, 4)
.gather(Gatherers.scan(() -> 0, Integer::sum))
.toList(); // [1, 3, 6, 10]
// sliding window of three
List<List<Integer>> windows = Stream.of(1, 2, 3, 4, 5)
.gather(Gatherers.windowSliding(3))
.toList(); // [[1,2,3],[2,3,4],[3,4,5]]
// mapConcurrent: map each element on a virtual thread with a bounded concurrency limit
List<String> bodies = urls.stream()
.gather(Gatherers.mapConcurrent(10, this::httpGet)) // at most 10 concurrent requests
.toList();
Before this, any stateful logic you needed inside a stream operation (running max, dedupe-by-key, batching) forced you to either leave the stream or cheat with an int[]/external variable — the exact smell we saw in the effectively-final section. gather turns that logic into a clean, composable, parallel-safe unit.
6) Lambda memory leaks and Serializable lambdas
When a lambda captures an instance field or an instance method, it actually captures the whole this (fields are reached through the enclosing this — the chapter's point). Now if that lambda is stored in a long-lived structure — a listener, a static cache, a CompletableFuture chain that never completes — then the entire host object (perhaps a heavy controller full of state) is never garbage-collected. This is a classic source of memory leaks in UIs and event buses.
class HeavyController {
private final byte[] cache = new byte[50_000_000];
void register(EventBus bus) {
bus.subscribe(e -> handle(e)); // handle is an instance method → this, hence cache, is pinned
}
}
Fix: write the lambda so it captures only what it needs (copy the field into a local first), or unsubscribe explicitly.
A lambda is not Serializable by default. If you force it with an intersection cast ((Runnable & Serializable) () -> ...), Java activates the heavy SerializedLambda machinery and the $deserializeLambda$ method, and its serialization is brittle (it depends on the synthetic method's name, which can change between compilations). You see this in distributed systems like Spark; never make a lambda Serializable without a genuine need.
7) Parallel streams vs virtual threads — the modern trap
The chapter said not to put blocking I/O in a parallel stream because it starves the shared commonPool. So what's the right way in modern Java?
parallelStream() is for CPU-bound work over in-memory data and still runs on ForkJoinPool.commonPool() — which is not virtual threads and was never built to block. For I/O-bound work (calling 500 HTTP services), the right tool is virtual threads (JDK 21): either Gatherers.mapConcurrent(n, ...), which runs directly on virtual threads, or structured concurrency (standard in JDK 25). These handle thousands of blocking tasks cheaply without starving any shared pool.
// common mistake: I/O on the shared commonPool → slows down the whole JVM
urls.parallelStream().map(this::httpGet).toList();
// the modern correct way (JDK 21+): virtual threads, a thousand blocking calls without starving any pool
urls.stream().gather(Gatherers.mapConcurrent(50, this::httpGet)).toList();
8) Two more small bites that draw blood
Collectors.groupingBy throws NullPointerException if the classifier returns null (even though HashMap allows a null key, groupingBy itself guards it with Objects.requireNonNull). This happens constantly on dirty real-world data: groupingBy(User::country) when some users have no country. Map null to a sentinel value (like "UNKNOWN") first.
opt.map(f) — if f produces null, instead of blowing up it silently returns Optional.empty(). This looks nice but is a silent bug: you think the value "was absent" when it was actually "present but mapped to null". And if f itself returns an Optional, map hands you Optional<Optional<X>>; there you must use flatMap to flatten one layer.
Because Function<T,R> declares no checked exception and a lambda can't throw more than its target SAM allows. The design reason is that functional interfaces are general and must work in parallel/async contexts, where the propagation path for a checked exception becomes ambiguous. In practice: try/catch inside the lambda and wrap to unchecked (UncheckedIOException), or an unchecked(ThrowingFn) adapter, or extract to a named method. Never sneaky-throw in library code, because it hides the error type from the caller.
supplier (empty container), accumulator (element → container), combiner (merge two containers in parallel), finisher (container → final result). Characteristics give the optimizer shortcuts: IDENTITY_FINISH means the finisher can be dropped because A is R; UNORDERED means order-preservation overhead can be skipped; CONCURRENT means the accumulator is thread-safe so multiple threads pour into one shared container and no combiner is needed. groupingByConcurrent has all three and beats groupingBy in parallel, but sacrifices encounter order.
true. This is vacuous truth: the proposition "all elements satisfy the predicate" is logically true over an empty set because there's no counterexample. anyMatch returns false on empty and noneMatch returns true on empty. Real bug: a validation check passes accidentally on an empty list; always handle the empty case separately.
parallelStream runs on the JVM-wide shared commonPool, sized for CPU-bound work at the number of cores, not for blocking. 500 blocking HTTP calls starve that pool and slow down every other parallel stream in the JVM (and internal JDK work), while the real parallelism is still capped at the core count. The correct replacement is virtual threads: urls.stream().gather(Gatherers.mapConcurrent(50, this::fetch)) or structured concurrency — they handle thousands of blocking tasks cheaply.
A Gatherer is a custom, possibly stateful intermediate operation invoked via Stream::gather; it's the mirror of collect but yields a Stream again (finalized in JEP 485, Java 24). Before it, stateful logic like a running total (scan), a sliding window (windowSliding), or controlled concurrent mapping (mapConcurrent) was either impossible or forced you to touch external state (int[]) that broke in parallel. A Gatherer turns that into a clean, composable unit.
When a lambda captures an instance field or method, it holds the enclosing this reference, so the entire host object stays alive. If that lambda is registered somewhere long-lived (an event bus, a static cache, a pending CompletableFuture), the host object — possibly megabytes of state — is never garbage-collected. Fix: copy only the needed variables into a local and capture those (not the field directly), or unregister explicitly. A non-capturing lambda doesn't have this problem because it's a stateless singleton.
reduce is designed for immutable reduction and assumes the accumulator returns a fresh value rather than mutating a shared container. Here everyone mutates one shared StringBuilder; in parallel that's a data race and a corrupted result, and the identity is just one instance shared across sub-tasks. The correct approach is mutable reduction with collect(Collectors.joining()) or a StringBuilder-based Collector built for exactly this, whose combiner merges the containers correctly.
.filter(p) before groupingBy removes elements before grouping, so keys whose members all failed the filter never appear in the map at all. filtering(p, downstream) (Java 9) filters as a downstream, so all keys remain but some end up with an empty list. If downstream code assumes all keys are present, the wrong choice means an NPE or missing data.
- Functional interfaces declare no checked exceptions; wrap to unchecked or extract to a method — never sneaky-throw.
- Every
Collectoris four functions + characteristics;CONCURRENT/UNORDERED/IDENTITY_FINISHchange performance andgroupingByConcurrentis needed for true parallelism. Spliteratorcharacteristics (SIZED/DISTINCT/SORTED) let the engine take shortcuts;unordered()speeds up parallel.allMatchon empty istrue(vacuous truth);findAnyis cheaper thanfindFirstin parallel.- Gatherers (Java 24) give custom stateful intermediate ops;
mapConcurrent+virtual threads is the right replacement for parallelStream on I/O. - A lambda capturing a field pins the whole
this→ memory leak.
- A lambda is ultimately an object with one method; the functional interface (SAM) is its target. Know
@FunctionalInterface,default/staticmethods, and the primitive variants (IntFunction, ...) for escaping boxing. - A lambda differs from an anonymous class in three ways:
thisbinding, shadowing, and compilation (invokedynamic+LambdaMetafactory, singleton cache for stateless ones). Captured locals must be effectively final; fields need not be. - A stream is a recipe, not the meal: lazy, single-use, element-by-element vertical. Know
map/filter/reduce/flatMap/mapMulti; three-argreduceneeds an associative combiner. - Collectors:
groupingBywith a downstream collector, the duplicate-key trap oftoMap, the difference betweenStream.toList()(unmodifiable) andCollectors.toList()(unspecified), andpartitioningBywhich always has both keys. - Parallel streams are powerful but perilous: shared common pool, no blocking I/O, zero shared state, and always with a benchmark. Default to sequential.
- Use
Optionalonly as a return type;orElsevsorElseGet,ofvsofNullable, and never for a field/parameter/collection.