People & Soft Skills · مهارتهای نرم و انسانی متوسطIntermediate ~98 دقیقه مطالعه~93 min read
مهارتهای نرم در محیط کار: ارتباط، تأثیرگذاری و رشدWorkplace Soft Skills: Communication, Influence & Growth
راهنمای عملی مهارتهای حرفهای محیط کار برای مهندسها — از نوشتن پیام، سند طراحی، ADR و postmortem تا مرور کد، جلسه، تخمین، مدیریت رو به بالا، تأثیرگذاری بدون اختیار، تعارض، رشد، اخلاق حرفهای و پایداری، همراه با قالب و جملهٔ آماده برای هر موقعیت.A hands-on guide to the professional skills that decide whether a strong engineer becomes senior — writing, design docs, ADRs, postmortems, code review, meetings, estimation, managing up, influence without authority, conflict, growth, ethics and sustainability, each with templates and scripts you can use.
یک صحنهٔ آشنا: دو مهندس در یک تیم. اولی الگوریتمها را بهتر میداند، سریعتر کد میزند و در بحثهای فنی معمولاً حق با اوست. دومی کندتر است، اما وقتی چیزی مینویسد بقیه میخوانند، وقتی تخمین میدهد کسی نگران نیست، و وقتی پیشنهادی میدهد تیم قبولش میکند. شش ماه بعد، دومی lead یک پروژهٔ بزرگ میشود و اولی هنوز همان ticketهای خودش را میگیرد و تحویل میدهد — و صادقانه نمیفهمد چرا.
این فصل دقیقاً دربارهٔ همان چیزی است که در آن فاصله اتفاق میافتد. و میخواهم از همان اول یک سوءتفاهم رایج را خراب کنم: اسم اینها را گذاشتهاند «مهارت نرم»، انگار که نرم یعنی مبهم، سلیقهای، یا چیزی که یا داری یا نداری. این غلط است. نوشتن یک design document که مرور میشود، دادن یک تخمین که بعداً به آبرویت لطمه نمیزند، مخالفت کردن در یک جلسه بدون تبدیل شدن به آدم منفی جلسه — اینها مهارتهای قابل آموزشاند، با ساختار مشخص، الگوی مشخص و جملات مشخص. دقیقاً مثل یک الگوی طراحی. و مثل هر مهارت فنی دیگری، با تمرین بهتر میشوند.
یک تفکیک مهم قبل از شروع: این فصل دربارهٔ انجام دادن کار است، نه گرفتن کار. هنر مصاحبه، ساختار پاسخ رفتاری و بانک سؤالهای آماده جای دیگریاند (فصلهای interview-craft و behavioral-interview-bank). اینجا فرض میکنیم استخدام شدهای و حالا باید در محیط واقعی بمانی، اثر بگذاری و رشد کنی.
۱) استدلال صادقانه: چرا از یک نقطه به بعد، سقف رشدت دیگر دانش فنی نیست · ۲) «سنیور» یعنی چه، و تفاوت مسیر IC با مسیر مدیریت · ۳) نوشتار بهعنوان پرقدرتترین اهرم یک مهندس: پیام، سؤال فنی، design doc، ADR، توضیح PR، commit، گزارش باگ، status update، postmortem · ۴) code review بهعنوان یک مهارت اجتماعی، از هر دو طرف · ۵) جلسهها، مخالفت سازنده و «نه» گفتن · ۶) تخمین، تعهد و مدیریت انتظار + مدیریت رو به بالا · ۷) تأثیرگذاری بدون اختیار رسمی و کار با غیرمهندسها · ۸) همکاری، تعارض و نردبان escalation · ۹) یادگیری و رشد بهعنوان یک عادت، نه یک حال خوب · ۱۰) قضاوت حرفهای و اخلاق · ۱۱) پایداری: burnout، مرزها، on-call، دورکاری و ۹۰ روز اول · ۱۲) یک «سیستمعامل» هفتگی برای خودت.
۱) استدلال صادقانه: سقف کجاست؟
بیایید بدون تعارف حساب کنیم. در سه چهار سال اول کار، تقریباً تمام رشد تو از یک منبع میآید: دانستن چیزهای بیشتر. زبان را بهتر یاد میگیری، framework را عمیقتر میفهمی، دیتابیس را میشناسی، الگوهای همزمانی را بلد میشوی. هر واحد دانشِ اضافه، مستقیماً به کیفیت کارت اضافه میکند. این دوره لذتبخش است چون رابطهاش خطی و منصفانه است.
بعد یک اتفاق میافتد: منحنی صاف میشود. نه چون یادگیری تمام شده — همیشه چیز بیشتری برای دانستن هست — بلکه چون گلوگاه عوض شده. الان کارهایی که میتوانی انجام دهی، محدود به دانش تو نیست؛ محدود به این است که چند نفر حاضرند به قضاوت تو اعتماد کنند، چقدر میتوانی کاری را که در سرت روشن است در سر بقیه هم روشن کنی، و آیا کسی حاضر است یک مسئلهٔ مبهم و بیصاحب را به تو بسپارد یا نه.
تصور کن یک خودرو با موتور ۵۰۰ اسببخار در یک کوچهٔ باریک شهری. اضافه کردن ۲۰۰ اسببخار دیگر عملاً هیچ سرعتی اضافه نمیکند، چون گلوگاه موتور نیست — جاده است. دانش فنی موتور توست. ارتباط، اعتماد و قضاوت، جادهاند. مهندسی که سالها فقط موتور را بزرگتر میکند و از جاده غافل است، احساس میکند «سخت کار میکنم ولی جایی نمیروم» — و دقیقاً درست احساس میکند.
مهم است این را بهعنوان یک انتقاد از خودت نخوانی. کسی به تو نگفته بود. برنامهٔ درسی دانشگاه، دورههای آنلاین و حتی مصاحبههای فنی، همه موتور را اندازه میگیرند. جاده را کسی درس نمیدهد و بعد در ارزیابی عملکرد، با جملات مبهمی مثل «باید بیشتر دیده شوی» یا «انتظار داریم impact بیشتری داشته باشی» به آدمها بازخورد میدهند که هیچ دستورالعمل عملی در آن نیست. هدف این فصل این است که همان جملات مبهم را به کارهای مشخص ترجمه کند.
«سنیور» دقیقاً یعنی چه؟
اگر نردبانهای شغلی مهندسی را در سازمانهای مختلف کنار هم بگذاری، با اسمهای متفاوت، تقریباً همه حول چهار محور یکسان میچرخند: دامنه (scope)، استقلال (autonomy)، ضریبدهی به دیگران و تحمل ابهام. هیچکدام دربارهٔ حجم دانش نیستند.
| محور | مهندس در حال رشد | مهندس سنیور | معنی عملی روزمره |
|---|---|---|---|
| دامنه | یک task یا یک کلاس | یک سیستم یا یک جریان کاری end-to-end | «این service مال کیست؟» جواب اسم توست |
| استقلال | مسئله برایش تعریف میشود | مسئله را خودش پیدا و تعریف میکند | کسی لازم نیست کارت را بشکند به قطعات کوچک |
| ضریبدهی | خروجی = کار خودش | خروجی = کار خودش + بهبود کار بقیه | review، mentor، ابزار، مستند، الگو |
| ابهام | با نیازمندی روشن خوب کار میکند | با نیازمندی مبهم شروع میکند و روشنش میکند | «معلوم نیست چه میخواهند» برایش شروع کار است نه بهانه |
| ریسک | ریسک فنی را میبیند | ریسک فنی را به زبان کسبوکار ترجمه میکند | مدیر بدون دانستن جزئیات، تصمیم درست میگیرد |
یک سنجهٔ ساده و بیرحم برای اینکه ببینی کجای این نردبان ایستادهای: چقدر انرژی از دیگران لازم است تا کار تو به سرانجام برسد؟ اگر برای هر کار باید یک نفر مسئله را برایت بشکند، وسطش چک کند و آخرش جمعش کند، تو مصرفکنندهٔ ظرفیت تیمی. اگر کاری را میگیری و آنچه برمیگردانی از کاری که تحویل گرفتی کاملتر است — مستند شده، تست شده، ریسکهایش اعلام شده — تو تولیدکنندهٔ ظرفیت تیمی. عنوان شغلی دیر یا زود خودش را با این عدد تطبیق میدهد.
یک الگوی شکست تکراری: مهندسی که واقعاً از همه فنیتر است، اما هیچکس دوست ندارد با او کار کند. reviewهایش تحقیرآمیز است، در جلسه حرف دیگران را قطع میکند، مستند نمینویسد چون «کد خودش گویاست»، و هر مخالفتی را یک نبرد میبیند. این آدم معمولاً برای مدتی نگه داشته میشود چون ارزشمند است — و بعد در اولین بازسازی سازمانی حذف میشود، چون هزینهاش روی بقیه از ارزش فردیاش بیشتر شده. مهارت فنی بالا هیچوقت مصونیت نمیآورد؛ فقط تاریخ انقضا را عقب میاندازد.
دو مسیر: IC و مدیریت
از یک سطحی به بعد (معمولاً بعد از سنیور) مسیر دوشاخه میشود. خیلیها این تصمیم را ناخودآگاه و بر اساس حقوق میگیرند و بعد چند سال ناراضیاند. بهتر است آگاهانه انتخابش کنی.
| مسیر IC (Staff/Principal) | مسیر مدیریت (EM/Director) | |
|---|---|---|
| اهرم اصلی | تصمیم فنی درست در نقاط پرریسک | ساختن و نگهداشتن تیمی که تصمیم درست بگیرد |
| واحد کار | معماری، استاندارد، سیستم بحرانی | آدم، فرایند، اولویت، بودجه |
| منبع رضایت | حل مسئلهٔ سخت، ساختن چیز ماندگار | رشد کردن آدمها، درست شدن یک تیم خراب |
| بدترین روز کاری | جلسههای پشتسرهم و صفر ساعت تمرکز | گفتگوی عملکرد سخت، خبر بد به تیم |
| ریسک شخصی | فاصله گرفتن از قدرت تصمیمگیری سازمانی | زنگزدن مهارت فنی، بیبازگشت شدن مسیر |
| اشتراک هر دو | نوشتن، تأثیرگذاری، مدیریت انتظار، تعارض | همانها — دقیقاً همانها |
نکتهٔ مهم و آرامشبخش: تمام محتوای این فصل برای هر دو مسیر لازم است. اگر IC بمانی، اینها ابزار تأثیرگذاریات هستند؛ اگر مدیر شوی، اینها شغل تو هستند. پس هیچ ساعتی که اینجا صرف میکنی، به یکی از دو مسیر قفل نمیشود.
«من این وضعیت را دیدهام و اولش برایم ناعادلانه به نظر میرسید. یک همکار داشتیم که بیشک عمیقترین دانش فنی تیم را داشت — هر باگ پیچیدهای که به بنبست میخورد، آخرش دست او حل میشد. اما وقتی نگاه کردم که خروجی تیم چطور شکل میگیرد، تفاوت را فهمیدم. کار او در سطح یک task تمام میشد: باگ حل میشد، هیچکس نمیفهمید چرا خراب شده بود، هیچ تستی اضافه نمیشد و دفعهٔ بعد باز خودش لازم بود. من از نظر فنی جلوتر از او نبودم، اما هر بار که چیزی را حل میکردم یک صفحه مینوشتم که چه شد، یک تست میگذاشتم که جلوی تکرارش را بگیرد، و در جلسهٔ تیمی سه دقیقه توضیح میدادم. نتیجه این شد که ریسکِ »فقط او بلد است« دور کارهای من نبود. وقتی سازمان میخواهد یک سیستم مهم را به کسی بسپارد، دنبال بیشترین دانش نمیگردد — دنبال کمترین ریسک میگردد. من در تصمیمی که گرفته شد، گزینهٔ کمریسکتر بودم. بعداً همین را با خود او هم در میان گذاشتم، بدون قضاوت، صرفاً بهعنوان چیزی که خودم یاد گرفته بودم.»
۲) نوشتار: پرقدرتترین مهارت یک مهندس
اگر فقط یک بخش از این فصل را جدی بگیری، همین باشد. دلیلش ساده و کمّی است: هر چیزی که میگویی یک بار و برای چند نفر شنیده میشود؛ هر چیزی که مینویسی میتواند دهها بار، توسط آدمهایی که هرگز ندیدهای، در ماههایی که آنجا نیستی خوانده شود. نوشتن تنها راهی است که تأثیر تو از حضور فیزیکیات جدا میشود.
و یک واقعیت ناخوشایند: در بیشتر سازمانها، تصمیمگیرندگان کد تو را نمیخوانند. آنها نوشتههای تو را میخوانند. کیفیت نوشتار تو، در عمل، رابط کاربریای است که سازمان از طریق آن با کیفیت مهندسی تو تماس میگیرد.
یک API خوب طراحی نمیکنی که «همه چیز را در بر بگیرد»؛ طراحی میکنی که مصرفکننده با کمترین تلاش به آنچه لازم دارد برسد. متن هم همین است. خواننده مصرفکنندهٔ توست و توجهش منبع کمیابش. هر جملهای که مجبورش میکنی بخواند و نتیجهای برایش ندارد، مثل یک فیلد اضافی و گیجکننده در response توست.
۲.۱ پیامی که آدمِ پرمشغله واقعاً روی آن اقدام میکند
قاعدهٔ اول: نتیجه اول (bottom line up front). اول بگو چه میخواهی، بعد توضیح بده چرا. اکثر مهندسها برعکس مینویسند، چون در ذهنشان دارند مسیر رسیدن به مسئله را بازگو میکنند؛ اما خواننده ابتدا باید بداند «این به من چه ربطی دارد و باید چه کار کنم».
قاعدهٔ دوم: یک درخواست در هر پیام. اگر سه چیز میخواهی، احتمال زیاد به یکی جواب میگیری و دو تا گم میشود.
قاعدهٔ سوم: مهلت و پیامدِ بیجوابی را بنویس. آدم پرمشغله بر اساس ضربالاجل اولویتبندی میکند، نه بر اساس ترتیب رسیدن پیام.
❌ نسخهٔ بد:
سلام. امیدوارم خوب باشی. داشتم روی سرویس گزارشگیری کار میکردم و
متوجه شدم که کوئریها روی جدول تراکنشها خیلی کند شدهاند و فکر
میکنم شاید به خاطر ایندکس باشد یا شاید حجم داده. البته ممکن است
مربوط به تغییرات هفتهٔ پیش هم باشد. نظرت چیست؟
✅ نسخهٔ خوب:
درخواست: تأیید تو برای اضافه کردن یک index روی transactions(account_id, created_at)
در محیط production — تا پایان روز چهارشنبه.
چرا: زمان پاسخ گزارش روزانه از ۴۰۰ms به ۹ ثانیه رفته و سه تیم گزارش کندی دادهاند.
علت: رشد جدول به ۸۰ میلیون رکورد، بدون index مناسب برای این الگوی کوئری.
ریسک: ساخت index حدود ۶ دقیقه طول میکشد، بهصورت CONCURRENTLY انجام میشود
و writeها بلاک نمیشوند. حدود ۲GB فضای دیسک اضافه میکند.
اگر تا چهارشنبه جواب نگیرم: با فرض تأیید پیش میروم و در کانال تیم اعلام میکنم.
آن خط آخر — «اگر جواب نگیرم چه میکنم» — یکی از مؤثرترین جملههایی است که یاد میگیری. هم فشار تصمیمگیری را از دوش خواننده برمیدارد، هم تو را از حالت «منتظر ماندم» خارج میکند.
پیامهایی که با «سلام، وقت داری؟» شروع میشوند و بعد منتظر جواب میمانند، ظاهراً مؤدبانهاند اما در عمل هزینه را به گیرنده منتقل میکنند: او باید حواسش را جمع کند، جواب بدهد، منتظر بماند، و تازه بعد بفهمد موضوع چه بوده. اگر تیم توزیعشده باشد، همین رفتوبرگشت میتواند یک روز کاری را بخورد. مؤدب بودن یعنی همهچیز را در یک پیام بگذاری، نه اینکه با احوالپرسی شروع کنی.
۲.۲ سؤال فنیای که آدمها دوست دارند جوابش را بدهند
پرسیدن سؤال ضعف نیست؛ بد پرسیدن است که به تو ضربه میزند. یک سؤال بد سه پیام میفرستد: «وقتم را برای فکر کردن نگذاشتم»، «انتظار دارم تو مسئله را از اول بسازی» و «اگر جوابت اشتباه بود هم متوجه نمیشوم». سؤال خوب دقیقاً برعکس.
ساختار چهارتایی که تقریباً همیشه جواب میدهد:
[هدف] دارم تلاش میکنم فایلهای آپلودشده را قبل از ذخیره در object storage
اسکن کنم تا فایلهای خراب رد شوند.
[تلاش] اینها را امتحان کردم: (۱) خواندن کامل در حافظه — روی فایل ۲GB با
OutOfMemoryError خوابید. (۲) استفاده از stream با buffer ۸KB — کار
میکند ولی برای فایل بزرگ ۹۰ ثانیه طول میکشد و timeout میخوریم.
[خطای دقیق] در حالت دوم، لاگ سرویس: "Read timed out after 60000 ms" در
متد StorageClient.upload، خط ۱۴۲.
[درخواست مشخص] سؤالم این است: آیا الگوی درست اینجا اسکن غیرهمزمان بعد از
آپلود است یا باید timeout را برای این مسیر بالا ببریم؟ اگر نمونهای
از الگوی مشابه در سرویس دیگری داریم، همان کافی است.
توجه کن که این متن برای گیرنده چه میکند: در ۳۰ ثانیه میفهمد مسئله چیست، میداند چه راههایی رفته شده (پس تکراری پیشنهاد نمیدهد)، خطای دقیق را میبیند (پس حدس نمیزند)، و میداند دقیقاً چه از او خواسته شده (پس میتواند در دو جمله جواب بدهد).
شایعترین شکست در سؤال پرسیدن این است که بهجای مسئلهٔ اصلی، دربارهٔ راهحلی که خودت انتخاب کردهای سؤال میکنی. مثلاً میپرسی «چطور در این regex کاراکتر آخر را بگیرم؟» در حالی که مسئلهٔ واقعی گرفتن پسوند فایل بوده و اصلاً regex لازم نبوده. همیشه یک جملهٔ «هدف نهاییام این است که…» به سؤالت اضافه کن. این یک جمله، بارها تو را از یک ساعت کار در مسیر غلط نجات میدهد.
قبل از پرسیدن باید تلاش کرده باشی: لاگ را خوانده باشی، مستند رسمی را نگاه کرده باشی، در repo جستجو کرده باشی. اما این قانون یک سر دیگر هم دارد که کمتر گفته میشود: تلاش بیپایان هم اشتباه است. یک سقف زمانی بگذار (مثلاً ۴۵ دقیقه برای یک مسئلهٔ متوسط، یا دو ساعت برای چیزی که هیچکس دیگری بلد نیست) و وقتی رسیدی، بپرس. کسی که سه روز روی چیزی گیر میکند که همکارش در دو دقیقه جواب میداد، مستقل نیست — فقط پرهزینه است. در بازخورد عملکرد، این را «قضاوت ضعیف» مینامند، نه «سختکوشی».
۲.۳ سند طراحی (design document) که واقعاً مرور میشود
بیشتر مهندسها فکر میکنند سند طراحی یعنی «توضیح چیزی که میخواهم بسازم». نه. سند طراحی یک ابزار تصمیمگیری است. هدفش این است که قبل از نوشتن کد، اختلافنظرها روی متن اتفاق بیفتد نه روی کدی که سه هفته وقت برده. اگر سند تو باعث نشود کسی سؤال سختی بپرسد، سند شکست خورده — حتی اگر همه تأییدش کنند.
تفاوت بنیادی با مستندسازی:
| سند طراحی | مستندات | |
|---|---|---|
| زمان نوشتن | قبل از ساخت | بعد از ساخت |
| مخاطب | تصمیمگیرندگان و همکاران فنی | کسی که میخواهد استفاده کند یا نگهداری کند |
| محور | چرا این گزینه، نه گزینههای دیگر | چطور کار میکند و چطور استفاده میشود |
| سرنوشت | بعد از تصمیم منجمد میشود (تاریخچه) | باید زنده بماند و بهروز شود |
| موفقیت یعنی | بحث سازنده و تصمیم شفاف | خواننده بدون پرسیدن کارش راه بیفتد |
قالب کوتاه و واقعی:
# طراحی: انتقال اعلانها به یک سرویس مستقل
## مسئله
اعلانها الان داخل سرویس سفارش تولید و ارسال میشوند. سه پیامد:
۱) کندی provider اعلان، ثبت سفارش را کند میکند (P99 از ۲۰۰ms به ۱.۸s در حادثهٔ ماه پیش).
۲) هر تیمی که اعلان لازم دارد باید در سرویس سفارش کد بزند — ماه گذشته ۴ PR از ۳ تیم.
۳) تست کردن سناریوهای اعلان نیازمند بالا آوردن کل سرویس سفارش است.
## محدودیتها
- بدون downtime برای ثبت سفارش.
- تیم دو نفر و شش هفته وقت دارد.
- الزام حسابرسی: هر اعلان ارسالشده باید تا ۹۰ روز قابل ردیابی باشد.
- زیرساخت پیامرسان موجود است و تیم با آن آشناست.
## گزینههایی که بررسی شد
الف) استخراج به سرویس مستقل با ارتباط رویدادمحور.
+ جداسازی کامل خطا، مقیاسپذیری مستقل، مالکیت روشن.
− سرویس جدید یعنی on-call جدید، هزینهٔ عملیاتی، پیچیدگی eventual consistency.
ب) نگه داشتن در همان سرویس ولی غیرهمزمان کردن با صف داخلی.
+ کمهزینهترین، دو هفته کار، بدون سرویس جدید.
− مالکیت هنوز مبهم است، مشکل «هر تیم در کد ما PR میزند» حل نمیشود.
ج) استفاده از سرویس اعلان آماده از بیرون.
+ سریعترین.
− الزام حسابرسی داده را از کنترل ما خارج میکند. رد شد.
## تصمیم
گزینهٔ (ب) در فاز اول، با مرز ماژول تمیز، و گزینهٔ (الف) در فاز دوم اگر
تعداد مصرفکنندگان از سه تیم بیشتر شد. دلیل: محدودیت ششهفتهای، و اینکه
مرز ماژول تمیز، مهاجرت بعدی را ارزان نگه میدارد.
## ریسکها و آنچه ما را غافلگیر میکند
- اگر صف پر شود، اعلانها با تأخیر میروند. کاهش: alert روی عمق صف بالای ۱۰هزار.
- ترتیب اعلانها تضمین نمیشود. با تیم محصول چک شد — قابل قبول است.
- اگر فاز دوم هرگز اتفاق نیفتد، وضعیت از امروز بدتر نیست ولی بهتر هم نیست.
## چه چیزی این تصمیم را باطل میکند
اگر تا پایان فصل بعد بیش از سه تیم مصرفکننده شوند، فاز دوم اجباری است.
بخش آخر — «چه چیزی این تصمیم را باطل میکند» — چیزی است که سندهای سطح سنیور را از بقیه جدا میکند. یعنی تو نهفقط تصمیم گرفتهای، بلکه شرایط بازبینیاش را هم از قبل مشخص کردهای.
هیچوقت یک سند مهم را مستقیم برای ده نفر نفرست. اول به یک یا دو نفر که احتمالاً بیشترین مخالفت را دارند بهصورت خصوصی بده و بگو: «هنوز خام است، میخواستم قبل از اینکه برای بقیه بفرستم نظرت را بدانم.» دو اتفاق میافتد: ضعفهای آشکار قبل از دیده شدن جمعی حذف میشوند، و آن آدم دیگر در جلسه احساس نمیکند تصمیم بدون او گرفته شده. بیشترِ اجماع، خارج از جلسه ساخته میشود.
نمودار زیر چرخهٔ عملی مرور یک سند طراحی را نشان میدهد · The practical review cycle of a design document.
flowchart TD
A[Draft: problem and constraints only] --> B[Share with 1-2 likely dissenters]
B --> C{Problem statement agreed?}
C -- No --> A
C -- Yes --> D[Add options and trade-offs]
D --> E[Broad review with a deadline]
E --> F{Blocking concerns raised?}
F -- Yes --> G[Address in writing, update options]
G --> E
F -- No --> H[Decision recorded and frozen]
H --> I[Write ADR and link from code]
اگر بنویسی «هر وقت فرصت کردید نگاهی بیندازید»، هیچکس نگاه نمیکند و دو هفته بعد هنوز منتظری. همیشه بنویس: «لطفاً تا پنجشنبه ساعت ۱۲ نظر بدهید؛ بعد از آن فرض میکنم موافقید و شروع میکنم.» این نه بیادبی است نه فشار — این احترام به وقت خودت و جلوگیری از فلج شدن تصمیم است.
۲.۴ ثبت تصمیم معماری (ADR)
ADR برادر کوچک و ماندگارِ سند طراحی است. سند طراحی میتواند بلند باشد و بعد فراموش شود؛ ADR یک فایل کوتاه است که کنار کد زندگی میکند و به یک سؤال جواب میدهد: «چرا اینجوری است؟» — همان سؤالی که مهندس بعدی، شش ماه بعد، با عصبانیت میپرسد.
قالب کلاسیک چهاربخشی، که تقریباً همهجا با کمی تفاوت همین است:
# ADR-0012: استفاده از idempotency key در API پرداخت
وضعیت: پذیرفتهشده — ۱۴۰۴/۰۵/۱۲
(وضعیتهای ممکن: پیشنهادی | پذیرفتهشده | منسوخ | جایگزینشده با ADR-XXXX)
## زمینه
کلاینت موبایل در شبکهٔ ضعیف، درخواست پرداخت را دوباره میفرستد. در سه ماه
گذشته ۱۴ مورد برداشت دوتایی گزارش شده که همه دستی برگردانده شدهاند.
تشخیص تکراری بودن از روی مبلغ و زمان قابل اتکا نیست چون کاربر ممکن است
واقعاً دو پرداخت یکسان انجام دهد.
## تصمیم
هر درخواست POST /payments باید هدر Idempotency-Key با یک UUID تولیدشده
توسط کلاینت داشته باشد. سرور نتیجهٔ اولین اجرا را ۲۴ ساعت نگه میدارد و
برای کلید تکراری همان پاسخ را برمیگرداند. نبود هدر با 400 رد میشود.
## پیامدها
+ برداشت دوتایی ناشی از retry شبکه حذف میشود.
+ کلاینت میتواند با خیال راحت retry کند — منطق resilience سادهتر میشود.
− همهٔ کلاینتها باید بهروزرسانی شوند. نسخهٔ قدیمی موبایل تا ۹۰ روز معاف است.
− نیاز به یک storage با TTL برای کلیدها و پایش نرخ برخورد کلید.
هر انتخابی ADR لازم ندارد. معیار عملی: اگر تغییر دادنِ این تصمیم شش ماه بعد چند هفته کار میبرد یا چند تیم را درگیر میکند، ADR بنویس. اگر با یک PR برمیگردد، ننویس. کتابخانهای از ۲۰۰ ADR که ۱۹۰تایش دربارهٔ نامگذاری متغیر است، کسی را نجات نمیدهد. و قاعدهٔ طلایی: ADR پذیرفتهشده را ویرایش نکن — ADR جدید بنویس و قبلی را «جایگزینشده» علامت بزن. ارزش اینها در تاریخچه است.
«ما یک لایهٔ cache روی سرویس جستجوی محصولات گذاشتیم و من طرفدار اصلیاش بودم. استدلالم درست بود — بار روی دیتابیس زیاد بود و cache آن را ۷۰ درصد کم کرد. چیزی که درست ندیده بودم این بود که تیم محصول در همان فصل قرار بود قیمتها را روزانه چند بار تغییر دهد. دو هفته بعد از انتشار، شکایت شروع شد که قیمت روی صفحهٔ محصول با قیمت سبد خرید فرق دارد. علتش invalidation ناقص cache من بود. کاری که کردم این بود: اول در همان روز در کانال تیم نوشتم که این مشکل از تصمیم من میآید و دارم رویش کار میکنم — نگذاشتم کسی وقتش را برای پیدا کردن مقصر بگذارد. بعد بهجای وصله زدن، ADR اصلی را باز کردم و دیدم در بخش »چه چیزی این تصمیم را باطل میکند« چیزی ننوشته بودم. یک ADR جدید نوشتم که قبلی را جایگزین میکرد: cache فقط برای فیلدهای غیرقیمتی، و قیمت همیشه زنده خوانده شود. اثر بلندمدتش این بود که از آن به بعد در هر سند طراحی یک بخش اجباری گذاشتیم برای اینکه چه تغییری در کسبوکار این طراحی را میشکند. آن اشتباه به تیم چیزی یاد داد که یک ماه بحث نظری یاد نمیداد.»
۲.۵ توضیح PR که مرور را سریع میکند
reviewer تو یک آدم پرمشغله است که وسط کار خودش، context را رها کرده و آمده سراغ کد تو. کیفیت مرور مستقیماً به این بستگی دارد که چقدر سریع میتواند وارد ذهن تو شود. توضیح PR این کار را میکند.
## چه چیزی
اعتبارسنجی شمارهٔ حساب مقصد در انتقال وجه، به لایهٔ دامنه منتقل شد.
## چرا
همین اعتبارسنجی در سه جا تکرار شده بود (controller انتقال، job دستهای،
و import فایل) و در job دستهای نسخهٔ قدیمی مانده بود — همان باگی که در
تیکت #4412 گزارش شد.
## چطور مرور کنی
- هستهٔ تغییر در AccountNumberPolicy است — از اینجا شروع کن.
- سه فایل دیگر فقط جایگزینی فراخوانیاند، سریع رد شو.
- تستهای جدید در AccountNumberPolicyTest سناریوهای مرزی را پوشش میدهند.
## آنچه عمداً انجام نشد
فرمت خروجی خطا را دست نزدم تا این PR کوچک بماند — در #4470 دنبال میشود.
## ریسک و بازگشت
تغییر رفتاری ندارد جز اینکه job دستهای حالا سختگیرتر است. اگر فایلهای
ورودی قدیمی رد شدند، با revert همین commit برمیگردد.
یک PR با ۸۰۰ خط تغییر معمولاً تأیید میشود، نه چون خوب است بلکه چون مرورش خستهکننده است. کیفیت بازخورد با اندازهٔ PR بهشدت افت میکند: در ۵۰ خط، نظر دربارهٔ منطق میگیری؛ در ۸۰۰ خط، یا سکوت میگیری یا نظر دربارهٔ نام متغیر. اگر مجبوری تغییر بزرگ بدهی، آن را به مجموعهای از PRهای پشتسرهم بشکن و در توضیح اولی نقشهٔ کل مسیر را بنویس. این احترام به وقت reviewer است و در ضمن بهترین راه برای گرفتن بازخورد واقعی روی طراحی.
۲.۶ commit و گزارش باگ
پیام commit برای آدمی نوشته میشود که شش ماه بعد با git blame روی یک خط عجیب ایستاده و میپرسد «چرا؟». پس بدنهٔ پیام باید چرایی را بگوید، نه چیزی که از diff پیداست.
Reject transfers to closed accounts in batch import
The batch import path used a stale copy of the account-status check that
did not include the CLOSED state, so an import file could create transfers
that later failed at settlement and had to be reversed by hand.
The check now lives in AccountNumberPolicy and is shared by all three entry
points. Behaviour of the online path is unchanged.
Refs: #4412
قواعد عملی: خط اول کوتاه (زیر ۷۲ کاراکتر)، وجه امری («Reject»، نه «Rejected» یا «Rejecting»)، خط خالی، بعد بدنهای که «چرا» را توضیح میدهد. جزئیات ابزاری و جریان کاری در فصل git-workflows است؛ اینجا فقط بخش ارتباطیاش را میگویم.
گزارش باگ هم دقیقاً همان منطق سؤال خوب را دارد، با یک تفاوت: باید قابل بازتولید باشد.
عنوان: انتقال وجه بین دو حساب ارزی یکسان، کارمزد تبدیل ارز میگیرد
محیط: staging، نسخهٔ 2.14.3، مرورگر بیربط (از طریق API هم بازتولید شد)
گامها:
۱) دو حساب با ارز EUR بساز.
۲) POST /transfers با مبلغ 100 از حساب اول به دوم.
۳) پاسخ را نگاه کن.
انتظار: fee برابر 0 چون تبدیلی انجام نمیشود.
واقعیت: fee برابر 1.5 و در فیلد feeType مقدار "FX" برمیگردد.
دامنه: فقط وقتی هر دو حساب غیر از ارز پایه باشند. با دو حساب ریالی رخ نمیدهد.
اثر: مشتری بابت خدماتی که انجام نشده پول میدهد — مسئلهٔ انطباق، نه فقط باگ.
شواهد: trace id مربوطه a7f3-… و لاگ FxRateResolver خط ۸۸.
جملهٔ «دامنه» — یعنی کِی رخ میدهد و کِی نمیدهد — بیشترین کمک را به کسی میکند که باید دیباگ کند، و اغلب نوشته نمیشود.
«یاد گرفتهام سند را دو لایه بنویسم. لایهٔ اول یک بخش کوتاه در ابتدای متن است که هیچ اصطلاح فنی ندارد و به سه سؤال جواب میدهد: چه چیزی الان مشکلساز است، چه پیشنهادی دارم، و اگر انجام ندهیم چه اتفاقی میافتد. همهاش پنج شش خط است. لایهٔ دوم بقیهٔ سند است که کاملاً فنی است و اصلاً سعی نمیکنم سادهاش کنم، چون مخاطبش مهندس است. یک بار سندی نوشتم برای اینکه چرا باید یک وابستگی قدیمی را ارتقا بدهیم. نسخهٔ اولش پر از جزئیات نسخه و breaking change بود و هیچ تصمیمی نگرفتند. بازنویسی کردم و بالایش نوشتم: »این کتابخانه دیگر وصلهٔ امنیتی نمیگیرد. اگر آسیبپذیری جدیدی منتشر شود، ما هیچ راهی جز خاموش کردن سرویس نداریم. کار لازم دو هفتهٔ یک نفر است.« همان هفته تأیید شد. نکته این نبود که سادهاش کردم، این بود که ریسک را به زبان تصمیم ترجمه کردم — هزینه، احتمال، و گزینهها.»
۲.۷ گزارش وضعیت (status update) که اعتماد میسازد
گزارش وضعیت، تنها ابزاری است که با آن آدمهایی که کارت را نمیبینند تصمیم میگیرند چقدر به تو اعتماد کنند. و بیشتر مهندسها آن را خراب میکنند، به یکی از این دو شکل: یا آنقدر مبهم است که چیزی نمیگوید («در حال پیشرفت، همه چیز خوب است»)، یا آنقدر جزئی است که خواننده باید خودش نتیجهگیری کند («ریفکتور کلاس X تمام شد، ۳ تست اضافه شد، دارم روی Y کار میکنم»).
گزارش خوب سه چیز دارد و همان سه چیز کافی است: آیا سر وقتیم؟، چه چیزی تغییر کرد؟، چه چیزی از تو میخواهم؟
پروژهٔ مهاجرت اعلانها — هفتهٔ ۳ از ۶
وضعیت: سبز، ولی با یک نکته.
پیشرفت: مسیر غیرهمزمان تا انتها کار میکند و در staging زیر بار
مصنوعی ۵ برابر تست شد. مصرفکنندهٔ اول (سرویس سفارش) وصل شده.
تغییر نسبت به هفتهٔ قبل: تیم پرداخت هم درخواست اعلان داده. این خارج
از دامنهٔ اولیه است؛ فعلاً نپذیرفتهام و بعد از فاز اول بررسی میکنیم.
ریسک: محیط تست provider اعلان هفتهٔ آینده دو روز خاموش است. اگر جای
دیگری برای تست پیدا نکنیم، تحویل دو روز عقب میافتد — هنوز داخل بافر است.
درخواست: فقط یک مورد — دسترسی به حساب sandbox تیم زیرساخت تا سهشنبه.
سه چیزی که این متن بهطور نامرئی انجام میدهد: (۱) وضعیت را با یک کلمه اعلام میکند تا مدیر لازم نباشد کل متن را بخواند؛ (۲) یک ریسک را قبل از تبدیل شدن به مشکل مطرح میکند — همین یک عادت، بیش از هر چیز دیگری اعتماد میسازد؛ (۳) درخواست را از توضیحات جدا میکند.
بدترین الگوی گزارشدهی این است که هفتهها بگویی همه چیز خوب است و بعد در آخرین هفته اعلام کنی که کار سه هفته عقب است. از دید خودت شاید منطقی باشد («امیدوار بودم جبرانش کنم»)، اما از دید مدیر و ذینفعان، تو یا اوضاع را نمیفهمیدی یا پنهان کردی — و هر دو بد است. یک تأخیرِ زود اعلامشده تقریباً همیشه قابل مدیریت است؛ همان تأخیر، دیر اعلامشده، یک بحران است. قاعده: بدترین خبر باید سریعترین خبر باشد.
قبل از فرستادن، متن را با چشم کسی بخوان که هیچچیز از هفتهٔ تو نمیداند. اسم کلاسها و تیکتها برای او بیمعنی است. اگر جملهای بدون دانستن جزئیات داخلی قابل فهم نیست، یا حذفش کن یا به اثر بیرونیاش ترجمهاش کن: «کش را invalidate کردیم» میشود «قیمتها دیگر با تأخیر نشان داده نمیشوند».
۲.۸ postmortem بدون سرزنش، اما همچنان مفید
بعد از هر حادثهٔ جدی، یک سند نوشته میشود. کیفیت این سند نشان میدهد سازمان بالغ است یا نه — و رفتار تو در آن جلسه نشان میدهد تو بالغ هستی یا نه.
«بدون سرزنش» (blameless) خیلی وقتها بد فهمیده میشود. معنیاش این نیست که «هیچکس اشتباه نکرده» یا «دربارهٔ آدمها حرف نزنیم». معنیاش این است: فرض میکنیم هر کسی با اطلاعات و ابزاری که آن لحظه داشت، منطقیترین کار را کرد؛ پس سؤال درست این است که چرا آن کارِ اشتباه در آن لحظه منطقی به نظر میرسید. این تغییر زاویه، تفاوت بین یک سند مفید و یک دادگاه است.
در صنعت هوانوردی، بعد از هر سانحه جعبهٔ سیاه را میخوانند تا بفهمند سیستم چطور اجازه داد این اتفاق بیفتد — نه اینکه کدام خلبان را جریمه کنند. اگر خلبانها میترسیدند، گزارش نمیدادند، و صنعت کور میشد. در نرمافزار هم دقیقاً همین است: اگر مهندسی که دستور اشتباه را اجرا کرده بترسد، دفعهٔ بعد تا جای ممکن پنهانش میکند و شما ده دقیقهٔ طلایی اول حادثه را از دست میدهید.
# Postmortem: قطعی جزئی سرویس پرداخت — ۹۴ دقیقه
## اثر
۹۴ دقیقه، حدود ۳۱٪ از درخواستهای پرداخت با خطای 503 برگشتند.
تخمین ۲٬۸۰۰ تراکنش ناموفق. هیچ دادهای از دست نرفت.
## خط زمانی (به وقت محلی)
14:02 — استقرار نسخهٔ 3.4.0 روی ۲ نود از ۶ نود (canary).
14:09 — نرخ خطا شروع به رشد. alert فعال نشد چون آستانه روی کل کلاستر بود.
14:31 — گزارش دستی از تیم پشتیبانی. شروع بررسی.
14:52 — علت مشکوک: نشت connection pool در نسخهٔ جدید.
15:12 — تصمیم به rollback.
15:36 — rollback کامل، نرخ خطا عادی.
## علت
نسخهٔ جدید در مسیر خطا، connection را به pool برنمیگرداند. زیر بار
عادی این نشت در حدود ۴۵ دقیقه pool را تمام میکرد.
## چرا زودتر ندیدیم (بخش مهم سند)
- alert نرخ خطا روی میانگین کل کلاستر بود، پس ۲ نود خراب از ۶ نود آن را
زیر آستانه نگه میداشت. سیستم پایش ما canary را نمیدید.
- تست بار موجود فقط مسیر موفق را میزد؛ مسیر خطا هرگز زیر بار نرفته بود.
- مهندس on-call به داشبورد canary دسترسی نداشت — دسترسی ماه قبل عوض شده بود.
## اقدامات (هر کدام با یک مالک و تاریخ)
۱) alert جداگانه بهازای گروه استقرار — تیم زیرساخت — تا ۲۰ام.
۲) اضافه کردن سناریوی خطا به تست بار — تیم پرداخت — تا ۲۷ام.
۳) بازبینی دسترسیهای on-call بهصورت فصلی — مدیر تیم — فصل بعد.
## آنچه خوب کار کرد
- rollback در ۲۴ دقیقه انجام شد چون اسکریپتش از قبل تست شده بود.
- ارتباط با پشتیبانی از همان دقیقهٔ ۳۵ برقرار بود و پیام یکسان داده شد.
بخش «چرا زودتر ندیدیم» ارزشمندترین قسمت سند است، چون همیشه چند لایه دفاعی باید همزمان شکسته شده باشند تا یک حادثه رخ دهد. و بخش «آنچه خوب کار کرد» را حذف نکن؛ اگر فقط شکستها را بنویسی، تیم یاد نمیگیرد کدام سرمایهگذاریها جواب دادهاند.
اگر علت ریشهای یک حادثه در سند نوشته شده «مهندس اشتباهاً دستور را روی production زد»، سند تمام نشده. سؤال بعدی این است: چرا ابزار اجازه داد؟ چرا محیطها شبیه هم بودند؟ چرا تأیید دوم لازم نبود؟ آدمها همیشه اشتباه میکنند — این یک ثابت طبیعت است، نه یافتهٔ تحقیق. سیستمی که با فرض بیخطا بودن آدمها ساخته شده، سیستم خرابی است.
«اولین کاری که میکنم این است که قبل از جلسه، خط زمانی و علت را خودم بنویسم و منتشر کنم. وقتی خودت زودتر و دقیقتر از همه توضیح میدهی چه شد، دو اتفاق میافتد: بحث از مرحلهٔ »چه کسی؟« رد میشود و مستقیم میرود سراغ »چطور جلویش را بگیریم؟«، و اعتماد به تو بهجای کم شدن، بیشتر میشود. در جلسه هم مراقبم که دو کار را نکنم: نه خودم را بیش از حد سرزنش کنم — چون این جلسه را ناراحتکننده میکند و بقیه بهجای تحلیل شروع میکنند به دلداری دادن — و نه دفاعی شوم. تمرکزم را میگذارم روی اینکه چرا سیستم اجازه داد این کد به production برسد: چه تستی نبود، چه alertای نبود، چه بازبینیای این را نگرفت. آخرین بار که این اتفاق افتاد، خروجی جلسه سه اقدام بود که هیچکدامشان »بیشتر دقت کن« نبود، و همان سه اقدام جلوی دو حادثهٔ مشابه در ماههای بعد را گرفت.»
۳) code review بهعنوان یک مهارت اجتماعی
مرور کد، تنها جایی است که در آن بهطور روزمره و مکتوب دربارهٔ کارِ یک همکار قضاوت میکنی. به همین دلیل، بیشترین آسیبِ روابط تیمی و بیشترین فرصتِ ساختن اعتماد، هر دو همینجاست.
محتوای فنی مرور — اینکه دنبال چه چیزی در کد بگردی، بوی بد کد، مرزهای انتزاع، تستپذیری — موضوع فصل clean-code-craft است و اینجا تکرارش نمیکنم. اینجا فقط دربارهٔ چطور گفتنش حرف میزنم، از هر دو طرف میز.
۳.۱ طرف مرورکننده: نقد بدون تخریب
مشکل بنیادی این است که متن، لحن ندارد. جملهٔ «چرا اینجا از stream استفاده نکردی؟» در ذهن تو یک کنجکاوی ساده است و در ذهن خواننده میتواند «تو حتی stream بلد نیستی» خوانده شود. راهحل، نرم کردن مصنوعی همه چیز نیست — راهحل این است که شدت هر نظر را صریح اعلام کنی.
| برچسب | معنی | نمونهٔ جمله | نویسنده باید چه کند |
|---|---|---|---|
blocking: |
تا حل نشود merge نمیکنم | «blocking: این مسیر خطا connection را نمیبندد و زیر بار pool تمام میشود.» | حتماً اصلاح یا استدلال متقابل |
issue: |
مشکل واقعی، ولی میشود بحث کرد | «issue: این تابع دو مسئولیت دارد و تستش را سخت میکند.» | اصلاح یا توافق روی زمان دیگر |
question: |
واقعاً نمیدانم، توضیح بده | «question: اگر ورودی null بیاید اینجا چه میشود؟ شاید جای دیگری گرفته میشود.» | فقط جواب بده |
suggestion: |
بهتر میشود، ولی سلیقهای نیست | «suggestion: میشود این شرط را به یک متد با نام معنادار برد.» | اختیاری، ولی جواب بده |
nit: |
جزئی، اصلاً لازم نیست انجام دهی | «nit: فاصلهٔ اضافه در خط ۲۲.» | آزادانه نادیده بگیر |
praise: |
این خوب بود | «praise: این تست مرزی دقیقاً همان چیزی است که ماه پیش گم کرده بودیم.» | هیچ |
این یک عادت کوچک با اثر بزرگ است. نویسنده در ۱۰ ثانیه میفهمد از ۱۷ نظری که گرفته، دو تا واقعاً جلوی merge را میگیرند و بقیه سلیقه یا کنجکاوی است. بدون این برچسبها، ۱۷ نظر همه به یک اندازه سنگین به نظر میرسند و حسِ «کل کارم رد شد» میدهند.
اگر در دور اول مرور، بیست نظر دربارهٔ نامگذاری و فرمت بگذاری و در دور دوم بگویی «راستی، فکر کنم کل این رویکرد باید عوض شود»، هم وقت نویسنده را سوزاندهای و هم اعتبار خودت را. اول کد را از بالا بخوان: آیا مسئلهٔ درست حل شده؟ آیا مرزها درستاند؟ آیا این تغییر جای درستی است؟ اگر جواب یکی از اینها نه است، فقط همان را بنویس و بقیه را نگه دار. سبک و فرمت هم اصولاً کار formatter و linter است، نه کار آدم.
یک الگوی آسیبزا: مرورکنندهای که از هر PR بهعنوان بهانهای برای نمایش دانستهٔ خودش استفاده میکند — «این را میشد با یک الگوی X خیلی زیباتر نوشت» روی PRای که یک باگ فوری را در سه خط درست کرده. نتیجه این است که آدمها از فرستادن PR به تو طفره میروند، تغییرها را بزرگتر و دیرتر میفرستند، و کیفیت کل تیم پایین میآید. قبل از هر نظر از خودت بپرس: «اگر این نظر را ننویسم، چه چیز بدی اتفاق میافتد؟» اگر جواب «هیچی، فقط سلیقهٔ من نیست» است، یا nit: بزن یا ننویس.
یک PR که سه روز منتظر مرور میماند، هزینههای زنجیرهای دارد: نویسنده context را از دست میدهد، branch از main دور میشود و conflict میگیرد، و مهمتر از همه، مهندس یاد میگیرد که PR بزرگتر بفرستد (چون هزینهٔ انتظار ثابت است). قاعدهٔ عملی که در تیمهای سالم دیده میشود: مرور اولیه در کمتر از یک روز کاری. اگر وقت کامل نداری، حداقل در ۱۵ دقیقه یک نگاه سطحبالا بینداز و بگو «طراحی برایم منطقی است، جزئیات را فردا صبح میبینم» — همین یک جمله نویسنده را از بلاتکلیفی درمیآورد.
۳.۲ طرف نویسنده: گرفتن نقد بدون دفاعی شدن
اولین چیزی که باید بپذیری، ناخوشایند است: مغز تو کد خودت را بخشی از خودت میداند. وقتی کسی مینویسد «این منطق اشتباه است»، همان بخشهایی از مغزت فعال میشود که موقع نقد شخصی. این طبیعی است و با اراده هم از بین نمیرود. کاری که میشود کرد، ساختن یک فاصلهٔ ساختگی است.
سه تکنیک عملی که کار میکنند:
یک) بازنویسی ذهنی جمله. هر نظر را قبل از خواندن، در ذهنت به این شکل ترجمه کن: «کد» بهجای «تو»، و «هنوز» به آخرش اضافه کن. «این تست ناقص است» میشود «این کد هنوز یک حالت مرزی را پوشش نمیدهد». همان اطلاعات، بدون بار.
دو) قانون بیست دقیقه. اگر یک نظر عصبانیات کرد، همان لحظه جواب نده. بیست دقیقه بگذار. تقریباً همیشه بعد از بیست دقیقه یا نظر را درست میبینی، یا میتوانی مخالفتت را بدون تیزی بنویسی. تقریباً هیچ پاسخی در code review آنقدر فوری نیست که ارزش سوزاندن یک رابطه را داشته باشد.
سه) تفکیک «اشتباه است» از «من جور دیگری مینوشتم». برای هر نظر تصمیم بگیر در کدام دسته است. اگر دستهٔ دوم است و هزینهٔ تغییر کم است، بیبحث انجامش بده؛ سرمایهٔ اجتماعیات را برای دستهٔ اول نگه دار. مهندسی که سر هر چیزی میجنگد، وقتی سر چیز مهمی میجنگد شنیده نمیشود.
و وقتی واقعاً باید پافشاری کنی، این ساختار جواب میدهد:
ممنون — نکتهٔ درستی است و اول با تو موافق بودم.
دلیل اینکه اینجوری نوشتم این بود که [محدودیت مشخص: مثلاً این متد
از یک scheduler بدون transaction صدا زده میشود].
با رویکرد پیشنهادی تو، در آن مسیر [پیامد مشخص] پیش میآید — این را
در تست X میشود دید.
اگر جای دیگری این محدودیت را ندیدهام، خوشحال میشوم بدانم و عوضش کنم.
سه ویژگی این متن: با تأیید شروع میشود (نه تعارف، بلکه واقعاً نشان میدهد نظر را خواندهای)، دلیل فنی و قابلبررسی میدهد نه سلیقه، و در را باز میگذارد. کسی که این را میخواند، نه احساس شکست میکند نه احساس نادیده گرفته شدن.
اگر بعد از دو رفتوبرگشت هنوز اختلاف هست، دور سوم تقریباً هرگز حلش نمیکند — فقط لحن را بدتر میکند. آنجا این جمله را بنویس: «فکر میکنم بهاندازهٔ کافی هر دو طرف را نوشتهایم و متن دارد ناکارآمد میشود. بیا ۱۵ دقیقه صحبت کنیم و اگر باز هم نرسیدیم، از [نفر سوم] بخواهیم قضاوت کند. هر کدام را انتخاب کند، من پیش میبرم.» جملهٔ آخر مهم است: از قبل تعهد میدهی که به نتیجه پایبندی، و این کل تنش را از فضا میگیرد.
«اول فرض میکنم نیت خوب است و مسئله را روی خودم میبندم نه روی او. یعنی بهجای اینکه در ذهنم بگویم »این آدم وسواسی است«، فرض میکنم شاید تیم استاندارد نوشتهشدهای ندارد و او دارد استاندارد ذهنی خودش را اعمال میکند. بعد خصوصی و کوتاه با او صحبت میکنم، نه در ترد PR — چون بحث دربارهٔ فرایند، وسط بحث دربارهٔ کد، همیشه بد پیش میرود. جملهام چیزی شبیه این است: »مرورهایت را جدی میگیرم و چند بار جلوی باگ من را گرفتهای. یک چیزی هست که میخواستم بگویم: تشخیص اینکه کدام نظرت جلوی merge را میگیرد و کدام سلیقه است برایم سخت است، و همین باعث میشود دورهای اضافه بزنیم. میشود امتحان کنیم که نظرها را با blocking و nit برچسب بزنیم؟« تقریباً همیشه استقبال میشود، چون طرف مقابل هم نمیخواسته کار را کند کند. اگر موضوع تکرارشونده و کلتیمی باشد، پیشنهاد میدهم یک formatter و linter مشترک بگذاریم تا کل دستهٔ سلیقهای از میز آدمها برداشته شود.»
«حواسم هست که دو هدف دارم که ممکن است با هم تعارض پیدا کنند: کیفیت این تغییر، و اینکه این آدم دفعهٔ بعد بهتر بنویسد و از فرستادن PR نترسد. اگر فقط هدف اول را ببینم، بیست نظر میگذارم و او فردا با ترس کد میزند. کاری که میکنم این است: اول یک چیز واقعی را که خوب بوده مینویسم — نه تعارف، مثلاً اینکه حالت خطا را در نظر گرفته. بعد بهجای اینکه راهحل درست را دیکته کنم، مسئله را به شکل سؤال مطرح میکنم: »اگر فردا یک نوع سوم پرداخت اضافه شود، چند جا باید عوض شود؟« این سؤال او را به همان نتیجهای میرساند که من رسیدهام، و این بار خودش رسیده. اگر دیدم موضوع بزرگتر از یک نظر متنی است، پانزده دقیقه با هم مینشینیم؛ چیزی که در متن ده پیام میشود، حضوری سه دقیقه است. و آخرش، اگر تغییر فوری لازم است، صریح میگویم کدام بخش باید همین حالا درست شود و کدام را میشود در یک تیکت بعدی گرفت.»
۴) جلسه، زمان همزمان، و «نه» گفتن
هر جلسه، هزینهٔ واقعی دارد: اگر شش نفر یک ساعت بنشینند، شش ساعت-نفر از ظرفیت مهندسی سازمان خرج شده — بهعلاوهٔ هزینهٔ پنهانی که کمتر دیده میشود، یعنی تکهتکه شدن روز. یک جلسه در ساعت ۱۱ صبح، ممکن است دو بلوک تمرکز را از بین ببرد، نه یکی.
۴.۱ جلسهای که ارزش هزینهاش را دارد
سه سؤال قبل از هر جلسهای که خودت میگذاری:
۱) تصمیم مشخصی هست که باید گرفته شود؟ اگر جواب «میخواهیم همه در جریان باشند» است، این یک نوشته است نه یک جلسه. ۲) چه کسی واقعاً لازم است؟ هر نفر اضافه، احتمال تصمیم گرفتن را کم میکند. برای تصمیم فنی، بیش از پنج نفر تقریباً همیشه یعنی جلسه به بحث تبدیل میشود. ۳) آیا شرکتکنندهها میتوانند از قبل آماده شوند؟ اگر متن یکصفحهای را از قبل بفرستی، جلسهٔ یکساعته اغلب بیستدقیقهای میشود.
و در دعوت، این سه خط را بنویس:
هدف: تصمیمگیری دربارهٔ اینکه فاز اول اعلانها با صف داخلی برود یا سرویس مستقل.
پیشمطالعه: سند طراحی (لینک) — بخش گزینهها، حدود ۵ دقیقه.
خروجی مورد انتظار: یک تصمیم ثبتشده و یک مالک برای ADR.
قبل از اینکه همه بروند، بلند بگو و بعد مکتوب کن: «پس تصمیم این شد که…، مالکش [اسم] است، و تا [تاریخ].» این سی ثانیه، نیمی از جلسههای تکراری سازمان را حذف میکند. تعداد شگفتآوری از جلسهها با این تصور تمام میشوند که «تصمیم گرفتیم»، در حالی که سه نفر سه برداشت متفاوت دارند. و اگر تو آن کسی باشی که همیشه این خلاصه را مینویسد، بهسرعت تبدیل میشوی به کسی که روایت رسمی تصمیمها را مینویسد — که یکی از ارزانترین شکلهای تأثیرگذاری است.
۴.۲ مخالفت سازنده در اتاق
مخالفت لازم است؛ تیمی که هیچکس در آن مخالفت نمیکند، تیم هماهنگ نیست، تیم ساکت است. اما شکل مخالفت همه چیز را تعیین میکند.
الگویی که تقریباً همیشه کار میکند، سهقدمی است: بازگویی، نگرانی مشخص، پیشنهاد جایگزین.
❌ «این کار نمیکند.»
❌ «من با این موافق نیستم، خیلی پیچیده است.»
✅ «بگذار مطمئن شوم درست فهمیدم: پیشنهاد این است که هر سرویس
جدول خودش را داشته باشد و همگامسازی با رویداد انجام شود. درست است؟
نگرانی من روی گزارشهای مالی است: الان یک join ساده جواب میدهد و
بعدش باید داده را از سه منبع جمع کنیم که با تأخیر همگام میشوند.
پیشنهادم این است که برای این یک مورد یک read model جداگانه
بسازیم، یا اگر خیلی گران است، دربارهٔ حد قابل قبول تأخیر تصمیم بگیریم.»
قدم اول (بازگویی) از همه مهمتر است و از همه بیشتر جا میافتد. وقتی حرف طرف مقابل را با کلمات خودت تکرار میکنی، سه چیز اتفاق میافتد: مطمئن میشوی با نسخهٔ واقعی پیشنهاد مخالفت میکنی نه با نسخهٔ ذهنی خودت؛ طرف مقابل احساس میکند شنیده شده و دفاعی نمیشود؛ و اگر بدفهمی وجود داشته، همانجا حل میشود و کل بحث لازم نمیشود.
جملههایی مثل «تو همیشه راه پیچیده را انتخاب میکنی» یا «این همان اشتباهی است که پارسال هم کردی» حتی اگر از نظر واقعیت درست باشند، بحث را از موضوع فنی به دفاع از هویت میبرند و از آن لحظه هیچکس دنبال جواب درست نیست. جمله را همیشه با «این طراحی…»، «این رویکرد…» شروع کن، نه با «تو…». همچنین دربارهٔ کسی که در اتاق نیست بد نگو — این عادت، سریعتر از هر چیزی اعتماد اطرافیانت را از بین میبرد، چون همه حساب میکنند که در غیاب آنها هم همین کار را میکنی.
۴.۳ حرف زدن وقتی کمسابقهترین آدم اتاقی
این وضعیت واقعی و سخت است. چند چیز که عملاً کمک میکند:
- زودتر حرف بزن، نه کاملتر. هر چه بیشتر صبر کنی تا حرفت «کامل» شود، ورود سختتر میشود. در پنج دقیقهٔ اول یک چیز بگو — حتی یک سؤال — تا سد شکسته شود.
- سؤال، ورودی امن است. «ببخشید، یک سؤال پایهای دارم: وقتی میگوییم consistency اینجا، منظور در سطح یک درخواست است یا کل سیستم؟» این نهتنها ضعف نیست، بلکه معمولاً چیزی است که سه نفر دیگر هم نمیدانستند و نپرسیدند. سؤال روشنکننده، ارزشمندترین مشارکت یک آدم تازه است.
- از داده استفاده کن، نه از اعتبار. وقتی سابقه نداری، «به نظرم این کند است» وزنی ندارد؛ «دیروز روی staging اندازه گرفتم، این مسیر ۹۰۰ms طول میکشد» وزن دارد. عدد، سابقه نمیخواهد.
- اگر حرفت را قطع کردند، آرام و بدون رنجش برگرد: «یک لحظه — جملهام را تمام کنم، دو جمله بیشتر نیست.» اگر جلسه را کسی اداره میکند، معمولاً کافی است یک بار این را بگویی.
- اگر ایدهات را کس دیگری تکرار کرد و اعتبارش را گرفت، مؤدبانه و بدون تلخی مالکیت را برگردان: «خوشحالم که همنظریم — همان چیزی است که چند دقیقه پیش گفتم. اگر موافقی من جزئیاتش را بنویسم و بفرستم.» تلخ نشو، چون در اتاق تو بازنده به نظر میرسی، نه او. یک راهکار تیمی خوب هم هست: در تیمهای سالم، آدمها برای هم این کار را میکنند — «این همان نکتهٔ [اسم] بود، بگذار او ادامه دهد.» اگر تو این کار را برای دیگران بکنی، معمولاً برایت جبران میشود.
«اول بررسی میکنم که واقعاً اطلاعاتی دارم که بقیه ندارند، یا فقط ترجیح متفاوتی دارم — این دو خیلی فرق میکنند. اگر فقط ترجیح است، معمولاً چیزی نمیگویم یا خیلی کوتاه میگویم و رد میشوم. اگر اطلاعات دارم، حتماً میگویم، حتی اگر همه موافق باشند، چون سکوت من در آن لحظه یعنی بعداً که مشکل پیش آمد بگویم »من میدانستم« — که بدترین حالت ممکن است. شکل گفتنش هم مهم است؛ بهجای اینکه تصمیم را رد کنم، ریسکش را روی میز میگذارم: »با تصمیم مشکلی ندارم و اگر همین باشد پیش میبرم. فقط میخواهم مطمئن شوم این ریسک را آگاهانه میپذیریم: با این طراحی، اگر provider کند شود کل مسیر ثبت سفارش کند میشود. اگر میدانیم و قبول داریم، من راضیام.« معمولاً یکی از دو چیز میشود: یا کسی میگوید »این را در نظر نگرفته بودیم« و تصمیم عوض میشود، یا واقعاً آگاهانه پذیرفته میشود و من با خیال راحت پیش میروم. آخرش هم آن نکته را در یادداشت جلسه یا ADR مکتوب میکنم — نه برای مچگیری، بلکه برای اینکه اگر شش ماه بعد اتفاق افتاد، تیم بداند تصمیم آگاهانه بوده و بتواند سریعتر واکنش نشان دهد.»
۴.۴ standup که واقعاً مفید است
standup معمولاً به یک مراسم گزارشدهی به مدیر تبدیل میشود که در آن هر نفر لیست کارهایش را میخواند و بقیه گوش نمیدهند. تست سادهاش این است: اگر همه نوبتی حرف بزنند و هیچکس به حرف دیگری واکنش نشان ندهد، این جلسه هیچ ارزشی تولید نکرده و باید به یک پیام متنی تبدیل شود.
هدف واقعی standup، همگام کردن است نه گزارش دادن. سه چیزی که ارزش گفتن دارند:
- چیزی که بقیه باید بدانند چون روی کارشان اثر دارد («امروز schema را عوض میکنم، اگر روی branch قدیمی هستید rebase کنید»).
- چیزی که در آن گیر کردهای و کسی ممکن است در دو دقیقه بازش کند.
- تغییری در برنامه یا ریسک («فکر میکردم امروز تمام میشود، حالا فکر میکنم پنجشنبه»).
بقیهاش را در ابزار مدیریت کار بنویس. و اگر بحثی بین دو نفر بالا گرفت، جملهٔ نجاتبخش این است: «این بحث خوبی است ولی فقط به دو نفر ربط دارد — بعد از standup ادامه دهیم.»
۴.۵ «نه» گفتن و پس زدن دامنه
مهندسها معمولاً بد «نه» میگویند: یا مطلقاً نه میگویند و منفی به نظر میرسند، یا بله میگویند و بعد تحویل نمیدهند — که خیلی بدتر است. راه سوم این است که بهجای رد کردن، هزینه را شفاف کنی و تصمیم را به صاحب اولویت برگردانی.
❌ «نه، وقت ندارم.»
❌ «باشه، سعی میکنم جا بدهم.» (و بعد هیچکدام سر وقت تمام نمیشود)
✅ «میتوانم این را انجام دهم. الان روی مهاجرت اعلانها هستم که تا
پانزدهم تعهد دادهام. این کار حدود سه روز میبرد، یعنی مهاجرت
میرود هجدهم. اگر این مهمتر است، من مشکلی ندارم — فقط باید با
تیم سفارش هماهنگ کنیم. کدام را ترجیح میدهی؟»
این جمله چند کار همزمان میکند: نشان میدهد همکار هستی نه مانع، تصویر واقعی ظرفیت را میدهد، پیامد را عددی میکند، و مهمتر از همه، تصمیم اولویت را به کسی میدهد که مالک اولویت است. تو دربارهٔ اولویت تصمیم نمیگیری؛ تو دربارهٔ واقعیت گزارش میدهی.
سه شکل دیگر «نه» که خوب کار میکنند:
- نهِ زمانی: «این هفته نه، از دوشنبه بله.»
- نهِ دامنهای: «نسخهٔ کاملش را نه، ولی میتوانم تا پنجشنبه نسخهای بدهم که فقط حالت اصلی را پوشش دهد. اگر جواب داد، بقیهاش را بعداً.»
- نهِ ارجاعی: «من الان نمیتوانم، ولی [همکار] این بخش را بهتر از من میشناسد و میتوانم وصلت کنم.»
وقتی به کاری بله میگویی، همان لحظه به یک کار دیگر — که معمولاً کار متعهدشدهٔ قبلی توست — نه گفتهای، فقط بیصدا. آدمهایی که هرگز نه نمیگویند، بهمرور به آدمهایی تبدیل میشوند که «کارهایش دیر میشود» و خودشان نمیفهمند چرا؛ از دید بیرون، تفاوتی بین «بیش از حد بله گفت» و «کند است» دیده نمیشود. شفاف بودن دربارهٔ ظرفیت، نهتنها بیادبی نیست، بلکه پیششرط قابلاتکا بودن است.
۵) تخمین، تعهد و مدیریت انتظار
هیچ موضوعی به اندازهٔ تخمین، رابطهٔ مهندسها با بقیهٔ سازمان را خراب نمیکند. و ریشهٔ ماجرا یک سوءتفاهم ساده است: مهندس فکر میکند دارد یک پیشبینی میدهد، و شنونده فکر میکند دارد یک تعهد میگیرد.
۵.۱ چرا بد تخمین میزنیم
سه دلیل ساختاری، که هیچکدام ربطی به تنبلی یا بیدقتی ندارند:
خطای برنامهریزی. ذهن آدم برای پیشبینی زمان، سناریوی موفق و بدون مانع را تصور میکند. حتی وقتی میدانی دفعهٔ قبل دو برابر طول کشید، باز هم دفعهٔ بعد خوشبینانه تخمین میزنی. این یک سوگیری شناختی شناختهشده است، نه یک ضعف شخصی.
کارِ نامرئی. وقتی میگویی «سه روز»، معمولاً داری زمان نوشتن کد اصلی را تخمین میزنی. اما بین شروع و «انجام شد» اینها هم هست: خواندن کد موجود، انتظار برای مرور، اصلاح بعد از مرور، تستهای شکسته در CI، هماهنگی با تیم دیگر، مستندسازی، استقرار، و رفع اولین مشکل بعد از استقرار. در بیشتر تیمها، زمان نوشتن کد کمتر از نیمی از زمان کل است.
فشار اجتماعی. وقتی کسی با لحن ناامید میپرسد «واقعاً سه هفته طول میکشد؟»، تخمینها بیاختیار کوچک میشوند. این نه صداقت است نه کمک؛ فقط بدهی را به آینده منتقل میکند.
اگر بپرسند از خانه تا فرودگاه چقدر طول میکشد، جواب صادقانه «۴۰ دقیقه» نیست — «بین ۳۵ دقیقه و ۹۰ دقیقه، بسته به ترافیک» است. و کسی که میخواهد پرواز را از دست ندهد، بر اساس عدد دوم برنامه میریزد. مهندسی هم همین است: یک عدد تنها، اطلاعات را حذف میکند. بازه، اطلاعات را نگه میدارد. کسی که همیشه یک عدد میدهد، در واقع دارد عدمقطعیت را پنهان میکند و ریسکش را به شنونده منتقل میکند بدون اینکه به او بگوید.
۵.۲ تخمینی که آبرو نمیبرد: بازه + فرض
قالب عملی، سه بخش دارد: بازه، فرضها، و آنچه بازه را تنگتر میکند.
تخمین: بین ۲ و ۴ هفته.
فرضهایی که این تخمین بر آنها بنا شده:
- API تیم حسابها همان چیزی است که در مستند آمده و تغییری لازم ندارد.
- محیط تست تا هفتهٔ آینده در دسترس است.
- مهاجرت داده لازم نیست چون رکوردهای قدیمی معافاند.
- من تماموقت روی این هستم و on-call نیستم.
اگر هر کدام از اینها غلط باشد، دوباره تخمین میزنم و همان روز خبر میدهم.
اگر عدد دقیقتر میخواهی: دو روز وقت بده تا مسیر یکپارچهسازی را
با یک نمونهٔ کوچک بزنم؛ بعد از آن میتوانم بازه را به ±۳ روز برسانم.
بخش آخر، کارآمدترین ابزار مذاکرهای است که در اختیار داری: بهجای بحث سر عدد، پیشنهاد میکنی عدمقطعیت را با کار کم کنی. این هم صادقانه است و هم حرفهای به نظر میرسد، چون هست.
| سطح اطمینان | چه زمانی این را میگویی | شکل بیان | استفادهٔ مجاز شنونده |
|---|---|---|---|
| حدس اولیه (t-shirt) | در جلسهٔ ایدهپردازی، بدون بررسی | «کوچک / متوسط / بزرگ — احتمالاً بزرگ» | فقط برای اولویتبندی، نه برنامهریزی |
| بازهٔ خام | مسئله را فهمیدهای، کد را ندیدهای | «بین ۲ و ۶ هفته» | برنامهریزی فصلی |
| بازهٔ مبتنی بر بررسی | یک spike یا نمونهٔ اولیه زدهای | «۳ هفته ±۳ روز، با این فرضها» | تعهد به ذینفع بیرونی |
| تعهد تحویل | کار تا ۸۰٪ پیش رفته | «پنجشنبه آماده است» | اعلام تاریخ به مشتری |
بیشترِ فاجعههای تخمینی از اینجا شروع میشوند که یک «حدس اولیه» در راهرو گفته میشود و دو هفته بعد بهعنوان تاریخ تعهد در یک اسلاید ظاهر میشود. راهحل، ندادن عدد نیست — راهحل، چسباندن برچسب به عدد است: «این یک حدس خام است، دقتش شاید دو برابر یا نصف باشد؛ اگر برایش برنامهریزی میکنی به من بگو تا یک روز رویش کار کنم و عدد بهتری بدهم.» این جمله را حفظ کن.
خیلیها یاد میگیرند تخمین را دو برابر کنند و بافر را پنهان نگه دارند. مشکل این است که اولاً کسی از بافر خبر ندارد پس نمیتواند تصمیم بهتری بگیرد، و ثانیاً کار طوری کش میآید که همان بافر را هم پر کند. بدتر از آن، وقتی کسی متوجه شود که تخمینهای تو همیشه با ضریب اطمینان پنهان میآید، از آن به بعد خودش عددت را نصف میکند — و تو مجبور میشوی بیشتر اغراق کنی. این مسابقه را هیچکس نمیبرد. بافر را بگذار، ولی صریح: «سه هفته کار است و یک هفته بافر برای چیزهایی که هنوز نمیدانیم.»
۵.۳ وقتی واقعیت عوض شد: مذاکرهٔ دوباره
تخمین یک قرارداد نیست؛ یک عکس فوری از دانستهٔ توست در لحظهای که هنوز کمترین اطلاعات را داشتی. وقتی اطلاعات عوض میشود، عدد هم باید عوض شود — و مسئولیت اعلامش با توست، نه با کسی که باید بپرسد.
مهمترین قانون: تأخیر را در لحظهای اعلام کن که فهمیدی، نه در لحظهای که مطمئن شدی. فاصلهٔ بین این دو معمولاً یک هفته است و همان یک هفته تفاوت بین «مدیریتشده» و «بحران» است.
یک خبر که بهتر است زودتر بگویم: فکر میکنم تحویل پانزدهم در خطر است
و احتمالش را حدود ۵۰-۵۰ میبینم.
چه چیزی عوض شد: فرض کرده بودم مهاجرت داده لازم نیست، ولی موقع کار
معلوم شد حدود ۴۰۰ هزار رکورد قدیمی هنوز فعالاند و باید منتقل شوند.
این حدود پنج روز کار اضافه است که در تخمین نبود.
گزینهها، به ترتیب ترجیح خودم:
۱) تاریخ را به بیستودوم ببریم. کامل و بدون بدهی فنی.
۲) پانزدهم برویم ولی فقط برای مشتریان جدید، و مهاجرت را هفتهٔ بعد
انجام دهیم. ریسک: دو مسیر موازی برای یک هفته.
۳) پانزدهم کامل برویم با کمک یک نفر دیگر از تیم برای مهاجرت.
ریسک: کار آن نفر عقب میافتد.
پیشنهاد من گزینهٔ ۲ است. تصمیم با توست — هر کدام را بگویی همان را
اجرا میکنم و امروز شروع میکنم.
این متن، تفاوت یک مهندس سنیور و یک مهندس در حال رشد را در یک صفحه نشان میدهد. مهندس در حال رشد یک مشکل تحویل میدهد؛ مهندس سنیور یک تصمیم با گزینههای قیمتگذاریشده تحویل میدهد. مدیری که این پیام را میگیرد، حتی اگر خبر بد باشد، اعتمادش به تو بیشتر میشود، نه کمتر.
«همان روز اعلام میکنم، حتی اگر هنوز عدد جدیدم دقیق نباشد. تجربهام این است که آدمها با خبر بد کنار میآیند، ولی با غافلگیری نه. جملهای که با آن شروع میکنم این است: »میخواهم زود بگویم که تاریخ در خطر است، هنوز عدد دقیق ندارم ولی تا فردا خواهم داشت.« بعد قبل از اینکه دوباره صحبت کنم، سه چیز را آماده میکنم: چه فرضی غلط از آب درآمد و چرا آن موقع منطقی بود، عدد جدید با بازه، و دو سه گزینهٔ واقعی با هزینه و ریسک هر کدام. مهم است که گزینهها را من بیاورم، چون اگر نیاورم، گزینهها را کس دیگری میسازد و معمولاً بدترینشان انتخاب میشود — مثل »همه آخر هفته کار کنند«. بعد از تحویل هم برمیگردم و مینویسم که این فرض غلط را دفعهٔ بعد چطور زودتر پیدا میکنم؛ در این مورد نتیجهاش این شد که قبل از هر تخمین بزرگ، نیم روز صرف نگاه کردن به دادههای واقعی production میکنم، چون دو بار پشت سر هم »حجم داده« بود که غافلگیرم کرد.»
«اول تلاش میکنم بفهمم پشت این حرف چیست، چون معمولاً یک محدودیت واقعی وجود دارد — یک تعهد به مشتری، یک رویداد، یا یک وابستگی. میپرسم: »چه چیزی باعث شده این تاریخ مهم باشد؟« جوابش کل مکالمه را عوض میکند. بعد صریح میگویم که عدد را نمیتوانم با اراده کم کنم، ولی میتوانیم دامنه را کم کنیم — و اینها دو چیز کاملاً متفاوتاند. معمولاً میگویم: »زمان و دامنه به هم قفلاند. اگر تاریخ ثابت است، بیا با هم تصمیم بگیریم چه چیزی از دامنه بیرون برود. من میتوانم بگویم اگر بخش گزارشگیری و پشتیبانی از فایل قدیمی را حذف کنیم، در نصف زمان قابل انجام است، و آن دو را در فاز بعد اضافه کنیم.« چیزی که هرگز نمیکنم این است که عدد را کم کنم بدون اینکه چیزی عوض شده باشد، چون آن یعنی همان تأخیر را با ماهها اضطراب و یک شکست علنی خریدهام. و اگر با وجود همهٔ اینها تصمیم گرفته شد که تاریخ فشرده بماند، مکتوب میکنم که با چه فرض و چه ریسکی داریم میرویم، بدون لحن طلبکارانه — فقط برای اینکه وقتی تصمیم گرفته میشود، هزینهاش هم مکتوب باشد.»
۶) مدیریت رو به بالا
«مدیریت رو به بالا» اسم بدی دارد و بوی چاپلوسی میدهد، ولی معنیاش کاملاً فنی است: مدیر تو یک سیستم است با ورودی محدود و ظرفیت پردازش محدود؛ کیفیت خروجیاش دربارهٔ تو، تابعی از کیفیت ورودیای است که تو میدهی. اگر ورودی ندهی، او با حدس و شنیدههای دیگران تصمیم میگیرد.
یک واقعیت که خیلیها دیر میفهمند: مدیرت کار تو را نمیبیند. کد را نمیخواند، در ترد PR نیست، نمیداند آن باگی که سه روز پیدایش نمیشد چقدر سخت بود. او یک تصویر بسیار کمرزولوشن از تو دارد که از چند منبع ساخته میشود: چیزهایی که خودت گفتهای، چیزهایی که دیگران دربارهٔ تو گفتهاند، و چند نتیجهٔ قابل مشاهده. سهم تو در آن تصویر، تنها بخشی است که کنترلش دستت است.
۶.۱ چه چیزی را باید بگویی
سه دستهٔ اطلاعات که مدیر واقعاً به آن نیاز دارد، و اکثر مهندسها نمیگویند:
- چیزی که باعث غافلگیری او در جمع میشود. اگر یک تاریخ میلغزد، یک مشتری ناراضی است، یا یک تصمیم فنی جنجالی گرفتهای، مدیر باید قبل از دیگران بداند. قاعدهٔ ساده: مدیرت هرگز نباید خبر بد مربوط به تو را از یک نفر سوم بشنود.
- چیزی که فقط او میتواند حلش کند. دسترسی، اولویت متضاد بین دو تیم، همکاری که پاسخ نمیدهد، ابزاری که بودجه میخواهد. آوردن اینها ضعف نیست؛ این دقیقاً کاری است که او برایش استخدام شده.
- چیزی که تصویر ناقصش را کامل میکند. یعنی کاری که کردهای و اثرش از بیرون دیده نمیشود.
۶.۲ دیده شدن، بدون فخرفروشی
مرز بین «کارم را قابل مشاهده کردم» و «خودم را تبلیغ کردم» جایی است که خیلیها از ترس عبور از آن، اصلاً حرکت نمیکنند. قاعدهٔ عملی برای پیدا کردن آن مرز ساده است: دربارهٔ اثر حرف بزن، نه دربارهٔ تلاش؛ و بهجای صفت، عدد بگذار.
❌ «این هفته خیلی سخت کار کردم و کلی چیز درست کردم.»
❌ «من مشکل کندی گزارشها را حل کردم.» (بیبافت، شبیه ادعا)
✅ «گزارش روزانه که سه تیم ازش شکایت داشتند از ۹ ثانیه به ۴۰۰ms
برگشت. علتش یک index گمشده بود. یک alert هم گذاشتم که اگر
دوباره از ۲ ثانیه رد شد قبل از شکایت کاربر بفهمیم.»
نسخهٔ سوم فخرفروشی نیست چون هیچ صفتی دربارهٔ خودت ندارد؛ فقط واقعیت را گزارش میکند و اتفاقاً همان واقعیت، تأثیرگذار است. اضافه کردن «چه کاری کردم که دیگر تکرار نشود» هم امضای مهندس سنیور است.
هر جمعه ده دقیقه بگذار و در یک فایل ساده بنویس: این هفته چه کردم، چه چیزی یاد گرفتم، چه چیزی گیر داشت. این «دفترچهٔ کار» است و برای خودت است. بعد ماهی یک بار، از آن، موارد قابل ارائه را به یک فایل دوم منتقل کن با ساختار «مسئله ← کاری که کردم ← اثر قابل اندازهگیری ← چه کسی تأیید میکند». این «سند دستاورد» است. شش ماه بعد که جلسهٔ ارزیابی میرسد، تفاوت بین کسی که این فایل را دارد و کسی که ندارد، حیرتآور است: حافظهٔ آدم فقط دو ماه اخیر را نگه میدارد و بقیهٔ سال ناپدید میشود. این کار برای مصاحبه هم بعداً به کارت میآید، ولی هدف اصلیاش این نیست.
۶.۳ جلسهٔ یکبهیک که هدر نمیرود
بیشتر جلسههای یکبهیک به گزارش وضعیت تبدیل میشوند — که اتلاف است، چون وضعیت را میشد نوشت. این جلسه تنها زمانی است که بهطور تضمینی توجه کامل مدیرت را داری. دستور جلسه را تو بنویس؛ اگر تو ننویسی، تبدیل میشود به هر چیزی که آن روز در ذهن او بوده.
دستور جلسهٔ این هفته (۳۰ دقیقه)
۱) موانع — ۱۰ دقیقه
تیم زیرساخت دو هفته است به درخواست دسترسی جواب نداده. میشود
مسیرش را باز کنی؟
۲) بازخورد — ۱۰ دقیقه
میخواستم دربارهٔ سند طراحی اعلانها نظرت را بدانم — بهخصوص
اینکه آیا بخش گزینهها بهاندازهٔ کافی روشن بود یا نه.
۳) مسیر رشد — ۱۰ دقیقه
میخواهم بدانم برای رسیدن به سطح بعدی، از دید تو چه چیزی هنوز
کم است. اگر یک مورد مشخص باشد که باید نشان بدهم، چه میگویی؟
بند سوم را حداقل هر دو ماه یک بار بیاور. جوابی که میگیری اغلب مبهم است («impact بیشتر») — همانجا فشار بده تا مشخص شود: «میشود یک مثال بزنی از کاری که اگر انجام داده بودم، این را نشان میداد؟» مبهم ماندن این جواب، شایعترین دلیل ماندن آدمها در یک سطح برای سالهاست.
یک باور رایج و پرهزینه: «اگر خوب کار کنم، بالاخره میبینند.» در تیمهای کوچک شاید. در سازمان بزرگ، تصمیم ارتقا در جلسهای گرفته میشود که تو در آن نیستی، توسط آدمهایی که بعضیشان تو را ندیدهاند، بر اساس چیزی که مدیرت میتواند دربارهٔ تو بگوید و مدرکی که برایش دارد. اگر مدرک را تو فراهم نکنی، معمولاً فراهم نمیشود. این بیعدالتی نیست، محدودیت اطلاعاتی است — و درمانش هم اطلاعات است، نه صبر.
۶.۴ درخواست پروژه، حقوق یا عنوان
سه اصل مشترک برای هر سه:
یک) درخواست را از ارزیابی جدا کن. جلسهٔ ارزیابی عملکرد، جای درخواست بد است چون تصمیمها معمولاً قبلش گرفته شدهاند. مکالمهٔ درست، دو تا سه ماه قبل از چرخهٔ تصمیمگیری اتفاق میافتد.
دو) شواهد را از قبل بده، نه در لحظه. یک صفحه بفرست و بگو در جلسهٔ بعدی دربارهٔ آن حرف بزنیم. مدیر باید بتواند حرف تو را در جلسهای که تو نیستی تکرار کند؛ کارت این است که آن حرف را برایش آماده کنی.
سه) یک درخواست مشخص بگو، نه یک نارضایتی کلی.
❌ «فکر میکنم حقوقم منصفانه نیست.»
✅ «میخواهم دربارهٔ رسیدن به سطح بعدی صحبت کنم و برنامهای برایش
داشته باشم. سه چیزی که فکر میکنم نشاندهندهٔ آن سطح است:
مهاجرت اعلانها را از تعریف مسئله تا استقرار خودم بردم، دو نفر
تازهوارد را onboard کردم که الان مستقل کار میکنند، و
postmortem حادثهٔ پرداخت و سه اقدام اصلاحیاش مال من بود.
سؤالم این است: از دید تو چه چیزی هنوز کم است؟ اگر شکاف مشخصی
هست، میخواهم شش ماه آینده را روی همان بگذارم.»
توجه کن که این متن، تقاضا نیست؛ دعوت به یک برنامهٔ مشترک است. مدیری که این را میشنود، حتی اگر جواب امروز نه باشد، تبدیل میشود به کسی که برای تو کار میکند. مذاکرهٔ حقوق در زمان استخدام، موضوع دیگری است و در فصل interview-craft آمده؛ اینجا دربارهٔ رشد در همان جای فعلی حرف میزنیم.
«اول فرض نمیکنم که مشکل از انصاف است؛ فرض میکنم مشکل از جریان اطلاعات است، چون در بیشتر مواردی که دیدهام همین بوده. سه کار میکنم. اول، شروع میکنم به نوشتن یک خلاصهٔ کوتاه دوهفتهای برای مدیرم — نه لیست کارها، بلکه چند خط دربارهٔ اثر: چه چیزی بهتر شد و چطور اندازهگیری میشود. دوم، کارم را جایی میگذارم که خودش دیده شود: بهجای اینکه نتیجهٔ یک بررسی را فقط به یک نفر بگویم، یک صفحه در فضای مشترک تیم مینویسم؛ بهجای اینکه یک الگوی خوب را فقط در کد خودم استفاده کنم، یک نمونه و یک توضیح کوتاه میگذارم تا بقیه هم استفاده کنند. سوم، صریح میپرسم. در یکبهیک میگویم »میخواهم مطمئن شوم تصویر تو از کاری که میکنم دقیق است؛ اگر جایی هست که فکر میکنی خروجی من کم بوده، ترجیح میدهم الان بدانم.« معمولاً یکی از دو چیز درمیآید: یا واقعاً خبر نداشته، که با نوشتن حل میشود؛ یا خبر داشته ولی چیز دیگری از من انتظار داشته، که آن هم اطلاعات ارزشمندی است. چیزی که سعی میکنم نکنم، رنجیدن بیصدا است، چون آن نه مشکل را حل میکند نه دیده میشود.»
۷) تأثیرگذاری بدون اختیار رسمی
بیشتر کارهای مهم در مهندسی، از کسی انجام میشود که هیچ اختیار رسمیای روی مجریانش ندارد. تو نمیتوانی به تیم دیگر دستور بدهی، نمیتوانی محصول را مجبور کنی، نمیتوانی به همکارت بگویی الگوی من را استفاده کن. پس یا یاد میگیری بدون اختیار تأثیر بگذاری، یا سقفت همانجاست.
۷.۱ چطور یک پیشنهاد فنی پذیرفته میشود
پیشنهاد خوب در جلسه پذیرفته نمیشود؛ قبل از جلسه پذیرفته میشود. اگر پیشنهادت اولین بار در جلسه شنیده شود، آدمها باید همزمان بفهمند، ارزیابی کنند و تصمیم بگیرند — و ذهن آدم در حالت غافلگیری، محافظهکار است. پس جواب پیشفرض «نه» یا «باید بیشتر فکر کنیم» است.
مسیری که جواب میدهد:
۱) مسئله را قبل از راهحل بفروش. اگر همه قبول ندارند مشکلی هست، بهترین راهحل دنیا هم پذیرفته نمیشود. اول داده جمع کن: چند بار این اتفاق افتاده، چقدر وقت خورده، چند نفر شکایت کردهاند. ۲) یکییکی حرف بزن. با سه چهار نفر کلیدی جداگانه صحبت کن. سؤال جادویی: «این را میخواهم پیشنهاد بدهم؛ بزرگترین اشکالش از دید تو چیست؟» با این یک سؤال، هم نقاط ضعف را قبل از افشای عمومی پیدا میکنی، هم آن آدم را از منتقد به مشارکتکننده تبدیل میکنی. ۳) کوچک شروع کن. یک پیشنهاد کوچک و قابل برگشت، صد برابر راحتتر از یک تغییر بزرگ پذیرفته میشود. «بگذار روی یک سرویس امتحان کنیم و بعد از یک ماه با داده تصمیم بگیریم» تقریباً همیشه بله میگیرد. ۴) راه بازگشت را بگو. ترس اصلی تصمیمگیرنده، گیر افتادن است. اگر بگویی چطور برمیگردیم، ریسک ذهنیاش نصف میشود. ۵) بگذار اعتبارش را دیگران هم داشته باشند. اگر پیشنهاد «مال تیم» شود نه «مال تو»، شانس اجرایش چند برابر است.
سرمایهٔ اجتماعی تو محدود است و با هر بار پافشاری خرج میشود. سالی دو یا سه بار میتوانی واقعاً روی چیزی بایستی و شنیده شوی؛ اگر هر هفته این کار را بکنی، به «آدمی که همیشه مخالف است» تبدیل میشوی و در آن یک موردی که واقعاً حیاتی است، صدایت شبیه بقیهٔ صداهاست. معیار عملی برای انتخاب تپه: آیا این تصمیم گران است برای برگشت؟ آیا به کاربر یا امنیت آسیب میزند؟ آیا سال بعد هنوز مهم است؟ اگر جواب هر سه نه است، نظرت را بگو، ثبتش کن، و رد شو.
۷.۲ انگیزههای آدمهای دیگر را بفهم
بیشتر «مقاومتهای غیرمنطقی» که در سازمان میبینی، از دید طرف مقابل کاملاً منطقیاند. او فقط بر اساس چیز دیگری ارزیابی میشود.
| نقش | بر اساس چه چیزی موفق شمرده میشود | ترس اصلیاش | چطور با او حرف بزنی |
|---|---|---|---|
| مدیر محصول | تحویل ارزش به کاربر در زمان مشخص | تاریخ از دست برود و توضیحی نداشته باشد | هزینه را به تاریخ و دامنه ترجمه کن، گزینه بده |
| تیم تضمین کیفیت | باگی به production نرسد | چیزی از دستش برود و مقصر شود | زودتر درگیرش کن، معیار پذیرش را با هم بنویس |
| عملیات / زیرساخت | پایداری و قابل پیشبینی بودن | تغییر ناگهانی نیمهشب بیدارش کند | برنامهٔ استقرار و rollback را از قبل نشان بده |
| پشتیبانی | کاهش تیکت و کاربر عصبانی | تغییری که موج تیکت بسازد | قبل از انتشار خبر بده و متن آماده بده |
| امنیت | نبود آسیبپذیری و انطباق | استثنایی که کسی ثبت نکند | زود بپرس، تصمیم را مکتوب کن |
وقتی این جدول را درونی کنی، جملهبندیات عوض میشود. به تیم عملیات نمیگویی «این معماری تمیزتر است»؛ میگویی «با این تغییر، وقتی نیمهشب این سرویس خطا بدهد، فقط همان سرویس درگیر است و در داشبورد دقیقاً پیداست کدام قسمت خراب شده».
۷.۳ حرف زدن با غیرمهندسها
قاعدهٔ اصلی: ریسک فنی را به سه چیز ترجمه کن — پول، زمان، یا احتمالِ اتفاق بد. هیچ مدیر غیرفنیای بر اساس «بدهی فنی زیاد شده» تصمیم نمیگیرد، ولی همه بر اساس این سه چیز تصمیم میگیرند.
❌ «کد سرویس سفارش پر از بدهی فنی است و coupling شدیدی دارد.»
✅ «هر تغییری در سرویس سفارش الان بهطور میانگین سه برابر تغییر
مشابه در سرویسهای دیگر طول میکشد، و در سه ماه گذشته دو حادثه
از همانجا آمده. اگر همینطور ادامه دهیم، سرعت تحویل قابلیتهای
جدید در این حوزه هر فصل بدتر میشود. پیشنهادم این است که
۲۰٪ از ظرفیت هر فصل را به این اختصاص دهیم، و هر فصل با عدد
زمان تحویل نشان دهم که جواب داده یا نه.»
و بحث همیشگی «چرا بازنویسی اینقدر گران است؟» — بهترین توضیحی که پیدا کردهام این است: سیستم فعلی، فقط کدی که میبینی نیست؛ چند سال تصمیمهای ریز است که هر کدام برای یک حالت واقعی گرفته شده و بیشترشان جایی نوشته نشده. بازنویسی یعنی همهٔ آن حالتها را دوباره کشف کنی، و تا کشفشان نکردهای نمیدانی وجود دارند. به همین دلیل بازنویسیها معمولاً از تخمین بیشتر طول میکشند و بهترین راه، جایگزینی تدریجی است: هر بار یک تکه بیرون کشیده شود و پشت همان رابط قبلی بنشیند، تا هیچ لحظهای «همه چیز نو و هیچ چیز کار نمیکند» نداشته باشیم.
وقتی فشار تاریخ میآید، وسوسه میشوی بگویی «یا کیفیت یا سرعت». این چارچوب معمولاً بازنده است، چون شنونده سرعت را انتخاب میکند و تو به آدم کندی که بهانه میآورد تبدیل میشوی. چارچوب بهتر، تفکیک بین بدهی آگاهانه و کار ناتمام است: میتوانی بگویی «میتوانیم تستهای خودکار این بخش را به بعد از انتشار موکول کنیم و دستی تست کنیم — این یک بدهی است با هزینهٔ مشخص و من تیکتش را میسازم. اما نمیتوانیم اعتبارسنجی ورودی را حذف کنیم، چون آن یک آسیبپذیری است نه یک بدهی.» با این تفکیک، از بحث احساسی خارج میشوی و مذاکره روی چیزهای مشخص انجام میشود.
وقتی تصمیمی خلاف نظرت گرفته شد و تو نظرت را کامل گفتهای، دو انتخاب داری: صادقانه پشتش بایست، یا آشکارا اعتراض کن. آنچه خرابکاری است، حالت سوم است: بله گفتن و بعد نیمبند اجرا کردن، در جلسههای دیگر طعنه زدن، یا منتظر شکست نشستن. اگر تصمیمی را قبول کردهای، در حضور دیگران از آن دفاع کن، حتی اگر مال تو نبوده — «تصمیم تیم این بود و دلیلش این» نه «من مخالف بودم ولی گفتند». آدمهایی که این را بلدند، سریعتر از همه به اتاق تصمیمگیری دعوت میشوند، چون قابل اتکا هستند.
«یک بار پیشنهاد دادم که لایهٔ دسترسی به داده را یکپارچه کنیم چون سه الگوی متفاوت در کد بود و هر تازهواردی گیج میشد. پیشنهاد رد شد، به این دلیل که آن فصل تعهد بیرونی داشتیم و ظرفیت نبود. اولین کاری که کردم این بود که دلیل رد شدن را دقیق بفهمم — پرسیدم آیا با خود ایده مخالفاند یا با زمانبندیاش. جواب »زمانبندی« بود، که کاملاً موضوع را عوض میکند. پس بهجای رها کردن یا اصرار، سه کار کردم: یک صفحه نوشتم که چه چیزی این تصمیم را در آینده عوض میکند، شروع کردم به جمع کردن داده — اینکه هر تازهوارد بهطور میانگین چقدر طول میکشد تا اولین تغییرش را در آن لایه بدهد — و در کارهای جاری خودم، هر جا دست میبردم همان الگوی هدف را استفاده میکردم تا نمونهٔ عملی وجود داشته باشد. فصل بعد که ظرفیت باز شد، پیشنهاد دوباره مطرح شد و این بار سه دقیقه طول کشید تا تأیید شود، چون هم داده داشت و هم نمونهٔ کارکرده. درسی که گرفتم این بود که »نه« اغلب یعنی »نه، الان نه، با این اطلاعات نه« و کار من این است که بفهمم کدام کلمه از آن جمله را میشود عوض کرد.»
۸) همکاری و تعارض
۸.۱ اختلاف در مقابل تعارض
اختلاف یعنی دو نفر دربارهٔ یک موضوع نظر متفاوت دارند؛ سالم است و اگر نباشد یعنی کسی فکر نمیکند. تعارض یعنی رابطه آسیب دیده و از آن به بعد، موضوع بهانه است. تفاوتشان در یک نشانه دیده میشود: در اختلاف، هر دو دنبال جواب درستاند و اگر دادهای بیاید نظرشان عوض میشود. در تعارض، دادهٔ جدید هیچ چیز را عوض نمیکند.
مکانیزم تبدیل اختلاف به تعارض تقریباً همیشه یکی است: انباشت. یک چیز کوچک آزارت میدهد، چیزی نمیگویی؛ دومی هم، سومی هم؛ و بار چهارم واکنشی نشان میدهی که با اندازهٔ آن اتفاق جور نیست و طرف مقابل گیج میشود. درمانش هم یک چیز است: زود و کوچک مطرح کن.
یک مشکل کوچک حلنشده با همکار، دقیقاً مثل یک میانبر در کد است: امروز ارزان است، فردا بهرهاش را میپردازی. مکالمهای که امروز دو دقیقه طول میکشد، شش ماه دیگر یک جلسهٔ سنگین با حضور مدیر است. تفاوتش با بدهی فنی این است که بهرهاش خیلی سریعتر مرکب میشود.
۸.۲ مطرح کردن یک مشکل، مستقیم و بدون جنگ
الگویی که در عمل جواب میدهد سه بخش دارد: رفتار مشخص و قابل مشاهده (نه صفت و نه تعمیم)، اثری که روی کار گذاشته (نه روی احساس تو بهعنوان اتهام)، و درخواست روشن.
❌ «تو هیچوقت به کسی خبر نمیدهی و همه چیز را خودت تصمیم میگیری.»
✅ «یک چیزی هست که میخواستم زودتر بگویم تا انباشته نشود. دیروز
ساختار جدول سفارشها عوض شد و من صبح فهمیدم، چون build من شکست
و دو ساعت دنبال دلیلش گشتم. درخواستم این است که برای تغییرهای
schema، قبلش یک پیام در کانال تیم بگذاریم. اگر جای بهتری برای
این هماهنگی میشناسی، من با هر روشی موافقم.»
سه نکتهٔ ظریف در این متن: با اعلام نیت شروع میشود («تا انباشته نشود») که به طرف مقابل میگوید این یک حمله نیست؛ بهجای «تو همیشه»، یک رویداد مشخص و قابل بررسی میآورد که نمیشود انکارش کرد؛ و در پایان، مالکیت راهحل را مشترک میکند. و مهمتر از همه: این مکالمه خصوصی است. نقد رفتار در جمع، حتی اگر درست باشد، تقریباً همیشه به دفاع سرسختانه ختم میشود.
۸.۳ استانداردهای متفاوت، و شخصیتهای سخت
همکاری که استانداردش پایینتر است. اول تفکیک کن: آیا نمیداند، یا نمیتواند، یا اولویتش نیست؟ برای «نمیداند» راهحل آموزش است و مرور کد جای خوبی است. برای «اولویتش نیست»، بحث فردی جواب نمیدهد — باید استاندارد را از سلیقهٔ تو به قاعدهٔ تیم تبدیل کنی: تعریف مکتوب «انجامشده»، بررسی خودکار در CI، الزام تست. ابزار و قاعده، بیطرفاند؛ آدم به آدم، هرگز.
همکاری که استانداردش بالاتر است و تو را کند میکند. این هم واقعی است. راهحل، توافق صریح روی سطح کیفیت بهازای نوع کار است: «برای این prototype که دو هفته دیگر دور انداخته میشود، تست واحد کامل لازم نداریم؛ برای سرویس پرداخت، بله.»
همکار سلطهجو در جلسه. بهجای رقابت بر سر فضا، ساختار را عوض کن: از قبل دستور جلسه بفرست، و در جلسه از ابزار ساده استفاده کن — «بگذار قبل از ادامه، نظر بقیه را هم بشنویم» یا «من دو دقیقه لازم دارم که فکرم را کامل بگویم». اگر تکرارشونده است، خصوصی مطرحش کن؛ خیلی از این آدمها اصلاً متوجه نیستند و با یک بازخورد مشخص واقعاً تغییر میکنند.
همکار ساکت. سکوت لزوماً موافقت نیست. مستقیم و بدون فشار دعوتش کن، ترجیحاً با سؤال مشخص نه سؤال باز: «تو بیشتر از همه با این کد کار کردهای — از دید تو این تغییر کجا میشکند؟» و اگر در جمع راحت نیست، نظرش را قبل یا بعد از جلسه بهصورت متنی بگیر.
۸.۴ نردبان escalation
escalation ذاتاً کار بدی نیست؛ زودهنگام یا پرشی انجام دادنش بد است. اصل کلی: همیشه از پایینترین پله شروع کن، در هر پله فرصت واقعی بده، و وقتی بالا میروی طرف مقابل غافلگیر نشود.
نمودار زیر مسیر بالا رفتن از پلهها را نشان میدهد · The escalation ladder, from a direct conversation up to a formal decision.
flowchart TD
A[Direct private conversation] --> B{Resolved?}
B -- Yes --> Z[Write down what you agreed]
B -- No --> C[Second try, in writing, with a clear ask]
C --> D{Resolved?}
D -- Yes --> Z
D -- No --> E[Tell the person you will raise it]
E --> F[Bring in a neutral peer or tech lead]
F --> G{Resolved?}
G -- Yes --> Z
G -- No --> H[Managers decide, with facts not adjectives]
| پله | چه زمانی | چه میگویی | خطای رایج |
|---|---|---|---|
| ۱ — گفتگوی مستقیم | همان هفتهای که اتفاق افتاد | رفتار مشخص + اثر + درخواست | صبر کردن تا انباشته شود |
| ۲ — تکرار مکتوب | وقتی بار اول اثر نداشت | همان حرف، کوتاهتر، با مهلت | تلختر شدن لحن |
| ۳ — اعلام قصد | قبل از بالا بردن | «فکر میکنم باید از X کمک بگیریم» | بالا بردن بدون خبر — این خیانت خوانده میشود |
| ۴ — نظر سوم بیطرف | اختلاف فنی حلنشده | «هر کدام را بگوید، من پیش میبرم» | انتخاب داوری که آشکارا طرفدار توست |
| ۵ — مدیران | اثر روی تحویل یا ایمنی | واقعیت و اثر، بدون صفت | داستانگویی بهجای گزارش |
وقتی به پلهٔ مدیران میرسی، جملهبندی همه چیز را تعیین میکند. «او همکاری نمیکند و کار را عمداً کند میکند» یک اتهام است و مدیر مجبور میشود از کسی دفاع کند. «سه هفته است منتظر بازبینی این تغییر هستم، دو بار پیگیری کردهام و تاریخ تحویل در خطر است — چطور میتوانیم راهش بیندازیم؟» یک واقعیت و یک درخواست است. اولی تو را طرفِ دعوا میکند، دومی تو را حلکنندهٔ مسئله. صفتها را حذف کن؛ فقط تاریخ، تعداد و اثر بگذار.
«اول مطمئن میشوم که مسئله واقعاً کیفیت است، نه تفاوت سلیقه — چند نمونهٔ مشخص جمع میکنم که بشود دربارهٔ آنها حرف زد، مثل باگی که به production رسید یا تغییری که سه بار برگشت خورد. بعد خصوصی و بدون حضور بقیه با خودش حرف میزنم و از موضع کنجکاوی شروع میکنم، نه اتهام: »دیدم در این چند تغییر تست اضافه نشده و بعدش دو باگ برگشت. میخواستم بدانم آیا چیزی سر راهت هست؟« خیلی وقتها جواب چیزی است که اصلاً حدس نمیزدم — فشار زمانی از جای دیگر، ندانستن ابزار تست، یا اینکه فکر میکرده در این بخش تست لازم نیست. اگر مسئله دانش است، مرور کد و یکی دو جلسهٔ کوتاه با هم حلش میکند. اگر مسئله استاندارد مشترک است، آن را از بحث فردی درمیآورم و به سطح تیم میبرم: تعریف مکتوب »انجامشده« و یک بررسی خودکار در CI، تا قاعده بیطرف باشد و به رابطهٔ من و او ربطی نداشته باشد. فقط اگر بعد از دو دور واقعی چیزی عوض نشد و اثرش روی تحویل تیم ادامه داشت، به مدیر میگویم — و قبلش به خودش میگویم که دارم مطرحش میکنم. چیزی که هرگز نمیکنم این است که در جمع تحقیرش کنم یا بیسروصدا کدش را دوباره بنویسم، چون هیچکدام مشکل را حل نمیکند و رابطه را هم خراب میکند.»
۹) یادگیری و رشد بهعنوان یک عادت
۹.۱ انتخاب آگاهانهٔ چیزی که یاد میگیری
اگر بدون برنامه یاد بگیری، بازار برایت انتخاب میکند — و بازار همیشه پرسروصداترین چیز را جلو میآورد، نه ماندگارترین را. یک تفکیک ساده که کل تصمیم را روشن میکند:
| لایه | مثال | نیمهعمر | چقدر وقت بگذار |
|---|---|---|---|
| اصول | همزمانی، شبکه، مدل داده، طراحی، تحلیل عملکرد | دهها سال | بیشترین — اینها سرمایهاند |
| اکوسیستم | زبان و کتابخانههای اصلی، ابزار build، پایگاه داده | ۵ تا ۱۰ سال | زیاد، ولی عمیق نه گسترده |
| ابزار روز | یک framework خاص، یک سرویس ابری خاص | ۲ تا ۴ سال | فقط تا حد نیاز کار |
| مد | آنچه این ماه سر زبانهاست | ماهها | فقط بخوان که بدانی چیست |
معیار عملی: قبل از هر یادگیری بزرگ بپرس «اگر این ابزار سه سال دیگر منسوخ شود، چه چیزی از این یادگیری برایم میماند؟» اگر جواب «هیچ» است، فقط بهاندازهٔ کار یاد بگیر.
۹.۲ وارد شدن سریع به یک codebase بزرگ
خواندن کد از فایل اول تا آخر، ناکارآمدترین روش ممکن است. روشی که جواب میدهد، ورود از رفتار بیرونی به داخل است:
۱) پروژه را بالا بیاور و یک درخواست واقعی بزن. تا وقتی چیزی روی ماشین تو اجرا نشده، هر خواندنی انتزاعی است.
۲) یک مسیر end-to-end را دنبال کن، از نقطهٔ ورود (کنترلر، مصرفکنندهٔ صف، دستور CLI) تا دیتابیس. یک مسیر کامل بیشتر از ده فایل پراکنده یاد میدهد.
۳) مدل داده را بکش. جدولها و روابطشان معمولاً صادقترین سند یک سیستماند؛ کد دروغ میگوید، schema کمتر.
۴) تاریخچه را بخوان. git log روی فایلهای مرکزی، و ADRها. تصمیمهای عجیب معمولاً دلیل داشتهاند.
۵) اولین تغییر کوچک را زود بده — حتی یک اصلاح یکخطی. اولین PR، بیشتر از یک هفته خواندن یاد میدهد چون کل زنجیرهٔ ساخت، تست و استقرار را لمس میکنی.
۶) نقشهات را بنویس و منتشر کن. همانطور که یاد میگیری، یک صفحه بنویس. این هم فهمت را محکم میکند، هم برای نفر بعدی میماند، و هم اولین کار قابل مشاهدهٔ توست.
وقتی چیزی را توضیح میدهی، مغزت مجبور میشود شکافها را پیدا کند — همان جاهایی که فکر میکردی میدانی ولی فقط آشنا بودی. به همین دلیل mentor کردن یک نفر تازهوارد، یا نوشتن یک صفحه دربارهٔ چیزی که تازه یاد گرفتهای، بازدهی یادگیریاش از خواندن یک کتاب دیگر بیشتر است. و mentoring هیچ عنوان رسمیای لازم ندارد؛ از یک مرور کد آموزنده و نیم ساعت جواب دادن به سؤالهای کسی که تازه آمده شروع میشود.
اگر تنها منبع بازخوردت مدیرت باشد، رشدت گروگان کیفیت آن یک نفر است — و مدیرها عوض میشوند. سه منبع دیگر بساز: مرور کد (مستقیمترین و مکررترین بازخورد فنی که میگیری)، دو یا سه همکار که مستقیماً از آنها بازخورد میخواهی، و سنجههای واقعی کارت. و وقتی بازخورد میخواهی، سؤال باز نپرس؛ «نظرت دربارهٔ کارم چیست؟» جواب مؤدبانه میگیرد. بپرس: «در آن جلسه، کدام قسمت توضیح من گنگ بود؟» — سؤال مشخص، جواب مشخص میگیرد.
تشخیص فلات. نشانههایش اینها هستند: ماههاست چیزی ننوشتهای که برایت سخت باشد، در مرورها هیچکس چیزی به تو یاد نمیدهد، و میتوانی کارِ هفتهٔ بعدت را بدون فکر پیشبینی کنی. راحتی لزوماً بد نیست — گاهی دورهای از تثبیت لازم است — ولی اگر بیش از دو سه فصل طول بکشد، مهارتت نسبت به بازار در حال کهنه شدن است. درمانش معمولاً عوض کردن شغل نیست؛ عوض کردن نوع کار است: یک حوزهٔ ناآشنا در همان سازمان، یک مسئلهٔ عملیاتی، یا نقش راهبری در یک پروژه.
۱۰) قضاوت حرفهای و اخلاق
۱۰.۱ پذیرفتن اشتباه، در جمع
این یکی از معدود جاهایی است که رفتار درست، هم اخلاقیتر و هم به نفع توست. الگو ساده است: سریع، مشخص، بدون بهانه، با اقدام.
من باعث قطعی امروز صبح شدم. تغییری که دیروز دادم connection را در
مسیر خطا برنمیگرداند. rollback انجام شد و از ساعت ۱۰:۲۰ سرویس
عادی است. اصلاح را با تست همین امروز میفرستم و در postmortem
مینویسم که چرا تست بار ما این را نگرفت.
معذرت میخواهم از تیم پشتیبانی که صبح سختی داشتند.
آنچه در این متن نیست بهاندازهٔ آنچه هست مهم است: نه توجیه، نه «ولی محیط تست خراب بود»، نه خودزنی نمایشی. یک بار، تمیز، و بعد تمرکز روی اصلاح. آدمی که اینطور اشتباهش را میپذیرد، اعتبارش بیشتر میشود نه کمتر — چون همه فهمیدهاند که اگر چیزی خراب شود، از او خواهند شنید.
۱۰.۲ «نمیدانم» گفتن بدون از دست دادن اعتبار
فرمولش این است: نمیدانم + آنچه میدانم + چطور میفهمم + کِی برمیگردم.
مطمئن نیستم؛ نمیخواهم حدس بزنم چون این عدد قرار است پایهٔ تصمیم شود.
چیزی که میدانم این است که مسیر همگام تا ۲۰۰ درخواست بر ثانیه تست
شده و مشکلی نداشته. برای بالاتر از آن باید اندازه بگیرم.
تا فردا ظهر با عدد واقعی برمیگردم.
مهندسهای باتجربه این را زیاد میگویند و اتفاقاً اعتبارشان از همین میآید؛ چون وقتی همان آدم میگوید «این را میدانم»، حرفش وزن دارد. کسی که هرگز نمیگوید نمیدانم، بهمرور کسی میشود که هیچ حرفش قابل اتکا نیست.
۱۰.۳ نه گفتن به میانبری که به کاربر آسیب میزند
اینجا جایی است که «مهارت نرم» تبدیل به ستون فقرات حرفهای میشود. تفکیک قبلی را به یاد بیاور: بدهی آگاهانه قابل مذاکره است، آسیب قابل مذاکره نیست. لاگ کردن رمز عبور، خاموش کردن اعتبارسنجی برای رد شدن از یک ددلاین، ارسال داده به جایی که مجوزش نیست، یا انتشار چیزی که میدانی دادههای مالی را خراب میکند — اینها دستهٔ دیگریاند.
این کار را انجام نمیدهم و میخواهم دلیلش را روشن بگویم. با این
تغییر، شمارهٔ کارت کاربر در لاگ ذخیره میشود و هر کسی که به لاگ
دسترسی دارد آن را میبیند. این یک انتخاب فنی نیست، یک نقض است.
راهحل جایگزینی که تا فردا آماده میکنم: فقط چهار رقم آخر لاگ شود،
که برای عیبیابی کافی است. اگر تاریخ فشار دارد، من میتوانم
پشتیبانی دستی این هفته را انجام دهم تا فشار کم شود.
سه ویژگی: صریح و بدون تهدید، دلیل به زبان پیامد نه اصول انتزاعی، و یک راه بیرونرفت که به طرف مقابل اجازه میدهد بدون باختن عقبنشینی کند. و اگر با وجود اینها فشار ادامه یافت، مکتوبش کن و بالا ببر — نه بهعنوان تهدید، بلکه چون تصمیمی در این اندازه باید مالک مشخص و ردپای مکتوب داشته باشد.
دسترسی مهندسی به دادههای واقعی، یک امانت است نه یک امتیاز. جستجوی کنجکاوانه در دادههای یک کاربر، برداشتن یک نمونهٔ داده روی لپتاپ شخصی «فقط برای دیباگ»، فرستادن یک لاگ حاوی داده به یک ابزار بیرونی، یا تعریف کردن جزئیات یک حادثه بیرون از سازمان — هر کدام از اینها میتواند در یک دقیقه، اعتبار چند ساله را از بین ببرد و پیامد حقوقی داشته باشد. قاعدهٔ ساده: اگر لازم است دادهای را ببینی، مسیر مجاز و ثبتشدهاش را استفاده کن؛ اگر مسیر مجازی نیست، همین خودش جواب است.
اعتبار حرفهای در طول سالها ساخته میشود و در چند دقیقه از بین میرود. جمع کوچکتر از آن است که فکر میکنی: همکاران امروز، مصاحبهکنندهها و مشتریهای فردا هستند. آنچه واقعاً میماند دو چیز است — اینکه آیا هر چه گفتی انجام دادی، و اینکه با کسی که قدرتی روی تو نداشت چطور رفتار کردی. هیچ پروژهٔ موفقی این دو را جبران نمیکند.
«اول جدا میکنم که این یک ریسک قابل پذیرش است یا یک آسیب. خیلی چیزها ریسکاند و تصمیمشان با کسبوکار است، نه با من — مثلاً کم بودن پوشش تست یک بخش کماهمیت. ولی چیزی که داده یا پول کاربر را در معرض خطر میگذارد، تصمیم من هم هست. اگر دستهٔ دوم بود، ریسک را به زبان پیامد میگویم نه به زبان فنی: چه اتفاقی میتواند بیفتد، با چه احتمالی، و هزینهاش چقدر است — چون تا وقتی طرف مقابل آن را بهصورت »یک نگرانی مهندسی« بشنود، در برابر تاریخ تحویل میبازد. بعد همیشه یک گزینهٔ سوم میآورم؛ آدمها معمولاً وقتی فشار میآورند که فکر میکنند فقط دو راه هست. مثلاً »با پرچم ویژگی فقط برای ۵٪ کاربران منتشر کنیم و تا سهشنبه بخش امن را کامل کنم.« اگر با وجود همهٔ اینها تصمیم به انتشار گرفته شد، همان روز یک ایمیل کوتاه و بدون لحن دفاعی مینویسم که چه ریسکی را داریم میپذیریم و چه کسی آن را پذیرفته، و اگر موضوع امنیتی یا انطباقی باشد تیم امنیت را در جریان میگذارم. این نه مچگیری است نه پوشاندن خودم؛ تصمیم در این اندازه باید مالک داشته باشد و مکتوب باشد.»
۱۱) پایداری: نگه داشتن خودت در بازی
۱۱.۱ فرسودگی، قبل از اینکه دیر شود
فرسودگی شغلی، «خستگی زیاد» نیست. خستگی با یک آخر هفته درست میشود؛ فرسودگی نمیشود. علائم واقعیاش اینهاست و معمولاً به همین ترتیب میآیند:
- بدبینی: هر پروژهٔ جدید از قبل بیفایده به نظر میرسد، و طعنه جای بحث را میگیرد.
- افت کارایی با تلاش ثابت: همان ساعتها را کار میکنی ولی خروجی نصف شده، و برای کارهای ساده هم انرژی شروع کردن نداری.
- بیتفاوتی به کیفیت: چیزی که قبلاً آزارت میداد، حالا برایت مهم نیست. این خطرناکترین نشانه است، چون شبیه آرامش به نظر میرسد.
- نشانههای بدنی و شناختی: خواب بد، تحریکپذیری، فراموشی، ناتوانی در تمرکز حتی در وقت آزاد.
آنچه واقعاً کمک میکند معمولاً «استراحت بیشتر» نیست، بلکه تغییر در سه چیز است: حجم، کنترل و معنا. اگر روی هیچکدام کنترلی احساس نمیکنی، همین را به مدیرت بگو — با مثال مشخص، نه با شکایت کلی: «سه ماه است همزمان روی دو پروژه و on-call هستم و هیچکدام درست پیش نمیرود. میخواهم یکی را کامل زمین بگذارم؛ کدام؟»
اگر شبها و آخر هفتهها جواب میدهی، دو اتفاق میافتد که هیچکدام به نفعت نیست. اول، انتظار بازتنظیم میشود: کاری که دیروز فداکاری بود، امروز حداقل انتظار است، و روزی که جواب ندهی، «تعهدش کم شده» خوانده میشود. دوم، مشکلات ساختاری پنهان میمانند؛ اگر تو هر بار سیستم را با ساعتهای اضافی نگه داری، هیچکس متوجه نمیشود که آن سیستم به یک نفر بیش از حد وابسته است و باید درست شود. قهرمانبازی، کوتاهمدت تشویق میگیرد و بلندمدت هم تو را میسوزاند و هم مشکل را نگه میدارد.
on-call سالم چند شرط ساده دارد که ارزش پافشاری روی آنها را دارد: هر alert که بیدارت میکند باید قابل اقدام باشد (اگر کاری نیست که در آن لحظه بکنی، آن alert باید حذف یا تبدیل به گزارش روزانه شود)، هر بیدارباش شبانه باید در جلسهٔ بعدی مرور شود، بعد از یک شب سخت باید صبحش را بخوابی نه اینکه در جلسه بنشینی، و هر کسی که کد را مینویسد باید نوبت on-call هم داشته باشد — این مؤثرترین انگیزهٔ نوشتن سیستم قابل اتکا است.
۱۱.۲ تمرکز، و کار مؤثر از راه دور
کار مهندسی به بلوکهای پیوسته نیاز دارد؛ چهار ساعت تکهتکهشده با سه جلسه، اصلاً معادل چهار ساعت نیست. چند کار که واقعاً اثر دارد: دو بلوک دو ساعته در تقویم رزرو کن و مثل جلسه با آن رفتار کن؛ جلسهها را در یک نیمهٔ روز جمع کن؛ اعلانها را در بلوک تمرکز خاموش کن و در وضعیتت بنویس کِی برمیگردی؛ و هر روز را با یک تصمیم شروع کن دربارهٔ اینکه «اگر فقط یک چیز امروز تمام شود، چه باشد».
در کار توزیعشده و بین چند منطقهٔ زمانی، قاعده عوض میشود: پیشفرض را ناهمگام بگذار. یعنی هر چیزی مکتوب و قابل بازیابی باشد، تصمیمها در کانال عمومی گرفته شوند نه در پیام خصوصی، و بهجای «یک لحظه وقت داری؟» یک پیام کامل با تمام بافت بفرستی که طرف مقابل هشت ساعت بعد بتواند بدون رفتوبرگشت جوابش را بدهد. کمی «بیشازحد اطلاعرسانی» در حالت دورکاری، تقریباً همیشه درست است — چون آنچه در دفتر بهطور اتفاقی شنیده میشد، از راه دور اصلاً وجود ندارد.
و یک نکتهٔ ظریف: حضور را با مشغول به نظر رسیدن اشتباه نگیر. سریع سبز شدن چراغ آنلاین، هیچ اعتمادی نمیسازد. آنچه اعتماد میسازد، قابل پیشبینی بودن است: پیامها در بازهٔ مشخصی جواب میگیرند، تعهدها سر وقت میرسند، و وضعیت کار بدون پرسیدن معلوم است.
۱۱.۳ نود روز اول در یک نقش جدید
| بازه | تمرکز اصلی | کارهای مشخص | نشانهٔ موفقیت |
|---|---|---|---|
| هفتهٔ ۱ | راهاندازی و آدمها | محیط توسعه کار کند · با هر همتیمی ۳۰ دقیقه گفتگو · واژهنامهٔ اصطلاحات داخلی را برای خودت بنویس | اولین PR کوچک merge شود |
| هفتهٔ ۲–۴ | فهم سیستم | یک مسیر end-to-end را دنبال کن · مدل داده را بکش · ADRها و postmortemهای اخیر را بخوان | بتوانی معماری را روی تخته برای یک تازهوارد بکشی |
| هفتهٔ ۵–۸ | مشارکت واقعی | یک قابلیت کامل از تعریف تا استقرار · شروع مرور کد دیگران · اولین on-call همراه با یک نفر باتجربه | بدون کمک روزانه کار میکنی |
| هفتهٔ ۹–۱۲ | بهبود بردن | یک نقطهٔ درد را که خودت دیدهای برطرف کن · مستند onboarding را با چیزهایی که گم بود کامل کن | تیم برای چیزی به تو مراجعه میکند |
| پایان ۹۰ روز | تنظیم انتظار | با مدیر: چه چیزی خوب رفت، از اینجا به بعد روی چه چیزی تمرکز کنم | یک برنامهٔ توافقشده برای فصل بعد |
تازهوارد بودن یک پنجرهٔ کوتاه است که در آن میتوانی هر سؤالی بپرسی بدون اینکه کسی تعجب کند. از آن استفاده کن و همهچیز را بپرس، بهخصوص «چرا اینطوری است؟». همزمان یک فایل بردار و هر چیزی را که برایت گیجکننده بود بنویس؛ سه ماه بعد این فایل به بهترین بازخورد ممکن دربارهٔ onboarding تیم تبدیل میشود و اولین کار قابل مشاهدهٔ توست. یک هشدار هم: در ماه اول، مشاهده کن و پیشنهاد بده، ولی از «در جای قبلی اینطور میکردیم» بهعنوان استدلال استفاده نکن — تا وقتی دلیل وضعیت فعلی را نفهمیدهای، پیشنهاد تغییر شنیده نمیشود.
«یاد گرفتهام به دو نشانه حساس باشم، چون این دو زودتر از خستگی میآیند. اول، وقتی متوجه میشوم دربارهٔ چیزهایی که قبلاً برایم مهم بود بیتفاوت شدهام — مثلاً کیفیت کدی که خودم مینویسم برایم اهمیتش را از دست داده. دوم، وقتی میبینم واکنش پیشفرضم به یک درخواست جدید، طعنه یا بدبینی است. اگر اینها را دیدم، اول بهجای اینکه سختتر کار کنم، سرعت را کم میکنم و علت را پیدا میکنم؛ معمولاً یکی از این سه است: حجم بیش از ظرفیت، نداشتن هیچ کنترلی روی کار، یا کاری که برایم بیمعنا شده. برای هر کدام کار متفاوتی لازم است. چیزی که یاد گرفتهام حتماً بکنم این است که زود و مشخص با مدیرم حرف بزنم، نه با شکایت کلی بلکه با یک درخواست عملی: »الان روی سه چیز همزمان هستم و هیچکدام درست پیش نمیرود؛ میخواهم یکی را کامل زمین بگذارم — کدام؟« تقریباً همیشه جواب گرفتهام، چون مدیرها معمولاً از حجم واقعی خبر ندارند. و یک مرز ثابت هم برای خودم دارم: خارج از on-call، شبها جواب نمیدهم، چون تجربهام نشان داده هر استثنا به قاعده تبدیل میشود.»
۱۲) یک سیستمعامل برای خودت
هیچکدام از اینها با یک تصمیم بزرگ عوض نمیشود؛ با یک ریتم عوض میشود. اینجا نسخهای است که میتوانی همین هفته اجرا کنی.
روزانه (حدود ۱۵ دقیقه، جز کار اصلی)
- صبح: یک جمله بنویس که «اگر فقط یک چیز امروز تمام شود، چه باشد».
- یک بلوک دو ساعتهٔ بدون اعلان، محافظتشده.
- مرور PRهای منتظر، حداکثر تا آخر روز کاری.
- اگر چیزی لغزید، همان روز بگو — نه فردا.
هفتگی (حدود ۳۰ دقیقه)
- جمعه: ده دقیقه دفترچهٔ کار — چه کردم، چه یاد گرفتم، کجا گیر داشتم.
- یک بهروزرسانی کوتاه برای مدیر یا ذینفعان: وضعیت، تغییر، ریسک، درخواست.
- یک بار «نه» یا «بله ولی با این هزینه» گفتن، بهجای بله بیفکر.
- یک چیز کوچک را برای بقیه بهتر کن: یک صفحه مستند، یک تست، یک ابزار.
ماهانه
- دفترچهٔ کار را به سند دستاورد تبدیل کن: مسئله ← اقدام ← اثر ← شاهد.
- از یک همکار یک بازخورد مشخص بخواه.
- بپرس: این ماه چه چیزی یاد گرفتم که سه سال دیگر هنوز به کارم میآید؟
فصلی
- با مدیر: شکاف تا سطح بعدی چیست، با مثال مشخص.
- یک تپه انتخاب کن که این فصل رویش میایستی — و بقیه را رها کن.
- خودارزیابی کوتاه زیر را دوباره جواب بده.
خودارزیابی (هر بار فقط بله/نه، صادقانه)
۱) آخرین باری که یک ریسک را قبل از تبدیل شدن به مشکل اعلام کردم کِی بود؟ ۲) اگر امروز غیب شوم، آیا کسی میتواند از روی نوشتههای من کارم را ادامه دهد؟ ۳) آخرین تخمینم بازه و فرض داشت، یا یک عدد تنها بود؟ ۴) در سه ماه گذشته، کار چند نفر بهخاطر من بهتر شد؟ ۵) آخرین باری که در جمع گفتم «نمیدانم» یا «اشتباه من بود» کِی بود؟ ۶) آیا مدیرم میتواند بدون حضور من، سه دستاورد مشخص من را نام ببرد؟ ۷) آخرین باری که یک اختلاف را زود و مستقیم مطرح کردم بهجای اینکه بگذارم انباشته شود؟ ۸) این فصل چیزی یاد گرفتم که سختم بود، یا فقط تکرار کردم؟
هر «نه» یک کار مشخص برای هفتهٔ آینده است — نه یک قضاوت دربارهٔ خودت.
«جواب صادقانهام دربارهٔ دانش فنی نیست، چون آن را میدانم چطور به دست بیاورم. چیزی که میخواهم روی آن کار کنم، تأثیرگذاری روی تصمیمهایی است که خارج از تیم من گرفته میشوند. الان میتوانم در تیم خودم یک تصمیم فنی را جا بیندازم، چون آدمها من را میشناسند و به قضاوتم اعتماد دارند. ولی وقتی موضوع بین چند تیم است، هنوز خوب نیستم — بهخصوص در فهمیدن اینکه هر طرف بر اساس چه چیزی ارزیابی میشود و پیشنهاد را چطور باید برای او قاب کنم. کاری که برایش میکنم مشخص است: بیشتر مینویسم، چون در مقیاس چند تیم فقط نوشته پخش میشود نه حضور؛ قبل از هر پیشنهاد بزرگ، جداگانه با آدمهای کلیدی حرف میزنم بهجای اینکه در جلسه رونمایی کنم؛ و سعی میکنم بهجای اصرار روی راهحل خودم، مسئله را طوری بگذارم که تیمهای دیگر هم مالکش شوند. سنجهای هم برای خودم دارم: اینکه چند بار تصمیمی که من شروعش کردهام، در نهایت بدون نام من پیش رفته است — آن برای من نشانهٔ موفقیت است نه از دست دادن اعتبار.»
از یک نقطه به بعد، سقف رشد تو دانش فنی نیست؛ ارتباط، اعتماد و قضاوت است — و اینها مهارتهای قابل آموزشاند، نه ویژگی شخصیتی. بنویس: پیام با نتیجه در ابتدا و یک درخواست، سؤال با هدف و تلاش و خطای دقیق، سند طراحی با گزینهها و ریسکها، ADR برای تصمیمهای گران، توضیح PR که مرور را سریع میکند، گزارش وضعیت که ریسک را زود میگوید، و postmortem که بهجای مقصر، لایههای شکسته را پیدا میکند. در مرور کد، شدت نظرت را برچسب بزن، طراحی را قبل از سبک بگو و سریع مرور کن؛ و در طرف دیگر، کد را از خودت جدا کن و بنبست را با نظر سوم بشکن. در جلسه، تصمیم و مالک و تاریخ را مکتوب کن، با بازگویی مخالفت کن، و بهجای «نه»، هزینه را شفاف کن و تصمیم را به مالک اولویت برگردان. در تخمین، بازه و فرض بده، سطح اطمینان را برچسب بزن، و تأخیر را در لحظهٔ فهمیدن اعلام کن، با گزینههای قیمتگذاریشده. مدیرت کارت را نمیبیند: دفترچهٔ کار و سند دستاورد بنویس، دستور جلسهٔ یکبهیک را خودت بگذار، و شکاف تا سطح بعدی را مشخص کن. تأثیر بدون اختیار از پیشسیمکشی، فهم انگیزهٔ طرف مقابل، شروع کوچک و انتخاب چند تپه در سال میآید؛ با غیرمهندسها، ریسک را به پول، زمان و احتمال ترجمه کن. تعارض را زود و کوچک و خصوصی مطرح کن، از پایینترین پلهٔ escalation شروع کن و صفتها را با تاریخ و عدد جایگزین کن. رشد را روی اصول سرمایهگذاری کن نه مد روز، حلقهٔ بازخوردت را به یک نفر وابسته نکن، و با یاد دادن یاد بگیر. قضاوت یعنی سریع اشتباهت را بپذیری، «نمیدانم» را با اعتبار بگویی، و بین بدهی آگاهانه و آسیب واقعی خط بکشی. و پایداری شرط همهٔ اینهاست: فرسودگی را از بدبینی و بیتفاوتی بشناس، مرز بگذار، تمرکز را محافظت کن و بهصورت ناهمگام و قابل پیشبینی کار کن. هیچکدام از اینها با یک تصمیم بزرگ به دست نمیآید — با ریتم هفتگی و چند عادت کوچک به دست میآید که در طول سالها روی هم جمع میشوند.
A scene you have probably watched play out. Two engineers on the same team. The first knows algorithms better, writes code faster, and is usually right in technical arguments. The second is slower, but when she writes something people read it, when she gives an estimate nobody worries, and when she proposes something the team adopts it. Six months later she is leading a large project and he is still picking up tickets and delivering them — and he honestly does not understand why.
This chapter is about exactly what happens in that gap. And let me break a common misconception right away: we call these "soft skills", as if soft meant vague, subjective, or something you either have or you don't. That is wrong. Writing a design document that actually gets reviewed, giving an estimate that does not damage your credibility later, disagreeing in a meeting without becoming the team's designated pessimist — these are teachable skills with clear structures, clear patterns and clear sentences. Exactly like a design pattern. And like any technical skill, they improve with deliberate practice.
One important boundary before we start: this chapter is about doing the job, not getting the job. Interview technique, behavioral answer structure and a bank of prepared stories live elsewhere (the interview-craft and behavioral-interview-bank chapters). Here we assume you have been hired, and now you have to survive, make an impact, and grow in a real environment.
- The honest argument: why past a certain point your ceiling is no longer technical knowledge · 2) What "senior" actually means, and the IC versus management fork · 3) Writing as an engineer's highest-leverage skill: messages, technical questions, design docs, ADRs, PR descriptions, commits, bug reports, status updates, postmortems · 4) Code review as a social skill, from both sides · 5) Meetings, productive disagreement, and saying no · 6) Estimation, commitments and expectations, plus managing up · 7) Influence without authority, and working with non-engineers · 8) Collaboration, conflict and the escalation ladder · 9) Learning and growth as a practice, not a mood · 10) Professional judgement and ethics · 11) Sustainability: burnout, boundaries, on-call, remote work and the first ninety days · 12) A weekly operating system you can run starting this week.
1) The honest argument: where is the ceiling?
Let's do the arithmetic without flattery. In your first three or four years, almost all of your growth comes from a single source: knowing more things. You learn the language better, understand the framework more deeply, get comfortable with the database, learn concurrency patterns. Every extra unit of knowledge converts directly into better work. That period is enjoyable because the relationship is linear and fair.
Then something happens: the curve flattens. Not because there is nothing left to learn — there is always more — but because the bottleneck has moved. What you are able to do is no longer limited by your knowledge; it is limited by how many people trust your judgement, how well you can make what is clear in your head clear in other people's heads, and whether anyone is willing to hand you an ambiguous, unowned problem.
Picture a 500-horsepower engine in a car crawling through narrow city streets. Adding another 200 horsepower adds essentially no speed, because the engine is not the constraint — the road is. Technical knowledge is your engine. Communication, trust and judgement are the road. An engineer who spends years making the engine bigger while ignoring the road ends up feeling "I work hard but I'm going nowhere" — and that feeling is completely accurate.
Do not read this as a criticism of yourself. Nobody told you. University curricula, online courses, even technical interviews all measure the engine. Nobody teaches the road, and then at performance review time people receive vague feedback like "you need more visibility" or "we expect more impact", which contains no actionable instruction whatsoever. The goal of this chapter is to translate those vague sentences into specific things you can do.
What does "senior" actually mean?
Put engineering career ladders from different organizations side by side and, under different names, they nearly all rotate around the same four axes: scope, autonomy, multiplying others, and tolerance for ambiguity. None of them is about the volume of what you know.
| Axis | Growing engineer | Senior engineer | What it means day to day |
|---|---|---|---|
| Scope | One task, one class | One system or one end-to-end flow | "Who owns this service?" The answer is your name |
| Autonomy | Problems are defined for them | Finds and defines the problem | Nobody has to break your work into pieces |
| Multiplication | Output = own work | Output = own work + everyone else's improved | Reviews, mentoring, tooling, docs, patterns |
| Ambiguity | Works well with clear requirements | Starts with unclear requirements and clarifies them | "Nobody knows what they want" is a starting point, not an excuse |
| Risk | Sees technical risk | Translates technical risk into business language | A manager can decide correctly without knowing the internals |
Here is a simple, unforgiving measure of where you stand on that ladder: how much energy from other people is required for your work to reach completion? If every task requires someone to decompose it for you, check in halfway, and tie it up at the end, you are a consumer of team capacity. If you take something in and what you hand back is more complete than what you received — documented, tested, risks flagged — you are a producer of team capacity. Job titles eventually adjust themselves to that number.
A repeating failure pattern: an engineer who genuinely is the most technically capable person there, but nobody wants to work with him. His reviews are condescending, he interrupts people in meetings, he does not write documentation because "the code speaks for itself", and he treats every disagreement as a battle. This person is usually kept around for a while because he is valuable — and then removed in the first reorganization, because his cost to everyone else finally exceeded his individual value. High technical skill never buys immunity; it only delays the expiry date.
Two tracks: IC and management
Past a certain level (usually after senior) the path forks. Many people make this decision unconsciously, based on salary, and spend the next few years unhappy. It is much better to choose deliberately.
| IC track (Staff/Principal) | Management track (EM/Director) | |
|---|---|---|
| Main lever | Getting high-risk technical decisions right | Building and keeping a team that decides well |
| Unit of work | Architecture, standards, critical systems | People, process, priorities, budget |
| Source of satisfaction | Solving hard problems, building something durable | Growing people, fixing a broken team |
| Worst day at work | Back-to-back meetings, zero focus hours | A hard performance conversation, delivering bad news |
| Personal risk | Drifting away from organizational decision power | Technical skills rusting, the move becoming one-way |
| What both share | Writing, influence, expectation setting, conflict | Exactly the same things |
The reassuring part: everything in this chapter is required on both tracks. If you stay an IC, these are your instruments of influence; if you become a manager, these are your job. So no hour you spend here is locked into one path.
"I have seen this, and at first it looked unfair to me. We had a colleague who unquestionably had the deepest technical knowledge on the team — every complex bug that hit a dead end eventually landed on his desk and got solved. But when I looked at how team output actually got produced, I understood the difference. His work ended at the boundary of a task: the bug was fixed, nobody understood why it had broken, no test was added, and next time he was needed again. I was not technically ahead of him, but every time I solved something I wrote a page explaining what happened, added a test that would catch a recurrence, and spent three minutes on it in the team meeting. The result was that the 'only he knows how' risk did not surround my work. When an organization decides who to trust with an important system, it is not looking for maximum knowledge — it is looking for minimum risk. In that decision, I was the lower-risk option. Later I shared exactly this with him, without judgement, just as something I had learned myself."
2) Writing: an engineer's highest-leverage skill
If you take only one section of this chapter seriously, make it this one. The reason is simple and quantitative: anything you say is heard once, by a few people; anything you write can be read dozens of times, by people you have never met, in months when you are not there. Writing is the only mechanism by which your impact detaches from your physical presence.
And an uncomfortable fact: in most organizations, decision-makers do not read your code. They read your writing. The quality of your writing is, in practice, the user interface through which the organization touches the quality of your engineering.
You do not design a good API to "cover everything"; you design it so the consumer reaches what they need with minimum effort. Text works the same way. Your reader is your consumer and their attention is the scarce resource. Every sentence you force them to read that produces no outcome for them is like a redundant, confusing field in your response payload.
2.1 A message a busy person will actually act on
Rule one: bottom line up front. Say what you want first, then explain why. Most engineers write it backwards, because in their head they are replaying the path that led them to the problem; but the reader first needs to know "what does this have to do with me and what am I supposed to do".
Rule two: one ask per message. If you want three things, you will most likely get one answer and lose two.
Rule three: state the deadline and what happens if nobody replies. Busy people prioritize by deadline, not by the order messages arrived.
❌ Bad version:
Hi, hope you're doing well. I've been working on the reporting service
and I noticed the queries on the transactions table have gotten really
slow, and I think it might be an index issue, or maybe the data volume.
It could also be related to last week's changes. What do you think?
✅ Good version:
Ask: your approval to add an index on transactions(account_id, created_at)
in production — by end of day Wednesday.
Why: the daily report went from 400ms to 9 seconds and three teams have
reported the slowdown.
Cause: the table grew to 80 million rows with no index matching this query
pattern.
Risk: index build takes about 6 minutes, runs CONCURRENTLY so writes are not
blocked, and adds roughly 2GB of disk.
If I don't hear back by Wednesday: I will proceed assuming approval and
announce it in the team channel.
That last line — "here is what I will do if you don't reply" — is one of the most effective sentences you will ever learn. It takes the decision pressure off the reader and it takes you out of the "I was waiting" state.
Messages that open with "Hi, got a minute?" and then wait for a reply look polite, but in practice they shift the cost onto the receiver: they have to context-switch, respond, wait, and only then discover what the topic even was. On a distributed team that single round trip can burn a working day. Being polite means putting everything in one message, not opening with a greeting.
2.2 A technical question people enjoy answering
Asking is not weakness; asking badly is what hurts you. A bad question sends three signals: "I didn't spend my own time thinking", "I expect you to reconstruct the problem from scratch", and "if your answer is wrong I won't notice". A good question does the exact opposite.
The four-part structure that almost always works:
[Goal] I'm trying to scan uploaded files before they are stored in object
storage, so corrupted files get rejected.
[Tried] I tried: (1) reading fully into memory — died with OutOfMemoryError
on a 2GB file. (2) streaming with an 8KB buffer — works, but takes
90 seconds for a large file and we hit a timeout.
[Exact error] In case 2, service log: "Read timed out after 60000 ms" in
StorageClient.upload, line 142.
[Specific ask] My question: is the right pattern here to scan asynchronously
after upload, or should we raise the timeout for this path? If we
already have a similar pattern in another service, pointing me at it
is enough.
Notice what this does for the receiver: in 30 seconds they know what the problem is, they know which paths have been tried (so they don't suggest a repeat), they see the exact error (so they don't guess), and they know precisely what is being asked (so they can answer in two sentences).
The most common failure in asking is to ask about the solution you already chose instead of the actual problem. You ask "how do I capture the last character in this regex?" when the real problem was extracting a file extension and no regex was needed at all. Always add one sentence: "what I'm ultimately trying to do is…". That single sentence will repeatedly save you an hour of work in the wrong direction.
Before asking you should have tried: read the log, checked official docs, searched the repo. But this rule has a second edge that gets mentioned far less often: endless trying is also wrong. Set a time box (say 45 minutes for a medium problem, or two hours for something nobody else knows either) and when you hit it, ask. Somebody who spends three days stuck on something a colleague would have answered in two minutes is not independent — they are just expensive. In performance feedback that is called "poor judgement", not "hard work".
2.3 A design document that actually gets reviewed
Most engineers think a design document means "an explanation of what I want to build". No. A design document is a decision-making instrument. Its purpose is to make disagreement happen on text, before it has to happen on three weeks of code. If your document does not cause anyone to ask a hard question, the document failed — even if everyone approved it.
The fundamental difference from documentation:
| Design document | Documentation | |
|---|---|---|
| Written | Before building | After building |
| Audience | Decision-makers and technical peers | Whoever will use or maintain it |
| Centered on | Why this option and not the others | How it works and how to use it |
| Fate | Frozen after the decision (a historical record) | Must stay alive and be kept current |
| Success means | Productive argument and a clear decision | The reader gets going without asking anyone |
A short, realistic template:
# Design: extracting notifications into a separate service
## Problem
Notifications are currently produced and sent inside the order service.
Three consequences:
1) Notification provider slowness slows order placement (P99 went from
200ms to 1.8s during last month's incident).
2) Every team that needs a notification writes code in the order service —
4 PRs from 3 teams last month.
3) Testing notification scenarios requires standing up the whole order service.
## Constraints
- No downtime for order placement.
- Two engineers, six weeks.
- Audit requirement: every sent notification must be traceable for 90 days.
- We already run a message broker and the team knows it.
## Options considered
A) Extract into a separate service with event-driven integration.
+ Full failure isolation, independent scaling, clear ownership.
− A new service means new on-call, operational cost, eventual consistency.
B) Keep it in place but make it asynchronous with an internal queue.
+ Cheapest, roughly two weeks, no new service.
− Ownership stays fuzzy, "every team writes code in our repo" is unsolved.
C) Use an off-the-shelf external notification service.
+ Fastest.
− Takes audit data out of our control. Rejected.
## Decision
Option (B) for phase one, behind a clean module boundary, with option (A)
in phase two if consumers exceed three teams. Rationale: the six-week
constraint, and a clean module boundary keeps the later migration cheap.
## Risks and what would surprise us
- If the queue backs up, notifications are delayed. Mitigation: alert on
queue depth above 10k.
- Notification ordering is not guaranteed. Checked with product — acceptable.
- If phase two never happens, we are no worse than today, but no better.
## What would invalidate this decision
If more than three consumer teams appear by the end of next quarter,
phase two becomes mandatory.
That final section — "what would invalidate this decision" — is what separates senior-level documents from the rest. It means you have not only made a decision, you have pre-specified the conditions under which it gets revisited.
Never send an important document straight to ten people. First give it privately to the one or two people most likely to object, and say: "this is still rough, I wanted your take before I send it out." Two things happen: the obvious weaknesses are removed before public exposure, and that person no longer feels in the meeting that a decision was made without them. Most consensus is built outside the meeting.
The diagram below shows the practical review cycle of a design document · چرخهٔ عملی مرور یک سند طراحی.
flowchart TD
A[Draft: problem and constraints only] --> B[Share with 1-2 likely dissenters]
B --> C{Problem statement agreed?}
C -- No --> A
C -- Yes --> D[Add options and trade-offs]
D --> E[Broad review with a deadline]
E --> F{Blocking concerns raised?}
F -- Yes --> G[Address in writing, update options]
G --> E
F -- No --> H[Decision recorded and frozen]
H --> I[Write ADR and link from code]
If you write "have a look whenever you get a chance", nobody looks, and two weeks later you are still waiting. Always write: "please comment by Thursday noon; after that I'll assume agreement and start." That is neither rude nor pushy — it is respect for your own time and it prevents decision paralysis.
2.4 The architecture decision record (ADR)
An ADR is the small, durable sibling of the design document. A design doc can be long and then forgotten; an ADR is a short file that lives next to the code and answers exactly one question: "why is it like this?" — the question the next engineer asks, irritably, six months from now.
The classic four-part format, which is nearly identical everywhere:
# ADR-0012: Idempotency keys on the payments API
Status: Accepted — 2025-08-03
(Possible statuses: Proposed | Accepted | Deprecated | Superseded by ADR-XXXX)
## Context
Mobile clients on weak networks retry payment requests. Over the past three
months we have had 14 reported double charges, all reversed by hand.
Detecting duplicates from amount plus timestamp is not reliable, because a
user may genuinely make two identical payments.
## Decision
Every POST /payments must carry an Idempotency-Key header containing a
client-generated UUID. The server retains the result of the first execution
for 24 hours and returns the same response for a repeated key. A missing
header is rejected with 400.
## Consequences
+ Double charges caused by network retries are eliminated.
+ Clients can retry safely, which simplifies their resilience logic.
− All clients must be updated. Older mobile versions are exempt for 90 days.
− Requires a TTL-backed store for keys and monitoring of key collision rate.
Not every choice needs an ADR. Practical test: if changing this decision six months from now would take weeks or involve several teams, write an ADR. If a single PR reverses it, don't. A library of 200 ADRs where 190 of them are about variable naming saves nobody. And the golden rule: never edit an accepted ADR — write a new one and mark the old one superseded. Their whole value is the history.
"We put a cache layer in front of the product search service and I was its main advocate. My reasoning was sound — database load was high and the cache cut it by 70 percent. What I had not properly considered was that the product team was about to start changing prices several times a day that quarter. Two weeks after release, complaints started: the price on the product page did not match the price in the cart. The cause was incomplete invalidation in my cache. What I did was this: on the same day I wrote in the team channel that this came from my decision and I was working on it — I did not let anyone spend time hunting for a culprit. Then instead of patching it, I reopened the original ADR and saw that I had written nothing under 'what would invalidate this decision'. I wrote a new ADR superseding it: cache only non-price fields, always read prices live. The longer-term effect was that from then on every design document in the team had a mandatory section on what business change would break the design. That mistake taught the team something a month of theoretical debate would not have."
2.5 A PR description that makes review fast
Your reviewer is a busy person who has abandoned their own context to come look at your code. The quality of the review depends directly on how quickly they can get inside your head. That is what the PR description is for.
## What
Destination account validation for transfers moved into the domain layer.
## Why
The same validation existed in three places (transfer controller, batch job,
file import) and the batch job had drifted to an older version — which is
exactly the bug reported in ticket #4412.
## How to review this
- The core of the change is AccountNumberPolicy — start there.
- The other three files are just call-site replacements, skim them.
- New tests in AccountNumberPolicyTest cover the boundary cases.
## What I deliberately did not do
I left the error response format alone to keep this PR small — follow-up
in #4470.
## Risk and rollback
No behaviour change except that the batch job is now stricter. If older
input files start getting rejected, reverting this commit restores the
previous behaviour.
An 800-line PR usually gets approved, not because it is good but because reviewing it is exhausting. Feedback quality collapses with PR size: at 50 lines you get comments about logic; at 800 lines you get either silence or comments about variable names. If you must make a large change, split it into a sequence of PRs and put the map of the whole path in the first description. That is respect for the reviewer's time and, incidentally, the best way to get real feedback on your design.
2.6 Commits and bug reports
A commit message is written for the person who, six months from now, is standing on a strange line with git blame open asking "why?". So the body must explain the why, not restate what the diff already shows.
Reject transfers to closed accounts in batch import
The batch import path used a stale copy of the account-status check that
did not include the CLOSED state, so an import file could create transfers
that later failed at settlement and had to be reversed by hand.
The check now lives in AccountNumberPolicy and is shared by all three entry
points. Behaviour of the online path is unchanged.
Refs: #4412
Practical rules: short subject line (under 72 characters), imperative mood ("Reject", not "Rejected" or "Rejecting"), blank line, then a body that explains why. The tooling and branching side of this lives in the git-workflows chapter; here I only care about the communication part.
A bug report follows exactly the same logic as a good question, with one addition: it must be reproducible.
Title: Transfer between two accounts in the same currency charges an FX fee
Environment: staging, version 2.14.3, browser irrelevant (also reproduced
via the API)
Steps:
1) Create two accounts in EUR.
2) POST /transfers for 100 from the first to the second.
3) Look at the response.
Expected: fee = 0, because no conversion takes place.
Actual: fee = 1.5 and feeType comes back as "FX".
Scope: only when both accounts are in a non-base currency. Does not happen
with two accounts in the base currency.
Impact: the customer is charged for a service that was not performed — this
is a compliance issue, not only a bug.
Evidence: trace id a7f3-… and FxRateResolver log line 88.
The "scope" line — when it happens and when it does not — is the single most useful thing for whoever has to debug it, and it is almost always missing.
"I have learned to write in two layers. The first layer is a short section at the top with no technical vocabulary that answers three questions: what is causing problems right now, what am I proposing, and what happens if we do nothing. That is five or six lines. The second layer is the rest of the document, fully technical, and I make no attempt to simplify it, because its audience is engineers. I once wrote a document arguing that we needed to upgrade an old dependency. The first version was full of version details and breaking-change lists, and no decision came out of it. I rewrote it and put this at the top: 'this library no longer receives security patches. If a new vulnerability is published, our only option is to turn the service off. The work is two weeks for one engineer.' It was approved that week. The point was not that I simplified it — it was that I translated the risk into the language of a decision: cost, likelihood, and options."
2.7 A status update that builds trust
The status update is the only instrument by which people who cannot see your work decide how much to trust you. And most engineers ruin it in one of two ways: either it is so vague it says nothing ("in progress, all good"), or so granular that the reader has to draw their own conclusions ("finished refactoring class X, added 3 tests, now working on Y").
A good update contains three things and three things are enough: are we on time?, what changed?, what do I need from you?
Notifications migration — week 3 of 6
Status: green, with one caveat.
Progress: the asynchronous path works end to end and was tested in staging
under 5x synthetic load. The first consumer (order service) is connected.
Change since last week: the payments team has also asked for notifications.
That is outside the original scope; I have declined for now and we will
look at it after phase one.
Risk: the notification provider's test environment is down for two days
next week. If we cannot find another way to test, delivery slips two days —
still inside the buffer.
Ask: exactly one — access to the infrastructure team's sandbox account by
Tuesday.
Three things this text does invisibly: (1) it states status in one word so the manager does not have to read everything; (2) it raises a risk before it becomes a problem — this single habit builds trust faster than anything else; (3) it separates the ask from the narrative.
The worst reporting pattern is saying everything is fine for weeks and then announcing in the final week that you are three weeks behind. From the inside it feels reasonable ("I was hoping to catch up"), but from the outside you either did not understand the situation or you concealed it — and both are bad. A delay reported early is almost always manageable; the same delay reported late is a crisis. Rule: the worst news should travel fastest.
Before sending, read it with the eyes of somebody who knows nothing about your week. Your class names and ticket numbers mean nothing to them. If a sentence is incomprehensible without internal detail, either cut it or translate it into its external effect: "we invalidated the cache" becomes "prices are no longer shown with a delay".
2.8 A blameless postmortem that is still useful
After every serious incident somebody writes a document. The quality of that document tells you whether the organization is mature — and your behaviour in that meeting tells everyone whether you are.
"Blameless" is widely misunderstood. It does not mean "nobody made a mistake" or "we won't talk about people". It means: we assume everyone did the most reasonable thing available given the information and tools they had at that moment, so the right question is why the wrong action looked reasonable at the time. That change of angle is the difference between a useful document and a trial.
After an accident, aviation investigators read the black box to understand how the system allowed this to happen — not to decide which pilot to fine. If pilots were afraid, they would stop reporting, and the industry would go blind. Software is identical: if the engineer who ran the wrong command is afraid, next time they will hide it as long as possible, and you will lose the golden first ten minutes of an incident.
# Postmortem: partial payments outage — 94 minutes
## Impact
For 94 minutes, roughly 31% of payment requests returned 503.
Estimated 2,800 failed transactions. No data was lost.
## Timeline (local time)
14:02 — version 3.4.0 deployed to 2 of 6 nodes (canary).
14:09 — error rate begins climbing. No alert fires: the threshold was
computed across the whole cluster.
14:31 — manual report from the support team. Investigation starts.
14:52 — suspected cause: connection pool leak in the new version.
15:12 — decision to roll back.
15:36 — rollback complete, error rate normal.
## Cause
On the error path, the new version failed to return connections to the pool.
Under normal load this leak exhausted the pool in about 45 minutes.
## Why we did not see it sooner (the important section)
- The error-rate alert averaged across the whole cluster, so 2 bad nodes out
of 6 kept it under the threshold. Our monitoring could not see a canary.
- The existing load test only exercised the success path; the error path had
never been run under load.
- The on-call engineer had no access to the canary dashboard — permissions
had changed the month before.
## Action items (each with an owner and a date)
1) Per-deployment-group alerting — infrastructure team — by the 20th.
2) Add error-path scenarios to the load test — payments team — by the 27th.
3) Quarterly review of on-call access — team manager — next quarter.
## What went well
- Rollback completed in 24 minutes because the script had been tested.
- Support was in the loop from minute 35 and everyone gave the same message.
The "why we did not see it sooner" section is the most valuable part of the document, because an incident always requires several layers of defence to fail at once. And do not delete "what went well"; if you only record failures, the team never learns which investments actually paid off.
If the root cause in your document reads "the engineer accidentally ran the command against production", the document is not finished. The next questions are: why did the tooling allow it? Why did the environments look identical? Why was a second confirmation not required? People always make mistakes — that is a constant of nature, not a research finding. A system built on the assumption of flawless humans is a broken system.
"The first thing I do is write and publish the timeline and the cause myself, before the meeting. When you explain what happened earlier and more precisely than anyone else, two things follow: the discussion skips the 'who?' stage and goes straight to 'how do we prevent it?', and trust in you goes up rather than down. In the meeting itself I am careful about two failure modes: not beating myself up excessively — because that makes the room uncomfortable and people switch from analysis to reassuring me — and not getting defensive. I focus on why the system let that code reach production: which test was missing, which alert was missing, which review would have caught it. The last time this happened, the meeting produced three action items and none of them was 'be more careful', and those three items prevented two similar incidents over the following months."
3) Code review as a social skill
Code review is the one place where you routinely, and in writing, pass judgement on a colleague's work. Which is exactly why it is simultaneously the largest source of damage to team relationships and the largest opportunity to build trust.
The technical content of review — what to look for in the code, code smells, abstraction boundaries, testability — is the subject of the clean-code-craft chapter and I will not repeat it here. This section is only about how to say it, from both sides of the table.
3.1 The reviewer's side: critique without damage
The fundamental problem is that text has no tone. "Why didn't you use a stream here?" is simple curiosity in your head and can read as "you don't even know streams" in theirs. The fix is not to artificially soften everything — the fix is to make the severity of each comment explicit.
| Label | Meaning | Example | What the author should do |
|---|---|---|---|
blocking: |
I will not merge until this is resolved | "blocking: this error path never closes the connection and the pool will exhaust under load." | Fix it, or make a counter-argument |
issue: |
A real problem, but debatable | "issue: this function has two responsibilities and that makes it hard to test." | Fix, or agree on a later time |
question: |
I genuinely don't know, explain | "question: what happens here if the input is null? Maybe it's handled upstream." | Just answer |
suggestion: |
This would be better, and it isn't just taste | "suggestion: this condition could move into a well-named method." | Optional, but reply |
nit: |
Trivial, feel free to ignore | "nit: extra whitespace on line 22." | Ignore freely |
praise: |
This was good | "praise: this boundary test is exactly what we were missing last month." | Nothing |
This is a small habit with an outsized effect. In ten seconds the author learns that out of 17 comments, two actually block the merge and the rest are taste or curiosity. Without the labels, all 17 land with equal weight and produce the feeling that the whole change was rejected.
If your first review round leaves twenty comments about naming and formatting, and your second round says "by the way, I think this whole approach should change", you have burned the author's time and your own credibility. Read the code top-down first: is the right problem being solved? Are the boundaries right? Is this change even in the right place? If the answer to any of those is no, write only that and hold the rest. Style and formatting are fundamentally a formatter and linter's job anyway, not a human's.
A damaging pattern: a reviewer who uses every PR as an excuse to display what they know — "this could have been written much more elegantly with pattern X" on a PR that fixed an urgent bug in three lines. The result is that people avoid sending you PRs, send larger and later changes, and the quality of the whole team drops. Before each comment ask yourself: "if I don't write this, what bad thing happens?" If the answer is "nothing, it just isn't my taste", either label it nit: or skip it.
A PR waiting three days for review has cascading costs: the author loses context, the branch drifts from main and picks up conflicts, and most importantly the engineer learns to send bigger PRs (because the waiting cost is fixed). The rule you see in healthy teams: first review inside one working day. If you don't have time for a full pass, at least spend 15 minutes on a high-level look and say "the design makes sense to me, I'll go through the details tomorrow morning" — that one sentence removes the author's uncertainty.
3.2 The author's side: receiving review without defensiveness
The first thing to accept is unpleasant: your brain treats your code as part of you. When someone writes "this logic is wrong", the same regions activate as under personal criticism. That is normal and willpower does not remove it. What you can do is manufacture some distance.
Three techniques that genuinely work:
One: mentally rewrite the sentence. Before reading a comment, translate it: "the code" instead of "you", and append "yet". "This test is incomplete" becomes "this code does not cover one boundary case yet". Same information, no charge.
Two: the twenty-minute rule. If a comment makes you angry, do not reply in that moment. Give it twenty minutes. Almost always, after twenty minutes you either see the comment as correct, or you can write your disagreement without an edge. Almost no reply in code review is urgent enough to be worth burning a relationship over.
Three: separate "this is wrong" from "I would have written it differently". Decide which category each comment belongs to. If it is the second and the change is cheap, just do it without debate; save your social capital for the first. An engineer who fights about everything is not heard when they fight about something important.
And when you genuinely should push back, this structure works:
Thanks — that's a fair point and I agreed with you at first.
The reason I wrote it this way is [specific constraint: e.g. this method is
called from a scheduler with no surrounding transaction].
With your suggested approach, that path would [specific consequence] — you
can see it in test X.
If there's a place where that constraint doesn't apply and I've missed it,
I'm happy to change it.
Three properties: it opens with acknowledgement (not politeness theatre — it shows you actually read the comment), it gives a technical, checkable reason rather than taste, and it leaves the door open. Whoever reads that feels neither defeated nor ignored.
If two round trips have not resolved a disagreement, a third almost never will — it only degrades the tone. Write this instead: "I think we've both stated our case and the thread is getting inefficient. Let's talk for 15 minutes, and if we still disagree let's ask [third person] to decide. Whichever way they go, I'll implement it." That last sentence matters: you commit to the outcome in advance, which drains all the tension out of the situation.
"First I assume good intent and frame the problem as mine rather than theirs. Instead of thinking 'this person is pedantic', I assume the team may not have a written standard and they are applying their own mental one. Then I talk to them privately and briefly — not in the PR thread, because a conversation about process in the middle of a conversation about code always goes badly. My sentence is roughly: 'I take your reviews seriously and you've caught real bugs of mine. There's one thing I wanted to raise: it's hard for me to tell which of your comments block the merge and which are preference, and that makes us do extra rounds. Could we try labelling comments as blocking and nit?' This is almost always welcomed, because the other person never wanted to slow things down either. If the issue is recurring and team-wide, I propose we adopt a shared formatter and linter, so the whole stylistic category is taken off human desks."
"I keep in mind that I have two goals that can conflict: the quality of this change, and this person writing better code next time without becoming afraid to send PRs. If I only optimize for the first, I leave twenty comments and tomorrow they code nervously. What I do instead: I start with one genuine thing that was good — not flattery, for example that they handled the error case. Then, rather than dictating the correct solution, I pose the problem as a question: 'if we add a third payment type tomorrow, how many places would need to change?' That question walks them to the same conclusion I reached, except this time they got there themselves. If I see the topic is bigger than a written comment, we sit together for fifteen minutes; something that takes ten messages in text takes three minutes in person. And at the end I am explicit about which part must be fixed now and which part can go into a follow-up ticket."
4) Meetings, synchronous time, and saying no
Every meeting has a real cost: six people for an hour is six person-hours of the organization's engineering capacity — plus the hidden cost that gets noticed less, which is the fragmentation of the day. A meeting at 11am can destroy two focus blocks, not one.
4.1 A meeting worth its cost
Three questions before any meeting you call:
- Is there a specific decision that has to be made? If the answer is "we want everyone to be informed", that is a document, not a meeting.
- Who is genuinely required? Each extra person lowers the probability of deciding. For a technical decision, more than five people almost always means the meeting becomes a discussion.
- Can attendees prepare in advance? Send a one-page document beforehand and the one-hour meeting frequently becomes twenty minutes.
And put these three lines in the invitation:
Goal: decide whether notifications phase one uses an internal queue or a
separate service.
Pre-read: design doc (link) — the options section, about 5 minutes.
Expected outcome: one recorded decision and an owner for the ADR.
Before people leave, say it out loud and then write it down: "so the decision is …, the owner is [name], by [date]." Those thirty seconds eliminate half of an organization's repeat meetings. A startling number of meetings end with everyone believing "we decided", while three people hold three different interpretations. And if you are the person who always writes that summary, you quickly become the person who writes the official record of decisions — one of the cheapest forms of influence there is.
4.2 Disagreeing productively in the room
Disagreement is necessary; a team where nobody disagrees is not aligned, it is quiet. But the shape of the disagreement determines everything.
The pattern that almost always works has three steps: restate, specific concern, alternative proposal.
❌ "That won't work."
❌ "I don't agree with this, it's too complicated."
✅ "Let me make sure I've got it: the proposal is that each service owns its
own table and we synchronize through events. Is that right?
My concern is the financial reports: today a simple join answers them,
and afterwards we'd be assembling data from three sources that sync with
a delay. I'd suggest we build a dedicated read model for that one case,
or, if that's too expensive, decide explicitly what staleness is
acceptable."
The first step, restating, is the most important and the most frequently skipped. When you repeat their proposal in your own words, three things happen: you make sure you are disagreeing with the real proposal rather than the version in your head; the other person feels heard and does not get defensive; and if there was a misunderstanding, it is resolved right there and the rest of the argument becomes unnecessary.
Sentences like "you always pick the complicated route" or "this is the same mistake you made last year", even if factually true, move the discussion from a technical question to a defence of identity, and from that moment nobody is looking for the right answer. Always start with "this design…", "this approach…", never with "you…". And do not speak badly about someone who is not in the room — that habit destroys the trust of everyone present faster than anything else, because they all silently calculate that you do it in their absence too.
4.3 Speaking up as the least senior person in the room
This situation is real and it is hard. A few things that actually help:
- Speak early, not completely. The longer you wait for your point to be "fully formed", the harder entry becomes. Say something in the first five minutes — even a question — to break the seal.
- A question is a safe entry. "Sorry, a basic question: when we say consistency here, do we mean within a single request or across the system?" That is not weakness; it is usually something three other people also did not know and did not ask. A clarifying question is the most valuable contribution a newcomer can make.
- Use data, not standing. When you have no tenure, "I think this is slow" carries no weight; "I measured this on staging yesterday, this path takes 900ms" does. A number needs no seniority.
- If you are interrupted, come back calmly and without resentment: "one moment — let me finish my sentence, it's only two more." If someone is chairing, saying it once is usually enough.
- If someone else repeats your idea and receives the credit, reclaim it politely and without bitterness: "glad we agree — that's what I suggested a few minutes ago. If you're on board I'll write up the details and send them." Do not get bitter, because in the room bitterness makes you look like the loser, not them. There is also a good team-level fix: on healthy teams people do this for each other — "that was [name]'s point, let them continue." If you do it for others, it usually comes back to you.
"First I check whether I actually have information the others lack, or just a different preference — those are very different situations. If it is only a preference, I usually say nothing or say it very briefly and move on. If I have information, I say it even when everyone agrees, because my silence in that moment means later, when the problem appears, I get to say 'I knew it' — which is the worst possible outcome. How I say it matters too; rather than rejecting the decision I put the risk on the table: 'I'm fine with the decision and if that's where we land I'll push it forward. I just want to make sure we're accepting this risk knowingly: with this design, if the provider gets slow, the whole order-placement path gets slow. If we know that and accept it, I'm satisfied.' One of two things usually happens: either somebody says 'we hadn't considered that' and the decision changes, or it is genuinely accepted with open eyes and I can proceed comfortably. Afterwards I record the point in the meeting notes or the ADR — not to build a case, but so that if it does happen six months later, the team knows the decision was deliberate and can react faster."
4.4 A standup that is actually useful
Standup usually degenerates into a reporting ritual where each person reads their task list and nobody listens. The simple test: if everyone speaks in turn and nobody reacts to anyone else, the meeting produced no value and should become a written message.
The real purpose of a standup is synchronization, not reporting. Three things are worth saying out loud:
- Something others need to know because it affects their work ("I'm changing the schema today, rebase if you're on an older branch").
- Something you are stuck on that someone might unblock in two minutes.
- A change in plan or risk ("I thought this would finish today, now I think Thursday").
Everything else belongs in the tracker. And when two people start going deep, the rescue sentence is: "this is a good discussion but it only involves the two of you — let's continue after standup."
4.5 Saying no and pushing back on scope
Engineers usually say no badly: either they refuse flatly and come across as negative, or they say yes and then fail to deliver, which is far worse. The third path is to make the cost explicit and hand the decision back to whoever owns priority.
❌ "No, I don't have time."
❌ "Sure, I'll try to squeeze it in." (and then nothing lands on time)
✅ "I can do this. Right now I'm on the notifications migration, which I've
committed to for the 15th. This would take about three days, which puts
the migration at the 18th. If this is more important, I'm fine with that —
we just need to align with the order team. Which do you prefer?"
That sentence does several things at once: it shows you are a collaborator rather than an obstacle, it gives a truthful picture of capacity, it quantifies the consequence, and most importantly it hands the prioritization decision to the person who owns prioritization. You do not decide priority; you report reality.
Three other forms of "no" that work well:
- The time no: "Not this week, yes from Monday."
- The scope no: "Not the full version, but I can give you something by Thursday that covers only the main case. If that works, we do the rest later."
- The referral no: "I can't right now, but [colleague] knows this area better than I do and I can connect you."
When you say yes to a piece of work, in that same moment you have said no to something else — usually something you had already committed to — just silently. People who never say no gradually become the people whose work "is always late", without understanding why; from the outside there is no visible difference between "over-committed" and "slow". Being transparent about capacity is not rudeness, it is a precondition for being reliable.
5) Estimation, commitments and expectations
Nothing damages the relationship between engineers and the rest of an organization more than estimation. And the root of it is one simple mismatch: the engineer thinks they are giving a forecast, and the listener thinks they are receiving a commitment.
5.1 Why engineers estimate badly
Three structural reasons, none of which is laziness or carelessness:
The planning fallacy. When predicting time, the human mind simulates the successful, obstacle-free version. Even when you know last time took twice as long, you will still be optimistic next time. This is a well-documented cognitive bias, not a personal defect.
Invisible work. When you say "three days", you are usually estimating the time to write the main code. But between "start" and "done" there is also: reading existing code, waiting for review, fixing what review found, broken CI tests, coordination with another team, documentation, deployment, and fixing the first problem after deployment. On most teams, writing code is less than half the total.
Social pressure. When someone asks "does it really take three weeks?" in a disappointed tone, estimates shrink involuntarily. That is neither honesty nor helpfulness; it just moves the debt into the future.
If someone asks how long it takes to get to the airport, the honest answer is not "40 minutes" — it is "between 35 and 90 minutes depending on traffic". And anyone who does not want to miss their flight plans against the second number. Engineering is the same: a single number destroys information, a range preserves it. Someone who always gives a single number is hiding uncertainty and transferring its risk to the listener without telling them.
5.2 An estimate that does not cost you credibility: range plus assumptions
The practical template has three parts: the range, the assumptions, and what would tighten the range.
Estimate: between 2 and 4 weeks.
Assumptions this rests on:
- The accounts team's API is what the docs say and needs no changes.
- The test environment is available by next week.
- No data migration is needed, because legacy records are exempt.
- I am full time on this and not on-call.
If any of those turns out to be wrong, I re-estimate and tell you the same day.
If you want a tighter number: give me two days to build a small end-to-end
integration spike, after which I can get the range down to ±3 days.
That last part is the most effective negotiating instrument you have: instead of arguing about the number, you propose reducing the uncertainty with work. That is both honest and professional-sounding, because it is professional.
| Confidence level | When you say it | How to phrase it | What the listener may use it for |
|---|---|---|---|
| Rough guess (t-shirt) | In an ideation meeting, no investigation | "small / medium / large — probably large" | Prioritization only, never planning |
| Raw range | You understand the problem, haven't seen the code | "between 2 and 6 weeks" | Quarterly planning |
| Investigated range | You've done a spike or prototype | "3 weeks ±3 days, with these assumptions" | Commitment to an external stakeholder |
| Delivery commitment | The work is ~80% done | "it will be ready Thursday" | Announcing a date to a customer |
Most estimation disasters begin with a "rough guess" said in a hallway that reappears two weeks later as a committed date on a slide. The fix is not refusing to give numbers — the fix is attaching a label to the number: "this is a raw guess, it could be off by 2x in either direction; if you're planning against it, tell me and I'll spend a day and give you a better one." Memorize that sentence.
Many people learn to double their estimates and keep the buffer secret. The problem is that first, nobody knows the buffer exists, so nobody can make a better decision with it, and second, the work expands to fill it anyway. Worse, once someone notices your estimates always carry a hidden safety factor, they start halving your numbers — and you start inflating them further. Nobody wins that race. Keep the buffer, but state it: "three weeks of work plus one week of buffer for what we don't know yet."
5.3 When reality changes: renegotiating
An estimate is not a contract; it is a snapshot of what you knew at the moment you knew least. When the information changes, the number must change too — and announcing it is your job, not the job of whoever thinks to ask.
The most important rule: raise a slip the moment you suspect it, not the moment you are certain. The gap between those two is usually a week, and that week is the difference between "managed" and "crisis".
Something I'd rather tell you early: I think the 15th is at risk, and I'd
put it at roughly 50-50.
What changed: I had assumed no data migration was needed, but it turns out
about 400,000 legacy records are still active and must be migrated. That is
roughly five days of work that was not in the estimate.
Options, in my order of preference:
1) Move the date to the 22nd. Complete, with no technical debt.
2) Ship on the 15th for new customers only, and migrate the following week.
Risk: two parallel paths for a week.
3) Ship fully on the 15th with a second engineer helping on the migration.
Risk: their current work slips.
My recommendation is option 2. The decision is yours — whichever you pick,
I'll start on it today.
This message demonstrates the difference between a senior engineer and a growing one on a single page. The growing engineer delivers a problem; the senior engineer delivers a decision with priced options. The manager who receives this ends up trusting you more, not less, even though the news is bad.
"I say it the same day, even if my new number isn't precise yet. My experience is that people cope with bad news but not with surprise. The sentence I open with is: 'I want to flag early that the date is at risk. I don't have an exact number yet but I will by tomorrow.' Then, before the follow-up conversation, I prepare three things: which assumption turned out to be wrong and why it was reasonable at the time, the new number as a range, and two or three realistic options with the cost and risk of each. It matters that I bring the options, because if I don't, someone else invents them and the worst one usually wins — something like 'everyone works the weekend'. After delivery I also go back and write down how I could have found that wrong assumption sooner; in this case the conclusion was that before any large estimate I now spend half a day looking at real production data, because twice in a row it was data volume that surprised me."
"First I try to understand what is behind it, because there is usually a real constraint — a customer commitment, an event, a dependency. I ask: 'what makes this date important?' The answer changes the entire conversation. Then I'm explicit that I cannot shrink the number by willpower, but we can shrink the scope — and those are two completely different things. I usually say: 'time and scope are locked together. If the date is fixed, let's decide together what comes out of scope. I can tell you that if we drop the reporting piece and legacy file support, it's achievable in half the time, and we add those in the next phase.' What I never do is lower the number when nothing has changed, because that just buys the same delay at the price of months of anxiety and a public failure. And if, despite all of that, the decision is to keep the compressed date, I write down what assumptions and risks we are proceeding with — with no accusatory tone, simply so that when a decision is made, its cost is on the record too."
6) Managing up
"Managing up" has an unfortunate name that smells of flattery, but the meaning is entirely mechanical: your manager is a system with limited inputs and limited processing capacity; the quality of its output about you is a function of the quality of the input you provide. If you provide none, they decide from guesswork and from what other people say.
A fact many people learn late: your manager does not see your work. They do not read the code, they are not in the PR thread, they have no idea how hard that bug was that took three days to find. They hold a very low-resolution picture of you, assembled from a few sources: things you have told them, things others have said about you, and a few visible outcomes. Your contribution to that picture is the only part you control.
6.1 What you should actually tell them
Three categories of information a manager genuinely needs, and most engineers withhold:
- Anything that could blindside them in public. If a date is slipping, a customer is unhappy, or you made a controversial technical decision, your manager should know before anyone else. Simple rule: your manager should never hear bad news about you from a third party.
- Anything only they can fix. Access, conflicting priorities between two teams, an unresponsive colleague, a tool that needs budget. Bringing these is not weakness; it is precisely what they were hired for.
- Anything that completes their incomplete picture. Meaning the work you did whose effects are not visible from outside.
6.2 Being visible without bragging
The line between "I made my work visible" and "I promoted myself" is one that many people, fearing to cross, never approach at all. The practical rule for finding it is simple: talk about effect, not effort; and replace adjectives with numbers.
❌ "I worked really hard this week and fixed a bunch of things."
❌ "I fixed the slow reports problem." (no context, reads as a claim)
✅ "The daily report that three teams were complaining about is back from
9 seconds to 400ms. The cause was a missing index. I also added an alert
so that if it crosses 2 seconds again we find out before a user complains."
The third version is not bragging because it contains no adjective about you; it just reports facts, and the facts happen to be impressive. Adding "and here is what I did so it doesn't recur" is the signature of a senior engineer.
Every Friday, spend ten minutes writing in a plain file: what I did this week, what I learned, what was blocked. That is your work journal and it is for you. Then once a month, promote the presentable items into a second file with the structure "problem → what I did → measurable effect → who can confirm it". That is your brag document. Six months later when review season arrives, the difference between someone who has that file and someone who does not is startling: human memory holds only the last two months and the rest of the year evaporates. It will also help you at interview time later, but that is not its main purpose.
6.3 A one-on-one that is not wasted
Most one-on-ones degenerate into status reporting — which is waste, because status could have been written. That meeting is the only time you are guaranteed your manager's full attention. You write the agenda; if you don't, it becomes whatever happened to be on their mind that day.
Agenda for this week (30 minutes)
1) Blockers — 10 min
The infrastructure team hasn't responded to the access request in two
weeks. Can you open that path?
2) Feedback — 10 min
I wanted your read on the notifications design doc — specifically
whether the options section was clear enough.
3) Growth — 10 min
I want to know what, in your view, is still missing for me to reach the
next level. If there's one specific thing I should demonstrate, what
would it be?
Bring item three at least every two months. The answer you get is often vague ("more impact") — push right there until it is specific: "can you give me an example of something that, if I had done it, would have demonstrated that?" Leaving that answer vague is the single most common reason people stay at one level for years.
A common and expensive belief: "if I do good work, they'll eventually see it." On a small team, maybe. In a large organization, a promotion decision is made in a meeting you are not in, by people some of whom have never met you, based on what your manager can say about you and what evidence they have for it. If you do not supply the evidence, it usually does not exist. That is not injustice, it is an information limit — and the cure is information, not patience.
6.4 Asking for a project, a raise or a title
Three principles that apply to all three:
One: separate the ask from the review. The performance review is a bad place to ask, because the decisions were usually made before it. The right conversation happens two to three months before the decision cycle.
Two: provide the evidence in advance, not in the moment. Send a page and propose discussing it next time. Your manager has to be able to repeat your case in a meeting you are not in; your job is to prepare that case for them.
Three: make one specific ask, not a general complaint.
❌ "I don't think my compensation is fair."
✅ "I'd like to talk about reaching the next level and to have a plan for it.
Three things I think point at that level: I took the notifications
migration from problem definition through to deployment, I onboarded two
new joiners who now work independently, and the payments postmortem and
its three action items were mine. My question is: what do you think is
still missing? If there's a specific gap, I want to spend the next six
months on it."
Note that this is not a demand; it is an invitation to a shared plan. A manager who hears this, even if today's answer is no, becomes someone working on your behalf. Salary negotiation at hiring time is a different subject and lives in the interview-craft chapter; here we are talking about growing where you already are.
"I don't start by assuming it's a fairness problem; I assume it's an information-flow problem, because in most cases I've seen, it was. I do three things. First, I start sending a short update every two weeks to my manager — not a task list, but a few lines about effect: what got better and how it is measured. Second, I put my work where it is visible on its own: instead of telling one person the result of an investigation, I write a page in the team's shared space; instead of using a good pattern only in my own code, I leave an example and a short explanation so others can use it. Third, I ask directly. In a one-on-one I say: 'I want to make sure your picture of what I'm doing is accurate; if there's somewhere you think my output has been light, I'd rather know now.' One of two things usually comes out: either they genuinely didn't know, which writing fixes, or they knew but expected something different from me, which is also valuable information. What I try not to do is resent it quietly, because that neither fixes the problem nor gets seen."
7) Influence without authority
Most important work in engineering is done by someone who has no formal authority over the people doing it. You cannot give orders to another team, you cannot compel product, you cannot tell a colleague to use your pattern. So either you learn to have influence without authority, or that is your ceiling.
7.1 How a technical proposal gets adopted
A good proposal is not accepted in the meeting; it is accepted before the meeting. If people hear your proposal for the first time in the room, they have to understand, evaluate and decide simultaneously — and a surprised mind is a conservative mind. So the default answer is "no" or "we need to think about it".
The path that works:
- Sell the problem before the solution. If people do not agree a problem exists, the best solution in the world will not be adopted. Gather data first: how many times has this happened, how much time did it cost, how many people complained.
- Talk to people one at a time. Speak to three or four key people separately. The magic question: "I'm going to propose this; what's the biggest flaw in it from your point of view?" That one question finds your weaknesses before public exposure and converts that person from critic into contributor.
- Start small. A small, reversible proposal is a hundred times easier to accept than a large change. "Let's try it on one service and decide with data after a month" almost always gets a yes.
- State the way back. A decision-maker's main fear is being trapped. If you explain how we would undo it, their perceived risk halves.
- Let other people own the credit. If the proposal becomes "the team's" rather than "yours", its chance of being implemented multiplies.
Your social capital is finite and every act of insistence spends some. Two or three times a year you can genuinely plant your feet and be heard; if you do it weekly you become "the person who always objects", and on the one occasion that truly matters your voice sounds like all the others. A practical test for choosing a hill: is this decision expensive to reverse? Does it harm users or security? Will it still matter next year? If all three answers are no, state your view, record it, and move on.
7.2 Understand what other people are measured on
Most "irrational resistance" you meet in an organization is completely rational from the other side. They are simply evaluated on something else.
| Role | What counts as success for them | Their main fear | How to talk to them |
|---|---|---|---|
| Product manager | Delivering user value on a known date | Missing the date with no explanation | Translate cost into date and scope, offer options |
| QA | No bug reaching production | Missing something and being blamed | Involve them early, write acceptance criteria together |
| Operations / infra | Stability and predictability | A sudden change waking them at night | Show the deployment and rollback plan up front |
| Support | Fewer tickets and fewer angry users | A change that creates a ticket wave | Tell them before release, give them prepared wording |
| Security | No vulnerabilities, compliance holds | An exception nobody recorded | Ask early, record the decision in writing |
Once you internalize this table, your phrasing changes. You do not tell the operations team "this architecture is cleaner"; you tell them "with this change, when this service fails at 2am, only that service is affected and the dashboard shows exactly which part broke".
7.3 Talking to non-engineers
The core rule: translate technical risk into one of three things — money, time, or the probability of something bad. No non-technical manager decides on the basis of "technical debt has increased", but everyone decides on the basis of those three.
❌ "The order service codebase is full of technical debt and heavily coupled."
✅ "Any change in the order service currently takes about three times as long
as an equivalent change in our other services, and two of the last three
months' incidents came from there. If we continue, our delivery speed in
this area gets worse every quarter. My proposal is to dedicate 20% of each
quarter's capacity to it, and each quarter I'll show with delivery-time
numbers whether it is working."
And the eternal "why is a rewrite so expensive?" — the best explanation I have found is this: the current system is not just the code you can see; it is several years of small decisions, each made for a real situation, and most of them written down nowhere. A rewrite means rediscovering all of those cases, and until you rediscover them you do not know they exist. That is why rewrites routinely exceed their estimates, and why the best approach is incremental replacement: pull out one piece at a time behind the existing interface, so there is never a moment where everything is new and nothing works.
When date pressure arrives, you are tempted to frame it as "either quality or speed". That framing usually loses, because the listener picks speed and you become the slow person making excuses. A better frame separates deliberate debt from unfinished work: you can say "we can defer automated tests for this area until after release and test it manually — that is a debt with a known cost and I'll file the ticket. But we cannot drop input validation, because that is a vulnerability, not a debt." With that distinction you leave the emotional argument and negotiate over specifics.
When a decision goes against you and you have made your case fully, you have two options: back it sincerely, or object openly. What is sabotage is the third state: saying yes and then implementing half-heartedly, making sideways remarks in other meetings, or waiting for it to fail. If you have accepted a decision, defend it in front of others even when it was not yours — "the team decided this, and here's why", not "I was against it but they insisted". People who can do this get invited into decision-making rooms faster than anyone, because they are dependable.
"I once proposed unifying our data access layer, because there were three different patterns in the codebase and every new joiner got lost. The proposal was rejected, on the grounds that we had an external commitment that quarter and no capacity. The first thing I did was understand the reason precisely — I asked whether they disagreed with the idea itself or with its timing. The answer was 'timing', which changes everything. So rather than dropping it or pushing harder, I did three things: I wrote a page describing what would change this decision in the future, I started collecting data — specifically how long it took each new joiner on average to make their first change in that layer — and in my own day-to-day work, wherever I touched the code, I used the target pattern so a working example existed. The next quarter, when capacity opened up, the proposal came back and took three minutes to approve, because it now had both data and a working example. The lesson I took was that 'no' usually means 'no, not now, not with this information', and my job is to work out which word of that sentence can be changed."
8) Collaboration and conflict
8.1 Disagreement versus conflict
Disagreement means two people hold different views on a topic; it is healthy, and its absence means nobody is thinking. Conflict means the relationship is damaged and from then on the topic is only a pretext. One symptom tells them apart: in a disagreement, both parties are looking for the right answer and new data changes their mind. In a conflict, new data changes nothing.
The mechanism that converts one into the other is almost always the same: accumulation. Something small bothers you and you say nothing; then a second thing, then a third; and on the fourth occasion you react in a way that does not match the size of the event, and the other person is bewildered. The cure is also one thing: raise it early and small.
A small unresolved problem with a colleague behaves exactly like a shortcut in code: cheap today, interest payable tomorrow. A conversation that takes two minutes today is a heavy meeting with a manager present six months from now. The difference from technical debt is that the interest compounds much faster.
8.2 Raising a problem directly, without a fight
The pattern that works in practice has three parts: specific observable behaviour (not an adjective, not a generalization), the effect on the work (not on your feelings, framed as an accusation), and a clear ask.
❌ "You never tell anyone anything and you decide everything on your own."
✅ "There's something I wanted to raise early so it doesn't pile up. The
orders table structure changed yesterday and I found out this morning
when my build broke and I spent two hours looking for the cause. My ask
is that for schema changes we post a message in the team channel first.
If you know a better place for that coordination, I'm happy with any
method."
Three subtleties there: it opens by stating intent ("so it doesn't pile up"), which tells the other person this is not an attack; instead of "you always", it brings one specific, checkable event that cannot be denied; and at the end it makes ownership of the solution shared. And most importantly: this conversation is private. Criticizing behaviour in public, even correct criticism, almost always ends in entrenched defensiveness.
8.3 Different standards, and difficult personalities
A colleague whose standards are lower. First distinguish: do they not know, can they not, or is it not their priority? For "does not know", the answer is teaching, and code review is a good place. For "not a priority", an individual conversation does not work — you have to convert the standard from your preference into a team rule: a written definition of done, an automated check in CI, a test requirement. Tooling and rules are impersonal; person-to-person never is.
A colleague whose standards are higher and who slows you down. This is also real. The answer is an explicit agreement about quality level per type of work: "for this prototype that gets thrown away in two weeks we don't need full unit tests; for the payments service, we do."
A dominating teammate in meetings. Rather than competing for airtime, change the structure: send the agenda beforehand, and in the room use simple tools — "before we continue, let's hear from the others" or "I need two minutes to finish my thought". If it is recurring, raise it privately; many of these people have no idea and genuinely change with one specific piece of feedback.
A silent teammate. Silence is not necessarily agreement. Invite them directly and without pressure, preferably with a specific question rather than an open one: "you've worked with this code more than anyone — where do you think this change breaks?" And if they are not comfortable in a group, collect their view in writing before or after the meeting.
8.4 The escalation ladder
Escalation is not inherently bad; escalating too early or skipping rungs is. The general principle: always start at the lowest rung, give a genuine chance at each rung, and never let the other person be surprised when you go up.
The diagram below shows the escalation ladder, from a direct conversation up to a formal decision · نردبان escalation، از گفتگوی مستقیم تا تصمیم رسمی.
flowchart TD
A[Direct private conversation] --> B{Resolved?}
B -- Yes --> Z[Write down what you agreed]
B -- No --> C[Second try, in writing, with a clear ask]
C --> D{Resolved?}
D -- Yes --> Z
D -- No --> E[Tell the person you will raise it]
E --> F[Bring in a neutral peer or tech lead]
F --> G{Resolved?}
G -- Yes --> Z
G -- No --> H[Managers decide, with facts not adjectives]
| Rung | When | What you say | Common mistake |
|---|---|---|---|
| 1 — Direct conversation | The same week it happened | Specific behaviour + effect + ask | Waiting until it accumulates |
| 2 — Written repeat | When the first attempt had no effect | The same point, shorter, with a deadline | Letting the tone get sharper |
| 3 — Announce intent | Before escalating | "I think we should get X involved" | Escalating without telling them — this reads as betrayal |
| 4 — Neutral third opinion | An unresolved technical disagreement | "Whichever way they go, I'll implement it" | Picking an arbiter who obviously favours you |
| 5 — Managers | Impact on delivery or safety | Facts and effect, no adjectives | Telling a story instead of reporting |
When you reach the manager rung, phrasing decides everything. "He isn't cooperating and is deliberately slowing the work down" is an accusation, and the manager is forced to defend someone. "I've been waiting three weeks for this change to be reviewed, I've followed up twice, and the delivery date is at risk — how can we unblock it?" is a fact plus a request. The first makes you a party to a dispute; the second makes you a problem-solver. Delete the adjectives; leave only dates, counts and effects.
"First I make sure the issue is really quality and not a difference in taste — I collect a few specific examples we can actually discuss, like a bug that reached production or a change that got reverted three times. Then I talk to them privately, with nobody else present, and I open from curiosity rather than accusation: 'I noticed the last few changes went in without tests and we had two regressions afterwards. I wanted to check whether something is getting in your way.' Very often the answer is something I would never have guessed — time pressure from elsewhere, not knowing the test tooling, or a belief that this area does not need tests. If the issue is knowledge, code review and one or two short pairing sessions fix it. If the issue is a shared standard, I take it out of the individual conversation and up to the team level: a written definition of done and an automated check in CI, so the rule is impersonal and has nothing to do with our relationship. Only if nothing changes after two genuine rounds and the impact on team delivery continues do I raise it with a manager — and I tell the person first that I am raising it. What I never do is embarrass them in public or quietly rewrite their code, because neither solves the problem and both damage the relationship."
9) Learning and growth as a practice
9.1 Choosing deliberately what to learn
If you learn without a plan, the market chooses for you — and the market always surfaces the loudest thing, not the most durable. A simple split clarifies the whole decision:
| Layer | Examples | Half-life | How much time to invest |
|---|---|---|---|
| Fundamentals | Concurrency, networking, data modelling, design, performance analysis | Decades | The most — these are capital |
| Ecosystem | Your main language and libraries, build tooling, your database | 5 to 10 years | A lot, but deep rather than wide |
| Current tools | A specific framework, a specific cloud service | 2 to 4 years | Only as far as the job requires |
| Fashion | Whatever is loud this month | Months | Read enough to know what it is |
Practical test: before any large learning investment, ask "if this tool is obsolete in three years, what part of this learning remains?" If the answer is "none", learn only as much as the work needs.
9.2 Getting into a large unfamiliar codebase quickly
Reading code from the first file to the last is the least efficient method possible. What works is entering from the outside behaviour inward:
- Get it running and issue one real request. Until something executes on your machine, all reading is abstract.
- Follow one end-to-end path, from the entry point (controller, queue consumer, CLI command) all the way to the database. One complete path teaches more than ten scattered files.
- Draw the data model. Tables and their relationships are usually the most honest document a system has; code lies, schemas lie less.
- Read the history.
git logon the central files, plus the ADRs. Strange decisions usually had reasons. - Make your first small change early — even a one-line fix. The first PR teaches you more than a week of reading, because it makes you touch the whole build, test and deploy chain.
- Write your map down and publish it. As you learn, write a page. It solidifies your understanding, it survives for the next person, and it is your first visible piece of work.
When you explain something, your brain is forced to find the gaps — the places where you thought you knew but were only familiar. That is why mentoring a new joiner, or writing a page about something you just learned, has a higher learning yield than reading another book. And mentoring requires no formal title; it starts with one instructive code review and half an hour answering the questions of somebody who just arrived.
If your manager is your only source of feedback, your growth is hostage to the quality of that one person — and managers change. Build three other sources: code review (the most direct and most frequent technical feedback you will ever get), two or three colleagues you explicitly ask, and real measurements of your work. And when you ask for feedback, do not ask an open question; "what do you think of my work?" earns a polite answer. Ask: "in that meeting, which part of my explanation was unclear?" A specific question gets a specific answer.
Recognising a plateau. The symptoms are: you have not written anything difficult in months, nobody teaches you anything in reviews, and you can predict next week's work without thinking. Comfort is not automatically bad — a consolidation period is sometimes necessary — but if it runs longer than two or three quarters, your skills are going stale relative to the market. The remedy is usually not changing jobs; it is changing the type of work: an unfamiliar domain inside the same organization, an operational problem, or a leadership role on a project.
10) Professional judgement and ethics
10.1 Owning a mistake in public
This is one of the few places where the ethical behaviour and the self-interested behaviour are identical. The pattern is simple: fast, specific, no excuses, with an action.
I caused this morning's outage. The change I made yesterday does not return
the connection on the error path. We rolled back and the service has been
normal since 10:20. I'll send the fix with a test today, and in the
postmortem I'll write up why our load test did not catch it.
Apologies to the support team, who had a rough morning.
What is absent from that message matters as much as what is present: no justification, no "but the test environment was broken", no performative self-flagellation. Once, cleanly, and then focus on the fix. Someone who owns a mistake this way ends up with more credibility, not less — because everyone now knows that if something breaks, they will hear it from that person.
10.2 Saying "I don't know" with credibility
The formula is: I don't know + what I do know + how I'll find out + when I'll come back.
I'm not sure, and I don't want to guess, because this number is going to be
the basis of a decision. What I do know is that the synchronous path has been
tested to 200 requests per second without problems. Above that I need to
measure.
I'll come back with a real number by tomorrow midday.
Experienced engineers say this often, and that is precisely where their credibility comes from; because when the same person says "this I know", it carries weight. Someone who never says "I don't know" gradually becomes someone whose statements can never be relied upon.
10.3 Refusing a shortcut that harms users
This is where "soft skills" become a professional backbone. Recall the earlier distinction: deliberate debt is negotiable, harm is not. Logging a password, disabling validation to make a deadline, sending data somewhere it is not permitted, or shipping something you know will corrupt financial records — these belong to a different category.
I'm not going to do this, and I want to be clear about why. With this change,
the customer's card number ends up in the logs and anyone with log access can
read it. That isn't a technical preference, it's a violation.
The alternative I can have ready by tomorrow: log only the last four digits,
which is enough for troubleshooting. If the date is under pressure, I can
take the manual support load this week to relieve it.
Three properties: direct and non-threatening, the reason stated as a consequence rather than an abstract principle, and an exit that lets the other person retreat without losing. And if the pressure continues despite that, put it in writing and escalate — not as a threat, but because a decision of that magnitude needs a named owner and a written trail.
Engineering access to real data is a trust, not a privilege. Idly browsing a user's data, taking a data sample onto a personal laptop "just for debugging", pasting a log containing personal data into an external tool, or recounting incident details outside the organization — any one of these can destroy years of credibility in a minute and carry legal consequences. The simple rule: if you need to see data, use the approved and audited path; if there is no approved path, that itself is the answer.
Professional reputation is built over years and destroyed in minutes. The community is smaller than you think: today's colleagues are tomorrow's interviewers and customers. Two things actually persist — whether you did what you said you would do, and how you treated people who had no power over you. No successful project compensates for either.
"First I separate whether this is an acceptable risk or a harm. Plenty of things are risks and the business owns the decision, not me — thin test coverage on a low-importance area, for instance. But something that puts user data or user money at risk is my decision too. If it's the second category, I state the risk as a consequence rather than in technical terms: what can happen, with what likelihood, and what it costs — because as long as the other side hears it as 'an engineering concern', it loses against a delivery date. Then I always bring a third option; people usually apply pressure because they believe there are only two paths. Something like 'let's ship behind a feature flag to 5% of users and I'll finish the secure path by Tuesday.' If, despite all of that, the decision is to ship, I write a short, non-defensive email the same day stating what risk we are accepting and who accepted it, and if it is a security or compliance matter I inform the security team. That is neither building a case against anyone nor covering myself; a decision of that size should have an owner and a written record."
11) Sustainability: staying in the game
11.1 Burnout, before it is too late
Burnout is not "being very tired". Tiredness is fixed by a weekend; burnout is not. Its real symptoms are these, and they usually arrive in this order:
- Cynicism: every new project looks pointless in advance, and sarcasm replaces discussion.
- Falling output at constant effort: you work the same hours but produce half as much, and even simple tasks have no activation energy.
- Indifference to quality: what used to bother you no longer does. This is the most dangerous sign, because it feels like calm.
- Physical and cognitive signs: poor sleep, irritability, forgetfulness, inability to focus even in free time.
What actually helps is usually not "more rest" but a change in three things: volume, control and meaning. If you feel you have no control over any of them, say exactly that to your manager — with a specific example rather than a general complaint: "for three months I've been on two projects plus on-call and none of them is going well. I want to put one down completely; which one?"
If you answer at night and on weekends, two things happen and neither is in your interest. First, expectations reset: what was a sacrifice yesterday is the minimum expectation today, and the day you do not answer is read as "less committed". Second, structural problems stay hidden; if you keep the system alive with extra hours every time, nobody discovers that the system depends far too heavily on one person and needs fixing. Heroics get short-term praise and long-term burn you out while preserving the problem.
Healthy on-call has a few simple conditions worth insisting on: every alert that wakes you must be actionable (if there is nothing to do in that moment, that alert should be deleted or turned into a daily report), every night page should be reviewed in the next team meeting, after a bad night you sleep the following morning rather than sitting in meetings, and everyone who writes the code takes a turn on-call — the single most effective incentive for building reliable systems.
11.2 Focus, and working effectively remotely
Engineering needs continuous blocks; four hours chopped up by three meetings is not equivalent to four hours at all. A few things that genuinely work: reserve two two-hour blocks in your calendar and treat them like meetings; cluster meetings into one half of the day; turn off notifications during a focus block and say in your status when you will be back; and start each day by deciding "if only one thing finishes today, what is it".
In distributed work across time zones, the rule changes: make asynchronous the default. That means everything is written and retrievable, decisions happen in public channels rather than private messages, and instead of "got a sec?" you send a complete message with all the context so the other person can answer eight hours later without a round trip. A degree of "over-communicating" while remote is almost always correct — because what was overheard incidentally in an office simply does not exist remotely.
And one subtlety: do not confuse presence with looking busy. A quick green dot builds no trust. What builds trust is predictability: messages get answered within a known window, commitments land on time, and the state of the work is visible without anyone asking.
11.3 The first ninety days in a new role
| Period | Main focus | Concrete actions | Sign it worked |
|---|---|---|---|
| Week 1 | Setup and people | Get the dev environment working · 30 minutes with each teammate · write your own glossary of internal terms | A first small PR merged |
| Weeks 2–4 | Understanding the system | Trace one end-to-end path · draw the data model · read the ADRs and recent postmortems | You can sketch the architecture on a whiteboard for a newcomer |
| Weeks 5–8 | Real contribution | One complete feature from definition to deployment · start reviewing others' code · first on-call shadowing an experienced engineer | You work without daily help |
| Weeks 9–12 | Making things better | Fix one pain point you found yourself · update the onboarding doc with what was missing | The team comes to you for something |
| End of 90 days | Resetting expectations | With your manager: what went well, what should I focus on from here | An agreed plan for the next quarter |
Being new is a short window in which you can ask anything without anyone being surprised. Use it and ask everything, especially "why is it like this?". At the same time, keep a file of everything that confused you; three months later that file becomes the best possible feedback on the team's onboarding, and your first visible contribution. One warning: in the first month, observe and suggest, but do not use "where I worked before we did it this way" as an argument — until you understand why the current situation exists, a proposal to change it will not be heard.
"I've learned to watch for two signals, because both arrive earlier than exhaustion does. First, when I notice I have become indifferent to things that used to matter to me — for example the quality of my own code stops mattering. Second, when my default reaction to a new request is sarcasm or cynicism. If I see those, rather than working harder I slow down and find the cause; it is usually one of three things: volume beyond capacity, no control over the work, or work that has stopped meaning anything to me. Each needs a different response. The one thing I've learned to always do is talk to my manager early and specifically, not as a general complaint but as a concrete request: 'I'm on three things at once and none of them is going well; I want to put one down completely — which one?' I have almost always gotten a result, because managers usually have no idea of the real load. And I keep one fixed boundary for myself: outside on-call I do not answer at night, because in my experience every exception turns into a rule."
12) An operating system for yourself
None of this changes with one big decision; it changes with a rhythm. Here is a version you can run starting this week.
Daily (about 15 minutes, outside the actual work)
- Morning: write one sentence — "if only one thing finishes today, what is it?"
- One protected two-hour block with no notifications.
- Clear pending reviews, by end of your working day at the latest.
- If something slipped, say it today — not tomorrow.
Weekly (about 30 minutes)
- Friday: ten minutes of work journal — what I did, what I learned, where I got stuck.
- One short update to your manager or stakeholders: status, change, risk, ask.
- Say "no", or "yes but at this cost", at least once instead of a thoughtless yes.
- Make one small thing better for everyone else: a page of docs, a test, a tool.
Monthly
- Promote the work journal into the brag document: problem → action → effect → witness.
- Ask one colleague for one specific piece of feedback.
- Ask: what did I learn this month that will still be useful in three years?
Quarterly
- With your manager: what is the gap to the next level, with a concrete example.
- Pick the one hill you will stand on this quarter — and let the rest go.
- Answer the short self-assessment below again.
Self-assessment (yes or no each time, honestly)
- When did I last raise a risk before it became a problem?
- If I disappeared today, could someone continue my work from what I have written?
- Did my last estimate have a range and stated assumptions, or was it a bare number?
- In the last three months, how many people's work got better because of me?
- When did I last say "I don't know" or "that was my mistake" in public?
- Could my manager name three specific accomplishments of mine without me present?
- When did I last raise a disagreement early and directly instead of letting it accumulate?
- Did I learn something hard this quarter, or did I just repeat myself?
Every "no" is one specific task for next week — not a verdict on yourself.
"My honest answer isn't about technical knowledge, because I know how to acquire that. What I want to work on is influencing decisions that are made outside my own team. Today I can get a technical decision adopted inside my team, because people know me and trust my judgement. But when something spans several teams I'm still not good at it — particularly at working out what each side is measured on and how I should frame a proposal for them. What I'm doing about it is concrete: I write more, because at multi-team scale only writing travels, not presence; before any large proposal I talk to key people individually instead of unveiling it in a meeting; and I try to frame the problem so other teams can own it, rather than insisting on my own solution. I even have a measure for myself: how often a decision I started ends up moving forward without my name attached to it. To me that is the sign it worked, not a loss of credit."
Past a certain point, your ceiling is not technical knowledge; it is communication, trust and judgement — and these are teachable skills, not personality traits. Write: messages with the bottom line up front and one ask, questions with goal, attempts and the exact error, design documents with options and risks, ADRs for expensive-to-reverse decisions, PR descriptions that make review fast, status updates that surface risk early, and postmortems that find broken layers instead of culprits. In code review, label the severity of your comments, discuss design before style, and review quickly; on the other side, separate the code from yourself and break deadlocks with a third opinion. In meetings, record the decision, the owner and the date, disagree by restating first, and instead of "no", make the cost explicit and hand the decision to whoever owns priority. In estimation, give a range and assumptions, label your confidence level, and raise a slip the moment you suspect it, with priced options. Your manager does not see your work: keep a work journal and a brag document, set your own one-on-one agenda, and make the gap to the next level specific. Influence without authority comes from pre-wiring, from understanding what other people are measured on, from starting small, and from choosing only a few hills a year; with non-engineers, translate risk into money, time and probability. Raise conflict early, small and privately, start at the lowest rung of the escalation ladder, and replace adjectives with dates and numbers. Invest your learning in fundamentals rather than fashion, do not let your feedback loop depend on one person, and learn by teaching. Judgement means owning a mistake fast, saying "I don't know" with credibility, and drawing a line between deliberate debt and real harm. And sustainability underwrites all of it: recognise burnout by cynicism and indifference, set boundaries, protect your focus, and work asynchronously and predictably. None of this arrives through one big decision — it arrives through a weekly rhythm and a handful of small habits that compound over years.