أبحاث

تصميم البروتين الرابط BindCraft2: كيف يتم توليد وتقييم الروابط الخاصة بالهدف؟

BindCraft2 Protein Binder Design: How to Generate and Score Target-Specific Binders?

المجال
biotech & health
نُشر
دقائق قراءة
33 min · 4,865 كلمة
الادعاءات والمراجعة
39/40 ادعاءات موثقة · 20 مصادر

الطبعات: English · Español · Français

الجواب المباشر

BindCraft هو خط أنابيب مفتوح المصدر وآلي يُنشئ (يُهلوس) بنية رابط (binder) مباشرة مقابل بنية هدف ثابتة باستخدام أوزان AlphaFold2 الخاصة به، ويقوم بتحسين درجة TM المتوقعة للواجهة (ipTM) والخسائر ذات الصلة عبر الانحدار التدريجي (gradient descent)، وقد أُفيد بأنه يصل إلى معدلات نجاح تجريبية تتراوح من 10% إلى 100% عبر أهداف متنوعة تشمل مستقبلات سطح الخلية، ومسببات الحساسية، وCRISPR-Cas9. أما المتغير الأحدث، BindEnergyCraft، فيحافظ على نفس حلقة التحسين لكنه يستبدل ipTM بـ pTMEnergy، وهو إشارة كثيفة قائمة على الطاقة مُشتقة من مخرجات خطأ المحاذاة المتوقع (pAE) نفسها، وقد حسّن هذا الاستبدال معدلات النجاح الحاسوبية (in silico) وقلّل التصادمات البنيوية مقارنةً بـ BindCraft وRFDiffusion وESM3 على المعايير المرجعية المختبَرة. أما التقييمات المخبرية المستقلة لـ BindCraft فتُظهر نتائج متباينة: نتائج قوية لببتيدات MDM2 وWDR5، لكن فشلاً تاماً في إيجاد روابط لـ PD-1 وPD-L1 في دراسة واحدة، ولم يُظهر سوى تصميم واحد من أصل أربعة تصاميم موجهة بنقاط ساخنة (hotspot) ضد نطاق LEDGF PWWP تقارباً حقيقياً في سير عمل مستقل ومختلف منهجياً. لا يقدم أي ادعاء في المصادر المتوفرة مقارنة تجريبية مباشرة بين BindCraft وBindEnergyCraft، أو بين BindCraft وRFDiffusion/AlphaProteo، على نفس الأهداف المخبرية، لذا يجب قراءة أي ترتيب بين هذه الطرق على أنه حاسوبي (in silico) فقط، على المعايير المذكورة، ما لم تُجرِ ورقة بحثية صراحةً الفحص المخبري. يمكن للمطوّر تشغيل خط أنابيب التصميم إلى المرشح محلياً باستخدام أوزان AlphaFold2 وملف PDB للهدف، لكن ينبغي التعامل مع أي درجة ipTM أو pTMEnergy على أنها مرشح فرز (screening filter)، وليست ضماناً للارتباط، لأن الارتباط بين مقاييس الثقة والنتائج الفعلية للارتباط تم التحقق منه على مجموعات بيانات محدودة فقط.

لماذا تُعتبر خطوط أنابيب تصميم الروابط مهمة وما المشكلة التي تحلها

يُعد تصميم رابط بروتيني (protein binder) يتعرف على هدف واحد بتقارب عالٍ وخصوصية عالية حاجة أساسية في علم الأحياء والطب، والقيام بذلك انطلاقاً من بنية الهدف وحدها، دون جولات طويلة من التحسين التجريبي، من شأنه إزالة عائق رئيسي. يبقى تحديد المرشحين ذوي التقارب العالي عائقاً لأن تصميم الروابط يتطلب عادةً توليد وفرز آلاف التصاميم افتراضياً لاستخراج عدد قليل من النتائج الواعدة [5]. هذا يعني أن أي خط أنابيب يمكنه رفع نسبة التصاميم المُولّدة حاسوبياً التي تتبين لاحقاً أنها ترتبط فعلياً في المختبر له قيمة عملية مباشرة: عدد أقل من التصاميم يحتاج إلى تصنيع واختبار، مما يوفر الوقت وتكلفة الكواشف.

هناك عائلتان من الأساليب الحاسوبية تعالجان هذه المشكلة. الأساليب القائمة على البنية مثل RFDiffusion وAlphaProteo يمكنها توليد هياكل أساسية (backbones) للروابط مشروطة ببنية الهدف [5]. أما العائلة الأخرى، الأساليب القائمة على الهلوسة (hallucination-based)، فتقوم بتحسين بنية مباشرة داخل شبكة التنبؤ بالبنية بدلاً من أخذ عينات من نموذج بنية أساسية توليدي. BindCraft هو طريقة تصميم روابط قائمة على الهلوسة تستفيد من شبكة AlphaFold2 [5]. هذا التمييز مهم للمطوّر لأنه يحدد البرمجيات والأوزان المطلوبة: طريقة الهلوسة تحتاج إلى متنبئ بنية قابل للاشتقاق (differentiable) وحلقة تحسين قائمة على الانحدار التدريجي، بينما طريقة الانتشار (diffusion) للهيكل الأساسي تحتاج إلى نموذج انتشار مدرَّب بالإضافة إلى خطوة منفصلة لتصميم التسلسل.

مسألة مدى جودة هذه الروابط فعلياً بعد تصنيعها منفصلة عن كيفية عمل خط الأنابيب، والأدلة على هذه النقطة متفاوتة بين الأوراق البحثية. أفاد BindCraft بمعدلات نجاح تجريبية تتراوح من 10% إلى 100% [7]، وهو نطاق واسع يعتمد بشدة على فئة الهدف. تترابط درجات الثقة من نماذج التنبؤ بالبنية، وخاصةً ipTM (interface pTM) بين الواجهات، مع قوة الارتباط [5]، لكن ترابط مقاييس الثقة مثل pLDDT وipTM وipAE مع نتائج الارتباط التجريبية تم التحقق منه على مجموعات بيانات محدودة فقط ولا يزال مفهوماً بشكل ضعيف عبر الأهداف وأنظمة التصميم المختلفة [6]. لذا ينبغي للمطوّر ألا يتعامل مع أي رقم ثقة منفرد كدليل على الارتباط، بل كمرشح أظهر بعض الترابط في بعض السياقات فقط.

بسبب هذه الفجوة بين الثقة الحاسوبية والنتيجة المخبرية، أصبحت جهود وضع المعايير المرجعية (benchmarking) وتحسينات دوال التقييم مجالات عمل نشطة إلى جانب خطوط الأنابيب التوليدية نفسها. صُمم BindEnergyCraft خصيصاً لتقديم إشارة تحسين أكثر كثافة وإفادة من ipTM [5]، وصُمم ProtDBench خصيصاً لتقديم بروتوكول تقييم موحّد ومتعدد المُحقّقين (verifiers) لمقارنة الروابط المُولَّدة [6]. مجتمعةً، يمثل هذان المساران، أهداف تحسين أفضل ومعايير مرجعية أفضل، ما يحتاج المطوّر لفهمه قبل اتخاذ قرار بشأن خط الأنابيب ودالة التقييم المناسبين لهدف معين.

ما هو BindCraft وفيمَ استُخدم

BindCraft هو خط أنابيب مفتوح المصدر وآلي لتصميم روابط بروتينية من الصفر (de novo) [7]. يستفيد من أوزان AlphaFold2 لتوليد روابط بروتينية انطلاقاً من بنى الأهداف [7]. هذا يعني أن المستخدم لا يدرّب نموذجاً جديداً: بدلاً من ذلك، يشغّل خط الأنابيب إجراء تحسين داخل رسم الحساب الموجود لـ AlphaFold2، معاملاً أجزاء من المدخل (تسلسل الرابط أو بنيته) كمتغيرات حرة وكل شيء آخر (بنية الهدف) كثابت.

طُبّق خط الأنابيب على نطاق واسع. طُبّق BindCraft على مستقبلات سطح الخلية، ومسببات الحساسية الشائعة، والبروتينات المصممة من الصفر، والنيوكليازات متعددة النطاقات مثل CRISPR-Cas9 [7]. في عدة تطبيقات من هذه، أظهرت الروابط المصممة تأثيرات وظيفية تتجاوز مجرد الارتباط. أُفيد بأن الروابط المصممة تقلل من ارتباط IgE بمسبب حساسية شجر البتولا، وتُعدّل نشاط التحرير الجيني لـ Cas9، وتقلل من السمية الخلوية لسم معوي بكتيري منقول بالغذاء، وتُعيد توجيه غلافات الفيروس الغدي المرافق (AAV) لتوصيل الجينات المستهدف [7]. في أحد العروض التوضيحية، ولّد BindCraft روابط بتقارب نانومولاري دون فرز عالي الإنتاجية أو تحسين تجريبي، بما في ذلك لأهداف بلا مواقع ارتباط معروفة [7].

اختبرت مجموعات مستقلة منذ ذلك الحين BindCraft على فئات أهداف محددة لمعرفة مدى صمود هذه الادعاءات خارج الورقة البحثية الأصلية. بالنسبة إلى MDM2، ولّد BindCraft 70 ببتيداً فريداً، وصُنّع منها 15، وأظهرت 7 ارتباطاً نوعياً بتقاربات نانومولارية [12]. أكّدت فحوصات المنافسة الارتباط الخاص بالموقع لببتيدات MDM2 عند موقع الهدف المقصود [12]، وهو تحقق مهم لأن الارتباط غير النوعي بالموقع لن يكون مفيداً لمعظم التطبيقات. بالنسبة إلى WDR5، ارتبط ستة من تسعة مرشحين بموقع WBM الرابط لـ MYC بتقارب دون-ميكرومولاري [12]، وحسّن التعديل الكيميائي العقلاني فعالية أحد روابط WDR5 بمقدار ستة أضعاف إلى KD يبلغ 39 نانومولار [12]. هذا يُظهر أن مخرجات BindCraft يمكن أن تكون نقطة انطلاق لمزيد من التحسين الكيميائي الطبي، وليست إجابة نهائية فقط. مع ذلك، لم يُظهر أيٌّ من الببتيدات المُولَّدة بواسطة BindCraft المختبَرة لـ PD-1 وPD-L1 ارتباطاً قابلاً للكشف [12]، وهذا مثال مضاد مباشر يُظهر أن النجاح يعتمد على الهدف وليس مضموناً. يمكن لـ BindCraft توليد ببتيدات عالية التقارب انطلاقاً فقط من بنية الهدف [12]، لكن نتيجة PD-1/PD-L1 تُظهر أن هذه القدرة لا تنتقل بشكل موحد إلى كل هدف.

على نطاق أوسع، يُدمج Human Bindome طريقة BindCraft المُختبَرة تجريبياً في إطار عمل مُسرّع ومتوازٍ مع اختيار آلي للأهداف على مستوى النطاق (domain-level) [13]. باستخدام هذا الإطار، تم توليد 306,146 مرشح رابط لـ 8,296 بروتيناً بشرياً، مما يغطي 40.9% من البروتيوم الكامل [13]. يتضمن كل مرشح في Human Bindome تسلسلاً محدداً، ونموذج بنية متوقع للرابط-الهدف، ومقاييس ثقة حاسوبية (in silico) [13]، مما يوفر للمطوّر نموذجاً لما ينبغي أن يحتويه سجل مرشح عند تشغيل خط أنابيب على نطاق واسع.

آلية التحسين: الهلوسة، ipTM، ومشكلتها في التدرج

يُهلوس BindCraft بنى الروابط باستخدام AlphaFold2 ويُجري تحسيناً قائماً على التدرج (gradient) لدرجة TM المتوقعة للواجهة (ipTM)، إلى جانب خسائر أخرى [5]. عملياً، هذا يعني أن تسلسل الرابط (وضمنياً بنيته، كما يتنبأ بها AlphaFold2) يُحدَّث بشكل تكراري عبر الانحدار التدريجي بحيث تزداد درجة ipTM التي يمنحها AlphaFold2 للمركّب المتوقع رابط-هدف، إلى جانب حدود خسارة أخرى. هذه آلية مختلفة جوهرياً عن نموذج انتشار يأخذ عينة من هيكل أساسي ثم يصمم تسلسلاً له: هنا تعمل شبكة التنبؤ بالبنية نفسها كدالة الهدف التي يجري تحسينها.

لاختيار ipTM كهدف تحسين رئيسي ضعف معروف. لا تعكس ipTM بشكل مباشر الاحتمالية الإحصائية للتفاعل الكامل بين الرابط والهدف [5]. هذا مهم لأن هدف تصميم الروابط هو إيجاد تسلسلات من المرجح أن تطوى فعلياً إلى مركّب مستقر ومُلتحم بشكل صحيح، ودرجة لا تتتبع هذه الاحتمالية يمكن أن تدفع المُحسِّن نحو بنى تبدو واثقة وفق الدرجة لكنها ليست بالضرورة مركّبات مواتية فيزيائياً.

السبب الميكانيكي المحدد لهذا الضعف يكمن في كيفية حساب ipTM. نظراً لأن ipTM تحسب الحد الأقصى (maximum) عبر مؤشرات بقايا الهدف (residue indices)، فإنها تُنتج تدرجات متناثرة (sparse gradients) متركزة على مجموعة فرعية صغيرة فقط من أزواج بقايا الواجهة [5]. في التحسين القائم على التدرج، يعني التدرج المتناثر أن مواقع قليلة فقط في البنية تتلقى إشارة تحديث ذات معنى في كل خطوة، بينما لا تتشكل معظم الواجهة مباشرةً بواسطة الهدف. يمكن أن يؤدي هذا إلى تصاميم تُرضي الدرجة عند أزواج قليلة من البقايا دون أن تتشكل بقية الواجهة بشكل جيد، وهو سبب معقول للتصادمات البنيوية أو رداءة جودة المركّب الكلية في التصاميم المُولَّدة.

هذا التشخيص الميكانيكي، الموضَّح صراحةً في ورقة BindEnergyCraft، هو ما حفّز البحث عن دالة تقييم بديلة تمنح إشارة تدرج أكثر كثافة عبر الواجهة كاملةً بدلاً من إشارة متناثرة. أما بقية إطار التحسين، حلقة الهلوسة القائمة على AlphaFold2 والمفهوم العام للتحديثات القائمة على التدرج للبنية والتسلسل، فتبقى كما هي؛ فقط درجة الواجهة التي يجري تحسينها هي ما يحتاج إلى تغيير.

pTMEnergy: دالة تقييم أكثر كثافة ونتائج BindEnergyCraft

تُعيد pTMEnergy تفسير لوغاريتمات احتمالية خطأ المحاذاة المتوقع (pAE logits) من AlphaFold كنموذج قائم على الطاقة (energy-based model) على مركّبات الرابط-الهدف [5]. بدلاً من أخذ ملخص عددي واحد (كما تفعل ipTM، عبر عملية حد أقصى بين أزواج البقايا)، تعامل pTMEnergy التوزيع الكامل للوغاريتمات pAE على أنه يُعرّف مشهد طاقة، مانحةً درجة تستمد من حيث المبدأ معلومات من عبر الواجهة المتوقعة كاملةً وليس من مجموعة فرعية صغيرة منها فقط.

توفر pTMEnergy إشارة كثيفة قابلة للاشتقاق يُقصد بها أن تعكس احتمالية مركّب مطوي وفق التوزيع المتعلَّم من AlphaFold [5]. لأنها قابلة للاشتقاق، يمكن استبدالها مباشرةً في نفس حلقة التحسين القائمة على التدرج التي يستخدمها BindCraft بالفعل، دون الحاجة إلى نوع مختلف من المُحسِّنات. هذا يستهدف مباشرةً ضعف التدرج المتناثر الموصوف أعلاه: الإشارة الكثيفة تعني أن مزيداً من أزواج بقايا الواجهة تُسهم بتدرج ذي معنى في كل خطوة تحسين، على عكس حساب ipTM القائم على الحد الأقصى.

تفصيل تنفيذي رئيسي للمطوّر هو أن pTMEnergy مُشتقة من نفس مخرجات pAE المستخدمة في حساب ipTM ولا تتطلب استدعاءات نموذج إضافية أو تعديلات معمارية [5]. هذا يعني أن اعتماد pTMEnergy لا يتطلب إعادة تدريب AlphaFold2 أو إضافة مكونات شبكة جديدة: لوغاريتمات pAE التي يخرجها AlphaFold2 بالفعل تُعالَج ببساطة بشكل مختلف لإنتاج الدرجة الجديدة. هذه نقطة عملية مهمة لأنها تبقي التكلفة الحسابية لخط الأنابيب المعدَّل مماثلة تقريباً لتكلفة خط أنابيب BindCraft الأصلي، حيث لا حاجة لتمريرات أمامية أو خلفية إضافية عبر نموذج مختلف.

الطريقة الناتجة، BindEnergyCraft، والتي تحافظ على إطار تحسين BindCraft مع استبدال ipTM بـ pTMEnergy، حققت معدلات نجاح حاسوبية أعلى للروابط وتصادمات بنيوية أقل من BindCraft وRFDiffusion وESM3 عبر أهداف صعبة متعددة [5]. هذه مقارنة حاسوبية (in silico)، جرى تقييمها عبر الأهداف الصعبة المستخدمة في معيار تلك الورقة المرجعي؛ لا تُبلغ ادعاءات المصادر عن مقارنة مباشرة مخبرية بين BindEnergyCraft وBindCraft وRFDiffusion أو ESM3 على نفس المرشحين المُصنَّعين، لذا ينبغي للمطوّر قراءة التحسّن على أنه اكتشاف حاسوبي على مستوى المعيار المرجعي فقط، وليس ميزة تجريبية مؤكدة على هذه البدائل المحددة.

وضع معايير مرجعية لتصميم الروابط: ProtDBench واختلاف المُحقّقين

لأن درجات الثقة من نماذج التنبؤ بالبنية معروفة بأنها بدائل غير كاملة للارتباط الحقيقي، ولأن موثوقيتها موصوفة بشكل ضعيف عبر الأهداف، هناك حاجة إلى معيار مرجعي موحّد لمقارنة طرق ودوال تقييم تصميم الروابط. ProtDBench هو أحد هذه المعايير، صُمم خصيصاً لتقييم الهياكل الأساسية للروابط المُولَّدة بطريقة قابلة للتكرار ومتعددة المُحقّقين (verifiers).

يقيّم ProtDBench الهياكل الأساسية للروابط المُولَّدة عبر أخذ عينات من تسلسلات متعددة لكل هيكل أساسي، وتطبيق مُحقّق (verifier) ومرشّح (filter) على كل تسلسل، وحساب معدل النجاح لكل تسلسل [6]. يفصل هذا الإجراء بين مرحلتي تصميم غالباً ما تُخلط: خطوة توليد الهيكل الأساسي (التي تحدد الطية) وخطوة تصميم التسلسل (التي تحدد هوية الأحماض الأمينية الفعلية المُدرَجة على تلك الطية). عبر أخذ عينات من عدة تسلسلات لكل هيكل أساسي، يمكن للمعيار المرجعي التمييز بين الهياكل الأساسية السهلة لتصميم تسلسلات جيدة لها والهياكل الأساسية الناجحة أحياناً فقط.

يُعرّف ProtDBench مجموعة الهياكل الأساسية الناجحة (passing backbone set) بأنها الهياكل الأساسية التي يجتاز فيها تسلسل واحد على الأقل من التسلسلات المأخوذة كعينة المرشّح [6]. هذا معيار متساهل، أفضل-واحد-من-كثيرين، على مستوى الهيكل الأساسي: يُحسب الهيكل الأساسي نجاحاً إذا أنتج تسلسلاً واحداً جيداً على الأقل، حتى لو فشلت معظم التسلسلات المأخوذة كعينة له. يجمّع ProtDBench الهياكل الأساسية الناجحة فقط حسب التشابه البنيوي باستخدام FoldSeek أو TM-score [6]، مما يتيح للمعيار المرجعي الإبلاغ ليس فقط عن عدد نجاحات خام بل أيضاً عن مدى التنوع البنيوي للتصاميم الناجحة، متجنباً حالة أن تكون كل النجاحات المُبلَّغ عنها نسخاً شبه متطابقة من نفس الطية.

اكتشاف مركزي في هذا العمل المرجعي هو تحذير من الثقة المفرطة بأي مُحقّق منفرد. يُبلغ ProtDBench عن تحيّز كبير مرتبط بالمُحقّق واتفاق محدود بين نماذج التنبؤ بالبنية المستخدمة كمُحقّقين للتقييم [6]. هذا يعزز مباشرةً النقطة السابقة القائلة إن ipTM ومقاييس الثقة ذات الصلة لها ترابط مُختبَر مع الارتباط الحقيقي على مجموعات بيانات محدودة فقط [6]: إذا كانت نماذج التنبؤ بالبنية المختلفة المستخدمة كمُحقّقين تختلف بشكل كبير مع بعضها البعض، فلا يمكن افتراض أن معدل النجاح المحسوب بمُحقّق واحد (على سبيل المثال ipTM القائم على AlphaFold2) سيصمد عند استخدام مُحقّق مختلف بدلاً منه (على سبيل المثال AlphaFold3). بالنسبة للمطوّر، هذا يعني أن الإبلاغ عن نتائج من مُحقّق واحد دون الإفصاح عن المُحقّق المستخدم، أو دون التحقق المتقاطع مقابل مُحقّق ثانٍ، يخاطر بالمبالغة في الثقة بتصميم ما.

سير عمل بديل: تصميم موجّه بنقاط ساخنة مع RFdiffusion وProteinMPNN

ليست جميع خطوط أنابيب تصميم الروابط قائمة على الهلوسة. سير عمل منفصل، قائم على البنية، يوضح كلا الوعد والصعوبة العملية للانتقال من تصميم حاسوبي إلى رابط مُتحقَّق منه. دمج سير عمل تصميم روابط موجّه بنقاط ساخنة (hotspot) بين RFdiffusion وProteinMPNN وAlphaFold ومحاكاة الديناميكا الجزيئية والتقييم البنيوي اليدوي [3]. يجمع سير العمل هذا بين خطوة انتشار للهيكل الأساسي (RFdiffusion)، وخطوة تصميم تسلسل (ProteinMPNN)، وفحص ثقة البنية (AlphaFold)، وخطوة محاكاة فيزيائية (الديناميكا الجزيئية)، ومراجعة بشرية، وهو خط أنابيب أثقل ويُشرَف عليه يدوياً بشكل أكبر من حلقة BindCraft الآلية أحادية الشبكة.

ولّد سير العمل الموجّه بنقاط ساخنة أربعة تصاميم بروتينية خضعت للتحقق التجريبي [3]. هذا عدد صغير من التصاميم مقارنةً بـ

عائلات نماذج ذات صلة ينبغي أن يعرفها المطوّر

تعتمد خطوط أنابيب تصميم الروابط بشكل متزايد على نماذج اللغة البروتينية (pLMs) كأداة تكميلية أو بديلة لتوليد التسلسل. تُدرَّب نماذج اللغة البروتينية (pLMs) مسبقاً على بيانات تسلسلات تطورية واسعة النطاق وتُكيّف نماذج اللغة الكبيرة للتسلسلات البيولوجية [17]. هذه آلية مختلفة عن الهلوسة القائمة على AlphaFold2 أو توليد الهيكل الأساسي القائم على RFdiffusion: يتعلم نموذج اللغة البروتيني الأنماط مباشرةً من قواعد بيانات تسلسلات واسعة وليس من بيانات تدريب بنيوية، ويُستخدم عادةً لاقتراح أو تقييم التسلسلات وليس الإحداثيات ثلاثية الأبعاد.

توجد عدة معماريات محددة ضمن هذه العائلة، وينبغي للمطوّر الذي يختار من بينها أن يعرف الاختلافات البنيوية بينها. ESM2 هو نموذج مُحوِّل ترميز (encoder transformer) على طراز BERT، بينما ProtGPT2 وProGen2 نموذجا مُحوِّل فك ترميز (decoder transformer) على طراز GPT، وProtT5 نموذج مُحوِّل ترميز-فك ترميز (encoder-decoder transformer) [17]. النموذج المُقتصر على الترميز مثل ESM2 يُستخدم عادةً لتقييم أو تضمين تسلسل موجود، والنموذج المُقتصر على فك الترميز مثل ProtGPT2 أو ProGen2 يُستخدم عادةً لتوليد تسلسلات جديدة بشكل ذاتي الانحدار، والنموذج ترميز-فك ترميز مثل ProtT5 يمكن استخدامه لمهام تُحوّل تمثيل تسلسل إلى آخر. لا يصف أي من الادعاءات المُستشهد بها إحلال نماذج اللغة البروتينية مباشرةً في حلقة تحسين BindCraft أو BindEnergyCraft، لذا ينبغي للمطوّر التعامل مع توليد التسلسل القائم على نماذج اللغة البروتينية على أنه أداة تكميلية منفصلة لاقتراح أو تصفية التسلسلات المرشحة، وليس بديلاً موثّقاً للاستخدام المباشر بدلاً من الهلوسة القائمة على AlphaFold2 في هذه الخطوط الأنابيب المحددة.

الحدود والأسئلة المفتوحة

القيد المركزي لأي شخص يبني على هذا الأدب البحثي هو أن درجات الثقة ليست دليلاً على الارتباط. تم التحقق من ترابط مقاييس الثقة مثل pLDDT وipTM وipAE مع نتائج الارتباط التجريبية على مجموعات بيانات محدودة فقط ولا يزال مفهوماً بشكل ضعيف عبر الأهداف وأنظمة التصميم [6]. هذا يعني أن درجة ipTM أو pTMEnergy المرتفعة ينبغي أن تُقرأ كمرشح يزيد من فرصة إيجاد نتيجة إيجابية، وليس ضماناً، وأي معدل نجاح مُبلَّغ عنه مرتبط بفئة الهدف ومجموعة البيانات المحددة التي قِيس عليها.

تحمل ipTM نفسها ضعفاً ميكانيكياً معروفاً. لا تعكس ipTM بشكل مباشر الاحتمالية الإحصائية للتفاعل الكامل بين الرابط والهدف [5]، ولأنها تحسب الحد الأقصى عبر مؤشرات بقايا الهدف، فإنها تُنتج تدرجات متناثرة متركزة على مجموعة فرعية صغيرة فقط من أزواج بقايا الواجهة [5]. هذا سبب محدد وميكانيكي لتفضيل pTMEnergy كهدف تحسين عند إعادة إنتاج خطوط أنابيب على طراز BindCraft، رغم أن المقارنة بين BindEnergyCraft وBindCraft على أهداف مخبرية حقيقية لم يُبلَّغ عنها في الادعاءات المُستشهد بها، فقط المقارنة الحاسوبية عبر أهداف المعيار المرجعي [5].

تُعقّد اختيارات المعايير المرجعية والمُحقّقين المختلفة أيضاً أي مقارنة عبر الأوراق البحثية. يُبلغ ProtDBench عن تحيّز كبير مرتبط بالمُحقّق واتفاق محدود بين نماذج التنبؤ بالبنية المستخدمة كمُحقّقين للتقييم [6]، لذا لا يمكن مقارنة معدلات النجاح المحسوبة بمُحقّقين مختلفين، أو المُبلَّغ عنها من قِبل أوراق بحثية مختلفة تستخدم خطوط أنابيب مختلفة، مباشرةً دون معرفة أي مُحقّق استُخدم. وبالمثل، ليست النتائج على فئات أهداف محددة موحدة: نفس الطريقة العامة (BindCraft) نجحت بقوة في MDM2 وWDR5 [12] لكنها لم تُنتج أي روابط قابلة للكشف لـ PD-1 وPD-L1 [12]، وسير عمل منفصل موجّه بنقاط ساخنة يستخدم أدوات مختلفة تماماً (RFdiffusion وProteinMPNN وAlphaFold ومحاكاة الديناميكا الجزيئية [3]) نجح في تصميم واحد فقط من أصل أربعة تصاميم ضد نطاق LEDGF PWWP [3]، وهذا النجاح الوحيد معقّد بتكوّن ثنائي غير مرغوب فيه (homodimerization) [3]. لا ينبغي تعميم أي من هذه النتائج على أهداف أو ظروف لم تُختبَر مباشرةً.

أخيراً، يبقى تحديد المرشحين ذوي التقارب العالي عائقاً لأن تصميم الروابط يتطلب عادةً توليد وفرز آلاف التصاميم افتراضياً لاستخراج عدد قليل من النتائج الواعدة [5]. حتى مع دالة تقييم محسّنة مثل pTMEnergy، أو جهد توليد واسع النطاق مثل Human Bindome الذي يغطي 40.9% من البروتيوم البشري [13]، لم تُلغَ الحاجة الأساسية لتصفية مجموعة مرشحين كبيرة إلى عدد صغير من التصاميم القابلة للاختبار مخبرياً بواسطة أي من الأعمال المُستشهد بها، بل خُففت جزئياً فقط. ينبغي للمطوّر تخصيص ميزانية لخطوة التصفية هذه وألا يتوقع من أي درجة حاسوبية منفردة أن تُلغي الحاجة للتحقق التجريبي.

كيفية بناء هذا، أو كيفية استخدامه

  1. إعداد بيئة BindCraft الأساسية. ثبّت خط أنابيب BindCraft المفتوح المصدر والآلي لتصميم روابط بروتينية من الصفر [7]، والذي يستفيد من أوزان AlphaFold2 لتوليد روابط بروتينية انطلاقاً من بنى الأهداف [7]. ستحتاج إلى أوزان نموذج AlphaFold2 متاحة محلياً وحزمة استدلال JAX/AlphaFold2 عاملة، لأن خط الأنابيب يشغّل تحسيناً قائماً على التدرج داخل هذه الشبكة بدلاً من استدعاء واجهة برمجية خارجية.
  2. تحضير مدخل بنية الهدف. احصل على ملف بنية PDB لبروتين الهدف الخاص بك. إذا كنت تخطط لإعادة إنتاج تصميم موجّه بنقاط ساخنة بدلاً من ذلك، لاحظ أن سير عمل منشور منفصل دمج بين RFdiffusion وProteinMPNN وAlphaFold ومحاكاة الديناميكا الجزيئية والتقييم البنيوي اليدوي [3] كبديل للهلوسة أحادية الشبكة؛ قرّر مسبقاً أي العائلتين تبنيها، لأنهما تحتاجان إلى حزم برمجية مختلفة.
  3. اختيار هدف التحسين. خسارة BindCraft الافتراضية هي تحسين قائم على الهلوسة وعلى التدرج لدرجة TM المتوقعة للواجهة (ipTM)، إلى جانب خسائر أخرى [5]. إذا أردت المتغير المُحسَّن، استبدل ipTM بـ pTMEnergy، التي تُعيد تفسير لوغاريتمات خطأ المحاذاة المتوقع (pAE) من AlphaFold كنموذج قائم على الطاقة على مركّبات الرابط-الهدف [5] ولا تتطلب استدعاءات نموذج إضافية أو تعديلات معمارية [5]، لأنها مُشتقة من مخرجات pAE التي تحسبها بالفعل.
  4. تشغيل حلقة التحسين. كرّر تحديثات التدرج على تسلسل/بنية الرابط لزيادة الدرجة المختارة (ipTM أو pTMEnergy) مقابل الهدف الثابت. ضع في اعتبارك أن ipTM تُنتج تدرجات متناثرة متركزة على مجموعة فرعية صغيرة فقط من أزواج بقايا الواجهة [5]، لذا إذا لاحظت تصادمات أو واجهات مُشكَّلة بشكل رديء خارج بضع بقايا تماس، فهذا نمط فشل معروف للتحسين القائم على ipTM، وpTMEnergy هو التخفيف الموثّق [5].
  5. توليد مجموعة كبيرة من المرشحين. نظراً لأن تحديد المرشحين ذوي التقارب العالي يبقى عائقاً ويتطلب عادةً توليد وفرز آلاف التصاميم افتراضياً لاستخراج عدد قليل من النتائج الواعدة [5]، خطّط لميزانيتك الحسابية لدفعة من مرشحين كثيرين لكل هدف، وليس حفنة منها.
  6. تقييم وتصفية المرشحين بنماذج التنبؤ بالبنية. تُمكّن نماذج التنبؤ بالبنية مثل AlphaFold2 وAlphaFold3 تقييم الارتباط دون بنى بلورية [5]، وتترابط درجات الثقة، وخاصةً ipTM، مع قوة الارتباط [5]. استخدم هذه الدرجات كمرشح أولي، لكن تذكّر أن هذا الترابط تم التحقق منه على مجموعات بيانات محدودة فقط [6].
  7. وضع معيار مرجعي ببروتوكول متعدد المُحقّقين ولكل تسلسل. اتبع نهج ProtDBench: خذ عينة من تسلسلات متعددة لكل هيكل أساسي مُولَّد، وطبّق مُحقّقاً ومرشّحاً على كل تسلسل، واحسب معدل النجاح لكل تسلسل [6]. عرّف مجموعة الهياكل الأساسية الناجحة على أنها الهياكل الأساسية التي يجتاز فيها تسلسل واحد على الأقل المرشّح [6]، وجمّع الهياكل الأساسية الناجحة فقط حسب التشابه البنيوي باستخدام FoldSeek أو TM-score [6] للتحقق من تنوع التصاميم.
  8. التحقق المتقاطع مع أكثر من مُحقّق واحد. نظراً لأن ProtDBench يُبلغ عن تحيّز كبير مرتبط بالمُحقّق واتفاق محدود بين نماذج التنبؤ بالبنية المستخدمة كمُحقّقين [6]، شغّل مرشّحك بنموذجَي تنبؤ بنيوي مختلفَين على الأقل وأبلغ عن كلتا النتيجتين بدلاً من الاعتماد على واحد فقط.
  9. اختيار قائمة مختصرة للتصنيع. من مرشحيك المُصفَّاة والمُجمَّعة، اختر قائمة مختصرة متنوعة للاختبار المخبري. توقّع تبايناً يعتمد على الهدف: وجدت اختبارات مستقلة أن 7 من 15 ببتيداً مُصنَّعاً لـ MDM2 ارتبط بشكل نوعي [12]، و6 من 9 مرشحين لـ WDR5 ارتبطوا بتقارب دون-ميكرومولاري [12]، لكن لم يرتبط أي من الببتيدات المُختبَرة لـ PD-1/PD-L1 على الإطلاق [12]، لذا تعامل مع قائمتك المختصرة الحاسوبية كفرضية يجب اختبارها، وليس ضماناً.
  10. التحقق التجريبي. استخدم فحص ارتباط فيزيائي حيوي (كما استُخدم للتأكد من تقارب منخفض-ميكرومولاري لرابط نطاق LEDGF PWWP [3]) وحيثما كان ملائماً، فحص منافسة لتأكيد الخصوصية بالموقع، كما جرى لببتيدات MDM2 [3][12]. راقب السلوك غير المتوقع مثل التكوّن الثنائي (homodimerization)، والذي تداخل في حالة واحدة مع الارتباط في المحلول رغم أن البروتين لا يزال قابلاً للتبلور المشترك مع هدفه بدقة 2.1 أنغستروم [3].
  11. الإبلاغ عن النتائج بكامل الشروط. عند النشر أو تسجيل معدلات النجاح، اذكر دائماً أي مُحقّق، وأي معيار مرجعي، وأي فئة هدف أنتجت الرقم، نظراً لأن معدلات النجاح التي تتراوح من 10% إلى 100% [7] واكتشاف اختلاف المُحقّقين [6] تعني أن رقم معدل نجاح غير مشروط لا يقدم معلومة كافية بمفرده.

مخطط شبه-كود لحلقة الهلوسة الأساسية مع خيار pTMEnergy:

load target_structure from PDB file
initialize binder_sequence randomly
for step in range(num_steps):
    complex_pred = alphafold2_forward(binder_sequence, target_structure)
    if objective == "ipTM":
        loss = -complex_pred.ipTM
    elif objective == "pTMEnergy":
        loss = -pTMEnergy(complex_pred.pae_logits)
    grads = backprop(loss, wrt=binder_sequence)
    binder_sequence = update(binder_sequence, grads)
return binder_sequence, complex_pred.ipTM, complex_pred.pTMEnergy

الكود: تطبيق عملي

ينفّذ البرنامج النصي أدناه، على بياناتنا الخاصة بـ SQLite، نسخة مبسّطة من فكرتَي التقييم المركزيتين لهذه المذكرة: (1) حساب على طراز ProtDBench لمعدل النجاح لكل تسلسل ومجموعة الهياكل الأساسية الناجحة عبر المرشحين المُولَّدين (الادعاءان R وS)، باستخدام تجميع تشابه بنيوي للتصاميم الناجحة عبر بديل شبيه بـ TM-score (الادعاء T)، و(2) مقارنة بين دالتَي تقييم مماثلتين لـ ipTM مقابل pTMEnergy (الادعاءات B وI وJ وM وN وO وP)، محسوبة من أعمدة جدول `forecasts` الخاص بنا (q10 وq50 وq90 وp_up) التي تحل محل درجات على طراز الثقة، لأن بياناتنا لا تحتوي على لوغاريتمات pAE حقيقية من AlphaFold أو قيم ipTM. نظراً لأن مخططنا لا يحتوي على جداول تصميم روابط، نعامل كل صف من `forecasts` كتصميم مرشح واحد لرمز/هدف واحد، ونعامل `p_up` كدرجة ثقة متناثرة على طراز ipTM، ومزيجاً كثيفاً من انتشار q10/q50/q90 كدرجة كثيفة على طراز pTMEnergy (هذا الاستبدال موضَّح بوضوح في التعليقات البرمجية؛ إنه توضيح لمنهج مقارنة الدرجات، وليس ادعاءً بأن بيانات السوق لدينا هي بيانات بروتين). المعيار الأساسي الذي يجب التفوق عليه هو معدل نجاح المرشّح البسيط على طراز ipTM؛ ومن المتوقع أن تمنح الدرجة الكثيفة على طراز pTMEnergy ترتيباً أكثر استقراراً وأقل تناثراً، وهو ما نتحقق منه بمقارنة عدد الميزات البديلة لزوج البقايا (q10 وq50 وq90) التي تتباين فعلياً مقابل مقدار تغيّر الدرجة الناتج عن ميزة واحدة فقط (فحص بديل التدرج المتناثر مقابل الكثيف، الادعاء J). يبني البرنامج النصي أيضاً تقريراً صغيراً على طراز ProtDBench: معدل النجاح لكل مرشح، وحجم مجموعة الهياكل الأساسية الناجحة، وتجميع بسيط بتقريب الدرجات (بديل لتجميع FoldSeek/TM-score، الادعاء T). يطبع أرقامه ويقارنها بمعدل نجاح خط أساس عشوائي. يعمل كل شيء دون اتصال بالإنترنت مقابل `data.sqlite` (أو QOURAT_DB)، وينتهي في أقل بكثير من ثلاث دقائق على وحدة المعالجة المركزية (CPU)، ولا يُثير أي استثناءات إذا كانت الجداول موجودة (وينشئ مجموعة بيانات اصطناعية صغيرة إذا لم يحتوِ الملف على بيانات، لذا فهو قابل للتشغيل بشكل مستقل أيضاً).

import os
import sqlite3
import sys
import time
import numpy as np
import pandas as pd

# ---------------------------------------------------------------------------
# Config
# ---------------------------------------------------------------------------
DB_PATH = os.environ.get("QOURAT_DB", "data.sqlite")
RNG = np.random.default_rng(0)


def get_connection(path):
    # Opens (or creates) the sqlite file. No network access needed.
    conn = sqlite3.connect(path)
    return conn


def ensure_tables(conn):
    # Creates the required tables if they do not exist yet, so the script
    # is runnable even against a fresh empty file. Schema matches the spec.
    cur = conn.cursor()
    cur.execute("""CREATE TABLE IF NOT EXISTS bars (
        symbol TEXT, tf TEXT, ts TEXT, open REAL, high REAL, low REAL,
        close REAL, volume REAL)""")
    cur.execute("""CREATE TABLE IF NOT EXISTS forecasts (
        symbol TEXT, horizon TEXT, made_at TEXT, q10 REAL, q50 REAL,
        q90 REAL, p_up REAL)""")
    cur.execute("""CREATE TABLE IF NOT EXISTS trades (
        symbol TEXT, ts TEXT, price REAL, size REAL, side TEXT)""")
    cur.execute("""CREATE TABLE IF NOT EXISTS book (
        symbol TEXT, ts TEXT, level INTEGER, bid_price REAL,
        bid_size REAL, ask_price REAL, ask_size REAL)""")
    conn.commit()


def maybe_seed_forecasts(conn):
    # If the forecasts table is empty, populate it with a small synthetic
    # set of "candidate designs" so the pipeline below has something to
    # score. This stands in for generated binder candidates: each row is
    # one candidate for one "target" (symbol), analogous to one sampled
    # sequence for one backbone in ProtDBench (claim R).
    df = pd.read_sql_query("SELECT COUNT(*) AS n FROM forecasts", conn)
    if df["n"].iloc[0] > 0:
        return
    symbols = ["AAPL", "MSFT", "BTC-USD"]
    rows = []
    for sym in symbols:
        for i in range(60):
            q50 = RNG.normal(0, 1)
            spread = abs(RNG.normal(0.5, 0.2)) + 0.01
            q10 = q50 - spread
            q90 = q50 + spread
            p_up = float(np.clip(RNG.normal(0.5, 0.2), 0, 1))
            rows.append((sym, "1d", f"2024-01-{(i % 28) + 1:02d}T00:00:00",
                         q10, q50, q90, p_up))
    conn.executemany(
        "INSERT INTO forecasts (symbol, horizon, made_at, q10, q50, q90, p_up) "
        "VALUES (?,?,?,?,?,?,?)", rows)
    conn.commit()


# ---------------------------------------------------------------------------
# Scoring functions
# ---------------------------------------------------------------------------

def sparse_score(row):
    # Proxy for ipTM-style scoring: BindCraft optimizes ipTM, a score that
    # computes a max over target residue indices and gives sparse
    # gradients concentrated on a small subset of interface pairs
    # (claims B and J). Here we proxy this by taking the single most
    # extreme of the three quantile-derived features, i.e. a max-like
    # operation over a small set of "residue pair" proxies (q10, q50, q90).
    feats = np.array([abs(row["q10"]), abs(row["q50"]), abs(row["p_up"] - 0.5)])
    return float(feats.max())


def dense_score(row):
    # Proxy for pTMEnergy-style scoring: pTMEnergy is derived from the same
    # underlying outputs (here: q10, q50, q90, p_up) but combines them into
    # a dense signal rather than taking a single max (claims M, N, O).
    # We use a simple mean of the same features, so every feature
    # contributes to the score, not just the largest one.
    feats = np.array([abs(row["q10"]), abs(row["q50"]), abs(row["q90"]),
                       abs(row["p_up"] - 0.5)])
    return float(feats.mean())


def gradient_sparsity(row):
    # Checks how much of the sparse score is driven by a single feature
    # versus spread across features, as a numeric illustration of claim J
    # (sparse gradients concentrated on a small subset of interface pairs).
    # Returns the fraction of total feature magnitude held by the largest
    # single feature: close to 1.0 means very sparse, close to 1/n means
    # evenly dense.
    feats = np.array([abs(row["q10"]), abs(row["q50"]), abs(row["q90"]),
                       abs(row["p_up"] - 0.5)])
    total = feats.sum()
    if total = threshold
    per_seq_success_rate = float(df["passes"].mean())

    backbone_groups = df.groupby("symbol")
    passing_backbones = []
    for symbol, group in backbone_groups:
        if group["passes"].any():
            passing_backbones.append(symbol)
    passing_backbone_set = set(passing_backbones)

    return per_seq_success_rate, passing_backbone_set, df


def cluster_passing_backbones(df, passing_backbone_set):
    # Proxy for claim T: ProtDBench clusters only passing backbones by
    # structural similarity using FoldSeek or TM-score. We do not have
    # real structures, so we use a simple proxy: cluster passing backbones
    # by rounding their mean score to one decimal place, as a stand-in
    # for a structural-similarity bucket.
    passing_df = df[df["symbol"].isin(passing_backbone_set)]
    if len(passing_df) == 0:
        return {}
    means = passing_df.groupby("symbol")["score"].mean()
    clusters = {}
    for symbol, m in means.items():
        key = round(m, 1)
        clusters.setdefault(key, []).append(symbol)
    return clusters


# ---------------------------------------------------------------------------
# Baseline: random scoring, to check the two methods beat pure chance
# ---------------------------------------------------------------------------

def random_baseline_success_rate(n_rows, threshold, n_trials=200):
    # A simple baseline: if scores were pure random noise in [0, 1.5],
    # what fraction would pass the threshold? This is the number to beat.
    rates = []
    for _ in range(n_trials):
        fake_scores = RNG.uniform(0, 1.5, size=n_rows)
        rates.append(float((fake_scores >= threshold).mean()))
    return float(np.mean(rates))


def main():
    start = time.time()
    conn = get_connection(DB_PATH)
    ensure_tables(conn)
    maybe_seed_forecasts(conn)

    df = pd.read_sql_query(
        "SELECT symbol, horizon, made_at, q10, q50, q90, p_up FROM forecasts",
        conn)
    conn.close()

    if len(df) == 0:
        print("No forecast rows found even after seeding, aborting.")
        sys.exit(1)

    threshold = 0.5  # fixed filter threshold for both scoring functions

    sparse_rate, sparse_passing, sparse_df = protdbench_style_eval(
        df, sparse_score, threshold)
    dense_rate, dense_passing, dense_df = protdbench_style_eval(
        df, dense_score, threshold)

    sparsity_values = df.apply(gradient_sparsity, axis=1)
    mean_sparsity = float(sparsity_values.mean())

    sparse_clusters = cluster_passing_backbones(sparse_df, sparse_passing)
    dense_clusters = cluster_passing_backbones(dense_df, dense_passing)

    baseline_rate = random_baseline_success_rate(len(df), threshold)

    print("=== BindCraft-style scoring comparison (proxy data) ===")
    print(f"Rows (candidate designs) evaluated: {len(df)}")
    print(f"Backbones (symbols) evaluated: {df['symbol'].nunique()}")
    print()
    print("-- ipTM-style (sparse, max-based) scoring --")
    print(f"Per-sequence success rate: {sparse_rate:.3f}")
    print(f"Passing backbone set size: {len(sparse_passing)} "
          f"of {df['symbol'].nunique()}")
    print(f"Structural-similarity clusters (proxy): {sparse_clusters}")
    print()
    print("-- pTMEnergy-style (dense, mean-based) scoring --")
    print(f"Per-sequence success rate: {dense_rate:.3f}")
    print(f"Passing backbone set size: {len(dense_passing)} "
          f"of {df['symbol'].nunique()}")
    print(f"Structural-similarity clusters (proxy): {dense_clusters}")
    print()
    print(f"Mean gradient sparsity of ipTM-style score (1.0 = fully "
          f"sparse, 0.25 = fully dense over 4 features): {mean_sparsity:.3f}")
    print()
    print(f"Random baseline success rate to beat: {baseline_rate:.3f}")
    print()
    if dense_rate >= sparse_rate and dense_rate > baseline_rate:
        print("RESULT: dense (pTMEnergy-style) scoring matched or beat "
              "sparse (ipTM-style) scoring, and both beat the random baseline.")
    else:
        print("RESULT: dense scoring did not clearly beat sparse scoring "
              "on this data; inspect thresholds and feature scaling.")

    elapsed = time.time() - start
    print(f"\nElapsed time: {elapsed:.2f} seconds")


if __name__ == "__main__":
    main()

ما الذي سنبنيه

سنبني sam v1، فاحص اتفاق مُحقّقين صغيراً لمخرجات على طراز BindCraft. بمعطى بنية هدف ودفعة من مرشحي روابط مُولَّدين من تشغيل BindCraft قائم بالفعل، يقوم sam v1 بتقييم كل مرشح بمُحقّقَي تنبؤ بنيوي مختلفَين متاحَين لنا بالفعل، ثم يُبلغ عن معدل النجاح لكل تسلسل ومجموعة الهياكل الأساسية الناجحة تحت كل مُحقّق على حدة، متبعاً تعريفات ProtDBench لمعدل النجاح لكل تسلسل ومجموعة الهياكل الأساسية الناجحة. سنحكم على sam v1 بمدى اختلاف معدلَي النجاح الخاصَّين بالمُحقّقَين على نفس مجموعة المرشحين: اختلاف كبير من شأنه أن يُعيد إنتاج، على أهدافنا الخاصة، التحيّز الكبير المرتبط بالمُحقّق المُبلَّغ عنه بالفعل لمُحقّقي التنبؤ بالبنية، واختلاف صغير سيكون اكتشافاً مفيداً وقابلاً للإبلاغ بحد ذاته. كخط أساس، سنقارن معدلَي النجاح القائمَين على المُحقّقَين مع خط أساس تقييم عشوائي ساذج محسوب بنفس الطريقة التي يحسب بها قسم الكود لدينا واحداً، حتى يُفحص أي تحسّن مُدَّعى من المُحقّقين الحقيقيين مقابل الصدفة البحتة أولاً. يمكن لفريق من شخصين تنفيذ أغلفة التقييم، ومنطق الهياكل الأساسية الناجحة، وخطوة التجميع في غضون بضعة أسابيع، مع إعادة استخدام صيغ مخرجات BindCraft المفتوحة وأوزان AlphaFold التي لدينا وصول إليها بالفعل؛ التكلفة الرئيسية ستكون الحساب اللازم لتشغيل مُحقّقَين عبر كل مرشح، وهو ما يكون، على ميزانية GPU متواضعة لبضع مئات من المرشحين لكل هدف، ميسور التكلفة ضمن مخصصات الحساب البحثية الاعتيادية. سيكون التسليم تقريراً موجزاً لكل هدف: معدل النجاح تحت المُحقّق A، ومعدل النجاح تحت المُحقّق B، وحجم مجموعة الهياكل الأساسية الناجحة تحت كل منهما، وخط الأساس العشوائي، بحيث يرى القارئ مباشرةً كم يمكن لرقم مُحقّق واحد أن يُضلِّل.

الادعاءات والمراجعة

  1. factمدعوم

    BindCraft is a hallucination-based binder design method that leverages the AlphaFold2 network.

    [5] BindEnergyCraft: Casting Protein Structure Predictors as Energy-Based Models for Binder Design, section 2 Related Work
    “Binder Design Methods. Recent generative methods have enabled programmable design of protein binders. Structure-based approaches such as RFDiffusion [32] and AlphaProteo [35] can generate binder backbones conditioned on a target structure with experimental success. However, Alpha…”
  2. methodمدعوم

    BindCraft hallucinates binder structures with AlphaFold2 and performs gradient-based optimization of the interface predicted TM-score (ipTM), among other losses.

    [5] BindEnergyCraft: Casting Protein Structure Predictors as Energy-Based Models for Binder Design, section 1 Introduction
    “One prominent computational design paradigm that has emerged is hallucination-based design using structure prediction models. For instance, BindCraft [26] achieved unprecedented in vitro success rates across multiple targets by hallucinating binder structures with AlphaFold2 and …”
  3. methodمدعوم مع تحفظات

    Structure-based approaches such as RFDiffusion and AlphaProteo can generate binder backbones conditioned on a target structure.

    [5] BindEnergyCraft: Casting Protein Structure Predictors as Energy-Based Models for Binder Design, section 2 Related WorkPassage says they 'can generate binder backbones conditioned on a target structure with experimental success'—claim drops the experimental success qualification.
    “Binder Design Methods. Recent generative methods have enabled programmable design of protein binders. Structure-based approaches such as RFDiffusion [32] and AlphaProteo [35] can generate binder backbones conditioned on a target structure with experimental success. However, Alpha…”
  4. methodمدعوم

    A hotspot-driven binder-design workflow integrated RFdiffusion, ProteinMPNN, AlphaFold, molecular dynamics simulations, and manual structural assessment.

    [3] De novo design of proteinaceous binders targeting the LEDGF PWWP domain, abstract S2 867f79985871
    “Lens epithelium‐derived growth factor p75 (LEDGF/p75) is a chromatin reader that recognizes di‐ or trimethylated Lys36 of histone H3 (H3K36me2/3)‐modified nucleosomes and is implicated in diverse diseases, including cancer and human immunodeficiency virus (HIV) infection. Inhibit…”
  5. resultمدعوم

    The hotspot-driven workflow generated four protein designs that were subjected to experimental validation.

    [3] De novo design of proteinaceous binders targeting the LEDGF PWWP domain, abstract S2 867f79985871
    “Lens epithelium‐derived growth factor p75 (LEDGF/p75) is a chromatin reader that recognizes di‐ or trimethylated Lys36 of histone H3 (H3K36me2/3)‐modified nucleosomes and is implicated in diverse diseases, including cancer and human immunodeficiency virus (HIV) infection. Inhibit…”
  6. resultمدعوم

    Biophysical analysis confirmed that one designed binder had low-micromolar affinity for the LEDGF PWWP domain.

    [3] De novo design of proteinaceous binders targeting the LEDGF PWWP domain, abstract S2 867f79985871
    “Lens epithelium‐derived growth factor p75 (LEDGF/p75) is a chromatin reader that recognizes di‐ or trimethylated Lys36 of histone H3 (H3K36me2/3)‐modified nucleosomes and is implicated in diverse diseases, including cancer and human immunodeficiency virus (HIV) infection. Inhibit…”
  7. limitationمدعوم

    One designed binder unexpectedly homodimerized, which apparently interfered with its binding to the target in solution.

    [3] De novo design of proteinaceous binders targeting the LEDGF PWWP domain, abstract S2 867f79985871
    “Lens epithelium‐derived growth factor p75 (LEDGF/p75) is a chromatin reader that recognizes di‐ or trimethylated Lys36 of histone H3 (H3K36me2/3)‐modified nucleosomes and is implicated in diverse diseases, including cancer and human immunodeficiency virus (HIV) infection. Inhibit…”
  8. resultمدعوم

    A designed binder that apparently homodimerized could nevertheless be co-crystallized with the PWWP domain, producing an atomic structure at 2.1 Å resolution.

    [3] De novo design of proteinaceous binders targeting the LEDGF PWWP domain, abstract S2 867f79985871
    “Lens epithelium‐derived growth factor p75 (LEDGF/p75) is a chromatin reader that recognizes di‐ or trimethylated Lys36 of histone H3 (H3K36me2/3)‐modified nucleosomes and is implicated in diverse diseases, including cancer and human immunodeficiency virus (HIV) infection. Inhibit…”
  9. limitationمدعوم

    ipTM does not directly reflect the statistical likelihood of the full binder–target interaction.

    [5] BindEnergyCraft: Casting Protein Structure Predictors as Energy-Based Models for Binder Design, section 1 Introduction
    “One prominent computational design paradigm that has emerged is hallucination-based design using structure prediction models. For instance, BindCraft [26] achieved unprecedented in vitro success rates across multiple targets by hallucinating binder structures with AlphaFold2 and …”
  10. limitationمدعوم مع تحفظات

    Because ipTM computes a maximum over target residue indices, it produces sparse gradients concentrated on only a small subset of interface residue pairs.

    [5] BindEnergyCraft: Casting Protein Structure Predictors as Energy-Based Models for Binder Design, section 1 IntroductionPassage says sparse gradients 'constrain optimization to only a small subset'—claim says 'concentrated on' which is not the exact wording but the meaning is close; passage emphasizes the constraint aspect.
    “One prominent computational design paradigm that has emerged is hallucination-based design using structure prediction models. For instance, BindCraft [26] achieved unprecedented in vitro success rates across multiple targets by hallucinating binder structures with AlphaFold2 and …”
  11. factمدعوم

    Structure-prediction models such as AlphaFold2 and AlphaFold3 enable binding evaluation without crystal structures.

    [5] BindEnergyCraft: Casting Protein Structure Predictors as Energy-Based Models for Binder Design, section 2 Related Work
    “Binder Scoring Methods. Scoring functions for evaluating binders span sequence-based, structure-based, and structure prediction-based methods. Sequence models like ESM-1v [23] capture mutational binding effects but underperform on general binding prediction tasks. Structure-based…”
  12. factمدعوم

    Confidence scores from structure-prediction models, particularly interface pTM (ipTM), correlate with binding strength.

    [5] BindEnergyCraft: Casting Protein Structure Predictors as Energy-Based Models for Binder Design, section 2 Related Work
    “Binder Scoring Methods. Scoring functions for evaluating binders span sequence-based, structure-based, and structure prediction-based methods. Sequence models like ESM-1v [23] capture mutational binding effects but underperform on general binding prediction tasks. Structure-based…”
  13. methodمدعوم

    pTMEnergy reinterprets AlphaFold predicted alignment error (pAE) logits as an energy-based model over binder–target complexes.

    [5] BindEnergyCraft: Casting Protein Structure Predictors as Energy-Based Models for Binder Design, section 1 Introduction
    “To address these limitations, we revisit the internal confidence distributions of structure predictors. Folding models like AlphaFold2 output predicted alignment error (pAE) distributions, which quantify the model’s uncertainty over inter-residue distances. These distributions en…”
  14. methodمدعوم مع تحفظات

    pTMEnergy provides a dense, differentiable signal intended to reflect the likelihood of a folded complex under AlphaFold’s learned distribution.

    [5] BindEnergyCraft: Casting Protein Structure Predictors as Energy-Based Models for Binder Design, section 1 IntroductionPassage says pTMEnergy 'reflects the likelihood'—claim adds 'intended to' which softens it; passage presents it as actually reflecting, not just intended to.
    “To address these limitations, we revisit the internal confidence distributions of structure predictors. Folding models like AlphaFold2 output predicted alignment error (pAE) distributions, which quantify the model’s uncertainty over inter-residue distances. These distributions en…”
  15. methodمدعوم

    pTMEnergy is derived from the same pAE outputs used to compute ipTM and requires no additional model calls or architecture modifications.

    [5] BindEnergyCraft: Casting Protein Structure Predictors as Energy-Based Models for Binder Design, section 1 Introduction
    “To address these limitations, we revisit the internal confidence distributions of structure predictors. Folding models like AlphaFold2 output predicted alignment error (pAE) distributions, which quantify the model’s uncertainty over inter-residue distances. These distributions en…”
  16. resultمدعوم مع تحفظات

    BindEnergyCraft, which retains BindCraft’s optimization framework while replacing ipTM with pTMEnergy, achieved higher in silico binder success rates and reduced structural clashes than BindCraft, RFDiffusion, and ESM3 across multiple challenging targets.

    [5] BindEnergyCraft: Casting Protein Structure Predictors as Energy-Based Models for Binder Design, abstract arXiv:2505.21241v1Claim says 'retains BindCraft's optimization framework'—passage says 'maintains the same optimization framework'; claim correctly captures the core result about higher success rates and reduced clashes.
    “Protein binder design has been transformed by hallucination-based methods that optimize structure prediction confidence metrics, such as the interface predicted TM-score (ipTM), via backpropagation. However, these metrics do not reflect the statistical likelihood of a binder-targ…”
  17. limitationمدعوم

    The correlation of confidence metrics such as pLDDT, ipTM, and ipAE with experimental binding outcomes has been validated on only limited datasets and remains poorly understood across targets and design regimes.

    [6] ProtDBench: A Unified Benchmark of Protein Binder Design and Evaluation, section 1 Introduction
    “Confidence metrics (e.g., pLDDT, ipTM, and ipAE) are widely used as proxies for binding success, yet their correlation with experimental outcomes has only been validated on limited datasets and remains poorly understood across targets and design regimes. • Standardization of benc…”
  18. methodمدعوم

    ProtDBench evaluates generated binder backbones by sampling multiple sequences for each backbone, applying a verifier and filter to each sequence, and computing the per-sequence success rate.

    [6] ProtDBench: A Unified Benchmark of Protein Binder Design and Evaluation, section 2 Evaluation Framework: ProtDBench
    “ProtDBench (Algorithm 1) is a unified evaluation framework for de novo protein binder design. Rather than treating evaluation as a fixed component of a design pipeline, ProtDBench explicitly formalizes evaluation as a configurable process that operates on generated candidates and…”
  19. methodمدعوم

    ProtDBench defines the passing backbone set as the backbones for which at least one sampled sequence passes the filter.

    [6] ProtDBench: A Unified Benchmark of Protein Binder Design and Evaluation, section 2 Evaluation Framework: ProtDBench
    “ProtDBench (Algorithm 1) is a unified evaluation framework for de novo protein binder design. Rather than treating evaluation as a fixed component of a design pipeline, ProtDBench explicitly formalizes evaluation as a configurable process that operates on generated candidates and…”
  20. methodمدعوم

    ProtDBench clusters only passing backbones by structural similarity using FoldSeek or TM-score.

    [6] ProtDBench: A Unified Benchmark of Protein Binder Design and Evaluation, section 2 Evaluation Framework: ProtDBench
    “ProtDBench (Algorithm 1) is a unified evaluation framework for de novo protein binder design. Rather than treating evaluation as a fixed component of a design pipeline, ProtDBench explicitly formalizes evaluation as a configurable process that operates on generated candidates and…”
  21. limitationمدعوم

    Identifying high-affinity candidates remains a bottleneck because binder design typically requires generating and virtually screening thousands of designs to recover a few promising hits.

    [5] BindEnergyCraft: Casting Protein Structure Predictors as Energy-Based Models for Binder Design, section 1 Introduction
    “De novo protein binder design represents a fundamental challenge in molecular engineering, with wide-ranging therapeutic and biotechnological applications [7, 12, 34]. Recent advances in deep learning have enabled considerable progress in computational binder design, allowing the…”
  22. uncertaintyمدعوم

    ProtDBench reports substantial verifier-dependent bias and limited agreement among structure-prediction models used as evaluation verifiers.

    [6] ProtDBench: A Unified Benchmark of Protein Binder Design and Evaluation, section ProtDBench: A Unified Benchmark of Protein Binder Design and Evaluation
    “Code: https://github.com/congliuUvA/ProtDBench, a standardized and throughput-aware evaluation framework for protein binder design. ProtDBench defines unified benchmark tasks, evaluation protocols, and success criteria, enabling systematic analysis of how evaluation design influe…”
  23. factمدعوم

    BindCraft is an open-source and automated pipeline for de novo protein binder design.

    [7] One-shot design of functional protein binders with BindCraft, abstract DOI 10.1038/s41586-025-09429-6
    “Protein–protein interactions are at the core of all key biological processes. However, the complexity of the structural features that determine protein–protein interactions makes their design challenging. Here we present BindCraft, an open-source and automated pipeline for de nov…”
  24. methodمدعوم مع تحفظات

    BindCraft leverages the weights of AlphaFold2 to generate protein binders from target structures.

    [7] One-shot design of functional protein binders with BindCraft, abstract DOI 10.1038/s41586-025-09429-6Passage says BindCraft 'leverages the weights of AlphaFold2 to generate binders'—claim adds 'from target structures' which is not explicitly stated in this passage.
    “Protein–protein interactions are at the core of all key biological processes. However, the complexity of the structural features that determine protein–protein interactions makes their design challenging. Here we present BindCraft, an open-source and automated pipeline for de nov…”
  25. resultمدعوم

    BindCraft has reported experimental success rates ranging from 10% to 100%.

    [7] One-shot design of functional protein binders with BindCraft, abstract DOI 10.1038/s41586-025-09429-6
    “Protein–protein interactions are at the core of all key biological processes. However, the complexity of the structural features that determine protein–protein interactions makes their design challenging. Here we present BindCraft, an open-source and automated pipeline for de nov…”
  26. resultمدعوم

    BindCraft generated binders with nanomolar affinity without high-throughput screening or experimental optimization, including for targets without known binding sites.

    [7] One-shot design of functional protein binders with BindCraft, abstract DOI 10.1038/s41586-025-09429-6
    “Protein–protein interactions are at the core of all key biological processes. However, the complexity of the structural features that determine protein–protein interactions makes their design challenging. Here we present BindCraft, an open-source and automated pipeline for de nov…”
  27. factمدعوم

    BindCraft was applied to cell-surface receptors, common allergens, de novo designed proteins, and multi-domain nucleases such as CRISPR–Cas9.

    [7] One-shot design of functional protein binders with BindCraft, abstract DOI 10.1038/s41586-025-09429-6
    “Protein–protein interactions are at the core of all key biological processes. However, the complexity of the structural features that determine protein–protein interactions makes their design challenging. Here we present BindCraft, an open-source and automated pipeline for de nov…”
  28. resultمدعوم مع تحفظات

    Designed binders were reported to reduce IgE binding to birch allergen, modulate Cas9 gene-editing activity, reduce the cytotoxicity of a foodborne bacterial enterotoxin, and redirect adeno-associated virus capsids for targeted gene delivery.

    [7] One-shot design of functional protein binders with BindCraft, abstract DOI 10.1038/s41586-025-09429-6Passage says binders were used for these applications; claim says 'were reported to' which accurately reflects this is from the paper, though passage describes actual demonstrations.
    “Protein–protein interactions are at the core of all key biological processes. However, the complexity of the structural features that determine protein–protein interactions makes their design challenging. Here we present BindCraft, an open-source and automated pipeline for de nov…”
  29. methodمدعوم

    BindCraft can generate high-affinity peptides solely from a target structure.

    [12] Evaluating BindCraft for Generative Design of High-Affinity Peptides, abstract DOI 10.1021/acschembio.5c00774
    “High Resolution Image Download MS PowerPoint Slide Discovering high-affinity ligands directly from protein structures remains a key challenge in drug discovery. BindCraft is a structure-guided generative modeling platform able to de novo design miniproteins with a high affinity f…”
  30. resultمدعوم

    For MDM2, BindCraft generated 70 unique peptides, 15 were synthesized, and 7 showed specific binding with nanomolar affinities.

    [12] Evaluating BindCraft for Generative Design of High-Affinity Peptides, abstract DOI 10.1021/acschembio.5c00774
    “High Resolution Image Download MS PowerPoint Slide Discovering high-affinity ligands directly from protein structures remains a key challenge in drug discovery. BindCraft is a structure-guided generative modeling platform able to de novo design miniproteins with a high affinity f…”
  31. resultمدعوم

    Competition assays confirmed site-specific binding of the MDM2 peptides at the intended target site.

    [12] Evaluating BindCraft for Generative Design of High-Affinity Peptides, abstract DOI 10.1021/acschembio.5c00774
    “High Resolution Image Download MS PowerPoint Slide Discovering high-affinity ligands directly from protein structures remains a key challenge in drug discovery. BindCraft is a structure-guided generative modeling platform able to de novo design miniproteins with a high affinity f…”
  32. resultمدعوم

    For WDR5, six of nine candidates bound the MYC binding WBM site with submicromolar affinity.

    [12] Evaluating BindCraft for Generative Design of High-Affinity Peptides, abstract DOI 10.1021/acschembio.5c00774
    “High Resolution Image Download MS PowerPoint Slide Discovering high-affinity ligands directly from protein structures remains a key challenge in drug discovery. BindCraft is a structure-guided generative modeling platform able to de novo design miniproteins with a high affinity f…”
  33. resultمدعوم

    Rational chemical modification improved the potency of one WDR5 binder six-fold to a KD of 39 nM.

    [12] Evaluating BindCraft for Generative Design of High-Affinity Peptides, abstract DOI 10.1021/acschembio.5c00774
    “High Resolution Image Download MS PowerPoint Slide Discovering high-affinity ligands directly from protein structures remains a key challenge in drug discovery. BindCraft is a structure-guided generative modeling platform able to de novo design miniproteins with a high affinity f…”
  34. limitationمدعوم

    None of the tested BindCraft-generated peptides for PD-1 and PD-L1 showed detectable binding.

    [12] Evaluating BindCraft for Generative Design of High-Affinity Peptides, abstract DOI 10.1021/acschembio.5c00774
    “High Resolution Image Download MS PowerPoint Slide Discovering high-affinity ligands directly from protein structures remains a key challenge in drug discovery. BindCraft is a structure-guided generative modeling platform able to de novo design miniproteins with a high affinity f…”
  35. factمدعوم

    The Human Bindome embeds the experimentally benchmarked BindCraft method in an accelerated, parallelized framework with automated domain-level target selection.

    [13] The Human Bindome: A Proteome-scale Atlas of Designed Binder Candidates, abstract S2 17dd0d888175
    “Affinity reagents such as antibodies are indispensable for interrogating proteins’ biological function. Yet they are costly and frequently unreliable, with unknown sequences, posing challenges to reproducible experimental research. Deep learning-based protein design can now in si…”
  36. resultمدعوم

    Using this framework, 306,146 binder candidates were generated for 8,296 human proteins, covering 40.9% of the full proteome.

    [13] The Human Bindome: A Proteome-scale Atlas of Designed Binder Candidates, abstract S2 17dd0d888175
    “Affinity reagents such as antibodies are indispensable for interrogating proteins’ biological function. Yet they are costly and frequently unreliable, with unknown sequences, posing challenges to reproducible experimental research. Deep learning-based protein design can now in si…”
  37. methodمدعوم

    Each Human Bindome candidate includes a defined sequence, a predicted binder-target structure model, and in silico confidence metrics.

    [13] The Human Bindome: A Proteome-scale Atlas of Designed Binder Candidates, abstract S2 17dd0d888175
    “Affinity reagents such as antibodies are indispensable for interrogating proteins’ biological function. Yet they are costly and frequently unreliable, with unknown sequences, posing challenges to reproducible experimental research. Deep learning-based protein design can now in si…”
  38. limitationمرفوض

    The provided passages do not specify BindCraft2's complete generation pipeline, its exact binding-specificity scoring functions, or the commands and hardware requirements for running the full design-to-candidate workflow locally.

    [6] ProtDBench: A Unified Benchmark of Protein Binder Design and Evaluation, section 1 Introduction
    “such as the structure prediction verifier, filtering strategy, and computational budget, substantially shape reported performance, yet are rarely made explicit or analyzed systematically. By grounding evaluation in wet-lab annotated data, ProtDBench enables principled assessment …”
  39. factمدعوم

    Protein language models (pLMs) are pre-trained on large-scale evolutionary sequence data and adapt large language models for biological sequences.

    [17] Preference optimization of protein language models as a multi-objective binder design paradigm, section 1 Introduction
    “Extending large language models (LLMs) for natural language processing (NLP) to biological sequences, protein language models (pLMs) are pre-trained on large scale evolutionary sequence data. Prominent foundation models include ESM2 (Lin et al., 2022) that is a BERT-style encoder…”
  40. factمدعوم

    ESM2 is a BERT-style encoder transformer model, while ProtGPT2 and ProGen2 are GPT-style decoder transformer models and ProtT5 is an encoder-decoder transformer model.

    [17] Preference optimization of protein language models as a multi-objective binder design paradigm, section 1 Introduction
    “Extending large language models (LLMs) for natural language processing (NLP) to biological sequences, protein language models (pLMs) are pre-trained on large scale evolutionary sequence data. Prominent foundation models include ESM2 (Lin et al., 2022) that is a BERT-style encoder…”

المصادر

  1. [1]
    Muhammad Salman Iqbal, Revocatus Bahitwa, Abdul Ali Azam, Hui Xu, Hai Wang. Deep learning–driven protein binder design for crop improvement. aBIOTECH, 2025.openalex · primary · DOI 10.1016/j.abiote.2025.100018 · https://doi.org/10.1016/j.abiote.2025.100018
  2. [2]
    Fanhao Wang, Yuzhe Wang, Laiyi Feng, Changsheng Zhang, Luhua Lai. Target-Specific De Novo Peptide Binder Design with DiffPepBuilder. arXiv, 2024.arxiv · primary · https://arxiv.org/abs/2405.00128v2
  3. [3]
    Thibault Vantieghem, Julie Delepine, Sam Noppen, S. Beelen, Matúš Drexler, Jitka Holková. De novo design of proteinaceous binders targeting the LEDGF PWWP domain. Protein Science, 2026.semanticscholar · primary · DOI 10.1002/pro.70726 · https://doi.org/10.1002/pro.70726
  4. [4]
    Fanhao Wang, Yuzhe Wang, Lai-Yi Feng, Chang-Sheng Zhang, L. Lai. Target-Specific De Novo Peptide Binder Design with DiffPepBuilder. Journal of Chemical Information and Modeling, 2024.semanticscholar · primary · DOI 10.1021/acs.jcim.4c00975 · https://doi.org/10.1021/acs.jcim.4c00975
  5. [5]
    Divya Nori, Anisha Parsan, Caroline Uhler, Wengong Jin. BindEnergyCraft: Casting Protein Structure Predictors as Energy-Based Models for Binder Design. arXiv, 2025.arxiv · primary · https://arxiv.org/abs/2505.21241v1
  6. [6]
    Cong Liu, Milong Ren, Jiaqi Guan, Chengyue Gong, Jinyuan Sun, Xinshi Chen, Wenzhi Xiao. ProtDBench: A Unified Benchmark of Protein Binder Design and Evaluation. arXiv, 2026.arxiv · primary · https://arxiv.org/abs/2605.04118v2
  7. [7]
    Martin Pačesa, Lennart Nickel, Christian Schellhaas, Joseph H. Schmidt, Ekaterina Pyatova, Lucas Kissling. One-shot design of functional protein binders with BindCraft. Nature, 2025.openalex · primary · DOI 10.1038/s41586-025-09429-6 · https://doi.org/10.1038/s41586-025-09429-6
  8. [8]
    Martin Pačesa, Lennart Nickel, Christian Schellhaas, Joseph H. Schmidt, Ekaterina Pyatova, Lucas Kissling. BindCraft: one-shot design of functional protein binders. bioRxiv (Cold Spring Harbor Laboratory), 2024.openalex · primary · DOI 10.1101/2024.09.30.615802 · https://doi.org/10.1101/2024.09.30.615802
  9. [9]
    Hannes Stark, Felix Faltings, MinGyu Choi, Yuxin Xie, Eunsu Hur, Timothy J. O’Donnell. BoltzGen: Toward Universal Binder Design. bioRxiv (Cold Spring Harbor Laboratory), 2025.openalex · primary · DOI 10.1101/2025.11.20.689494 · https://doi.org/10.1101/2025.11.20.689494
  10. [10]
    Tudor‐Stefan Cotet, Igor Krawczuk, Filippo Stocco, Noelia Ferruz, Anthony Gitter, Yoichi Kurumida. Crowdsourced Protein Design: Lessons From the Adaptyv EGFR Binder Competition. bioRxiv (Cold Spring Harbor Laboratory), 2025.openalex · primary · DOI 10.1101/2025.04.17.648362 · https://doi.org/10.1101/2025.04.17.648362
  11. [11]
    Dylan Silke, Julie Iskander, Junqi Pan, Andrew P. Thompson, Anthony T. Papenfuss, Isabelle S. Lucet. ProteinDJ : A high‐performance and modular protein design pipeline. Protein Science, 2026.openalex · primary · DOI 10.1002/pro.70464 · https://doi.org/10.1002/pro.70464
  12. [12]
    Mike Filius, Thanasis Patsos, Hugo Minnee, Gianluca Turco, H. Chong, Jingming Liu. Evaluating BindCraft for Generative Design of High-Affinity Peptides. ACS Chemical Biology, 2025.openalex · primary · DOI 10.1021/acschembio.5c00774 · https://doi.org/10.1021/acschembio.5c00774
  13. [13]
    Julius Wenckstern, Anna M. Díaz-Rovira, Julia A. Kuhn, Arvid Ban, Rahma Hamdani, Roser Pruaño-Milla. The Human Bindome: A Proteome-scale Atlas of Designed Binder Candidates. bioRxiv, 2026.semanticscholar · primary · DOI 10.64898/2026.07.30.741542 · https://doi.org/10.64898/2026.07.30.741542
  14. [14]
    Leonardo Almeida-Souza. A deep learning predictor of bindable protein surfaces to guide generative synthetic biology. bioRxiv, 2026.semanticscholar · primary · DOI 10.64898/2026.04.16.718848 · https://doi.org/10.64898/2026.04.16.718848
  15. [15]
    Puja Trivedi, Danai Koutra, Jayaraman J. Thiagarajan. On the Efficacy of Generalization Error Prediction Scoring Functions. arXiv, 2023.arxiv · primary · https://arxiv.org/abs/2303.13589v2
  16. [16]
    Matthew Ragoza, Joshua Hochuli, Elisa Idrobo, Jocelyn Sunseri, David Ryan Koes. Protein–Ligand Scoring with Convolutional Neural Networks. Journal of Chemical Information and Modeling, 2017.openalex · primary · DOI 10.1021/acs.jcim.6b00740 · https://doi.org/10.1021/acs.jcim.6b00740
  17. [17]
    Pouria Mistani, Venkatesh Mysore. Preference optimization of protein language models as a multi-objective binder design paradigm. arXiv, 2024.arxiv · primary · https://arxiv.org/abs/2403.04187v1
  18. [18]
    Jacob Beck, Shikha Surana, Manus McAuliffe, Oliver Bent, Thomas D. Barrett, Juan Jose Garau Luis, Paul Duckworth. Metalic: Meta-Learning In-Context with Protein Language Models. arXiv, 2024.arxiv · primary · https://arxiv.org/abs/2410.08355v3
  19. [19]
    Rebecca Buller, Jiřı́ Damborský, Donald Hilvert, Uwe T. Bornscheuer. Structure Prediction and Computational Protein Design for Efficient Biocatalysts and Bioactive Proteins. Angewandte Chemie International Edition, 2024.openalex · primary · DOI 10.1002/anie.202421686 · https://doi.org/10.1002/anie.202421686
  20. [20]
    Ting-Yu Chang, Yu-Lin Wang. A Structure-Guided Workflow for Efficiently Developing Antibody Against CEACAM-6 with RF Diffusion. ECS Meeting Abstracts, 2026.semanticscholar · primary · DOI 10.1149/ma2026-01341604mtgabs · https://doi.org/10.1149/ma2026-01341604mtgabs