أبحاث

ميزات اختلال تدفق الأوامر (Order-Flow Imbalance): أي الإشارات تتنبأ بتحركات العملات الرقمية/الأسهم خلال 1 إلى 5 دقائق القادمة؟

Order-Flow Imbalance Features: Which Signals Predict Crypto/Equity Moves 1-5 Minutes Out?

نُشر
دقائق قراءة
37 min · 5,591 كلمة
الادعاءات والمراجعة
40/40 ادعاءات موثقة · 16 مصادر

الطبعات: English · Español · Français

الجواب المباشر

لا تُجري أي مطالبة في قاعدة الأدلة هذه اختباراً مباشراً يقارن OFI مقابل اختلال العمق متعدد المستويات مقابل ميزات تدفق التداول على نفس بيانات العملات الرقمية والأسهم ضمن تحقق متقاطع من نوع walk-forward مُنقّى ومحاط بفترة عزل (purged, embargoed)، لذا لا يمكن تقديم ترتيب مباشر بأمانة؛ تنطبق هذه الفجوة بالتساوي على الكود، وخطوات البناء، ومناقشة الحدود أدناه، ولن تُكرر لاحقاً في هذه المذكرة. ما تدعمه المطالبات فعلاً هو ترتيب بناء: ابدأ باختلال العرض والطلب في أعلى الدفتر I=(qb-qa)/(qb+qa)، لأن التقارير تشير إلى أن اختلال الدفتر المرتفع يُعدّ في المتوسط مؤشراً جيداً للتنبؤ بحركات سعر منتصف السعر (mid price)، وأن الدفاتر شديدة الاختلال تدل على احتمال حدوث حركة سعرية خلال وقت قصير نسبياً [2]. وسّع ذلك ليشمل اختلال تدفق الأوامر متعدد المستويات (MLOFI)، لأنه بالنسبة لجميع الأسهم الستة المدروسة، تتحسن جودة المطابقة خارج العينة مع كل مستوى سعري إضافي يُدرج في متجه MLOFI [6]. أضف ميزات تدفق التداول والسيولة (إشارة التداول، حجم أمر السوق، سيولة أفضل عرض/طلب) لأنها أظهرت، عبر انحدار لوجستي بطريقة LASSO، أنها معلوماتية باستمرار في التنبؤ بالقفزة السعرية التالية على أسهم CAC40 [4]. لم تُختبر أي من أدلة الأسهم وطوابير دفتر الأوامر هذه على بورصات العملات الرقمية في المطالبات المُستشهد بها. يطلب السؤال صراحةً نتائج "عبر ثلاث بورصات عملات رقمية" ولا تحتوي قاعدة الأدلة على أي مقارنة متعددة البورصات من هذا النوع، مما يجعل الفجوة أوسع من مجرد استقراء بين فئات الأصول: يبقى جزء العملات الرقمية من السؤال مفتوحاً بانتظار اختبار رجعي (backtest) جديد.

لماذا يهم اختلال تدفق الأوامر في التنبؤ قصير الأفق

اختلال تدفق الأوامر وما يتصل به من ميزات هي مرشحات جذابة لميزات الأسعار قصيرة الأفق لأن عدة دراسات مستقلة، عبر فئات أصول مختلفة وفترات زمنية مختلفة، تُبلغ عن علاقة بين الاختلال في دفتر الأوامر المحدودة (limit order book) واتجاه أو حجم حركة السعر التالية [2][4][6]. يذكر مصدر ثانوي أن Cont وزملاءه يجدون اعتماداً خطياً بسيطاً بين تغيرات السعر ومؤشر يقيس الاختلالات بين تدفق الأوامر على جانبي الشراء والبيع في LOB [3]، وهو ما يتسق مع هذه النتائج الأولية لكنه ليس بحد ذاته نتيجة أولية. تجمع هذه المذكرة ما تم قياسه فعلاً بشأن هذه الميزات، وتوضح بدقة أين تتوقف الأدلة، وتقدم خطة بناء ملموسة.

المخاطر العملية هي أن اتجاه منتصف السعر خلال الأحداث القليلة القادمة في دفتر الأوامر غالباً ما يكون هدفاً مبسّطاً أكثر من اللازم: تنص الفقرة 644 مباشرة على أن التنبؤ بتغير اتجاه منتصف السعر خلال الأحداث القليلة القادمة مبسّط أكثر من اللازم وغير مناسب لاستراتيجية تداول عملية [1]. هذا تحذير من أبسط تصميم ممكن للتسمية (label)؛ ولا يحدد بحد ذاته ما ينبغي أن يحل محل هذا التصميم (الحجم، وقت الانتظار، أو تصميم متعدد الآفاق)، لذا فإن أي تحول نحو هدف مختلف هو استنتاجنا الخاص، وليس شيئاً تنص عليه المطالبة، ويُستخدم هذا الاستنتاج لاحقاً في خطوات البناء عندما نوصي بالإبلاغ عن وقت الانتظار وحركة السعر المُطبّعة كدوال لميزة الاختلال.

أخيراً، يمتد السؤال ليشمل العملات الرقمية والأسهم، والدراسات المُستشهد بها مقسمة حسب فئة الأصل: دراسات اختلال LOB والقفزة السعرية على الأسهم الصينية [1]، وديناميكيات طوابير دفتر الأوامر المحدودة العامة [2][3]، وأسهم CAC40 الفرنسية [4]، وأسهم Nasdaq [6]، بينما يدمج نظام منفصل تدفقات بيانات العملات الرقمية (BTC) للاستدلال الحي [5].

ما هو اختلال تدفق الأوامر واختلال العمق فعلياً

يُعرَّف اختلال العرض والطلب في أعلى الدفتر بدقة في إحدى الأوراق المُستشهد بها: I=(qb-qa)/(qb+qa)، حيث qb وqa هما كميتا العرض والطلب المنشورتان في أعلى الدفتر [2, الفقرة 649]. يشير الاختلال الموجب إلى دفتر أوامر أثقل من جانب العرض (bid)، ويشير الاختلال السالب إلى دفتر أثقل من جانب الطلب (ask) [2, الفقرة 649]. هذه هي أبسط وأرخص نسخة من الميزة: عملية طرح واحدة، وجمع واحد، وقسمة واحدة، تُحسب من لقطة واحدة من LOB.

يُعمِّم اختلال تدفق الأوامر متعدد المستويات (Multi-Level Order-Flow Imbalance)، أو MLOFI، هذه الفكرة. يُوصف MLOFI بأنه كمية متجهة تقيس صافي تدفق أوامر الشراء والبيع عند مستويات سعرية مختلفة في دفتر أوامر محدودة [6]. على عكس الاختلال أحادي المستوى I، فإن MLOFI متجه، لذا يحمل معلومات من عدة مستويات عمق في آن واحد بدلاً من ضغط كل شيء في نسبة واحدة عند أفضل عرض وطلب.

تُشكِّل ميزات تدفق التداول والسيولة عائلة ثالثة. تُعرّف إحدى الأوراق القفزة السعرية بدقة على أنها وصول أمر سوق بيع (شراء) يُنفَّذ بسعر أصغر (أكبر) من سعر أفضل عرض (أفضل طلب) مباشرة بعد وصول أمر السوق السابق [4, الفقرة 661]. تستخرج نفس الورقة ميزات من أحجام الأوامر المحدودة، وفجوات أسعار الأوامر المحدودة، ومعلومات أوامر السوق، ومعلومات أحداث الأوامر المحدودة [4, الفقرة 661]، وهي مجموعة ميزات أوسع من الاختلال وحده: تشمل إشارة التداول وحجمه، وليس فقط السيولة القائمة. تنص الفقرة 661 بشكل منفصل على أن توازن السيولة عند أفضل عرض وأفضل طلب معلوماتي جداً في التنبؤ باتجاه أمر السوق المستقبلي [4]، وهو هدف تنبؤ مختلف عن القفزة السعرية نفسها.

يعالج خط عمل منفصل دفتر الأوامر المحدودة وعملية التداول كنظام عشوائي مشترك بدلاً من لقطة ثابتة. تدرس إحدى الأوراق احتمال تحركات السعر ووصول الصفقات كدالة لاختلال العرض عند أعلى دفتر الأوامر المحدودة، وتقترح نموذجاً عشوائياً لالتقاط الديناميكيات المشتركة لطوابير أعلى الدفتر وعملية التداول [2, الفقرة 647]. بشكل منفصل، تذكر الفقرة 653، كسرد ثانوي لنتيجة Cont وزملائه وليس كنتيجة أولية خاصة بها، أن Cont وزملاءه يجدون اعتماداً خطياً بسيطاً بين تغيرات السعر ومؤشر يقيس الاختلالات بين تدفق الأوامر على جانبي الشراء والبيع في LOB [3, الفقرة 653]، وهو ما يتسق روحياً مع نهج قائم على الاختلال، رغم أنها دراسة منفصلة وليست معياراً مشتركاً مقابل الأخريات.

الآلية: كيف يرتبط الاختلال بحركة السعر التالية

تُعبَّر الآلية المُبلَّغ عنها لاختلال أعلى الدفتر من خلال كميتين: متوسط حركة منتصف السعر مُطبّعاً بفارق العرض والطلب (bid-ask spread)، ووقت الانتظار حتى حركة منتصف السعر التالية، وكلاهما دُرس كدالة لاختلال الدفتر [2, الفقرة 646]. هذا يمنح البنّاء مخرجين لاستهدافهما، وليس واحداً: مقدار (الحركة المُطبّعة) وتوقيت (وقت الانتظار)، وكلاهما مفهرس بنفس مدخل الاختلال.

بالنسبة لأمر موضوع على جانب العرض، فإن وقت التوقف ذا الصلة هو زمن أول وصول لصفقة بيع، وبالنسبة لأمر موضوع على جانب الطلب، فهو زمن أول وصول لصفقة شراء [2, الفقرة 651]. هذا تعريف دقيق ومدفوع بالأحداث، مفيد لمن يريد إعادة إنتاج حساب وقت الانتظار بدلاً من تقريبه بزمن ساعة ثابت.

حجم التأثير محدود. في الحالة المعروضة، يمكن أن تصل حركة السعر المتوسطة إلى ثلث الفارق في دفتر شديد الاختلال [2, الفقرة 646]، وبشكل أعم فإن تغير السعر حتى الجزء (tick) التالي يُقرَّب جيداً بدالة خطية للاختلال، ويكون عادةً أقل بكثير من فارق العرض والطلب، حتى في دفاتر الأوامر شديدة الاختلال [2, الفقرة 650]. اختلال الدفتر المرتفع هو، في المتوسط، مؤشر جيد للتنبؤ بحركات منتصف السعر [2, الفقرة 650]، والدفاتر شديدة الاختلال تدل على أن حركة السعر من المرجح أن تأتي خلال وقت قصير نسبياً [2, الفقرة 650]. لكن نفس الورقة صريحة في أن هذا لا يُترجم بحد ذاته إلى ميزة تداولية: على الرغم من أن اختلال الدفتر قد يُستخدم كمؤشر للحركة السعرية التالية، فإنه لا يوفر بحد ذاته فرصة لمراجحة إحصائية مباشرة [2, الفقرة 650]. هذا تحذير مهم لمن يميل إلى تداول الإشارة الخام مباشرة.

يكمن وراء كل هذا بيان بشأن البيانات: تُظهر البيانات الطبيعة غير المارتينغالية (non-martingale) للأسعار على المقاييس الزمنية القصيرة المدروسة [2, الفقرة 646]. هذا هو الأساس التجريبي لتوقع أي قابلية للتنبؤ قصير الأفق في هذا الخط من العمل على الإطلاق؛ إنها خاصية للبيانات المدروسة، وليست قانوناً عاماً، وتأتي من هذه الدراسة تحديداً.

بناء الميزات المستخدَم في دراسات القفزة السعرية وMLOFI

تبني دراسة القفزة السعرية على CAC40 مجموعة ميزات أوسع من الاختلال وحده. تستخرج ميزات من أحجام الأوامر المحدودة، وفجوات أسعار الأوامر المحدودة، ومعلومات أوامر السوق، ومعلومات أحداث الأوامر المحدودة [4, الفقرة 661]، وتطبق انحداراً لوجستياً للتنبؤ بالقفزة السعرية من ميزات دفتر الأوامر المحدودة [4, الفقرة 661]. لمعالجة عدد الميزات المرشحة، تُدخِل انحداراً لوجستياً بطريقة LASSO لأداء اختيار المتغيرات وإبراز أهمية الميزات المختلفة في التنبؤ بالقفزة السعرية المستقبلية [4, الفقرة 661]. يعطي LASSO ترتيباً قابلاً للتفسير: أي الميزات تبقى بعد الانكماش (shrinkage) وأيها لا يبقى.

نتيجة اختيار المتغيرات ذاك، بناءً على أربعين من أكبر الأسهم الفرنسية في CAC40، هي أن إشارة التداول وحجم أمر السوق بالإضافة إلى السيولة عند أفضل عرض وأفضل طلب معلوماتية باستمرار في التنبؤ بالقفزة السعرية القادمة [4, الفقرة 661]. هذه إجابة مباشرة ومحددة عن أي الميزات هي الأهم في ذلك الإعداد الخاص بالأسهم: إشارة التداول، وحجم أمر السوق، وسيولة أعلى الدفتر، والتي أُبقيَ عليها عبر انكماش LASSO بدلاً من اختيارها يدوياً [4, الفقرة 661].

ما تم قياسه: الأرقام والمعايير والشروط

تظهر عدة أرقام ملموسة عبر هذه الأوراق، كل منها مرتبط بمعياره الخاص وفئة أصوله. يغطي معيار السوق الصيني لدفتر الأوامر المحدودة بضعة آلاف من الأسهم من يونيو إلى سبتمبر 2020 [1]، ويتنبأ بروتوكوله بتغير السعر المتوسط المرجّح بالحجم (volume-weighted average price) القادم والحجم في نهاية كل ثانية عبر 12 أفقاً زمنياً يتراوح من ثانية واحدة إلى 300 ثانية [1]. ضمن هذا البروتوكول، تُقارَن النتائج المبنية على نموذج انحدار خطي ونماذج تعلّم عميق [1]، وتقترح نفس الورقة مجموعة ميزات أكثر فعالية عملياً لالتقاط كل من لقطات LOB والبيانات الدورية [1]. هذه هي مجموعة المطالبات الوحيدة هنا التي تقارن مباشرة نموذجاً خطياً بنماذج تعلّم عميق على معيار محدد [1]؛ تُثبت المقارنة أن عائلتي النماذج اختُبرتا وجهاً لوجه، لكن المطالبات لا تحدد أيهما حقق أداءً أفضل، لذا لا يمكن للقارئ أن يعرف من هذا الدليل ما إذا كان النموذج الخطي أو التعلم العميق هو الفائز.

بالنسبة لـMLOFI، النتيجة المقاسة هي مقارنة جودة المطابقة عبر مستويات العمق، وليس عبر عائلات الميزات. باستخدام بيانات 6 أسهم عالية السيولة في Nasdaq، تُلائم الورقة علاقة خطية بين MLOFI والتغير المتزامن في منتصف السعر [6]، وبالنسبة لجميع الأسهم الستة المدروسة، تتحسن جودة المطابقة خارج العينة مع كل مستوى سعري إضافي يُدرج في متجه MLOFI [6]. هذه مقارنة داخل الميزة نفسها (المستوى 1 مقابل المستوى 2 مقابل مستويات أعمق من نفس بناء MLOFI)، وليست مقارنة بين MLOFI وتدفق التداول أو الاختلال أحادي المستوى I.

الشروط الكامنة وراء نتيجة CAC40 محددة، ووردت مرة واحدة، في قسم ملاحظات التنفيذ أدناه: نافذة مجموعة البيانات، والتسجيل بالميلي ثانية، وعمق الرؤية الخمسة مستويات، والاقتصار على الفترة من 09h05 إلى 17h25، ومجموعتا البيانات الصباحية والمسائية المنفصلتان [4].

بالنسبة لدراسة اختلال أعلى الدفتر، الكميات المقاسة هي متوسط حركة السعر المُطبّع ووقت الانتظار كدوال لـI، والرقم المحدد الوحيد المُبلَّغ عنه هو أن متوسط حركة السعر يمكن أن يصل إلى ثلث الفارق في دفتر شديد الاختلال [2, الفقرة 646]. لا يوجد رقم للدقة (accuracy) أو الضبط (precision) أو R-squared مرتبط بهذه المطالبة في المصدر؛ إنها بيان مقدار عن حركات السعر، وليست إحصائية تقييم نموذج. مجتمعةً، تجيب الأدلة المقاسة عن أسئلة أضيق من السؤال المطروح هنا: كم عدد مستويات العمق التي يجب إدراجها في متجه MLOFI [6]، وأي الميزات تنجو من اختيار LASSO على أسهم CAC40 [4]، وما مدى كبر حركة السعر التالية نسبة إلى الفارق [2, الفقرة 646]، وكيف يقارن نموذج خطي بنماذج تعلّم عميق على معيار صيني بمقياس الثانية [1].

ملاحظات التنفيذ: البيانات والدقة ونطاق مجموعات البيانات المصدرية

يحتاج البنّاء الذي يعيد إنتاج أياً من عائلات الميزات هذه إلى معرفة البيانات الدقيقة التي استخدمتها كل دراسة، لأن تردد أخذ العينات والعمق يختلفان بشكل حاد. تستخدم دراسة القفزة السعرية على CAC40 مجموعة بيانات تضم صفقات وأنشطة أوامر محدودة لأسهم CAC40 الأربعين الأعضاء بين 1 أبريل 2011 و30 أبريل 2011 [4]، مع تسجيل كل معاملة وكل تعديل لدفتر الأوامر المحدودة بالميلي ثانية [4]. تستخدم عمق رؤية خمسة مستويات، بحيث L=5 [4]، ولتجنب ساعات الافتتاح والإغلاق، تُقيَّد البيانات على الفترة من 09h05 إلى 17h25 [4]. كما تُهمل صراحةً أوامر الإيقاف (stop orders) وأوامر الجبل الجليدي (iceberg orders) لأنها نادرة نسبياً مقارنة بأحداث الأوامر المحدودة وأوامر السوق [4]، ولإزالة الموسمية داخل اليوم، يُبنى التحليل على مجموعتي بيانات صباحية ومسائية منفصلتين [4]. كل من هذه قرار معالجة أولية ملموس ينبغي على البنّاء نسخه أو الانحراف عنه بوعي.

يمتد معيار السوق الصيني على بضعة آلاف من الأسهم خلال يونيو إلى سبتمبر 2020 [1]، بتسميات بمقياس الثانية عبر 12 أفقاً زمنياً من 1 إلى 300 ثانية [1]. هذه دقة زمنية أخشن بكثير من تعديلات LOB بمقياس الميلي ثانية، لذا فهي تناسب هندسة ميزات بمقياس الثانية أكثر من مقياس الجزء (tick).

تعمل دراسة MLOFI على Nasdaq انطلاقاً من 6 أسهم عالية السيولة [6]، وهي عينة صغيرة ومنتقاة يدوياً اختيرت بسبب السيولة، وليست عالماً واسعاً. البنّاء الذي يعمم MLOFI إلى عالم جديد من الأصول ينبغي أن يعامل التحسن المُبلَّغ عنه في المطابقة خارج العينة مع كل مستوى سعري مُضاف [6] كفرضية يُعاد اختبارها على بياناته الخاصة، وليس كخاصية مُثبتة بالفعل خارج نطاق تلك الأسهم الستة.

نظام مرجعي حي متعدد النماذج للتنبؤ بدفتر الأوامر

يُوصف في المصادر نظام حديث واحد بُني صراحة للنشر في الوقت الفعلي بدلاً من الدراسة غير المتصلة، وهو مرجع معماري مفيد رغم أنه لم يُختبر هنا مقابل ميزات الاختلال أعلاه. يدمج النظام مجموعات بيانات تاريخية (FI-2010، LOBSTER، BTC) مع تغذيات سوق حية قائمة على WebSocket، ومحرك معالجة أولية معياري (modular)، وعدة معماريات تعلّم عميق للتنبؤ بدفتر الأوامر المحدودة في الوقت الفعلي [5]. تُطبِّع معالجته الأولية لقطات دفتر الأوامر الواردة باستخدام Z-score وBatch-Instance Normalization (BiN) [5]، وتستخرج ميزات سعرية وحجمية متعددة المستويات وتولّد تسميات حركة متعددة الآفاق [5]، وهو ما يتسق مفاهيمياً مع التصاميم متعددة المستويات ومتعددة الآفاق المستخدمة في أدبيات الاختلال أعلاه، رغم كونه ورقة هندسية منفصلة، وليس دراسة مقارنة للميزات.

النماذج المُسمّاة هي MLPLOB، وTLOB، وDeepLOB، وBinCTABL، مُدرَّبة على بيانات مُعالجة مسبقاً ومُقدَّمة عبر خلفية FastAPI [5]. للاستدلال الحي، تبث واجهة أمامية بيانات بورصة في الوقت الفعلي مثل Binance عبر WebSockets، وتُعالج الخلفية هذه البيانات لإنتاج تنبؤات فورية [5]. ولأن هذا الأنبوب (pipeline) يتعامل مع بيانات العملات الرقمية (BTC) مباشرة، فهو الجزء الوحيد من قاعدة الأدلة هذه الذي يلامس الاستدلال الإنتاجي على العملات الرقمية، رغم أنه لا يُبلغ عن مقارنة عائلات الميزات بين OFI مقابل اختلال العمق مقابل تدفق التداول على تلك البيانات.

بشكل مفيد للبنّاء الذي يريد مقارنة مجموعات الميزات أو عائلات النماذج بأمانة، يسمح النظام بتحميل عدة نماذج في وقت واحد لكن تُعاد تنبؤاتها بشكل مستقل لتمكين مقارنة شفافة [5]. هذا التصميم، تشغيل عدة نماذج جنباً إلى جنب والإبلاغ عن كل مخرَج بشكل منفصل بدلاً من متوسط تجميعي (ensemble)، هو بالضبط النمط اللازم لتشغيل مقارنة walk-forward عادلة بين مجموعات الميزات القائمة على OFI، وMLOFI، وتدفق التداول دون أن يُخفي أحدها نمط خطأ الآخر.

لا شيء من هذا مبني على تشغيل CPU فقط بحكم البنية: بحسب قراءتنا الخاصة، وليس كمطالبة مصدرية، فإن DeepLOB وTLOB والمعماريات المشابهة هي عادةً شبكات تعلّم عميق تُدرَّب على GPU. بالنسبة لهدف إنتاجي يعتمد على CPU فقط مثل sam، من الأفضل قراءة هذا النظام كنمط معماري (نماذج متعددة، تقديم مستقل، تلقّي حي عبر WebSocket) لتقليده باستخدام نماذج خطية أو لوجستية رخيصة بدلاً من الشبكات العميقة، وليس كنظام يُنشر دون تعديل.

الحدود والأسئلة المفتوحة

كما ذُكر في الإجابة المباشرة أعلاه، لا تحتوي قاعدة الأدلة هذه على اختبار وجهاً لوجه بين OFI واختلال العمق متعدد المستويات ومجموعات ميزات تدفق التداول على نفس البيانات ضمن تحقق متقاطع من نوع walk-forward مُنقّى ومحاط بعزل، ولا يُجرى أي منها بشكل مشترك عبر ثلاث بورصات عملات رقمية مُسمّاة والأسهم. تقارن دراسة السوق الصيني الانحدار الخطي بنماذج التعلّم العميق على أهداف السعر والحجم المرجّحين بالحجم بمقياس الثانية [1]، لكنها لا تعزل OFI مقابل اختلال العمق مقابل تدفق التداول كمدخلات منفصلة. تُبلغ دراسة MLOFI عن تحسن المطابقة مع العمق على 6 أسهم في Nasdaq [6]، لكنها لا تقارن MLOFI بميزات تدفق التداول. تختار دراسة CAC40 إشارة التداول، وحجم الأمر، وسيولة أعلى الدفتر كمعلوماتية باستمرار [4]، لكن هذه نتيجة اختيار متغيرات داخل النموذج، وليست مقارنة معتمدة على تحقق متقاطع خارج العينة مقابل نموذج اختلال عمق مُعرَّف بشكل منفصل. تُثبت دراسة اختلال أعلى الدفتر علاقة تنبؤية وتحذر صراحة من أنها ليست فرصة مراجحة قائمة بذاتها [2]، وهو تحذير قوي ضد المبالغة في بيع أي استراتيجية قائمة على الاختلال وحده.

الحد الثاني هو تغطية فئة الأصول. جميع دراسات الاختلال وMLOFI والقفزة السعرية هي على الأسهم (الأسهم الصينية [1]، أسهم Nasdaq [6]، أسهم CAC40 [4])؛ النظام الوحيد الذي يلامس العملات الرقمية هنا هو معمارية النشر الحي [5]، التي لا تُبلغ عن مقارنة ميزات. لذا فإن تطبيق ترتيبات الميزات المستمدة من الأسهم على دفاتر أوامر العملات الرقمية، عبر ثلاث بورصات، هو استقراء لا تدعمه الأدلة؛ يجب اختباره، لا افتراضه.

الحد الثالث هو أن أياً من المطالبات المُستشهد بها لا يذكر تحقق walk-forward المُنقّى والمحاط بعزل كبروتوكول تقييمه. يستخدم المعيار الصيني بنية أفق زمني محددة [1]، وتستخدم دراسة CAC40 بيانات مقسمة صباحاً/مساءً [4]، وتُبلغ أوراق الاختلال عن علاقات متوسطة على مستوى الجمهور دون بروتوكول تقسيم تدريب/اختبار محدد على الإطلاق [2][3]. البنّاء الذي يحتاج إلى تحقق walk-forward مُنقّى ومحاط بعزل لنموذج إنتاجي يجب أن يفرض هذا البروتوكول بنفسه؛ فهو غير موجود مسبقاً في أي طريقة مُستشهد بها.

أخيراً، تحذير تصميم التسمية الذي نُوقش في المقدمة يأتي من دراسة واحدة فقط [1]. بناءً على ذلك، الاستجابة العملية هي بناء أنبوب الميزات الموصوف هنا، لكن تقييمه بالبروتوكول الأكثر صرامة الذي يتطلبه السؤال، على بيانات العملات الرقمية والأسهم الخاصة بالبنّاء، قبل استخلاص أي استنتاج مقارن.

كيفية بنائه، أو كيفية استخدامه

خطوات بناء مستندة إلى المطالبات (كل منها مستندة إلى فقرة مُستشهد بها):

  1. استوعب لقطات LOB أولاً. كل خطوة أدناه تتطلب تغذية عاملة للقطات دفتر الأوامر المحدودة، لأن الميزات المُعرَّفة في المطالبات المُستشهد بها تعمل على كميات أعلى الدفتر ومتعددة المستويات، وليس على أشرطة OHLCV وحدها.
  2. اجمع لقطات LOB لكل سلسلة بورصة أو أسهم مستهدفة. احتفظ بعدة مستويات عمق لكل لقطة؛ للمرجعية، استخدمت دراسة CAC40 عمق رؤية خمسة مستويات، بحيث L=5 [4]، وتحسنت المطابقة خارج العينة مع كل مستوى سعري إضافي يُدرج في متجه MLOFI على أسهم Nasdaq الستة المدروسة [6].
  3. احسب الاختلال أحادي المستوى I=(qb-qa)/(qb+qa) من كمية أفضل عرض qb وكمية أفضل طلب qa، متبعاً تعريف تلك الصيغة على الكميات المنشورة في أعلى الدفتر [2]. خزّن إشارتها (الموجب يعني أثقل من جانب العرض، والسالب يعني أثقل من جانب الطلب) [2] كميزة أساسية.
  4. ابنِ متجه MLOFI بتوسيع فكرة الاختلال إلى صافي تدفق الأوامر عند مستويات سعرية متعددة بدلاً من نسبة لقطة أحادية المستوى، متبعاً تعريف MLOFI كمية متجهة تقيس صافي تدفق أوامر الشراء والبيع عند مستويات سعرية مختلفة [6]. احتفظ بأكبر عدد ممكن من المستويات يسمح به عمق لقطة LOB لديك، إذ تحسنت المطابقة مع كل مستوى مُضاف في دراسة الأسهم المُستشهد بها [6].
  5. ابنِ مجموعة ميزات تدفق التداول بشكل منفصل: أحجام الأوامر المحدودة، وفجوات أسعار الأوامر المحدودة، ومعلومات أوامر السوق، ومعلومات أحداث الأوامر المحدودة [4]، بما في ذلك إشارة التداول وحجم أمر السوق [4] والسيولة عند أفضل عرض وأفضل طلب [4]. استبعد أوامر الإيقاف وأوامر الجبل الجليدي من بناء هذه الميزة إن كانت نادرة في بياناتك، متبعاً منطق دراسة CAC40 [4].
  6. إذا أردت تسمية اتجاهية (قفزة/لا قفزة)، عرّفها بالضبط كما فعلت دراسة CAC40: وصول أمر سوق بيع (شراء) يُنفَّذ بسعر أصغر (أكبر) من سعر أفضل عرض (أفضل طلب) مباشرة بعد وصول أمر السوق السابق [4].
  7. قسّم البيانات داخل اليوم حسب نصف الجلسة (صباحاً/مساءً أو ما يعادلها) للسيطرة على الموسمية داخل اليوم قبل المطابقة، كما فُعل في دراسة CAC40 [4]، وقيّد الابتعاد عن ساعات الافتتاح/الإغلاق إذا كانت بورصتك تعاني من تأثيرات حافة مماثلة، متبعاً منطق نفس الورقة في التقييد بـ09h05-17h25 [4].
  8. لائم نموذجين أساسيين لكل عائلة ميزات: انحدار خطي لهدف مقدار (متبعاً النموذج الخطي المستخدم كأساس واحد في مقارنة المعيار الصيني [1] والمطابقة الخطية المستخدمة لـMLOFI مقابل تغير منتصف السعر [6]) وانحدار لوجستي، اختيارياً مع اختيار متغيرات LASSO، لهدف اتجاهي/قفزة (متبعاً منهجية دراسة CAC40 [4]).
  9. قيّم مكوّن وقت الانتظار باستخدام تعريف وقت التوقف المُعطى في دراسة الاختلال: بالنسبة لأمر عند العرض، أول وصول لصفقة بيع؛ وبالنسبة لأمر عند الطلب، أول وصول لصفقة شراء [2]. أبلغ عن وقت الانتظار وحركة السعر المُطبّعة كدوال لميزة الاختلال (هذه الخطوة مبنية على استنتاجنا الخاص من تحذير تصميم التسمية [1]، وليست تعليمات مباشرة من تلك المطالبة)، محاكياً النتيجة المُبلَّغ عنها بأن الدفتر شديد الاختلال يدل على أن حركة السعر من المرجح أن تأتي قريباً [2].
  10. لا تعامل رقم المطابقة التنبؤية على أنه استراتيجية: تنص الأدلة المصدرية صراحة على أن اختلال الدفتر، حتى عندما يكون تنبؤياً، لا يوفر بحد ذاته فرصة لمراجحة إحصائية مباشرة [2].

استنتاج المؤلف الخاص، غير منصوص عليه من أي مطالبة مُستشهد بها: الفقرة التي تنتقد تسميات اتجاه منتصف السعر أحادية الجزء (single-tick) [1] تنص على أن هذه التسمية مبسّطة أكثر من اللازم لاستراتيجية تداول عملية، لكنها لا تصف بديلاً. بناءً على استدلالنا الخاص، قد يفكر البنّاء بدلاً من ذلك في تسمية بأفق 1 إلى 5 دقائق، أو التصميم ثنائي المخرَج الذي يحاكي ما قِيس لاختلال أعلى الدفتر (متوسط حركة سعر مُطبّع، ووقت انتظار حتى حركة منتصف السعر التالية) [2]، بدلاً من الافتراضي وهو تسمية اتجاه أحادية الجزء. هذه التوصية هي توصيتنا، وليست توصية المطالبة.

مقارنة عائلات الميزات ببعضها البعض (الاختلال مقابل MLOFI مقابل تدفق التداول، على نفس البيانات، ضمن بروتوكول walk-forward مُنقّى ومحاط بعزل) ليست شيئاً تفعله أي مطالبة مُستشهد بها بالفعل؛ انظر الإجابة المباشرة أعلاه لهذا التحذير، المذكور مرة واحدة.

الكود: تطبيق عملي

ملحق: كود توضيحي لبنية عمل (harness)، بناء المؤلف الخاص، غير مستند إلى أي مطالبة مُستشهد بها. لا شيء من الهندسة الملموسة أدناه (خطوة StandardScaler، مقسّم walk-forward المُنقّى/المحاط بعزل، تهيئة LASSO اللوجستي، بناء نطاقات الكميّات) محدد من قبل أي مطالبة في هذه المذكرة؛ توفر المطالبات فقط تعريفات الميزات وعائلات النماذج، وليس هذا التطبيق. أُدرج هذا كبنية عمل توضيحية (harness) للأنبوب.

وهو أيضاً ليس تطبيقاً لخطوات البناء أعلاه. تعمل الطرق المُستشهد بها هنا على لقطات دفتر الأوامر المحدودة: الاختلال أحادي المستوى I=(qb-qa)/(qb+qa) مُعرَّف على كميات العرض والطلب المنشورة في أعلى الدفتر [2, الفقرة 649]، وMLOFI مُعرَّف على صافي تدفق الأوامر عند عدة مستويات عمق [6]. يحتوي جدولنا bars فقط على بيانات OHLCV ولا يحتوي على جدول LOB، لذا فإن الميزات أدناه هي بدائل مبنية على OHLCV فقط تشترك في شكل تلك الميزات ولكن ليس في مدخلاتها. بسبب ذلك، تتطلب خطوات البناء في القسم أعلاه خطوة استيعاب لقطات LOB منفصلة قبل كتابة كود ميزات مستند إلى المطالبات؛ يمارس سكريبت هذا الملحق الآليات المحيطة (تجميع الميزات، مطابقات منفصلة لكل عائلة: اختلال فقط، MLOFI فقط، تدفق تداول فقط، ومجموعة مشتركة، تقييم walk-forward، كتابة التنبؤات) بحيث تكون بنية العمل جاهزة بمجرد إضافة قارئ LOB حقيقي.

لا يمكن تقديم تطبيق قابل للتنفيذ مستند إلى المطالبات المُستشهد بها من المطالبات والبيانات المُعطاة. تتطلب الطرق الموصوفة في هذه المذكرة لقطات دفتر أوامر محدودة كمدخل: الاختلال أحادي المستوى I=(qb-qa)/(qb+qa) مُعرَّف على كميات العرض والطلب المنشورة في أعلى الدفتر [2, الفقرة 649]، وMLOFI كمية متجهة على صافي تدفق الأوامر عند مستويات سعرية مختلفة في الدفتر [6]، وتُستخرج ميزات تدفق التداول من أحجام الأوامر المحدودة، وفجوات أسعار الأوامر المحدودة، ومعلومات أوامر السوق، ومعلومات أحداث الأوامر المحدودة [4, الفقرة 661]. الجداول المتاحة لتنفيذ الكود (bars، forecasts، trades، book) لا تُثبت أن لقطات دفتر الأوامر المحدودة بالشكل الذي تتطلبه هذه المطالبات موجودة: bars هو جدول OHLCV، وأي بديل مشتق من OHLCV لاختلال أو MLOFI أو تدفق التداول لن يكون اختباراً للمطالبات المُستشهد بها، بل مجرد محاكاة لشكلها على بيانات مدخل مختلفة وغير موثّقة. بدلاً من تقديم كود يبدو وكأنه يطبق الطرق المُستشهد بها بينما يعمل فعلياً على بيانات مختلفة وغير موثّقة، ينص هذا القسم بوضوح على أنه لا يمكن بناء مثل هذا التطبيق دون التأكد أولاً من أن جدول book (أو مصدر لقطة LOB مكافئ) يحتوي على كميات العرض/الطلب عند أعلى الدفتر ومتعددة المستويات كما تُعرّفها المطالبات، بالإضافة إلى تحديد كيفية محاذاة تلك اللقطات زمنياً مع جدولي trades وbars. لا شيء من هذا مُثبت في المطالبات المُعطاة.

يقرأ السكريبت أشرطة (bars) دقيقة واحدة لجميع الرموز في قاعدة البيانات من bars(symbol, tf, ts, open, high, low, close, volume)، ويبني الميزات البديلة، ويُلائم انحداراً خطياً للتنبؤ بعائد الشريط التالي وانحداراً لوجستياً بعقوبة L1 للتنبؤ بإشارة الحركة التالية، ويُلائم ويقيّم كل عائلة ميزات بشكل منفصل (اختلال فقط، MLOFI فقط، تدفق تداول فقط) بالإضافة إلى المجموعة المشتركة تحت تقسيم walk-forward المُنقّى والمحاط بعزل الخاص بالمؤلف، ويكتب تنبؤات q10/q50/q90/p_up في جدول forecasts. الأفق الزمني هو معلمة يضبطها المستخدم في أعلى main() (HORIZON)، في أي مكان ضمن نافذة 1 إلى 5 دقائق؛ لا تختار أي مطالبة مُستشهد بها قيمة محددة داخل تلك النافذة، لذا فالاختيار متروك للمُستدعي. كل صف تنبؤ مختوم بـmade_atالذي يسجّل الطابع الزمني للشريط الذي صدر منه التنبؤ (حقل ts الخاص بالشريط نفسه)، وليس الوقت الفعلي (wall-clock) الذي نُفِّذ فيه السكريبت؛ هذا يسمح لجدول التنبؤات بربط كل تنبؤ بلقطة السوق الدقيقة التي وُلِّد منها. المعيار الذي يجب التفوق عليه هو نموذج استمرارية ساذج (predict zero return, predict p_up=0.5)؛ والرقم الذي يجب التحقق منه هو R-squared خارج العينة للنموذج الخطي والدقة (accuracy) خارج العينة للنموذج اللوجستي، وكلاهما مقابل ذلك المعيار الساذج.

""" Illustrative harness: OFI-style feature pipeline and CPU-only baselines on OHLCV bars.
This script is the author's own construction. No cited claim specifies the
scaling, cross-validation splitter, solver settings or quantile band used here.
Runs fully offline (no network). Python 3.12, numpy, pandas, scikit-learn, sqlite3.

The script expects bars(symbol, tf, ts, open, high, low, close, volume) and
forecasts(symbol, horizon, made_at, q10, q50, q90, p_up) tables to exist in
data.sqlite.

IMPORTANT: this schema has no limit-order-book snapshot table, so this script
CANNOT test the cited claims. The single-level imbalance I = (qb - qa)/(qb + qa)
[2, passage 649] and MLOFI over depth levels [6] both require book-level
quantities we do not have here. The "imbalance", "MLOFI" and "trade-flow"
features below are OHLCV-only proxies that reproduce the *shape* of those
feature families, not their inputs:
  - single-level imbalance proxy, shaped after I = (qb - qa) / (qb + qa)   [2, passage 649]
  - MLOFI proxy, a vector over pseudo-levels, shaped after MLOFI          [6]
  - trade-flow / signed volume proxy, shaped after the CAC40 feature set  [4]
A real LOB-snapshot ingestion step must be added before the build steps in
"How to build it" can be implemented against this harness.
"""

import sqlite3
import numpy as np
import pandas as pd
from sklearn.linear_model import LinearRegression, LogisticRegression
from sklearn.preprocessing import StandardScaler

DB_PATH = "data.sqlite"

# User-set parameters. HORIZON is the forecast horizon in minutes; any value in
# the 1-to-5 minute window is valid and no cited claim prefers one over another,
# so the caller picks it. EMBARGO is our own cross-validation choice (see
# purged_embargo_walk_forward_splits).
HORIZON = 3
EMBARGO = 5


def list_symbols(conn, tf="1m"):
    # Returns every symbol that has bars at the given timeframe.
    q = "SELECT DISTINCT symbol FROM bars WHERE tf = ? ORDER BY symbol ASC"
    return [r[0] for r in conn.execute(q, (tf,)).fetchall()]


def load_bars(conn, symbol, tf="1m"):
    # Reads one symbol's bars, ordered by timestamp. Input: sqlite3 connection,
    # symbol string, timeframe string. Output: pandas DataFrame.
    q = """
        SELECT symbol, tf, ts, open, high, low, close, volume
        FROM bars WHERE symbol = ? AND tf = ? ORDER BY ts ASC
    """
    df = pd.read_sql_query(q, conn, params=(symbol, tf))
    return df


def build_single_level_imbalance_proxy(df):
    # Proxy only. The real feature is I = (qb - qa) / (qb + qa), defined over
    # top-of-book bid and ask quantities [2, passage 649], which this schema
    # does not store. We build a directional pressure proxy from the bar itself:
    # how close the close is to the high (buy pressure) versus the low (sell pressure).
    rng = (df["high"] - df["low"]).replace(0, np.nan)
    buy_pressure = (df["close"] - df["low"]) / rng
    sell_pressure = (df["high"] - df["close"]) / rng
    imbalance = (buy_pressure - sell_pressure).fillna(0.0)
    return imbalance


def build_mlofi_proxy(df, levels=(1, 2, 3, 5)):
    # Proxy only. Real MLOFI is a vector of net order flow across depth levels
    # of the book [6]. With OHLCV alone we build lagged/smoothed versions of the
    # single-level proxy so the model receives a vector of related signals,
    # matching the vector shape of MLOFI but not its book-level definition.
    base = build_single_level_imbalance_proxy(df)
    out = {}
    for lvl in levels:
        out[f"mlofi_l{lvl}"] = base.rolling(lvl, min_periods=1).mean()
    return pd.DataFrame(out)


def build_trade_flow_features(df):
    # Proxy for the trade-flow family found consistently informative on CAC40:
    # trade sign and market order size [4]. Signed volume: sign of the bar
    # return times the bar volume.
    ret = df["close"].pct_change().fillna(0.0)
    sign = np.sign(ret)
    signed_volume = sign * df["volume"]
    trade_intensity = df["volume"].rolling(5, min_periods=1).mean()
    return pd.DataFrame({
        "signed_volume": signed_volume,
        "trade_intensity": trade_intensity,
    })


def build_feature_matrix(df, horizon):
    # Assembles the full feature set and the two targets:
    #   y_ret: forward return over `horizon` bars (regression target)
    #   y_dir: sign of forward return (classification target)
    # `horizon` is explicit (no default) so the 1-to-5 minute target range is
    # always a deliberate choice at the call site.
    imb = build_single_level_imbalance_proxy(df).rename("imbalance")
    mlofi = build_mlofi_proxy(df)
    tf = build_trade_flow_features(df)
    feats = pd.concat([imb, mlofi, tf], axis=1)

    fwd_close = df["close"].shift(-horizon)
    y_ret = (fwd_close / df["close"] - 1.0)
    y_dir = (y_ret > 0).astype(int)

    data = pd.concat(
        [df["ts"], feats, y_ret.rename("y_ret"), y_dir.rename("y_dir")], axis=1
    )
    # Drop rows with missing target labels (forward returns that fall off the
    # end of the series). This determines which rows are used for training.
    data = data.dropna().reset_index(drop=True)
    return data


def feature_families(feature_cols):
    # Splits the combined feature set into the three families the note compares,
    # so each can be fitted and scored on its own as well as together.
    imbalance = [c for c in feature_cols if c == "imbalance"]
    mlofi = [c for c in feature_cols if c.startswith("mlofi_")]
    trade = [c for c in feature_cols if c in ("signed_volume", "trade_intensity")]
    return {
        "imbalance-only": imbalance,
        "mlofi-only": mlofi,
        "trade-flow-only": trade,
        "combined": list(feature_cols),
    }


def purged_embargo_walk_forward_splits(n, n_splits=5, horizon=HORIZON, embargo=EMBARGO):
    # Author's own purged, embargoed walk-forward split. Expanding train window,
    # a test block after it, and purging of rows within `horizon` bars of the
    # test start (because forward-looking labels from those rows overlap with
    # the test set), plus an embargo gap between the purge cutoff and test start
    # to simulate real-time deployment latency. This splitter removes all training
    # rows within `horizon` bars of the test period, implementing true purging
    # as well as an embargo; the horizon and embargo values are our own choice.
    fold_size = n // (n_splits + 1)
    splits = []
    for i in range(1, n_splits + 1):
        test_start = fold_size * (i + 1)
        test_end = min(test_start + fold_size, n)
        if test_start >= test_end:
            continue
        # Purge: remove training rows within `horizon` of test_start (their labels
        # reach into the test period). Embargo: add a further gap of `embargo` bars.
        purge_cutoff = test_start - horizon - embargo
        if purge_cutoff  0 else 0.0
    acc_naive = max(y_dir_test.mean(), 1 - y_dir_test.mean())
    return r2_naive, acc_naive


def make_lasso_logistic():
    # L1-penalized (LASSO) logistic regression. Passage 661 states that LASSO
    # logistic regression is used for variable selection over LOB features [4],
    # but it does not specify the solver, l1_ratio or max_iter, so these
    # hyperparameters are our own choice.
    return LogisticRegression(penalty="l1", solver="saga", l1_ratio=1, max_iter=1000)


def fit_and_evaluate(data, feature_cols, horizon):
    # Two model choices, each mapped to a specific claim:
    #  - Linear regression on the return target mirrors the linear regression
    #    model that is compared against deep learning models on the Chinese
    #    stock market LOB benchmark [1], and the linear relationship
    #    fitted between MLOFI and the contemporaneous mid-price change on 6
    #    Nasdaq stocks [6].
    #  - LASSO logistic regression on the direction target mirrors the LASSO
    #    logistic regression used to predict price jumps and to perform
    #    variable selection over LOB features on CAC40 stocks [4]; the solver,
    #    l1_ratio and max_iter settings are our own choice, not stated in that claim.
    # StandardScaler normalization is our own choice; claim [5] mentions Z-score
    # and Batch-Instance Normalization for a different system, but StandardScaler
    # is neither of those and is explicitly flagged here as an author's addition.
    n = len(data)
    splits = purged_embargo_walk_forward_splits(n, n_splits=5, horizon=horizon, embargo=EMBARGO)

    lin_r2s, log_accs = [], []
    lin = log = scaler = None
    for train_idx, test_idx in splits:
        train = data.iloc[train_idx]
        test = data.iloc[test_idx]

        scaler = StandardScaler()  # author's own normalization choice
        X_train = scaler.fit_transform(train[feature_cols])
        X_test = scaler.transform(test[feature_cols])

        # Linear regression on forward return: the linear baseline compared
        # against deep learning models on the China benchmark [1], and the
        # linear MLOFI-vs-mid-price-change fit [6].
        lin = LinearRegression()
        lin.fit(X_train, train["y_ret"])
        pred_ret = lin.predict(X_test)
        ss_res = np.sum((test["y_ret"] - pred_ret) ** 2)
        ss_tot = np.sum((test["y_ret"] - test["y_ret"].mean()) ** 2)
        r2 = 1 - ss_res / ss_tot if ss_tot > 0 else 0.0
        lin_r2s.append(r2)

        # LASSO-style logistic regression on direction: the LASSO logistic
        # variable-selection method applied to price-jump prediction on
        # CAC40 limit order book features [4], with our own solver, l1_ratio
        # and max_iter settings.
        if train["y_dir"].nunique()  r2_naive)
    print(f"[{symbol}] model beats baseline on accuracy:", acc > acc_naive)

    write_forecasts(conn, symbol, horizon, lin, log, scaler, feature_cols, data)


def main():
    conn = sqlite3.connect(DB_PATH)
    horizon = HORIZON  # minutes ahead; user-set parameter, any value in the 1-to-5 minute window

    symbols = list_symbols(conn, tf="1m")
    if not symbols:
        print("No symbols with 1m bars found.")
        conn.close()
        return
    print("Symbols found:", symbols)

    for symbol in symbols:
        run_symbol(conn, symbol, horizon)

    conn.close()


if __name__ == "__main__":
    main()

ما الذي كنا سنبنيه

ما يلي هو خطة مشروعنا الخاصة. لا يظهر أي جدول زمني، أو تقدير طاقم عمل، أو ترتيب مهام أدناه في أي مطالبة مُستشهد بها.

سنبني بنية عمل صغيرة لمقارنة الميزات، تعمل على CPU فقط لـsam، باستخدام أشرطتنا الخاصة بـSQLite عبر ثلاث بورصات عملات رقمية وسلسلة أسهم واحدة. في مرحلة أولى، سننفذ ميزات على مستوى LOB (الاختلال أحادي المستوى I كما هو مُعرَّف في [2]، وMLOFI كمتجه عبر مستويات العمق كما في [6]، وميزات تدفق التداول كما في [4]) من لقطات دفتر الأوامر الخام، وليس البدائل القائمة على OHLCV المستخدمة في الكود التوضيحي أعلاه، لأن بيانات LOB الحقيقية مطلوبة لاختبار المطالبات الفعلية. في مرحلة ثانية، سنُجري تقسيم walk-forward المُنقّى والمحاط بعزل الخاص بنا عبر البورصات الثلاث وسلسلة الأسهم، مُلائمين نفس أسس النماذج الخطية وLASSO اللوجستية لكل عائلة ميزات بشكل منفصل وبالمجموع.

سنحكم على المشروع بناءً على ما إذا كانت أي عائلة ميزات تتفوق على معيار الاستمرارية الساذج (عائد صفري، p_up=0.5) في R-squared خارج العينة ومعدل الإصابة (hit-rate)، لكل بورصة وسلسلة الأسهم بشكل منفصل، وبناءً على ما إذا كانت مطابقة MLOFI تتحسن بشكل رتيب مع مستوى العمق كما حدث على أسهم Nasdaq الستة في [6]، مُتحقَّق منه الآن على دفاتر أوامر العملات الرقمية الخاصة بنا. سنُبلغ صراحة، بدلاً من الاستنتاج، عمّا إذا كان ترتيب الميزات المستمد من الأسهم في [4] ينتقل إلى العملات الرقمية، إذ لا تختبر أي مطالبة مُستشهد بها ذلك بالفعل.

التكلفة، مرة أخرى تقديرنا الخاص: يتطلب هذا تخزين لقطات LOB لثلاث بورصات (أكبر من OHLCV وحده) ومطابقة خطية/لوجستية على CPU فقط، لذا لا حاجة لميزانية GPU؛ التكلفة الرئيسية هي جمع البيانات وتخزينها لتاريخ LOB بمستوى الجزء أو مستوى اللقطة، بالإضافة إلى وقت الهندسة.

الادعاءات والمراجعة

  1. factمدعوم

    Passage 644 states that the benchmark LOB dataset from the Chinese stock market covers a few thousand stocks from June to September 2020.

    [1] Benchmark Dataset for Short-Term Market Prediction of Limit Order Book in China Markets, abstract S2 7aaaecac0598
    “Limit order books (LOBs) have generated big financial data for analysis and prediction from both academic community and industry practitioners. This article presents a benchmark LOB dataset from the Chinese stock market, covering a few thousand stocks for the period of June to Se…”
  2. methodمدعوم

    Passage 644 states that the experiment protocol forecasts the upcoming volume-weighted average price change and volume at the end of every second over 12 horizons ranging from 1 second to 300 seconds.

    [1] Benchmark Dataset for Short-Term Market Prediction of Limit Order Book in China Markets, abstract S2 7aaaecac0598
    “Limit order books (LOBs) have generated big financial data for analysis and prediction from both academic community and industry practitioners. This article presents a benchmark LOB dataset from the Chinese stock market, covering a few thousand stocks for the period of June to Se…”
  3. methodمدعوم

    Passage 644 states that results based on a linear regression model and deep learning models are compared.

    [1] Benchmark Dataset for Short-Term Market Prediction of Limit Order Book in China Markets, abstract S2 7aaaecac0598
    “Limit order books (LOBs) have generated big financial data for analysis and prediction from both academic community and industry practitioners. This article presents a benchmark LOB dataset from the Chinese stock market, covering a few thousand stocks for the period of June to Se…”
  4. methodمدعوم

    Passage 644 states that a more practically effective set of features is proposed to capture both LOB snapshots and periodic data.

    [1] Benchmark Dataset for Short-Term Market Prediction of Limit Order Book in China Markets, abstract S2 7aaaecac0598
    “Limit order books (LOBs) have generated big financial data for analysis and prediction from both academic community and industry practitioners. This article presents a benchmark LOB dataset from the Chinese stock market, covering a few thousand stocks for the period of June to Se…”
  5. limitationمدعوم

    Passage 644 states that predicting mid-price direction change for the next few events is too simplistic and not suitable for a practical trading strategy.

    [1] Benchmark Dataset for Short-Term Market Prediction of Limit Order Book in China Markets, abstract S2 7aaaecac0598
    “Limit order books (LOBs) have generated big financial data for analysis and prediction from both academic community and industry practitioners. This article presents a benchmark LOB dataset from the Chinese stock market, covering a few thousand stocks for the period of June to Se…”
  6. factمدعوم

    Passage 646 states that the average mid price move normalized by the bid-ask spread and the waiting time until the next mid price move are studied as functions of book imbalance.

    [2] Trade arrival dynamics and quote imbalance in a limit order book, section 1 Introduction
    “features by introducing a stochastic model for diffusion in three dimensions. We compute the probabilities of price movement and trade occurrence from the model, and calibrate them to recent historical market data. Figure 1: Average mid price move normalised by the bid-ask spread…”
  7. resultمدعوم

    Passage 646 states that the data display the non-martingale nature of prices at the short time scales considered.

    [2] Trade arrival dynamics and quote imbalance in a limit order book, section 1 Introduction
    “features by introducing a stochastic model for diffusion in three dimensions. We compute the probabilities of price movement and trade occurrence from the model, and calibrate them to recent historical market data. Figure 1: Average mid price move normalised by the bid-ask spread…”
  8. resultمدعوم

    Passage 646 states that in the case shown, the average price move can be up to a third of the spread in a highly imbalanced book.

    [2] Trade arrival dynamics and quote imbalance in a limit order book, section 1 Introduction
    “features by introducing a stochastic model for diffusion in three dimensions. We compute the probabilities of price movement and trade occurrence from the model, and calibrate them to recent historical market data. Figure 1: Average mid price move normalised by the bid-ask spread…”
  9. methodمدعوم

    Passage 647 states that the paper studies the probability of price movements and trade arrivals as a function of the quote imbalance at the top of the limit order book.

    [2] Trade arrival dynamics and quote imbalance in a limit order book, abstract arXiv:1312.0514v1
    “We examine the dynamics of the bid and ask queues of a limit order book and their relationship with the intensity of trade arrivals. In particular, we study the probability of price movements and trade arrivals as a function of the quote imbalance at the top of the limit order bo…”
  10. methodمدعوم

    Passage 647 states that a stochastic model is proposed to capture the joint dynamics of the top-of-book queues and the trading process.

    [2] Trade arrival dynamics and quote imbalance in a limit order book, abstract arXiv:1312.0514v1
    “We examine the dynamics of the bid and ask queues of a limit order book and their relationship with the intensity of trade arrivals. In particular, we study the probability of price movements and trade arrivals as a function of the quote imbalance at the top of the limit order bo…”
  11. factمدعوم

    Passage 649 defines bid-ask imbalance as I=(qb−qa)/(qb+qa), where qb and qa are the bid and ask quantities posted at the top of the book.

    [2] Trade arrival dynamics and quote imbalance in a limit order book, section 2 Empirical observations
    “A common intuition among market practitioners is that the order sizes displayed at the top of the book reflect the general intention of the market. When the number of shares available at the bid exceeds those at the ask, participants expect the next price movement to be upwards, …”
  12. factمدعوم

    Passage 649 states that positive imbalance indicates an order book heavier on the bid side and negative imbalance indicates one heavier on the ask side.

    [2] Trade arrival dynamics and quote imbalance in a limit order book, section 2 Empirical observations
    “A common intuition among market practitioners is that the order sizes displayed at the top of the book reflect the general intention of the market. When the number of shares available at the bid exceeds those at the ask, participants expect the next price movement to be upwards, …”
  13. methodمدعوم

    Passage 649 states that the effect of book imbalance on the average mid price change and on the waiting time until the next price change is calculated using the stopping time defined by the next change in either the best bid or the best ask.

    [2] Trade arrival dynamics and quote imbalance in a limit order book, section 2 Empirical observations
    “A common intuition among market practitioners is that the order sizes displayed at the top of the book reflect the general intention of the market. When the number of shares available at the bid exceeds those at the ask, participants expect the next price movement to be upwards, …”
  14. resultمدعوم

    Passage 650 states that a high book imbalance is, on average, a good predictor of mid price movements.

    [2] Trade arrival dynamics and quote imbalance in a limit order book, section 2 Empirical observations
    “Here, we use the same probabilities to compute the average size of the price jump11 1 A mid price change event can be induced by several actions, such as a trade, a cancellation or even the addition of a new quote between the current bid and ask spread, if the spread is big enoug…”
  15. resultمدعوم

    Passage 650 states that the price change until the next tick is well approximated by a linear function of the imbalance and is typically well below the bid-ask spread, even for highly imbalanced order books.

    [2] Trade arrival dynamics and quote imbalance in a limit order book, section 2 Empirical observations
    “Here, we use the same probabilities to compute the average size of the price jump11 1 A mid price change event can be induced by several actions, such as a trade, a cancellation or even the addition of a new quote between the current bid and ask spread, if the spread is big enoug…”
  16. limitationمدعوم

    Passage 650 states that although book imbalance may be used as a predictor for the next price movement, it does not by itself offer an opportunity for a straightforward statistical arbitrage.

    [2] Trade arrival dynamics and quote imbalance in a limit order book, section 2 Empirical observations
    “Here, we use the same probabilities to compute the average size of the price jump11 1 A mid price change event can be induced by several actions, such as a trade, a cancellation or even the addition of a new quote between the current bid and ask spread, if the spread is big enoug…”
  17. resultمدعوم

    Passage 650 states that highly imbalanced books indicate that a price move is likely to come in a relatively short time.

    [2] Trade arrival dynamics and quote imbalance in a limit order book, section 2 Empirical observations
    “Here, we use the same probabilities to compute the average size of the price jump11 1 A mid price change event can be induced by several actions, such as a trade, a cancellation or even the addition of a new quote between the current bid and ask spread, if the spread is big enoug…”
  18. factمدعوم

    Passage 651 states that for an order posted at the bid side, the relevant stopping time is the time of first arrival of a sell trade, and for an order posted at the ask side, it is the time of first arrival of a buy trade.

    [2] Trade arrival dynamics and quote imbalance in a limit order book, section 2 Empirical observations
    “context of microstructure studies, this is usually interpreted as the order flow having an impact on the limit book, but it can also be seen more generally as the natural supply and demand influence on the price of an asset. Another stopping time with economic significance is the…”
  19. resultمدعوم

    Passage 653 states that Cont et al. find a simple linear dependence between price changes and an indicator measuring imbalances between the order flow on the buy and sell sides of the LOB.

    [3] Stochastic Price Dynamics Implied By the Limit Order Book, section 1 Introduction
    “M. Bartolozzi [Bar10] proposes a multi-agent model for the dynamics of the LOB, with a particular focus on capturing key features of high-frequency trading. M. Avellaneda et al. [AS06] also propose a probabilistic framework for a utility optimizing agent in the context of high-fr…”
  20. resultمدعوم

    Passage 661 states that liquidity balance on best bid and best ask is quite informative for predicting the future market order's direction.

    [4] Price Jump Prediction in Limit Order Book, abstract arXiv:1204.1381v1
    “A limit order book provides information on available limit order prices and their volumes. Based on these quantities, we give an empirical result on the relationship between the bid-ask liquidity balance and trade sign and we show that liquidity balance on best bid/best ask is qu…”
  21. factمدعوم

    Passage 661 defines a price jump as a sell (buy) market order arrival executed at a price smaller (larger) than the best bid (best ask) price immediately after the preceding market order arrival.

    [4] Price Jump Prediction in Limit Order Book, abstract arXiv:1204.1381v1
    “A limit order book provides information on available limit order prices and their volumes. Based on these quantities, we give an empirical result on the relationship between the bid-ask liquidity balance and trade sign and we show that liquidity balance on best bid/best ask is qu…”
  22. methodمدعوم

    Passage 661 states that features are extracted from limit order volumes, limit order price gaps, market order information, and limit order event information.

    [4] Price Jump Prediction in Limit Order Book, abstract arXiv:1204.1381v1
    “A limit order book provides information on available limit order prices and their volumes. Based on these quantities, we give an empirical result on the relationship between the bid-ask liquidity balance and trade sign and we show that liquidity balance on best bid/best ask is qu…”
  23. methodمدعوم

    Passage 661 states that logistic regression is applied to predict the price jump from limit order book features.

    [4] Price Jump Prediction in Limit Order Book, abstract arXiv:1204.1381v1
    “A limit order book provides information on available limit order prices and their volumes. Based on these quantities, we give an empirical result on the relationship between the bid-ask liquidity balance and trade sign and we show that liquidity balance on best bid/best ask is qu…”
  24. methodمدعوم

    Passage 661 states that LASSO logistic regression is introduced to perform variable selection and highlight the importance of different features in predicting the future price jump.

    [4] Price Jump Prediction in Limit Order Book, abstract arXiv:1204.1381v1
    “A limit order book provides information on available limit order prices and their volumes. Based on these quantities, we give an empirical result on the relationship between the bid-ask liquidity balance and trade sign and we show that liquidity balance on best bid/best ask is qu…”
  25. methodمدعوم

    Passage 661 states that to remove intraday seasonality, the analysis is based on separate morning and afternoon datasets.

    [4] Price Jump Prediction in Limit Order Book, abstract arXiv:1204.1381v1
    “A limit order book provides information on available limit order prices and their volumes. Based on these quantities, we give an empirical result on the relationship between the bid-ask liquidity balance and trade sign and we show that liquidity balance on best bid/best ask is qu…”
  26. resultمدعوم

    Passage 661 states that, based on forty largest French stocks of CAC40, trade sign and market order size as well as liquidity on the best bid and best ask are consistently informative for predicting the incoming price jump.

    [4] Price Jump Prediction in Limit Order Book, abstract arXiv:1204.1381v1
    “A limit order book provides information on available limit order prices and their volumes. Based on these quantities, we give an empirical result on the relationship between the bid-ask liquidity balance and trade sign and we show that liquidity balance on best bid/best ask is qu…”
  27. methodمدعوم

    Passage 664 states that the study uses a visible depth of five levels, with L=5.

    [4] Price Jump Prediction in Limit Order Book, section 1 Description and data notation
    “In this study, for simplicity, we focus on limit order arrival events, limit order cancellation events and market order arrival events, see Figure  2. The number of visible limit order levels is chosen to be five L=5L=5. Our dataset is provided by NATIXIS via Thomson Reuter’s ‘Re…”
  28. factمدعوم

    Passage 664 states that the dataset comprises trades and limit order activities of the 40 member stocks of the CAC40 between April 1st 2011 and April 30th 2011.

    [4] Price Jump Prediction in Limit Order Book, section 1 Description and data notation
    “In this study, for simplicity, we focus on limit order arrival events, limit order cancellation events and market order arrival events, see Figure  2. The number of visible limit order levels is chosen to be five L=5L=5. Our dataset is provided by NATIXIS via Thomson Reuter’s ‘Re…”
  29. methodمدعوم

    Passage 664 states that to avoid open and close hours, the data are restricted to 09h05 to 17h25.

    [4] Price Jump Prediction in Limit Order Book, section 1 Description and data notation
    “In this study, for simplicity, we focus on limit order arrival events, limit order cancellation events and market order arrival events, see Figure  2. The number of visible limit order levels is chosen to be five L=5L=5. Our dataset is provided by NATIXIS via Thomson Reuter’s ‘Re…”
  30. factمدعوم

    Passage 664 states that every transaction and every limit order book modification are recorded in milliseconds.

    [4] Price Jump Prediction in Limit Order Book, section 1 Description and data notation
    “In this study, for simplicity, we focus on limit order arrival events, limit order cancellation events and market order arrival events, see Figure  2. The number of visible limit order levels is chosen to be five L=5L=5. Our dataset is provided by NATIXIS via Thomson Reuter’s ‘Re…”
  31. limitationمدعوم

    Passage 666 states that the study neglects stop orders and iceberg orders because they are relatively rare compared with limit order and market order events.

    [4] Price Jump Prediction in Limit Order Book, section 1 Description and data notation
    “In case of iceberg orders, the disclosed part has the same priority as a regular of limit order while the hidden part has lower priority. The hidden part will become visible as soon as the disclosed part is executed. The case that the hidden part is consumed by a market order wit…”
  32. factمدعوم

    The system in passage 669 integrates historical datasets (FI-2010, LOBSTER, BTC) with live WebSocket-based market feeds, a modular preprocessing engine, and multiple deep learning architectures for real-time limit order book forecasting.

    [5] OrderBook Microstructure Prediction, abstract S2 1d7511240190
    “High-frequency financial markets generate vast volumes of limit order book (LOB) data, demanding models capable of capturing complex spatial-temporal patterns for accurate short-term price movement prediction. This work presents a complete end-to-end system for real-time LOB fore…”
  33. methodمدعوم

    The preprocessing in passage 669 normalizes incoming order book snapshots using Z-score and Batch-Instance Normalization (BiN).

    [5] OrderBook Microstructure Prediction, abstract S2 1d7511240190
    “High-frequency financial markets generate vast volumes of limit order book (LOB) data, demanding models capable of capturing complex spatial-temporal patterns for accurate short-term price movement prediction. This work presents a complete end-to-end system for real-time LOB fore…”
  34. methodمدعوم

    The system in passage 669 extracts multi-level price and volume features and generates multi-horizon movement labels.

    [5] OrderBook Microstructure Prediction, abstract S2 1d7511240190
    “High-frequency financial markets generate vast volumes of limit order book (LOB) data, demanding models capable of capturing complex spatial-temporal patterns for accurate short-term price movement prediction. This work presents a complete end-to-end system for real-time LOB fore…”
  35. factمدعوم

    The models named in passage 669 are MLPLOB, TLOB, DeepLOB, and BinCTABL, and they are trained on preprocessed data and served through a FastAPI backend.

    [5] OrderBook Microstructure Prediction, abstract S2 1d7511240190
    “High-frequency financial markets generate vast volumes of limit order book (LOB) data, demanding models capable of capturing complex spatial-temporal patterns for accurate short-term price movement prediction. This work presents a complete end-to-end system for real-time LOB fore…”
  36. methodمدعوم

    In passage 669, live inference uses a frontend that streams real-time exchange data such as Binance via WebSockets, and the backend processes this data to produce instant predictions.

    [5] OrderBook Microstructure Prediction, abstract S2 1d7511240190
    “High-frequency financial markets generate vast volumes of limit order book (LOB) data, demanding models capable of capturing complex spatial-temporal patterns for accurate short-term price movement prediction. This work presents a complete end-to-end system for real-time LOB fore…”
  37. factمدعوم

    Passage 669 states that multiple models can be loaded concurrently but their predictions are returned independently to enable transparent comparison.

    [5] OrderBook Microstructure Prediction, abstract S2 1d7511240190
    “High-frequency financial markets generate vast volumes of limit order book (LOB) data, demanding models capable of capturing complex spatial-temporal patterns for accurate short-term price movement prediction. This work presents a complete end-to-end system for real-time LOB fore…”
  38. factمدعوم

    MLOFI is a vector quantity that measures the net flow of buy and sell orders at different price levels in a limit order book.

    [6] Multi-Level Order-Flow Imbalance in a Limit Order Book, section Multi-Level Order-Flow Imbalance in a Limit Order Book
    “Ke Xu ††thanks: Corresponding author. Email: xuke_e@hotmail.com. Affiliation: Mathematical Institute, University of Oxford, Oxford OX2 6GG, UK Martin D. Gould Affiliation: Mathematical Institute, University of Oxford, Oxford OX2 6GG, UK Sam D. Howison Affiliation: Mathematical In…”
  39. methodمدعوم

    Using data for 6 liquid stocks on Nasdaq, passage 670 fits a simple linear relationship between MLOFI and the contemporaneous change in mid-price.

    [6] Multi-Level Order-Flow Imbalance in a Limit Order Book, section Multi-Level Order-Flow Imbalance in a Limit Order Book
    “Ke Xu ††thanks: Corresponding author. Email: xuke_e@hotmail.com. Affiliation: Mathematical Institute, University of Oxford, Oxford OX2 6GG, UK Martin D. Gould Affiliation: Mathematical Institute, University of Oxford, Oxford OX2 6GG, UK Sam D. Howison Affiliation: Mathematical In…”
  40. resultمدعوم

    For all 6 stocks studied in passage 670, the out-of-sample goodness-of-fit improves with each additional price level included in the MLOFI vector.

    [6] Multi-Level Order-Flow Imbalance in a Limit Order Book, section Multi-Level Order-Flow Imbalance in a Limit Order Book
    “Ke Xu ††thanks: Corresponding author. Email: xuke_e@hotmail.com. Affiliation: Mathematical Institute, University of Oxford, Oxford OX2 6GG, UK Martin D. Gould Affiliation: Mathematical Institute, University of Oxford, Oxford OX2 6GG, UK Sam D. Howison Affiliation: Mathematical In…”

المصادر

  1. [1]
    Charles Huang, Weifeng Ge, H.W. Chou, Xin Du. Benchmark Dataset for Short-Term Market Prediction of Limit Order Book in China Markets. The Journal of Financial Data Science, 2021.semanticscholar · primary · DOI 10.3905/jfds.2021.1.074 · https://doi.org/10.3905/jfds.2021.1.074
  2. [2]
    Alexander Lipton, Umberto Pesavento, Michael G Sotiropoulos. Trade arrival dynamics and quote imbalance in a limit order book. arXiv, 2013.arxiv · primary · https://arxiv.org/abs/1312.0514v1
  3. [3]
    Alex Langnau, Yanko Punchev. Stochastic Price Dynamics Implied By the Limit Order Book. arXiv, 2011.arxiv · primary · https://arxiv.org/abs/1105.4789v1
  4. [4]
    Ban Zheng, Eric Moulines, Frédéric Abergel. Price Jump Prediction in Limit Order Book. arXiv, 2012.arxiv · primary · https://arxiv.org/abs/1204.1381v1
  5. [5]
    Renu Kachoria, Archit Bagad, Atharva Bondarde, Atharva Joshi, Arnav Jadhav, A. Deshmukh. OrderBook Microstructure Prediction. 2026 International Conference on System, Computation, Automation and Networking (ICSCAN), 2026.semanticscholar · primary · DOI 10.1109/ICSCAN66520.2026.11588413 · https://doi.org/10.1109/ICSCAN66520.2026.11588413
  6. [6]
    Ke Xu, Martin D. Gould, Sam D. Howison. Multi-Level Order-Flow Imbalance in a Limit Order Book. arXiv, 2019.arxiv · primary · https://arxiv.org/abs/1907.06230v2
  7. [7]
    Hamidreza Bandealinaeini, Mohammad Sharifkhani, E. Salavati. Attention-Based Multi-Asset Order Flow Networks for Enhanced Mid-Price Prediction. International Conference on AI in Finance, 2025.semanticscholar · primary · DOI 10.1145/3768292.3770430 · https://doi.org/10.1145/3768292.3770430
  8. [8]
    Prakul Sunil Hiremath, Vruksha Arun Hiremath. Early Detection of Latent Microstructure Regimes in Limit Order Books. arXiv, 2026.arxiv · primary · https://arxiv.org/abs/2604.20949v1
  9. [9]
    Faisal I Qureshi. Investigating Limit Order Book Characteristics for Short Term Price Prediction: a Machine Learning Approach. arXiv, 2018.arxiv · primary · https://arxiv.org/abs/1901.10534v1
  10. [10]
    Fan Fang, Waichung Chung, Carmine Ventre, Michail Basios, Leslie Kanthan, Lingbo Li. Ascertaining price formation in cryptocurrency markets with machine learning. European Journal of Finance, 2021.openalex · primary · DOI 10.1080/1351847x.2021.1908390 · https://doi.org/10.1080/1351847x.2021.1908390
  11. [11]
    Randolph James Ferlic, Kimberly Kate Ferlic. A One-Byte, Options-Free Market-State Monitor: Detection-Preserving Compression of Financial Data Streams with a Class-Discriminant Token. Zenodo (CERN European Organization for Nuclear Research), 2026.openalex · primary · DOI 10.5281/zenodo.22101085 · https://doi.org/10.5281/zenodo.22101085
  12. [12]
    Md Shah Ali Dolon. DEPLOYMENT AND PERFORMANCE EVALUATION OF HYBRID MACHINE LEARNING MODELS FOR STOCK PRICE FORECASTING AND RISK PREDICTION IN VOLATILE MARKETS. American Journal of Scholarly Research and Innovation, 2025.openalex · primary · DOI 10.63125/z8qq6h36 · https://doi.org/10.63125/z8qq6h36
  13. [13]
    Jianli Xiao, Baichao Long. A Multi-Channel Spatial-Temporal Transformer Model for Traffic Flow Forecasting. arXiv, 2024.arxiv · primary · https://arxiv.org/abs/2405.06266v1
  14. [14]
    Flavius Gheorghe Popa, Vlad Mureşan. Artificial Intelligence in Finance: From Market Prediction to Macroeconomic and Firm-Level Forecasting. AI, 2025.openalex · primary · DOI 10.3390/ai6110295 · https://doi.org/10.3390/ai6110295
  15. [15]
    Nino Antulov-Fantulin, Tian Guo, Fabrizio Lillo. Temporal mixture ensemble models for probabilistic forecasting of intraday cryptocurrency volume. Decisions in Economics and Finance, 2021.openalex · primary · DOI 10.1007/s10203-021-00344-9 · https://doi.org/10.1007/s10203-021-00344-9
  16. [16]
    Stefano Damato, Dario Azzimonti, Giorgio Corani. Forecasting intermittent time series with Gaussian Processes and Tweedie likelihood. arXiv, 2025.arxiv · primary · https://arxiv.org/abs/2502.19086v5