<?xml version="1.0" encoding="utf-8"?><feed xmlns="http://www.w3.org/2005/Atom" ><generator uri="https://jekyllrb.com/" version="4.4.1">Jekyll</generator><link href="https://thakicloud.github.io/feed.xml" rel="self" type="application/atom+xml" /><link href="https://thakicloud.github.io/" rel="alternate" type="text/html" /><updated>2026-07-21T17:12:59+09:00</updated><id>https://thakicloud.github.io/feed.xml</id><title type="html">Thaki Cloud Tech Blog | ThakiCloud | 다키클라우드 기술 블로그</title><subtitle>Thaki Cloud (ThakiCloud, 다키클라우드, thaki cloud, THAKI CLOUD, ثاكي كلاود)는 AI/ML Engineering, LLMOps, DevOps 분야의 최신 기술과 실무 경험을 공유하는 전문 기술 블로그입니다. 머신러닝 모델 운영, 쿠버네티스, 클라우드 인프라, AI 엔지니어링 커리어, 인공지능 기술 블로그, 다키클라우드 개발 팀의 깊이 있는 인사이트를 제공합니다. مدونة تقنية متخصصة في هندسة الذكاء الاصطناعي والحوسبة السحابية.</subtitle><author><name>{&quot;name&quot;=&gt;nil, &quot;avatar&quot;=&gt;nil, &quot;bio&quot;=&gt;nil, &quot;location&quot;=&gt;&quot;Seoul, Korea&quot;, &quot;email&quot;=&gt;&quot;info@thakicloud.co.kr&quot;, &quot;uri&quot;=&gt;nil, &quot;home&quot;=&gt;nil, &quot;links&quot;=&gt;[{&quot;label&quot;=&gt;&quot;Website&quot;, &quot;icon&quot;=&gt;&quot;fas fa-fw fa-link&quot;, &quot;url&quot;=&gt;&quot;https://thakicloud.co.kr&quot;}, {&quot;label&quot;=&gt;&quot;GitHub&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-github&quot;, &quot;url&quot;=&gt;&quot;https://github.com/thakicloud&quot;}, {&quot;label&quot;=&gt;&quot;LinkedIn&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-linkedin&quot;, &quot;url&quot;=&gt;&quot;https://www.linkedin.com/company/thakicloud&quot;}, {&quot;label&quot;=&gt;&quot;X&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-x-twitter&quot;, &quot;url&quot;=&gt;&quot;https://x.com/thakicloud&quot;}]}</name><email>info@thakicloud.co.kr</email></author><entry xml:lang="ar"><title type="html">158 مهارة و24 وكيلًا في مكوّن إضافي واحد: كيف يروّض هيكل حتمي انفجار الوكلاء</title><link href="https://thakicloud.github.io/ar/dev/agentops/agent-plugin-158-skills-deterministic-flow/" rel="alternate" type="text/html" title="158 مهارة و24 وكيلًا في مكوّن إضافي واحد: كيف يروّض هيكل حتمي انفجار الوكلاء" /><published>2026-07-21T00:00:00+09:00</published><updated>2026-07-21T00:00:00+09:00</updated><id>https://thakicloud.github.io/ar/dev/agentops/agent-plugin-158-skills-deterministic-flow</id><content type="html" xml:base="https://thakicloud.github.io/ar/dev/agentops/agent-plugin-158-skills-deterministic-flow/"><![CDATA[<p><img src="/assets/images/agent-plugin-158-skills-deterministic-flow-hero.png" alt="تصور تجريدي لوحدات مهارات عديدة تتقارب في خط أنابيب عمودي مرتّب واحد" /></p>

<h2 id="نظرة-عامة">نظرة عامة</h2>

<p>كل من بنى نظام وكلاء جادًّا يصطدم بالمفارقة نفسها. إضافة مزيد من المهارات والوكلاء يبدو أنه سيجعل النظام أذكى، لكنه غالبًا يفعل العكس. فبمجرد تجاوز بضع عشرات من المهارات، يبدأ الوكيل بالحيرة حول أي مهارة يستخدم ومتى، وبمجرد وجود عدة وكلاء، يعالجون المهمة نفسها بطرق مختلفة أو ينحرف ترتيب وصيغة المخرجات في كل تشغيل. ترتفع القدرة بينما تنخفض اتساقية النتائج.</p>

<p>يُعدّ المكوّن الإضافي مفتوح المصدر <strong>Digital Marketing Pro</strong> حالة مثيرة تتصدى لهذه المفارقة مباشرة. فهو يجمع بين 158 مهارة و24 وكيلًا متخصصًا (وثائق المستودع تذكر 25، والتغريدة الأصلية قالت 24) ويحافظ مع ذلك على اتساق إنتاج الملفات نفسها بالترتيب نفسه في كل مرة. السر ليس نموذجًا أذكى بل تدفّق استراتيجية مثبّت في 12 جزءًا، أي هيكل حتمي. يحلّل هذا المقال ليس أداة التسويق نفسها بل تصميم هندسة الوكلاء بداخلها. ما البنية التي تصمد حتى حين تنفجر المهارات عددًا، وكيف يتصل ذلك المبدأ بمنصة الوكلاء التي تبنيها ThakiCloud.</p>

<p>سبب أهمية هذه الحالة للمطوّرين واضح. فهي تُظهر، بشيفرة مفتوحة المصدر ملموسة، لماذا يفشل الأمل الساذج بأن “أنشئ الكثير من المهارات” كثيرًا في الممارسة، وما الذي يوقف ذلك الفشل.</p>

<h2 id="ما-هو-هذا-المكوّن-الإضافي">ما هو هذا المكوّن الإضافي</h2>

<p>Digital Marketing Pro مكوّن إضافي تسويقي مفتوح المصدر صادر برخصة MIT. غرضه الظاهري مساعدة الوكالات والفرق التسويقية الداخلية على إنتاج مستندات تسويقية باتساق عبر علامات تجارية عديدة. ووفقًا لوصف المستودع، يستهدف الوكالات التي تتعامل مع ما بين 50 و200 علامة تجارية للعملاء، مُمرّرًا كل علامة عبر التدفّق نفسه المكوّن من 12 جزءًا لإنتاج الملفات نفسها بالترتيب نفسه.</p>

<p>من حيث الأرقام، المكوّن كبير نسبيًا. لديه 158 مهارة و24 وكيلًا متخصصًا، وتدفّق استراتيجية من 12 جزءًا مُوسّع إلى 61 خطوة تفصيلية. وفوق ذلك يقع الاستعداد للمادة 50 من قانون الذكاء الاصطناعي الأوروبي، وميزات AEO/GEO (تحسين محرّكات الإجابة) لست منصّات بما فيها Google AI Mode، ودعم Cowork الذي يحفظ الحالة على مستوى الفريق.</p>

<p>ما يستحق الملاحظة هو هدف التثبيت. فالمكوّن ليس مقيّدًا بـ Claude Code وحده؛ إنه يُثبَّت عبر عدة أوقات تشغيل للوكلاء منها Cowork وCodex وCursor وCopilot CLI وAntigravity. بعبارة أخرى، صُمّمت حزمة واحدة من المهارات والوكلاء لتعمل عبر عدة أطر (harnesses). وهذا قرار تصميمي مهم بما يكفي لتناوله على حدة أدناه.</p>

<p>باختصار، تحت مظهر “أداة تسويق”، يحمل هذا المكوّن إجابة واحدة عن كيفية تنظيم حزمة كبيرة من المهارات والوكلاء وتنفيذها باتساق.</p>

<h2 id="هيكل-حتمي-يروّض-انفجار-المهارات">هيكل حتمي يروّض انفجار المهارات</h2>

<p>الفكرة الجوهرية لهذا المكوّن أنه لا يترك المهارات الـ158 والوكلاء الـ24 يتعاونون بحرية. بل يُجبر كل مهمة على المرور عبر تدفّق استراتيجية مثبّت في 12 جزءًا. ينتج كل جزء مخرجًا محدّدًا بترتيب محدّد، وهناك قواعد اعتماد صريحة بين الأجزاء. لا يُشغّل جزء لاحق إلا حين تكون نتيجة الجزء السابق موجودة، وتبقى أسماء ملفات النتائج وترتيبها متطابقة حتى مع تغيّر العلامة التجارية.</p>

<p>تتضح أهمية ذلك إذا تخيّلت العكس. لو اختار 24 وكيلًا بحرية المهارة التي “تبدو الأفضل” وشغّلوا بترتيب حر، لاختلف تكوين وصيغة المخرجات بين علامة وأخرى. قد تحصل علامة على تحليل المنافسين أولًا، وقد تتخطّى أخرى تلك الخطوة كليًا. وإذا كانت الوكالة تدير 200 عميل، يصبح هذا التباين سريعًا فوضى غير قابلة للتدقيق. يقلّل التدفّق المكوّن من 12 جزءًا هذه الحرية عمدًا لرفع متوسط الجودة والاتساق.</p>

<p>يوضّح المخطط أدناه بشكل مبسّط كيف يقيّد هذا الهيكل الحتمي حرية المهارات والوكلاء.</p>

<pre><code class="language-mermaid">flowchart TB
    A[طلب مهمة&lt;br/&gt;علامة X] --&gt; B[دخول التدفّق الثابت&lt;br/&gt;من 12 جزءًا]
    B --&gt; C[كل جزء: مخرج محدّد&lt;br/&gt;بترتيب محدّد]
    C --&gt; D{اختيار المهارة المناسبة للجزء&lt;br/&gt;من بين 158}
    D --&gt; E{إسناد دور&lt;br/&gt;من بين 24 وكيلًا}
    E --&gt; F[تطبيق قواعد الاعتماد&lt;br/&gt;الصريحة بين الأجزاء]
    F --&gt; G[الملفات نفسها بالترتيب نفسه&lt;br/&gt;اتساق مستقل عن العلامة]
    G --&gt; H[محفظة مستندات&lt;br/&gt;قابلة للتدقيق]
</code></pre>

<p>الدرس هنا لا علاقة له بالتسويق. طريقة حماية الجودة مع نمو المهارات والوكلاء ليست جعل النموذج أذكى بل تخفيض التصميم الحر إلى ملء هيكل مُتحقّق منه. يمتلك هيكل حتمي الصيغة والترتيب والاعتماديات، بينما يملأ النموذج المحتوى داخل ذلك الهيكل فقط. سواء كانت 158 مهارة أو 500، فما دام الهيكل يمسك درجات الحرية، تبقى النتيجة قابلة للتنبّؤ.</p>

<h2 id="ماذا-يعني-التثبيت-عبر-ست-أوقات-تشغيل">ماذا يعني التثبيت عبر ست أوقات تشغيل</h2>

<p>تصميم آخر يستحق الملاحظة هو أن هذا المكوّن يُثبَّت عبر عدة أوقات تشغيل للوكلاء. Claude Code وCursor وCodex وCopilot CLI كلٌّ منها إطار مختلف. مُطالبات النظام لديها مختلفة، وأساليب تعريف الأدوات مختلفة، ونماذج الأذونات مختلفة. وأن تُصمَّم حزمة المهارات والوكلاء نفسها لتعمل فوقها جميعًا يعني أن القدرة تراكمت في المهارات، لا في الإطار.</p>

<p>هذا التمييز مهم عمليًا. لو كانت معرفة سير عمل تسويقي محشوّة في ملفات إعداد أداة معيّنة أو مُطالبة نظامها، لعنى تبديل الأداة إعادة بناء كل شيء. وبالعكس، حين تعيش المعرفة في حزمة مهارات قابلة للنقل، يبقى الإطار رفيعًا وتُعاد المهارات عبر الأدوات. تثبيت Digital Marketing Pro عبر أوقات التشغيل ممارسة لمبدأ “إطار رفيع، مهارات سميكة” على نطاق تجاري.</p>

<p>بالطبع لدعم عدة أوقات تشغيل في آن تكلفة. فلأن كل وقت تشغيل يحمّل المهارات ويستدعيها بطريقة مختلفة قليلًا، قد يترك التصميم على القاسم المشترك ميزات فريدة لوقت تشغيل معيّن غير مستغلّة. ومع ذلك، إعطاء الأولوية لقابلية النقل توجّه معقول يحرّر أصول المهارات من الارتباط بأداة ويجعلها تصمد أطول.</p>

<h2 id="الأثر-على-منتجات-thakicloud">الأثر على منتجات ThakiCloud</h2>

<p>ما يجعل هذه الحالة مثيرة أنها تعالج مشكلة تشبه بشكل لافت ما تبنيه ThakiCloud بـ<strong>Paxis</strong>. Paxis هي السحابة الأصلية للوكلاء من ThakiCloud، وتتعامل مع المهارات والأدوات والسياسات وسجلات التدقيق كموارد من الدرجة الأولى. يختار مسخّر المهارات المهارة المناسبة من بين أكثر من 960 مهارة عبر BM25، ويشغّلها في صندوق رمل معزول، ويمرّر كل إجراء عبر بوابات السياسة وسجلات التدقيق.</p>

<p>المشكلة نفسها التي حلّها Digital Marketing Pro بترويض 158 مهارة عبر تدفّق من 12 جزءًا، تحلّها Paxis على نطاق أكبر. فبمجرد تجاوز المهارات 960، يصل سؤال “أي مهارة ومتى” إلى نطاق لا يستطيع إنسان تحديده يدويًا، فيحلّ اختيار المهارات المعتمد على BM25 محلّ ذلك الهيكل. فبدلًا من استدعاء أي مهارة بحرية، تُطرح فقط المهارات الأوثق صلة بالطلب كمرشّحين، ما يقلّل درجات الحرية. هذا المبدأ نفسه الذي منع به التدفّق من 12 جزءًا الترتيب الحر، لكن بدل تدفّق ثابت يتحكم في الحرية عبر اختيار قائم على الاسترجاع.</p>

<p>كذلك، تأكيد المكوّن على الاستعداد للمادة 50 من قانون الذكاء الاصطناعي الأوروبي وإنتاج مستندات قابلة للتدقيق يتوافق مع تعامل Paxis مع سجلات التدقيق وبوابات السياسة كموارد من الدرجة الأولى. ففي بيئات العملاء حيث تهمّ التنظيمات والتدقيق، يجب أن تكون قادرًا على تتبّع “ما الذي أُنتج، وبأي ترتيب، وعلى أي أساس.” التدفّق الحتمي وسجلات التدقيق هما المحوران اللذان يصنعان هذه القابلية للتتبّع، وتوفّرهما Paxis على مستوى المنصة. فمهما كدّست من مهارات، ولأن بوابات السياسة وسجلات التدقيق تسجّل كل إجراء، يمكن تشغيل أصل مهارات كبير بأمان حتى في بيئات منظّمة.</p>

<p>أخيرًا، قابلية النقل عبر أوقات التشغيل تتوافق مع الاتجاه الذي تسعى إليه ThakiCloud. فتصميم يعيد استخدام أصل مهارات عبر الأطر بدل ربطه بأداة معيّنة هو السبب نفسه لتعامل Paxis مع المهارات كموارد من الدرجة الأولى. حين تتراكم القدرة في المهارات لا في الإطار، تبقى الأصول التي بنيتها حتى مع تغيّر الأداة.</p>

<h2 id="القيود-والاعتراضات">القيود والاعتراضات</h2>

<p>من المهم عدم المبالغة في قراءة هذه الحالة. التدفّق الثابت من 12 جزءًا يضحّي بالمرونة مقابل الاتساق. فالحاجة الاستثنائية التي تخرج عن التدفّق القياسي، مثل مهمة غير مهيكلة مطلوبة لعلامة معيّنة فقط، قد تُعالج بشكل محرج داخل هذا الهيكل أو لا تُعالج أصلًا. الهيكل الحتمي قوي للعمل المجمّع القابل للتكرار، لكنه يصبح قيدًا للعمل ذي الاستثناءات الإبداعية الكثيرة.</p>

<p>رقم 158 مهارة نفسه يستحق قراءة متأنّية. فكثرة المهارات تعني كثرة أهداف الصيانة، وما إذا كانت كل مهارة متحقَّقًا منها فعلًا ومُحدّثة مسألة منفصلة. الرقم لا يضمن الجودة. وكم عدد المهارات الأساسية التي يستدعيها التدفّق فعلًا، وكم مرة تُستخدم البقية، أمر يصعب تأكيده من وثائق المستودع وحدها [تقدير].</p>

<p>كذلك، يحلّل هذا المقال مبادئ تصميم المكوّن، لا الجودة الفعلية لمخرجاته التسويقية. فإنتاج تدفّق حتمي لمستندات متسقة مسألة مختلفة عمّا إذا كانت تلك المستندات تؤدّي إلى نتائج تسويقية حقيقية. ما نأخذه من هذه الحالة ليس النتيجة التسويقية بل النمط الهندسي في ترويض حزمة كبيرة من المهارات والوكلاء بهيكل حتمي.</p>

<h2 id="المصادر">المصادر</h2>

<ul>
  <li>المستودع: <a href="https://github.com/indranilbanerjee/digital-marketing-pro">github.com/indranilbanerjee/digital-marketing-pro</a></li>
  <li>المصدر الأصلي: <a href="https://x.com/hjguyhan/status/2079315207579660557">تغريدة @tom_doerr</a></li>
</ul>]]></content><author><name>{&quot;name&quot;=&gt;nil, &quot;avatar&quot;=&gt;nil, &quot;bio&quot;=&gt;nil, &quot;location&quot;=&gt;&quot;Seoul, Korea&quot;, &quot;email&quot;=&gt;&quot;info@thakicloud.co.kr&quot;, &quot;uri&quot;=&gt;nil, &quot;home&quot;=&gt;nil, &quot;links&quot;=&gt;[{&quot;label&quot;=&gt;&quot;Website&quot;, &quot;icon&quot;=&gt;&quot;fas fa-fw fa-link&quot;, &quot;url&quot;=&gt;&quot;https://thakicloud.co.kr&quot;}, {&quot;label&quot;=&gt;&quot;GitHub&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-github&quot;, &quot;url&quot;=&gt;&quot;https://github.com/thakicloud&quot;}, {&quot;label&quot;=&gt;&quot;LinkedIn&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-linkedin&quot;, &quot;url&quot;=&gt;&quot;https://www.linkedin.com/company/thakicloud&quot;}, {&quot;label&quot;=&gt;&quot;X&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-x-twitter&quot;, &quot;url&quot;=&gt;&quot;https://x.com/thakicloud&quot;}]}</name><email>info@thakicloud.co.kr</email></author><category term="dev" /><category term="agentops" /><category term="AgentOps" /><category term="Skills" /><category term="MultiAgent" /><category term="ClaudeCode" /><category term="Plugins" /><category term="Determinism" /><category term="Paxis" /><category term="AIAgents" /><summary type="html"><![CDATA[يجمع المكوّن الإضافي مفتوح المصدر Digital Marketing Pro بين 158 مهارة و24 وكيلًا متخصصًا دون أن ينهار. السر هيكل حتمي: تدفّق ثابت من 12 جزءًا. نحلّل التصميم ونبيّن كيف تحوّل Paxis من ThakiCloud المبدأ نفسه إلى منتج.]]></summary></entry><entry xml:lang="ar"><title type="html">وضع قارئ الشاشة في Claude Code: سطر واحد يفتح البرمجة الطرفية بالذكاء الاصطناعي للجميع</title><link href="https://thakicloud.github.io/ar/dev/claude-code-screen-reader-accessibility/" rel="alternate" type="text/html" title="وضع قارئ الشاشة في Claude Code: سطر واحد يفتح البرمجة الطرفية بالذكاء الاصطناعي للجميع" /><published>2026-07-21T00:00:00+09:00</published><updated>2026-07-21T00:00:00+09:00</updated><id>https://thakicloud.github.io/ar/dev/claude-code-screen-reader-accessibility</id><content type="html" xml:base="https://thakicloud.github.io/ar/dev/claude-code-screen-reader-accessibility/"><![CDATA[<p><img src="/assets/images/claude-code-screen-reader-accessibility-hero.png" alt="تصور تجريدي لطرفية أُعيد تنظيمها إلى تدفق خطي نظيف من النص" /></p>

<h2 id="نظرة-عامة">نظرة عامة</h2>

<p>تطورت أدوات البرمجة الطرفية المعتمدة على الذكاء الاصطناعي في معظمها نحو ملء الشاشة بشكل جميل: مؤشرات دوران حية، فروقات ملوّنة، نوافذ أذونات محاطة بإطارات، ومؤشرات تقدّم تُعاد رسمتها مع تحرك المؤشر. بالنسبة للمستخدمين المبصرين، تُعدّ هذه الكثافة البصرية ميزة. أما بالنسبة للمطوّر الذي يقرأ الطرفية بقارئ شاشة بدلًا من عينيه فإنها تعمل بالعكس. الشاشة التي تُعاد رسمتها باستمرار تجعل من الصعب على قارئ الشاشة أن يقرر ما هو الجديد فعلًا، وتُقرأ الإطارات والحركات كضجيج بلا ترتيب.</p>

<p>يتصدى Claude Code الآن لهذه المشكلة مباشرة بوضع قارئ شاشة. سطر واحد، <code class="language-plaintext highlighter-rouge">claude --ax-screen-reader</code>، يحوّل واجهة الطرفية البصرية إلى نص خطي بسيط. فبدلًا من العرض المزخرف، يطبع أسطرًا موسومة بالترتيب حتى تستطيع قارئات الشاشة مثل VoiceOver وNVDA وJAWS القراءة من الأعلى إلى الأسفل بشكل طبيعي. يستعرض هذا المقال بالضبط ما يغيّره الوضع، وكيف يعمل، ولماذا تُعدّ إمكانية الوصول لواجهات الوكلاء مشكلة يجب أن تتبنّاها منظومة التطوير بأكملها الآن.</p>

<p>يبدو الأمر علمًا صغيرًا، لكن التغيير يوسّع الإجابة عن سؤال حقيقي: من يستطيع فعلًا استخدام وكيل ذكاء اصطناعي طرفي؟ إنه موضوع تصطدم به ThakiCloud باستمرار أثناء بناء سحابة أصلية للوكلاء، لذا نتناوله ليس كملاحظة عن ميزة فحسب، بل من منظور تصميم الواجهة.</p>

<h2 id="ما-هو-وضع-قارئ-الشاشة">ما هو وضع قارئ الشاشة</h2>

<p>تتعامل جلسة Claude Code العادية مع الطرفية كأنها لوحة رسم. تحرّك المؤشر، وتمسح الأسطر التي طبعتها بالفعل وتعيد رسمها، وتُظهر التقدّم كحركة حية. هذا مثالي لمن يمسح الشاشة بعينيه، لكنه أسوأ مُدخل ممكن لقارئ الشاشة. على قارئ الشاشة أن يقرر ما يقرأه في كل مرة تتغير فيها ذاكرة العرض، وحين تُعاد رسم الشاشة في كل إطار فإنه يميل إلى تكرار المحتوى نفسه أو إغفال المخرجات الجديدة المهمة تمامًا.</p>

<p>يغيّر وضع قارئ الشاشة نموذج العرض نفسه. فبدلًا من إعادة رسم الشاشة، يلحق المعلومات الجديدة كأسطر مفردة موسومة بالترتيب. عند تشغيل أداة مثلًا، تأتي علامات صريحة مثل طلب إذن، وإشعار بتشغيل الأداة، ونتيجة، كنص. يقرأ قارئ الشاشة هذا النص الخطي من الأعلى إلى الأسفل، فتكتمل متابعة المحادثة كاملة والموافقة على أذونات الأدوات ومراجعة المخرجات بالصوت وحده.</p>

<p>يوضّح المخطط أدناه بشكل مبسّط كيف يتفرّع مساران للعرض.</p>

<pre><code class="language-mermaid">flowchart TB
    A[بدء جلسة Claude Code] --&gt; B{هل وضع قارئ الشاشة&lt;br/&gt;مُفعّل؟}
    B --&gt;|الوضع العادي| C[إعادة رسم اللوحة&lt;br/&gt;حركة المؤشر والعرض]
    B --&gt;|--ax-screen-reader| D[مخرجات نص خطي&lt;br/&gt;إلحاق أسطر موسومة]
    C --&gt; E[معلومات عالية الكثافة&lt;br/&gt;للمستخدمين المبصرين]
    D --&gt; F[قارئ الشاشة يقرأ&lt;br/&gt;من الأعلى للأسفل]
    F --&gt; G[محادثة وموافقات ومراجعة&lt;br/&gt;تكتمل بالصوت]
    D --&gt; H[جرس الطرفية&lt;br/&gt;عند الحاجة للانتباه]
</code></pre>

<p>الفكرة ليست “إعطاء معلومات أقل” بل “إعطاء المعلومات نفسها كنص مرتّب.” فبدلًا من تجريد المعنى، يزيل الزخرفة البصرية ويوفّر تدفّق مخرجات رتيبًا ومتوقّعًا يمكن لقارئ الشاشة الوثوق به.</p>

<h2 id="كيفية-التفعيل-وطريقة-العمل">كيفية التفعيل وطريقة العمل</h2>

<p>هناك طريقتان لتشغيل وضع قارئ الشاشة. لتفعيله لجلسة واحدة، مرّر العلم عند الإطلاق.</p>

<div class="language-bash highlighter-rouge"><div class="highlight"><pre class="highlight"><code>claude <span class="nt">--ax-screen-reader</span>
</code></pre></div></div>

<p>هذا العلم موجود فعلًا في Claude Code المثبّت. يُظهره فحص مخرجات المساعدة:</p>

<div class="language-bash highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="nv">$ </span>claude <span class="nt">--help</span> | <span class="nb">grep </span>ax-screen
  <span class="nt">--ax-screen-reader</span>                    Render screen-reader friendly output
</code></pre></div></div>

<p>لتطبيقه افتراضيًا على كل جلسة تُبدأ من الصدفة، اضبط متغيّر البيئة.</p>

<div class="language-bash highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="nb">export </span><span class="nv">CLAUDE_AX_SCREEN_READER</span><span class="o">=</span>1
</code></pre></div></div>

<p>الآن تستخدم أي جلسة Claude Code تُفتح في تلك الصدفة مخرجات ملائمة لقارئ الشاشة دون علم منفصل. وفقًا للوثائق الرسمية، يعمل هذا الوضع على Claude Code الإصدار v2.1.181 وما بعده، وترفض الإصدارات الأقدم العلم <code class="language-plaintext highlighter-rouge">--ax-screen-reader</code> بخطأ.</p>

<p>هناك تفاصيل سلوكية مدروسة أيضًا. في وضع قارئ الشاشة، يقرع Claude Code جرس الطرفية عندما يحتاج انتباه المستخدم. وبالتحديد، يُقرع الجرس عند انتهاء أداة استغرقت أكثر من خمس ثوانٍ، للإشارة إلى انتهاء مهمة طويلة دون الحاجة للنظر إلى الشاشة. لا يستطيع مستخدم قارئ الشاشة التأكد بصريًا من وصول نتيجة بعد إطلاق أمر، لذا تُنشئ هذه الإشارة الصوتية إيقاعًا للتفاعل.</p>

<p>هناك إعداد منفصل للمستخدمين ضعاف البصر الذين يعتمدون على مكبّر شاشة.</p>

<div class="language-bash highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="nb">export </span><span class="nv">CLAUDE_CODE_ACCESSIBILITY</span><span class="o">=</span>1
</code></pre></div></div>

<p>ضبط هذا يُبقي مؤشر الطرفية الأصلي ظاهرًا. تكبّر مكبّرات الشاشة مثل Zoom في macOS الشاشة باتباع موضع المؤشر، فإذا أخفت أداة المؤشر فقد المكبّر تركيزه. يكشف هذا الإعداد عن المؤشر ليتمكّن المكبّر من تتبّع موضع المستخدم بدقة.</p>

<p>إذًا ينقسم دعم إمكانية الوصول إلى ثلاثة مسارات: مخرجات نص خطي لقارئات الشاشة، وجرس طرفية للانتباه، وإبقاء المؤشر ظاهرًا للمكبّرات. يستهدف كل منها تقنية مساعِدة مختلفة ويمكن تفعيله بشكل مستقل عبر متغيّرات البيئة.</p>

<h2 id="لماذا-يهمّ-الآن">لماذا يهمّ الآن</h2>

<p>السبب الأول لأهمية هذه الميزة هو أن وكلاء الذكاء الاصطناعي الطرفيين يصبحون بسرعة أداة أساسية للمطوّرين. تحدث قراءة الشيفرة وإصلاحها وتشغيل الأوامر ومراجعة النتائج داخل هذه الأدوات بشكل متزايد. إذا غابت إمكانية الوصول عن هذا التدفق، فلن يتمكن المطوّرون المكفوفون أو ضعاف البصر من استخدام أدوات الإنتاجية نفسها التي يستخدمها زملاؤهم. مهما كانت الأداة قادرة، فإن ضيق باب الوصول إلى تلك القدرة يجعلها بالنسبة لبعض المطوّرين وكأنها غير موجودة.</p>

<p>السبب الثاني هو أن هذه الميزة انطلقت من طلب المجتمع. رُفعت في المستودع العام قضايا تطلب دعم NVDA وJAWS، وتحوّل ذلك الطلب إلى إصدار فعلي. غالبًا ما تُؤجَّل ميزات إمكانية الوصول إلى “لاحقًا”، لذا فإن حالة رفع طلب مستخدم للأولوية مرجع جيد. إمكانية الوصول ليست حاجة خاصة لفئة ضيقة؛ إنها محور تصميمي يحدّد نطاق الأشخاص الذين يمكنهم استخدام الأداة.</p>

<p>السبب الثالث أن هذا النهج يؤكد حقيقة قديمة: النص الخطي واجهة متينة. فتدفّق النص المرتّب والموسوم والمتوقّع ليس جيدًا لقارئات الشاشة فحسب. إنه سهل التسجيل، وسهل التمرير عبر الأنابيب، وسهل التحليل للأتمتة. ليس من قبيل الصدفة أن يكون وضع مخرجات بُني لإمكانية الوصول مواتيًا أيضًا للبرمجة النصية والتدقيق.</p>

<h2 id="الأثر-على-منتجات-thakicloud">الأثر على منتجات ThakiCloud</h2>

<p>تُشغّل ThakiCloud سحابة أصلية للوكلاء تُسمى <strong>Paxis</strong>. تتعامل Paxis مع المهارات والأدوات والسياسات وسجلات التدقيق كموارد من الدرجة الأولى: يختار مسخّر المهارات المهارة المناسبة من بين العديد ويشغّلها في صندوق رمل معزول، ممرّرًا كل إجراء عبر بوابات السياسة وسجلات التدقيق. وكلما اتسع السطح الذي يتفاعل فيه الوكيل مع الناس، أصبح سؤال ما إذا كان ذلك السطح “متاحًا للجميع” محورًا تصميميًا أساسيًا لا إضافةً.</p>

<p>الدرس من وضع قارئ الشاشة في Claude Code واضح. إمكانية الوصول لواجهة وكيل، بمعزل عن جعل الشاشة جميلة، تعتمد على قدرتك على تقديم المعلومات نفسها كنص خطي موسوم. منصة مثل Paxis تتعامل بالفعل مع سجلات التدقيق وبوابات السياسة كموارد من الدرجة الأولى في موقع بنيوي جيد هنا. فلأن كل إجراء للوكيل مُسجّل بالفعل كحدث موسوم، فإن إعادة تشكيل تدفّق الأحداث ذلك إلى مخرجات خطية يقرأها البشر ليس بناءً لخط عرض جديد كليًا بقدر ما هو إبراز لسجلات مُهيكلة تملكها أصلًا.</p>

<p>تُظهر هذه الحالة أيضًا أن المخرجات القابلة للوصول والمخرجات الملائمة للأتمتة تنبعان من الجذر نفسه. النص الذي يستطيع قارئ الشاشة قراءته هو أيضًا نص يستطيع جامع السجلات تحليله ومسار التدقيق حفظه. وبالنظر إلى مدى تأكيد ThakiCloud على قابلية الملاحظة والتدقيق في منصة الوكلاء لديها، فإن تصميم واجهة خطية قابلة للوصول إلى جانبهما يحقّق الهدفين معًا. فبدلًا من اعتبار الواجهة الغنية والنص القابل للوصول نقيضين، يعرض النهج الأفضل كلا التمثيلين على أساس مشترك هو تدفّق أحداث مُهيكل.</p>

<h2 id="القيود-والاعتراضات">القيود والاعتراضات</h2>

<p>من المهم عدم المبالغة في تقدير هذه الميزة. وضع قارئ الشاشة خط بداية لإمكانية الوصول لا خط نهاية. فإخراج نص خطي لا يجعل كل تفاعل مريحًا تلقائيًا، وفهم كتلة شيفرة طويلة أو فرق معقّد بالصوت وحده يظل مهمة مرهقة إدراكيًا. ويظل استيعاب السياق الكامل لإعادة هيكلة كبيرة دون شاشة صعبًا حتى مع هذا الوضع.</p>

<p>تختلف أيضًا إشارة الانتباه المعتمدة على جرس الطرفية باختلاف البيئة. فبعض محاكيات الطرفية مضبوطة لتحويل الجرس إلى ومضة بصرية أو لإسكاته تمامًا، فقد لا تصل إشارة الجرس كما يُقصد. يحتاج المستخدمون إلى ضبط إعدادات طرفياتهم للحصول على أفضل تجربة.</p>

<p>أخيرًا، وجود وضع لإمكانية الوصول يختلف عن التحقّق منه جيدًا في الممارسة. يحتاج مطوّرون مكفوفون فعليون إلى استخدامه عبر قارئات شاشة وسير عمل متنوعة على مدى طويل، مراكمين ملاحظات قبل أن تظهر الحواف الخشنة وتُصقل. وبالنظر إلى أن هذا الوضع يعمل أول مرة في v2.1.181، فإنه لا يزال في بدايته، مع مجال واسع للتحسين. ومع ذلك، فإن تضمين ميزة كهذه في التوزيع الافتراضي هو بحد ذاته إشارة ذات معنى إلى توجّه للتعامل مع إمكانية الوصول الآن لا لاحقًا.</p>

<h2 id="المصادر">المصادر</h2>

<ul>
  <li>وثائق إمكانية الوصول في Claude Code: <a href="https://code.claude.com/docs/en/accessibility">code.claude.com/docs/en/accessibility</a></li>
  <li>قضية طلب الميزة (NVDA/JAWS): <a href="https://github.com/anthropics/claude-code/issues/11002">anthropics/claude-code #11002</a></li>
  <li>المصدر الأصلي: <a href="https://x.com/hjguyhan/status/2079435394727416168">تغريدة @ClaudeDevs</a></li>
</ul>]]></content><author><name>{&quot;name&quot;=&gt;nil, &quot;avatar&quot;=&gt;nil, &quot;bio&quot;=&gt;nil, &quot;location&quot;=&gt;&quot;Seoul, Korea&quot;, &quot;email&quot;=&gt;&quot;info@thakicloud.co.kr&quot;, &quot;uri&quot;=&gt;nil, &quot;home&quot;=&gt;nil, &quot;links&quot;=&gt;[{&quot;label&quot;=&gt;&quot;Website&quot;, &quot;icon&quot;=&gt;&quot;fas fa-fw fa-link&quot;, &quot;url&quot;=&gt;&quot;https://thakicloud.co.kr&quot;}, {&quot;label&quot;=&gt;&quot;GitHub&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-github&quot;, &quot;url&quot;=&gt;&quot;https://github.com/thakicloud&quot;}, {&quot;label&quot;=&gt;&quot;LinkedIn&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-linkedin&quot;, &quot;url&quot;=&gt;&quot;https://www.linkedin.com/company/thakicloud&quot;}, {&quot;label&quot;=&gt;&quot;X&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-x-twitter&quot;, &quot;url&quot;=&gt;&quot;https://x.com/thakicloud&quot;}]}</name><email>info@thakicloud.co.kr</email></author><category term="dev" /><category term="ClaudeCode" /><category term="Accessibility" /><category term="ScreenReader" /><category term="AICoding" /><category term="DeveloperProductivity" /><category term="Paxis" /><category term="InclusiveDev" /><summary type="html"><![CDATA[أضاف Claude Code وضع قارئ الشاشة الذي يستبدل واجهة الطرفية البصرية بنص خطي بسيط. إليك ما يغيّره الأمر `claude --ax-screen-reader` فعليًا، وكيف يعمل، ولماذا تهمّ إمكانية الوصول لواجهات الوكلاء منصات مثل ThakiCloud.]]></summary></entry><entry xml:lang="ar"><title type="html">استبدل واجهة OpenAI Realtime بسطر واحد: hugging-voice، حزمة صوتية مفتوحة تشغّلها بنفسك</title><link href="https://thakicloud.github.io/ar/llmops/hugging-voice-open-realtime-voice-self-hosted/" rel="alternate" type="text/html" title="استبدل واجهة OpenAI Realtime بسطر واحد: hugging-voice، حزمة صوتية مفتوحة تشغّلها بنفسك" /><published>2026-07-21T00:00:00+09:00</published><updated>2026-07-21T00:00:00+09:00</updated><id>https://thakicloud.github.io/ar/llmops/hugging-voice-open-realtime-voice-self-hosted</id><content type="html" xml:base="https://thakicloud.github.io/ar/llmops/hugging-voice-open-realtime-voice-self-hosted/"><![CDATA[<p><img src="/assets/images/hugging-voice-open-realtime-voice-self-hosted-hero.png" alt="خط معالجة صوتي لحظي مفتوح تشغّله بنفسك" /></p>

<p>كُتب هذا المقال للمهندسين الذين أرادوا إضافة وكيل صوتي لكنهم ترددوا أمام الارتباط بمزوّد واحد وتكلفة واجهة OpenAI Realtime، ولمسؤولي البنية التحتية الذين يوازنون ما إذا كان الصوت الحواري قابلاً للتشغيل على حزمتهم الخاصة. باختصار، إن تصميم العرض التجريبي <a href="https://huggingface.co/spaces/HuggingFaceM4/hugging-voice">hugging-voice</a> من Hugging Face والمكتبة التي تعمل تحته، <a href="https://github.com/huggingface/speech-to-speech">speech-to-speech</a>، بسيط وعملي في آن معاً. فهو يفتح خط المعالجة الصوتي اللحظي بمراحله الأربع كمصدر مفتوح، بينما يغلّف الطرف الخارجي بالواجهة نفسها التي يقدمها OpenAI Realtime. لذا إن كان لديك بالفعل كود مكتوب لعميل OpenAI اللحظي، فيمكنك الانتقال إلى حزمتك الخاصة بتغيير سطر واحد فقط: العنوان الذي يشير إليه الخادم. لا نستشهد بأرقام الأداء إلا ضمن النطاق الذي نشره المشروع، ونوضّح مسبقاً أنها ليست أرقاماً قسناها بأنفسنا.</p>

<h2 id="نظرة-عامة">نظرة عامة</h2>

<p>خلال العام الماضي، لم يعد الصوت الحواري ميزة جانبية لروبوتات الدردشة النصية، بل صار فئة منتجات قائمة بذاتها. يتكلم المستخدم، ويتوقع أن يعود الرد فوراً، بالطريقة التي يتدفق بها حوار بشري. المشكلة أن المسار التجاري لتلبية هذا التوقع تقارب فعلياً نحو حفنة من الخدمات المغلقة مثل واجهة OpenAI Realtime. مريحة، نعم، لكن حركة الصوت تُحاسَب عادةً بالثانية لا بالرمز، وتخرج البيانات من نطاقك، ويرتبط كل من النموذج والأصوات بالمزوّد.</p>

<p>يأتي hugging-voice كمثال معاكس لهذا الاتجاه. عنوانه الفرعي يقولها مباشرة: “صوت لحظي مفتوح يمكنك فعلاً تشغيله بنفسك”. الفكرة الجوهرية هي أن الرحلة كاملة، أي تحويل الصوت الوارد إلى الميكروفون إلى نص، وإرساله إلى نموذج لغوي، وإعادة الرد صوتاً، مفتوحة كخط معالجة يمكن استبدال كل مكوّن فيه. بالنسبة لنا نحن الذين نخدم النماذج في بيئات محلية وسيادية، يعني هذا أنه صار هناك تطبيق مرجعي ملموس لسؤال “هل يمكن تشغيل الصوت اللحظي على مجموعتنا الخاصة؟”</p>

<h2 id="ما-هو-hugging-voice-وما-هو-speech-to-speech">ما هو hugging-voice وما هو speech-to-speech</h2>

<p>لنحدد المصطلحات أولاً: hugging-voice هو العرض التجريبي (Space) الذي يمكنك التحدث إليه مباشرة في المتصفح، والمحرك الذي يعالج الصوت داخله فعلياً هو مكتبة speech-to-speech. تقسّم المكتبة الوكيل الصوتي اللحظي إلى أربع مراحل: كشف النشاط الصوتي (VAD)، وتحويل الكلام إلى نص (STT)، ونموذج لغوي (LLM)، وتحويل النص إلى كلام (TTS). تعمل كل مرحلة في خيط منفصل وتُوصل بطوابير، بحيث يتدفق خرج مرحلة إلى المرحلة التالية بثاً حياً. يظهر نص جزئي قبل أن ينهي المستخدم كلامه، ويبدأ توليد الكلام على الكلمات الأولى قبل أن يكمل النموذج الجملة، ما يقلّص زمن الاستجابة المُدرَك.</p>

<div class="mermaid">
flowchart TB
    A["دخل الميكروفون<br />تدفق صوتي لحظي"] --&gt; B["كشف الكلام VAD<br />Silero VAD v5"]
    B --&gt; C["التعرّف على الكلام STT<br />Parakeet TDT · Whisper"]
    C --&gt; D["توليد الرد LLM<br />OpenAI-compatible API · vLLM · llama.cpp"]
    D --&gt; E["تركيب الكلام TTS<br />Qwen3-TTS · Kokoro"]
    E --&gt; F["خرج السماعة<br />تشغيل ببثّ حي"]
    G["خادم WebSocket<br />متوافق مع OpenAI Realtime"] -.- B
    G -.- C
    G -.- D
    G -.- E
</div>

<p>في هذا المخطط، خادم WebSocket على اليمين هو سلاح المشروع الحقيقي. مجرد لصق أربع مراحل معاً ليس جديداً. ما يميّز speech-to-speech هو أنه يغلّف خط المعالجة بأكمله بنقطة نهاية WebSocket متوافقة مع بروتوكول OpenAI Realtime. هذا يتيح لعميل OpenAI لحظي قائم أن يتصل بهذا الخادم كأنه OpenAI نفسه. ويجدر بالذكر أن هذه الحزمة ليست لعبة تجريبية: فهي تشغّل البنية التحتية الصوتية اللحظية لروبوتات Reachy Mini من Hugging Face في الإنتاج.</p>

<h2 id="الانتقال-بسطر-واحد">الانتقال بسطر واحد</h2>

<p>هذا هو الجزء الذي يسمّيه المشروع نفسه “الترحيل بسطر واحد”. الشيء الوحيد الذي على عميل كان يستخدم OpenAI Realtime تغييره هو عنوان الاتصال. في ما يلي مثال عميل Python الذي توثّقه وثائق المشروع.</p>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="kn">from</span> <span class="n">openai</span> <span class="kn">import</span> <span class="n">OpenAI</span>

<span class="n">client</span> <span class="o">=</span> <span class="nc">OpenAI</span><span class="p">(</span>
    <span class="n">base_url</span><span class="o">=</span><span class="sh">"</span><span class="s">http://localhost:8765/v1</span><span class="sh">"</span><span class="p">,</span>
    <span class="n">websocket_base_url</span><span class="o">=</span><span class="sh">"</span><span class="s">ws://localhost:8765/v1</span><span class="sh">"</span><span class="p">,</span>
    <span class="n">api_key</span><span class="o">=</span><span class="sh">"</span><span class="s">not-needed</span><span class="sh">"</span><span class="p">,</span>
<span class="p">)</span>

<span class="k">with</span> <span class="n">client</span><span class="p">.</span><span class="n">realtime</span><span class="p">.</span><span class="nf">connect</span><span class="p">(</span><span class="n">model</span><span class="o">=</span><span class="sh">"</span><span class="s">local</span><span class="sh">"</span><span class="p">)</span> <span class="k">as</span> <span class="n">conn</span><span class="p">:</span>
    <span class="n">conn</span><span class="p">.</span><span class="nf">send</span><span class="p">({</span>
        <span class="sh">"</span><span class="s">type</span><span class="sh">"</span><span class="p">:</span> <span class="sh">"</span><span class="s">session.update</span><span class="sh">"</span><span class="p">,</span>
        <span class="sh">"</span><span class="s">session</span><span class="sh">"</span><span class="p">:</span> <span class="p">{</span>
            <span class="sh">"</span><span class="s">type</span><span class="sh">"</span><span class="p">:</span> <span class="sh">"</span><span class="s">realtime</span><span class="sh">"</span><span class="p">,</span>
            <span class="sh">"</span><span class="s">instructions</span><span class="sh">"</span><span class="p">:</span> <span class="sh">"</span><span class="s">You are a helpful assistant.</span><span class="sh">"</span><span class="p">,</span>
            <span class="sh">"</span><span class="s">audio</span><span class="sh">"</span><span class="p">:</span> <span class="p">{</span>
                <span class="sh">"</span><span class="s">input</span><span class="sh">"</span><span class="p">:</span> <span class="p">{</span>
                    <span class="sh">"</span><span class="s">turn_detection</span><span class="sh">"</span><span class="p">:</span> <span class="p">{</span>
                        <span class="sh">"</span><span class="s">type</span><span class="sh">"</span><span class="p">:</span> <span class="sh">"</span><span class="s">server_vad</span><span class="sh">"</span><span class="p">,</span>
                        <span class="sh">"</span><span class="s">interrupt_response</span><span class="sh">"</span><span class="p">:</span> <span class="bp">True</span><span class="p">,</span>
                    <span class="p">}</span>
                <span class="p">}</span>
            <span class="p">},</span>
        <span class="p">}</span>
    <span class="p">})</span>

    <span class="k">for</span> <span class="n">event</span> <span class="ow">in</span> <span class="n">conn</span><span class="p">:</span>
        <span class="nf">print</span><span class="p">(</span><span class="n">event</span><span class="p">.</span><span class="nb">type</span><span class="p">)</span>
</code></pre></div></div>

<p>ما يستحق الانتباه هو أن <code class="language-plaintext highlighter-rouge">base_url</code> و<code class="language-plaintext highlighter-rouge">websocket_base_url</code> يشيران إلى خادم محلي، وأن <code class="language-plaintext highlighter-rouge">api_key</code> غير مطلوب فعلياً. أما التعليمات المُمرَّرة عبر <code class="language-plaintext highlighter-rouge">session.update</code>، وكشف الأدوار المعتمد على VAD في جهة الخادم، ومقاطعة الرد أثناء بثّه، فكلها تتبع المخطط نفسه الذي يتبعه OpenAI Realtime. بعبارة أخرى، يكاد كود التطبيق لا يتغير، ولا ينتقل سوى الخلفية من واجهة خارجية إلى خادمك الخاص. وللفرق التي تقلق من الارتباط بمزوّد واحد، فإن توافق الواجهة هذا يمثّل بمفرده أكبر قيمة عملية.</p>

<h2 id="التثبيت-والتشغيل">التثبيت والتشغيل</h2>

<p>مسار إقامة الخادم مختصر بالقدر نفسه. يغطي التثبيت الافتراضي المسار اللحظي القياسي دفعة واحدة.</p>

<div class="language-bash highlighter-rouge"><div class="highlight"><pre class="highlight"><code>pip <span class="nb">install </span>speech-to-speech
</code></pre></div></div>

<p>يستخدم الإعداد الافتراضي Parakeet TDT لـ STT، وواجهة متوافقة مع OpenAI لـ LLM، وQwen3-TTS لـ TTS. وإن احتجت خلفية محددة، فثبّتها بالإضافات.</p>

<div class="language-bash highlighter-rouge"><div class="highlight"><pre class="highlight"><code>pip <span class="nb">install</span> <span class="s2">"speech-to-speech[kokoro]"</span>
pip <span class="nb">install</span> <span class="s2">"speech-to-speech[faster-whisper]"</span>
</code></pre></div></div>

<p>يبدو تشغيل الخادم هكذا. يطلق هذا الأمر خادماً متوافقاً مع OpenAI Realtime عبر WebSocket محلي.</p>

<div class="language-bash highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="nb">export </span><span class="nv">OPENAI_API_KEY</span><span class="o">=</span>...
speech-to-speech
</code></pre></div></div>

<p>النقطة المثيرة هنا هي أن مرحلة LLM يمكن أن تعمل محلياً بالكامل. في ما يلي مثال يقيم نموذجاً من فئة Gemma 4 بواسطة llama.cpp، ويجعل speech-to-speech يشير إلى تلك النقطة المحلية.</p>

<div class="language-bash highlighter-rouge"><div class="highlight"><pre class="highlight"><code>llama-server <span class="nt">-hf</span> ggml-org/gemma-4-E4B-it-GGUF <span class="nt">-np</span> 2 <span class="nt">-c</span> 65536

speech-to-speech <span class="se">\</span>
    <span class="nt">--model_name</span> <span class="s2">"ggml-org/gemma-4-E4B-it-GGUF"</span> <span class="se">\</span>
    <span class="nt">--responses_api_base_url</span> <span class="s2">"http://127.0.0.1:8080/v1"</span> <span class="se">\</span>
    <span class="nt">--responses_api_api_key</span> <span class="s2">""</span>
</code></pre></div></div>

<p>على أجهزة Mac بمعالج Apple Silicon يمكنك تفعيل الإعدادات المحسّنة وربط نموذج mlx، وعلى العكس، إن نقص عتاد GPU المحلي، يمكنك توجيه خلفية LLM إلى Hugging Face Inference Providers.</p>

<div class="language-bash highlighter-rouge"><div class="highlight"><pre class="highlight"><code>speech-to-speech <span class="se">\</span>
    <span class="nt">--local_mac_optimal_settings</span> <span class="se">\</span>
    <span class="nt">--model_name</span> <span class="s2">"mlx-community/Qwen3-4B-Instruct-2507-bf16"</span>
</code></pre></div></div>

<p>إن القدرة على نقل خط المعالجة نفسه عبر الطيف كله، من التشغيل المحلي الكامل إلى تفويض الاستدلال السحابي، ببضعة أعلام (flags)، هي فضيلة في التصميم.</p>

<h2 id="استبدال-الوحدات-خلفيات-stt-وllm-وtts">استبدال الوحدات: خلفيات STT وLLM وTTS</h2>

<p>سبب قراءة هذا المشروع كبنية مرجعية لا مجرد عرض تجريبي هو أن خلفية كل مرحلة قابلة للاستبدال. باختصار:</p>

<table>
  <thead>
    <tr>
      <th>المرحلة</th>
      <th>الخلفية الافتراضية</th>
      <th>البدائل</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>VAD</td>
      <td>Silero VAD v5</td>
      <td>مدمجة فقط</td>
    </tr>
    <tr>
      <td>STT</td>
      <td>Parakeet TDT</td>
      <td>Whisper, Faster Whisper, Paraformer</td>
    </tr>
    <tr>
      <td>LLM</td>
      <td>واجهة متوافقة مع OpenAI</td>
      <td>Transformers, mlx-lm, vLLM, llama.cpp</td>
    </tr>
    <tr>
      <td>TTS</td>
      <td>Qwen3-TTS</td>
      <td>Kokoro, Pocket TTS, ChatTTS, MMS</td>
    </tr>
  </tbody>
</table>

<p>وهناك أيضاً أربعة أوضاع تشغيل. الوضع الافتراضي realtime هو WebSocket الذي يتكلم بروتوكول OpenAI Realtime؛ ويرتبط local مباشرة بالميكروفون والسماعة؛ ويتبادل الوضعان websocket وsocket صوت PCM الخام عبر WebSocket وTCP على التوالي. ويمكن تحديد اللغة أو تركها للكشف التلقائي. إن القدرة على مزج دقة STT، وجودة LLM وتكلفته، وطابع TTS وزمن استجابته بما يناسب متطلباتك، هي بالضبط درجة الحرية التي لا تمنحها واجهة مغلقة.</p>

<h2 id="زمن-الاستجابة-وحالة-واقعية">زمن الاستجابة، وحالة واقعية</h2>

<p>في الوكيل الصوتي، يهمّ زمن الاستجابة بقدر ما تهمّ الدقة. يشعر الناس بانقطاع الحوار إذا تأخر الرد بضع مئات من الأجزاء من الثانية فقط. يستهدف <a href="https://huggingface.co/blog/cerebras-gemma4-voice-ai">العرض التجريبي للصوت اللحظي</a> الذي نشرته Hugging Face بالتعاون مع Cerebras مشكلة زمن الاستجابة مباشرة. يستخدم الإعداد Parakeet من Nvidia لـ STT، ونموذج Gemma 4 من Google DeepMind يعمل على استدلال Cerebras للنموذج اللغوي، وQwen3-TTS من Alibaba لـ TTS. والهدف هو خفض زمن استجابة مرحلة LLM باستدلال Cerebras فائق السرعة كي يتدفق الحوار بطبيعية تضاهي التحدث إلى إنسان. غير أنه، ضمن ما استطاع هذا المقال التحقق منه، لم تُنشر أرقام محددة بالأجزاء من الثانية، لذا نتحفظ عن مقارنة كمية.</p>

<p>وهناك دليل إنتاجي أيضاً. تشغّل هذه الحزمة روبوتات Reachy Mini المذكورة آنفاً، وتذكر Hugging Face أن أكثر من 9,000 روبوت منتشرة بالفعل في الميدان. وسبب هوس المشروع بزمن الاستجابة هو أن سرعة الاستجابة في البيئات المدمجة هي ما يجعل التفاعل “يبدو حياً”. وقد تناولنا سابقاً ميزانية زمن الاستجابة للوكلاء الصوتيين من زاوية خدمة GPU، فإن كنت تفكر في كيفية توزيع زمن الاستجابة عبر كل مرحلة من خط المعالجة، نقترح قراءة <a href="/ar/llmops/voice-agent-latency-budget-gpu-serving/">ميزانية زمن استجابة الوكيل الصوتي وخدمة GPU</a> إلى جانب هذا المقال.</p>

<h2 id="دلالات-على-منتجات-thakicloud">دلالات على منتجات ThakiCloud</h2>

<p>يتشابك خط المعالجة هذا بطبيعية مع منتجينا كليهما.</p>

<p>من منظور ai-platform، فإن مرحلة LLM في speech-to-speech لا تستهلك في النهاية سوى نقطة نهاية متوافقة مع OpenAI، لذا يمكنك وضع خدمة vLLM من ThakiCloud مباشرة في ذلك المكان. ويصبح نموذجا STT وTTS حِملين منفصلين يشغلان GPU، وقوّتنا تكمن بالضبط في تحميل مثل هذه الأحمال الاستدلالية غير المتجانسة على مجموعة واحدة معاً، بجدولة GPU عبر Kueue وعزلها بين المستأجرين. وحركة الصوت تميل إلى تراكم التكلفة بالثانية، لذا تعمل تنافسية تكلفة الوحدة للخدمة الذاتية بقوة خاصة، وتلائم هذه الحزمة المفتوحة المتطلبات المحلية والسيادية حيث يجب ألا تغادر البيانات المبنى. وللعملاء الذين تمثّل لهم واجهة صوتية مغلقة عبئاً، يمكننا أن نقدّم خيار “تشغيلها على مجموعتك الخاصة عبر الواجهة نفسها”.</p>

<p>ومن منظور Paxis، الصوت قناة دخل وخرج جديدة تُلحق بالوكيل. Paxis هو مستوى تحكّم Agent-Native Cloud يعمل فوق ai-platform ويتعامل مع Skills وTools وPolicies وAudit Logs كموارد من الدرجة الأولى، وواجهة speech-to-speech المتوافقة مع OpenAI Realtime تسهّل إضافة قناة الصوت هذه إلى تنسيق الوكلاء القائم. فيمكن تكوين مسار يصبح فيه أمر منطوق دخلاً للوكيل، وتعود فيه نتيجة تنفيذ مهارة صوتاً، مع مرورها عبر بوابات السياسات وسجلات التدقيق. الخدمة الذاتية منخفضة التكلفة (ai-platform) تصنع اقتصاديات الوكلاء الصوتيين، وفوقها يتعامل مستوى التحكّم Agent-Native (Paxis) مع الصوت كقناة بأمان.</p>

<h2 id="خاتمة">خاتمة</h2>

<p>الرسالة التي يبعثها hugging-voice واضحة. لم يعد الصوت اللحظي مضطراً للاتكاء وحده على واجهات مغلقة لبضعة مزوّدين؛ وعند الحاجة، يمكنك إبقاء الواجهة كما هي ونقل الخلفية وحدها إلى بنيتك التحتية الخاصة. هذا التصميم، باختيار كل مرحلة من VAD إلى TTS وتغيير سطر واحد فقط لدى العميل، يخفض بشدة الحاجز أمام الفرق التي تدرس الخدمة الذاتية. وإن أردت رؤيته بنفسك، فتحدّث إليه مباشرة في <a href="https://huggingface.co/spaces/HuggingFaceM4/hugging-voice">العرض التجريبي (Space)</a>، أو ابدأ بـ <code class="language-plaintext highlighter-rouge">pip install speech-to-speech</code> من <a href="https://github.com/huggingface/speech-to-speech">مستودع GitHub</a>.</p>]]></content><author><name>{&quot;name&quot;=&gt;nil, &quot;avatar&quot;=&gt;nil, &quot;bio&quot;=&gt;nil, &quot;location&quot;=&gt;&quot;Seoul, Korea&quot;, &quot;email&quot;=&gt;&quot;info@thakicloud.co.kr&quot;, &quot;uri&quot;=&gt;nil, &quot;home&quot;=&gt;nil, &quot;links&quot;=&gt;[{&quot;label&quot;=&gt;&quot;Website&quot;, &quot;icon&quot;=&gt;&quot;fas fa-fw fa-link&quot;, &quot;url&quot;=&gt;&quot;https://thakicloud.co.kr&quot;}, {&quot;label&quot;=&gt;&quot;GitHub&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-github&quot;, &quot;url&quot;=&gt;&quot;https://github.com/thakicloud&quot;}, {&quot;label&quot;=&gt;&quot;LinkedIn&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-linkedin&quot;, &quot;url&quot;=&gt;&quot;https://www.linkedin.com/company/thakicloud&quot;}, {&quot;label&quot;=&gt;&quot;X&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-x-twitter&quot;, &quot;url&quot;=&gt;&quot;https://x.com/thakicloud&quot;}]}</name><email>info@thakicloud.co.kr</email></author><category term="llmops" /><category term="speech-to-speech" /><category term="realtime-voice" /><category term="VoiceAI" /><category term="OpenAIRealtime" /><category term="STT" /><category term="TTS" /><category term="LLM-serving" /><category term="LLMOps" /><category term="on-prem" /><category term="self-hosting" /><summary type="html"><![CDATA[يغلّف مشروع hugging-voice من Hugging Face ومحركه speech-to-speech خط معالجة صوتي لحظي كامل، من كشف النشاط الصوتي إلى STT وLLM وTTS، خلف واجهة WebSocket متوافقة مع OpenAI Realtime. غيّر سطراً واحداً في عنوان الخادم لدى عميلك لتنتقل إلى بنيتك التحتية الخاصة. فككنا التصميم من منظور التشغيل والخدمة.]]></summary></entry><entry xml:lang="ar"><title type="html">لم يخترق Hugging Face بشرٌ بل وكيل ذكاء اصطناعي ذاتي: عندما صار خط معالجة البيانات سطح الهجوم</title><link href="https://thakicloud.github.io/ar/news/huggingface-agentic-ai-breach/" rel="alternate" type="text/html" title="لم يخترق Hugging Face بشرٌ بل وكيل ذكاء اصطناعي ذاتي: عندما صار خط معالجة البيانات سطح الهجوم" /><published>2026-07-21T00:00:00+09:00</published><updated>2026-07-21T00:00:00+09:00</updated><id>https://thakicloud.github.io/ar/news/huggingface-agentic-ai-breach</id><content type="html" xml:base="https://thakicloud.github.io/ar/news/huggingface-agentic-ai-breach/"><![CDATA[<p><img src="/assets/images/huggingface-agentic-ai-breach-hero.png" alt="صورة تجريدية لسرب من الوكلاء الذاتيين يتسلل إلى خط بيانات" /></p>

<p>الخبر الذي هزّ التسلسلات الزمنية في نهاية الأسبوع لم يكن نموذجًا جديدًا ولا معيارًا جديدًا، بل إشعارًا بأن Hugging Face، مركز منظومة الذكاء الاصطناعي المفتوحة، قد اُخترق. وما لفت الانتباه أكثر هو من فعل ذلك. فبحسب الشركة، لم يجلس قرصان بشري ليكتب الأوامر طوال الليل، بل قاد إطار وكيل ذكاء اصطناعي ذاتي الهجوم من أوله إلى آخره.</p>

<p>إن اختراق شركة تبيع النماذج على يد نموذج يشكّل حكاية لافتة. لكن هدف هذه المقالة ليس استهلاك تلك المفارقة. فبالنسبة لشركة مثل ThakiCloud تتعامل مع النماذج والبيانات فوق بنية تحتية للعملاء، فإن العمل الحقيقي هو التمييز بهدوء بين المكان الذي دخل منه الهجوم بالضبط وما الذي تأكد. ونقطة الدخول هنا لم تكن ثغرة يوم صفري براقة، بل الشيء الذي نلمسه كل يوم: مجموعة بيانات.</p>

<h2 id="ماذا-حدث">ماذا حدث</h2>

<p>كشف Hugging Face عن الاختراق في تدوينة يوم الخميس 16 يوليو 2026. جاء ذلك بعد أن أكدت الشركة في وقت سابق من ذلك الأسبوع وصولًا غير مصرّح به إلى مجموعات بيانات وبيانات اعتماد داخلية، واحتوت التسلل. وبحسب رواية الشركة، بدأ التسلل في خط معالجة البيانات، حيث استخدم المهاجم مجموعة بيانات خبيثة واحدة لفتح مسارَي تنفيذ للتعليمات البرمجية.</p>

<p>هذا هو الهيكل المؤكد: قاده وكيل ذاتي، وكانت نقطة الدخول مجموعة بيانات، وأدّت ثغرتان إلى تنفيذ التعليمات البرمجية. أما التفاصيل المتبقية فتختلف نقاط تركيزها من منصة إلى أخرى، لذا يجب قراءة الحقائق المؤكدة بمعزل عن التقارير الثانوية.</p>

<h2 id="مسار-الهجوم-خط-معالجة-البيانات-كان-سطح-الهجوم">مسار الهجوم: خط معالجة البيانات كان سطح الهجوم</h2>

<p>الجوهر هو أسلوب الدخول. رفع المهاجم مجموعة بيانات خبيثة إلى Hugging Face Hub. وفي اللحظة التي مرّت فيها تلك المجموعة عبر خط المعالجة، انطلقت ثغرتان تباعًا. الأولى مسار محمّل بيانات بتنفيذ عن بُعد، والثانية حقن قوالب أثناء تحليل إعداد مجموعة البيانات. وكلتاهما انتهتا إلى تنفيذ تعليمات برمجية اعتباطية.</p>

<p>قد تبدو فكرة أن مجموعة بيانات يمكنها تشغيل التعليمات البرمجية غريبة، لكن الممارسين يعرفون هذا الخطر جيدًا. فكثير من محمّلات البيانات تثق في نصوص التحميل من المستودعات البعيدة وتنفّذها، وتعرض حقول الإعداد كقوالب. تلك المرونة، المصممة للراحة، تصبح قناة تنفيذ في اللحظة التي تلتقي فيها بمدخل يعبر حدّ الثقة.</p>

<p>وما تلا تأمين تنفيذ التعليمات البرمجية كان سلسلة اختراق نموذجية. رفع المهاجم صلاحياته بوصول على مستوى العقدة، وجمع بيانات اعتماد السحابة والعناقيد، وتحرك أفقيًا إلى عدة عناقيد داخلية خلال عطلة نهاية الأسبوع. كان الدخول نقطة واحدة، لكن من اللحظة التي منحت فيها تلك النقطة صلاحيات التنفيذ، انتشر الأمر تلقائيًا.</p>

<pre><code class="language-mermaid">flowchart TB
    A[المهاجم: يرفع مجموعة بيانات خبيثة] --&gt; B[خط معالجة مجموعات البيانات]
    B --&gt; C1["الثغرة 1&lt;br/&gt;remote-code dataset loader"]
    B --&gt; C2["الثغرة 2&lt;br/&gt;dataset config template injection"]
    C1 --&gt; D[تنفيذ تعليمات برمجية اعتباطية RCE]
    C2 --&gt; D
    D --&gt; E[الحصول على وصول بمستوى العقدة]
    E --&gt; F[جمع بيانات اعتماد السحابة والعناقيد]
    F --&gt; G[تحرك أفقي إلى العناقيد الداخلية]
    G --&gt; H["إطار وكيل ذاتي&lt;br/&gt;آلاف الإجراءات عبر سرب من الصناديق الرملية قصيرة العمر"]
</code></pre>

<h2 id="وزن-القول-إن-وكيلًا-ذاتيًا-قاد-الهجوم">وزن القول إن وكيلًا ذاتيًا قاد الهجوم</h2>

<p>الجزء الجديد في هذا الحادث ليس الأدوات بل مقعد القيادة. وصف Hugging Face الحملة بأنها “إطار وكيل ذاتي ينفّذ آلاف الإجراءات الفردية عبر سرب من الصناديق الرملية قصيرة العمر، مع قناة قيادة وتحكم تنتقل بنفسها فوق خدمات عامة”. فبدلًا من تدخل بشري في كل خطوة، تولّى الوكيل الاستطلاع والتنفيذ والتحرك في سلسلة متصلة.</p>

<p>المشكلة التي يطرحها هذا البناء على المدافعين هي السرعة والحجم. فالمهاجم البشري لديه حدود مادية من التعب وسرعة الكتابة، أما سرب الوكلاء فيلقي بآلاف المحاولات على التوازي وينتقل إلى التالية فور فشل خطوة. واستخدام الصناديق الرملية قصيرة العمر ثم التخلص منها يمحو مراسي الاكتشاف، وقناة القيادة والتحكم التي تنتقل عبر خدمات عامة تُبطل قوائم الحظر.</p>

<p>دار هامش مثير للاهتمام في التقارير الثانوية. فمع تطور الاستجابة، عندما حاول الفريق تسليم التحليل الجنائي إلى نماذج تجارية متقدمة (GPT، Claude)، يُقال إن حواجز الأمان اعتبرت حمولات الاستغلال وآثار القيادة والتحكم هجمات ورفضت التعاون، فواصل الفريق الاكتشاف والتحليل بنموذج من فئة GLM 5.2 [تقديري]. تأتي هذه التفصيلة من بعض المنصات لا من الإشعار الرسمي لـ Hugging Face، لذا من الأسلم عدم قراءتها كحقيقة مؤكدة. لكن بصرف النظر عن دقتها، فإن التوتر نفسه، حيث لا يستطيع المدافع استخدام أداة بسبب سياسة أمانها، جدير بالتسجيل بوصفه أمرًا قد يتكرر.</p>

<h2 id="ما-الذي-كان-آمنًا-وما-لا-يزال-قيد-التحقيق">ما الذي كان آمنًا وما لا يزال قيد التحقيق</h2>

<p>كلما كان الحادث أسهل للمبالغة، وجب رسم الحدود بوضوح أكبر. قال Hugging Face إنه أغلق مسارات تنفيذ التعليمات البرمجية المعرّضة، وطرد المهاجم، وأعاد بناء العقد المخترقة، وأبطل جميع بيانات الاعتماد المتأثرة وبدّلها. وأضاف أنه لم يجد دليلًا على العبث بالنماذج العامة أو مجموعات البيانات الموجهة للمستخدمين أو Spaces، وأن سلسلة توريد البرمجيات لديه، بما فيها صور الحاويات والحزم المنشورة، تم التحقق من نظافتها.</p>

<p>كان إجراء المستخدمين توصية احترازية. نصحت الشركة المستخدمين بتبديل رموز الوصول ومراجعة نشاط الحساب الأخير. وهنا تمييز مهم. تلك التوصية ليست تأكيدًا على تسرّب رموز المستخدمين بالجملة، بل تدبير أمان محافظ نظرًا لطبيعة حادث سُرقت فيه بيانات اعتماد داخلية. أما ما إذا كانت بيانات الشركاء أو العملاء قد تأثرت فكان، حتى وقت الكشف، لا يزال قيد التحقيق.</p>

<p>باختصار، المؤكد هو الاختراق الداخلي وسرقة بيانات الاعتماد، ووجود ثغرتين في البيانات، والاحتواء والتبديل السريعان. وما يبقى مفتوحًا هو ما إذا كانت بيانات الشركاء والعملاء قد تأثرت، وتأكيد بعض التفاصيل في التقارير الثانوية (العدد الدقيق للإجراءات، وحكاية رفض النموذج). خلط المؤكد بغير المؤكد يجعل الحادث يبدو أكبر أو أصغر مما هو عليه.</p>

<h2 id="منظور-thakicloud-التعامل-مع-معالجة-البيانات-بوصفها-حدّ-ثقة">منظور ThakiCloud: التعامل مع معالجة البيانات بوصفها حدّ ثقة</h2>

<p>الدرس الذي يقدمه هذا الحادث لشركة بنية تحتية واضح. مجموعة البيانات ليست ملفًا سلبيًا بل مدخلًا نشطًا يمكنه تنفيذ التعليمات البرمجية في اللحظة التي تُعالَج فيها. لذا ننظر إلى هذا عبر عدستين.</p>

<p><strong>عبر عدسة ai-platform</strong>، منصة ai-platform من ThakiCloud هي بنية تحتية للذكاء الاصطناعي وتعلم الآلة متعددة المستأجرين قائمة على K8s. في مثل هذه البيئة، يجب التعامل مع تحميل البيانات ومعالجتها الأولية بوصفها مدخلًا من خارج حدّ الثقة لا من داخله. وعمليًا، يعني ذلك تشغيل مهام معالجة البيانات في حاويات معزولة بأدنى صلاحيات، وحجب المخرج الشبكي افتراضيًا، وفصل بيانات اعتماد العقدة والسحابة بحيث لا تلمسها أحمال العمل مباشرة. إن انتشار هذا الاختراق من الوصول بمستوى العقدة إلى سرقة بيانات الاعتماد يُظهر مجددًا لماذا يجب أن يكون عزل التنفيذ وفصل بيانات الاعتماد افتراضًا لا خيارًا. وهذا أيضًا سبب ارتفاع الطلب على الذكاء الاصطناعي المحلي والسيادي: فكلما بقيت البيانات والتنفيذ داخل حدود العميل، صغُر نطاق انفجار مثل هذه الهجمات على خط المعالجة.</p>

<p><strong>عبر عدسة Paxis</strong>، يتداخل هذا الحادث تمامًا مع نموذج التهديد الذي صُممت له سحابة أصلية للوكلاء منذ البداية. Paxis هي سحابة ThakiCloud الأصلية للوكلاء، وتعتبر تشغيل المهارات والأدوات في صناديق رملية معزولة وتمرير كل إجراء عبر بوابة سياسة وسجل تدقيق مبادئ من الدرجة الأولى. إن إلقاء المهاجم آلاف الإجراءات بسرب وكلاء ذاتي يثبت بالضبط لماذا يلزم بناء يفحص سلوك الوكيل بالسياسة قبل التنفيذ ويسجله في سجل تدقيق بعد التنفيذ. ولمواجهة نمط هجوم يستخدم صناديق رملية قصيرة العمر ثم يتخلص منها، يجب على المدافع أيضًا عزل كل تنفيذ، وتحديد نطاق صلاحياته صراحةً، وترك أثر تدقيق قابل للعكس. التنفيذ المعزول مع السياسة والتدقيق ليس ترفًا في عصر الوكلاء بل حدًا أدنى من المتطلبات.</p>

<p>تتكامل العدستان. تضيّق ai-platform نطاق الانفجار في طبقة البنية التحتية لمعالجة البيانات، بينما تفحص Paxis كل إجراء في طبقة التحكم لسلوك الوكيل. في هجوم كهذا، حيث الدخول خط بيانات والانتشار وكيل ذاتي، يلزم الدفاع في الطبقتين لكسر السلسلة.</p>

<h2 id="الحدود-والاعتراضات">الحدود والاعتراضات</h2>

<p>تجنبًا للثقة المفرطة في استنتاجات هذه المقالة، ينبغي توضيح بضعة أمور. أولًا، لا تزال تفاصيل الحادث قيد الاستقرار. التفاصيل الملونة مثل العدد الدقيق للإجراءات، ونطاق سرقة بيانات الاعتماد، وحكاية رفض النموذج التجاري، تعتمد بشدة على التقارير الثانوية ويجب تمييزها عن الحقائق المؤكدة في الإشعار الرسمي.</p>

<p>ثانيًا، سرديتنا الدفاعية لا تعني الأمان الكامل. العزل والسياسة والتدقيق مبادئ تصميم تقلّص نطاق الانفجار، لا سحرًا يزيل الثغرات نفسها. ثغرات مثل تنفيذ التعليمات البرمجية عن بُعد في محمّل بيانات أو الحقن في تحليل الإعداد يجب أن تستمر ملاحقتها وترقيعها على مستوى الشيفرة، والعزل هو خط الدفاع الثاني الذي يحتوي الضرر عند انطلاق مثل تلك الثغرة.</p>

<p>ثالثًا، المبالغة في تقدير هجمات الوكلاء الذاتيين خطرة أيضًا. لم يكن السبب الجذري لهذا الاختراق ذكاءً اصطناعيًا متطورًا بل ثغرتين مألوفتين سمحتا لمدخل يعبر حدّ الثقة بتنفيذ التعليمات البرمجية. لم يكن الوكيل سوى الأتمتة التي استغلت تلك الثغرتين بسرعة واتساع أكبر. لذا تبقى أولوية الاستجابة في الأساسيات: فصل المدخلات غير الموثوقة عن صلاحيات التنفيذ، وفصل بيانات الاعتماد عن أحمال العمل، وجعل كل تنفيذ قابلًا للرصد.</p>

<p>سيبقى احتواء Hugging Face السريع وكشفه الشفاف مثالًا جيدًا على الاستجابة. وما يبقى من واجب علينا بسيط: التعامل مع مجموعات البيانات بوصفها شيفرة لا ملفات، وجعل كل إجراء للوكيل موضوعًا للفحص والتدقيق.</p>

<h2 id="المصادر">المصادر</h2>

<ul>
  <li><a href="https://huggingface.co/blog/security-incident-july-2026">Security incident disclosure, July 2026 (مدونة Hugging Face الرسمية)</a></li>
  <li><a href="https://www.helpnetsecurity.com/2026/07/20/hugging-face-breached-by-autonomous-ai-agent/">Hugging Face breached by autonomous AI agent (Help Net Security)</a></li>
  <li><a href="https://www.bleepingcomputer.com/news/security/hugging-face-breach-autonomous-ai-agent-system-internal-datasets-credentials/">Hugging Face warns an autonomous AI agent hacked its network (BleepingComputer)</a></li>
  <li><a href="https://thehackernews.com/2026/07/worlds-largest-ai-model-repository.html">World’s Largest AI Model Repository Hugging Face Breached by Autonomous AI Agent (The Hacker News)</a></li>
  <li>تقارير ثانوية (العدد الدقيق للإجراءات وحكاية رفض النموذج تقارير منقولة لا حقائق مؤكدة): Cryptobriefing, Undercode Testing</li>
</ul>]]></content><author><name>{&quot;name&quot;=&gt;nil, &quot;avatar&quot;=&gt;nil, &quot;bio&quot;=&gt;nil, &quot;location&quot;=&gt;&quot;Seoul, Korea&quot;, &quot;email&quot;=&gt;&quot;info@thakicloud.co.kr&quot;, &quot;uri&quot;=&gt;nil, &quot;home&quot;=&gt;nil, &quot;links&quot;=&gt;[{&quot;label&quot;=&gt;&quot;Website&quot;, &quot;icon&quot;=&gt;&quot;fas fa-fw fa-link&quot;, &quot;url&quot;=&gt;&quot;https://thakicloud.co.kr&quot;}, {&quot;label&quot;=&gt;&quot;GitHub&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-github&quot;, &quot;url&quot;=&gt;&quot;https://github.com/thakicloud&quot;}, {&quot;label&quot;=&gt;&quot;LinkedIn&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-linkedin&quot;, &quot;url&quot;=&gt;&quot;https://www.linkedin.com/company/thakicloud&quot;}, {&quot;label&quot;=&gt;&quot;X&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-x-twitter&quot;, &quot;url&quot;=&gt;&quot;https://x.com/thakicloud&quot;}]}</name><email>info@thakicloud.co.kr</email></author><category term="news" /><category term="security" /><category term="huggingface" /><category term="ai-agent" /><category term="supply-chain" /><category term="sandbox" /><category term="dataset-security" /><category term="news" /><category term="thakicloud" /><summary type="html"><![CDATA[في يوليو 2026 كشف Hugging Face عن اختراق داخلي قاده وكيل ذكاء اصطناعي ذاتي. كانت نقطة الدخول مجموعة بيانات خبيثة واحدة، وأدت ثغرتان في خط معالجة مجموعات البيانات إلى تنفيذ التعليمات البرمجية. نفصل ما تأكد عمّا لا يزال قيد التحقيق، ونشرح لماذا يجب التعامل مع معالجة البيانات بوصفها حدّ ثقة.]]></summary></entry><entry xml:lang="ar"><title type="html">لماذا لم يتفوق ضبط حجم الدفعة (batch size) الديناميكي على الجدولة الثابتة: تحليل صادق لفشل حلقة التحكم (control loop) في خدمة vLLM متعددة المستأجرين (multi-tenant)</title><link href="https://thakicloud.github.io/ar/research/agent-dynamic-batch-tuning-vllm/" rel="alternate" type="text/html" title="لماذا لم يتفوق ضبط حجم الدفعة (batch size) الديناميكي على الجدولة الثابتة: تحليل صادق لفشل حلقة التحكم (control loop) في خدمة vLLM متعددة المستأجرين (multi-tenant)" /><published>2026-07-21T00:00:00+09:00</published><updated>2026-07-21T00:00:00+09:00</updated><id>https://thakicloud.github.io/ar/research/agent-dynamic-batch-tuning-vllm</id><content type="html" xml:base="https://thakicloud.github.io/ar/research/agent-dynamic-batch-tuning-vllm/"><![CDATA[<p>إذا كنت تدير مجموعة (cluster) لخدمة vLLM على Kubernetes تتقاسم فيها عدة جهات مستأجرة (tenants) وحدات معالجة رسوميات (GPU) عالية الأداء مثل H200، وكنت تعتقد أن التحكم الثابت في القبول (static admission control) الخاص بـ Kueue “كافٍ”، فهذا المقال موجه إليك. بل يستحق القراءة أكثر إذا كنت تفترض أن “قيام عميل نموذج لغوي كبير (LLM agent) بضبط معاملات الخدمة (serving parameters) في الوقت الفعلي أمر أفضل دائماً بلا شك.” فهذه الورقة البحثية تُظهر أن هذا الافتراض ليس صحيحاً دائماً، وتفعل ذلك بأسباب محددة تماماً.</p>

<h2 id="الإشكالية-القبول-ثابت-بينما-الحمل-فعلي-في-الوقت-الحقيقي">الإشكالية: القبول ثابت بينما الحمل فعلي في الوقت الحقيقي</h2>

<p>أصبحت مشكلة خدمة الاستدلال (inference) لنماذج لغوية كبيرة لعدة مستأجرين على وحدات معالجة رسوميات مشتركة تحدياً تشغيلياً شائعاً الآن. تتقاسم عدة جهات مستأجرة، تختلف فيما بينها من حيث تقطّع حركة المرور (traffic burstiness) وطول التسلسلات (sequence length) ومستويات اتفاقية مستوى الخدمة (SLO)، مسرّعاً واحداً عالي السعة مثل H200 في الوقت ذاته. في المنصات القائمة على Kubernetes، عادة ما يقرر مجدوِل قائم على الحصص والطوابير (quota and queue based scheduler) مثل Kueue أي الأعمال يُسمح لها بالدخول إلى المجمع (pool)، لكن المعاملات الخاصة بزمن الخدمة، مثل حجم الدفعة (batch size) والعدد الأقصى للتسلسلات المتزامنة (max concurrent sequences)، وهي التي تحدد كيفية تقاسم الطلبات التي دخلت بالفعل لسعة المحرك المحدودة، تظل عادة مثبتة بشكل ثابت وقت النشر (deployment).</p>

<p>تطرح هذه الورقة البحثية السؤال الذي يترتب طبيعياً على ذلك. في مجمع H200 متعدد المستأجرين تديره Kueue، حين تتشارك عدة جهات مستأجرة خادم استدلال vLLM واحداً، هل يستطيع عميل نموذج لغوي كبير يراقب قياسات ذاكرة وحدة معالجة الرسوميات وزمن الاستجابة عن بعد (telemetry) في الوقت الفعلي، ويعيد ضبط حجم الدفعة والعدد الأقصى للتسلسلات المتزامنة لكل مستأجر عبر الإنترنت (online)، أن يحسّن فعلياً جبهة باريتو (Pareto frontier) للإنتاجية (throughput) وزمن الاستجابة عند المئين 99 (p99 latency) والتكلفة، مقارنة بالتحكم الثابت في القبول لدى Kueue؟ ولكي يجيب المؤلفون عن هذا السؤال، صمموا بروتوكولاً تجريبياً كاملاً يُفترض تشغيله على وحدة H200 واحدة، لكنهم يوضحون منذ البداية أن التنفيذ نفسه لم يتم لعدم توفر الوصول إلى سياق مجموعة Kubernetes المستهدفة. لذلك، لا يظهر في هذه الورقة أي رقم واحد لمعدل إنتاجية أو زمن استجابة قِيس على عتاد حقيقي. وبدلاً من ذلك، تقدم الورقة ثلاثة أشياء: صياغة رياضية لحلقة التحكم (control loop) الخاصة بمسألة الضبط الديناميكي لحجم الدفعة والتزامن (concurrency) عبر الإنترنت، وبروتوكولاً قابلاً للتكرار (reproducible) يمكن تنفيذه فور استعادة الوصول إلى المجموعة، ومحاكاة قائمة انتظار (queuing simulation) حتمية تختبر بنية ثلاث سياسات تحكم مختلفة تحت الضغط.</p>

<h2 id="صياغة-مسألة-الضبط-كحلقة-تحكم">صياغة مسألة الضبط كحلقة تحكم</h2>

<p>تصوغ الورقة مسألة الضبط الديناميكي عبر الإنترنت لحجم الدفعة والتزامن الخاصين بكل مستأجر باعتبارها مسألة تحكم مغلقة الحلقة (closed loop control) في زمن منفصل (discrete time)، ذات بنية شبيهة بعملية قرار ماركوف (Markov decision process). تتكون الحالة (state) من عمق طابور كل مستأجر، ونسبة استخدام ذاكرة وحدة معالجة الرسوميات، ومئينات زمن الاستجابة الأخيرة (p50/p95/p99)، والحد الأقصى الحالي للتزامن. أما الفعل (action) فهو زيادة أو نقصان محدودة النطاق في تزامن كل مستأجر. وتُصاغ دالة المكافأة (reward function) بطرح عقوبة على تجاوز زمن الاستجابة عند المئين 99 لاتفاقية مستوى الخدمة، وبند التكلفة، من الإنتاجية، بحيث تُحسَّن الإنتاجية وزمن الاستجابة والتكلفة معاً ضمن نطاق آمن. وبناءً على هذه الصياغة، يقترح المؤلفون أيضاً طريقة دمج العميل في عملية النشر الفعلية. فالعميل لا يحل محل قبول المهام (job admission) الخاص بـ Kueue، بل يعمل كعامل جانبي (sidecar) داخله، ويضبط فقط طريقة تقاسم الأعمال التي قُبلت بالفعل لسعة المحرك. والنقطة الجوهرية أن الرافعة التي يتم التحكم بها ليست معامل max_num_seqs الداخلي لمحرك vLLM، بل تزامن القبول (admission concurrency) الخاص بكل مستأجر من جهة العميل. ولأن إصدارات vLLM الإنتاجية لا تُعرِّض معاملات المحرك كمقبض (knob) يمكن تغييره في الوقت الفعلي دون إعادة تشغيل، فإن هذا التصميم يعكس القيد العملي القائل بأن النقطة الوحيدة القابلة فعلياً للضبط هي عدد الطلبات المتزامنة الواردة إلى الخادم.</p>

<p><img src="/assets/images/posts/research/agent-dynamic-batch-tuning-vllm/fig-control-loop.png" alt="Control Loop: Agent-Driven Tuning Architecture" />
<em>بنية حلقة التحكم التي يراقب فيها عميل نموذج لغوي كبير قياسات وحدة معالجة الرسوميات وعمق طابور كل مستأجر، وينتج قيمة ضبط محدودة النطاق للتزامن، ليعيد من خلالها ضبط تشكيلة الدفعة (batch configuration). هذا رسم تخطيطي مفاهيمي للبنية، وليس نتيجة مقيسة فعلياً على عتاد.</em></p>

<h2 id="التحذير-الذي-كشفته-المحاكاة-التحكم-الديناميكي-الساذج-كان-أسوأ-من-الثابت">التحذير الذي كشفته المحاكاة: التحكم الديناميكي الساذج كان أسوأ من الثابت</h2>

<p>لسد الفراغ الناتج عن عدم تنفيذ البروتوكول التجريبي، بنى المؤلفون محاكاة لقائمة انتظار (queuing simulation) في زمن منفصل، ثابتة البذرة (seed fixed)، مكتوبة بالكامل باستخدام مكتبة Python القياسية فقط. وعلى نموذج مبسّط يتشارك فيه مستأجران مجمعاً محدوداً من الفتحات (slots)، شغّلوا ثلاث سياسات على مدى 20 بذرة عشوائية (seed) و1800 ثانية لكل منها، وأخذوا المتوسط. تُثبّت السياسة الثابتة الحد الأقصى لتزامن كل مستأجر عند 4. أما السياسة الديناميكية الساذجة فتراقب نسبة استخدام ذاكرة وحدة معالجة الرسوميات وزمن الاستجابة عند المئين 99 ضمن نافذة مدتها 6 ثوانٍ كل 3 ثوانٍ، وتضبط المستأجرَين معاً بزيادة أو نقصان مقداره 1. أما سياسة العميل المتمايز (differentiated agent) فهي نموذج بديل (surrogate model) يحاكي استدلالاً أكثر تطوراً للعميل، من خلال مراعاة نسبة التراكم (backlog ratio) لكل مستأجر بشكل مستقل.</p>

<p>جاءت النتائج مخالفة للتوقعات. فقد كانت السياسة الثابتة الأفضل في كل من الإنتاجية (0.686 طلب/ثانية) وعدد الطلبات المُسقَطة (بمتوسط 1.9 طلب)، بينما انخفضت إنتاجية السياسة الديناميكية الساذجة إلى 0.515 طلب/ثانية مع إسقاط 309.5 طلباً في المتوسط. وكان أداء نموذج العميل المتمايز البديل أفضل من السياسة الديناميكية الساذجة (إنتاجية 0.601 طلب/ثانية، وإسقاط 154.7 طلباً)، لكنه ظل غير قادر على مجاراة السياسة الثابتة، كما ارتفع زمن الاستجابة عند المئين 99 في كلا الشكلين الديناميكيين مقارنة بالسياسة الثابتة (11.75 ثانية)، إذ بلغ 12.79 ثانية و13.25 ثانية على التوالي.</p>

<p><img src="/assets/images/posts/research/agent-dynamic-batch-tuning-vllm/fig-throughput-dropped.png" alt="Throughput vs. Dropped Requests by Policy" />
<em>سجّلت السياسة الثابتة أعلى إنتاجية وأقل عدد من الطلبات المُسقَطة، بينما كان أداء كلا الشكلين الديناميكيين ضعيفاً في المحاكاة. هذه نتائج محاكاة قائمة انتظار حتمية بمتوسط 20 بذرة عشوائية، وليست قيماً مقيسة فعلياً على وحدة معالجة رسوميات.</em></p>

<p><img src="/assets/images/posts/research/agent-dynamic-batch-tuning-vllm/fig-p99-latency.png" alt="p99 Latency by Policy" />
<em>كان زمن الاستجابة عند المئين 99 للسياسة الثابتة هو الأدنى، ولم يتمكن أي من المتحكمين الديناميكيين من تقليل زمن الاستجابة في الذيل (tail latency) ضمن نطاق هذه المحاكاة. هذه نتائج محاكاة بمتوسط 20 بذرة عشوائية، وليست قِيماً مقيسة فعلياً على عتاد.</em></p>

<p>يتحقق المؤلفون من أن هذه النتيجة تعكس الديناميكيات الفعلية للنموذج وليست مجرد خطأ برمجي (bug)، ويتتبعون السبب عبر أربع خطوات. أولاً، بما أن زمن الخدمة (service time) يتبع توزيعاً أسياً (exponential distribution)، فإن الذيل ثقيل بطبيعته أصلاً. فتوزيع أسي بمتوسط 2.5 ثانية يُنتج زمن استجابة عند المئين 99 يقارب 11.5 ثانية حتى دون أي انتظار في الطابور على الإطلاق، وهو رقم لا يختلف كثيراً عن الـ 11.75 ثانية التي سجّلتها السياسة الثابتة. أي أن معظم زمن الاستجابة في الذيل الملاحَظ لا ينبع من الازدحام، بل من تباين زمن الخدمة نفسه. ثانياً، عتبة التقييد (throttle threshold) مضبوطة عند 6 ثوانٍ، أي ضعف نافذة الـ 3 ثوانٍ، وهي أقل بكثير من زمن الاستجابة الجوهري عند المئين 99 البالغ 11.5 ثانية. وزمن الاستجابة عند المئين 99 المقدَّر ضمن نافذة قصيرة، اعتماداً على عدد قليل من الطلبات المكتملة فقط، يحمل ضوضاء (noise) كبيرة وينحاز نحو هذا الذيل الجوهري، لذا يتجاوز العتبة كثيراً حتى في حالات الحمل الخفيف فعلياً. ثالثاً، ولأن هذا التقييد الكاذب (false positive throttle) يخفض المستأجرَين معاً، فإن المستأجر غير المزدحم يُقيَّد أيضاً مع المستأجر المزدحم، وتحت مهلة زمنية صارمة (hard timeout) مدتها 3 ثوانٍ، يتحول كل تقييد غير ضروري مباشرة إلى إسقاط للطلب. رابعاً، ورغم أن قاعدة التوسّع (scaling rule) تسمح بالتعافي، فإن تكلفة التقييد والإسقاط أكبر بشكل غير متناظر من مكسب التوسّع، وبالتالي تنخفض الإنتاجية الصافية.</p>

<h2 id="متطلبات-التصميم-المستخلصة-من-الفشل">متطلبات التصميم المستخلصة من الفشل</h2>

<p>لا يعني هذا التشخيص أن الضبط الديناميكي عديم الفائدة في حد ذاته. بل يعني أن اجتماع إشارة عالية الضوضاء ومنحازة نحو الذيل، وتقييد يربط المستأجرين معاً، ومهلة زمنية صارمة، يمكن أن يجعل التحكم الساذج بعتبة ثابتة أسوأ فعلياً من التحكم الثابت. من هنا، يستخلص المؤلفون أربعة متطلبات تصميم يجب أن يستوفيها أي متحكم عميل فعّال. يجب تقدير إشارة الازدحام بشكل متين اعتماداً على سجل قياسات أطول أمداً وتقديرات ثقة (confidence estimates)، بدلاً من مئينات خام ضمن نافذة قصيرة. ويجب أن تكون القرارات متمايزة لكل مستأجر، لا مبنية على إشارة عامة تجمع المستأجرين معاً. كما يجب مراعاة التكلفة غير المتناظرة بين التقييد والتوسّع بشكل صريح، والاستفادة من سياق إضافي لا يمكن التعبير عنه بعتبة رقمية ثابتة، مثل التعرف على أنماط التدفقات المفاجئة (burst patterns) أو بيانات وصفية خاصة باتفاقية مستوى الخدمة لكل مستأجر. والواقع أن استعادة نموذج العميل المتمايز البديل لنحو نصف الخسارة مقارنة بالمتحكم الساذج، تشير إلى أن هذا الاتجاه قد يكون صحيحاً بالفعل. غير أن المؤلفين يوضحون بصراحة أنه، بما أن كلاً من درجة الترابط ونوع الإشارة تغيّرا في آن واحد، فلا يمكن لهذه التجربة وحدها أن تفصل أيّهما ساهم في هذا التعافي.</p>

<h2 id="ما-تتركه-هذه-الورقة-للشركة-والمجتمع-والعلم">ما تتركه هذه الورقة للشركة والمجتمع والعلم</h2>

<p>بالنسبة إلى ThakiCloud، ما تتركه هذه الدراسة الآن ليس توفيراً مؤكداً في التكلفة، بل بروتوكولاً تجريبياً قابلاً للتكرار وقابلاً للتطبيق مباشرة على مجموعتنا من H200 متعددة المستأجرين، وتحذيراً محدداً من أن نهجاً ساذجاً قد يؤدي في الواقع إلى خسارة. وقد اكتمل الهيكل التجريبي (harness) نفسه بالفعل، بحيث يكفي، فور استعادة الوصول إلى المجموعة، استبدال دالة قرار عميل نموذج لغوي كبير فعلية داخل هذا البروتوكول. وعلى نطاق أوسع، كلما تطورت منهجيات تقليل هدر الموارد في مجمعات وحدات معالجة الرسوميات المشتركة، انفتح مسار عملي أمام المؤسسات الصغيرة أيضاً للوفاء باتفاقيات مستوى الخدمة على مجموعات مشتركة، وهو ما يفيد كفاءة الطاقة والتكلفة في بنية الاستدلال التحتية بشكل عام. أما من الناحية العلمية، فإن الإسهام يكمن في الصياغة الصريحة لحلقة تحكم يراقب فيها عميل نموذج لغوي كبير قياسات النظام في الوقت الفعلي ويعيد ضبط معاملات الخدمة الفائقة (serving hyperparameters) عبر الإنترنت، وفي التشخيص الكمي لأنماط فشلها من خلال محاكاة قابلة للتكرار. وتُعد هذه الورقة أيضاً جزءاً من سلسلة الأبحاث التي واصلتها ThakiCloud على نفس الأساس القائم على Kueue ووحدات معالجة الرسوميات. فمقارنة بـ ABJ-Gate، التي تجدول ميزانية أخذ العينات لنموذج الحَكَم (judge model)، وبـ Attested Confidential Sovereign Inference، التي تتناول التصديق عن بعد (remote attestation) لبيئة تنفيذ موثوقة (TEE) عند لحظة القبول، وبورقة “Escalate or Act?” التي تتناول قرارات ثنائية بين التصعيد أو التحرك إزاء حوادث منفصلة نادرة، تتميز هذه الورقة عنها جميعاً بتناولها مسألة مختلفة نوعياً، وهي التحكم المتعدد الأهداف في الوقت الفعلي الذي يُعاد ضبطه باستمرار على مدى ثوانٍ.</p>

<h2 id="القيود">القيود</h2>

<p>القيود التي تكشفها هذه الورقة عن نفسها واضحة. وأهمها أنه لا يوجد أي تحقق فعلي على عتاد حقيقي بأي شكل من الأشكال، إذ لم يُنفَّذ البروتوكول التجريبي بسبب تعذر الوصول إلى سياق المجموعة المستهدفة، ولم يُقَس أي رقم في الورقة على وحدة معالجة رسوميات فعلية. كما تبسّط المحاكاة الرموز (tokens) وذاكرة التخزين المؤقت KV إلى مجرد عدد فتحات عددي واحد (single scalar slot count)، دون أن تنمذج ديناميكيات مرحلتَي التعبئة المسبقة وفك الترميز (prefill/decode) أو ضغط الذاكرة، بل إن إشارة استخدام الذاكرة نفسها ليست سوى قيمة بديلة (proxy) لإشغال الفتحات، لا قياساً فعلياً لذاكرة الجهاز. كما اقتصر اختبار أنواع المستأجرين على نمطين اصطناعيين فقط لحركة المرور، لذا قد لا تعمم أنماط الفشل المكتشفة على مزيج آخر من حركة المرور. وسياسة العميل داخل المحاكاة أيضاً نموذج بديل مكتوب يدوياً وليس نموذجاً لغوياً كبيراً فعلياً، ولأن درجة الترابط ونوع الإشارة تغيّرا معاً في آن واحد، فلا يوجد حتى الآن دليل مُتحقَّق منه على أن عميلاً فعلياً قائماً على نموذج لغوي كبير سيستوفي متطلبات التصميم المستخلصة في القسم الرابع. ويوضح المؤلفون أيضاً بصراحة أنهم اكتفوا بالإبلاغ عن المتوسط والانحراف المعياري عبر 20 بذرة عشوائية، دون إجراء أي اختبار للدلالة الإحصائية (statistical significance). ولهذه الأسباب، تضع هذه الورقة نفسها لا كنتيجة مُتحقَّق منها، بل كصياغة رياضية وبروتوكول ومحاكاة تحذيرية.</p>

<p>يمكن الاطلاع على صفحة تفاصيل الورقة البحثية من هنا: <a href="https://huggingface.co/datasets/thaki-AI/daily-paper-2026-07-21-agent-dynamic-batch-tuning-vllm">https://huggingface.co/datasets/thaki-AI/daily-paper-2026-07-21-agent-dynamic-batch-tuning-vllm</a></p>]]></content><author><name>{&quot;name&quot;=&gt;nil, &quot;avatar&quot;=&gt;nil, &quot;bio&quot;=&gt;nil, &quot;location&quot;=&gt;&quot;Seoul, Korea&quot;, &quot;email&quot;=&gt;&quot;info@thakicloud.co.kr&quot;, &quot;uri&quot;=&gt;nil, &quot;home&quot;=&gt;nil, &quot;links&quot;=&gt;[{&quot;label&quot;=&gt;&quot;Website&quot;, &quot;icon&quot;=&gt;&quot;fas fa-fw fa-link&quot;, &quot;url&quot;=&gt;&quot;https://thakicloud.co.kr&quot;}, {&quot;label&quot;=&gt;&quot;GitHub&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-github&quot;, &quot;url&quot;=&gt;&quot;https://github.com/thakicloud&quot;}, {&quot;label&quot;=&gt;&quot;LinkedIn&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-linkedin&quot;, &quot;url&quot;=&gt;&quot;https://www.linkedin.com/company/thakicloud&quot;}, {&quot;label&quot;=&gt;&quot;X&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-x-twitter&quot;, &quot;url&quot;=&gt;&quot;https://x.com/thakicloud&quot;}]}</name><email>info@thakicloud.co.kr</email></author><category term="research" /><category term="vllm" /><category term="kueue" /><category term="gpu-scheduling" /><category term="multi-tenant-serving" /><category term="llm-agents" /><category term="dynamic-batching" /><category term="inference-cost-optimization" /><category term="h200" /><category term="queuing-simulation" /><category term="control-loop" /><summary type="html"><![CDATA[هل يتفوق عميل (agent) يراقب قياسات وحدة معالجة الرسوميات (GPU telemetry) ويضبط حجم الدفعة (batch size) والتزامن (concurrency) في الوقت الفعلي على التحكم الثابت في القبول (admission control) الخاص بـ Kueue؟ أجابت المحاكاة (simulation) بالنفي، وتتبعت السبب بدقة.]]></summary></entry><entry xml:lang="en"><title type="html">The Best AI That Fits Your Graphics Card</title><link href="https://thakicloud.github.io/en/comics/best-local-ai-by-vram-size/" rel="alternate" type="text/html" title="The Best AI That Fits Your Graphics Card" /><published>2026-07-21T00:00:00+09:00</published><updated>2026-07-21T00:00:00+09:00</updated><id>https://thakicloud.github.io/en/comics/best-local-ai-by-vram-size</id><content type="html" xml:base="https://thakicloud.github.io/en/comics/best-local-ai-by-vram-size/"><![CDATA[<p>Someone posted a ranking of the best local AI model for each VRAM tier of your graphics card. A local model is one that runs inside your own machine instead of being shipped off to the cloud. Phone-tier 4GB gets Bonsai, 12GB gets Gemma, 36GB gets Qwen, and so on: you just pick whatever your VRAM already holds. The quiet punchline is that you are not renting somebody’s GPU by the token. You are dropping the model onto a card that is already in your drawer. Paxis and Metis pull out a calculator and go to war over the chart.</p>

<p><img src="/assets/images/posts/comics/best-local-ai-by-vram-size/strip.png" alt="The Best AI That Fits Your Graphics Card" /></p>

<blockquote>
  <p>Source: <a href="https://x.com/hjguyhan/status/2079223629368463776">RT @jun_song: Best Local AI models by VRAM size (7/18)</a> · twitter</p>
</blockquote>

<h2 id="what-this-means-for-thakicloud">What this means for ThakiCloud</h2>

<p>What the chart is really about is sovereignty: keeping the model, the data, and the infrastructure under your control instead of someone else’s. On-prem is simply the version of that where the whole thing runs inside your own facility rather than a rented datacenter. That is the exact seat ThakiCloud sells. Metis matches a model to the size of the GPU you actually have, and Paxis runs agents on top of it to get real work done. Because the model lives in your rack, the meter does not tick per token. The cloud still has its place. The point is to ask which card is already in your drawer before you rent another.</p>

<hr />

<p><em>An auto-generated comic riffing on this week’s industry news.</em></p>]]></content><author><name>{&quot;name&quot;=&gt;nil, &quot;avatar&quot;=&gt;nil, &quot;bio&quot;=&gt;nil, &quot;location&quot;=&gt;&quot;Seoul, Korea&quot;, &quot;email&quot;=&gt;&quot;info@thakicloud.co.kr&quot;, &quot;uri&quot;=&gt;nil, &quot;home&quot;=&gt;nil, &quot;links&quot;=&gt;[{&quot;label&quot;=&gt;&quot;Website&quot;, &quot;icon&quot;=&gt;&quot;fas fa-fw fa-link&quot;, &quot;url&quot;=&gt;&quot;https://thakicloud.co.kr&quot;}, {&quot;label&quot;=&gt;&quot;GitHub&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-github&quot;, &quot;url&quot;=&gt;&quot;https://github.com/thakicloud&quot;}, {&quot;label&quot;=&gt;&quot;LinkedIn&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-linkedin&quot;, &quot;url&quot;=&gt;&quot;https://www.linkedin.com/company/thakicloud&quot;}, {&quot;label&quot;=&gt;&quot;X&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-x-twitter&quot;, &quot;url&quot;=&gt;&quot;https://x.com/thakicloud&quot;}]}</name><email>info@thakicloud.co.kr</email></author><category term="comics" /><category term="local-ai" /><category term="vram" /><category term="on-prem" /><category term="sovereignty" /><category term="cost" /><category term="thakicloud" /><summary type="html"><![CDATA[We sized the model to the card we already owned. The only thing that starved was the cloud invoice.]]></summary><media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://thakicloud.github.io/assets/images/posts/comics/best-local-ai-by-vram-size/strip.png" /><media:content medium="image" url="https://thakicloud.github.io/assets/images/posts/comics/best-local-ai-by-vram-size/strip.png" xmlns:media="http://search.yahoo.com/mrss/" /></entry><entry xml:lang="en"><title type="html">158 Skills and 24 Agents in One Plugin: How a Deterministic Skeleton Tames Agent Explosion</title><link href="https://thakicloud.github.io/en/dev/agentops/agent-plugin-158-skills-deterministic-flow/" rel="alternate" type="text/html" title="158 Skills and 24 Agents in One Plugin: How a Deterministic Skeleton Tames Agent Explosion" /><published>2026-07-21T00:00:00+09:00</published><updated>2026-07-21T00:00:00+09:00</updated><id>https://thakicloud.github.io/en/dev/agentops/agent-plugin-158-skills-deterministic-flow</id><content type="html" xml:base="https://thakicloud.github.io/en/dev/agentops/agent-plugin-158-skills-deterministic-flow/"><![CDATA[<p><img src="/assets/images/agent-plugin-158-skills-deterministic-flow-hero.png" alt="Abstract visualization of many skill modules converging into one ordered vertical pipeline" /></p>

<h2 id="overview">Overview</h2>

<p>Anyone who has built a serious agent system runs into the same paradox. Adding more skills and agents feels like it should make the system smarter, but often it does the opposite. Once you pass a few dozen skills, the agent starts getting confused about which skill to use when, and once there are several agents, they handle the same task differently or the order and format of outputs drift every run. Capability goes up while consistency of results goes down.</p>

<p>The open-source marketing plugin <strong>Digital Marketing Pro</strong> is an interesting case that tackles this paradox head on. It bundles 158 skills and 24 specialist agents (the repo docs list 25; the original tweet said 24) and still keeps the consistency of producing the same files in the same order every time. The trick is not a smarter model but a strategy flow fixed into 12 parts, a deterministic skeleton. This article dissects not the marketing tool itself but the agent engineering design inside it. What structure survives even when skills explode in number, and how that principle connects to the agent platform ThakiCloud is building.</p>

<p>Why this case matters to developers is clear. It shows, in concrete open-source code, why the naive hope of “just make lots of skills” so often fails in practice, and what stops that failure.</p>

<h2 id="what-the-plugin-is">What the Plugin Is</h2>

<p>Digital Marketing Pro is an open-source marketing plugin released under the MIT license. Its surface purpose is to help agencies and in-house marketing teams produce marketing documents consistently across many brands. According to the repo description, it targets agencies handling between 50 and 200 client brands, running every brand through the same 12-part flow to produce the same files in the same order.</p>

<p>By the numbers, the plugin is fairly large. It has 158 skills, 24 specialist agents, and a 12-part strategy flow expanded into 61 detailed steps. On top of that sit EU AI Act Article 50 readiness, AEO/GEO (answer engine optimization) features for six platforms including Google AI Mode, and Cowork support that persists state at the team level.</p>

<p>Worth noting is the install target. The plugin is not tied to Claude Code alone; it installs across multiple agent runtimes including Cowork, Codex, Cursor, Copilot CLI, and Antigravity. In other words, a single bundle of skills and agents is designed to work across many harnesses. This is an important enough design decision to treat separately below.</p>

<p>In short, beneath the appearance of a “marketing tool,” this plugin holds one answer to how you organize and consistently execute a large bundle of skills and agents.</p>

<h2 id="a-deterministic-skeleton-tames-the-skill-explosion">A Deterministic Skeleton Tames the Skill Explosion</h2>

<p>The core insight of this plugin is that it does not let the 158 skills and 24 agents collaborate freely. Instead it forces every task through a strategy flow fixed into 12 parts. Each part produces a defined output in a defined order, and there are explicit dependency rules between parts. A later part runs only when the earlier result exists, and the names and order of the result files stay identical even as the brand changes.</p>

<p>Why this matters becomes clear if you imagine the opposite. If 24 agents freely picked the skill that “looked best” and ran in free order, the composition and format of outputs would differ per brand. One brand might get competitor analysis first; another might skip that step entirely. If an agency manages 200 clients, this variance quickly becomes unauditable chaos. The 12-part flow deliberately reduces that freedom to raise average quality and consistency.</p>

<p>The flow below is a simplified view of how this deterministic skeleton constrains the freedom of skills and agents.</p>

<pre><code class="language-mermaid">flowchart TB
    A[Task request&lt;br/&gt;Brand X] --&gt; B[Enter fixed&lt;br/&gt;12-part flow]
    B --&gt; C[Each part: defined output&lt;br/&gt;defined order]
    C --&gt; D{Select the part-appropriate&lt;br/&gt;skill from 158}
    D --&gt; E{Assign a role&lt;br/&gt;from 24 agents}
    E --&gt; F[Apply explicit&lt;br/&gt;inter-part dependency rules]
    F --&gt; G[Same files, same order&lt;br/&gt;brand-independent consistency]
    G --&gt; H[Auditable&lt;br/&gt;document portfolio]
</code></pre>

<p>The lesson here has nothing to do with marketing. The way to protect quality as skills and agents grow is not to make the model smarter but to demote free design into filling in a validated skeleton. A deterministic structure owns the format, order, and dependencies, while the model fills only the content inside that skeleton. Whether there are 158 skills or 500, as long as the skeleton holds the degrees of freedom in check, the result stays predictable.</p>

<h2 id="what-it-means-to-install-across-six-runtimes">What It Means to Install Across Six Runtimes</h2>

<p>Another design worth watching is that this plugin installs across multiple agent runtimes. Claude Code, Cursor, Codex, and Copilot CLI are each a different harness. Their system prompts differ, their tool definition styles differ, and their permission models differ. That the same skill and agent bundle is designed to run on top of all of them means the capability was accumulated in the skills, not in the harness.</p>

<p>This distinction matters in practice. If the knowledge of a marketing workflow were baked into a specific tool’s config files or system prompt, switching tools would mean rebuilding everything. Conversely, when the knowledge lives in a portable bundle of skills, the harness stays thin and the skills are reused across tools. Digital Marketing Pro’s cross-runtime install is a case of practicing this “thin harness, fat skills” principle at a commercial scale.</p>

<p>Of course supporting many runtimes at once has a cost. Because each runtime loads and calls skills slightly differently, designing to the common denominator can leave a specific runtime’s unique features underused. Even so, prioritizing portability is a reasonable direction that frees skill assets from tool lock-in and lets them survive longer.</p>

<h2 id="implications-for-thakicloud-products">Implications for ThakiCloud Products</h2>

<p>What makes this case interesting is that it deals with a problem strikingly similar to what ThakiCloud is building with <strong>Paxis</strong>. Paxis is ThakiCloud’s agent-native cloud, treating Skills, Tools, Policies, and Audit Logs as first-class resources. A skill harness selects the right skill among more than 960 skills via BM25, runs it in an isolated sandbox, and passes every action through policy gates and audit logs.</p>

<p>The exact problem Digital Marketing Pro solved by taming 158 skills with a 12-part flow, Paxis solves at larger scale. Once skills pass 960, “which skill to use when” reaches a scale a human cannot specify by hand, so BM25-based skill selection replaces that skeleton. Instead of freely calling any skill, only the skills most relevant to the request are surfaced as candidates, reducing the degrees of freedom. This is the same principle by which the 12-part flow blocked free order, except that instead of a fixed flow it controls freedom through retrieval-based selection.</p>

<p>Also, the plugin’s emphasis on EU AI Act Article 50 readiness and auditable document output aligns with Paxis treating audit logs and policy gates as first-class. In customer environments where regulation and auditing matter, you must be able to trace “what was produced, in what order, on what basis.” A deterministic flow and audit logs are the two axes that create this traceability, and Paxis provides them at the platform level. No matter how many skills you stack, because policy gates and audit logs record every action, a large skill asset can be operated safely even in regulated environments.</p>

<p>Finally, cross-runtime portability matches the direction ThakiCloud aims for. A design that reuses a skill asset across harnesses rather than binding it to a specific tool is the same reason Paxis treats skills as first-class resources. When capability is accumulated in the skills rather than the harness, the assets you have built remain even as the tool changes.</p>

<h2 id="limitations-and-counterpoints">Limitations and Counterpoints</h2>

<p>It is important not to over-read this case. The fixed 12-part flow sacrifices flexibility in exchange for consistency. An exceptional need that departs from the standard flow, such as an unstructured task required only for a specific brand, may be handled awkwardly within this skeleton or not at all. A deterministic skeleton is powerful for repeatable bulk work but becomes a shackle for work with many creative exceptions.</p>

<p>The number 158 skills itself deserves careful reading. Having many skills means having many maintenance targets, and whether each skill is actually validated and kept current is a separate matter. A number does not guarantee quality. How many core skills the 12-part flow actually calls, and how often the rest are used, is hard to confirm from the repo docs alone [estimate].</p>

<p>Also, this article analyzes the plugin’s design principles, not the actual quality of its marketing output. That a deterministic flow produces consistent documents is a different matter from whether those documents lead to real marketing results. What we take from this case is not the marketing outcome but the engineering pattern of taming a large bundle of skills and agents with a deterministic skeleton.</p>

<h2 id="sources">Sources</h2>

<ul>
  <li>Repository: <a href="https://github.com/indranilbanerjee/digital-marketing-pro">github.com/indranilbanerjee/digital-marketing-pro</a></li>
  <li>Original source: <a href="https://x.com/hjguyhan/status/2079315207579660557">@tom_doerr tweet</a></li>
</ul>]]></content><author><name>{&quot;name&quot;=&gt;nil, &quot;avatar&quot;=&gt;nil, &quot;bio&quot;=&gt;nil, &quot;location&quot;=&gt;&quot;Seoul, Korea&quot;, &quot;email&quot;=&gt;&quot;info@thakicloud.co.kr&quot;, &quot;uri&quot;=&gt;nil, &quot;home&quot;=&gt;nil, &quot;links&quot;=&gt;[{&quot;label&quot;=&gt;&quot;Website&quot;, &quot;icon&quot;=&gt;&quot;fas fa-fw fa-link&quot;, &quot;url&quot;=&gt;&quot;https://thakicloud.co.kr&quot;}, {&quot;label&quot;=&gt;&quot;GitHub&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-github&quot;, &quot;url&quot;=&gt;&quot;https://github.com/thakicloud&quot;}, {&quot;label&quot;=&gt;&quot;LinkedIn&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-linkedin&quot;, &quot;url&quot;=&gt;&quot;https://www.linkedin.com/company/thakicloud&quot;}, {&quot;label&quot;=&gt;&quot;X&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-x-twitter&quot;, &quot;url&quot;=&gt;&quot;https://x.com/thakicloud&quot;}]}</name><email>info@thakicloud.co.kr</email></author><category term="dev" /><category term="agentops" /><category term="AgentOps" /><category term="Skills" /><category term="MultiAgent" /><category term="ClaudeCode" /><category term="Plugins" /><category term="Determinism" /><category term="Paxis" /><category term="AIAgents" /><summary type="html"><![CDATA[The open-source marketing plugin Digital Marketing Pro bundles 158 skills and 24 specialist agents without collapsing. The trick is a deterministic skeleton: a fixed 12-part flow. We dissect the design and show how ThakiCloud's Paxis productizes the same principle.]]></summary></entry><entry xml:lang="en"><title type="html">Claude Code Screen Reader Mode: One Flag That Opens Terminal AI Coding to Everyone</title><link href="https://thakicloud.github.io/en/dev/claude-code-screen-reader-accessibility/" rel="alternate" type="text/html" title="Claude Code Screen Reader Mode: One Flag That Opens Terminal AI Coding to Everyone" /><published>2026-07-21T00:00:00+09:00</published><updated>2026-07-21T00:00:00+09:00</updated><id>https://thakicloud.github.io/en/dev/claude-code-screen-reader-accessibility</id><content type="html" xml:base="https://thakicloud.github.io/en/dev/claude-code-screen-reader-accessibility/"><![CDATA[<p><img src="/assets/images/claude-code-screen-reader-accessibility-hero.png" alt="Abstract visualization of a terminal reorganized into a clean linear stream of text" /></p>

<h2 id="overview">Overview</h2>

<p>Terminal-based AI coding tools have mostly evolved toward filling the screen beautifully: live spinners, color-coded diffs, boxed permission prompts, and progress indicators that redraw as the cursor moves around. For sighted users, that visual density is a strength. For a developer who reads the terminal with a screen reader rather than with their eyes, it works in reverse. A screen that constantly redraws makes it hard for the screen reader to decide what is actually new, and boxes and animations get narrated as orderless noise.</p>

<p>Claude Code now tackles this head on with a screen reader mode. A single line, <code class="language-plaintext highlighter-rouge">claude --ax-screen-reader</code>, turns the visual terminal UI into plain, linear text. Instead of ornate rendering, it prints labeled lines in order so that screen readers like VoiceOver, NVDA, and JAWS can read top to bottom naturally. This article walks through exactly what the mode changes, how it works, and why the accessibility of agent interfaces is a problem the whole development ecosystem should own right now.</p>

<p>It looks like a small flag, but the change widens the answer to a real question: who can actually use a terminal AI agent? It is a theme ThakiCloud keeps running into while building an agent-native cloud, so we cover it not just as a feature note but from an interface design perspective.</p>

<h2 id="what-the-screen-reader-mode-is">What the Screen Reader Mode Is</h2>

<p>A normal Claude Code session treats the terminal like a canvas. It moves the cursor, erases lines it already printed and redraws them, and shows progress as a live animation. That is optimal for someone scanning the screen with their eyes, but it is the worst possible input for a screen reader. The screen reader must decide what to read every time the buffer changes, and when the screen redraws every frame it tends to repeat the same content or miss the important new output entirely.</p>

<p>Screen reader mode changes the rendering model itself. Instead of redrawing the screen, it appends new information as labeled single lines in order. When a tool runs, for example, explicit labels such as a permission request, a tool-running notice, and a result come through as text. The screen reader simply reads that linear text top to bottom, so following the whole conversation, approving tool permissions, and reviewing output can all be completed by sound alone.</p>

<p>The flow below is a simplified view of how the two rendering paths diverge.</p>

<pre><code class="language-mermaid">flowchart TB
    A[Claude Code session starts] --&gt; B{Screen reader mode&lt;br/&gt;enabled?}
    B --&gt;|Normal mode| C[Canvas redraw&lt;br/&gt;cursor moves, re-render]
    B --&gt;|--ax-screen-reader| D[Linear text output&lt;br/&gt;append labeled lines]
    C --&gt; E[High-density info&lt;br/&gt;for sighted users]
    D --&gt; F[Screen reader reads&lt;br/&gt;top to bottom]
    F --&gt; G[Conversation, approvals, review&lt;br/&gt;completed by sound]
    D --&gt; H[Terminal bell&lt;br/&gt;when attention needed]
</code></pre>

<p>The point is not “give less information” but “give the same information as ordered text.” Instead of stripping meaning, it removes visual flourish and provides a monotone, predictable output stream that a screen reader can trust.</p>

<h2 id="how-to-enable-it-and-how-it-works">How to Enable It and How It Works</h2>

<p>There are two ways to turn on screen reader mode. To enable it for a single session, pass the flag at launch.</p>

<div class="language-bash highlighter-rouge"><div class="highlight"><pre class="highlight"><code>claude <span class="nt">--ax-screen-reader</span>
</code></pre></div></div>

<p>This flag genuinely exists in the installed Claude Code. Checking the help output shows it:</p>

<div class="language-bash highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="nv">$ </span>claude <span class="nt">--help</span> | <span class="nb">grep </span>ax-screen
  <span class="nt">--ax-screen-reader</span>                    Render screen-reader friendly output
</code></pre></div></div>

<p>To apply it by default to every session started from a shell, set the environment variable.</p>

<div class="language-bash highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="nb">export </span><span class="nv">CLAUDE_AX_SCREEN_READER</span><span class="o">=</span>1
</code></pre></div></div>

<p>Now any Claude Code session opened in that shell uses screen-reader-friendly output without a separate flag. According to the official docs, this mode works on Claude Code v2.1.181 and later, and earlier versions reject the <code class="language-plaintext highlighter-rouge">--ax-screen-reader</code> flag with an error.</p>

<p>There are thoughtful behavioral details too. In screen reader mode, Claude Code rings the terminal bell when it needs the user’s attention. In particular, the bell rings when a tool that ran longer than five seconds finishes, signaling the end of a long task without needing to look at the screen. A screen reader user cannot visually confirm when a result has arrived after kicking off a command, so this audible signal creates a rhythm for the interaction.</p>

<p>There is a separate setting for low-vision users who rely on a screen magnifier.</p>

<div class="language-bash highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="nb">export </span><span class="nv">CLAUDE_CODE_ACCESSIBILITY</span><span class="o">=</span>1
</code></pre></div></div>

<p>Setting this keeps the native terminal cursor visible. Screen magnifiers like macOS Zoom magnify the screen by following the cursor position, so if a tool hides the cursor the magnifier loses focus. This setting exposes the cursor so the magnifier can accurately track where the user is.</p>

<p>So the accessibility support splits three ways: linear text output for screen readers, a terminal bell for attention, and cursor persistence for magnifiers. Each targets a different assistive technology and can be enabled independently through environment variables.</p>

<h2 id="why-it-matters-now">Why It Matters Now</h2>

<p>The first reason this feature matters is that terminal AI agents are quickly becoming a core tool for developers. Reading code, fixing it, running commands, and reviewing results increasingly happen inside these tools. If accessibility is missing from that flow, blind or low-vision developers cannot use the same productivity tools their peers use. No matter how capable the tool is, if the door to that capability is narrow, for some developers it does not exist.</p>

<p>The second reason is that this feature started from a community request. Issues asking for NVDA and JAWS support were filed in the public repository, and that demand turned into an actual release. Accessibility features are often pushed to “later,” so a case where a user request raised the priority is a good reference. Accessibility is not a niche special need; it is a design axis that determines the range of people who can use the tool.</p>

<p>The third is that this approach reaffirms an old truth: linear text is a robust interface. A text stream that is clearly ordered, labeled, and predictable is not only good for screen readers. It is easy to log, easy to pipe, and easy to parse for automation. It is no coincidence that an output mode built for accessibility turns out to also favor scripting and auditing.</p>

<h2 id="implications-for-thakicloud-products">Implications for ThakiCloud Products</h2>

<p>ThakiCloud operates an agent-native cloud called <strong>Paxis</strong>. Paxis treats Skills, Tools, Policies, and Audit Logs as first-class resources: a skill harness selects the right skill among many and runs it in an isolated sandbox, passing every action through policy gates and audit logs. As the surface where an agent interacts with people grows, the question of whether that surface is “accessible to anyone” becomes a basic design axis rather than an add-on.</p>

<p>The lesson from Claude Code’s screen reader mode is clear. Accessibility of an agent interface, separate from making the screen beautiful, comes down to whether you can present the same information as linear, labeled text. A platform like Paxis that already treats audit logs and policy gates as first-class is structurally well positioned here. Because every agent action is already recorded as a labeled event, reconstructing that event stream into human-readable linear output is less about building a whole new rendering pipeline and more about surfacing structured logs you already have.</p>

<p>This case also shows that accessible output and automation-friendly output share the same root. Text a screen reader can read is also text a log collector can parse and an audit trail can preserve. Given how much ThakiCloud emphasizes observability and audit on its agent platform, designing an accessible linear interface alongside them satisfies both goals at once. Rather than treating a rich UI and accessible text as opposites, the better approach renders both representations on the common foundation of a structured event stream.</p>

<h2 id="limitations-and-counterpoints">Limitations and Counterpoints</h2>

<p>It is important not to overstate this feature. Screen reader mode is a starting line for accessibility, not the finish. Outputting linear text does not automatically make every interaction comfortable, and understanding a long code block or a complex diff by sound alone is still a cognitively heavy task. Grasping the full context of a large refactor without a screen remains hard even with this mode.</p>

<p>The attention signal that relies on the terminal bell also varies by environment. Some terminal emulators are configured to convert the bell into a visual flash or to silence it entirely, so the bell signal may not arrive as intended. Users need to tune their own terminal settings for the best experience.</p>

<p>Finally, the fact that an accessibility mode exists is different from the fact that it has been thoroughly validated in practice. Real blind developers need to use it across various screen readers and workflows over a long period, accumulating feedback before the rough edges surface and get smoothed. Given that this mode first works in v2.1.181, it is still early, with plenty of room for improvement ahead. Even so, the fact that such a feature is included in the default distribution is itself a meaningful signal of a direction to handle accessibility now rather than later.</p>

<h2 id="sources">Sources</h2>

<ul>
  <li>Claude Code accessibility docs: <a href="https://code.claude.com/docs/en/accessibility">code.claude.com/docs/en/accessibility</a></li>
  <li>Feature request issue (NVDA/JAWS): <a href="https://github.com/anthropics/claude-code/issues/11002">anthropics/claude-code #11002</a></li>
  <li>Original source: <a href="https://x.com/hjguyhan/status/2079435394727416168">@ClaudeDevs tweet</a></li>
</ul>]]></content><author><name>{&quot;name&quot;=&gt;nil, &quot;avatar&quot;=&gt;nil, &quot;bio&quot;=&gt;nil, &quot;location&quot;=&gt;&quot;Seoul, Korea&quot;, &quot;email&quot;=&gt;&quot;info@thakicloud.co.kr&quot;, &quot;uri&quot;=&gt;nil, &quot;home&quot;=&gt;nil, &quot;links&quot;=&gt;[{&quot;label&quot;=&gt;&quot;Website&quot;, &quot;icon&quot;=&gt;&quot;fas fa-fw fa-link&quot;, &quot;url&quot;=&gt;&quot;https://thakicloud.co.kr&quot;}, {&quot;label&quot;=&gt;&quot;GitHub&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-github&quot;, &quot;url&quot;=&gt;&quot;https://github.com/thakicloud&quot;}, {&quot;label&quot;=&gt;&quot;LinkedIn&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-linkedin&quot;, &quot;url&quot;=&gt;&quot;https://www.linkedin.com/company/thakicloud&quot;}, {&quot;label&quot;=&gt;&quot;X&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-x-twitter&quot;, &quot;url&quot;=&gt;&quot;https://x.com/thakicloud&quot;}]}</name><email>info@thakicloud.co.kr</email></author><category term="dev" /><category term="ClaudeCode" /><category term="Accessibility" /><category term="ScreenReader" /><category term="AICoding" /><category term="DeveloperProductivity" /><category term="Paxis" /><category term="InclusiveDev" /><summary type="html"><![CDATA[Claude Code added a screen reader mode that swaps its visual terminal UI for plain, linear text. Here is what `claude --ax-screen-reader` actually changes, how it works, and why accessibility of agent interfaces matters for platforms like ThakiCloud.]]></summary></entry><entry xml:lang="en"><title type="html">Switch off the OpenAI Realtime API in one line: hugging-voice, an open voice stack you run yourself</title><link href="https://thakicloud.github.io/en/llmops/hugging-voice-open-realtime-voice-self-hosted/" rel="alternate" type="text/html" title="Switch off the OpenAI Realtime API in one line: hugging-voice, an open voice stack you run yourself" /><published>2026-07-21T00:00:00+09:00</published><updated>2026-07-21T00:00:00+09:00</updated><id>https://thakicloud.github.io/en/llmops/hugging-voice-open-realtime-voice-self-hosted</id><content type="html" xml:base="https://thakicloud.github.io/en/llmops/hugging-voice-open-realtime-voice-self-hosted/"><![CDATA[<p><img src="/assets/images/hugging-voice-open-realtime-voice-self-hosted-hero.png" alt="An open realtime voice pipeline you run yourself" /></p>

<p>This post is for engineers who wanted to add a voice agent but hesitated at the lock-in and cost of the OpenAI Realtime API, and for infrastructure owners weighing whether conversational voice can be served on their own stack. The short version is that the design of Hugging Face’s demo <a href="https://huggingface.co/spaces/HuggingFaceM4/hugging-voice">hugging-voice</a> and the library beneath it, <a href="https://github.com/huggingface/speech-to-speech">speech-to-speech</a>, is both simple and practical. It opens the entire four-stage realtime voice pipeline as open source, while wrapping the outside in the very same interface as OpenAI Realtime. So if you already have code written against an OpenAI realtime client, you can move onto your own stack by changing a single line: the address the server points to. We only cite performance figures within the range the project has published, and we note up front that these are not numbers we benchmarked ourselves.</p>

<h2 id="overview">Overview</h2>

<p>Over the past year, conversational voice has stopped being a side feature of text chatbots and become a product category of its own. Users speak, and they expect the reply to come back instantly, the way a human conversation flows. The problem is that the commercial path to meeting that expectation has effectively converged on a handful of closed services like the OpenAI Realtime API. Convenient, yes, but voice traffic tends to be billed by the second rather than by the token, data leaves your walls, and both the model and the voices are tied to the provider.</p>

<p>hugging-voice arrives as a counterexample to that trend. The subtitle of the Space says it directly: “An Open Realtime Voice You Can Actually Run Yourself.” The core idea is that the whole round trip, turning the voice arriving at the microphone into text, sending it to a language model, and turning the reply back into speech, is opened as a pipeline whose every component can be swapped. For those of us who serve models in on-prem and sovereign environments, it means there is now a concrete reference implementation for the question “can realtime voice run on our own cluster?”</p>

<h2 id="what-hugging-voice-and-speech-to-speech-are">What hugging-voice and speech-to-speech are</h2>

<p>To settle the terms first: hugging-voice is the demo Space where you can speak to it directly in the browser, and the engine that actually processes the voice inside it is the speech-to-speech library. The library splits a realtime voice agent into four stages: voice activity detection (VAD), speech-to-text (STT), a language model (LLM), and text-to-speech (TTS). Each stage runs in a separate thread and is connected by queues, so the output of one stage streams into the next. A partial transcript appears before the user has finished speaking, and speech synthesis begins on the opening words before the model has finished the sentence, which cuts perceived latency.</p>

<div class="mermaid">
flowchart TB
    A["Microphone input<br />realtime audio stream"] --&gt; B["VAD speech detection<br />Silero VAD v5"]
    B --&gt; C["STT recognition<br />Parakeet TDT · Whisper etc."]
    C --&gt; D["LLM response<br />OpenAI-compatible API · vLLM · llama.cpp"]
    D --&gt; E["TTS synthesis<br />Qwen3-TTS · Kokoro etc."]
    E --&gt; F["Speaker output<br />streaming playback"]
    G["OpenAI Realtime compatible<br />WebSocket server"] -.- B
    G -.- C
    G -.- D
    G -.- E
</div>

<p>In this diagram, the WebSocket server on the right is the project’s real weapon. Simply gluing four stages together is not new. What sets speech-to-speech apart is that it wraps the entire pipeline in a WebSocket endpoint compatible with the OpenAI Realtime protocol. That lets an existing OpenAI realtime client connect to this server as if it were OpenAI itself. It is worth noting that this stack is not an experimental toy: it runs the realtime voice infrastructure of Hugging Face’s Reachy Mini robots in production.</p>

<h2 id="switching-over-in-one-line">Switching over in one line</h2>

<p>This is the part the project itself calls “one-line migration.” The only thing a client that used OpenAI Realtime has to change is the connection address. Below is the Python client example the project documentation offers.</p>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="kn">from</span> <span class="n">openai</span> <span class="kn">import</span> <span class="n">OpenAI</span>

<span class="n">client</span> <span class="o">=</span> <span class="nc">OpenAI</span><span class="p">(</span>
    <span class="n">base_url</span><span class="o">=</span><span class="sh">"</span><span class="s">http://localhost:8765/v1</span><span class="sh">"</span><span class="p">,</span>
    <span class="n">websocket_base_url</span><span class="o">=</span><span class="sh">"</span><span class="s">ws://localhost:8765/v1</span><span class="sh">"</span><span class="p">,</span>
    <span class="n">api_key</span><span class="o">=</span><span class="sh">"</span><span class="s">not-needed</span><span class="sh">"</span><span class="p">,</span>
<span class="p">)</span>

<span class="k">with</span> <span class="n">client</span><span class="p">.</span><span class="n">realtime</span><span class="p">.</span><span class="nf">connect</span><span class="p">(</span><span class="n">model</span><span class="o">=</span><span class="sh">"</span><span class="s">local</span><span class="sh">"</span><span class="p">)</span> <span class="k">as</span> <span class="n">conn</span><span class="p">:</span>
    <span class="n">conn</span><span class="p">.</span><span class="nf">send</span><span class="p">({</span>
        <span class="sh">"</span><span class="s">type</span><span class="sh">"</span><span class="p">:</span> <span class="sh">"</span><span class="s">session.update</span><span class="sh">"</span><span class="p">,</span>
        <span class="sh">"</span><span class="s">session</span><span class="sh">"</span><span class="p">:</span> <span class="p">{</span>
            <span class="sh">"</span><span class="s">type</span><span class="sh">"</span><span class="p">:</span> <span class="sh">"</span><span class="s">realtime</span><span class="sh">"</span><span class="p">,</span>
            <span class="sh">"</span><span class="s">instructions</span><span class="sh">"</span><span class="p">:</span> <span class="sh">"</span><span class="s">You are a helpful assistant.</span><span class="sh">"</span><span class="p">,</span>
            <span class="sh">"</span><span class="s">audio</span><span class="sh">"</span><span class="p">:</span> <span class="p">{</span>
                <span class="sh">"</span><span class="s">input</span><span class="sh">"</span><span class="p">:</span> <span class="p">{</span>
                    <span class="sh">"</span><span class="s">turn_detection</span><span class="sh">"</span><span class="p">:</span> <span class="p">{</span>
                        <span class="sh">"</span><span class="s">type</span><span class="sh">"</span><span class="p">:</span> <span class="sh">"</span><span class="s">server_vad</span><span class="sh">"</span><span class="p">,</span>
                        <span class="sh">"</span><span class="s">interrupt_response</span><span class="sh">"</span><span class="p">:</span> <span class="bp">True</span><span class="p">,</span>
                    <span class="p">}</span>
                <span class="p">}</span>
            <span class="p">},</span>
        <span class="p">}</span>
    <span class="p">})</span>

    <span class="k">for</span> <span class="n">event</span> <span class="ow">in</span> <span class="n">conn</span><span class="p">:</span>
        <span class="nf">print</span><span class="p">(</span><span class="n">event</span><span class="p">.</span><span class="nb">type</span><span class="p">)</span>
</code></pre></div></div>

<p>The parts worth noticing are that <code class="language-plaintext highlighter-rouge">base_url</code> and <code class="language-plaintext highlighter-rouge">websocket_base_url</code> point to a local server and that <code class="language-plaintext highlighter-rouge">api_key</code> is essentially not needed. The instructions passed through <code class="language-plaintext highlighter-rouge">session.update</code>, server-side VAD turn detection, and interrupting a response mid-stream all follow the same schema as OpenAI Realtime. In other words, the application code barely changes, and only the backend moves from an external API to your own server. For teams worried about vendor lock-in, this interface compatibility is on its own the biggest practical value.</p>

<h2 id="install-and-run">Install and run</h2>

<p>The path to standing up a server is just as brief. The default install covers the standard realtime path in one shot.</p>

<div class="language-bash highlighter-rouge"><div class="highlight"><pre class="highlight"><code>pip <span class="nb">install </span>speech-to-speech
</code></pre></div></div>

<p>The default configuration uses Parakeet TDT for STT, an OpenAI-compatible API for the LLM, and Qwen3-TTS for TTS. If you need a specific backend, install it with extras.</p>

<div class="language-bash highlighter-rouge"><div class="highlight"><pre class="highlight"><code>pip <span class="nb">install</span> <span class="s2">"speech-to-speech[kokoro]"</span>
pip <span class="nb">install</span> <span class="s2">"speech-to-speech[faster-whisper]"</span>
</code></pre></div></div>

<p>Running the server looks like this. The command launches an OpenAI Realtime compatible server over a local WebSocket.</p>

<div class="language-bash highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="nb">export </span><span class="nv">OPENAI_API_KEY</span><span class="o">=</span>...
speech-to-speech
</code></pre></div></div>

<p>The interesting point here is that the LLM stage can run fully local. Below is an example that stands up a Gemma 4 class model with llama.cpp and has speech-to-speech point at that local endpoint.</p>

<div class="language-bash highlighter-rouge"><div class="highlight"><pre class="highlight"><code>llama-server <span class="nt">-hf</span> ggml-org/gemma-4-E4B-it-GGUF <span class="nt">-np</span> 2 <span class="nt">-c</span> 65536

speech-to-speech <span class="se">\</span>
    <span class="nt">--model_name</span> <span class="s2">"ggml-org/gemma-4-E4B-it-GGUF"</span> <span class="se">\</span>
    <span class="nt">--responses_api_base_url</span> <span class="s2">"http://127.0.0.1:8080/v1"</span> <span class="se">\</span>
    <span class="nt">--responses_api_api_key</span> <span class="s2">""</span>
</code></pre></div></div>

<p>On Apple Silicon Macs you can turn on optimized settings and attach an mlx model, and conversely, if local GPU is short, you can point the LLM backend at Hugging Face Inference Providers.</p>

<div class="language-bash highlighter-rouge"><div class="highlight"><pre class="highlight"><code>speech-to-speech <span class="se">\</span>
    <span class="nt">--local_mac_optimal_settings</span> <span class="se">\</span>
    <span class="nt">--model_name</span> <span class="s2">"mlx-community/Qwen3-4B-Instruct-2507-bf16"</span>
</code></pre></div></div>

<p>Being able to move the same pipeline across the spectrum, from fully local execution to delegated cloud inference, with a handful of flags is a virtue of the design.</p>

<h2 id="modular-swaps-stt-llm-tts-backends">Modular swaps: STT, LLM, TTS backends</h2>

<p>The reason this project reads as a reference architecture rather than a mere demo is that the backend of each stage can be swapped. In summary:</p>

<table>
  <thead>
    <tr>
      <th>Stage</th>
      <th>Default backend</th>
      <th>Alternatives</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>VAD</td>
      <td>Silero VAD v5</td>
      <td>built-in only</td>
    </tr>
    <tr>
      <td>STT</td>
      <td>Parakeet TDT</td>
      <td>Whisper, Faster Whisper, Paraformer</td>
    </tr>
    <tr>
      <td>LLM</td>
      <td>OpenAI-compatible API</td>
      <td>Transformers, mlx-lm, vLLM, llama.cpp</td>
    </tr>
    <tr>
      <td>TTS</td>
      <td>Qwen3-TTS</td>
      <td>Kokoro, Pocket TTS, ChatTTS, MMS</td>
    </tr>
  </tbody>
</table>

<p>There are also four run modes. The default, realtime, is the WebSocket that speaks the OpenAI Realtime protocol; local attaches directly to the microphone and speaker; and websocket and socket exchange raw PCM audio over WebSocket and TCP respectively. Language can be specified or left to auto-detection. Being able to combine STT accuracy, LLM quality and cost, and TTS timbre and latency to fit your own requirements is exactly the degree of freedom a closed API does not give you.</p>

<h2 id="latency-and-a-real-world-case">Latency, and a real-world case</h2>

<p>In a voice agent, latency matters as much as accuracy. People feel a conversation break when the reply is even a few hundred milliseconds late. The <a href="https://huggingface.co/blog/cerebras-gemma4-voice-ai">realtime voice demo</a> that Hugging Face published together with Cerebras targets this latency problem head-on. The configuration uses Nvidia’s Parakeet for STT, Google DeepMind’s Gemma 4 running on Cerebras inference for the language model, and Alibaba’s Qwen3-TTS for TTS. The goal is to drive down the LLM stage’s response time with Cerebras’ ultra-fast inference so that the conversation flows as naturally as talking to a person. That said, within what this article could confirm, no specific millisecond figures were published, so we hold off on a quantitative comparison.</p>

<p>There is production evidence too. This stack drives the Reachy Mini robots mentioned earlier, and Hugging Face states that more than 9,000 robots are already deployed in the field. The reason the project obsesses over latency is that in embedded settings, responsiveness is what makes an interaction “feel alive.” We have separately covered the latency budget of voice agents from a GPU serving angle, so if you are thinking about how to allocate latency across each stage of the pipeline, we suggest reading <a href="/en/llmops/voice-agent-latency-budget-gpu-serving/">Voice agent latency budget and GPU serving</a> alongside this.</p>

<h2 id="implications-for-thakiclouds-products">Implications for ThakiCloud’s products</h2>

<p>This pipeline meshes naturally with both of our products.</p>

<p>Through the ai-platform lens, the LLM stage of speech-to-speech ultimately just consumes an OpenAI-compatible endpoint, so you can drop ThakiCloud’s vLLM serving straight into that slot. The STT and TTS models each become separate GPU-occupying workloads, and our strength is exactly in loading such heterogeneous inference workloads onto one cluster together, queuing GPUs with Kueue and isolating them across tenants. Voice traffic tends to accumulate cost by the second, so the unit-cost competitiveness of self-serving works especially strongly, and this open stack fits on-prem and sovereign requirements where data must not leave the premises. To customers for whom a closed voice API is a burden, we can offer the option of “running it on your own cluster through the same interface.”</p>

<p>Through the Paxis lens, voice is a new input and output channel attached to an agent. Paxis is an Agent-Native Cloud control plane that runs on top of ai-platform and treats Skills, Tools, Policies, and Audit Logs as first-class resources, and the OpenAI Realtime compatible interface of speech-to-speech makes it easy to layer this voice channel onto existing agent orchestration. A flow where a spoken instruction becomes the agent’s input and the result of a skill execution returns as speech can be composed while passing through policy gates and audit logs. Low-cost self-serving (ai-platform) creates the economics of voice agents, and on top of it the Agent-Native control plane (Paxis) handles voice as a channel safely.</p>

<h2 id="closing">Closing</h2>

<p>The message hugging-voice sends is clear. Realtime voice no longer has to lean solely on a few providers’ closed APIs; when you need to, you can keep the interface as is and move only the backend onto your own infrastructure. This design, picking each stage from VAD through TTS and changing just one line in the client, sharply lowers the barrier for teams considering self-serving. If you want to see it for yourself, speak to it directly in the <a href="https://huggingface.co/spaces/HuggingFaceM4/hugging-voice">demo Space</a>, or start with <code class="language-plaintext highlighter-rouge">pip install speech-to-speech</code> from the <a href="https://github.com/huggingface/speech-to-speech">GitHub repository</a>.</p>]]></content><author><name>{&quot;name&quot;=&gt;nil, &quot;avatar&quot;=&gt;nil, &quot;bio&quot;=&gt;nil, &quot;location&quot;=&gt;&quot;Seoul, Korea&quot;, &quot;email&quot;=&gt;&quot;info@thakicloud.co.kr&quot;, &quot;uri&quot;=&gt;nil, &quot;home&quot;=&gt;nil, &quot;links&quot;=&gt;[{&quot;label&quot;=&gt;&quot;Website&quot;, &quot;icon&quot;=&gt;&quot;fas fa-fw fa-link&quot;, &quot;url&quot;=&gt;&quot;https://thakicloud.co.kr&quot;}, {&quot;label&quot;=&gt;&quot;GitHub&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-github&quot;, &quot;url&quot;=&gt;&quot;https://github.com/thakicloud&quot;}, {&quot;label&quot;=&gt;&quot;LinkedIn&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-linkedin&quot;, &quot;url&quot;=&gt;&quot;https://www.linkedin.com/company/thakicloud&quot;}, {&quot;label&quot;=&gt;&quot;X&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-x-twitter&quot;, &quot;url&quot;=&gt;&quot;https://x.com/thakicloud&quot;}]}</name><email>info@thakicloud.co.kr</email></author><category term="llmops" /><category term="speech-to-speech" /><category term="realtime-voice" /><category term="VoiceAI" /><category term="OpenAIRealtime" /><category term="STT" /><category term="TTS" /><category term="LLM-serving" /><category term="LLMOps" /><category term="on-prem" /><category term="self-hosting" /><summary type="html"><![CDATA[Hugging Face's hugging-voice and its engine speech-to-speech wrap a full realtime voice pipeline, from VAD through STT, LLM, and TTS, behind an OpenAI Realtime compatible WebSocket. Change one line in your client's base URL and you move onto your own infrastructure. We took the design apart from a serving perspective.]]></summary></entry><entry xml:lang="en"><title type="html">Hugging Face Wasn’t Breached by a Human, but by an Autonomous AI Agent: When the Dataset Pipeline Became the Attack Surface</title><link href="https://thakicloud.github.io/en/news/huggingface-agentic-ai-breach/" rel="alternate" type="text/html" title="Hugging Face Wasn’t Breached by a Human, but by an Autonomous AI Agent: When the Dataset Pipeline Became the Attack Surface" /><published>2026-07-21T00:00:00+09:00</published><updated>2026-07-21T00:00:00+09:00</updated><id>https://thakicloud.github.io/en/news/huggingface-agentic-ai-breach</id><content type="html" xml:base="https://thakicloud.github.io/en/news/huggingface-agentic-ai-breach/"><![CDATA[<p><img src="/assets/images/huggingface-agentic-ai-breach-hero.png" alt="Abstract image of an autonomous agent swarm infiltrating a data pipeline" /></p>

<p>The story that shook timelines over the weekend was not a new model or a new benchmark. It was a notice that Hugging Face, the center of the open AI ecosystem, had been breached. What drew even more attention was who did it. According to the company, no human hacker typed commands through the night. An autonomous AI agent framework drove the attack from start to finish.</p>

<p>A company that sells models getting hit by a model makes for a striking narrative. But the point of this post is not to enjoy the irony. For a company like ThakiCloud that handles models and data on top of customer infrastructure, the real work is to soberly separate exactly where the attack entered and what has been confirmed. And the entry point here was not some flashy zero-day. It was the thing we touch every day: a dataset.</p>

<h2 id="what-happened">What Happened</h2>

<p>Hugging Face disclosed the breach in a blog post on Thursday, July 16, 2026. It came after the company had already confirmed unauthorized access to internal datasets and credentials earlier that week and had contained the intrusion. By the company’s account, the intrusion began in the data-processing pipeline, where the attacker used a single malicious dataset to open two code-execution paths.</p>

<p>That is the confirmed skeleton: an autonomous agent drove it, the entry point was a dataset, and two vulnerabilities led to code execution. The remaining details are emphasized differently across outlets, so confirmed facts and secondary reporting should be read apart.</p>

<h2 id="the-attack-path-the-dataset-pipeline-was-the-attack-surface">The Attack Path: The Dataset Pipeline Was the Attack Surface</h2>

<p>The essence is the entry method. The attacker uploaded a malicious dataset to the Hugging Face Hub. The moment that dataset passed through the processing pipeline, two vulnerabilities fired in sequence. One was a remote-code dataset loader path; the other was a template injection while parsing the dataset configuration. Both ultimately resolved into arbitrary code execution.</p>

<p>The idea that a dataset can run code may sound unfamiliar, but practitioners know the risk well. Many dataset loaders trust and execute loading scripts from remote repositories and render configuration fields as templates. That flexibility, built for convenience, becomes an execution channel the moment it meets input that crosses a trust boundary.</p>

<p>What followed once code execution was secured was a textbook breach chain. The attacker escalated with node-level access, harvested cloud and cluster credentials, and moved laterally into several internal clusters over the weekend. The entry was a single point, but from the moment that point granted execution privileges, the spread propagated automatically.</p>

<pre><code class="language-mermaid">flowchart TB
    A[Attacker: uploads malicious dataset] --&gt; B[Dataset processing pipeline]
    B --&gt; C1["Vulnerability 1&lt;br/&gt;remote-code dataset loader"]
    B --&gt; C2["Vulnerability 2&lt;br/&gt;dataset config template injection"]
    C1 --&gt; D[Arbitrary code execution RCE]
    C2 --&gt; D
    D --&gt; E[Node-level access obtained]
    E --&gt; F[Cloud and cluster credentials harvested]
    F --&gt; G[Lateral movement into internal clusters]
    G --&gt; H["Autonomous agent framework&lt;br/&gt;thousands of actions across a swarm of short-lived sandboxes"]
</code></pre>

<h2 id="the-weight-of-saying-an-autonomous-agent-drove-it">The Weight of Saying an Autonomous Agent Drove It</h2>

<p>The novel part of this incident is not the tooling but the cockpit. Hugging Face described the campaign as “an autonomous agent framework executing many thousands of individual actions across a swarm of short-lived sandboxes, with self-migrating command-and-control staged on public services.” Instead of a human intervening at each step, the agent handled reconnaissance, execution, and movement in a continuous chain.</p>

<p>The problem this structure poses for defenders is speed and scale. A human attacker has physical limits of fatigue and typing speed, but an agent swarm throws thousands of attempts in parallel and moves to the next one the instant a step fails. Using and discarding short-lived sandboxes erases the anchors for detection, and command-and-control that migrates across public services defeats blocklists.</p>

<p>One interesting side note circulated in secondary reporting. As the response unfolded, when the team tried to hand forensics to commercial frontier models (GPT, Claude), safety guardrails reportedly recognized the exploit payloads and command-and-control artifacts as attacks and refused to cooperate, so the team continued detection and analysis with a GLM 5.2-class model [estimated]. This detail comes from some outlets rather than Hugging Face’s official notice, so it is safer not to read it as settled fact. Regardless of its accuracy, though, the tension itself, where a defender cannot use a tool because of its safety policy, is worth recording as something that may recur.</p>

<h2 id="what-was-safe-and-what-is-still-under-investigation">What Was Safe and What Is Still Under Investigation</h2>

<p>The easier an incident is to exaggerate, the clearer the boundaries must be drawn. Hugging Face said it closed the vulnerable code-execution paths, evicted the attacker, rebuilt the compromised nodes, and revoked and rotated all affected credentials. It added that it found no evidence of tampering with public models, user-facing datasets, or Spaces, and that its software supply chain, including container images and published packages, was verified clean.</p>

<p>The user action was a precautionary recommendation. The company advised users to rotate access tokens and review recent account activity. There is an important distinction here. That recommendation is not a confirmation that user tokens were leaked en masse, but a conservative safety measure given the nature of an incident where internal credentials were harvested. Whether partner or customer data was affected was, as of the disclosure, still under investigation.</p>

<p>In short, what is confirmed is the internal breach and credential theft, the existence of two dataset vulnerabilities, and the swift containment and rotation. What remains open is whether partner and customer data was affected, and the confirmation of some details in secondary reporting (the exact action count, the model-refusal anecdote). Mixing the confirmed with the unconfirmed makes an incident look bigger or smaller than it is.</p>

<h2 id="the-thakicloud-view-treating-dataset-processing-as-a-trust-boundary">The ThakiCloud View: Treating Dataset Processing as a Trust Boundary</h2>

<p>The lesson this incident offers an infrastructure company is clear. A dataset is not a passive file but an active input that can execute code the moment it is processed. So we look at this through two lenses.</p>

<p><strong>Through the ai-platform lens</strong>, ThakiCloud’s ai-platform is a K8s-based multi-tenant AI/ML infrastructure. In such an environment, dataset loading and preprocessing must be treated as input from outside the trust boundary, not inside it. Concretely, this means running dataset-processing jobs in least-privilege isolated containers, blocking network egress by default, and separating node and cloud credentials so workloads cannot touch them directly. That this breach spread from node-level access to credential theft shows again why execution isolation and credential separation must be a default, not an option. This is also why demand for on-prem and sovereign AI is high: the more data and execution stay inside the customer boundary, the smaller the blast radius of such pipeline attacks.</p>

<p><strong>Through the Paxis lens</strong>, this incident overlaps exactly with the threat model that an Agent-Native Cloud is designed for in the first place. Paxis is ThakiCloud’s Agent-Native Cloud, and it treats running skills and tools in isolated sandboxes and passing every action through a policy gate and audit log as first-class principles. That the attacker threw thousands of actions with an autonomous agent swarm proves precisely why a structure that screens agent behavior with policy before execution and records it in an audit log after execution is necessary. To counter an attack pattern that uses and discards short-lived sandboxes, the defender too must isolate each execution, explicitly scope its permissions, and leave a reversible audit trail. Isolated execution plus policy-and-audit is not a luxury of the agent era but a minimum requirement.</p>

<p>The two lenses complement each other. ai-platform narrows the blast radius at the infrastructure layer of dataset processing, while Paxis screens each action at the control layer of agent behavior. In an attack like this one, where entry is a data pipeline and the spread is an autonomous agent, defense at both layers is needed to break the chain.</p>

<h2 id="limits-and-counterpoints">Limits and Counterpoints</h2>

<p>To avoid overconfidence in this post’s conclusions, a few things should be made clear. First, the details of the incident are still being settled. Colorful details like the exact action count, the scope of credential theft, and the commercial-model refusal anecdote lean heavily on secondary reporting and must be distinguished from the confirmed facts of the official notice.</p>

<p>Second, our defensive narrative does not mean complete safety. Isolation and policy-and-audit are design principles that shrink the blast radius, not magic that eliminates the vulnerabilities themselves. Vulnerabilities like remote code execution in a dataset loader or injection in config parsing must continue to be found and patched at the code level, and isolation is the second line of defense that contains the damage when such a vulnerability fires.</p>

<p>Third, overrating autonomous-agent attacks is also risky. The root cause of this breach was not sophisticated AI but two familiar vulnerabilities that let input crossing a trust boundary execute code. The agent was merely the automation that exploited those vulnerabilities faster and wider. So the priority for response still lies in the fundamentals: separating untrusted input from execution privileges, detaching credentials from workloads, and making every execution observable.</p>

<p>Hugging Face’s swift containment and transparent disclosure will stand as a good response example. What remains our homework is simple: treat datasets as code rather than files, and make every agent action a subject of screening and audit.</p>

<h2 id="sources">Sources</h2>

<ul>
  <li><a href="https://huggingface.co/blog/security-incident-july-2026">Security incident disclosure, July 2026 (Hugging Face official blog)</a></li>
  <li><a href="https://www.helpnetsecurity.com/2026/07/20/hugging-face-breached-by-autonomous-ai-agent/">Hugging Face breached by autonomous AI agent (Help Net Security)</a></li>
  <li><a href="https://www.bleepingcomputer.com/news/security/hugging-face-breach-autonomous-ai-agent-system-internal-datasets-credentials/">Hugging Face warns an autonomous AI agent hacked its network (BleepingComputer)</a></li>
  <li><a href="https://thehackernews.com/2026/07/worlds-largest-ai-model-repository.html">World’s Largest AI Model Repository Hugging Face Breached by Autonomous AI Agent (The Hacker News)</a></li>
  <li>Secondary reporting (the exact action count and model-refusal anecdote are cited reporting, not confirmed fact): Cryptobriefing, Undercode Testing</li>
</ul>]]></content><author><name>{&quot;name&quot;=&gt;nil, &quot;avatar&quot;=&gt;nil, &quot;bio&quot;=&gt;nil, &quot;location&quot;=&gt;&quot;Seoul, Korea&quot;, &quot;email&quot;=&gt;&quot;info@thakicloud.co.kr&quot;, &quot;uri&quot;=&gt;nil, &quot;home&quot;=&gt;nil, &quot;links&quot;=&gt;[{&quot;label&quot;=&gt;&quot;Website&quot;, &quot;icon&quot;=&gt;&quot;fas fa-fw fa-link&quot;, &quot;url&quot;=&gt;&quot;https://thakicloud.co.kr&quot;}, {&quot;label&quot;=&gt;&quot;GitHub&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-github&quot;, &quot;url&quot;=&gt;&quot;https://github.com/thakicloud&quot;}, {&quot;label&quot;=&gt;&quot;LinkedIn&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-linkedin&quot;, &quot;url&quot;=&gt;&quot;https://www.linkedin.com/company/thakicloud&quot;}, {&quot;label&quot;=&gt;&quot;X&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-x-twitter&quot;, &quot;url&quot;=&gt;&quot;https://x.com/thakicloud&quot;}]}</name><email>info@thakicloud.co.kr</email></author><category term="news" /><category term="security" /><category term="huggingface" /><category term="ai-agent" /><category term="supply-chain" /><category term="sandbox" /><category term="dataset-security" /><category term="news" /><category term="thakicloud" /><summary type="html"><![CDATA[In July 2026 Hugging Face disclosed an internal breach driven by an autonomous AI agent. The entry point was a single malicious dataset, and two vulnerabilities in the dataset-processing pipeline led to code execution. We separate what is confirmed from what is still under investigation, and explain why dataset processing must be treated as a trust boundary.]]></summary></entry></feed>