Executive Summary & Core Abstract
The advent of self-improving Large Language Model (LLM) agents presents a fundamental challenge to auditability and regulatory compliance within sensitive domains, particularly credit pipelines where adaptive rule changes must be transparent and traceable. Ravil Akhtyamov's seminal work, "The Harness as the Only Mutable Surface," posits a novel scientific discovery: self-evolution in LLM agents is reviewable and compliant only if its mutability is strictly confined to the runtime harness—encompassing instruction text, tool-call logic, and primitive composition—while the foundational model weights remain immutable. This mechanism ensures that every agentic adaptation is a discrete, auditable differential change, intrinsically linked to a defined cause and a verifiable test, thereby preserving the artefact a supervisor reviews. This approach directly addresses the issue where an agent rewriting itself otherwise obliterates the necessary audit trail for regulatory scrutiny.
Structured Overview
-
Fundamental Scientific Discovery: The core insight lies in the principle of "harness-bounded self-evolution," which posits that LLM agents can adapt to evolving operational rules without sacrificing accountability by restricting their self-modification capabilities to the declarative runtime environment. This framework facilitates the creation of a dual-loop engine featuring an admission gate, which records a hash-chained audit trail prior to any deployment. The underlying mechanism is a constrained evolutionary process that allows for dynamic adaptation (e.g., to supervisory re-interpretations) within a defined, auditable surface, thereby maintaining reviewability and compliance by ensuring changes are treated as explicit, tested diffs rather than opaque, unversioned alterations to the underlying model.
-
Empirical Benchmark & Technical Breakthrough: Through rigorous simulation, the proposed gated dual-loop engine demonstrated significant improvements in maintaining compliance and performance stability. Across three families of supervisory re-interpretation at varying severities, with 10 seeds each, the admission gate successfully filtered candidate changes, admitting only 144 out of 7,449 proposed modifications. Crucially, none of these admitted changes worsened error on held-out historical data. The system demonstrably restored the false-positive rate to the oracle level, without introducing an increase in missed flags, across all low- and mid-severity scenarios. In stark contrast, an unbounded system, employing a simpler check based on recent trace errors, admitted 309 harmful changes and resulted in missed flags exceeding 10% in 49 of 90 runs, indicating a severe loosening of screening criteria. Furthermore, the gate consistently rejected every candidate when evaluated against pre-shift labels, reinforcing that rule re-interpretation must be explicitly encoded as a relabeling of historical data for proper processing.
-
Global Significance & Practical Takeaway: This research offers a critical pathway toward the responsible and compliant deployment of self-evolving AI agents, particularly in high-stakes environments such as credit scoring. By mapping its mechanisms to provisions like the EU AI Act's requirements for high-risk systems and highlighting the current exclusion of agentic AI from April 2026 US model-risk guidance, the study underscores an urgent need for robust regulatory frameworks. The practical takeaway is profound: future AI systems designed for autonomous adaptation must incorporate explicit, auditable constraints on their self-modification capabilities, ensuring transparency and accountability. The unambiguous directional trend established is a secular deceleration of harmful, unreviewed agentic changes and an acceleration of auditable, compliant adaptations, thereby enhancing the trustworthiness and deployability of advanced AI within regulated sectors.
Theoretical Foundation & Governing Principles
The burgeoning capabilities of large language model (LLM) agents to self-improve and adapt to evolving operational contexts present a profound challenge to established paradigms of oversight and regulatory compliance, particularly within high-stakes environments such as credit pipelines. Traditional software development and deployment mandates that every modification to a system artefact be reviewable, traceable to a cause, accompanied by tests, and subject to formal approval. Unbounded self-evolution, however, wherein an agent directly modifies its underlying model weights or core architectural primitives without external auditable diffs, fundamentally eradicates this required transparency, rendering the system unsuitable for regulated domains.
Ravil Akhtyamov's seminal work addresses this critical bottleneck by positing that self-evolution can only be reconciled with compliance if it is strictly confined to the "runtime harness." This foundational theoretical model introduces the harness as the *only mutable surface*, defining it as the agent's externalized instruction text, its tool-call logic, and the composition of its operational primitives. Crucially, the underlying LLM model weights ($W$) remain fixed throughout this process. This architectural constraint transforms opaque, internal weight adjustments into explicit, auditable modifications ($\Delta H$) of the harness configuration, where each $\Delta H$ can be reviewed and formally tested.
The governing mechanism is a dual-loop engine incorporating a "measured admission gate" ($G$). This gate acts as a robust control layer, evaluating every candidate harness modification proposed by the agent. A modification $H_{t+1}' = H_t + \Delta H$ is subjected to rigorous scrutiny against held-out historical data and pre-defined compliance thresholds. The gate's decision function, $G(\Delta H, E_{t}', E_{t+1}', C)$, where $E$ represents error metrics and $C$ denotes compliance criteria, determines acceptance. The core mathematical principle enforced by the gate is the maintenance or restoration of critical performance metrics without detrimental side effects. Specifically, the system aims to ensure that the false-positive rate (FPR) of the agent's decisions is restored to an oracle level ($FPR_{oracle}$) without an unacceptable increase in missed flags (false negatives, $FNR_{baseline}$).
$$ \text{If } G(\Delta H, E_t, E_{t+1}', C) = \text{ACCEPT, then } H_{t+1} = H_t + \Delta H \text{ such that } FPR(H_{t+1}) \le FPR_{oracle} \land FNR(H_{t+1}) \le FNR_{baseline} $$
$$ \text{Else } H_{t+1} = H_t $$
Upon acceptance, a hash-chained record of the new harness state is immutably written before deployment, providing an unbroken audit trail. This rigorously contrasts with unbounded systems, which, relying merely on recent trace performance, may admit changes that ultimately degrade performance (e.g., 309 harmful changes in simulation, as observed in the comparative analysis) and undermine regulatory adherence. The gate's ability to reject candidates evaluated on pre-shift labels further underlines its role in enforcing a clear delineation between system adaptation and compliance integrity, directly mapping to provisions within the EU AI Act for high-risk credit scoring and addressing limitations in current model-risk guidance.
Empirical Findings & Research Attribution
Empirical Validation of Compliance-Bounded Self-Evolution
Empirical investigations into compliance-bounded self-evolution of LLM agents within credit pipelines reveal significant advantages of confining adaptive changes to the runtime harness, mediated by a measured admission gate. The theoretical premise, which posits that self-evolution is reviewable and controllable only if model weights remain fixed and adaptations are presented as diffs with clear causal links and tests, is strongly supported by simulation results. A dual-loop engine incorporating such an admission gate was rigorously evaluated in a simulated environment, utilizing a simulated agent and a seeded-search proposer rather than full language models to isolate the mechanism's core properties.
Across three distinct families of supervisory re-interpretation, each tested at three severity levels with 10 seeds per cell, the gated loop demonstrated robust control over agentic changes. Specifically, the admission gate permitted 144 out of 7,449 candidate changes. Crucially, none of these admitted changes worsened error on held-out historical data. Furthermore, the system successfully restored the false-positive rate to its oracle level in every low- and mid-severity scenario, without concurrently raising the incidence of missed flags. This empirically validates the gate's capacity to facilitate beneficial adaptations while maintaining critical performance thresholds.
In stark contrast, a comparative evaluation replacing the admission gate with a less stringent check—akin to those in unbounded systems that only consider fewer errors visible in recent traces—yielded detrimental outcomes. This unbounded approach admitted 309 harmful changes and resulted in missed flags remaining above 10% in 49 of 90 runs. The observed reduction in false positives under this unbounded condition was directly attributable to a loosening of the screening mechanism, confirming that less rigorous admission criteria lead to the propagation of undesirable agent behaviors. When evaluated against pre-shift labels, the gate consistently rejected every candidate, underscoring the necessity for re-interpretations to be formally encoded as rules that relabel historical data. While parametric and scope shifts were repaired locally, structural shifts necessitated primitive replacement, with the gate's fixed tolerance blocking the correct replacement in half the seeds at the highest structural severity.
Lead Authors & Principal Investigators: Ravil Akhtyamov
Primary University/Institute affiliations: Academic Research Consortium
Publishing Journal or Venue: arXiv Preprint Repository (Category: quant-ph/physics, 2610.10629)
Canonical Link / DOI: arXiv:2610.10629
Experimental/Computational Methodology: The research employed a dual-loop engine designed to confine LLM agent self-evolution to a runtime harness, with adaptations screened by an admission gate that writes a hash-chained record prior to deployment. The system was evaluated in simulation using a simulated agent and a seeded-search proposer, rather than full language models. Experiments involved three families of supervisory re-interpretation, each tested at three severities, with 10 seeds per condition. Performance was measured by the number of admitted candidate changes, error rates on held-out history, and false-positive/missed-flag rates. A comparative baseline involved replacing the gate with a less rigorous check based on recent trace errors.
Key Scientific Insights & Future Horizons
Core Takeaways
- Fundamental Mechanism: The seminal contribution of Akhtyamov (2026) establishes a novel paradigm for agentic AI evolution, positing the "harness"—comprising instruction text, tool-call logic, and primitive composition—as the sole mutable surface for self-improving LLM agents within high-stakes environments like credit pipelines. This dual-loop engine, fortified by a measured admission gate that mandates hash-chained recordkeeping prior to deployment, ensures that every adaptation is a reviewable, auditable diff with a clear causal link and an attached test, critically preserving artifact integrity against the destructive nature of unbounded self-rewriting. By confining self-evolution, the system enables controlled adaptation without sacrificing transparency or the ability for supervisors to review changes.
- Real-World Value: This controlled evolutionary framework directly addresses the pressing need for transparent, auditable, and compliant AI systems in regulated industries. By strictly confining agentic adaptation, the mechanism enables practical self-improvement while maintaining regulatory adherence, particularly pertinent to provisions of the EU AI Act concerning high-risk credit scoring and highlighting the current exclusion of agentic AI from the April 2026 US model-risk guidance. It offers a verifiable path to restore critical metrics like false-positive rates to oracle levels without compromising missed flags, as evidenced by simulation admitting 144 safe changes out of 7,449 candidates compared to 309 harmful ones admitted by unbounded systems.
Applications & Future Outlook
The proposed compliance-bounded self-evolution mechanism holds profound implications for financial services, regulatory technology (RegTech), and any domain where AI adaptability must converge with stringent governance. In credit underwriting, it enables LLM agents to autonomously adjust to evolving supervisory interpretations or market shifts while guaranteeing full auditability and accountability, mitigating systemic risk by ensuring every change is defensible. Scientifically, it opens avenues for developing "governed AI" frameworks, extending to other high-consequence applications such as medical diagnostics, autonomous systems, or critical infrastructure management, where explainability and provable safety of change are paramount. Future research must rigorously explore the gate's performance under extreme structural shifts where its fixed tolerance currently impedes correct primitive replacement, as observed in half the high-severity structural cases. Further challenges include scaling the "seeded-search proposer" to complex, real-world harness modifications, developing automated methods for encoding supervisory re-interpretations into actionable rule updates that relabel history, and thoroughly validating the framework with actual language models rather than simulated agents to fully characterize their inherent biases and emergent properties within this constrained evolutionary path.
Akhtyamov, R. (2026). The Harness as the Only Mutable Surface: Compliance-Bounded Self-Evolution of LLM Agents in Credit Pipelines, with a Measured Admission Gate. arXiv Preprint Repository. arXiv:2610.10629.
सार संक्षेप एवं कार्यकारी सारांश
स्व-सुधार करने वाले वृहद भाषा मॉडल (एलएलएम) एजेंटों का आगमन संवेदनशील क्षेत्रों में, विशेषकर ऋण प्रणालियों में, जहाँ अनुकूली नियम परिवर्तन पारदर्शी और पता लगाने योग्य होने चाहिए, अंकेक्षणीयता और विनियामक अनुपालन के समक्ष एक मूलभूत चुनौती प्रस्तुत करता है। रविल अख्त्यामौव का प्रतिष्ठित कार्य, "हारनेस ही एकमात्र परिवर्तनीय सतह है" (The Harness as the Only Mutable Surface), एक नवीन वैज्ञानिक खोज का प्रतिपादन करता है: एलएलएम एजेंटों में स्व-विकास की समीक्षा तभी की जा सकती है और वह तभी अनुपालक होगा जब उसकी परिवर्तनीयता रनटाइम हारनेस—जिसमें निर्देश पाठ, टूल-कॉल तर्क और प्रारंभिक संघटन शामिल हैं—तक ही सख्ती से सीमित हो, जबकि मूलभूत मॉडल भार अपरिवर्तनीय रहते हैं। यह तंत्र सुनिश्चित करता है कि प्रत्येक एजेंटिक अनुकूलन एक पृथक, अंकेक्षणीय विभेदी परिवर्तन है, जो एक परिभाषित कारण और एक सत्यापन योग्य परीक्षण से आंतरिक रूप से जुड़ा हुआ है, जिससे उस कलाकृति को संरक्षित रखा जा सके जिसकी कोई पर्यवेक्षक समीक्षा करता है। यह दृष्टिकोण सीधे उस समस्या का समाधान करता है जहाँ एक एजेंट स्वयं को फिर से लिखकर विनियामक जांच के लिए आवश्यक अंकेक्षण निशान को अन्यथा मिटा देता है।
संरचित अवलोकन
-
मौलिक वैज्ञानिक खोज: मूल अंतर्दृष्टि "हारनेस-बद्ध स्व-विकास" के सिद्धांत में निहित है, जो यह प्रतिपादित करता है कि एलएलएम एजेंट अपनी स्व-संशोधन क्षमताओं को घोषणात्मक रनटाइम परिवेश तक सीमित करके जवाबदेही का त्याग किए बिना विकसित हो रहे परिचालन नियमों के अनुकूल हो सकते हैं। यह ढांचा एक दोहरी-लूप इंजन के निर्माण को सुगम बनाता है, जिसमें एक प्रवेश द्वार होता है, जो किसी भी परिनियोजन से पहले एक हैश-श्रृंखलित अंकेक्षण निशान को रिकॉर्ड करता है। अंतर्निहित तंत्र एक संयमित विकासवादी प्रक्रिया है जो एक परिभाषित, अंकेक्षणीय सतह के भीतर गतिशील अनुकूलन (उदाहरण के लिए, पर्यवेक्षी पुनर्व्याख्याओं के लिए) की अनुमति देती है, जिससे समीक्षा क्षमता और अनुपालन बना रहता है, यह सुनिश्चित करके कि परिवर्तनों को अंतर्निहित मॉडल में अपारदर्शी, असंस्मरणीकृत परिवर्तनों के बजाय स्पष्ट, परीक्षित अंतर के रूप में माना जाता है।
-
अनुभवजन्य बेंचमार्क एवं तकनीकी सफलता: कठोर अनुकरण के माध्यम से, प्रस्तावित गेटेड दोहरी-लूप इंजन ने अनुपालन और प्रदर्शन स्थिरता बनाए रखने में महत्वपूर्ण सुधार प्रदर्शित किए। विभिन्न तीव्रताओं पर पर्यवेक्षी पुनर्व्याख्या की तीन श्रेणियों में, प्रत्येक में 10 सीड्स के साथ, प्रवेश द्वार ने उम्मीदवार परिवर्तनों को सफलतापूर्वक फ़िल्टर किया, और 7,449 प्रस्तावित संशोधनों में से केवल 144 को ही स्वीकार किया। महत्वपूर्ण रूप से, इनमें से किसी भी स्वीकृत परिवर्तन ने रखे गए ऐतिहासिक डेटा पर त्रुटि को खराब नहीं किया। प्रणाली ने स्पष्ट रूप से गलत-सकारात्मक दर को ऑरेकल स्तर तक बहाल कर दिया, सभी कम- और मध्यम-तीव्रता वाले परिदृश्यों में छूटे हुए फ़्लैग में वृद्धि किए बिना। इसके विपरीत, एक अबाधित प्रणाली, जो हाल की ट्रेस त्रुटियों पर आधारित एक सरल जांच का उपयोग करती है, ने 309 हानिकारक परिवर्तनों को स्वीकार किया और 90 में से 49 परीक्षणों में छूटे हुए फ़्लैग 10% से अधिक हो गए, जो स्क्रीनिंग मानदंडों में गंभीर ढिलाई का संकेत देता है। इसके अतिरिक्त, द्वार ने प्रत्येक उम्मीदवार को लगातार अस्वीकार कर दिया जब पूर्व-शिफ्ट लेबलों के विरुद्ध मूल्यांकन किया गया, इस बात पर बल देते हुए कि उचित प्रसंस्करण के लिए नियम पुनर्व्याख्या को ऐतिहासिक डेटा के पुनर-लेबलिंग के रूप में स्पष्ट रूप से एन्कोड किया जाना चाहिए।
-
वैश्विक महत्व एवं व्यावहारिक निष्कर्ष: यह शोध स्व-विकसित एआई एजेंटों के जिम्मेदार और अनुपालक परिनियोजन की दिशा में एक महत्वपूर्ण मार्ग प्रदान करता है, विशेषकर क्रेडिट स्कोरिंग जैसे उच्च जोखिम वाले वातावरण में। यूरोपीय संघ एआई अधिनियम की उच्च-जोखिम प्रणालियों के लिए आवश्यकताओं जैसे प्रावधानों के साथ इसके तंत्रों का मानचित्रण करके और अप्रैल 2026 के यूएस मॉडल-जोखिम मार्गदर्शन (US model-risk guidance) से एजेंटिक एआई के वर्तमान बहिष्कार को उजागर करके, यह अध्ययन मजबूत विनियामक ढाँचों की तत्काल आवश्यकता को रेखांकित करता है। व्यावहारिक निष्कर्ष गहरा है: स्वायत्त अनुकूलन के लिए डिज़ाइन किए गए भविष्य के एआई सिस्टम को अपनी स्व-संशोधन क्षमताओं पर स्पष्ट, अंकेक्षणीय बाधाओं को शामिल करना चाहिए, जिससे पारदर्शिता और जवाबदेही सुनिश्चित हो। स्थापित अस्पष्ट दिशात्मक प्रवृत्ति हानिकारक, असंमीक्षित एजेंटिक परिवर्तनों का एक धर्मनिरपेक्ष मंदन और अंकेक्षणीय, अनुपालक अनुकूलन का त्वरण है, जिससे विनियमित क्षेत्रों के भीतर उन्नत एआई की विश्वसनीयता और परिनियोजन क्षमता बढ़ती है।
सैद्धांतिक आधार एवं मूलभूत वैज्ञानिक सिद्धांत
वृहद भाषा मॉडल (LLM) एजेंटों की स्वतः सुधार करने और विकसित हो रहे परिचालन संदर्भों के अनुरूप ढलने की बढ़ती क्षमताएँ, विशेषकर क्रेडिट पाइपलाइनों जैसे उच्च-दांव वाले वातावरणों में, पर्यवेक्षण और नियामक अनुपालन के स्थापित प्रतिमानों के समक्ष एक गहन चुनौती प्रस्तुत करती हैं। पारंपरिक सॉफ्टवेयर विकास और परिनियोजन यह अनिवार्य करते हैं कि किसी प्रणाली कलाकृति में प्रत्येक संशोधन की समीक्षा की जा सके, उसके कारण का पता लगाया जा सके, परीक्षणों के साथ हो, और औपचारिक अनुमोदन के अधीन हो। हालाँकि, असीमित स्वतः-विकास, जिसमें एक एजेंट बाहरी लेखापरीक्षण योग्य अंतर के बिना अपने अंतर्निहित मॉडल भार या मुख्य स्थापत्य संबंधी आदिमों को सीधे संशोधित करता है, इस आवश्यक पारदर्शिता को मौलिक रूप से समाप्त कर देता है, जिससे प्रणाली विनियमित क्षेत्रों के लिए अनुपयुक्त हो जाती है।
रविल अख्तियामोव का मूलभूत कार्य इस महत्वपूर्ण अवरोध को यह प्रतिपादित करके संबोधित करता है कि स्वतः-विकास का अनुपालन के साथ सामंजस्य तभी स्थापित किया जा सकता है जब इसे "रनटाइम हार्नेस" तक कड़ाई से सीमित रखा जाए। यह मूलभूत सैद्धांतिक मॉडल हार्नेस को एकमात्र परिवर्तनशील सतह के रूप में प्रस्तुत करता है, इसे एजेंट के बाह्यीकृत अनुदेश पाठ, इसकी उपकरण-कॉल तर्क, और इसके परिचालन आदिमों की संरचना के रूप में परिभाषित करता है। महत्वपूर्ण रूप से, अंतर्निहित LLM मॉडल भार ($W$) इस पूरी प्रक्रिया के दौरान स्थिर रहते हैं। यह स्थापत्य संबंधी बाधा अपारदर्शी, आंतरिक भार समायोजन को हार्नेस विन्यास के स्पष्ट, लेखापरीक्षण योग्य संशोधनों ($\Delta H$) में परिवर्तित करती है, जहाँ प्रत्येक $\Delta H$ की समीक्षा और औपचारिक रूप से परीक्षण किया जा सकता है।
नियंत्रक क्रियाविधि एक दोहरी-लूप इंजन है जिसमें एक "मापा गया प्रवेश द्वार" ($G$) शामिल है। यह द्वार एक सुदृढ़ नियंत्रण परत के रूप में कार्य करता है, जो एजेंट द्वारा प्रस्तावित प्रत्येक प्रत्याशी हार्नेस संशोधन का मूल्यांकन करता है। एक संशोधन $H_{t+1}' = H_t + \Delta H$ को संचित ऐतिहासिक डेटा और पूर्व-निर्धारित अनुपालन सीमाओं के विरुद्ध कठोर जाँच के अधीन किया जाता है। द्वार का निर्णय फलन, $G(\Delta H, E_{t}', E_{t+1}', C)$, जहाँ $E$ त्रुटि मेट्रिक्स का प्रतिनिधित्व करता है और $C$ अनुपालन मानदंडों को दर्शाता है, स्वीकृति निर्धारित करता है। द्वार द्वारा प्रवर्तित मूल गणितीय सिद्धांत हानिकारक दुष्प्रभावों के बिना महत्वपूर्ण प्रदर्शन मेट्रिक्स का रखरखाव या पुनर्स्थापन है। विशेष रूप से, प्रणाली का उद्देश्य यह सुनिश्चित करना है कि एजेंट के निर्णयों की मिथ्या-सकारात्मक दर (FPR) एक ओरेकल स्तर ($FPR_{oracle}$) पर बहाल हो, छूटे हुए फ्लैग (मिथ्या-नकारात्मक, $FNR_{baseline}$) में अस्वीकार्य वृद्धि के बिना।
$$ \text{If } G(\Delta H, E_t, E_{t+1}', C) = \text{ACCEPT, then } H_{t+1} = H_t + \Delta H \text{ such that } FPR(H_{t+1}) \le FPR_{oracle} \land FNR(H_{t+1}) \le FNR_{baseline} $$
$$ \text{Else } H_{t+1} = H_t $$
स्वीकृति पर, नई हार्नेस स्थिति का एक हैश-श्रृंखलाबद्ध अभिलेख परिनियोजन से पहले अपरिवर्तनीय रूप से लिखा जाता है, जो एक अटूट लेखापरीक्षण मार्ग प्रदान करता है। यह असीमित प्रणालियों के साथ कठोरता से विपरीत है, जो केवल हाल के ट्रेस प्रदर्शन पर निर्भर करते हुए, ऐसे परिवर्तनों को स्वीकार कर सकते हैं जो अंततः प्रदर्शन को कम करते हैं (उदाहरण के लिए, सिमुलेशन में 309 हानिकारक परिवर्तन, जैसा कि तुलनात्मक विश्लेषण में देखा गया) और नियामक अनुपालन को कमजोर करते हैं। पूर्व-शिफ्ट लेबलों पर मूल्यांकन किए गए प्रत्याशियों को अस्वीकार करने की द्वार की क्षमता, प्रणाली अनुकूलन और अनुपालन अखंडता के बीच एक स्पष्ट सीमांकन लागू करने में इसकी भूमिका को और रेखांकित करती है, जो उच्च-जोखिम वाले क्रेडिट स्कोरिंग के लिए EU AI अधिनियम के प्रावधानों के साथ सीधे मेल खाती है और वर्तमान मॉडल-जोखिम मार्गदर्शन में सीमाओं को संबोधित करती है।
अनुसंधान के मुख्य निष्कर्ष एवं प्रामाणिक संदर्भ
अनुपालन-परिसीमित स्वतः-विकास का आनुभविक सत्यापन
क्रेडिट पाइपलाइनों के भीतर एलएलएम एजेंटों के अनुपालन-परिसीमित स्वतः-विकास की आनुभविक जाँचें, रनटाइम हार्नेस तक अनुकूली परिवर्तनों को सीमित करने के महत्वपूर्ण लाभों को उजागर करती हैं, जो एक मापित प्रवेश द्वार द्वारा मध्यस्थ होते हैं। सैद्धांतिक आधारशिला, जो यह मानती है कि स्वतः-विकास की समीक्षा और उसे नियंत्रित किया जा सकता है केवल तभी जब मॉडल भार स्थिर रहें और अनुकूलन को स्पष्ट कारणात्मक संबंध और परीक्षणों के साथ भिन्नता के रूप में प्रस्तुत किया जाए, इसका सिमुलेशन परिणामों द्वारा दृढ़ता से समर्थन किया जाता है। ऐसे प्रवेश द्वार को समाहित करने वाले दोहरे-लूप इंजन का सिमुलेटेड वातावरण में कठोरता से मूल्यांकन किया गया, तंत्र के मूल गुणों को पृथक करने के लिए पूर्ण भाषा मॉडल के बजाय एक सिमुलेटेड एजेंट और एक सीडेड-सर्च प्रस्तावक का उपयोग करते हुए।
पर्यवेक्षी पुनः-व्याख्या के तीन भिन्न परिवारों में, प्रत्येक को तीन गंभीरता स्तरों पर, प्रति सेल 10 सीड के साथ परीक्षण किया गया, गेटेड लूप ने एजेंटिक परिवर्तनों पर सुदृढ़ नियंत्रण प्रदर्शित किया। विशेष रूप से, प्रवेश द्वार ने `7,449 उम्मीदवार परिवर्तनों में से 144 को` अनुमति दी। निर्णायक रूप से, इनमें से किसी भी स्वीकृत परिवर्तन ने रोक कर रखे गए ऐतिहासिक डेटा पर त्रुटि को बदतर नहीं किया। इसके अतिरिक्त, प्रणाली ने निम्न- और मध्यम-गंभीरता वाले प्रत्येक परिदृश्य में गलत-सकारात्मक दर को उसके ओरेकल स्तर पर सफलतापूर्वक बहाल किया, साथ ही छूटे हुए फ़्लैग की घटनाओं को बढ़ाए बिना। यह आनुभविक रूप से गेट की लाभकारी अनुकूलन को सुगम बनाने की क्षमता को सत्यापित करता है, जबकि महत्वपूर्ण प्रदर्शन थ्रेशोल्ड को बनाए रखता है।
इसके विपरीत, प्रवेश द्वार को एक कम कड़े चेक—जो अबाधित प्रणालियों में उन चेकों के समान है जो केवल हालिया ट्रेस में दिखाई देने वाली कम त्रुटियों पर विचार करते हैं—से बदलने वाले एक तुलनात्मक मूल्यांकन ने हानिकारक परिणाम दिए। इस अबाधित दृष्टिकोण ने `309 हानिकारक परिवर्तनों को` स्वीकार किया और इसके परिणामस्वरूप `90 में से 49 रनों` में छूटे हुए फ़्लैग `10% से ऊपर` रहे। इस अबाधित स्थिति के तहत गलत-सकारात्मक में देखी गई कमी सीधे स्क्रीनिंग तंत्र के ढीले पड़ने के कारण थी, इस बात की पुष्टि करती है कि कम कड़े प्रवेश मानदंड अवांछनीय एजेंट व्यवहारों के प्रसार की ओर ले जाते हैं। प्री-शिफ्ट लेबलों के विरुद्ध मूल्यांकन करने पर, गेट ने लगातार हर उम्मीदवार को अस्वीकार कर दिया, जो ऐतिहासिक डेटा को पुनः लेबल करने वाले नियमों के रूप में पुनः-व्याख्याओं को औपचारिक रूप से एन्कोड करने की आवश्यकता पर बल देता है। जबकि पैरामीट्रिक और स्कोप शिफ्ट को स्थानीय रूप से ठीक किया गया, संरचनात्मक शिफ्टों के लिए आदिम प्रतिस्थापन की आवश्यकता पड़ी, जिसमें गेट की निश्चित सहिष्णुता ने उच्चतम संरचनात्मक गंभीरता पर आधे बीजों में सही प्रतिस्थापन को अवरुद्ध कर दिया।
प्रमुख लेखक एवं प्रधान अन्वेषक: रविल अख्त्यामौव
प्राथमिक विश्वविद्यालय/संस्थान संबद्धताएँ: अकादमिक अनुसंधान संघ
प्रकाशन पत्रिका या स्थान: arXiv प्रीप्रिंट रिपोजिटरी (श्रेणी: quant-ph/physics, 2610.10629)
प्रामाणिक लिंक / डीओआई: arXiv:2610.10629
प्रायोगिक/अभिकलनात्मक कार्यप्रणाली: इस शोध में एक दोहरे-लूप इंजन का उपयोग किया गया था जिसे एलएलएम एजेंट के स्वतः-विकास को रनटाइम हार्नेस तक सीमित करने के लिए डिज़ाइन किया गया था, जिसमें अनुकूलनों को एक प्रवेश द्वार द्वारा जांचा जाता है, जो परिनियोजन से पहले एक हैश-श्रृंखलित रिकॉर्ड लिखता है। प्रणाली का सिमुलेशन में मूल्यांकन एक सिमुलेटेड एजेंट और एक सीडेड-सर्च प्रस्तावक का उपयोग करते हुए किया गया, पूर्ण भाषा मॉडल के बजाय। प्रयोगों में पर्यवेक्षी पुनः-व्याख्या के तीन परिवार शामिल थे, प्रत्येक को तीन गंभीरताओं पर परीक्षण किया गया, प्रति स्थिति 10 सीड के साथ। प्रदर्शन को स्वीकृत उम्मीदवार परिवर्तनों की संख्या, रोक कर रखे गए इतिहास पर त्रुटि दरों, और गलत-सकारात्मक/छूटे हुए-फ़्लैग दरों से मापा गया। एक तुलनात्मक बेसलाइन में गेट को हालिया ट्रेस त्रुटियों पर आधारित एक कम कड़े चेक से बदलना शामिल था।
प्रमुख वैज्ञानिक निष्कर्ष एवं भावी संभावनाएँ
प्रमुख निष्कर्ष
- मौलिक तंत्र: अख्तियामोव (2026) का मूलभूत योगदान अभिकर्मी एआई के विकास के लिए एक नवीन प्रतिमान स्थापित करता है, जिसमें 'हार्नेस' (harness)—निर्देश पाठ, उपकरण-कॉल तर्क और प्रारंभिक संघटन से युक्त—को क्रेडिट पाइपलाइनों जैसे उच्च-दांव वाले वातावरण में स्व-सुधार करने वाले एलएलएम अभिकर्मियों के लिए एकमात्र परिवर्तनीय सतह के रूप में प्रस्तुत किया गया है। यह दोहरी-लूप इंजन, एक मापा प्रवेश द्वार द्वारा सुदृढ़ है जो परिनियोजन से पहले हैश-श्रृंखलित रिकॉर्डकीपिंग अनिवार्य करता है, यह सुनिश्चित करता है कि प्रत्येक अनुकूलन एक समीक्षणीय, लेखापरीक्षणीय अंतर हो जिसमें एक स्पष्ट कार्य-कारण संबंध और एक संलग्न परीक्षण हो, जिससे अप्रतिबंधित स्व-पुनर्लेखन की विनाशकारी प्रकृति के विरुद्ध कलाकृति की अखंडता गंभीर रूप से संरक्षित रहती है। स्व-विकास को सीमित करके, यह प्रणाली पारदर्शिता या पर्यवेक्षकों द्वारा परिवर्तनों की समीक्षा करने की क्षमता का त्याग किए बिना नियंत्रित अनुकूलन को सक्षम बनाती है।
- वास्तविक-विश्व महत्व: यह नियंत्रित विकासवादी ढाँचा विनियमित उद्योगों में पारदर्शी, लेखापरीक्षणीय और अनुपालक एआई प्रणालियों की तत्काल आवश्यकता को सीधे संबोधित करता है। अभिकर्मी अनुकूलन को सख्ती से सीमित करके, यह क्रियाविधि नियामक अनुपालन बनाए रखते हुए व्यावहारिक स्व-सुधार को सक्षम बनाती है, जो विशेष रूप से उच्च जोखिम वाले क्रेडिट स्कोरिंग से संबंधित यूरोपीय संघ के एआई अधिनियम के प्रावधानों के लिए प्रासंगिक है और अप्रैल 2026 के अमेरिकी मॉडल-जोखिम दिशानिर्देशों से अभिकर्मी एआई के वर्तमान बहिष्करण को उजागर करती है। यह मिथ्या-सकारात्मक दरों जैसे महत्वपूर्ण मेट्रिक्स को ओरैकल स्तरों तक बहाल करने के लिए एक सत्यापन योग्य मार्ग प्रदान करता है, बिना छूटे हुए फ़्लैग्स से समझौता किए, जैसा कि सिमुलेशन द्वारा प्रमाणित है, जिसमें 7,449 उम्मीदवारों में से 144 सुरक्षित परिवर्तनों को स्वीकार किया गया, जबकि अप्रतिबंधित प्रणालियों द्वारा 309 हानिकारक परिवर्तनों को स्वीकार किया गया।
अनुप्रयोग एवं भविष्य की संभावनाएँ
प्रस्तावित अनुपालन-बाधित स्व-विकास क्रियाविधि के वित्तीय सेवाओं, नियामक प्रौद्योगिकी (रेगटेक) और किसी भी ऐसे क्षेत्र के लिए गहन निहितार्थ हैं जहाँ एआई अनुकूलनशीलता को कठोर शासन के साथ संरेखित होना चाहिए। क्रेडिट अंडरराइटिंग में, यह एलएलएम अभिकर्मियों को विकसित होती पर्यवेक्षी व्याख्याओं या बाजार परिवर्तनों के अनुरूप स्वायत्त रूप से समायोजित करने में सक्षम बनाता है, जबकि पूर्ण लेखापरीक्षण क्षमता और जवाबदेही की गारंटी देता है, जिससे यह सुनिश्चित करके प्रणालीगत जोखिम कम होता है कि प्रत्येक परिवर्तन रक्षा योग्य है। वैज्ञानिक रूप से, यह 'शासित एआई' ढाँचों के विकास के लिए मार्ग प्रशस्त करता है, जिसे चिकित्सा निदान, स्वायत्त प्रणालियों या महत्वपूर्ण अवसंरचना प्रबंधन जैसे अन्य उच्च-परिणामी अनुप्रयोगों तक विस्तारित किया जा सकता है, जहाँ व्याख्यात्मकता और परिवर्तन की सिद्ध सुरक्षा सर्वोपरि है। भविष्य के शोध को अत्यधिक संरचनात्मक परिवर्तनों के तहत गेट के प्रदर्शन का कठोरता से अन्वेषण करना चाहिए, जहाँ इसकी निश्चित सहनशीलता वर्तमान में सही प्रारंभिक प्रतिस्थापन को बाधित करती है, जैसा कि आधे उच्च-गंभीरता वाले संरचनात्मक मामलों में देखा गया है। आगे की चुनौतियों में 'सीडेड-सर्च प्रपोज़र' को जटिल, वास्तविक-विश्व हार्नेस संशोधनों तक स्केल करना, पर्यवेक्षी पुनर्-व्याख्याओं को कार्रवाई योग्य नियम अपडेट्स में एन्कोड करने के लिए स्वचालित विधियों का विकास करना जो इतिहास को पुनः लेबल करते हैं, और इस विवश विकासवादी मार्ग के भीतर उनके अंतर्निहित पूर्वाग्रहों और उद्गामी गुणों को पूरी तरह से चित्रित करने के लिए सिम्युलेटेड अभिकर्मियों के बजाय वास्तविक भाषा मॉडल के साथ ढाँचे का पूरी तरह से सत्यापन करना शामिल है।
अख्तियामोव, आर. (2026). द हार्नेस ऐज़ द ओनली म्यूटैबल सरफेस: कंप्लायंस-बाउंडेड सेल्फ-इवोल्यूशन ऑफ एलएलएम एजेंट्स इन क्रेडिट पाइपलाइंस, विथ अ मेज़र्ड एडमिशन गेट. arXiv प्रीप्रिंट रिपॉजिटरी. arXiv:2610.10629.
💬 Comments