文章背景与核心概要
在实际的网络安全应用中,传统方法在评估“持续学习框架”(即允许大语言模型通过记忆或检索进行自我改进而无需重新训练的系统)时往往捉襟见肘。标准基准测试通常存在陈旧、缺乏代表性或资源稀缺的问题,导致从业者很难判断某个学习框架是否真正有效。
本文引入了一种基于扩展假设的新型评估框架。作者没有依赖静态标签,而是利用更强大的“教师”模型为配备了学习框架的“学生”模型提供稀疏采样的纠正。随后,根据学生模型随时间向教师模型性能靠拢的有效性来对该框架进行评分。研究人员证明,这种“相对教师的提升”可以作为真实性能改进的可靠代理指标,即使在没有标注标准答案(gold standards)的情况下依然适用。
Evaluating Agentic Learning Harness Capabilities Without Labels via the Scaling Hypothesis
arXiv: 2608.13608
Authors: Aryan Luthra, Kshitij Jain, Siddharth Arya, Bobby Filar, Anna Bertiger
Published: August 11, 2026
Venue: Accepted at CAMLIS 2025 (Conference on Applied Machine Learning for Information Security)
arXiv: 2608.13608
Authors: Aryan Luthra, Kshitij Jain, Siddharth Arya, Bobby Filar, Anna Bertiger
Published: August 11, 2026
Venue: Accepted at CAMLIS 2025 (Conference on Applied Machine Learning for Information Security)
Summary
In operational cybersecurity, traditional methods for evaluating "Continual Learning Harnesses"—systems that allow LLMs to improve via memory or retrieval without retraining—often fail. Standard benchmarks are frequently stale, unrepresentative, or scarce, making it difficult for practitioners to determine if a harness is actually effective.
This paper introduces a novel evaluation framework grounded in the scaling hypothesis. Instead of relying on static labels, the authors use a stronger "teacher" model to provide sparsely sampled corrections to a "student" model equipped with a learning harness. The harness is then scored based on how effectively the student converges toward the teacher’s performance over time. The researchers demonstrate that this "teacher-relative lift" serves as a reliable proxy for true performance improvements, even in the absence of labeled gold standards.
Summary
In operational cybersecurity, traditional methods for evaluating "Continual Learning Harnesses"—systems that allow LLMs to improve via memory or retrieval without retraining—often fail. Standard benchmarks are frequently stale, unrepresentative, or scarce, making it difficult for practitioners to determine if a harness is actually effective.
This paper introduces a novel evaluation framework grounded in the scaling hypothesis. Instead of relying on static labels, the authors use a stronger "teacher" model to provide sparsely sampled corrections to a "student" model equipped with a learning harness. The harness is then scored based on how effectively the student converges toward the teacher’s performance over time. The researchers demonstrate that this "teacher-relative lift" serves as a reliable proxy for true performance improvements, even in the absence of labeled gold standards.
Key Findings
- Evaluation Framework: The proposed method allows for end-to-end evaluation of learning harnesses without requiring labeled benchmarks.
- Validation: Improvement relative to a teacher model correlates strongly with improvement on held-out gold standard datasets, validating the approach.
- Limitations of Current Methods: The study confirms that using "LLM-as-a-judge" between models of similar power provides no usable signal for evaluating performance gains.
- Practical Implications: These results suggest that if human experts provide high-precision, sparse corrections, a teacher-sized model can be effectively improved using the same harness architecture.
Key Findings
- Evaluation Framework: The proposed method allows for end-to-end evaluation of learning harnesses without requiring labeled benchmarks.
- Validation: Improvement relative to a teacher model correlates strongly with improvement on held-out gold standard datasets, validating the approach.
- Limitations of Current Methods: The study confirms that using "LLM-as-a-judge" between models of similar power provides no usable signal for evaluating performance gains.
- Practical Implications: These results suggest that if human experts provide high-precision, sparse corrections, a teacher-sized model can be effectively improved using the same harness architecture.
Access & Resources
- Full Paper: View PDF
- Experimental HTML: View HTML
- TeX Source: Download Source
- License: Creative Commons Attribution 4.0 International

Access & Resources
- Full Paper: View PDF
- Experimental HTML: View HTML
- TeX Source: Download Source
- License: Creative Commons Attribution 4.0 International
Metadata
- Subjects: Artificial Intelligence (cs.AI); Cryptography and Security (cs.CR); Machine Learning (cs.LG)
- DOI: https://doi.org/10.48550/arXiv.2608.13608
Metadata
- Subjects: Artificial Intelligence (cs.AI); Cryptography and Security (cs.CR); Machine Learning (cs.LG)
- DOI: https://doi.org/10.48550/arXiv.2608.13608