Research Reality Cards
Dense arXiv abstracts distilled into 3-bullet builder takeaways. No hype, no padding.
Quantifying the Memorization-to-Generalization Transition: Scaling Laws and Phase Structure in Grokking
Not provided in the abstract
The study quantitatively characterizes the memorization-to-generalization boundary in hyperparameter space, revealing that data complexity is the dominant factor influencing the transition.
The power-law scaling relation for generalization onset time is given by: T_grok ∝ H^{-0.27} D^{-2.04} η^{-0.50} λ^{-0.64}, with R^2 = 0.732.
The study focuses on a specific architecture (two-hidden-layer MLPs) and modular arithmetic, which may limit generalizability to other architectures or tasks.
Automating Quadratic Unconstrained Binary Optimization (QUBO) Formulation Generation from Natural Language
Not provided
The proposed multi-agent framework achieves 68% accuracy in generating QUBO formulations, outperforming a baseline by 22%.
Achieved 68% accuracy on QUBOBench.
The framework's performance may depend on the quality of natural language input and the need for substantial domain expertise.
A Multi-Stage Rule-Chaining Framework for Compositional and Interpretable Cognitive Reasoning
The proposed framework achieved over 95% accuracy in cognitive generalization tasks by integrating three complementary solvers for rule discovery, pattern composition, and structural abstraction.
Training passed for 995 tasks out of 1000, achieving strong coverage across deterministic, compositional, and abstract categories.
The paper does not specify the exact conditions under which the framework was trained, which may affect reproducibility.
Understanding LoRA Rank Trade-offs in Diffusion Model Fine-Tuning
Author1, Author2, Author3 +2 more
Moderate LoRA ranks (specifically rank 4) are most efficient for diffusion model fine-tuning, achieving the best FID score with lower adaptation costs.
Rank 4 achieves the best DDPM FID score of 124.1380.
The study is limited to specific datasets and fixed optimization settings, which may not generalize to all scenarios.
Probabilistic Focal Search: Accelerating Bounded-Suboptimal Search via Lower-Bound Advancement
Not provided in the abstract
PFS can reduce node expansions by about 90% or more in scenarios where long f_min plateaus delay useful FOCAL admissions.
PFS outperforms traditional Focal Search (FS) in benchmarks like N-Puzzle and TSP, especially when FOCAL admission is a bottleneck.
The probabilistic factor's effectiveness is domain- and bound-dependent, which may affect reproducibility across different problem types.
Halo: Improving forecast accuracy through heteroscedastic estimation
Halo improves point estimate accuracy in forecasting by integrating a scale parameter estimation alongside the location parameter, achieving significant reductions in MSE and MAE across multiple models and markets.
Halo improves MSE by 2.6% to 16.5% and MAE by 1.7% to 11.0% in 28 of 30 model-market-metric comparisons.
The paper does not specify potential limitations regarding the reproducibility of the results across different datasets or forecasting scenarios.
GRADE: Single-Frame Generative Radar Depth Estimation Under Visual Degradation
Not provided in the content
GRADE achieves a mean absolute error (MAE) of 0.303 m in clear scenes and 0.313 m under smoke, outperforming existing baselines.
Trained and evaluated on ~95K frames across 12 buildings, achieving an MAE of 0.303 m in clear scenes.
The method's performance may vary significantly under different environmental conditions not covered in training.
MHE-Former: Multi-Hypothesis Transformers via Entropy Maximization for 3D Mesh Recovery
Not provided in the abstract
MHE-Former achieves state-of-the-art performance in accuracy and diversity for 3D mesh recovery by utilizing a multi-hypothesis approach and context-aware hypothesis selection.
The framework demonstrates significant improvements in both accuracy and diversity across multiple datasets.
The abstract does not specify any limitations or reproducibility concerns.
Think Before You Link: Rarity, Reasoning, and Retrieval in Multilingual Entity Linking
Not provided in the abstract
The proposed framework improves state-of-the-art accuracy by 6.9% overall and by up to 23.3% on rare-entity slices in multilingual entity linking.
The combination of reasoning and retrieval outperforms both methods individually, achieving a 6.9% overall improvement.
The paper does not specify the authors or provide detailed experimental setups, which may hinder reproducibility.
An Autonomous GeoAI Agent for Arctic Eco-Navigation
Samira At, Author 2, Author 3 +2 more
The framework integrates multiple criteria for safer and more socially responsible Arctic navigation while maintaining human control over value judgments.
The system employs a human-in-the-loop, multi-agent approach for eco-navigation, enhancing decision-making with ecological considerations.
The reliance on specialized agents for data acquisition may limit reproducibility in different contexts.
AcFlow: Controlling Text-to-Image Diffusion Transformers via Learned Conditional Activation Flow
Nove1yst, Author2, Author3 +2 more
AcFlow achieves the best style-content trade-off among evaluated baselines, significantly improving style alignment while allowing for fine-grained control over image generation.
AcFlow attains a style-content alignment score of 0.5365/0.2860, outperforming the best baseline score of 0.4397/0.2684.
The method may require extensive training data for optimal performance and may not generalize well to all unseen concepts.
Byzantine-Robust Federated Fire Detection with a Rotating Coordinator
Not specified in the provided content
The proposed semi-decentralized Byzantine-robust federated learning method achieves comparable accuracy and detection speed to fixed-server methods while eliminating single points of failure.
Model updates are compressed up to 10 times with only a small loss in balanced accuracy.
The reproducibility of the results may be limited by the availability of the curated indoor fire-detection dataset and the complexity of the semi-decentralized method.
Zero-shot rib design: merging training-free generative prior with topology optimization
Not specified in the provided content
The framework achieved statistically significant compliance reductions in structural designs by combining generative gradients with physics-based topology optimization.
38 of 49 prompt-domain combinations achieved compliance reductions up to -31.5% mechanical and -23.0% thermoelastic.
The paper does not specify the authors or provide detailed reproducibility guidelines for the methodology.
NCP-ArchPreview Technical Report: Moving towards Latent Space Language Models through Next Concept Prediction
Not specified in the provided content
NCP-ArchPreview achieves a significant performance improvement over OLMo-3-7B by utilizing only 51.3% of the total training tokens and 85% of the standard computation.
Outperforms OLMo-3-7B by 2.45 points on the downstream macro-average, with a notable 5.99-point gain on GSM8K.
The paper does not specify authors or provide detailed reproducibility instructions.
Subagents vs Agent Skills: Executing Reusable Knowledge for Long-Horizon Agentic Tasks
Author1, Author2, Author3 +2 more
Subagent execution outperforms agent-skill execution for long-horizon tasks when skill packages have clear input-output contracts and procedural knowledge.
Subagent execution shows improved performance compared to agent-skill execution, particularly as task horizons grow.
The additional communication overhead may complicate the implementation and evaluation of subagent execution.
OpenDiscoveryTrace: Process Traces for Evaluating AI Scientist Workflows
Not specified in the provided content
OpenDiscoveryTrace provides a comprehensive dataset of AI scientific agent trajectories that allows for the evaluation of reasoning processes, not just final outputs.
The dataset includes 558 complete AI scientific agent trajectories across various scientific tasks.
The paper does not specify authors or detailed methodologies, which may hinder reproducibility.
CMNIE: An Information Extraction Benchmark for Chinese Military News
Not specified in the provided content
CMNIE provides a comprehensive dataset with 13,000 instances and unified annotations for event triggers, arguments, entities, and relations, facilitating improved information extraction in Chinese military news.
The dataset includes manual annotations for 7 event types, 10 argument roles, 7 entity types, and 8 relation types.
The challenge of exact matching of event-argument spans and the performance of zero-shot LLMs in identifying relevant semantic units without matching gold span boundaries.
Adaptive Entangled Game Modules in Artificial General Intelligence
Not provided in the abstract
The integration of adaptive entangled game modules into AGI architectures can significantly enhance the efficiency and robustness of AI systems compared to conventional ANN-based approaches.
Adaptive entangled game modes explain 82-94% of observed decision patterns in stock market data.
The empirical analysis is based on specific stock market data, which may limit generalizability to other domains.
Gradland: On Phenomenal Experience, Differentiated Across Many Dimensions
Not provided
The paper demonstrates that the first-order structure of physical interactions, represented by gradients or Jacobians, can effectively account for various aspects of phenomenal experience.
Introduces two measures of Jacobian structure: effective rank and cohesion, based on Kirchhoff complexity.
The idealized nature of the model may limit its applicability to real-world scenarios.
Artificial Intelligence Algorithms for the Detection of Pathologies Related to Lung Cancer through Image Analysis using Convolutional Neural Networks and Data Augmentation: a systematic mapping of the literature
The use of AI and deep learning can significantly improve the accuracy and efficiency of lung cancer diagnosis through automated image analysis.
96 articles were reviewed, demonstrating high sensitivity and specificity in AI-based lung cancer detection.
Current limitations include the lack of standardized data and concerns regarding model explainability and patient privacy.
Multilingual in Name Only? Cultural and Linguistic Weaknesses of LLMs in Urdu
Not provided in the content
Current multilingual LLMs exhibit significant weaknesses in grammar, semantics, coherence, and cultural relevance when generating content in Urdu.
Generated a corpus of 93 Urdu stories using three contemporary LLMs and annotated errors under a nine-label taxonomy.
The study highlights that cultural and context errors largely remain unresolved despite few-shot prompting.
Meta-Learning for Data-Efficient Plant Growth Estimation via Vision Transformers and Fuzzy Clustering
Not provided in the abstract
The proposed framework significantly improves plant growth estimation performance using meta-learning techniques in scenarios with limited labeled data.
Second-order meta-learning methods like MAML++ outperform classical baselines in few-shot learning scenarios.
The impact of intra-cluster support selection is limited and varies depending on the dataset.
Rethinking Handwritten Character Recognition
Author1, Author2, Author3 +2 more
GraphemeNet outperforms published baselines on fourteen benchmarks across eight writing systems by effectively integrating stroke-level geometric regularity into its architecture.
Achieved state-of-the-art performance on fourteen benchmarks.
The architecture's reliance on script-specific geometric regularities may limit its applicability to scripts not represented in the training data.
M3-Former: Multimodal Transformer with Mixture-of-Experts for Long-Term Vessel Trajectory Prediction
Not provided in the abstract
M3-Former reduces Average Displacement Error (ADE) and Final Displacement Error (FDE) by 4.4% and 5.1%, respectively, compared to the strongest baseline in 4-hour prediction tasks.
Achieved a 4.4% reduction in Average Displacement Error (ADE) in long-term trajectory prediction.
The paper does not specify the authors or provide detailed implementation guidelines, which may hinder reproducibility.
Data-Efficient Language Modeling: From Frontier Advancement to Principle-Guided Model Improvement
Qiushi Engine
The research demonstrated that organizing experience around contextual dependencies significantly enhances data-efficient learning and model performance.
Overall performance improved from 42.02 to 42.25 across two generations.
The findings indicate that recovering familiar performance does not guarantee generalization to unseen inputs, raising concerns about reproducibility.
Osprey: Target-agnostic Pre-training Makes Stronger Drafters in Speculative Decoding
Author1, Author2, Author3 +2 more
Osprey improves mean acceptance length by 16.1% to 22.7% across different target models while increasing tokens per second by 17.5%.
Mean acceptance length improved by 22.7% for the 229B MiniMax-M2.5.
The method requires careful adaptation for each target, which may complicate reproducibility.
StochBench: A Domain-Specific Benchmark for Stochastic Processes in Lean
Not provided
StochBench introduces a comprehensive benchmark for stochastic processes that better represents domain-specific applied mathematics compared to existing benchmarks.
The Opus 4.8-based agent achieved a 34.9% proof rate (157/450) under a 15-minute per-problem limit.
The benchmark may not fully capture the complexity of all stochastic processes, potentially limiting its applicability.
Lensless Gaze Is Not Private by Default: Auditing Identity Leakage Across Disclosure Surfaces
Not specified in the provided content
The study demonstrates that privacy claims for lensless sensing must be empirically tested at disclosure boundaries, revealing significant identity leakage through various representations.
Simulated lensless measurements yield 96.7% top-1 identification accuracy.
The study focuses on a known-gallery closed-set identification protocol, which may not generalize to unknown or varying optical keys.
X-CoSD: Communication-Efficient Cross-Vocabulary Collaborative Speculative Decoding
Not provided in the abstract
X-CoSD and its enhanced variant X-CoSD-E significantly improve token generation speed while maintaining generation quality comparable to that of the server LLM.
X-CoSD reduces communication load by requiring distribution transmission only for the common-vocabulary region.
The paper does not specify the authors or provide detailed experimental setups, which may affect reproducibility.
AgenticGen: Reward-Guided Agentic Video Generation for Advertising
Not provided in the abstract
AgenticGen improves click-through rate (CTR) by 2.72%, conversion rate (CVR) by 2.63%, and advertising value (Advv) by 9.61% compared to the SFT baseline in online A/B tests.
Improved CTR by 2.72%, CVR by 2.63%, and Advv by 9.61% over the SFT baseline.
The specific authors and detailed methodology are not provided, which may hinder reproducibility.
M2LG-DG: A Multi-modal Local-Global Domain Generalization Framework for Cross-site Major Depressive Disorder Classification
Not provided in the abstract
M2LG-DG achieves an AUC of 69.48% for cross-site major depressive disorder classification, outperforming the closest comparison method by 2.18 percentage points.
Achieved an AUC of 69.48% on four held-out REST-meta-MDD sites.
The paper does not specify the authors or provide details on the reproducibility of the results.
Auditable Emergency Triage for Maternal and Newborn Care in India
Not provided in the abstract
The new triage system improved recall from 0.565 to 0.810 and F1 from 0.606 to 0.702 while allowing clinical experts to independently add rules without causing regressions.
The system triaged 152,421 patient queries and flagged 28,535 (18.7%) as emergencies.
The initial system was opaque, making it difficult to analyze mistakes and requiring costly evaluations for prompt changes.
When Do Options Help? Policy Necrosis and Redundant Coverage in Option-Critic
Not provided
The study reveals that adding more options does not improve performance due to issues like policy necrosis and ineffective termination rules.
The chance of all options failing in the same state drops from 59% to 4% with additional options.
The termination rule learned by option-critic contributes nothing to performance, leading to potential exploration issues.
Multi-granularity Adaptive Hypergraph Representation Learning via Granular-ball
Not provided
MGHRL significantly outperforms baseline models on benchmark datasets by effectively capturing high-order relationships through adaptive granular hypergraph generation.
MGHRL introduces an Adaptive Granular Hypergraph Generation strategy that captures high-order relationships based on the graph's topological structure.
The paper does not specify the reproducibility of the results across different types of graphs or datasets.
Capsule Lens: Locating and Tracking Concept Geometry in Model Representations
Not provided
Capsule Lens successfully matches concept geometry with a trackable geometric form, validated on held-out samples across various models.
The framework demonstrated the ability to locate concept geometry and analyze geometric characteristics across different models.
The study may face challenges in rigorous validation of the proposed hypotheses and the generalizability of findings across all model types.
AhaBench: Do Agents Learn from Prior Experience? A Benchmark for Long-Horizon Continual Learning
Not specified in the provided content
The study demonstrates that models that effectively utilize visible support do not necessarily exhibit the same learning improvements across different evaluation conditions.
Claude Opus 4.6 achieved an aggregate Post-Experience Score of 64.3 and a Learning Lift of +25.8.
The evaluation conditions may not fully capture the complexities of real-world scenarios, potentially affecting reproducibility.
AutoFyn Technical Report: Non-Parametric Expert Iteration for Long-Horizon Agents
Not specified in the provided content
AutoFyn outperformed existing coding agents in various domains, achieving higher scores in the 2026 International Mathematical Olympiad and building the top-ranked agent on the Spider 2.0 dbt benchmark.
Produced 16 maintainer-confirmed vulnerability advisories across multiple projects.
The reliance on explicit interfaces for persistent state may limit generalizability and reproducibility across different tasks.
Integrating Unimodal and Vision-Language Representations in Latent Space for Multi-Label Chest X-Ray Classification
Not provided
The proposed framework achieves a mean AUROC of 0.840 and an mAP of 0.467 by effectively integrating RAD-DINO and BioViL-T representations.
The best-performing model achieves a mean AUROC of 0.840.
The study has only been evaluated internally on MIMIC-CXR-JPG, raising concerns about generalizability to other healthcare data.
Damage-Aware Bandit Pruning for Vision and Language Transformers
Author1, Author2, Author3 +2 more
The proposed bandit pruning method reduces degradation in transformer models compared to traditional budgeted greedy selection methods across multiple datasets.
In 28 comparisons, 23 bootstrap confidence intervals exclude zero, indicating significant improvements.
The method's effectiveness may vary across different models and datasets, which could affect reproducibility.
SWORD: Wikidata-based Distortions Reveal Hidden Cross-Lingual Inconsistencies in LLM Factual Error Rejection
Not provided in the abstract
The SWORD benchmark reveals that LLMs achieve higher accuracy on semantically plausible distortions than on nonsensical substitutions, indicating a reliance on distributional familiarity over factual verification.
Cross-lingual performance gaps reach up to 28 percentage points in (East) Asian languages when models are presented with distorted statements.
The paper does not provide specific details on the reproducibility of the distortion generation process or the models tested.
CriticGen: Generation-Aware Evaluation as Actionable Feedback
Not provided in the abstract
CriticGen improves the quality of evaluation rubrics, leading to a 73.17% improvement in answer quality with a 93.28% non-degradation rate.
Achieves a Pearson correlation of 0.9556 and a Spearman correlation of 0.9560.
The abstract does not specify limitations or reproducibility concerns.
When Does Memory Help? A Cost-Aware Evaluation of Long-Term Memory in Tool-Using LLM Agents
Not specified in the provided content
The introduction of the MERIT benchmark demonstrates that memory significantly improves dependent-task success rates for LLM agents, achieving scores from 0.55 to 1.00.
Memory lifts dependent-task success from a leak-verified floor of 0.00 to 0.55-1.00 across 23,440 scored episodes.
The unpredictability of embedding retrieval performance and the variability in task success based on memory implementation may hinder reproducibility.
HB-PVI: A Hierarchical Bayesian Personalization and Value-of-Information Framework for Complex Activity Recognition
Not provided in the abstract
HB-PVI achieved utility-optimal performance in 199 of 216 cost-threshold settings, advocating for a population-first deployment policy when personalization gains are minimal.
One-step EVSI was zero at every decision state, leading to a 100% reduction in labeling while maintaining a posterior mean F1 loss of 0.00217.
The study's findings may be limited by the specific cohort size (47 participants) and the context of the MUSIC-CAR dataset.
Evidence-Order Calibration for Selective Visual Reasoning under Progressive Loss of Question-Critical Evidence
Not provided in the abstract
The study demonstrates that adding evidence-order supervision to a binary cross-entropy loss reduces the evidence monotonicity violation rate (EMVR) from 0.330 to 0.303.
The evidence-order supervision approach reduced EMVR by 0.027, indicating improved reliability in visual reasoning.
The results on AUROC, Brier, and AURC differences are statistically inconclusive, raising concerns about the reproducibility of the findings.
Beyond Right and Wrong: Evaluating Second-order Social Reasoning in Large Language Models
Not provided in the abstract
The study reveals that current LLMs overpredict negative sanctions in social norm violations compared to human judgments, particularly as social distance increases.
The release of the NormReact dataset, which includes 450 norm violation scenarios annotated for emotions and behavioral responses.
The findings suggest that LLMs may misrepresent social regulation, indicating a need for further evaluation in norm-sensitive applications.
The microscope is the mask: privileged views and labels from a cryo-ET forward model
Not provided in the content
The CARNIVAL model outperforms a state-of-the-art model by utilizing forward model-based paired views and privileged information for protein annotation tasks.
CARNIVAL was evaluated without finetuning and showed improved performance on classification and detection tasks in real tomograms.
The reliance on simulated data and specific forward model assumptions may limit generalizability to all real-world scenarios.
MedProb: Probing Internal Representations of Vision-Language Models for Medical Question Answering
Not provided in the abstract
MedProb recovers more answer-relevant signals than prompting and outperforms medical VLMs and agentic systems in multiple-choice Med-VQA tasks.
MedProb shows improved performance across multiple datasets, recovering substantially more relevant signals than traditional prompting methods.
The study does not consistently demonstrate that medical adaptation improves linear decodability across all tested models.
Adapting from Downturns: Prediction of Long-Term Conversational-Skill Development in Mental-Health Crisis Counselors
Author1, Author2, Author3 +2 more
The study demonstrates that identifying and analyzing how counselors adapt to challenging conversational moments can predict their long-term improvement in steering conversations toward positive outcomes.
The counselor-adaptation approach yields better results than baseline methods that rely solely on conversation transcripts.
The prediction task is challenging and may require extensive longitudinal data for accurate results.
How Much Does Corpus Choice Change Dependency-Distance Estimates?
Not provided in the content
Dependency-distance estimates vary significantly across different treebanks, with treebank choice accounting for approximately 29% of variance in estimates.
Substituting one treebank for another reversed nearly 40% of pairwise language orderings.
The study indicates that cross-treebank agreement is moderate at best, raising concerns about the reliability of dependency-distance estimates derived from single corpora.
Memory as transformation: LETHE, a self-referential gan-inspired architecture
LETHE successfully implements a self-referential system for audio processing that evolves parameters without external datasets or supervision.
The active generator is necessary for parametric evolution, with a consistent result of Δc_{22}=0.000 across all 15 ablation sessions.
The system's reliance on a closed configuration and lack of external datasets may limit reproducibility and generalization.
Evidence Integration in Large Language Models
Not specified in the provided content
The study presents a distributional theory of evidence integration in LLMs, demonstrating that evidence can shift the distribution of initial answers based on receiver properties rather than just trust in the evidence source.
Confirmed predictions over ten million trials across twelve LLMs from four families and eight domains.
The paper does not specify the authors, which may limit reproducibility and further exploration of the findings.
FailSAE: Towards Interpretable Failure Prediction for Vision-Language Models via Sparse Autoencoders
Not provided
The proposed framework outperforms existing methods in failure prediction while providing improved interpretability.
The framework shows improved performance in failure prediction compared to evaluated baselines.
The paper does not specify the authors or provide detailed experimental setups, which may hinder reproducibility.
Step Back to Move Forward: Reflection-Aware Preference Optimization for Visual Generation
Not provided in the abstract
RA-GRPO significantly improves the alignment of generative models with human preferences, particularly in mitigating reward hacking and enhancing generalization.
RA-GRPO outperforms existing methods in T2I and T2V tasks, demonstrating improved semantic faithfulness and visual realism.
The paper does not specify potential limitations or reproducibility concerns.
Quantum-Assisted Memory-Efficient Training for Parameter-Intensive Wi-Fi-Based Human Activity Recognition
Not provided in the abstract
Q-MET achieves a 90% to 95% reduction in trainable parameters compared to conventional deep learning training while maintaining or exceeding classification accuracy.
Achieves 75% to 85% model sparsity with less than 2% loss in classification accuracy.
The abstract does not provide specific details on the reproducibility of the quantum-assisted approach.
A Data Fusion Framework for Grounding Aerospace Surrogate Model via Experimental Wind-Tunnel Observations
Not provided in the abstract
The proposed correction framework allows a CFD-trained deep learning surrogate to adapt using limited experimental data, significantly improving agreement with experimental measurements without retraining the surrogate.
At Mach 0.85, the grounded surrogate achieves agreement with measurements within 2.3-2.7% of the measured Cp range.
The method relies on a limited experimental dataset, which may affect the generalizability of the results.
Spectral-Target Physical Latent Structuring for JEPA-Style World Models
Not provided in the abstract
The introduction of a Fourier auxiliary head significantly improves planning success rates in dynamic environments by enforcing physically-informed structuring of the latent space.
The auxiliary head leads to substantial improvements in planning success rates, particularly in dynamic environments, and is especially impactful in low-data regimes.
The paper does not specify the authors or provide detailed experimental setups, which may hinder reproducibility.
ProToMEx: Rapid, Interpretable Explanations via Structured Representations
Not provided
ProToMEx achieves explanations of comparable fidelity to SHAP and LIME while being 30-40x faster in generating local explanations.
ProToMEx is ~30-40x faster than SHAP and LIME over standardised tabular datasets.
The paper does not specify limitations or reproducibility concerns.
Joint Alignment and Distillation for Video Generation via Sample-Guided Distribution Matching
Not specified in the provided content
The proposed DM-Align framework synergizes gradient directions to enhance distillation quality and preference alignment without the need for multi-step reward evaluation.
The framework consistently outperforms standalone variants and sequential two-stage pipelines in comprehensive experiments.
The paper does not specify authors or detailed experimental setups, which may hinder reproducibility.
From Matching Models to Recruiting Agents: A Systematized Narrative Review of AI Recruitment Systems, Evaluation, and Governance
Not specified in the provided content
The paper introduces a staged mapping from evaluation evidence to the strongest defensible claim, emphasizing the need for evidence-grounded and auditable recruitment systems.
Analyzed 40 representative works and identified persistent gaps in behavioral labels and privacy evaluation.
Behavioral labels confound exposure, preference, and qualification, limiting external validity.
Harbor Adapters and Harbor-Index: Infrastructure and a Curated Meta-Dataset for Large-Scale Agentic Evaluation
Not specified in the provided content
The authors developed a unified evaluation infrastructure that ports over 80 benchmarks for agent evaluation and introduced a curated set of 82 high-quality tasks for comprehensive agentic evaluation.
Conducted a large-scale evaluation of 8 models across 54 benchmarks, with the strongest model achieving a 28.0% pass rate.
No evaluated model-harness configuration exceeds a 30% pass rate, indicating potential challenges in achieving reliable performance.