arXivLabs
PulseAugur coverage of arXivLabs — every cluster mentioning arXivLabs across labs, papers, and developer communities, ranked by signal.
25 day(s) with sentiment data
-
New paper proposes 'horizon residual' to analyze AI agent failure on long tasks
A new paper proposes a method called "trajectory-induced degradation" to better understand why AI agents fail on long-horizon tasks. The authors argue that existing benchmarks don't sufficiently explain failure modes, w…
-
New research identifies "visual lock-in" in vision-language models
Researchers have identified a phenomenon called "visual lock-in" in vision-language models, where outdated verbal descriptions can lead to incorrect decisions about current scenes. This occurs when changes in the model'…
-
Study finds overall ability, not discrete skills, dominates student math performance
A new study utilizing Bernoulli Mixture Models on a large dataset of 119,034 students in the United Kingdom has investigated the assumptions behind personalized learning systems regarding mathematical competence. The re…
-
New paper offers practical lessons for building self-serve entity resolution systems
A new paper details practical lessons learned from building a self-serve entity resolution (ER) system. The research highlights that no single matching algorithm is universally effective, recommending a pipeline that tr…
-
New CoRAS method optimizes image sensing with adaptive rate control
Researchers have introduced Conformalized Rate-Adaptive Sensing (CoRAS), a novel method designed to optimize the collection of measurements for high-resolution imaging systems. CoRAS adaptively determines the acquisitio…
-
Extended Reality proposed as mediation layer for human-robot control
A new paper proposes Extended Reality (XR) as a crucial mediation layer for situated human control in human-robot teaming. The authors argue that XR can enhance a human collaborator's ability to understand, shape, autho…
-
New SHarD harness distributes AI coding agent security controls
A new paper introduces SHarD (Secure Harness Distribution), a system designed to distribute security controls for AI coding agents. The research demonstrates that common security measures like OS sandboxing, skill scann…
-
Research report details noise-shaped coefficients in discrete polynomial Fourier extension
This research report delves into the intricacies of noise-shaped one-bit coefficients within normalized discrete polynomial Fourier extension. It explores error analysis for first-order Sigma-Delta quantization, establi…
-
Temporal Convolutional Networks for Trajectory Inpainting
Researchers have developed a Temporal Convolutional Network (TCN) designed to reconstruct missing segments in trajectory data. This model utilizes symmetric dilation, allowing it to consider both past and future observa…
-
Paper details best practices for ML-driven geospatial map production
A new paper outlines best practices for creating large-scale geospatial map products using machine learning and Earth observation data. The paper addresses challenges in the end-to-end pipeline, from data preprocessing …
-
arXiv report details peer review data nuances and best practices
A new report published on arXiv examines three key nuances within peer review data: paper versions, score versions, and input formats. The study characterizes the differences among these variants and quantifies their im…
-
New co-learning framework tackles missing data in multi-modal classification
Researchers have developed a novel co-learning framework to address the challenge of missing modalities in multi-modal classification tasks. This framework is designed to handle situations where any subset of data modal…
-
New research probes LLM context attribution reliability with overlapping knowledge
A new research paper introduces an evaluation protocol and benchmark dataset to assess the reliability of context attribution methods for large language models (LLMs). The study highlights that current attribution metho…
-
New CallBench benchmark evaluates dual-goal coordination in phone call assistants
Researchers have introduced CallBench, a new benchmark designed to evaluate the dual-goal coordination capabilities of phone call assistants. This benchmark comprises 50,000 multi-turn dialogues across six scenarios, in…
-
New research explores learning distributions from multiple data providers
A new research paper published on arXiv introduces a theoretical framework for learning distributions from multiple, potentially overlapping data providers. The study focuses on a stylized model where a learner aims to …
-
New grayscale level set framework speeds up image segmentation
A new grayscale level set framework for image segmentation has been developed, addressing challenges in segmenting images with multiple degradations. This framework theoretically demonstrates that length regularization …
-
New variational model enhances weak-boundary image segmentation
Researchers have developed a new variational model for image segmentation, specifically designed to handle homogeneous structures with weak or ambiguous boundaries. This model, based on the Cahn-Hilliard equation, integ…
-
Paper explores emergent behavior in financial markets using formal methods
A new paper explores emergent behavior in financial markets, drawing parallels between complex systems and the formal methods community. The research identifies and structures sources of complexity within electronic fin…
-
New algorithm speeds up real-time crowd tracking by 6.5x
Researchers have developed a new incremental assignment algorithm designed to improve the efficiency of real-time crowd tracking. This algorithm exploits the block-sparse structure of cost matrices in dense crowd scenar…
-
New benchmark assesses citation trustworthiness in AI-generated legal reports
Researchers have developed LegalCiteTrust, a new benchmark designed to evaluate the trustworthiness of citations within Chinese long-form legal research reports. This benchmark assesses reports across three dimensions: …