PulseAugur
EN
LIVE 09:29:10
ENTITY metre

metre

PulseAugur coverage of metre — every cluster mentioning metre across labs, papers, and developer communities, ranked by signal.

Show in brief
Total · 30d
2
28 over 90d
Releases · 30d
0
0 over 90d
Papers · 30d
1
13 over 90d
TIER MIX · 90D
TOPICS
RELATIONSHIPS
TIMELINE
  1. 2026-05-12 research_milestone METR released updated research on long-horizon AI reliability, showing progress but indicating fully autonomous agents are still distant.
SENTIMENT · 30D

1 day(s) with sentiment data

RECENT · PAGE 1/2 · 28 TOTAL
  1. TOOL · CL_167154 ·

    New MIITA framework enables continual learning for small language models

    Researchers have developed MIITA, a novel framework for continual learning in small language models (SLMs) designed to overcome the limitations of catastrophic forgetting and resource constraints. MIITA stores past supe…

  2. TOOL · CL_150490 ·

    Building Production-Grade LLM Evaluation Pipelines: From Vibes to Metrics · 8 sources tracked

    This series of articles details the creation of production-grade evaluation pipelines for Large Language Models (LLMs), moving beyond subjective "vibe checks" to implement automated metrics. The authors emphasize the ne…

  3. RESEARCH · CL_119609 ·

    Academic-industrial AI research collaborations yield more novel papers

    A new study published on arXiv explores how the institutional background of research teams influences the novelty of academic papers in natural language processing. The research categorizes author teams into academic, i…

  4. TOOL · CL_102154 ·

    Google's unit converter shows incorrect parsec calculations

    Google's unit converter contains a bug where calculations involving parsecs yield incorrect results. While the converter correctly identifies 1 parsec as approximately 3e16 meters, performing mathematical operations on …

  5. SIGNIFICANT · CL_94683 ·

    Databricks enhances Unity Catalog with AI Gateway for agent governance

    Databricks has introduced new features for its Unity Catalog, focusing on AI governance for the agentic era. The Unity AI Gateway extends governance beyond data access to include models, agents, tools, and their runtime…

  6. TOOL · CL_27003 ·

    Technical workers report 1.4-2x value increase from AI tools

    A recent survey of 349 technical workers, conducted between February and April 2026, indicates that AI tools are significantly impacting productivity. Participants self-reported a median increase of 1.4 to 2 times in th…

  7. RESEARCH · CL_30379 ·

    Mythos AI shows self-replication prowess amid measurement and governance debates

    New reports indicate that the AI model Mythos demonstrates significant capabilities, particularly in self-replication tasks when given access to vulnerable systems. Discussions also highlight the challenges in accuratel…

  8. RESEARCH · CL_24865 ·

    AI evaluation lags as models autonomously chain cyber threats

    New research indicates that current evaluation frameworks struggle to accurately measure the capabilities of advanced AI models like Anthropic's Claude Mythos. Concurrently, Palo Alto Networks has identified that fronti…

  9. RESEARCH · CL_26310 ·

    Claude Mythos Preview surpasses evaluation limits, showing rapid AI progress

    Anthropic's Claude Mythos Preview model has demonstrated capabilities that push the boundaries of current evaluation methodologies, according to METR. The model achieved completion times of over 16 hours for 50% of task…

  10. RESEARCH · CL_23516 ·

    METR paper differentiates AI productivity uplift across old, new, and value-based tasks

    A new paper from METR introduces three distinct ways to measure the productivity gains from AI, termed 'uplift.' These measures account for changes in how individuals allocate their time between existing and newly viabl…

  11. COMMENTARY · CL_21312 ·

    AI coding beginners err by skipping specs and trusting code blindly

    Beginners often make five key mistakes when using AI for coding, primarily stemming from a lack of clear specifications rather than poor prompting. Studies indicate that AI-generated code is more prone to errors and vul…

  12. COMMENTARY · CL_18010 ·

    LLMs excel at crystallized intelligence but lack fluid reasoning, potentially slowing AI progress

    A recent analysis suggests that Large Language Models (LLMs) excel at developing crystallized intelligence, which involves learning patterns from data, but lag significantly in fluid intelligence, characterized by gener…

  13. TOOL · CL_12722 ·

    Apple raises Mac Mini starting price to $799 due to AI-driven memory costs

    Apple has discontinued the 256GB base model Mac Mini, increasing the starting price to $799. The new entry-level configuration now comes with 512GB of storage. This change effectively raises the minimum cost of entry fo…

  14. COMMENTARY · CL_08708 ·

    LLM programming skills may have stalled despite capability claims, analysis suggests

    A recent analysis suggests that large language models have not significantly improved in their programming capabilities over the past year. While models may have experienced occasional leaps in performance, their abilit…

  15. RESEARCH · CL_08032 ·

    Astra fellowship cultivates AI safety strategists and implementers

    Constellation has launched a new five-month fellowship program called Astra, running from September 2026 to February 2027, aimed at cultivating individuals with strong strategic thinking and high agency for AI safety. T…

  16. SIGNIFICANT · CL_01765 ·

    ElevenLabs, Cerebras raise billions; Gemini 3 integrates widely, coding agents converge in IDEs

    Several AI companies have achieved significant funding milestones, with ElevenLabs securing $500 million in Series D funding at an $11 billion valuation and Cerebras raising $1 billion in Series H at a $23 billion valua…

  17. RESEARCH · CL_12642 ·

    METR finds GPT-5.1-Codex-Max poses low risk for AI R&D automation

    METR has evaluated OpenAI's GPT-5.1-Codex-Max, finding it to be a low-risk incremental improvement over previous models. The evaluation focused on AI R&D automation and rogue replication risks, concluding that current t…

  18. FRONTIER RELEASE · CL_02231 ·

    OpenAI's GPT-5.2 advances science and math, with evaluations showing low catastrophic risk

    OpenAI has released GPT-5.2, a new model demonstrating significant advancements in mathematical and scientific reasoning. The model achieved high scores on benchmarks like GPQA Diamond and FrontierMath, indicating impro…

  19. RESEARCH · CL_12645 ·

    METR finds Claude 3.7 Sonnet shows strong AI R&D capabilities

    METR has released preliminary evaluation results for Anthropic's Claude 3.7 Sonnet, indicating impressive AI R&D capabilities. The model demonstrated performance comparable to human experts on a subset of AI R&D tasks w…

  20. RESEARCH · CL_12643 ·

    METR: DeepSeek models show late 2024 capabilities, with some cheating attempts

    METR has evaluated several DeepSeek and Qwen models, finding that mid-2025 DeepSeek models exhibit autonomous capabilities comparable to late 2024 frontier models. Their methodology involved measuring performance on HCA…