Announcements tracker
What tech authorities announced
1,064 announcements from official sources, each one linking straight to the document that issued it. We record what was announced and who announced it. We do not rewrite it.
- arXiv — cs.AI preprintsInternational5 Oct 2026
MintFlow: Minimal Trajectory Intervention for Constrained Flow Matching
arXiv:2610.02260v1 Announce Type: new Abstract: Flow matching models excel at generative modeling, and many downstream applications require their samples to satisfy prescribed constraints, such as observed measurements…
- arXiv — cs.AI preprintsInternational5 Oct 2026
Fast Models, Slow Evidence: A Paired and Self-Audited Evaluation of System-1 Decision Models for LLM Agent Harnesses
arXiv:2610.02267v1 Announce Type: new Abstract: Agent harnesses make many small, typed decisions per task: which model to call, which tool to use, whether retrieved text is relevant, whether an input carries an…
- arXiv — cs.AI preprintsInternational5 Oct 2026
The AI Risk Observatory: What Can We Learn from AI Disclosures in Annual Reports About Societal Resilience?
arXiv:2610.02281v1 Announce Type: new Abstract: Societal resilience research relies on access to useful and actionable data, which motivates our main research question: Can annual reports, processed at scale with LLMs,…
- arXiv — cs.AI preprintsInternational5 Oct 2026
Keep It CALM: Analyzing the Limits of Global Unsafety in Text-to-Image Generation
arXiv:2610.02300v1 Announce Type: new Abstract: Training-free safeguards for text-to-image generation often rely on a reusable safety signal, such as an unsafe direction or global toxic subspace, applied broadly across…
- arXiv — cs.AI preprintsInternational5 Oct 2026
Choosing Before Acting: Comparative Value Estimation for Long-Horizon Tool-Use Agents
arXiv:2610.02330v1 Announce Type: new Abstract: Large language models (LLMs) rely on long-horizon tool invocation sequences for complex tasks, where each invocation can alter the task state and condition subsequent…
- arXiv — cs.AI preprintsInternational5 Oct 2026
World Editing: Intervening on Executable Worlds at Increasing Depth
arXiv:2610.02331v1 Announce Type: new Abstract: Interactive world models are increasingly capable of generating environments and acting within them, yet deliberately editing an existing executable world remains…
- arXiv — cs.AI preprintsInternational5 Oct 2026
A Multi Method Importance and Performance Efficiency Analysis of Topological Metrics for Natural Visibility Graph Based Cyber Attack Detection
arXiv:2610.02342v1 Announce Type: new Abstract: Natural Visibility Graph (NVG) based analysis characterizes network traffic through topological descriptors reflecting different structural properties. However, not all…
- arXiv — cs.AI preprintsInternational5 Oct 2026
DeReAct: Decomposed Reasoning and Acting for Reliable AI Agents
arXiv:2610.02351v1 Announce Type: new Abstract: ReAct-based agents typically rely on a single LLM policy to propose actions, interact with the environment, and decide when a task is complete. This coupling makes action…
- arXiv — cs.AI preprintsInternational5 Oct 2026
Traversing the Satisfaction-Diversity Frontier in Text-to-Image Diffusion
arXiv:2610.02372v1 Announce Type: new Abstract: Text-to-image generation enables users to explore several images generated from the same prompt. For these generated images to be useful, each one must reflect the user's…
- arXiv — cs.AI preprintsInternational5 Oct 2026
THPL: A Vision-to-Language Decision Support Framework for Rainbow Trout Feeding Management in RAS
arXiv:2610.02378v1 Announce Type: new Abstract: In Recirculating Aquaculture Systems (RAS), precision feeding is critical for minimizing costs and improving fish welfare. However, existing methods lack cognitive…
- arXiv — cs.AI preprintsInternational5 Oct 2026
FlashSinkhorn 2: Block-Sparse Entropic Optimal Transport
arXiv:2610.02395v1 Announce Type: new Abstract: Streaming GPU solvers for entropic optimal transport (EOT), such as FlashSinkhorn, avoid storing the dense kernel but still evaluate all $n\times m$ point pairs in every…
- arXiv — cs.AI preprintsInternational5 Oct 2026
When Terminal-Agent Training Stalls: Demystifying Data Generation and Verification Challenge
arXiv:2610.02405v1 Announce Type: new Abstract: Using a frontier model like Claude Opus as a meta-agent to generate terminal tasks and verifiers for RL training is increasingly common. Yet a runnable Docker image and…
- arXiv — cs.AI preprintsInternational5 Oct 2026
Reinforcement Learning Techniques for the Optimization of Target Polarization in Nuclear Physics Scattering Experiments
arXiv:2610.02452v1 Announce Type: new Abstract: The operation of dynamically polarized targets in nuclear physics experiments relies on continuous tuning of the microwave frequency to compensate for radiation damage and…
- arXiv — cs.AI preprintsInternational5 Oct 2026
Tropical Reinforcement Learning
arXiv:2610.02478v1 Announce Type: new Abstract: Reinforcement learning for large language models typically maximizes expected return, adding up the probabilities of all successful trajectories. However, the classical…
- arXiv — cs.AI preprintsInternational5 Oct 2026
MEA: A Reward-Driven Multi-Agent System for Faithful Model Explanations
arXiv:2610.02480v1 Announce Type: new Abstract: Recent years have seen the employment of a plethora of machine learning (ML) models in high-stakes domains, but they remain largely opaque to the practitioners who act on…
- arXiv — cs.AI preprintsInternational5 Oct 2026
What Does a Token Cost? A Mixture-of-Agents Measurement of Sufficient Per-Token Compute
arXiv:2610.02491v1 Announce Type: new Abstract: Large language models spend the same amount of computation on every token they generate, regardless of how difficult each token is to produce. Methods such as speculative…
- arXiv — cs.AI preprintsInternational5 Oct 2026
Right Order, Wrong Scale: Auditing LLM Judges for Occupational AI Measurement
arXiv:2610.02492v1 Announce Type: new Abstract: LLM judges are increasingly used to assess whether AI outputs meet workplace requirements, but agreement on response rankings does not establish agreement on acceptance…
- arXiv — cs.AI preprintsInternational5 Oct 2026
"I just assumed that it would translate": examining MT risk awareness among healthcare staff with abbreviations as a use case
arXiv:2610.02496v1 Announce Type: new Abstract: In the UK, public healthcare staff report turning to machine translation (MT) - predominantly Google Translate (GT) - to communicate with patients across language…
- arXiv — cs.AI preprintsInternational5 Oct 2026
HXAI: Hierarchical Privacy-Preserving Explainable AI in Distributed Energy Systems
arXiv:2610.02504v1 Announce Type: new Abstract: Balancing electricity demand and supply is increasingly difficult due to the inherent intermittency of renewable power generation and the stochastic power consumption.…
- arXiv — cs.AI preprintsInternational5 Oct 2026
World Action Modeling with Progressive Visual Planning
arXiv:2610.02508v1 Announce Type: new Abstract: World action models (WAMs) have emerged as a promising paradigm for robotic control by jointly predicting future visual dynamics and actions from an initial observation…
- arXiv — cs.AI preprintsInternational5 Oct 2026
On-Premises Multi-Course RAG Tutoring for Business Education: Hardware-Software Trade-offs in a Campus AI Tutor
arXiv:2610.02510v1 Announce Type: new Abstract: Campus AI tutors based on retrieval-augmented generation (RAG) must ground answers in assigned course materials while keeping textbooks and student dialogue on…
- arXiv — cs.AI preprintsInternational5 Oct 2026
Hypothesis-guided discovery of cognitive algorithms via program refinement
arXiv:2610.02523v1 Announce Type: new Abstract: Developing cognitive models of algorithmic reasoning from behavioral data is a central problem in cognitive science that challenges current methods. Traditional approaches…
- arXiv — cs.AI preprintsInternational5 Oct 2026
Learning What to Investigate Next: Meta-Reasoning for Long-Horizon Research Agents
arXiv:2610.02525v1 Announce Type: new Abstract: Long-horizon research agents must decide both how to investigate and what to investigate next as evidence accumulates. This is hard to learn because such decisions are…
- arXiv — cs.AI preprintsInternational5 Oct 2026
How To Train Your World Model: Fine-tuning vs RAG for LM-based World Modeling
arXiv:2610.02542v1 Announce Type: new Abstract: World models (WMs) simulate the transition dynamics of environments, enabling agents to plan over the consequences of their actions. In text-based environments,…
- arXiv — cs.AI preprintsInternational5 Oct 2026
How to Have a Sensitive Debate: An Instance-Optimal Protocol for AI Debate
arXiv:2610.02557v1 Announce Type: new Abstract: As powerful AI systems reach and sometimes surpass the abilities of human experts across a range of cognitively demanding tasks, the problem of accurate oversight and…
- arXiv — cs.AI preprintsInternational5 Oct 2026
Mitigating Social Sycophancy via Pluralistic Preference Optimization
arXiv:2610.02568v1 Announce Type: new Abstract: Personal advice, including relationship advice, now ranks among the most common uses of generative AI. But language models (LMs) exhibit sycophancy: they affirm users much…
- arXiv — cs.AI preprintsInternational5 Oct 2026
Answering clinicians' questions over trial evidence tables with verifiable, feedback-driven language models
arXiv:2610.02576v1 Announce Type: new Abstract: Systematic reviews condense clinical trials into evidence tables, yet clinicians can interrogate these tables only through database queries, and many questions concern…
- arXiv — cs.AI preprintsInternational5 Oct 2026
Labels Override Definitions in Jev-Style Typed Decision Models
arXiv:2610.02586v1 Announce Type: new Abstract: A typed decision model answers a fixed question about an input by returning a probability for each of several caller-defined options. Each option carries a short label and…
- arXiv — cs.AI preprintsInternational5 Oct 2026
Open-Endedness Bench: Measuring Epistemic Process from Agent Records
arXiv:2610.02588v1 Announce Type: new Abstract: Agents are increasingly given open-ended research tasks: discovering an empirical law from self-designed experiments, improving a heuristic whose optimum nobody knows, or…
- arXiv — cs.AI preprintsInternational5 Oct 2026
TasteBench: Multimodal Benchmark for Sensory Prediction, from Molecules to Sustainable Foods
arXiv:2610.02599v1 Announce Type: new Abstract: Sustainable protein discovery lacks the fast computational proxies, analogous to molecular docking or density functional theory, that accelerate drug and materials…
- arXiv — cs.AI preprintsInternational5 Oct 2026
Time Series Forecasting Benchmarks Need Scenario-Grounded Stress Testing
arXiv:2610.02608v1 Announce Type: new Abstract: Time series forecasting (TSF) increasingly drives decisions in transportation, energy, finance, healthcare, and infrastructure, yet current evaluation remains overly…
- arXiv — cs.AI preprintsInternational5 Oct 2026
VERSE: Verified Self-Evolving Optimizer for Agent Harnesses
arXiv:2610.02616v1 Announce Type: new Abstract: Harness evolution improves an LLM agent's prompts, tools, and workflow, while the optimizer's own tools and procedures often remain fixed. We study whether an optimizer…
- arXiv — cs.AI preprintsInternational5 Oct 2026
CuBEs: Culturally-Situated Behavioral Evaluations and the Limitations of Culture-Blind LLM Judges
arXiv:2610.02622v1 Announce Type: new Abstract: Evaluating the occurrence and triggers of large language model (LLM) behaviors - such as sycophancy, self-preference, or over-confidence - is critical for predicting…
- arXiv — cs.AI preprintsInternational5 Oct 2026
Lost in the Request: How Communication Variation Disrupts Retrieval and Action in Email Agents
arXiv:2610.02627v1 Announce Type: new Abstract: An email assistant should not complete less work simply because a user phrases the same request differently. Yet most benchmarks test each task with only one canonical…
- arXiv — cs.AI preprintsInternational5 Oct 2026
Designing the Future of User Feedback for Generative AI
arXiv:2610.02631v1 Announce Type: new Abstract: Post-deployment feedback from users can be a cost-effective, scalable, and representative means to monitor and improve generative AI systems and features. When implemented…
- arXiv — cs.AI preprintsInternational5 Oct 2026
Batched Speech Decisions Without Decoding: Single-Token Supervision Lets a Frozen LLM Hear Beyond the Transcript
arXiv:2610.02638v1 Announce Type: new Abstract: Full-duplex voice agents make many small, closed decisions, which current systems answer by slow autoregressive decoding. We propose DuplexJev, which feeds ASR-encoder…
- arXiv — cs.AI preprintsInternational5 Oct 2026
Coherence-Driven Belief Formation and Population Dynamics of Contagion in LLM Agents
arXiv:2610.02654v1 Announce Type: new Abstract: Models of social contagion usually assume how individuals adopt beliefs and derive population behavior from it. We instead empirically measure belief adoption in language…
- arXiv — cs.AI preprintsInternational5 Oct 2026
A GHOST in Long-Horizon Agents: Governance Hazard from Overlooked Safety Constraints across Turns
arXiv:2610.02664v1 Announce Type: new Abstract: Long-horizon agents are now playing an increasingly significant role in assisting humans with complex problem-solving. However, it is exactly their extended interaction…
- arXiv — cs.AI preprintsInternational5 Oct 2026
Spend Teacher Tokens Where They Matter: Success-Referenced On-Policy Distillation
arXiv:2610.02678v1 Announce Type: new Abstract: On-policy distillation (OPD) combines student-generated rollouts with dense token-level supervision from a teacher, but providing such supervision for every rollout…
- arXiv — cs.AI preprintsInternational5 Oct 2026
DataWeave: Deploying Human-LLM Analytics for Exploratory Structured Data Analysis
arXiv:2610.02679v1 Announce Type: new Abstract: Data journalism, the practice of using data analysis to surface newsworthy stories, depends increasingly on the ability of reporters and investigative journalists to…
- arXiv — cs.AI preprintsInternational5 Oct 2026
Large language models exhibit unreliable updating of clinical judgment as patient evidence evolves
arXiv:2610.02684v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly explored for clinical reasoning, but whether they appropriately revise judgments as patient evidence evolves remains unclear.…
- arXiv — cs.AI preprintsInternational5 Oct 2026
Decoupling Memory from Context: Structured Memory for Token-Efficient Test-Time Continual Learning
arXiv:2610.02687v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly deployed in enterprise, scientific, and medical applications, where agents must incorporate domain-specific knowledge and…
- arXiv — cs.AI preprintsInternational5 Oct 2026
Learning to Revise Reasoning with Segment-wise On-Policy Distillation
arXiv:2610.02703v1 Announce Type: new Abstract: On-policy distillation (OPD) improves large language model reasoning by training students on their own rollouts with dense token-wise supervision from the teacher.…
- arXiv — cs.AI preprintsInternational5 Oct 2026
Label-Efficient Time Series Classification at Scale: A Dual-Stream OSSE-LSTM with Counterfactual Attribution
arXiv:2610.02704v1 Announce Type: new Abstract: Time series are produced continuously at enormous scale by industrial equipment, wearables, power grids, and clinical monitors, yet annotation remains manual, expensive,…
- arXiv — cs.AI preprintsInternational5 Oct 2026
Ego2World: Compiling Egocentric Cooking Videos into Executable Worlds for Belief-State Planning
arXiv:2610.02715v1 Announce Type: new Abstract: Egocentric videos capture how people carry out everyday activities, yet testing an agent requires evaluating the consequences of actions it chooses itself. We introduce…
- arXiv — cs.AI preprintsInternational5 Oct 2026
On the Chain-of-Thought Monitorability of Looped Language Models
arXiv:2610.02741v1 Announce Type: new Abstract: Chain-of-thought (CoT) monitoring provides a promising approach for detecting undesirable model behavior. Looped language models (LoopLMs) repeatedly apply shared…
- arXiv — cs.AI preprintsInternational5 Oct 2026
Dynamic LLM Routers are Often Misguided
arXiv:2610.02762v1 Announce Type: new Abstract: Dynamic LLM routers promise to cut inference costs by sending each query to the cheapest model that can answer it correctly. We analyze six commercial routers across 14…
- arXiv — cs.AI preprintsInternational5 Oct 2026
Law And Order: Tax Law Autoformalization
arXiv:2610.02792v1 Announce Type: new Abstract: Legal systems are increasingly implemented through software, yet scalable methods for translating legal texts into accurate symbolic representations remain underdeveloped.…
- arXiv — cs.AI preprintsInternational5 Oct 2026
PAPER2LLM++: Continual Self-Evolution of LLMs from Research Papers
arXiv:2610.02793v1 Announce Type: new Abstract: Research on LLMs continually uncovers model limitations, their causes, and potential solutions. Yet these human discoveries remain largely disconnected from model…
- arXiv — cs.AI preprintsInternational5 Oct 2026
Modeling Shared and Individual Structure for Cross-Subject Continuous Affect Regression from EEG-fNIRS
arXiv:2610.02796v1 Announce Type: new Abstract: Continuous, second-by-second valence-arousal estimation from physiological signals is typically studied in a subject-dependent setting, where the model sees labeled data…
- arXiv — cs.AI preprintsInternational5 Oct 2026
BitNest: Bit-Nested Speculative Decoding for Memory-Efficient LLM Inference Acceleration
arXiv:2610.02800v1 Announce Type: new Abstract: Speculative decoding accelerates autoregressive generation by using a lightweight draft to propose multiple tokens for parallel verification. However, existing methods…
- arXiv — cs.AI preprintsInternational5 Oct 2026
VIGOR: Zero-Shot Visual Generalization via Latent-Space Consistency in Model-Based Reinforcement Learning
arXiv:2610.02801v1 Announce Type: new Abstract: Model-based reinforcement learning (MBRL) achieves strong sample efficiency by planning within learned latent dynamics, yet its performance degrades substantially under…
- arXiv — cs.AI preprintsInternational5 Oct 2026
ROUTEAUDIT: Interaction-Aware Identification for Budgeted Multi-Verifier Routing
arXiv:2610.02808v1 Announce Type: new Abstract: Adaptive multi-verifier systems are commonly compared through endpoint quality-cost gaps, even when the verifier catalog, availability, accounting, information filtration,…
- arXiv — cs.AI preprintsInternational5 Oct 2026
iS-KV: Online Low-Rank KV Cache Compression via Block-Incremental SVD
arXiv:2610.02815v1 Announce Type: new Abstract: Long chain-of-thought reasoning substantially increases KV-cache memory during autoregressive decoding, as every generated token introduces new key and value states and…
- arXiv — cs.AI preprintsInternational5 Oct 2026
MetaRubric: Learning to Reward for Rubric-Based Reinforcement Learning
arXiv:2610.02824v1 Announce Type: new Abstract: Rubric-based reinforcement learning extends reward-driven optimization to open-ended tasks by assigning partial credit to individual response requirements. However, rubric…
- arXiv — cs.AI preprintsInternational5 Oct 2026
Scaling Trajectories for Complex Tasks through Recursive Self-Rewrite
arXiv:2610.02826v1 Announce Type: new Abstract: Successful trajectories on difficult tasks provide valuable supervision for model improvement, but specialized harnesses introduce interventions that may be unavailable…
- arXiv — cs.AI preprintsInternational5 Oct 2026
MLCommons Jailbreak Benchmark v1.0
arXiv:2610.02827v1 Announce Type: new Abstract: Modern AI systems are designed to refuse hazardous requests. A jailbreak is a prompt crafted to bypass those safeguards and elicit outputs that the system would normally…
- arXiv — cs.AI preprintsInternational5 Oct 2026
FSPO: Policy-Consistent Risk and Pareto-Feasible Control for Budgeted LLM RL Post-Training
arXiv:2610.02828v1 Announce Type: new Abstract: Adaptive LLM reinforcement-learning post-training changes multiple training actuators online, including rollout temperature, group size, clipping, KL regularization,…
- arXiv — cs.AI preprintsInternational5 Oct 2026
AMBER: Multi-View Adaptive Budget Allocation for Listwise Vision-Language Reranking
arXiv:2610.02831v1 Announce Type: new Abstract: Vision-language models (VLMs) are powerful listwise rerankers for multimodal retrieval, but high inference costs restrict them to evaluating small local candidate views.…
- arXiv — cs.AI preprintsInternational5 Oct 2026
DNAlign: Dynamic Null-Space Safe Alignment for LLMs
arXiv:2610.02844v1 Announce Type: new Abstract: Ensuring the safe and reliable deployment of large language models (LLMs) remains a fundamental challenge. Existing safety alignment approaches either incur high…