arXiv · Computation and Language · 13 Aug 2026 · paper
Large language model evaluations typically focus on performance under nominal conditions, creating an illusion of capability where models comfortably walk a narrow, highly optimized generat…
arXiv · Computation and Language · 13 Aug 2026 · paper
Large language model (LLM) agents integrated with external tools are vulnerable to indirect prompt injections embedded in environmental states. However, existing studies largely rely on man…
arXiv · Computation and Language · 13 Aug 2026 · paper
The deployment of large language models (LLMs) in mental health contexts raises questions about the relationship between clinical safety and environmental cost. In this paper, we examine th…
arXiv · Computation and Language · 13 Aug 2026 · paper
Conflicting objectives are general in RL alignment, and training on them data-efficiently is hard. Training a safety guard with RL means optimizing two objectives that conflict: catch real…
arXiv · Machine Learning · 13 Aug 2026 · paper
Many sequential decision problems offer qualitatively different ways of influencing the environment: some interventions act immediately, whereas others induce persistent effects that contin…
arXiv · Machine Learning · 13 Aug 2026 · paper
Natural-language-based scenario generation offers an intuitive means of describing rare and complex driving interactions, yet it is still uncertain whether training with language-structured…
arXiv · Machine Learning · 13 Aug 2026 · paper
IoT firmware vulnerability detection remains challenging due to heterogeneous firmware ecosystems, resource-constrained platforms, and limitations in existing benchmarks. Many datasets are…
arXiv · Machine Learning · 13 Aug 2026 · paper
Hamilton-Jacobi (HJ) reachability provides a mathematically rigorous framework for safe control of dynamical systems, but its practical application is bottlenecked by the computational comp…
arXiv · Machine Learning · 13 Aug 2026 · paper
Intrusion detection systems (IDS) and automated systems for detecting and reporting cyber threats, are commonly handled via supervised machine learning methods. Though effective, these mode…
arXiv · Machine Learning · 13 Aug 2026 · paper
Assessing catheter and tube placement on chest X-rays is safety-critical yet tedious and error-prone. Current deep learning methods either classify placement globally -- losing track of whi…
arXiv · Machine Learning · 13 Aug 2026 · paper
Cyberattack detection in electric vehicle charging infrastructure is complicated by legitimate post-activation revisions to requested energy and departure time. Charging manipulation attack…
arXiv · Machine Learning · 13 Aug 2026 · paper
Modern stochastic predictors can model rich, multi-modal outcome distributions. However, this expressive power comes with challenges in ensuring robust predictions $-$ a critical requiremen…
arXiv · Computation and Language · 13 Aug 2026 · paper
Reinforcement-learning (RL) post-training equips multimodal large reasoning models (MLRMs) with exploratory chains of thought (CoT), substantially improving visual reasoning. However, we fi…
arXiv · Machine Learning · 13 Aug 2026 · paper
Graph Neural Networks (GNNs) serve as the backbone for high-stakes applications in Machine-Learning-as-a-Service (MLaaS). Still, their black-box deployment exposes them to Model Extraction…
arXiv · Machine Learning · 13 Aug 2026 · paper
Offline reinforcement learning (RL) offers considerable promise for optimizing ICU treatment decisions, yet standard evaluation metrics Mean Squared Error (MSE) and Fitted Q-Evaluation (FQE…
arXiv · Artificial Intelligence · 13 Aug 2026 · paper
Autonomous AI agents increasingly execute actions, invoke tools, and operate on protected resources with limited human oversight. Existing authentication and authorization mechanisms establ…
arXiv · Artificial Intelligence · 13 Aug 2026 · paper
Cloud-based Large Language Models (LLMs) can perform autonomous penetration-testing sub-tasks such as Linux privilege escalation, but raise security, privacy, and sovereignty concerns. Loca…
arXiv · Artificial Intelligence · 13 Aug 2026 · paper
System Instructions in Large Language Models (LLMs) are commonly used to enforce safety policies, define agent behavior, and protect sensitive operational context in agentic AI applications…
arXiv · Machine Learning · 13 Aug 2026 · paper
The integration of artificial intelligence (AI) agents into web browsers introduces security challenges that go beyond traditional web application threat models. Prior work has identified p…
arXiv · Artificial Intelligence · 13 Aug 2026 · paper
Recent advances in AI applications have raised growing concerns about the need for ethical guidelines and regulations to mitigate the risks posed by these technologies. In this paper, we pr…
arXiv · Machine Learning · 13 Aug 2026 · paper
Large Language Models (LLMs) can achieve strong performance on everyday coding tasks, but they can fail on complex tasks that require non-trivial reasoning about program semantics. Finding…
arXiv · Artificial Intelligence · 13 Aug 2026 · paper
After an older adult misses a medication dose, a software agent can send another reminder and an embodied agent can bring the medication. Yet neither explains whether the person forgot, is…
arXiv · Machine Learning · 13 Aug 2026 · paper
Electronic health-record interoperability is a boundary problem: legacy systems, generative models, terminology services, identity systems, and human reviewers may each expose rich internal…
arXiv · Artificial Intelligence · 13 Aug 2026 · paper
Artificial Intelligence (AI) is rapidly transforming organizations, raising a fundamental organizational and economic question: when will a human employee be replaced by AI? We present an a…
arXiv · Computation and Language · 13 Aug 2026 · paper
We introduce SteeringSafety, a benchmark for evaluating representation steering methods across nine safety perspectives spanning 18 datasets. While prior work highlights the general capabil…
arXiv · Artificial Intelligence · 13 Aug 2026 · paper
To engineer AGI, we should first capture the essence of intelligence in a species-agnostic form that can be evaluated, while being sufficiently general to encompass diverse paradigms of int…
arXiv · Artificial Intelligence · 13 Aug 2026 · paper
Large language model (LLM) agents can execute long-horizon scientific workflows, but their numerical outputs are difficult to trust: agents lose context, game verification checks, and can p…
arXiv · Artificial Intelligence · 13 Aug 2026 · paper
LLM agents increasingly rely on third-party skills, using natural-language descriptions for selection and instruction bodies for planning. This progressive-disclosure design exposes two seq…
arXiv · Computation and Language · 13 Aug 2026 · paper
Evaluating security vulnerability detection tools requires benchmark datasets with vulnerability-inducing commits (VICs) - the commits that first introduce vulnerabilities into codebases. V…
arXiv · Machine Learning · 13 Aug 2026 · paper
Recent research in machine and deep learning has shown the potential of learningbased motion planning approaches to improve the driving behavior of automated vehicles, especially in complex…
arXiv · Artificial Intelligence · 13 Aug 2026 · paper
With the increasing complexity of cyber assaults in cloud environments, adaptable security solutions are needed that can support real-time detection and autonomous response. In this paper,…
arXiv · Artificial Intelligence · 13 Aug 2026 · paper
Algorithm registers have been championed as a means of providing transparency on the use of algorithms in public services. Yet potential publics differ in their expectations of what should…
arXiv · Machine Learning · 13 Aug 2026 · paper
In high-stakes applications, reliable confidence estimates are as important as the predictions themselves. Confidence calibration ensures that predicted probabilities reflect the likelihood…
arXiv · Machine Learning · 13 Aug 2026 · paper
Graph Neural Networks (GNNs) achieve strong predictive performance on graph-structured data across domains such as chemistry, biology, and network analysis, yet they provide no intrinsic ex…
arXiv · Artificial Intelligence · 13 Aug 2026 · paper
Medical devices are becoming more software-intensive, connected, and AI-enabled. Their development requires risk-management evidence aligned with ISO 14971 and, for software, IEC 62304. Thi…
arXiv · Machine Learning · 13 Aug 2026 · paper
Extreme events in air transport, such as severe arrival delays and abnormal air times, cause cascading network disruptions with substantial operational, economic, and safety costs. Such eve…
arXiv · Artificial Intelligence · 13 Aug 2026 · paper
Governments increasingly fund indigenous foundation models to strengthen national AI capability, digital sovereignty, and multilingual computing. Assessing the progress of such national eco…
arXiv · Artificial Intelligence · 13 Aug 2026 · paper
Methods to increase the resilience of systems to cyber-attacks become increasingly important. Control-flow monitoring provides a principled basis to ensure integrity and detect possible ano…
arXiv · Computation and Language · 13 Aug 2026 · paper
Generating actionable financial advice from business records demands that models integrate numerical reasoning, domain knowledge, and sound judgment, while avoiding recommendations that cou…
arXiv · Artificial Intelligence · 13 Aug 2026 · paper
Binary code representation learning is a fundamental problem in software security and reverse engineering. Existing methods mainly learn function-level embeddings that capture coarse-graine…
arXiv · Machine Learning · 13 Aug 2026 · paper
Many reinforcement learning systems, from fleet management to traffic signal control, must serve an objective that changes dynamically after deployment, and retraining a policy for each new…
arXiv · Computation and Language · 13 Aug 2026 · paper
Persuasion is a core dynamic of natural language communication, shaping how large language models (LLMs) update beliefs, resolve disagreements, and reach decisions. As LLMs increasingly deb…
arXiv · Artificial Intelligence · 13 Aug 2026 · paper
The convergence of artificial intelligence (AI), Industrial Internet of Things, cyber-physical systems, and advanced robotics is reshaping manufacturing faster than engineering curricula ca…
arXiv · Computation and Language · 13 Aug 2026 · paper
Large Language Models (LLMs) are increasingly integrated into software engineering workflows, helping developers write, debug, test, and maintain code. While prompt wording and structure ar…
arXiv · Artificial Intelligence · 13 Aug 2026 · paper
AI agents are rapidly improving in cybersecurity capabilities when the source code is available for analysis, yet much of the software most consequential to cybersecurity, including malware…
arXiv · Artificial Intelligence · 13 Aug 2026 · paper
Modern end-to-end driving agents can achieve high average performance yet still violate basic traffic rules that a human driver would never miss. The reason is structural: they learn statis…
arXiv · Artificial Intelligence · 13 Aug 2026 · paper
Large language models are being proposed as agents in scientific workflows, in domains where no downstream verifier exists. Such deployment assumes the model can distinguish reliable scient…
arXiv · Artificial Intelligence · 13 Aug 2026 · paper
Long-running agents periodically compact their context, replacing the transcript with a model-generated summary.Recent work shows that dropping a standing safety constraint during compactio…
arXiv · Machine Learning · 13 Aug 2026 · paper
This paper proposes a contextual quality-diversity evolutionary reinforcement-learning controller, CQD-ERL, for the supervisory control of a tropical, water-cooled chiller plant and its ass…
arXiv · Artificial Intelligence · 13 Aug 2026 · paper
The dominant paradigm treats AI safety as a property to be instilled during model training via RLHF, DPO, or Constitutional AI. We argue this is structurally insufficient for autonomous age…
arXiv · Artificial Intelligence · 13 Aug 2026 · paper
LLMs are increasingly pervasive in the security environment, with limited measures of their effectiveness, which limits trust and usefulness to security practitioners. Here, we present an o…
arXiv · Artificial Intelligence · 13 Aug 2026 · paper
Dynamic Master Logic (DML) provides a hierarchical framework for representing system behavior by linking functional objectives to underlying structural elements. However, DML construction t…
arXiv · Computation and Language · 13 Aug 2026 · paper
AI models have achieved remarkable success across diverse domains, yet the mechanisms underlying their capabilities and the risks they may pose remain poorly understood. As AI development b…
arXiv · Artificial Intelligence · 13 Aug 2026 · paper
Aligned large language models (LLMs) are expected to exhibit safety behavior based on the content of the user request: they should refuse unsafe requests and comply with safe ones. However,…
arXiv · Artificial Intelligence · 13 Aug 2026 · paper
Digital twins are increasingly used to monitor and simulate the behavior of cyber-physical systems. Even with skilled operators, interpreting anomalies detected within digital twin pipeline…
arXiv · Artificial Intelligence · 13 Aug 2026 · paper
Safety alignment in large language models is often treated as a distributed property of the entire network, yet its practical brittleness suggests that refusal behavior may be concentrated…
arXiv · Machine Learning · 13 Aug 2026 · paper
Hit-to-lead optimization requires iterative design of hit analogs across competing potency, selectivity, physicochemical, pharmacokinetic, safety, and synthetic constraints. We present SABL…
arXiv · Computation and Language · 13 Aug 2026 · paper
Construction standards are critical for building safety and sustainability. Existing standard application workflows rely on keyword-based document retrieval and manual cross-clause interpre…
arXiv · Machine Learning · 13 Aug 2026 · paper
We argue that a single structural fact organizes a wide range of phenomena in contemporary AI safety: a semantic safety constraint (e.g., the agent does not escape its sandbox) is an off-su…
arXiv · Machine Learning · 13 Aug 2026 · paper
Activation steering modifies a language model by adding a learned direction to its hidden activations, enabling targeted behavioral changes without retraining. While effective, steering oft…