arXiv · Computation and Language · 13 Aug 2026 · paper
Mixture-of-Experts (MoE) models have become a dominant architecture for large-scale AI services, yet deploying them over geo-distributed heterogeneous edge servers remains challenging. When…
arXiv · Computation and Language · 13 Aug 2026 · paper
Serving Mixture-of-Experts (MoE) large language models across distributed edge servers is bottlenecked by the cross-server expert transmission. The existing approaches mainly focus on how t…
arXiv · Computation and Language · 13 Aug 2026 · paper
A key capability for video understanding is reliably linking subjects to events across time, yet whether Video Large Language Models (VideoLLMs) actually achieve this remains unclear. In th…
arXiv · Computation and Language · 13 Aug 2026 · paper
Large language model (LLM)-based educational assistants often provide direct answers offering little incentive for students to explore or engage with course materials. We present BLADE (Bet…
arXiv · Computation and Language · 13 Aug 2026 · paper
Political stance detection in LLMs has long been dominated by closed-ended, multiple-choice political survey questions---originally designed for humans, and thus lacks the realism and nuanc…
arXiv · Computation and Language · 13 Aug 2026 · paper
Large language model evaluations typically focus on performance under nominal conditions, creating an illusion of capability where models comfortably walk a narrow, highly optimized generat…
arXiv · Computation and Language · 13 Aug 2026 · paper
Subword tokenizers represent many common words twice in space-using writing systems, once with a leading space and once without. The two entries have separate embeddings in models, so occur…
arXiv · Computation and Language · 13 Aug 2026 · paper
Reinforcement learning (RL) search agents commonly model retrieval as free-form natural-language query generation and optimize multi-turn interactions using final-answer rewards. Current st…
arXiv · Computation and Language · 13 Aug 2026 · paper
The Unigram tokenizer uses an elegant representation which makes it straightforward to edit vocabularies, but its training is comparatively heavy and complex. We introduce MinGram (Minimali…
arXiv · Computation and Language · 13 Aug 2026 · paper
The performance of LLM-based agents is jointly shaped by their base models and the harnesses that mediate their interaction with the environment. Because different models exhibit distinct b…
arXiv · Computation and Language · 13 Aug 2026 · paper
Experience-based self-evolution enables language-model agents to improve their behavior by accumulating and updating experience at test time, yet existing evaluations often assume recurring…
arXiv · Computation and Language · 13 Aug 2026 · paper
Discourse comprehension in complex documents often involves continuously posing and resolving Questions Under Discussion (QUDs). While QUD frameworks have so far focused on text, scientific…
arXiv · Computation and Language · 13 Aug 2026 · paper
Although large language models (LLMs) show promise for generating educational feedback, it remains unclear how feedback should be designed (e.g., tone and information coverage) to support a…
arXiv · Computation and Language · 13 Aug 2026 · paper
This paper presents a multifunctional speech synthesis system that integrates voice cloning and emotion control speech synthesis within a unified framework. The goal of this work is to addr…
arXiv · Computation and Language · 13 Aug 2026 · paper
Recent work shows superior performance when using large language models (LLMs) as formalizers instead of as end-to-end solvers for symbolic reasoning problems. Given the problem description…
arXiv · Computation and Language · 13 Aug 2026 · paper
Creative agents still lack an effective way to learn from high-quality human films, limiting their ability to produce cinematic-grade videos. A key challenge is the absence of a structured…
arXiv · Computation and Language · 13 Aug 2026 · paper
Large language models can extract richer signals from financial news than fixed sentiment lexicons, and recent work has explored feeding such signals into portfolio construction. We study a…
arXiv · Computation and Language · 13 Aug 2026 · paper
We present RT-SEMamba, a fully causal speech enhancement (SE) model built upon causal time-frequency Mamba blocks. Unlike Transformer-based architectures that rely on a growing key-value ca…
arXiv · Computation and Language · 13 Aug 2026 · paper
Large language model (LLM) agents integrated with external tools are vulnerable to indirect prompt injections embedded in environmental states. However, existing studies largely rely on man…
arXiv · Computation and Language · 13 Aug 2026 · paper
The deployment of large language models (LLMs) in mental health contexts raises questions about the relationship between clinical safety and environmental cost. In this paper, we examine th…
arXiv · Computation and Language · 13 Aug 2026 · paper
Long-form song generation models continue to improve in duration, structural integrity, and acoustic complexity, making reliable aesthetic rewards increasingly important for aligning these…
arXiv · Computation and Language · 13 Aug 2026 · paper
Recent advances in zero-shot text-to-speech (TTS) have substantially improved speech quality and voice cloning fidelity. However, many zero-shot TTS systems still depend on audio prompt tra…
arXiv · Computation and Language · 13 Aug 2026 · paper
The computability of real numbers and functions using Turing Machines has been a central area of theoretical computer science since the mid-20th century. In the late 20th century, it was sh…
arXiv · Computation and Language · 13 Aug 2026 · paper
Public procurement involves the allocation of substantial financial resources; therefore, continuous oversight through audits, controls, and monitoring mechanisms is essential. However, sta…
arXiv · Computation and Language · 13 Aug 2026 · paper
We present the first systematic study of Massive activations (MAs) in layer-interleaved HLA LLMs and uncover two architecture-aligned morphologies: MAs consistently spike immediately before…
arXiv · Computation and Language · 13 Aug 2026 · paper
While retrieval-augmented generation (RAG) has proven effective at giving LLMs access to external knowledge, mainstream dense-retrieval implementations remain inherently limited in handling…
arXiv · Computation and Language · 13 Aug 2026 · paper
The same event can be reported from different perspectives depending on the experiences, background, and beliefs of the writer or speaker. A variety of NLP areas engage with perspectives, s…
arXiv · Computation and Language · 13 Aug 2026 · paper
Regional dialectal variation poses a fundamental challenge to natural language processing (NLP) in Bangla, where over 240 million speakers communicate across diverse regional variants that…
arXiv · Computation and Language · 13 Aug 2026 · paper
Large language models (LLMs) may generate fluent but incorrect answers, making uncertainty quantification important for reliable question answering. However, heuristic uncertainty scores ca…
arXiv · Computation and Language · 13 Aug 2026 · paper
Small Language Models (SLMs) have emerged as a more efficient alternative to traditional Large Language Models (LLMs), offering promising potential in resource-constrained scenarios. Existi…
arXiv · Computation and Language · 13 Aug 2026 · paper
Turning a research idea into a complete paper requires more than text generation: the system must retrieve literature, design and execute experiments, revise claims according to evidence, p…
arXiv · Computation and Language · 13 Aug 2026 · paper
Training large language models on limited hardware is increasingly a scheduling problem across GPU compute, host memory, PCIe transfer, and storage bandwidth. Existing offloading systems re…
arXiv · Computation and Language · 13 Aug 2026 · paper
Long-running conversational agents increasingly rely on a memory system to avoid resending the whole conversation each turn, yet how much that costs to serve has received little systematic…
arXiv · Computation and Language · 13 Aug 2026 · paper
The Seungjeongwon Ilgi, a UNESCO Memory of the World record, is only 37.4% translated, and the most conspicuous failure mode in automatic translation is the person name -- a misread name co…
arXiv · Computation and Language · 13 Aug 2026 · paper
Gated attention is an effective approach to mitigate attention sinks and enhance the representational capacity of attention. To further extend its effectiveness-efficiency Pareto frontier,…
arXiv · Computation and Language · 13 Aug 2026 · paper
Aggressive quantization disproportionately harms multilingual capability: in the sub-4B INT3 GPTQ regime, we measure 2-4x larger perplexity degradation on non-English languages than on Engl…
arXiv · Computation and Language · 13 Aug 2026 · paper
Self-correction is particularly useful when a failure constrains the next repair. Coding agents benefit from this property because compilers, tests, and execution traces turn many failures…
arXiv · Computation and Language · 13 Aug 2026 · paper
When a language model answers an interventional question, the computation it must perform depends on the type of evidence the query requires. We report a decoupling in how a transformer org…
arXiv · Computation and Language · 13 Aug 2026 · paper
Multimodal Large Language Models (MLLMs) exhibit strong generalization and reasoning abilities due to large-scale multimodal pre-training. However, fine-tuning these models on downstream ta…
arXiv · Computation and Language · 13 Aug 2026 · paper
Financial text is produced and interpreted within a market environment, yet financial text classifiers almost always receive text alone. We study whether financial time series are useful as…
arXiv · Computation and Language · 13 Aug 2026 · paper
Diffusion Large Language Models (dLLMs) have emerged as a competitive alternative to autoregressive language models, offering the potential for substantially faster inference through parall…
arXiv · Computation and Language · 13 Aug 2026 · paper
As users increasingly turn to Large Language Models (LLMs) for information and advice on political matters, particularly during election periods, the political preferences expressed by thes…
arXiv · Computation and Language · 13 Aug 2026 · paper
Audio transcription is a critical bottleneck in language documentation. While multilingual Automatic Speech Recognition (ASR) models like Whisper offer solutions, field linguists often lack…
arXiv · Computation and Language · 13 Aug 2026 · paper
In medical imaging, the clinical value of Computed Tomography (CT) lies not only in depicting current disease status, but crucially in enabling longitudinal comparison of serial scans to de…
arXiv · Computation and Language · 13 Aug 2026 · paper
In the context of CCSK, a reversible extension of CCS, we study different notions of bisimilarity (strong/weak, forward-only/reversible) and highlight their differences and commonalities. I…
arXiv · Computation and Language · 13 Aug 2026 · paper
Group alignment adapts a language model to a demographic group to produce responses that reflect the group's opinions, values, and preferences. Sycophancy, a well-documented by-product of a…
arXiv · Computation and Language · 13 Aug 2026 · paper
Large Language Model-powered agents are increasingly used in the workplace via human-artificial intelligence (AI) collaboration. In this new era of work, it is important to understand the k…
arXiv · Computation and Language · 13 Aug 2026 · paper
Low-resource automatic speech recognition (ASR) commonly relies on cross-lingual transfer, where models are adapted from higher-resource donor languages. However, selecting donors remains c…
arXiv · Computation and Language · 13 Aug 2026 · paper
Online communities increasingly provide spaces where survivors of sexual violence can share their experiences and seek support. Although prior research has examined stigma and social suppor…
arXiv · Computation and Language · 13 Aug 2026 · paper
The lack of diversity in LM content is widely attributed to the alignment process, but how and where exactly in the pipeline this collapse begins is unknown. We argue that output homogeneit…
arXiv · Computation and Language · 13 Aug 2026 · paper
Prior white-box studies show that large language models can retain latent traces of target knowledge after unlearning, even when the knowledge is no longer expressed in their outputs. Howev…
arXiv · Computation and Language · 13 Aug 2026 · paper
Embodied agents are increasingly built as systems around foundation models, where performance depends not only on model weights but also on the skills, context, action interfaces, and execu…
arXiv · Computation and Language · 13 Aug 2026 · paper
Sign-language research for resource-constrained languages is often limited by the cost of dense linguistic labels such as glosses, temporal boundaries, and sign order. Broadcast news offers…
arXiv · Machine Learning · 13 Aug 2026 · paper
Anchor-based pointwise LLM reranking scores each candidate against a shared reference passage to recover cross-document context at pointwise cost. We study when this actually helps, using G…
arXiv · Computation and Language · 13 Aug 2026 · paper
Conflicting objectives are general in RL alignment, and training on them data-efficiently is hard. Training a safety guard with RL means optimizing two objectives that conflict: catch real…
arXiv · Machine Learning · 13 Aug 2026 · paper
Harness plays a critical role in large language model agent performance, and building a high-performing harness requires substantial expert effort. Therefore, recent research has increasing…
arXiv · Machine Learning · 13 Aug 2026 · paper
Image-source-method (ISM)-based room impulse response (RIR) simulation is a useful and physically interpretable tool for acoustic scene modeling, but full-order ISM becomes computationally…
arXiv · Machine Learning · 13 Aug 2026 · paper
Although mixture-of-experts, MoE, models have been increasingly adopted to scale large language models with moderate computation cost, it remains challenging to deploy MoE inference over re…
arXiv · Machine Learning · 13 Aug 2026 · paper
Large kernel depthwise convolutions achieve strong performance but suffer from significant degradation as kernel size grows due to irregular memory access from gather-based computation; whi…
arXiv · Computation and Language · 13 Aug 2026 · paper
Proposal. Long context can replay history, but it does not decide which completed observations deserve authority. MMLA formalizes a bounded resident memory between transient context and slo…