arXiv · Computation and Language · 13 Aug 2026 · paper
Diffusion Large Language Models (dLLMs) have emerged as a competitive alternative to autoregressive language models, offering the potential for substantially faster inference through parall…
arXiv · Machine Learning · 13 Aug 2026 · paper
Large kernel depthwise convolutions achieve strong performance but suffer from significant degradation as kernel size grows due to irregular memory access from gather-based computation; whi…
arXiv · Machine Learning · 13 Aug 2026 · paper
Multi-agent reinforcement learning (MARL) provides a promising solution for cooperative target tracking in networks of autonomous underwater vehicles (AUVs). However, existing methods still…
arXiv · Machine Learning · 13 Aug 2026 · paper
Credit risk detection, particularly mitigating individual fraud, is crucial for maintaining the stability of digital financial ecosystems. Accurately identifying credit fraud among billions…
arXiv · Machine Learning · 13 Aug 2026 · paper
Text-to-image generative models have advanced rapidly, with modern Diffusion Transformer architectures producing images that are increasingly difficult to distinguish from human-created art…
arXiv · Machine Learning · 13 Aug 2026 · paper
AlphaFold3-class structure predictors pair a representational trunk, which processes sequence and context, with a diffusion module, which generates atomic coordinates. How biological inform…
arXiv · Machine Learning · 13 Aug 2026 · paper
The parity problem--deciding whether the number of ones in a binary vector is odd or even--remains challenging for standard neural networks due to linear inseparability and the need for glo…
arXiv · Machine Learning · 13 Aug 2026 · paper
Coarse-grid numerical solvers can substantially reduce the computational cost of time-dependent PDE simulation, but under-resolution often degrades both the trajectory and the spatial fidel…
arXiv · Machine Learning · 13 Aug 2026 · paper
Assortment optimization is a fundamental problem in revenue management, typically addressed using parametric choice models such as the multinomial logit (MNL) and its variants. While these…
arXiv · Machine Learning · 13 Aug 2026 · paper
Accurate solar irradiance forecasting is essential for the reliable integration of photovoltaic power into modern electricity grids. All-sky imagers (ASI) provide high-resolution observatio…
arXiv · Machine Learning · 13 Aug 2026 · paper
Most image colorization systems operate in $Lab$ space by predicting chroma ($ab$) while preserving an input-derived luminance channel ($L$). While effective on standard benchmarks, this fi…
arXiv · Machine Learning · 13 Aug 2026 · paper
Fine-tuning pre-trained robot policies with reinforcement learning (RL) often inherits the bottlenecks introduced by pre-training with behavioral cloning (BC), which produces narrow action…
arXiv · Machine Learning · 13 Aug 2026 · paper
Most modern bridge-diffusion methods achieve finite-time transport by specifying an interpolation, Schrodinger-bridge, or stochastic-control objective and then learning the associated score…
arXiv · Artificial Intelligence · 13 Aug 2026 · paper
Incremental Learning (IL) aims to learn new tasks while preserving previously acquired knowledge. Integrating the zero-shot learning capabilities of pre-trained vision-language models into…
arXiv · Artificial Intelligence · 13 Aug 2026 · paper
Aerial vision-language navigation (VLN) requires an embodied agent to integrate visual evidence over time, plan future actions, and determine when it has reached a navigation goal under par…
arXiv · Machine Learning · 13 Aug 2026 · paper
Graph Neural Networks (GNNs) achieve strong predictive performance on graph-structured data across domains such as chemistry, biology, and network analysis, yet they provide no intrinsic ex…
arXiv · Artificial Intelligence · 13 Aug 2026 · paper
Video diffusion transformers are costly to sample: every denoising step applies self-attention over a long 3D token sequence, a quadratic cost that dominates as resolution and duration grow…
arXiv · Artificial Intelligence · 13 Aug 2026 · paper
Quantum generative adversarial networks (QGANs) have attracted increasing attention for image generation using parameterized quantum circuits. Existing amplitude-based approaches face two k…
arXiv · Artificial Intelligence · 13 Aug 2026 · paper
Proprietary text-to-image diffusion models are increasingly distributed as hosted services and downloadable checkpoints, making their intellectual property (IP) protection an increasingly c…
arXiv · Artificial Intelligence · 13 Aug 2026 · paper
Videos captured through glass often contain reflections that degrade visual quality and interfere with downstream vision tasks. Although single-image reflection removal has been extensively…
arXiv · Machine Learning · 13 Aug 2026 · paper
Generative semantic segmentation exposes structured predictions as images, but direct color decoding is susceptible to color drift and boundary mixing, whereas latent-feature decoders that…
arXiv · Artificial Intelligence · 13 Aug 2026 · paper
Text-to-image (T2I) models are increasingly asked to illustrate literary and cultural content, yet we cannot measure how well an image renders the meaning of a poem. The task is many-sided:…
arXiv · Machine Learning · 13 Aug 2026 · paper
Long-tailed classification poses a reliability challenge because models trained on imbalanced data are unevenly reliable across frequent and underrepresented classes. While existing methods…
arXiv · Computation and Language · 13 Aug 2026 · paper
We study the problem of lossless text compression, motivated by the rapid growth in the collection and storage of digital textual data - including plain text, source code, and structured fo…
arXiv · Machine Learning · 13 Aug 2026 · paper
Multimodal retrieval and classification across different types of media, spanning text, images,video and audio, has traditionally relied on dual-encoder models that align visual and textual…
arXiv · Artificial Intelligence · 13 Aug 2026 · paper
Diffusion language models (DLMs) update many tokens in parallel, yet practical decoders often use a fixed denoising horizon. Many predictions stabilize early, but blockwise decoding continu…
Hugging Face · 23 Jul 2026
Berkeley AI Research · 01 Jul 2026 · paper
Congratulations to the Berkeley Artificial Intelligence Research (BAIR) Lab class of 2026! This year, BAIR celebrates another remarkable group of Ph.D. graduates whose curiosity, creativity…
NVIDIA Developer · 12 Jun 2026
Developers building real-time AI—such as chat assistants, copilots, and agentic workflows—are often constrained by token-by-token generation speed. This...
OpenAI · 21 Apr 2026 · release software
ChatGPT Images 2.0 introduces a state-of-the-art image generation model with improved text rendering, multilingual support, and advanced visual reasoning.
Berkeley AI Research · 20 Apr 2026 · paper
GRASP is a new gradient-based planner for learned dynamics (a “world model”) that makes long-horizon planning practical by (1) lifting the trajectory into virtual states so optimization is…
OpenAI · 16 Apr 2026 · comunicato aziendale
The updated Codex app for macOS and Windows adds computer use, in-app browsing, image generation, memory, and plugins to accelerate developer workflows.
Hugging Face · 05 Mar 2026 · release software
Hugging Face · 03 Mar 2026
Hugging Face · 03 Feb 2026
Hugging Face · 20 Jan 2026 · release software
VentureBeat · 13 Jan 2026 · intervista
Salesforce on Tuesday launched an entirely rebuilt version of Slackbot , the company's workplace assistant, transforming it from a simple notification tool into what executives describe as…
OpenAI · 16 Dec 2025 · comunicato aziendale
The new ChatGPT Images is powered by our flagship image generation model, delivering more precise edits, consistent details, and image generation up to 4× faster. The upgraded model is roll…
OpenAI · 22 Sep 2025 · tutorial
SchoolAI uses GPT-4.1, image generation, and TTS to power safe, teacher-guided AI tools for over 1 million classrooms, improving engagement, oversight, and personalized learning.
Hugging Face · 11 May 2025
OpenAI · 24 Apr 2025 · recensione
Watch hands-on demos of the lastest in ChatGPT for Business: o3, image generation, enhanced memory, and internal knowledge.
OpenAI · 23 Apr 2025 · release software
Our latest image generation model is now available in the API via ‘gpt-image-1’—enabling developers and businesses to build professional-grade, customizable visuals directly into their own…
OpenAI · 16 Apr 2025 · comunicato aziendale
OpenAI o3 and OpenAI o4-mini combine state-of-the-art reasoning with full tool capabilities—web browsing, Python, image and file analysis, image generation, canvas, automations, file search…
OpenAI · 25 Mar 2025 · release software
At OpenAI, we have long believed image generation should be a primary capability of our language models. That’s why we’ve built our most advanced image generator yet into GPT‑4o. The result…
OpenAI · 25 Mar 2025 · comunicato aziendale
4o image generation is a new, significantly more capable image generation approach than our earlier DALL·E 3 series of models. It can create photorealistic output. It can take images as inp…
Hugging Face · 09 Dec 2024
OpenAI · 23 Oct 2024 · comunicato aziendale
We’ve simplified, stabilized, and scaled continuous-time consistency models, achieving comparable sample quality to leading diffusion models, while using only two sampling steps.
Hugging Face · 22 Oct 2024
Hugging Face · 30 Jul 2024
OpenAI · 20 Jun 2024 · comunicato aziendale
Diffusion models have significantly advanced the fields of image, audio, and video generation, but they depend on an iterative sampling process that causes slow generation.
Hugging Face · 12 Jun 2024
OpenAI · 15 Feb 2024 · comunicato aziendale
We explore large-scale training of generative models on video data. Specifically, we train text-conditional diffusion models jointly on videos and images of variable durations, resolutions…
Hugging Face · 03 Feb 2024
Hugging Face · 04 Jan 2024
Hugging Face · 03 Oct 2023
Hugging Face · 29 Sep 2023
Hugging Face · 13 Sep 2023 · release software
Hugging Face · 27 Jul 2023
Hugging Face · 14 Jul 2023
Hugging Face · 26 Jun 2023