WisPaper
WisPaper
Search
Features
Resources
Pricing
Download
Workspace
Blog
No more endless PDFs. Discover the core value of the latest top-tier research in one article.
User Shared
Trends
The Spike, the Sparse and the Sink: Anatomy of Massive Activations and Attention Sinks
Neural Thickets: Diverse Task Experts Are Dense Around Pretrained Weights
Training Language Models via Neural Cellular Automata
IndexCache: Accelerating Sparse Attention via Cross-Layer Index Reuse
RubiCap: Rubric-Guided Reinforcement Learning for Dense Image Captioning
Scalable Training of Mixture-of-Experts Models with Megatron Core
When Does Sparsity Mitigate the Curse of Depth in LLMs
Autonomous Agents Coordinating Distributed Discovery Through Emergent Artifact Exchange
Theory of Code Space: Do Code Agents Understand Software Architecture?
Understanding the Emergence of Seemingly Useless Features in Next-Token Predictors
SCAN: Structured Capability Assessment and Navigation for LLMs
Towards a Neural Debugger for Python
Mixture-of-Depths Attention
Why AI systems don't learn and what to do about it: Lessons on autonomous learning from cognitive science
Online Experiential Learning for Language Models
Attention Residuals
Efficient Exploration at Scale
V-JEPA 2.1: Unlocking Dense Features in Video Self-Supervised Learning
Representation Learning for Spatiotemporal Physical Systems
PRISM: Demystifying Retention and Interaction in Mid-Training
Overton Pluralistic Reinforcement Learning for Large Language Models
R&D-Agent-Quant: A Multi-Agent Framework for Data-Centric Factors and Model Joint Optimization
Exclusive Self Attention
M$^2$RNN: Non-Linear RNNs with Matrix-Valued States for Scalable Language Modeling
Can AI Agents Agree?
Memento-Skills: Let Agents Design Agents
BloombergGPT: A Large Language Model for Finance
Towards Prospective Identification of Optimal Scan Parameters for CT-guided Interventions -A Phantom Study
GradMem: Learning to Write Context into Memory with Test-Time Gradient Descent
Hyperagents
←
1
...
21
22
23
...
614
→