Blog

告别冗长 PDF,一文读懂最新顶会与核心期刊的创新价值。

High-fidelity collisional quantum gates with fermionic atoms
Shape anisotropy governs organization of active rods: Swarming, turbulence, flocking, and jamming.
Evaluating Large Language Models in Scientific Discovery
Evaluating Large Language Models in Scientific Discovery
StepSearch: Igniting LLMs Search Ability via Step-Wise Proximal Policy Optimization
食品 3D 打印的发展及挑战
LLMs are not (consistently) Bayesian: Quantifying internal (in)consistencies of LLMs' probabilistic beliefs
Shape anisotropy governs organization of active rods: Swarming, turbulence, flocking, and jamming.
ReSeek: A Self-Correcting Framework for Search Agents with Instructive Rewards
Using vision language models for safety hazard identification in construction
CMDPAD: A Chinese multimodal dynamic personality and affect dataset for affect prediction in conversations $
Direct Reasoning Optimization: Constrained RL with Token-Level Dense Reward and Rubric-Gated Constraints for Open-ended Tasks
A Comprehensive Survey on Agent Skills: Taxonomy, Techniques, and Applications
PRISM: Pareto-Efficient Retrieval over Intent-Aware Structured Memory for Long-Horizon Agents
Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control
Improving RL Exploration for LLM Reasoning through Retrospective Replay
Cruxeval: A benchmark for code reasoning, understanding and execution
Deep learning-based embryo assessment of static images can reduce the time to live birth in <i>in vitro</i> fertilization
ProcMEM: Learning Reusable Procedural Memory from Experience via Non-Parametric PPO for LLM Agents
Mechanisms of pile-soil stress and deformation in excavations under the coupled effects of excavation disturbance and extreme rainfall infiltration
Useful for Exploration, Risky for Precision: Evaluating AI Tools in Academic Research
MemSkill: Learning and Evolving Memory Skills for Self-Evolving Agents
It Takes Two: Your GRPO Is Secretly DPO
Shape anisotropy governs organization of active rods: Swarming, turbulence, flocking, and jamming.
GRAPE: Let GPRO Supervise Query Rewriting by Ranking for Retrieval
Towards Self-Evolving Agentic Literature Retrieval
Composite Signal Monitoring for Reward Hacking Detection: A Case Study with Dense Gold Evaluation
amt-10-393-2017
UIEC2Net 21年v2
UIECLIP24v2