Blog
告别冗长 PDF,一文读懂最新顶会与核心期刊的创新价值。
High-fidelity collisional quantum gates with fermionic atoms Shape anisotropy governs organization of active rods: Swarming, turbulence, flocking, and jamming. Evaluating Large Language Models in Scientific Discovery Evaluating Large Language Models in Scientific Discovery StepSearch: Igniting LLMs Search Ability via Step-Wise Proximal Policy Optimization LLMs are not (consistently) Bayesian: Quantifying internal (in)consistencies of LLMs' probabilistic beliefs Shape anisotropy governs organization of active rods: Swarming, turbulence, flocking, and jamming. ReSeek: A Self-Correcting Framework for Search Agents with Instructive Rewards Using vision language models for safety hazard identification in construction CMDPAD: A Chinese multimodal dynamic personality and affect dataset for affect prediction in conversations $ Direct Reasoning Optimization: Constrained RL with Token-Level Dense Reward and Rubric-Gated Constraints for Open-ended Tasks A Comprehensive Survey on Agent Skills: Taxonomy, Techniques, and Applications PRISM: Pareto-Efficient Retrieval over Intent-Aware Structured Memory for Long-Horizon Agents Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control Improving RL Exploration for LLM Reasoning through Retrospective Replay Cruxeval: A benchmark for code reasoning, understanding and execution Deep learning-based embryo assessment of static images can reduce the time to live birth in <i>in vitro</i> fertilization ProcMEM: Learning Reusable Procedural Memory from Experience via Non-Parametric PPO for LLM Agents Mechanisms of pile-soil stress and deformation in excavations under the coupled effects of excavation disturbance and extreme rainfall infiltration Useful for Exploration, Risky for Precision: Evaluating AI Tools in Academic Research MemSkill: Learning and Evolving Memory Skills for Self-Evolving Agents It Takes Two: Your GRPO Is Secretly DPO Shape anisotropy governs organization of active rods: Swarming, turbulence, flocking, and jamming. GRAPE: Let GPRO Supervise Query Rewriting by Ranking for Retrieval Towards Self-Evolving Agentic Literature Retrieval Composite Signal Monitoring for Reward Hacking Detection: A Case Study with Dense Gold Evaluation