|
arxiv.org
|
<font color="#0080FF">Sentient Agent as a Judge: Evaluating Higher-Order Social Cognition in Large Language Models</font>
|
Анализировать url
|
|
x.com
|
X
|
Анализировать url
|
|
arxiv.org
|
<font color="#0080FF">RLVER: Reinforcement Learning with Verifiable Emotion Rewards for Empathetic Agents</font>
|
Анализировать url
|
|
x.com
|
X
|
Анализировать url
|
|
arxiv.org
|
<font color="#0080FF">RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents</font>
|
Анализировать url
|
|
x.com
|
X
|
Анализировать url
|
|
arxiv.org
|
<font color="#0080FF">Think Fast and Slow: Step-Level Cognitive Depth Adaptation for LLM Agents</font>
|
Анализировать url
|
|
x.com
|
X
|
Анализировать url
|
|
arxiv.org
|
<font color="#0080FF">Encouraging Divergent Thinking in Large Language Models through Multi-Agent Debate</font>
|
Анализировать url
|
|
openreview.net
|
<font color="#0080FF">Apathetic or Empathetic? Evaluating LLMs' Emotional Alignments with Humans</font>
|
Анализировать url
|
|
openreview.net
|
<font color="#0080FF">On the Humanity of Conversational AI: Evaluating the Psychological Portrayal of LLMs</font>
|
Анализировать url
|
|
arxiv.org
|
<font color="#0080FF">Too Good to be Bad: On the Failure of LLMs to Role-Play Villains</font>
|
Анализировать url
|
|
x.com
|
X
|
Анализировать url
|
|
preprints.org
|
<font color="#0080FF">Deep Research: A Systematic Survey</font>
|
Анализировать url
|
|
x.com
|
X
|
Анализировать url
|
|
arxiv.org
|
<font color="#0080FF">The Hunger Game Debate: On the Emergence of Over-Competition in Multi-Agent Systems</font>
|
Анализировать url
|
|
x.com
|
X
|
Анализировать url
|
|
arxiv.org
|
<font color="#0080FF">Social Welfare Function Leaderboard: When LLM Agents Allocate Social Welfare</font>
|
Анализировать url
|
|
x.com
|
X
|
Анализировать url
|
|
openreview.net
|
<font color="#0080FF">From Poetry Parties to Poetry Critics: Benchmarking Classical Chinese Poetry Generation and Evaluation</font>
|
Анализировать url
|
|
arxiv.org
|
<font color="#0080FF">Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs</font>
|
Анализировать url
|
|
x.com
|
X
|
Анализировать url
|
|
arxiv.org
|
<font color="#0080FF">Thoughts Are All Over the Place: On the Underthinking of o1-Like LLMs</font>
|
Анализировать url
|
|
x.com
|
X
|
Анализировать url
|
|
arxiv.org
|
<font color="#0080FF">DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning</font>
|
Анализировать url
|
|
x.com
|
X
|
Анализировать url
|
|
arxiv.org
|
<font color="#0080FF">Critical Tokens Matter: Token-Level Contrastive Estimation Enhances LLM's Reasoning Capability</font>
|
Анализировать url
|
|
x.com
|
X
|
Анализировать url
|
|
arxiv.org
|
<font color="#0080FF">Don't Get Lost in the Trees: Streamlining LLM Reasoning by Overcoming Tree Search Exploration Pitfalls</font>
|
Анализировать url
|
|
x.com
|
X
|
Анализировать url
|
|
arxiv.org
|
<font color="#0080FF">Trust, But Verify: A Self-Verification Approach to Reinforcement Learning with Verifiable Rewards</font>
|
Анализировать url
|
|
x.com
|
X
|
Анализировать url
|
|
arxiv.org
|
<font color="#0080FF">Two Experts Are All You Need for Steering Thinking: Reinforcing Cognitive Effort in MoE Reasoning Models Without Additional Training</font>
|
Анализировать url
|
|
x.com
|
X
|
Анализировать url
|
|
dx.doi.org
|
<font color="#0080FF">The First Few Tokens Are All You Need: An Efficient and Effective Unsupervised Prefix Fine-Tuning Method for Reasoning Models</font>
|
Анализировать url
|
|
x.com
|
X
|
Анализировать url
|
|
arxiv.org
|
<font color="#0080FF">Dancing with Critiques: Enhancing LLM Reasoning with Stepwise Natural Language Self-Critique</font>
|
Анализировать url
|
|
x.com
|
X
|
Анализировать url
|
|
arxiv.org
|
<font color="#0080FF">Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains</font>
|
Анализировать url
|
|
x.com
|
X
|
Анализировать url
|
|
arxiv.org
|
<font color="#0080FF">CDE: Curiosity-Driven Exploration for Efficient Reinforcement Learning in Large Language Models</font>
|
Анализировать url
|
|
x.com
|
X
|
Анализировать url
|
|
arxiv.org
|
<font color="#0080FF">Can't See the Forest for the Trees: Benchmarking Multimodal Safety Awareness for Multimodal LLMs</font>
|
Анализировать url
|
|
arxiv.org
|
<font color="#0080FF">Insight Over Sight: Exploring the Vision-Knowledge Conflicts in Multimodal LLMs</font>
|
Анализировать url
|
|
arxiv.org
|
<font color="#0080FF">Chain-of-Jailbreak Attack for Image Generation Models via Step by Step Editing</font>
|
Анализировать url
|
|
arxiv.org
|
<font color="#0080FF">VisFactor: Benchmarking Fundamental Visual Cognition in Multimodal Large Language Models</font>
|
Анализировать url
|
|
arxiv.org
|
<font color="#0080FF">BatonVoice: An Operationalist Framework for Enhancing Controllable Speech Synthesis with Linguistic Intelligence from LLMs</font>
|
Анализировать url
|
|
x.com
|
X
|
Анализировать url
|
|
arxiv.org
|
<font color="#0080FF">VISTA: Enhancing Vision-Text Alignment in MLLMs via Cross-Modal Mutual Information Maximization</font>
|
Анализировать url
|
|
x.com
|
X
|
Анализировать url
|
|
arxiv.org
|
<font color="#0080FF">GPT4Video: A Unified Multimodal Large Language Model for lnstruction-Followed Understanding and Safety-Aware Generation</font>
|
Анализировать url
|
|
arxiv.org
|
<font color="#0080FF">Macaw-LLM: Multi-Modal Language Modeling with Image, Audio, Video, and Text Integration</font>
|
Анализировать url
|
|
github.com
|
Project
|
Анализировать url
|
|
arxiv.org
|
<font color="#0080FF">Is ChatGPT A Good Translator? Yes With GPT-4 As The Engine</font>
|
Анализировать url
|
|
arxiv.org
|
<font color="#0080FF">Not All Countries Celebrate Thanksgiving: On the Cultural Dominance in Large Language Models</font>
|
Анализировать url
|
|
arxiv.org
|
<font color="#0080FF">Refuse Whenever You Feel Unsafe: Improving Safety in LLMs via Decoupled Refusal Training</font>
|
Анализировать url
|
|
x.com
|
X
|
Анализировать url
|
|
arxiv.org
|
<font color="#0080FF">GPT-4 is too Smart to be Safe: Stealthy Chat with LLMs via Cipher</font>
|
Анализировать url
|
|
arxiv.org
|
<font color="#0080FF">Exploring Human-Like Translation Strategy with Large Language Models</font>
|
Анализировать url
|
|
arxiv.org
|
<font color="#0080FF">Improving Machine Translation with Human Feedback: An Exploration of Quality Estimation as a Reward Model</font>
|
Анализировать url
|
|
arxiv.org
|
<font color="#0080FF">Document-Level Machine Translation with Large Language Models</font>
|
Анализировать url
|
|
arxiv.org
|
<font color="#0080FF">Learning to Remember Translation History with a Continuous Cache</font>
|
Анализировать url
|
|
arxiv.org
|
<font color="#0080FF">Neural Machine Translation with Reconstruction</font>
|
Анализировать url
|
|
github.com
|
code
|
Анализировать url
|
|
arxiv.org
|
<font color="#0080FF">Context Gates for Neural Machine Translation</font>
|
Анализировать url
|
|
github.com
|
code
|
Анализировать url
|
|
arxiv.org
|
<font color="#0080FF">Modeling Coverage for Neural Machine Translation</font>
|
Анализировать url
|
|
github.com
|
code
|
Анализировать url
|