Original research
The work behind the intelligence.
Explore the papers behind retrieval, reasoning, and evaluation, from research foundations to FrontierFinance.
Preprint · 2026
FrontierFinance: A Challenging Benchmark for Measuring Frontier Intelligence of Finance Agents
An open benchmark of 220 financial research questions with expert-written evaluation criteria.
ACM CAIS · 2026
OpaqueToolsBench: Learning Nuances of Tool Behavior Through Interaction
Learning tool behavior from execution feedback when documentation is incomplete.
ARXIV · 2025
Lost in the Maze: Overcoming Context Limitations in Long-Horizon Agentic Search
Keeping long research trajectories focused through lightweight context management.
Preprint · 2024
Promptriever: Instruction-Trained Retrievers Can Be Prompted Like Language Models
Retrieval that follows instructions about relevance, not just a search query.
EMNLP · 2024
Dancing in Chains: Reconciling Instruction Following and Faithfulness in Language Models
Studying the tension between following instructions and remaining faithful to evidence.
EMNLP · 2024
RAG-QA Arena: Evaluating Domain Robustness for Long-Form Retrieval-Augmented Question Answering
Evaluating long-form retrieval-augmented answers across domains.
ACL FINDINGS · 2024
Selective ‘Selective Prediction’: Reducing Unnecessary Abstention in Vision-Language Reasoning
Reducing unnecessary abstention in visual reasoning tasks.
NAACL · 2024
When Does Monolingual Data Help Multilingual Translation: The Role of Domain and Model Scale
Examining when monolingual data helps translation across languages.
TACL · 2024
Lost in the Middle: How Language Models Use Long Contexts
Understanding how the position of evidence affects long-context performance.
ACL FINDINGS · 2024
Deal, or No Deal (or Who Knows)? Forecasting Uncertainty in Conversations Using Large Language Models
Forecasting uncertainty in conversations with language models.
NATURE MACHINE INTELLIGENCE · 2023
Improving Wikipedia Verifiability With AI
Helping identify and improve the evidence supporting Wikipedia claims.
CVPR · 2023
Cross-Domain Image Captioning With Discriminative Finetuning
Adapting image captioning across domains.
ICLR · 2023
Can Discrete Information Extraction Prompts Generalize Across Language Models?
Testing whether extraction prompts transfer between language models.
Bring the question behind your next decision.
See how Samaya handles a research question relevant to your team.