Original research

The work behind the intelligence.

Explore the papers behind retrieval, reasoning, and evaluation, from research foundations to FrontierFinance.

Preprint · 2026

FrontierFinance: A Challenging Benchmark for Measuring Frontier Intelligence of Finance Agents

An open benchmark of 220 financial research questions with expert-written evaluation criteria.

ACM CAIS · 2026

OpaqueToolsBench: Learning Nuances of Tool Behavior Through Interaction

Learning tool behavior from execution feedback when documentation is incomplete.

ARXIV · 2025

Lost in the Maze: Overcoming Context Limitations in Long-Horizon Agentic Search

Keeping long research trajectories focused through lightweight context management.

Preprint · 2024

Promptriever: Instruction-Trained Retrievers Can Be Prompted Like Language Models

Retrieval that follows instructions about relevance, not just a search query.

EMNLP · 2024

Dancing in Chains: Reconciling Instruction Following and Faithfulness in Language Models

Studying the tension between following instructions and remaining faithful to evidence.

EMNLP · 2024

RAG-QA Arena: Evaluating Domain Robustness for Long-Form Retrieval-Augmented Question Answering

Evaluating long-form retrieval-augmented answers across domains.

ACL FINDINGS · 2024

Selective ‘Selective Prediction’: Reducing Unnecessary Abstention in Vision-Language Reasoning

Reducing unnecessary abstention in visual reasoning tasks.

NAACL · 2024

When Does Monolingual Data Help Multilingual Translation: The Role of Domain and Model Scale

Examining when monolingual data helps translation across languages.

TACL · 2024

Lost in the Middle: How Language Models Use Long Contexts

Understanding how the position of evidence affects long-context performance.

ACL FINDINGS · 2024

Deal, or No Deal (or Who Knows)? Forecasting Uncertainty in Conversations Using Large Language Models

Forecasting uncertainty in conversations with language models.

NATURE MACHINE INTELLIGENCE · 2023

Improving Wikipedia Verifiability With AI

Helping identify and improve the evidence supporting Wikipedia claims.

CVPR · 2023

Cross-Domain Image Captioning With Discriminative Finetuning

Adapting image captioning across domains.

ICLR · 2023

Can Discrete Information Extraction Prompts Generalize Across Language Models?

Testing whether extraction prompts transfer between language models.

Bring the question behind your next decision.

See how Samaya handles a research question relevant to your team.