Andrew Drozdov

About me

I’m a research scientist at Databricks working on scaling intelligence through search and for search. I design and train search methods that integrate with agents, helping them find information, reason across sources, and solve complex tasks. Recent projects include KARL, Instructed Retriever, AutoIndex, and FreshStack.

I received my PhD from UMass Amherst CICS, advised by Andrew McCallum and Mohit Iyyer, and my MS in computer science from NYU, where I worked with Samuel Bowman and Kyunghyun Cho. Previously, I worked at Google and IBM.

I’ve reviewed more than 100 papers at AI, information retrieval, and NLP conferences and served as an area chair and senior area chair.

Say hello: andrew.drozdov@databricks.com

Updates

Virginia Tech Frontier AI Seminar, hosted by Tu Vu: “Three Reasons You Should Train Search Agents.”
SIGIR LLMUP Workshop keynote: “Training Adaptive Search Agents for Dynamic Environments.”
Databricks blog: “3x Faster Search: Parallel Test-Time Scaling with Instructed-Retriever-1.”
NYU Fundamentals of Machine Learning guest lecture, hosted by Kyunghyun Cho: “Grounded Reasoning in Real-World ML Systems.”
Databricks blog: “Instructed Retriever: Unlocking System-Level Reasoning in Search Agents.”

All updates →

Publications

  1. AutoIndex: Learning Representation Programs for Retrieval
    Sam O’Nuallain, Nithya Rajkumar, Ramya Narayanasamy, Hanna Jiang, Shreyas Chaudhari, and Andrew Drozdov
    arXiv 2026
  2. Can QPP Choose the Right Query Variant? Evaluating Query Variant Selection for RAG Pipelines
    Negar Arabzadeh, Andrew Drozdov, Michael Bendersky, and Matei Zaharia
    In SIGIR 2026
  3. KARL: Knowledge Agents via Reinforcement Learning
    Jonathan D. Chang, Andrew Drozdov, Shubham Toshniwal, Owen Oertell, Alexander Trott, Jacob Portes, Abhay Gupta, Pallavi Koppol, Ashutosh Baheti, Sean Kulinski, Ivan Zhou, Irene Dea, Krista Opsahl-Ong, Simon Favreau-Lessard, Sean Owen, Jose Javier Gonzalez Ortiz, Arnav Singhvi, Xabi Andrade, Cindy Wang, Kartik Sreenivasan, Sam Havens, Jialu Liu, Peyton DeNiro, Wen Sun, Michael Bendersky, and Jonathan Frankle
    arXiv 2026
  4. A State-of-the-Art SQL Reasoning Model using RLVR
    Alnur Ali, Ashutosh Baheti, Jonathan Chang, Ta-Chung Chi, Brandon Cui, Andrew Drozdov, Jonathan Frankle, Abhay Gupta, Pallavi Koppol, Sean Kulinski, Jonathan Li, Dipendra Misra, Krista Opsahl-Ong, Jose Javier Gonzalez Ortiz, Matei Zaharia, and Yue Zhang
    arXiv 2025
  5. FreshStack: Building Realistic Benchmarks for Evaluating Retrieval on Technical Documents
    Nandan Thakur, Jimmy Lin, Sam Havens, Michael Carbin, Omar Khattab, and Andrew Drozdov
    arXiv 2025
  6. The Second Tutorial on Retrieval-Enhanced Machine Learning: Synthesis and Opportunities
    Fernando Diaz, Andrew Drozdov, To Eun Kim, Alireza Salemi, and Hamed Zamani
    In SIGIR 2025
  7. Drowning in Documents: Consequences of Scaling Reranker Inference
    Mathew Jacob, Erik Lindgren, Matei Zaharia, Michael Carbin, Omar Khattab, and Andrew Drozdov
    arXiv 2024
  8. Retrieval-Enhanced Machine Learning: Synthesis and Opportunities
    Fernando Diaz, Andrew Drozdov, To Eun Kim, Alireza Salemi, and Hamed Zamani
    SIGIR-AP 2024
  9. Unlocking Natural Language Generalization with Adaptive Retrieval-based Methods
    Andrew Drozdov
    2024
  10. ReDMM: Retrieval Driven Memory Manager
    Andrew Drozdov, Andrew McCallum, and Mohit Iyyer
    2024
  11. Multistage collaborative knowledge distillation from a large language model for semi-supervised sequence generation
    Jiachen Zhao, Wenlong Zhao, Andrew Drozdov, Benjamin Rozonoyer, Md Arafat Sultan, Jay Yoon Lee, Mohit Iyyer, and Andrew McCallum
    In ACL 2024
  12. PaRaDe: Passage Ranking using Demonstrations with LLMs
    Andrew Drozdov, Honglei Zhuang, Zhuyun Dai, Zhen Qin, Razieh Rahimi, Xuanhui Wang, Dana Alon, Mohit Iyyer, Andrew McCallum, Donald Metzler, and Kai Hui
    In EMNLP (Findings) 2023
  13. kNN-LM Does Not Improve Open-ended Text Generation
    Shufan Wang, Yixiao Song, Andrew Drozdov, Aparna Garimella, Varun Manjunatha, and Mohit Iyyer
    In EMNLP 2023
  14. Compositional Semantic Parsing with Large Language Models
    Andrew Drozdov, Nathanael Schärli, Ekin Akyürek, Nathan Scales, Xinying Song, Xinyun Chen, Olivier Bousquet, and Denny Zhou
    In ICLR 2023
  15. You can’t pick your neighbors, or can you? When and how to rely on retrieval in the kNN-LM
    Andrew Drozdov, Shufan Wang, Razieh Rahimi, Andrew McCallum, Hamed Zamani, and Mohit Iyyer
    In EMNLP (Findings) 2022
  16. Inducing and Using Alignments for Transition-based AMR Parsing
    Andrew Drozdov, Jiawei Zhou, Radu Florian, Andrew McCallum, Tahira Naseem, Yoon Kim, and Ramon Fernandez Astudillo
    In NAACL 2022
  17. Improved Latent Tree Induction with Distant Supervision via Span Constraints
    Zhiyang Xu, Andrew Drozdov, Jay Yoon Lee, Tim O’Gorman, Subendhu Rongali, Dylan Finkbeiner, Shilpa Suresh, Mohit Iyyer, and Andrew McCallum
    In EMNLP 2021
  18. Unsupervised Parsing with S-DIORA: Single Tree Encoding for Deep Inside-Outside Recursive Autoencoders
    Andrew Drozdov, Subendhu Rongali, Yi-Pei Chen, Tim O’Gorman, Mohit Iyyer, and Andrew McCallum
    In EMNLP 2020
  19. The Impact of Preprint Servers in the Formation of Novel Ideas
    Swarup Satish, Zonghai Yao, Andrew Drozdov, and Boris Veytsman
    In EMNLP (Workshop on Scholarly Document Processing) 2020
  20. Unsupervised Labeled Parsing with Deep Inside-Outside Recursive Auto-Encoders
    Andrew Drozdov, Patrick Verga, Yi-Pei Chen, Mohit Iyyer, and Andrew McCallum
    In EMNLP 2019
  21. Unsupervised Latent Tree Induction with Deep Inside-Outside Recursive Autoencoders
    Andrew Drozdov, Patrick Verga, Mohit Yadav, Mohit Iyyer, and Andrew McCallum
    In NAACL 2019
  22. Do latent tree learning models identify meaningful structure in sentences?
    Adina Williams, Andrew Drozdov, and Samuel R. Bowman
    TACL 2018
  23. Emergent Communication in a Multi-Modal, Multi-Step Referential Game
    Katrina Evtimova, Andrew Drozdov, Douwe Kiela, and Kyunghyun Cho
    In ICLR 2018
  24. The Coadaptation Problem when Learning How and What to Compose
    Andrew Drozdov, and Samuel R. Bowman
    In ACL (Workshop on Representation Learning for NLP) 2017