Agentic search policies
Agentic search systems learn or specify how to query, inspect evidence, continue reasoning, and stop under external budgets.
Reason and act
- reasonUnresolved question
- actSearch query
- observeRetrieved result
- repeatUpdated state
- stopAnswer or abstain
- ReAct
- Self-Ask
Outcome RL
- sampleSearch trajectory
- scoreFinal answer reward
- optimizePolicy update
- auditSearch cost and faithfulness
- Search-R1
- ReSearch
- R1-Searcher
Process and value RL
- stateEvidence state
- valueRetrieve or stop
- rewardProcess quality
- selectNext action
- terminateBounded result
- StepSearch
- Q-RAG
- DeepRAG
- HiPRAG
Integrated control
- emitControl token
- retrieveExternal evidence
- resolveConflict and utility
- citeEvidence selection
- answerFinal output
- GRIP
- Knowledgeable-R1
- RAG-RL