Hunter-RAG by Arc

Agentic retrieval

Give documents a filesystem.

Your agent already knows how to work a codebase. Hunter-RAG lets it work your documents the same way.

How?
agent~/papers
>
search_chunks("children's diagnostic accuracy lung ultrasound highest specificity")→ 5 chunks · #33 #32 #24 #35 …
read_chunk(32)→ table rows, no header row
read_chunk(24)→ a different table (adults)
read_chunk(35)→ trailing rows and notes

Pleural effusion, 74.5%.

74.5% is the adult study's value. The chunk it came from had lost its table.

Diagnostic accuracy of point-of-care ultrasound for pulmonary tuberculosis: A systematic review
Bigio et al. · PLoS ONE 2021 · doi 10.1371/journal.pone.0251236 · CC BY 4.0
section · sec_008
Diagnostic accuracy of lung ultrasound findings in adults.
Montuori 2019Sensitivity (95% CI)Specificity (95% CI)
Subpleural nodule72.5 (58.3–84.1)66.7 (52.1–79.2)
Lung consolidation78.4 (64.7–88.7)35.3 (22.4–49.9)
Pleural effusion19.6 (9.8–33.1)74.5 (60.4–85.7)
table · tab_004
section · sec_012
Diagnostic accuracy of lung ultrasound findings in children.
Heuvelings 2019Sensitivity (95% CI)Specificity (95% CI)
Interrupted pleural line78.4 (70.2–85.3)26.7 (14.6–41.9)
Consolidation45.6 (36.7–54.8)53.3 (37.9–68.3)
Pleural gap52.8 (43.7–61.8)57.8 (42.2–72.3)
table · tab_005
ranked by similarity to the question
#22
#23
#24
#25
#26
#27
#28
#29
#30
#31
#32
#33
#34
#35
  • medical_PMC8104425
  • section
  • sec_012 Children
  • table
  • tab_004 …lung ultrasound findings in adults
  • tab_005 …lung ultrasound findings in children
  • figure
  • fig_000
every unit has an address: medical_PMC8104425:table:tab_005
hunting…
sec_000sec_007sec_008sec_010sec_011sec_012tab_000tab_003tab_004tab_005tab_006fig_000fig_001

References

  1. 1medical_PMC8104425:table:tab_005row: Pleural effusion · Specificity 91.1 (78.8–97.5)
  2. 2medical_PMC8104425:section:sec_012“…in Heuvelings 2019 [37] is shown in Table 6.”

medical_PMC8104425:table:tab_004 · adults · inspected, not cited

Not a claim that the answer is right: the evidence shows where it came from, so you can check.

973 documents · 51,104 units
section · 32,407table · 10,092figure · 4,094code_block · 2,430equation · 2,081
medical_PMC8104425 › figure › fig_000
PRISMA study flowchart.no image extracted
the page, rendered as an image
vault/notes
  • lus-children.mdstatusconfirmedsource[[PMC8104425]]topiclung ultrasound→ [[PMC8104425]]
  • lus-first-pass.mdstatussupersededtopiclung ultrasound→ [[lus-children]]superseded · still searchable
  • protocol-imaging.mdstatusdraftprojectTB screening→ [[lus-children]]new · cites lus-children
SELECT * FROM doc_meta WHERE key = 'status'frontmatter, queryable · wikilinks, followable · no embeddings · no serverNo server is not offline: compiling cards and answering use the model you choose. Matching a task to a card does not.

The retrieval pipeline

chunkcompile to units
embed
vector indexunit store
top-k
rerank
stuff the promptthe agent navigates

Example questions

Financeannual reports · 10-K · earnings releases“Which segment's revenue grew fastest last year, and what does management say drove it?”table · segment resultssection · MD&Acalculatecite
Legal & regulatoryregulations · contracts · policies“What must a provider of a high-risk system do before it reaches the market, and where are the exceptions?”section · the articletrace_source → annexsection · definitionscite
Medicine & life sciencessystematic reviews · trial reports · labels“In the children's data, which finding had the highest specificity?”table · drop the adults'section · “shown in Table 6”table · commitcite
Engineering & operationsmanuals · specifications · procedures“What torque does the procedure give for step 4, and when does it change?”section · the proceduretable · torque valuesfootnote · conditionscite
Researchpapers · preprints · technical reports“Which equation defines the loss, and which ablation hurts the result most?”equation · the losstable · ablationscompare_unitscite

Some questions spend more tokens: reading the whole table or clause is what quality costs on questions like these.

Enterprise RAG, ready

In your environmentThe store and the agent run where your documents are. Bring any model: hosted, through your gateway, or on your own hardware.
Auditable by constructionEvery answer keeps its evidence, as unit addresses, and the full trajectory that found it.
Code you can readThe navigation core is public. The full source is open to licensed customers for audit.
Deployed with youArc's engineers deploy it into your systems and support it in production.
A chunk agent, given matching reading tools0%
Hunter-RAG0%

78 questions over six public documents (research papers, a medical review, an earnings release, a contract, software docs). Same model for both.

Not on every question: in some cases, on prose or on code, chunk retrieval can do as well or better.

Why not just grep?

Ten questions that count, group and filter across 106 notes.

A coding agent with grep and file tools0 steps
The same agent with Hunter-RAG in reach0 steps

Both: 10 of 10 correct.

Looking up a single phrase in a small corpus, grep is as good. Counting, grouping and comparing across a corpus is where it reads every file.

  • orientpeek_document · open_document
  • locatefind_units · grep_units · grep_within_unit
  • inspectpeek_unit
  • readexpand_unit · read_unit · read_table · read_figure
  • relateget_neighbors · trace_source · compare_units
  • computecalculate · commit_computation · cdu_sql
  • citecite_units

The right paper. The right row. The wrong table.

Agents got active. The documents stayed in pieces.

What if the document stayed whole, and the agent read it as it is?

Every part keeps its shape, its type and an address.

Chunking becomes compiling. Embedding and reranking leave the loop.

Search finds candidates. Hunting commits evidence.

The answer arrives with its evidence, not after it.

Where the answer is a number in a table, a clause in a regulation, a step in a procedure.

Built to run where your documents are.

Same model. Same questions. Only the shape of the documents changed.

One document or ten thousand. Query them like a database.

Parsed on the spot. When a figure doesn't come through, the agent reads the page.

Grep finds a phrase. Ask it to count, and it reads every file.

Your notes become a knowledge base. No embeddings. No vector database. No server.Arc's own knowledge base runs on Hunter-RAG. Every Arc agent session starts with a card it compiled.

About twenty primitives, built for units, not for text.

Stop feeding agents fragments. Let them hunt.

The built-in agent

pip install hunter-ragcdu your.pdf

Ask; /cite shows the units it read, /trace every step.

Your own agent

Give Claude Code, Codex or any agent that runs a shell the Hunter-RAG skill. It hunts with the same primitives and cites unit addresses.

Enterprise and industrial deployments, built and supported by Arc's engineers.

How it works →