Public records search assistant
Search results retain page-level provenance and publication context.
- For
- Journalists and civic research organizations
- Solves
- Large public record collections are hard to search meaningfully.
- Delivers
- Cited search results and research folders
- Built in
- about 5 weeks of creation time, MVP in 5 days
- Investment
- $10,000 for the MVP, $39,500 for the full product
- Run it
- Inside your business, or as part of your offer to clients
What it does
For journalists and civic research organizations, turn legitimately accessible records and document metadata into cited search results and research folders.
- Index documents.
- Recognize scanned text.
- Filter dates and agencies.
- Retrieve passages.
- Preserve document context.
- Export citations.
What goes in, what comes out
- Legitimately accessible records
- Document metadata
AI drafts, people review. Source-linked assistant and administrator console.
- Cited search results
- Research folders
How it works
The workflow
- InStart with
Legitimately accessible records and document metadata
- 1
Add an approved collection
- 2
Assign source owners and access rules
- 3
Test representative questions
- 4
Let users ask questions
- 5
Retrieve supporting passages
- 6
Answer or request clarification
- 7
Hand off unresolved cases with their context
- OutFinish with
Cited search results and research folders
AI does the heavy lifting, people stay in charge
Retrieve permitted passages and generate answers constrained to those sources. Use structured rules for transactional facts. Detect missing context and refuse to invent unsupported details. Store reviewer corrections for evaluation and controlled knowledge updates.
What your team sees
Key screens: Collection search, passage viewer, citation export. Give end users a simple search or conversation surface with short answers and expandable citations. Administrators get source status, unanswered questions and handoff queues. Show the source date beside relevant answers. Keep conversation context available to the staff member receiving an escalation. In this product, the first view is collection search, followed by passage viewer and citation export.
Accounts and administration
Source ownership, document permissions, freshness checks, conversation history, human handoff, feedback, test questions, usage limits and access logs.
Integrations and data access
Official publications, agency document stores and approved service workflows. Approved knowledge repositories, websites, service desks and staff messaging systems. Validate access inheritance and use read-only ingestion for the initial deployment. These are candidate integration categories, not verified supported connectors.
How we build it
We build with our own AI software development factory, so most implementations take days to a few weeks of creation time, not months. You see working software at every step, and exact timing depends on availability.
- 1
Scoping call
Day 1Thirty minutes on your process, your data and how you want to run it: for your own team, or for your clients. You get a fixed scope and price for the MVP.
- 2
MVP
5 daysOne buyer segment, one recurring use case; first modules: index documents; recognize scanned text. Manual review in the loop. Built by our AI software factory.
- 3
Paid pilot
6 daysAccounts, roles, review states, audit trail and the first integration, hardened for two to three paying pilot customers.
- 4
Full product
2 weeksRemaining modules: retrieve passages; preserve document context; export citations. Self-serve onboarding, billing, monitoring and the wider integration set.
- 5
Run and improve
MonthlyWe host, monitor and improve it for a fixed monthly fee, or hand it over to your team. How the retainer works.
Why we start with an MVP
An MVP, or minimum viable product, is the smallest version that your users can actually work with. It is not a cheap version of the full solution. It is a test, built to answer the questions that decide whether the rest is worth building.
- Pick the riskiest assumption. Here: will journalists and civic research organizations use it to solve "large public record collections are hard to search meaningfully"?
- Build only what tests it. One team, one use case, a few core modules. People do the rest by hand for now.
- Run a paid pilot. Restrict the assistant to one collection and test answered, ambiguous and unanswerable questions.
- Measure, then decide. Track relevant retrieval and citation accuracy. Then expand, change course or stop, with evidence instead of opinions.
MVP scope for this solution. Begin with journalists and civic research organizations and one recurring use case. Build the first two modules: index documents; recognize scanned text. Provide operator assistance for the third module: filter dates and agencies. Deliver cited search results and research folders through a manual review queue. Perform other necessary full-scope functions manually during the pilot. Include all applicable access, accuracy and professional-review controls from the start.
After the MVP. After paid pilots establish value, automate the remaining modules: retrieve passages; preserve document context; export citations. Add one validated source integration, reusable customer configuration and recurring delivery. Expand to additional teams, document formats or languages only after testing the new scope.
What the build depends on. Permission-filtered retrieval, document versioning, a question evaluation set, staff handoff and a source update process. Reliability depends on source quality and scope.
Investment
A planning range to start the conversation, not a quote. You pay per phase, so you can stop after the MVP.
- Phase 1
MVP
One buyer segment, one recurring use case; first modules: index documents; recognize scanned text. Manual review in the loop.
- Phase 2
Paid pilot
Accounts, roles, review states, audit trail and the first integration, hardened for two to three paying pilot customers.
- Phase 3
Full product
Remaining modules: retrieve passages; preserve document context; export citations. Self-serve onboarding, billing, monitoring and the wider integration set.
Indicative total, MVP to full product$39,500about 5 weeks of creation time · start with the MVP from $10,000
Running costs per month
A rough indication of monthly hosting and AI model costs once it is live, not tested. Real costs depend on usage, file sizes and the models chosen.
| Stage | Hosting and infrastructure | AI usage | Total per month |
|---|---|---|---|
| MVP and paid pilotabout 3 customers | $50–$100 | $60–$120 | $110–$220 |
| Full productabout 50 customers | $190–$380 | $530–$1,050 | $720–$1,430 |
Run it or resell it
For your own team
Journalists and civic research organizations run it inside the business: legitimately accessible records and document metadata in, cited search results and research folders out, reviewed by your people.
As part of your offer
Agencies, consultancies and software companies can offer it to their own clients under their brand. We build and maintain it; you sell and deliver it.
Your brand, or this one
Run it under your own brand, or start from this concept style.
- primary
#359127 - accent
#9e54c9 - surface
#e6f1e4 - ink
#22201e
- Headings
- Playfair Display
- Text
- Source Sans 3
- Voice
- Plain-spoken, neutral, accountable
Selling it to your own clients: the go-to-market playbook
Pricing to test
Test USD 500-2,000 setup plus USD 150-600 monthly for one defined source collection and usage allowance. Price multi-location deployments and specialist support separately. Validate willingness to pay; these are hypotheses.
Message to test
Public records search assistant for journalists and civic research organizations. Search results retain page-level provenance and publication context. Demonstrate the claim through a searchable public-record collection demo.
Where to find buyers
Investigative journalism communities
Lead magnet
A searchable public-record collection demo
The first 30 days
- Week 1: interview five prospective buyers in this segment: journalists and civic research organizations. Ask to see a recent example of the problem and their current process.
- Week 2: prepare this demonstration using authorized or synthetic material: a searchable public-record collection demo.
- Week 3: present it through investigative journalism communities and seek one narrowly scoped paid pilot.
- Week 4: review relevant retrieval, citation accuracy, total delivery effort and a concrete renewal decision before increasing scope.
Paid pilot
Restrict the assistant to one collection and test answered, ambiguous and unanswerable questions. Run supervised use before wider rollout. Measure correctness, escalation quality and staff effort. For this solution, use legitimately accessible records and document metadata and evaluate cited search results and research folders. Agree success thresholds with the buyer before starting; collect a baseline for relevant retrieval, citation accuracy. A positive signal is payment and repeat use with acceptable quality and delivery cost, not a favorable demo reaction alone.
Success metrics
Relevant retrieval, citation accuracy
Retention and expansion
Review unanswered questions and source freshness monthly. Expand to another source collection or team only after the existing assistant meets its agreed accuracy and handoff criteria.
Why clients would pick it
A maintained domain knowledge collection, realistic evaluation questions, useful escalation paths and integrations in the customer’s daily work. For this solution, build around search results retain page-level provenance and publication context. This advantage requires execution and accumulated customer trust; the base model alone is not a defensible asset.
Alternatives and positioning
Manual search, static FAQs, general chat tools and support or intranet suites. Differentiate on this specific proposed advantage: search results retain page-level provenance and publication context. Test it against the buyer's current method on the same task. Competitor coverage and uniqueness have not been established.
Main delivery costs
Document ingestion, retrieval and generation, source maintenance, support, evaluation and staff time handling unresolved cases.
Safeguards
Preserve official source versions, accessibility and audit records. Confirm agency-specific procurement, records and data handling requirements during discovery. Validate source access and reviewer availability during the pilot. Maintain customer-level access, data deletion controls and a record of final approvals.