Practical AI assistants built around your own documents, procedures, and operational knowledge
We build retrieval-augmented generation (RAG) chatbots and vector-database assistants for organisations that need faster access to the right information without forcing staff to dig through folders, PDFs, spreadsheets, or scattered internal notes.
That includes work for mining, oil and gas, and healthcare contexts where the real value is not a flashy chatbot demo. It is giving staff a safer, faster way to ask questions against the material they already rely on, with document grounding, traceable references, and a tighter fit to the way the organisation actually works.
Discuss a RAG chatbot project
Where this fits best
- Mining and oil and gas teams with operational documents, manuals, procedures, shutdown plans, or technical reference material that staff need to search quickly
- Healthcare teams that need easier access to internal knowledge, process notes, or structured reference material without relying on tribal knowledge
- Businesses with large document sets where keyword search is not enough and staff need answers in plain English with supporting context
- Teams that want AI support, but need the assistant grounded in their own material rather than generic internet answers
What we actually build
- Document-aware chat interfaces: ask questions in plain language and get answers grounded in uploaded internal content
- Vector database retrieval pipelines: chunking, embedding, indexing, and retrieval workflows tuned to the source material
- Multi-file ingestion: support for PDFs, spreadsheets, and other operational reference documents
- Traceable answers: answers can point back to the relevant document or section instead of pretending certainty
- Domain shaping: prompts, retrieval rules, and interface behaviour adjusted to the operating context, not left as generic AI defaults
- Practical integration paths: deploy as a standalone internal tool or connect it to an existing workflow where that makes more sense
Why this matters in operational environments
In mining, oil and gas, and healthcare settings, people often need answers quickly, but the underlying knowledge is spread across long documents, older files, and systems that were not built for conversational access. A well-built RAG assistant can reduce that friction without pretending to replace judgement, training, or process discipline.
The point is to shorten the path from question to grounded answer, make onboarding easier for newer staff, and reduce time lost to document hunting, inconsistent handover knowledge, or repeated basic questions.
Recent example
Recent work has included building RAG vector-database chatbots for oil and gas, mining, and healthcare use cases where staff needed a better way to query internal material. That kind of project typically involves document ingestion, chunking, embeddings, retrieval logic, prompt shaping, and interface work so the assistant is actually useful in the target environment instead of being a generic chatbot wrapper.
Enterprise-scale engineering document estates
The same pattern scales well beyond a single document set. For engineering-driven organisations sitting on document estates measured in terabytes — drawing registers, specifications, standards, project archives, handover packs, decades of reports — we scope and build retrieval over the whole estate rather than a cherry-picked corpus.
- PostgreSQL + pgvector architecture: open, proven retrieval infrastructure that runs in your environment, fits existing database operations, and avoids vendor lock-in on the search layer
- Staged ingestion at scale: start with a pilot corpus to prove answer quality, then scale out ingestion with deduplication, OCR for scanned documents, and format handling for the long tail of legacy files
- Access-aware retrieval: answers respect who is allowed to see what, so the assistant doesn't become a way around document permissions
- Flexible engagement: senior delivery on a time-and-materials basis, scoped in stages — no big-bang program commitment before the approach has proven itself on your own material
Our own mining document assistant, MineAction, is the working base for this pattern — the architecture extends from a single-team tool to estate-wide search as the corpus grows.
Discuss an enterprise document estate
Onyx: self-hosted open-source enterprise search and RAG
Where an organisation prefers a ready-made platform over a fully custom build, we now deploy and operate Onyx — an open-source (MIT-licensed) enterprise search and RAG system — on your own infrastructure. It provides a production chat interface, hybrid search, and answers grounded in your source material with citations, with no per-user licence fees and without handing your documents to a third-party cloud service.
- Connectors to where your knowledge already lives: SharePoint and Microsoft 365, Teams, file shares, websites, and many other sources, indexed into one searchable system
- Grounded, cited answers: every answer references the source document or page, and assistants can be scoped so they only answer from a chosen set of material
- Self-hosted and model-flexible: runs in your environment with local embeddings and a choice of local or cloud answering models, so sensitive content can stay on-premises
- Permission-aware and observable: retrieval can respect existing access controls, with ingestion progress you can monitor
- Custom extension where it counts: we add the bespoke pieces a platform doesn't cover out of the box — for example extracting legacy Outlook PST/OST email archives, or connectors to in-house systems
Whether the better fit is Onyx or a custom PostgreSQL + pgvector build depends on your sources, scale, and control requirements. We help you choose the right approach rather than pushing a single path, and we can stand up a working Onyx instance on a representative slice of your own documents so you can judge answer quality before committing.
What a sensible first version usually looks like
- Choose one document set: start with the material people actually struggle to search today
- Build a grounded prototype: ingestion, vector search, and a constrained chat interface
- Test with real questions: compare answer quality against the kinds of questions staff already ask
- Improve trust and fit: tighten retrieval, answer format, references, and interface behaviour
- Expand only after it proves useful: add more document sets, workflows, or integration points once the first use case works
Thinking about a RAG chatbot for your own document set?
We can help scope a useful first version around one operational problem area instead of overbuilding too early.
Request scoping options
Need the next step, not another generic read?
Discuss a software bottleneck
·
Book an IHTMaps workflow review
·
Request a website quote
Best fit for inherited systems, spreadsheet-heavy workflows, internal tools, inspection processes, and websites that need better enquiry flow or calmer technical ownership.
Call 0432 000 583
if you want to talk through the current bottleneck directly.
E-business card (QR ready) for conferences and in-person shares. · Site map
Copyright © 2026
Industrial Hypertext - Software Development Perth, Western Australia | All rights reserved