How many
~30 minimum, 100–300 is the sweet spot for a trial.
A short guide to assembling documents for a trial run. A well-chosen sample shows BDS at its best — sharp search, a clean entity graph, and useful entity profiles. Two things matter most: keep the set to one subject, and send enough documents to be representative.
~30 minimum, 100–300 is the sweet spot for a trial.
All about one case, matter, or topic — not a mix.
PDF (including scanned), Word, Excel, text, HTML.
BDS learns the entities and relationships within your documents — which works best when the documents belong together. Before indexing, BDS automatically measures how cohesive your set is and reports it as Tight, Medium, or Diffuse; a diffuse set is flagged with a recommendation to split it.
Everything centers on a single subject:
Several unrelated subjects in one set:
Have more than one subject? Send them as separate sets — BDS keeps each as its own project, which is exactly how it performs best.
More is not always better for a trial. Search quality plateaus early; the entity graph keeps enriching up to a few hundred documents. This range shows the full picture without a long run.
Works, but the entity graph and search have little to draw on — fine for a quick look.
The sweet spot. Search sharpens by ~60 documents; relationships keep building to ~250.
Handled fine, but for a trial it mostly adds processing time, not new insight.
Mixed formats and mixed document lengths in one set are fine. You do not need to convert or OCR anything — BDS reads scanned pages for you.
Privileged, regulated, or otherwise sensitive material belongs in production, not the trial. BDS is on-premise software: in a real deployment it runs on your own hardware, so your live corpus is processed in place and never leaves your control. The evaluation is a capability-and-fit check on a sample; production is where your confidential documents stay home.
From your sample set, a BDS evaluation produces:
Hybrid semantic + keyword search across every document, with grounded answers that cite their sources.
The people, organizations, and other entities in your set, with how they connect and co-occur across documents.
A concise, sourced summary card for each significant entity, drawn from your own documents.
The Tight / Medium / Diffuse cohesion result and a sense of how well the set size covers the subject.
One subject per set, roughly 100–300 documents, any mix of common formats, sent as-is — scanned PDFs included. Keep privileged material for an on-premise production deployment, and give us a representative sample for the trial.