Early Case Intelligence and Culling
Analytics, deduplication, email threading, and communications mapping applied before full review begins. Modern ECA can reduce active review volume by 60 to 80 percent. GDF's process makes every reduction decision documented and defensible.
What This Solves
After collection, you are looking at several hundred thousand documents. Most of them are not relevant. A substantial portion are exact duplicates sitting in multiple custodians' inboxes. Hundreds of email threads are present in fragmented form, with the same conversation appearing across a dozen custodians at different stages of the reply chain. Review all of that document by document and you have weeks of attorney time spent on material the court will never see.
Early case intelligence (ECA) is the phase where GDF applies analytics and reduction techniques before a single document goes to an attorney for review. The goal is to move only the right data forward: unique, potentially responsive documents, in a form reviewers can navigate efficiently. The process is not about cutting corners. Every reduction decision is documented, methodology is logged, and the parameters used to exclude documents can be reproduced and explained. Modern ECA, done properly, reduces review volume by 60 to 80 percent without losing responsive material.
Deduplication: Global and Custodial
Hash-based deduplication is the foundation of volume reduction. GDF computes cryptographic hash values (MD5 and SHA-256) for every document in the collection. Two files with identical hash values are exact duplicates, byte for byte. One copy moves to review. The duplicate is suppressed but not destroyed: GDF maintains a cross-reference log mapping every suppressed duplicate to its canonical copy, so that custodian-specific context is preserved if needed.
The choice between global and custodial deduplication matters, and it needs to be agreed upon with counsel before processing begins. Global deduplication removes a document from review if it appears anywhere in the collection, regardless of which custodian held it. A single copy proceeds to review. Custodial deduplication removes duplicates only within each custodian's own data set. If the same document appears in five custodians' email, all five copies proceed to review (one per custodian), because the fact that each person held the document may itself be relevant.
GDF documents which deduplication methodology was applied, the parameters used, the hash algorithm, and the resulting document counts at each stage. This documentation goes into the processing log and is available for opposing counsel review or court submission if challenged.
Email Threading and Inclusive Email Analysis
An email thread starts as a single message. By the time discovery is collected, that same conversation may exist as dozens of individual messages, each held by a different custodian at a different point in the reply chain. Every message in a thread is technically a separate document, but reviewing each separately is redundant: later messages in the thread quote all prior content.
Email threading identifies the most inclusive email in each conversation thread. The inclusive email is the latest message in the chain that contains all prior content within it. If a reviewer reads the inclusive email, they have seen everything said in that thread. Earlier messages (non-inclusive) are flagged as thread members but can be set aside unless a specific review decision requires looking at earlier branch points separately.
Threading also handles replies, forwards, and divergent branches. GDF's threading analysis identifies where a thread splits (when someone forwards a message and starts a parallel conversation) and ensures that each branch has its own inclusive email identified. The result is a structured view of every conversation, not a flat list of individual messages.
Near-Duplicate Detection
Exact hash deduplication removes byte-identical files. Near-duplicate detection handles the next layer: documents that are substantially similar but not identical. A contract with a single clause changed, a draft email with one paragraph removed, or a report with an updated figure, each is a different document by hash but may require only a fraction of the review time if identified as a near-duplicate of something already reviewed.
GDF's near-duplicate analysis computes similarity scores between documents and groups near-duplicates into clusters. Reviewers can see the pivot document for each cluster, review the differences within the cluster, and apply consistent review decisions across clustered documents without reading each one independently. The similarity threshold used and the resulting clusters are documented in the processing log.
Communications Mapping and Data Visualization
Communications mapping analyzes the email and message metadata across the entire collection to build a structured picture of who communicated with whom, how frequently, and during which time periods. This analysis serves two purposes. First, it helps counsel identify the most active communicators in the relevant time period, which often focuses document review on the custodians and relationships that matter most to the theory of the case. Second, it identifies unexpected communications: messages between parties who should not have been in contact, or communication patterns that are anomalous relative to the surrounding time period.
GDF delivers communications maps as visual outputs: network graphs showing communication volume and frequency between custodians, timeline charts showing communication spikes around key events, and frequency tables exportable to Excel or PDF. These visualizations support attorney strategy sessions and can be adapted for presentation to clients or co-counsel.
Search Strategy Development
Keyword searches in eDiscovery are more consequential than most attorneys appreciate. An overly broad search returns tens of thousands of irrelevant hits that inflate review costs. An overly narrow search misses responsive documents and creates production gaps that become motion practice. GDF works with counsel to develop a search strategy grounded in the facts of the matter.
The process starts with a review of the case narrative: who are the key players, what were the key events, what documents does counsel need to find? GDF translates that narrative into a structured set of search terms with field-level targeting (searching sender fields differently than body text), date range parameters, custodian scoping, and proximity operators for multi-word combinations. GDF then runs the terms against the post-processing collection and reports hit counts so that counsel can see the impact of each term before committing to it. Terms that return an unusually high or low hit count relative to the matter's facts trigger review and refinement.
The final search protocol is documented in writing and becomes part of the matter record, supporting disclosure of search methodology under Rule 26 and demonstrating reasonable, proportionate collection efforts.
GDF's Early Case Intelligence Process
Ingestion and Indexing
Collected data is loaded into GDF's processing environment. Every file is extracted, metadata is captured, and the full text is indexed for search. Containers (ZIP archives, PST files, MSG files with attachments) are exploded into their component items while family relationships are preserved. The ingestion log records file counts, document counts, processing exceptions, and any items requiring special handling.
Analytics and Email Threading
GDF runs email threading across the full data set, identifying inclusive emails and building a structured conversation hierarchy for every thread. Near-duplicate clustering identifies groups of substantially similar documents and designates a pivot document for each cluster. Communications mapping generates custodian interaction data and timeline analytics. All analytical outputs are reviewed with counsel before any documents are excluded from review.
Deduplication: Global or Custodial
GDF applies hash-based deduplication using MD5 and SHA-256 algorithms, using the methodology specified by counsel (global or custodial). Family integrity is preserved throughout: if a parent document is suppressed as a duplicate, its attachments are suppressed with it; if a parent is active, all attachments remain active in the review set. The deduplication log maps every suppressed document to its canonical copy and records hash values for both.
Communications Mapping and Anomaly Identification
Communications analysis outputs are presented to counsel: network graphs of custodian interactions, volume timelines with annotations for key case events, and tables of high-frequency communicators. GDF identifies communication patterns that warrant closer attention, including unexpected communication between parties, communication spikes around key dates, and custodians with low collection volume relative to their expected role in the matter.
Search Strategy Development and Testing
GDF develops a structured search protocol in collaboration with counsel, including keyword terms, field targeting, date filters, custodian scoping, and proximity operators. Each term is tested against the collection before the protocol is finalized, with hit counts reported so that counsel can see the impact of individual terms. The agreed protocol is documented in writing before it is applied to generate the final review set.
Documented Reduction Decisions and Reporting
Before any data moves to the review platform, GDF produces an ECA summary report documenting each reduction step: documents received, documents after deduplication, documents after threading suppression, documents after search culling, and final review set count. The report includes the reduction percentages at each stage, the parameters used, and a summary of analytical findings. This report supports Rule 26(f) planning discussions and provides a record of reasonable, proportionate discovery steps.
Family Integrity and Defensibility
Family integrity is one of the most important constraints in ECA. A "family" in eDiscovery means a parent document and its attachments, or an email and its embedded images and forwarded attachments. Courts expect that families travel together through review and production: you cannot produce a contract without its exhibits, or an email without the spreadsheet that was attached to it. GDF's processing preserves family relationships at every reduction step.
When a non-inclusive email is suppressed by threading, its attachments are suppressed with it. When that email's inclusive copy proceeds to review, the attachments from all thread members are associated with it. No attachment is orphaned from its parent. When deduplication suppresses a document, the family of the active copy receives any unique attachments from suppressed copies. These rules are applied consistently, logged, and documented.
Every ECA engagement produces a defensibility log: a record of who made each reduction decision, what parameters were used, when the decision was applied, and what the resulting document counts were at each stage. If opposing counsel challenges the completeness of the review set, GDF can produce this log to demonstrate that the reduction methodology was reasonable, documented, and reproducible.
What GDF Delivers
- ECA summary report with before/after document counts at each reduction stage
- Deduplication log with hash values and custodian cross-reference map
- Email thread index identifying inclusive emails and thread members
- Near-duplicate cluster report with similarity scores and pivot document identification
- Communications mapping outputs: network graphs, timeline charts, and frequency tables
- Documented search protocol with per-term hit counts
- Defensibility log suitable for disclosure to opposing counsel or court submission
- Review-ready data set loaded into counsel's preferred review platform
Last reviewed and updated: April 2026
Hash-Based Deduplication
- MD5 and SHA-256 hash computation
- Global or custodial methodology
- Family integrity preserved
- Full cross-reference log maintained
Email Threading
- Inclusive email identification
- Thread branch handling
- Non-inclusive suppression with logging
- Reply, forward, and divergent path analysis
Near-Duplicate Detection
- Similarity scoring across the data set
- Cluster grouping with pivot document
- Documented similarity threshold
- Reduces redundant review within clusters
Communications Mapping
- Custodian network graphs
- Volume timelines with case event annotations
- Anomaly identification
- Exportable to Excel or PDF
Early Case Intelligence Consultation
GDF can run ECA on collected data within days of engagement. Contact us to discuss your matter, data volume, and review timeline.
Related Services
Review Platform
GDF's in-house review platform with flat-rate feature access, AI-assisted review queues, privilege workflows, and production handoff, all in one environment.
Learn MoreManaged Review and Expert Services
GDF-staffed review operations with project management, quality control, privilege review, and expert declarations for matters that require hands-on support.
Learn MoreCloud and SaaS Collections
Direct-source collection from Microsoft 365, Google Workspace, Slack, and other platforms to ensure ECA begins with complete and defensible input data.
Learn MoreReduce Review Volume Before Review Begins
GDF's early case intelligence services apply analytics, deduplication, and communications mapping to move only the right documents into review. Contact us to discuss your matter and data volume.