Internal · Working draft
Firefly + Mangrove · Sep 2026
Insight Exposure & Risk Sciences Group
Solutions research & validation
Evaluating secure AI infrastructure and storage for Insight's two core use cases: case material summarization and document search/organization. All options assessed against ZDR (non-negotiable), HIPAA, and subpoena-risk requirements. Updated Sep 2026 with Harvey.ai as a potential all-in-one platform.
Secure AI environment
Storage solutions
Document pipeline
Architecture
Discovery
Ruled out
At a glance
| Solution | What it is | Custom agents / workflows | ZDR | HIPAA | Data residency | Effort | Cost | Subpoena exposure |
| Harvey.ai | Managed AI platform built for law firms and litigation. Document vault, workflow agents, and report generation in one product. 2,400+ legal orgs use it | Yes. Agent Builder (no-code), 500+ pre-built agents, scheduled recurring agents | Yes (all providers) | Unconfirmed | US/EU/AU options | Low | ~$1.2-2K/seat/mo | Customer data logically separated; ZDR across all models |
| LibreChat (self-hosted) | Open-source chat + agent platform with built-in RAG. Model-agnostic (plug in any LLM). Used at Shopify (10K daily users). Polished UI, scoped agents, hybrid search | Yes. Scoped agents with custom instructions, document sets, and model assignments. Chain agents for multi-step workflows. Fully customizable | Yes w/ Claude ZDR, Bedrock, or local models; configurable per agent | Enclosed; BAA only if using external LLM API | Cloudflare, DigitalOcean, Railway, AWS, GCP, or any Docker host | Low-Med | Hosting + API usage | Self-hosted; nothing leaves your infrastructure |
| Claude API + ZDR | Direct API access to Anthropic's Claude models. You build the application layer; Claude handles the reasoning. Best-in-class for summarization and structured output | Build your own. API provides the reasoning engine; you build the agent logic and UI | Yes | BAA avail. | Anthropic infra | Low-Med | Usage-based | Nothing stored (except flagged content, up to 2yr) |
| Claude on Bedrock | Same Claude models, but accessed through Amazon Web Services. Data stays in your AWS account. Good fit if Insight already has or wants AWS infrastructure | Via Bedrock Agents (AWS managed). Visual workflow builder, built-in knowledge bases | Yes (configurable) | Yes | Your AWS account | Medium | Usage-based | Your AWS account = your control |
| Claude on Vertex AI | Same Claude models through Google Cloud Platform. Data stays in your GCP project. Good fit if Insight already has or wants Google infrastructure | Via Vertex AI Agent Builder (GCP managed). Similar to Bedrock Agents | Unconfirmed toggle | Yes | GCP region select | Medium | Usage-based | Google as processor |
| Open-source LLM (self-hosted) | Run Llama, Mistral, or other open models on your own hardware. Maximum control but meaningfully behind Claude on complex summarization and citation accuracy | Pair with LibreChat or build custom. No built-in agent framework | Yes (full control; no external calls) | Enclosed; air-gapped capable | Cloudflare, AWS, GCP, or any GPU-capable host | High | GPU infra (~$500-2K/mo) | Nothing leaves your servers |
| Claude Teams/Enterprise | Anthropic's consumer product (what Insight uses now). Chat interface with 30-day data retention. Fine for non-confidential work, not for privileged materials | 8 custom skills built in Track 1. Limited to Claude's built-in Projects feature | No (30-day) | N/A | Anthropic | Done | $30/user/mo | 30-day window |
Managed platform — all-in-one option
What it is: AI platform built specifically for legal and professional services. Handles document storage + querying (Vault), workflow automation + report generation (Agents), and custom process builders (Agent Builder) in one managed platform. 2,400+ legal organizations, 200K+ professionals.
Why this matters for Insight: Instead of stitching together Claude API + Onyx + Box + custom middleware, Harvey may deliver the entire closed-circuit system out of the box — with security designed for "if this leaked, it's a legal event" materials.
Security posture:
- ZDR across all model providers — contractually requires zero data retention and prohibits model training on customer data (Security)
- SOC 2 Type II, ISO 27001, ISO 27701, ISO 42001, GDPR, CCPA
- First AI company to certify under EU-US Data Privacy Framework
- SAML SSO, audit logs, IP allow-listing, data lifecycle management, ethical walls enforcement
- Logical data separation per customer; independent pen testing by Schellman, NCC Group, Bishop Fox
- Data sovereignty options: US, EU/Switzerland, or Australia
Product lines that map to Insight's needs:
- Vault — secure document repository (up to 100K docs/vault). Natural-language querying across all content. "Review Tables" extract and compare key data points from thousands of documents. 96% key-term extraction accuracy. DMS integrations (iManage, SharePoint, Google Drive). Maps to Focus Area 2 + storage layer
- Agents — end-to-end task execution with cited, review-ready output. Scheduled agents for recurring reports. Multi-format support (documents, images, video, audio in; reports, presentations out). 700K+ daily tasks, 5.8M+ documents analyzed daily. Maps to Focus Area 1
- Agent Builder — no-code custom workflow builder. Chain steps: document ingestion → clause extraction → comparison → risk flagging → output. 500+ pre-built use case agents. Insight could build their own summarization workflows
Microsoft integration (Insight is a Microsoft shop):
- Available as an agent inside Microsoft 365 Copilot
- Outlook add-in for email summarization and drafting
- Word integration (reported 60% reduction in NDA review times)
- SharePoint integration for file sync
- Teams transcripts → Vault — secure path from Teams recordings into the AI system for processing. Could address the Fireflies.ai concern (exec team won't use Fireflies because of recording anxiety)
Models underneath:
- Multi-model: Claude (via Bedrock), GPT-4.1/o3 (via Azure), Gemini 2.5 Pro (via Vertex), Mistral
- Auto-routes to best model per task, or user can select
- ZDR enforced across all providers — not just one model vendor
Meeting / transcript handling:
- Harvey mobile app can record depositions and client meetings (up to 2 hrs), transcripts saved to Vault with auto speaker labels
- Video ingestion: MP4, MOV, AVI, WebM (up to 2 hrs per file) — litigation teams bulk-upload deposition videos
- Not a general meeting bot — does not join Zoom/Teams/Google Meet calls automatically like Fireflies. Focused on legal recordings, not all-hands meetings
- But: Teams transcripts can flow into Vault via Microsoft integration — the recording happens in Teams (where the exec team is already comfortable), then Harvey processes the transcript securely
Pricing reality check: Harvey does not publish pricing. Enterprise-only, demo required. Third-party estimates: ~$1,200–$2,000/seat/month at small scale (20-50 seats minimum). Annual contracts typically $50K–$300K+. A credit-metered (pay-as-you-go) option also exists. This is a significant budget conversation — needs a sales call to get real numbers for Insight's 45-person team.
Open questions: HIPAA BAA availability is not explicitly documented (certifications list SOC 2 / ISO / GDPR / CCPA but not HIPAA). Must confirm directly with Harvey. Also: 100K document cap per vault could be a constraint for large litigation cases. And no public data export documentation — lock-in risk needs evaluation.
Harvey changes the framing of this entire project. Two paths forward:
Path A: Harvey as the primary platform
- One vendor covers AI + document storage + workflow automation + Microsoft integration
- Security and ZDR are built in, not DIY. Certifications already in place
- Designed specifically for "if this leaked, it's a legal event" sensitivity level
- Dramatically less infrastructure to build and maintain
- Could fully replace Claude for confidential work (and potentially all AI use)
- Risk: pricing, vendor lock-in, 100K doc/vault limit, unconfirmed HIPAA BAA
Path B: Build-your-own stack (LibreChat + LLM APIs + Box/S3 + custom middleware)
- Maximum control over every layer
- No vendor lock-in; swap any component
- Requires significantly more engineering, maintenance, and security ops
- Multiple vendors to manage, each with their own compliance posture
- Insight would need dedicated technical capacity (or us) to maintain
Path C: Hybrid
- Harvey for confidential/privileged work (case materials, litigation support)
- Claude Teams for non-confidential creative/research work (tiers 1-2)
- But: Jen's concern about "bouncing between two systems" is valid. If Harvey covers enough general use, a clean single-platform migration may be better
Internal note (Maiya/Jen): Concern that Claude has "too much gray area" for enforcing guardrails in the AI policy. If Harvey has all the functionality Insight needs, a full migration before they get too deep into Claude could be cleaner than maintaining two parallel systems. The team is using Claude minimally right now — mostly to build comfort with AI. A well-executed Harvey onboarding + training could get them there instead.
Open-source platform — build-your-own with full control
Open-source chat and agent platform with built-in RAG, model-agnostic backend, and a polished user-facing UI similar to ChatGPT/Claude. Used internally at Shopify (~10,000 daily users), actively contributed to. Recently acquired by ClickHouse (major analytical database vendor). Runs in Docker with a Postgres backend.
- Agent-based architecture: create scoped agents around specific topics or data sets. An "HR agent" for policies, a "case research agent" for litigation materials, a "library agent" for scientific PDFs. Users pick an agent from a dropdown and chat within that scope
- Custom workflows: build and customize agents with specific instructions, document scopes, and model assignments. Chain agents for multi-step processes (ingest documents, extract key data, generate formatted reports). Comparable to Harvey's Agent Builder but fully under your control
- Built-in RAG with hybrid search: vector search + keyword search in a single system. Hybrid search matters for scientific and technical content where exact terminology is as important as semantic meaning
- Model agnostic: plug in Claude API, OpenAI, Gemini, Mistral, or any model via OpenRouter. Can restrict specific agents to specific models for cost control (cheap model for basic tasks, Claude for complex summarization)
- MCP server support: LibreChat can act as an MCP server, so staff using Claude Desktop or Claude Code on their own machines can still query the shared document base without switching tools
- Data stays in Postgres: all vectorized content lives in a database you control. That same database can feed automation, report generation, and other tools beyond the chat interface. No vendor lock-in on the data layer
- Hosting options: self-host on DigitalOcean, Railway, Cloudflare, or any Docker-capable infrastructure. Also available via hosting partners
- Proof of concept speed: 1-2 developers, roughly one week to stand up a working demo with real documents (per Seb Barre, Shopify)
Security considerations:
- Self-hosted means you own encryption, access controls, and compliance at the infrastructure layer
- No built-in HIPAA certification, but fully enclosed environment with no external data processors if paired with local models
- Advantage over proprietary platforms: no vendor data access, no vendor subpoena exposure, no per-seat toll on API access to your own data
How it fits with Insight:
- LibreChat as the user-facing layer (chat UI + agent routing) on top of whichever LLM fits the task (Claude, GPT, Gemini, open-source models via OpenRouter) and a Postgres vector DB as the document index
- Could replace or complement Onyx: LibreChat handles the full stack (UI + RAG + agent logic) where Onyx is primarily a RAG engine with a basic chat interface
- The agent model maps well to Insight's needs: different sensitivity tiers, different document collections, different access levels per role
- Non-technical staff get a familiar chat interface. Technical staff can tap into the same data via MCP from their own tools
Source: Seb Barre (Engineering, Shopify), via Chris Bryce. Shopify runs LibreChat at scale internally. Seb confirmed it handles the exact pattern we're looking at: scoped document collections, role-based agent access, hybrid search, model routing. "It's a pretty rock solid platform."
LLM providers and deployment options
If Harvey doesn't work out (pricing, HIPAA, fit), Insight needs a closed internal AI system per their AI Governance Policy. ZDR is non-negotiable for any tool touching tier 3+ confidential data. These are the viable deployment models for a build-your-own approach.
What it is: Claude via the Messages API at api.anthropic.com with Zero Data Retention enabled at the org level. Must be requested through Anthropic sales.
ZDR guarantees: Prompts and responses are not stored at rest after the API response returns. Data is never used for model training (contractual commitment).
- ZDR covers: Messages API, Token Counting, prompt caching, streaming, tool use, structured outputs, PDF input, citations
- ZDR does NOT cover: Claude Console/playground, Batch API (29-day retention), Files API, Code Execution, Managed Agents
- Fable 5 and Mythos 5 require 30-day retention — not available under ZDR. Opus, Sonnet, Haiku all eligible
- Certifications: SOC 2 Type II, ISO 27001:2022, ISO 42001:2023, HIPAA BAA available
Subpoena caveat: If a session is flagged for a policy violation, Anthropic may retain inputs/outputs for up to 2 years regardless of ZDR. Whether this creates privilege exposure is a legal question for Insight's counsel. Under normal operation with ZDR, there is nothing stored to produce.
What it is: AWS Bedrock is Amazon's managed service for accessing foundation models (Claude, Llama, Mistral, etc.) through an API within your own AWS account. Instead of calling Anthropic's API directly, requests go through AWS infrastructure. AWS is the data processor — Anthropic never sees the prompts or responses.
Same Claude models — Opus, Sonnet, Haiku are all available on Bedrock. Same quality, same capabilities. You just access them through AWS instead of Anthropic.
ZDR guarantees: Bedrock does not store prompts or completions by default. You can enforce this org-wide with AWS Service Control Policies so no one on the team can accidentally turn logging on.
- Data stays in your AWS account with guaranteed isolation
- PrivateLink available — traffic between your app and Bedrock never touches the public internet
- HIPAA-eligible, SOC 1/2/3, ISO 27001, FedRAMP
- Billing goes through your AWS account, not Anthropic
- If subpoenaed, the target is your AWS account (which you control), not a third-party vendor
Bedrock vs. Anthropic API — when to choose which: Bedrock adds infrastructure control (your data, your account, your network, your policies). The tradeoff is more setup and AWS ecosystem commitment. For Insight, Bedrock is the stronger option if they already have or are willing to stand up AWS infrastructure. If not, the Anthropic API with ZDR gets them 90% of the way with less overhead.
Claude on Vertex AI with Google Cloud as data processor. Data residency controls available. Cannot confirm an explicit ZDR toggle equivalent to Bedrock's. HIPAA BAA available, SOC 1/2/3, ISO 27001, FedRAMP.
Run Llama 3, Mistral, etc. on Insight's own infrastructure. Maximum data control. No subpoena risk from a vendor.
- Model quality gap: meaningfully behind Claude on complex summarization, citation accuracy, structured outputs
- Significant infrastructure and maintenance burden
- Could make sense for embeddings/search (competitive) while using Claude API for summarization
Insight already deployed Claude for Teams. The product interface retains data for 30 days. Fine for public research (tiers 1-2), not for confidential case materials. Correctly positioned as a creative thinking tool. The closed system needs to be built on the API.
Document search / RAG layer
Open-source (MIT), fully self-hosted. 40+ connectors (Dropbox, Google Drive, Slack), pluggable LLM backend, hybrid search (vector + keyword). Can use Claude API (ZDR) for answer generation.
- Full data residency control — nothing leaves Insight's infra if paired with local embeddings
- Supports air-gapped deployment with zero outbound traffic (ITAR/FedRAMP/CMMC environments)
- Agentic RAG with deep research mode for multi-step retrieval
- SSO (SAML/OIDC), RBAC, credential encryption, audit logging. Yearly pen tests (results under NDA)
- Free community edition; enterprise support available
Security caveats for litigation use:
- No HIPAA certification — Onyx Cloud is SOC 2 Type II, but no BAA pathway documented for either Cloud or self-hosted
- Encryption is DIY on self-hosted — you must configure disk encryption, database encryption, and TLS yourself
- Telemetry is opt-out, not opt-in — default behavior sends aggregated data + daily version heartbeat to Onyx. Must set
DISABLE_TELEMETRY=true before loading any data
- Document chunks are sent to the configured LLM — if pointed at Claude API with ZDR, that's fine. But Onyx itself caches docs in its local vector DB and Postgres — subpoena resistance depends on your infra's encryption, not Onyx's
- A fully MIT-licensed fork exists at
onyx-dot-app/onyx-foss — worth auditing
Bottom line: Onyx provides the plumbing, not the compliance. Self-hosted with telemetry disabled, connected to Claude API (ZDR), on encrypted infrastructure you control = defensible architecture. But you own encryption, access controls, audit trails, and subpoena resistance at the infrastructure layer.
Managed RAG on AWS. S3 for documents, managed embeddings + vector store. HIPAA-eligible, ACL filtering. Best if Insight moves document storage to AWS. Lower maintenance than Onyx but requires AWS ecosystem commitment.
SOC 2 Type II, HIPAA compliant, single-tenant VPC option. First-year costs typically $300K–$1M+. At 45 people, likely overkill for Insight's budget.
Current state
Dropbox was chosen when Insight was 5 people. Now at 45. No security hardening. Some Box usage. Documents scattered with inconsistent naming and rampant duplication.
Key insight (Maiya): Dropbox itself is the data leak risk, not Claude. Any AI connector needs a secure tunnel to Dropbox, but the Dropbox environment itself needs an audit.
Dropbox security assessment
What it provides: AES-256 at rest, TLS in transit. SOC 2 Type II, ISO 27001/27017. HIPAA BAA available (must request). SSO, MFA, remote wipe, audit logging.
Where it falls short for litigation materials:
- Not zero-knowledge — Dropbox holds the encryption keys. They can technically access file contents. For privileged materials, this is real exposure
- 84% compliance on legal requests — Dropbox hands over data in most cases when served with legal process. Opposing counsel could subpoena Dropbox directly
- 2024 Dropbox Sign breach — unauthorized access to Sign (HelloSign) exposing emails, hashed passwords, API keys, OAuth tokens. Core storage wasn't hit, but demonstrates the attack surface
- Limited data residency — US-based by default, EU requires 250+ seats
Bottom line: Dropbox meets baseline enterprise security standards but is not designed for the level of confidentiality litigation materials require. The combination of Dropbox-held encryption keys + willingness to comply with third-party legal requests = subpoena risk for privileged case materials.
Storage alternatives comparison
| Solution | Encryption model | Subpoena resistance | HIPAA | AI/RAG integration | ~Cost (45 users/mo) |
| Dropbox Business * | AES-256, Dropbox holds keys | Low — 84% compliance rate | BAA avail. | Dash eliminated (no ZDR) | ~$900-1,400 |
| Box Enterprise + KeySafe | AES-256 + customer-managed keys (AWS/GCP KMS) | Good — revoke key = ciphertext only | Yes (Enterprise+) | Best — mature API, Box AI built in | ~$2,000-2,500 |
| Tresorit | Zero-knowledge E2E (AES-256 + RSA-4096, client-side) | Best — provider cannot decrypt | BAA avail. | Weakest — limited S3-compat API | ~$1,000-1,300 |
| Egnyte | AES-256 + customer-managed keys (Azure KV / AWS HSM) | Moderate — with customer keys | Yes | REST API available | ~$900-1,200 |
| AWS S3 + KMS | Customer-managed or client-side encryption | Best — with client-side enc. | Yes | Best — full programmatic access | ~$200-500 + ops |
| Azure Blob + Key Vault | Customer-managed keys via Key Vault | Best — ciphertext only w/o keys | Yes | Good — full API | ~$200-500 + ops |
* Dropbox is currently in use. Not recommending it for confidential/privileged materials due to the security gaps above, but it remains the active platform. Any migration would be phased — Dropbox stays operational during transition.
Detailed evaluations
Why it leads: Closest Dropbox replacement UX-wise, minimizing migration pain. KeySafe lets Insight manage their own encryption keys via AWS KMS or Google Cloud KMS — Box cannot decrypt without Insight's key service being active. Best API ecosystem for AI/RAG integration of any file storage option.
- SOC 2 Type II, HIPAA BAA (Enterprise+), FedRAMP, ISO 27001
- With KeySafe: if you revoke key access, Box holds ciphertext it cannot decrypt. A subpoena to Box yields nothing usable
- Without KeySafe: same exposure as Dropbox (Box can produce file contents)
- Mature REST API, extensive SDKs, Box AI built in — strongest connector story for RAG pipelines
- KeySafe is a significant add-on cost (custom pricing)
For Insight: The practical choice. Familiar UX for 45 non-technical users, strongest API for AI integration, and customer-managed keys that actually solve the subpoena problem. The cost premium over Dropbox is justified by the security upgrade and AI integration capabilities.
Zero-knowledge, end-to-end encryption. Files encrypted client-side before upload (AES-256 + RSA-4096). Tresorit never holds keys. Swiss jurisdiction (stronger privacy law than US).
- A subpoena to Tresorit yields metadata only (account info, timestamps, file names) — not file contents
- SOC 2, HIPAA BAA available, GDPR
- Critical tradeoff: S3-compatible API with significant limitations. Zero-knowledge means the server can't read files, so server-side AI integration is structurally impossible. Any RAG system needs client-side decryption first
- ~$19-24/user/mo; Enterprise custom-quoted
For Insight: If subpoena resistance is the dominant concern above all else and AI integration can wait or be done client-side. The zero-knowledge model inherently limits what server-side integrations can do — a fundamental tradeoff, not a fixable gap.
Popular in life sciences and regulated industries. AES-256, customer-managed keys via Azure Key Vault/AWS CloudHSM. SOC 2 Type II, HIPAA BAA, ISO 27001. Hybrid cloud + on-prem sync. REST API available but less mature than Box.
- ~$20/user/mo (Business). Enterprise custom
- Subpoena resistance moderate — depends on customer-managed key configuration
You control everything. SSE-KMS with customer-managed keys, or client-side encryption for true zero-knowledge. Full programmatic access — best possible foundation for AI/RAG pipelines. HIPAA-eligible, SOC 2, FedRAMP.
- ~$0.023/GB/mo + requests. ~$150-300/mo for 5TB storage
- No end-user interface out of the box — need to build or buy a file management UI
- Requires dedicated DevOps/security capacity
- Maximum control, maximum operational burden
Purpose-built for law firms with ethical walls, matter-based access controls, and audit trails. NetDocuments offers three-layer encryption including customer-managed keys. Both carry SOC 2, ISO 27001.
- ~$30-50+/user/mo plus $10K-50K+ implementation
- HIPAA BAA availability unconfirmed — needs direct verification
- Insight is a scientific consulting firm that does litigation support, not a law firm. The ethical wall and matter management features add complexity they likely don't need
Platform alternative: Harvey Vault
If Harvey is selected (Path A), the storage question simplifies significantly. Harvey Vault handles the privileged/confidential document layer with built-in natural-language querying, ZDR across all model providers, and DMS integrations. Dropbox or Box would remain only for general business files (non-confidential). See the Secure AI environment tab for the full Harvey evaluation and the Architecture tab for how this changes the system design.
Context strategy for Claude (build-your-own path)
Ben's Track 1 work never addressed how Claude accesses files, local file storage, or team-wide consistency. This is a gap we fill in discovery:
- Structured folder hierarchy connected to Claude (or migrated to S3/secure alternative)
- Per-project document sets vs. company-wide reference library — different access patterns, different solutions
- Team-wide Claude Projects configuration (which docs always available vs. loaded per-case)
- Secure connector architecture: how data moves from storage to the LLM and back
- Storage security audit — permissions, sharing, external access, encryption at rest
Why PDF → Markdown conversion matters
Insight's documents are primarily PDFs — many scanned, with mixed layouts (text + handwriting + stamps), Bates numbers, depositions, medical records. Raw PDFs are expensive and inefficient for LLMs. A 50-page scanned deposition fed as a PDF image burns tokens on layout parsing and often produces worse results than clean structured text.
Converting to Markdown first gives us:
- Token efficiency — Markdown is 3-10x more compact than raw OCR text or PDF image tokens. At scale across thousands of case documents, this directly impacts API costs
- Better LLM comprehension — Claude processes structured text (headings, lists, tables) far more accurately than unstructured OCR dumps
- Searchability — Markdown files are directly indexable by RAG systems. Vector embeddings on clean text produce better retrieval
- Metadata preservation — Bates numbers, page references, exhibit numbers encoded as structured Markdown (headers, footnotes, YAML front matter)
- Version control — Markdown diffs are human-readable. Re-processed or corrected documents are trackable
This is a prerequisite for both Focus Areas. Should be one of the first things built and tested in discovery.
OCR + conversion options
| Solution | Accuracy | Output | HIPAA | Self-hosted | Cost |
| Marker (open-source) | Good | PDF → Markdown directly | N/A (self-hosted) | Yes | Free |
| Google Document AI | 96-99% (structured) | JSON/text (needs MD step) | BAA available | No (cloud) | $0.60-1.50/1K pages |
| Amazon Textract | High (forms); ~71% handwriting | JSON/text (needs MD step) | Yes | No (AWS) | Usage-based |
| Adobe PDF Services | Good (preserves structure) | JSON/text (needs MD step) | Available | No (cloud) | Usage-based |
| Tesseract (open-source) | Moderate (clean docs only) | Plain text (needs MD step) | N/A (self-hosted) | Yes | Free |
- Step 1: Deduplicate the archive first — save OCR costs on redundant docs
- Step 2: OCR via commercial service (Textract if on AWS, Document AI otherwise) or Marker for direct PDF→MD
- Step 3: Post-process into clean Markdown with Bates numbers, page numbers, and document metadata preserved as YAML front matter
- Step 4: Index into the RAG system (LibreChat's Postgres vector DB, Onyx, or Bedrock Knowledge Bases)
Test early: Accuracy varies dramatically between clean typed docs and degraded scanned depositions. Get sample files from Insight for each document type and test before committing to a pipeline.
AI Governance Policy — gap analysis
Insight's AI Governance Policy (V1.0, June 2026) is thorough for Public AI Tool use. But there are gaps between the policy and their current Claude deployment that we should flag.
Section 9 requires AI Conversations to be retained for one (1) year and says retention controls "shall be set to a period consistent with this retention requirement." But Claude for Teams has a fixed 30-day retention that cannot be configured. These are in direct conflict.
The Approved Tools list (Appendix A) lists Claude on a "Pro License" — no mention of Claude for Teams as a distinct deployment with different retention characteristics.
Risk: If Insight is using Claude for Teams for any project-related conversations, those conversations are being auto-deleted at 30 days, violating their own 1-year retention policy. Under a litigation hold, this could be a serious compliance issue.
The policy explicitly states: "If the Company develops a Closed Internal AI System at a later time, additional provisions addressing that system will be added to this Policy." Those provisions don't exist yet. That's exactly what SOW3 is building toward.
The policy also says privileged information in a Closed Internal AI System "will require separate consideration with Legal Oversight. This scenario will be covered in a future version of this policy."
Until the closed system exists and the policy is updated, all confidential/privileged materials are prohibited from any AI tool Insight currently has.
- 5-tier data classification with clear boundaries
- Never/With Restriction/Always framework is practical and enforceable
- Case materials are always Prohibited in Public AI Tools — no loopholes
- Citation verification and scientific attribution requirements
- AI note-taker rules (pause for privilege discussion, internal-only at rollout)
- AI Conversations classified as Company Documents under Record Retention Policy
- Litigation hold and discovery readiness provisions
- "The Policy is a floor, not a ceiling" — encourages conservatism
Secure connector architecture
The policy already defines the architecture we're proposing (Section 2, "Closed Internal AI System"): a system "hosted and operated under the Company's direct control, accessed through a ZDR agreement" that "includes programs or systems that function entirely on an individual user's device or that sends information for private cloud or server computation, but is not visible or retained by the providing service."
This means yes — the document store can and should be separate from the AI system, with secure connectors between them. This is the standard pattern for ZDR-compliant AI systems:
Layer 1: Document store (Box + KeySafe, or AWS S3 + KMS)
- Encrypted at rest with customer-managed keys
- Insight controls who can access, and can revoke keys to render data unreadable
Layer 2: Middleware / connector (custom app, running in Insight's infra)
- Pulls documents from the store via secure API
- Handles OCR, chunking, Markdown conversion
- Sends structured text to the LLM endpoint
- This is code we build and Insight controls — no third-party data processor
Layer 3: LLM endpoint (Claude API with ZDR)
- Processes the request, returns the response, retains nothing
- All traffic over TLS; optionally via AWS PrivateLink if using Bedrock
Layer 4: Chat UI + RAG / search index (LibreChat, self-hosted)
- User-facing chat interface with scoped agents per document collection and sensitivity tier
- Built-in hybrid search (vector + keyword) with Postgres vector DB backend
- Handles search queries, retrieves relevant chunks, passes to LLM for answer generation
- Nothing leaves Insight's environment for the search or chat layer
- Also exposes MCP servers for staff using Claude Desktop/Code directly
No data leaks between layers: Each connection is TLS-encrypted. The LLM has ZDR (nothing retained). The document store has customer-managed keys (nothing readable without Insight's keys). The connector layer is code Insight controls. The search index is self-hosted. A subpoena to any single vendor produces nothing usable.
Recommended architecture — Path A (Harvey)
If Harvey checks out on HIPAA, pricing, and fit — this dramatically simplifies the architecture:
For case material summarization (Focus Area 1):
- Harvey Agents handle end-to-end summarization with cited, review-ready output
- Agent Builder lets Insight create custom summarization workflows matching their exact template requirements
- ZDR enforced across all underlying models (Claude, GPT, Gemini) automatically
For document search & organization (Focus Area 2):
- Harvey Vault as the document repository with natural-language querying
- Review Tables for structured extraction and comparison across thousands of documents
- DMS integration for existing file stores (SharePoint, Google Drive)
For file storage:
- Vault handles the AI-accessible document layer (confidential/privileged materials)
- Existing Dropbox (or upgraded Box) remains for general business file storage (non-confidential)
- Migration: privileged materials move into Vault; general business files stay where they are
For meeting transcripts:
- Teams recordings flow securely into Vault via Microsoft integration
- Harvey processes transcripts for summarization, action items, follow-ups
- Exec team stays in Teams (familiar), never needs a separate recording tool
- Addresses the Fireflies.ai concern directly
What changes for the team:
- Claude Teams could be deprecated for all work if Harvey covers tiers 1-2 adequately
- Single platform means one set of guardrails to enforce, one compliance posture to audit
- Training effort concentrated on one tool instead of teaching "use Claude for X, Harvey for Y"
What this means for our role: Shifts from infrastructure build (months) to platform deployment + custom workflow development + data migration + training. Still significant scope, but lower ongoing maintenance burden for Insight.
Recommended architecture — Path B (build-your-own)
If Harvey doesn't work out, the build-your-own approach satisfies ZDR + HIPAA + subpoena requirements across all layers:
For case material summarization (Focus Area 1):
- Claude API with ZDR enabled — best-in-class at structured summarization, citation extraction, and template formats
- Custom application layer: ingest case PDFs → OCR → Markdown → structured chunks to Claude via API
- No data stored on Anthropic's side; all case materials remain in Insight's controlled environment
For document search & organization (Focus Area 2):
- LibreChat as the user-facing platform (chat UI + agent routing + built-in RAG with hybrid search). Scoped agents per document collection, role-based access, familiar chat interface for non-technical staff
- Postgres vector database backend for embeddings and document index (all data stays on Insight's infrastructure)
- Model-agnostic LLM layer: plug in Claude API (ZDR), OpenAI, Gemini, or open-source models via OpenRouter. Can assign different models to different agents based on cost and capability needs
- MCP server capability means staff using Claude Desktop or Code can still query the shared document base
- Deduplication and structured reorganization of storage as a prerequisite
- Alternative: Onyx (self-hosted) or AWS Bedrock Knowledge Bases if LibreChat's RAG capabilities prove insufficient for Insight's scale
For file storage:
- Migrate from Dropbox to Box Enterprise + KeySafe (customer-managed keys, best API for AI integration)
- Or: AWS S3 + customer-managed KMS if Insight is willing to invest in infrastructure
- Structured folder hierarchy with per-project isolation and role-based access
For automation and reporting:
- LibreChat's Postgres backend means the same document data that powers search can feed report generation, custom workflows, and other automation. One pipeline serves multiple use cases
- Agent-built microsites for recurring reports (one-time token cost vs. regenerating via chat every time). Pattern validated at Shopify
What stays on Claude Teams (current):
- Public research, creative brainstorming, non-confidential writing — tiers 1-2 in their governance policy
- The 8 existing skills from Track 1 (already working)
Positioning for Josh
Frame as "initial findings from the dev team." Key messages:
- The scope is bigger than just picking an LLM — storage, security, and document infrastructure are equally critical
- Harvey.ai changes the conversation — a managed platform built specifically for legal-grade AI, with ZDR baked in across all model providers. We need to evaluate it against building custom
- We're doing more research before recommending a direction; the Harvey demo/pricing conversation is the immediate next step
- Mangrove brings an engineering team — whether it's Harvey deployment + custom workflow development or building a custom stack with LibreChat, we cover the technical execution
- LibreChat (new finding): validated open-source platform used at Shopify (10K daily users). If Harvey doesn't fit, LibreChat gives Insight a polished chat + agent UI with built-in RAG, model-agnostic backend, and full data control. Proof of concept in ~1 week
- The current Claude deployment has gray areas the AI policy doesn't address cleanly (retention conflict, no closed-system provisions). Harvey may solve this structurally
- We take subpoena risk as seriously as they do — that's why we're evaluating ZDR at every layer, not just the model
Harvey.ai evaluation (immediate next step)
- Request Harvey demo — get real pricing for 45 users and credit-metered option
- Confirm HIPAA BAA — not listed on their public security page, must verify directly
- Ask about data export — what happens if Insight wants to leave? Can they extract everything from Vault?
- Test Vault with sample case documents — does 96% key-term extraction hold for scientific/litigation docs?
- Evaluate Agent Builder for Insight's specific summarization templates and citation requirements
- Confirm Teams → Vault transcript pipeline works for their Microsoft licensing tier
- Understand onboarding: how long from contract to team-wide deployment?
- Ask about the 100K document/vault limit — is this per project or org-wide? Can they scale?
- Subpoena question for Harvey legal: what does Harvey produce in response to third-party legal process?
Questions for Insight (Phase 1)
- Get sample documents from each type (deposition, interrogatory, MSDS, medical record, past report) for testing
- Confirm exact output format/template for case summaries — what does a "good" manual summary look like?
- Map the citation format requirements — how do they currently cite back to source pages?
- Understand access frequency: how often is the scientific library searched vs. per-project case materials?
- Clarify the tier 3-4 confidentiality gray areas the team is nervous about
- Audit current Dropbox structure: folder hierarchy, permissions, sharing settings, external access
- Identify the new Technical Director's technical background and infrastructure preferences
- Confirm Insight's IT infrastructure: do they have or would they provision AWS/cloud accounts?
- Legal review: have Insight's attorneys assessed the subpoena risk of different AI deployment models?
- Volume sizing: typical project doc count, total library size, concurrent users
- Appetite for platform vs. build: Would Insight prefer a managed platform (Harvey) or a custom-built stack? Budget range for annual AI tooling spend?
Internal validation (before presenting to client)
- Harvey demo + pricing — this is the gating decision before investing deeper in build-your-own
- Contact Anthropic sales about ZDR enablement timeline and process (fallback if Harvey doesn't work)
- Prototype: test Claude API on a sample case summary with mock data matching Insight's format
- Prototype: stand up LibreChat locally in Docker and test agent-scoped document search against sample scientific PDFs (1-2 devs, ~1 week for working demo)
- Prototype: stand up Onyx locally as a comparison if LibreChat's RAG isn't sufficient
- Price out AWS infrastructure for a Bedrock-based approach at Insight's estimated volume
Talk to Seb Barre (via Chris Bryce) for build-your-own perspective Done Sep 8 — LibreChat recommendation, Shopify internal validation, hybrid RAG guidance. See notes above
- Talk to Sebastian from Bureau and Ben Fox for additional perspectives
- Get Box Enterprise + KeySafe pricing for 45 users (fallback storage option)
- Test Marker on sample litigation PDFs to evaluate PDF→Markdown quality
Eliminated solutions
Solutions evaluated and eliminated during research. Kept here so we don't re-evaluate them.
| Solution | Category | Reason eliminated | Date |
| Dropbox Dash | Document search | No ZDR — 30-day data retention via OpenAI. Non-negotiable fail on ZDR for any tier 3+ data. | Aug 2026 |