Legal Automation Breaks Down on Sensitive Documents
Legal and accounting professionals are facing challenges in automating document workflows, particularly when dealing with sensitive data and regulatory compliance. The desire for speed and efficiency clashes with the need for auditability, deterministic processes, and human oversight to avoid errors and maintain accountability. Current 'AI agent' solutions often fall short due to inaccuracies and a lack of transparency.
SOURCES (60)
“Add a filter, hit save... filter gone. Hit save and run, wheel spins, stays in edit mode, doesn't run. Changes not saved. Never had these issues before, suddenly constant the last two months. EDIT: It is also an issue where I hit edit, change nothing,…”
Local Meetup Event description Build a working document extraction flow: classify mixed packets, extract key fields, validate results, and handle harder unstructured documents. Please bring a laptop so you can participa…
“Harvey allows processing over 10,000 documents and will happily put them in a grid and run prompts on each one. It’s also got visual workflow and agent builders, and extensive metrics and reporting, along with integrations into third party tools like iManage or other legal specific tools. It’s got “better” data residency and governance controls as well if you’re a large firm.”
“It sounds like it has very good document review and organizational capabilities. And like maybe you can make a big case file for each case, as opposed to just plugging individual folders into CoWork? I saw that Everlaw (which my firm uses) now has a Claude connector, but I haven't had a chance to test it out yet.”
“It works very well. The conditions are based on data points ("properties") If it doesn't "see" our pre-approved positions in the doc, the clause flag goes up. I sometimes get incorrect flags, because "payments" or something similar comes up more than once in the doc. All I have to do is untag the clause it selected, highlight the correct clause and tag it as the "payments" clause.”
“I think you could build this for yourself in Claude code. Most of the services have MCP and or APIs to get this data. Gong is the hardest and official MCP gives summaries vs transcripts but if you can get an API key there are third party ones that can help. If you really want something OOTB there are a ton of startups claiming to do this from a lot of sources. My experience is they struggle due to lack of context about your app, maker, users, etc and get to much wrong. This is where I think home”
““Structurally” is the word I keep getting stuck on. A template comparison tells you the indemnity section is there — not that what’s in it actually meets the playbook. Does the structural pass catch most of it for you, or does the substantive gap still come down to someone reading closely? Other thing I can’t get a read on: when one does slip through, what does it actually cost? Renegotiation, a write-off, something that never gets tallied? Trying to work out whether it’s a nuisance or a real nu”
“The breakdown is almost always process, not reviewer, because human memory is a terrible checklist. The fix that actually works: a written playbook tied to contract type, reviewed against the document clause by clause, not just scanned. Second reviewer helps but only if they're checking the playbook too, not just reading cold. Some tools, Irys One is one, let you compare a draft against a template to surface gaps structurally. Volume compounds missing-clause risk, so batch similar contracts”
“the risk framing is the right way to think about it. efficiency is nice but the real argument for proper workflow software at your scale is that it removes reliance on individuals catching things in email. we run a team managing multiple parallel engagements and the shift with TaxDome wasn't that we got faster, it was that things stopped quietly falling through the cracks. every task has a clear owner, every client interaction is logged, and nothing moves forward until the right step is comp”
“I recently received rejection of all my edits to an SOW on the basis that the edits weren't made to a prior similar SOW executed with this provider over a year ago and thus I must accept the original language as is, since it was previously acceptable - and every single rejected edit had a bubble comment with this identical explanation. I mean, this is just lazy counterparty, either unwilling to budge, or failing to find another explanation for rejecting your changes (whether legal or technic”
“Whatever tool of software you select - you’ll need to configure it to your requirements. Smartsheet offers workflows but you can theoretically do similar with Excel with smarts including AI (most PPM tools are now playing in this space). I’ve also come across Wisegrid which is packaged comparable to current known competitors. MS project and Jira are options for Schedule management. Remember, tools are to support processes and hopefully to optimise efficiency and productivity - but do they really”
“I didn’t mean to be an ass! I’m a patent litigator so I’m more interested in the ip and litigation suite right now. Going to do more benchmarking (batched sonnett 5 soon). Ive used the bench as a way of tuning my own harness tools (not through prompt engineering)”
“I think the easiest way is to identify the owner of the critical information you are lacking in this report and create a shorter workflow for this information. If you standard workflow take a week but you have only 2 people holding the info you're lacking ? place a meeting the same morning or the afternoon before to have the freshest info and add a slide at the end of your presentation. As far as I understand your problem seems to mostly be the worklow. I don't think a tool will help you”
“Old school company reporting in. All the MS docs are password protected and nobody writes down the password for that particular project.”
“Those uploaded documents we sanitized as needed Is it a manual process? Do you have your offline agents do it?”
“We did not give Claude access to our local folders. Everything is done through Projects and uploading documents to the Project knowledge bases and when running the Projects. Those uploaded documents we sanitized as needed”
“How are you performing the "sanitized uploaded info" step? E.g. if Claude goes off to find something and takes a peak in the wrong folder?”
“I'm at an organization with a user group that prefers to get PDF reports, instead of logging into a web service to use an interactive dashboard, for good reason. I think this way of data consumption will continue, and I'm trying to lead our efforts in switching our reporting tool. We currently use AWS Quicksight which just barely gets the job done and has been having several technical issues, so I'm in the early stages of exploring our options. Paginated reporting and PDF emailing ar”
“I've had a good experience with Ironclad and Juro. Both are relatively straightforward to roll out compared to some of the larger enterprise platforms, and they're well suited for lean legal teams.”
“This is not a system fix it is a policy fix. You have a policy that all documents marked sensitive need permission from the document owner to be shared, otherwise it is a policy violation with punishment up to termination, etc.”
“Hot take from a banking DevSecOps team: treating prompt injection as something the model vendor should fix is a dead end. The real issue is that the context window has no provenance. The model cannot tell user instructions from a poisoned README or a tool response. Until that changes, the practical mitigations look a lot like classic supply chain controls: pin your dependencies, verify what you fetch, restrict what each component is allowed to do. Anyone mapping this to SLSA-style controls yet?”
“Hello, So I am not a data person at all so if I miss speak or something like that please dont judge. So pretty much where I work we are moving from to different databases for employee info. From Company A I recieved the data as ASCII files with some BLB files and D files. Company B needs the file as a PDF. From what I understand the TXT files kinda work as the how the data is put back together and the BLB is the data its self and the D file is the key. IDK if thats right (probably not). I am won”
“Why is this an issue? Are you afraid that they will steal your work or take it out of context?”
“Importante usar um filtro para não virar ruído. Por exemplo, não adianta apenas automatizar alertas. É necessário acompanhar os "alertas passivos" (Google Alerts + newsletters especializadas) pra não ser pego de surpresa e n tentar acompanhar os reguladores direto. Uma outra pratica, é: só vou na fonte primária (lei/decisão/orientação) quando o mesmo tema aparece em 2 fontes independentes na mesma semana. Importante tbm lembrar que toda mudança passa por uma pergunta: isso muda um dado”
“Integrations guy here, this is one of the most common problems I see, the cause is that no system is defined as the source of truth for each field as people already stated here. The fix that works at small scale: pick one owner per data type, make every other system read-only for that field, and add one daily reconciliation check that compares totals and flags drift instead of you finding it manually”
“This is the way: make the document simple to update, short to the point, and makes the employees job easier by removing ambiguity.”
“The problem with documentation is it tends to become a moment in time, and often misses steps that the person writing them just assumes everyone else knows. So you need to treat them as an organic thing that needs constant review so you can be sure it's still relevant, and makes sense to people who don't already know the process. Obviously, the less you can allow your processes to organically change, the better, but someone will always find ways to improve something, so you need to ensur”
“Pay-As-You-Go. You are charged only when you actively use the platform for analysis, summaries, or work product.”
“Thanks for the quick response! And it's always cool hearing from the developer. I appreciate you clarifying the bulk creation feature. I'll use it next time. Is there a way to track unique scans, like could I tell if someone scanned it 100 times? It's a unique use case again so no worries if there isn't a way currently.”
“I would use: “We can share our standard security pack under NDA. Custom questionnaires are scoped after commercial review, and larger requests may be billable.” That leaves room to waive the fee for a deal you already know is worth the time.”
“Transparency first: I'm an engineer researching how small defense suppliers handle export controlled data, might eventually build tooling in this space. Nothing to sell and nothing to link, the question is the point. The setups I keep seeing are 10 to 50 person manufacturers, M365, one IT person or an MSP, and a folder of ITAR-stamped drawings from a prime. Under ITAR a non-US person opening those files counts as an export even inside the US, and AI assistants add a second version of the sam”
“Whoever is doing the exports should be the person doing the matching in QBO. It will cause a mess if not.”
“page history restores content, database history restores structure (formulas included). they're just separated and I couldn't find the database one”
“OpenRefine already does most of this for free, and HubSpot or Salesforce have native dedupe built in. The gap might be cleaning the CSV before it ever touches a CRM.”
Or they could act like a sysadmin and write it themselves.
“File retrieval is a pretty well solved problem. But it's going to required involved set up, whether you host a local LM or use a cloud provider like AWS or Foundry search.”
“The confidentiality piece is the right lens to lead with. For a searchable knowledge base of decades of legal advice and submissions, you're basically building a private firm memory system. I'd separate the question into two parts: the tool and the data path. On the tool side, AnythingLLM is a reasonable starting point but it gets fiddly because it's stitching together a lot of moving parts (embedding model, vector DB, chat model, chunking, retrieval). For 35 years of documents, the”
““Please combine these 700 PDF files into one giant markdown and then provide a detailed summary.””
“Yes, you are right about that. I see a github repo without any code in it, also there is a issues section on the website - maybe asking the author directly would provide some answers regarding privacy concerns.”
“I always clean names from files I feed to AI. It is more difficult if the docs are in pdf form, you need to print, redact and rescan again”
“Claude Legal, CoCounsel and Eve all are capable of them as far as I know. When I was taking certification classes last year, if I gave chatgpt the exact information I needed formatted into a citation it worked just fine. But I wouldn't trust any of the LLMs to do anything more than format what you provide and you need to be sure of what you're providing.”
“Do you integrate Westlaw with Claude in any sense? For example, having Claude suggest search string, bulk downloading full text of the search hits, and handing them to Claude?”
“Data engineers and BI analysts often have a broken handoff process. Engineer builds the pipeline. Analyst builds the report. Nobody owns the middle. Here's the data layer checklist I now run before any Power BI handoff: DATASET HEALTH □ All tables have descriptions □ Column data types explicitly set (not auto-detected) □ Nulls handled — not passed through raw □ Unnecessary columns removed (only what the report needs) □ Date table exists, complete, no gaps, marked as Date Table REFRESH &”
“How do you handle data mapping for regional or smaller ocean carriers that don't support standard EDI/API feeds. When you get irregular line-item formats or custom commercial invoices do you rely on tech to map the data or do have data entry staff manually fix ? Just starting working for a shipping company as their CFO so I need to get up to speed quickly on the issues within freight forwarding, logistics etc. Thanks submitted by /u/Realestate_Uno [link] [comments]”
“how are you catching stage history, snapshotting every row each poll or only writing a record when the stage field changes? notion won't give you the transition itself, you're diffing consecutive pulls to reconstruct it. anything that flips twice inside the same 15 minutes just gets flattened to the latest value.”
“ok, so a competitor or potential competitor could do a search "how to do XYZ" and without your data it would not know how to answer how to do XYZ but with your data it could then provide an answer that would give your competitor information on how to better compete with you. Is this an accurate characterization of the risk that you see?”
“Exact same Microsoft product interface I’m not sure how close your requirements are. But I’m 99% sure you’re looking for something like iManage or NetDocuments. Look into those. Like 70% of biglaw is using either of these tools. Has version control, collaboration features, etc. integrates pretty well with the whole Microsoft suite of products”
“I have consulted at many companies (including several starts ups) and the most reoccurring thing I've seen hold them back is poor documentation. I'll elaborate on that at the end of the post. Now I'm preparing my own start up and I don't want to fall into the same trap. I'm very experienced with documentation on the engineering and product development side of things, however my background is in mechanical engineering/medtech and not in business or finance so what I want to as”
“You really need a proper consultant. I know another person suggested Clio but with your volume of data, this isn't a straightforward decision. You need to consider which document storage platform to use e.g. NetDocuments might be better than SharePoint for your use case, and what CMS integrates with that. Clio is a new popular CMS but there are others like Smokeball. I don't think SharePoint is fit for purpose for your organisation's scale.”
“Yeah, great if everyone had the tools they need, but estimate input is needed from such a wide range of people that spreadsheets are often the only common piece of software”
“Something to consider and reflect upon. If you have an agreed and approved project plan and schedule then why do you need to recreate that same information in a different format? You're creating an unnecessary administration overhead for yourself, or you have to consider as the PM you're not communicating clearly and consciously enough for everyone to understand the who, what and when of your delivery schedule. What is in your RAID log that is not in your project plan or schedule? Are yo”
“On your text-layer question: yeah, that's the classic screwup. If a tool just draws a black rectangle over the words, the text is still sitting in the file underneath and anyone can copy it out or pull it from the content stream. Real redaction removes the characters from the file structure, not just covers them, and flattening alone often isn't enough either. It's exactly why the old-guard firms still print, Sharpie, and re-scan — a rasterized image has nothing underneath to recover”
“Appreciate this info. It is helpful. If committing to Microsoft is there any CMS that integrate well? For example, can an attorney work on a document either locally or on sharepoint or on OneDrive and that document auto stores to the CMS. Or vice versa - can an attorney work on a document inside a CMS and it look/feel like they are working on Microsoft products?”
