They didn’t tell thelawyers it was AI
Would you send this to a client? That's all the product validation needed.
In 2022, Winston Weinberg and Gabriel Pereyra pulled about a hundred landlord-tenant questions off r/legaladvice, wrote a chain-of-thought prompt before that technique had a name, and ran the answers through GPT-3. Then they handed the results to three landlord-tenant attorneys. Weinberg describes it on No Priors:
“And we gave it to three landlord-tenant attorneys. And we just said, nothing about AI. We just said, here is a question that a potential client asked, and here is an answer. Would you send this answer without any edits to that client? Would you be fine with that? You know, is that ethical? Is it a good enough answer to send? And 86 out of 100 was yes.”
That was the whole product. No software, no interface, no demo. And the model was bad. In a longer interview Weinberg describes the GPT-3 API of that moment: no instruction tuning, a neon green box covering your output, and if you forgot the punctuation at the end of your prompt it would autocomplete your sentence instead of answering it. His summary: “it was really rough.”
Some things I found awesome:
The questions weren't theirs. Real people on Reddit describing their actual apartment problems, tagged and taken as-is.
The AI label was removed on purpose. Weinberg's words: “We didn't say anything about ai, nothing.” In 2022, telling a practicing attorney that an answer came from a language model would have probably sunk the whole convo anyway.
The question had liability in it. Not "is this good" or "is this impressive." Thumbs up meant you would send it to a client with zero edits. Thumbs down meant tell me what you'd change. I suppose that prices in malpractice risk, professional ethics, and the reviewer's own name, which no accuracy score does.
Another curious thing to me: Weinberg was a securities and antitrust litigator at O'Melveny & Myers. He makes the point in that interview that a securities guy knows essentially nothing about landlord-tenant law in a given jurisdiction, even though people assume lawyers know all of it. So he picked a domain where he was not qualified to grade the output.
Like so many things in AI, I always think about what being wrong costs you. In the Reddit landlord scenario and the testers not much! Once it ramps up, it can cost a lot. But you can prove things out so fast nowadays. Free questions, a public API, three friends, no product. Harvey is a great success story in a lot of ways but good on them for their scrappy-ness.
