Google Discovery Violations in Viacom v. YouTube

This post is part of Revisiting Litigation Alleging Google Discovery Violations.

Viacom International, Inc. v. YouTube, Inc. – docket.  1:07-CV-02103-LLS (S.D.N.Y.).

Filed March 13, 2007.  First filing as to discovery violations: March 18, 2010.

Case allegation: “YouTube’s website purports to be a forum for users to share their own original ‘user generated’ video content. In reality, however, a vast amount of that content consists of infringing copies of Plaintiffs’ copyrighted works.”  Complaint.

Case disposition: Summary judgment granted for defendant based on the safe-harbor provisions of the Digital Millenium Copyright Act.  Viacom appealed.  The parties announced a settlement in March 2014 and announced that no money changed hands.

Viacom’s Motion for Summary Judgment at heading “Defendants Cannot Walk Away from Their Contemporaneous Internal Documents.”

As to YouTube’s knowledge of copyright infringement: “The internal emails and memoranda of YouTube’s founders and Google’s senior executives discussed above make a compelling and indisputable record of Defendants’ intent to use infringing videos clips to build the YouTube business.”  Viacom’s Motion, pages 5 to through 21, offer a dozen quotes supporting this contention, such as “we’re hosting copyrighted content” and “we need views, [but] I’m a little concerned with the recent Supreme Court ruling on copyright content.”  “We’re going to have a tough time defending the fact that we’re not liable for the copyrighted material on the site because we didn’t put it up when one of the co-founders is blatantly stealing content from other sites and trying to get everyone to see it.”

As to YouTube founder documents: “Defendants fail[ed] to preserve and produce many key documents.”  “Almost none of … the internal emails and memoranda of YouTube’s founders … were produced by Google or YouTube, which claims they were all lost.”  “Chad Hurley, a founder and YouTube’s Chief Executive from its inception to today, revealed for the first time [at] his deposition that he ‘lost all’ of his YouTube emails for the key time period of this case.”

As to documents from Google CEO Eric Schmidt: He “claims to use and email from ‘probably 30’ different computers … [y]et [his] search for responsive materials [pertaining to YouTube’s policies and practices and Google’s acquisition] ‘yielded 19 documents.’  Schmidt explained: ‘[i]t has been my practice for 30 years to not retain my emails unless asked specifically. … ‘It was my practice to delete or otherwise cause the emails that I had read to go away as quickly as possible’.”

On witness memories: Viacom remarks on “the ostensible memory failures of their key executives when deposed.”  YouTube founder Chad Hurley “developed serial amnesia” in his deposition.  Google go-founder Larry Page “essentially disclaimed memory on any topic relevant to this litigation, even including, for example, whether he was in favor of Google’s acquisition of YouTube, even though it was Google’s largest corporate transaction to date and viewed as transformative to its business.”   (See excerpted deposition transcript.)  “This Court can decide whether these key executives and witnesses behaved with the level of candor and respect for the legal process that this Court has a right to expect from senior executives of important public companies.”

Viacom never filed a motion for sanctions for spoliation.

***

Deposition of Google then-CEO Eric Schmidt including his explanation why he had just 19 emails pertaining to YouTube (“[i]t has been my practice to not keep my e-mails”) as well as claiming not to remember numerous facts about YouTube and Google’s strategy in video.

Deposition of Google co-founder Larry Page with analysis of his recollection.  By my count, Page said he did remember when asked about 132 aspects of Google and YouTube practices.

Deposition of YouTube co-founder Chad Hurley (parts 1, 2, 3).  Hurley had no difficulty recalling most aspects of product strategy, negotiation strategy, distribution strategy, and other topics.  However he was unable to recall multiple subjects that were sensitive in light of the litigation: the basis of Google’s valuation of YouTube (part 1 pages 29-32), instances of entire movies uploaded to YouTube (part 2 pages 12-16), copyright enforcement policies (part 2 pages 19-20 and 26-30), and why YouTube removed ads from Watch Pages (part 2 pages 23-24).

Google’s Opposition to Plaintiffs’ Motions for Partial Summary Judgment.  Makes no mention of Viacom’s allegations of documents inexplicably lost, of emails deleted, or witness’ memories.

Order of June 23, 2010 granted Google’s motion for summary judgment without ruling on any of the discovery issues raised in Viacom’s motion.

Eric Schmidt’s Missing Emails and Memory Problems

In Viacom v. Google, Viacom took the deposition of Google then-CEO Eric Schmidt.  In sworn testimony, Page explained why he had just 19 emails pertaining to YouTube, then Google’s largest acquisition.  (Source: Hohengarten declaration ¶266)  Schmidt also repeatedly reported  not remembering events and facts relating to YouTube.

Viacom remarked in associated briefing: “This Court can decide whether these key executives and witnesses behaved with the level of candor and respect for the legal process that this Court has a right to expect from senior executives of important public companies.”

Schmidt’s testimony (as excerpted by Viacom).

As to document retention: “[Some] people over time either delete or lose some of that e-mail. It has been my practice for 30 years to not retain my e-mails unless asked specifically.”  “[I]t has been my practice to not keep my e-mails.”  “Q: And is this on some sort of automatic system where they are deleted in the ordinary course over some ordinary period of time?  A: Depending on the e-mail system and the company and so forth, the answer would vary.”  As to document retention while at Google: “It was my practice to delete or otherwise cause the e-mails that I had read to go away as quickly as possible.  Q: Within days?  A: Yes.”

Separately, Schmidt claimed he didn’t remember multiple aspects of Google’s strategy in video.  A representative example:

Q: [Y]ou are aware, I assume, that the acquisition agreement contains an indemnification provision relating to copyright lawsuits?  … A: Yes. Q: And was that discussed by the board …?  A: Yes.  Q: And do you remember that discussion, sir?  A: No.”

Other subjects Schmidt didn’t remember.

Larry Page’s Bad Memory of YouTube and Google Video

In Viacom v. Google copyright litigation, Viacom took the deposition of Google co-founder Larry Page.  In a sworn deposition, Page reported remembering very little about YouTube and Google video.

Viacom remarked in associated briefing: “This Court can decide whether these key executives and witnesses behaved with the level of candor and respect for the legal process that this Court has a right to expect from senior executives of important public companies.”

Page’s full deposition is available in parts 1, 2, 3.

In the list below, I summarize the substantive questions asked of Page, and quote his verbatim response to each.  By my count, there were 132 distinct instances in which Page could not recall the answer to a question.  (In this  count and list, I excluded questions about whether Page received documents, excluded questions about why the reasons why he doesn’t remember, and excluded or grouped most repeat questions.)  From my line-by-line review of Page’s deposition:

The 2008 Walker Memo

In recent competition proceedings against Google, multiple courts have discussed whether Google staff were forthright in their remarks (not to mention whether the company preserved and produced documents in accordance with court rules).  See Google Discovery Violations.  Discussions often turn back to a memo sent by Google General Counsel Kent Walker as well as Bill Coughran (then SVP of Engineering at Google).  In the dockets for these cases, this document appears in unformatted monospace with broken spacing between paragraphs, and in image form without indexing or searchability.  I am therefore reposting it here in plain text.

Sent: September 16, 2008 2:09 PM
Subject: Business communications in a complicated world
Confidential/Please Do Not forward

Googlers –

As you know, Google continues to be in the midst of several significant legal and regulatory matters, including government reviews of our deal with Yahoo!, various copyright, patent, and trademark lawsuits, and lots of other claims. Given our continuing commitment to developing revolutionary products and doing disruptive things, we’re going to keep facing these kinds of challenges. So we’ve got two requests of you and one change to announce.

First, please write carefully and thoughtfully. We’re an email and instant-messaging culture. we conduct much of our work online. We believe that information is good. But anything you write can become subject to review in legal discovery, misconstrued, or taken out of context, and may be used against you or us in ways you wouldn’t expect. Writing stuff that’s sarcastic, speculative, or not fully informed inevitably creates problems in litigation. In your communications, please avoid stating legal conclusions. Speculation about whether something might breach a complex contract, or whether it might violate a law somewhere in the world, is often wrong and rarely helpful. So please do think twice before you write about hot topics, don’t comment before you have all the facts, and direct questions regarding continuing litigation holds and any legal and/or regulatory matters involving Google to the friendly (albeit lawyerly) folks at [redacted]@google.com

Second, remember that these same rules apply not just to Gmail but also to Google Talk and all other forms of electronic communication (for example wiki’s, doc’s, spreadsheets, etc.). We end up reviewing millions of pages of these communications as part of producing documents in regulatory and litigation matters — and we’re working together to streamline and simplify that process.

To help avoid inadvertent retention of instant messages, we have decided to make “off the record” the Google corporate default setting for Google Talk. We’ll also be providing this option to our Google Apps enterprise customers. You should see this new default setting taking effect over the next few days. You will still be able to save Talk conversations that are useful to you — but please remember that “on the record” conversations become part of your (more or less) permanent record and are added to Google’s long-term document storehouse. If you’ve received notice that you’re subject to a litigation hold, and you must chat regarding matters covered by that hold, please make sure that those chats are “on the record”.

Finally, remember that even when you’re “off the record”, your chat partner may be recording the conversation, so always take care with what you write. Thanks for your help and understanding on this. Let one of us know if you have any questions.

Bill Coughran

Kent Walker

See also required “Communicate with Care” training.

See also Antitrust Basics for Search Team (March 2011), recommending word choice for Google employees discussing competition matters.

Impact of GitHub Copilot on code quality

Jared Bauer summarizes results of a study I suggested this spring.  202 developers were randomly assigned GitHub Copilot, while the others were instructed not to use AI tools.  The participants were asked to complete a coding task.  Developers with GitHub Copilot had 56% greater likelihood of passing all unit tests.  Other developers evaluated code to assess quality and readability.  Code from developers with GitHub Copilot was rated better on readability, maintainability, and conciseness.  All these differences were statistically significant.

The Effect of Microsoft Copilot in a Multi-lingual Context with Donald Ngwe

We tested Microsoft Copilot in multilingual contexts, examining how Copilot can facilitate collaboration between colleagues with different native languages.

First, we asked 77 native Japanese speakers to review a meeting recorded in English. Half the participants had to watch and listen to the video. The other half could use Copilot Meeting Recap, which gave them an AI meeting summary as well as a chatbot to answer questions about the meeting.

Then, we asked 83 other native Japanese speakers to review a similar meeting, following the same script, but this time held in Japanese by native Japanese speakers. Again, half of participants had access to Copilot.

For the meeting in English, participants with Copilot answered 16.4% more multiple-choice questions about the meeting correctly, and they were more than twice as likely to get a perfect score.  Moreover, in comparing accuracy between the two scenarios, people listening to a meeting in English with Copilot achieved 97.5% accuracy, slightly more accurate than people listening to a meeting in their native Japanese using standard tools (94.8%). This is a statistically significant difference (p<.05). The changes are small in percentage point terms because the baseline accuracy is so high, but Copilot closed 38.5% of the gap to perfect accuracy for those working in their native language (p<0.10) and closed 84.6% of the gap for those working in (non-native) English (p<.05).

 

Summary from Jaffe et al, Generative AI in Real-World Workplaces, July 2024.

Impact of M365 Copilot on Legal Work at Microsoft

Teams at Microsoft often reflect on how Copilot helps.  I try to help these teams both by measuring Copilot usage in the field (as they do their ordinary work) and in lab experiments (idealized versions of their tasks in environments where I can better isolate cause and effect).  This month I ran an experiment with CELA, Microsoft’s in-house legal department.  Hossein Nowbar, Chief Legal Officer and Corporate Vice President, summarized the findings in a post at LinkedIn:

Recently, we ran a controlled experiment with Microsoft’s Office of the Chief Economist, and the results are groundbreaking. In this experiment, we asked legal professional volunteers on our team to complete three realistic legal tasks and randomly granted Copilot to some participants. Individuals with Copilot completed the tasks 32% faster and with 20.3% greater accuracy!

Copilot isn’t just a tool; it’s a game-changer, empowering our team to focus on what truly matters by enhancing productivity, elevating work quality, and, most importantly, reclaiming time.

All findings statistically significant at P<0.05.

Full results.

Early LLM-based Tools for Enterprise Information Workers Likely Provide Meaningful Boosts to Productivity

Early LLM-based Tools for Enterprise Information Workers Likely Provide Meaningful Boosts to Productivity. Microsoft Research Report – AI and Productivity Team. With Alexia Cambon, Brent Hecht, Donald Ngwe, Sonia Jaffe, Amy Heger, Mihaela Vorvoreanu, Sida Peng, Jake Hofman, Alex Farach, Margarita Bermejo-Cano, Eric Knudsen, James Bono, Hardik Sanghavi, Sofia Spatharioti, David Rothschild, Daniel G. Goldstein, Eirini Kalliamvakou, Peter Cihon, Mert Demirer, Michael Schwarz, and Jaime Teevan.

This report presents the initial findings of Microsoft’s research initiative on “AI and Productivity”, which seeks to measure and accelerate the productivity gains created by LLM-powered productivity tools like Microsoft’s Copilot. The many studies summarized in this report, the initiative’s first, focus on common enterprise information worker tasks for which LLMs are most likely to provide significant value. Results from the studies support the hypothesis that the first versions of Copilot tools substantially increase productivity on these tasks. This productivity boost usually appeared in the studies as a meaningful increase in speed of execution without a significant decrease in quality. Furthermore, we observed that the willingness-to-pay for LLM-based tools is higher for people who have used the tools than those who have not, suggesting that the tools provide value above initial expectations. The report also highlights future directions for the AI and Productivity initiative, including an emphasis on approaches that capture a wider range of tasks and roles.

Studies I led that are included within this report:

Randomized Controlled Trials for Microsoft Copilot for Security with James Bono, Sida Peng, Roberto Rodriguez, and Sandra Ho. updated March 29, 2024.

Randomized Controlled Trials for Microsoft Copilot for Security. SSRN Working Paper 4648700. With James Bono, Sida Peng, Roberto Rodriguez, and Sandra Ho.

We conducted randomized controlled trials (RCTs) to measure the efficiency gains from using Security Copilot, including speed and quality improvements. External experimental subjects logged into a M365 Defender instance created for this experiment and performed four tasks: Incident Summarization, Script Analyzer, Incident Report, and Guided Response. We found that Security Copilot delivered large improvements on both speed and accuracy. Copilot brought improvements for both novices and security professionals.

(Also summarized in What Can Copilot’s Earliest Users Teach Us About Generative AI at Work? at “Role-specific pain points and opportunities: Security.” Also summarized in AI and Productivity Report at “M365 Defender Security Copilot study.”)

Sound Like Me: Findings from a Randomized Experiment with Donald Ngwe

Sound Like Me: Findings from a Randomized Experiment. SSRN Working Paper 4648689. With Donald Ngwe.

A new version of Copilot for Microsoft 365 includes a feature to let Outlook draft messages that “Sound Like Me” (SLM) based on training from messages in a user’s Sent Items folder. We sought to evaluate whether SLM lives up to its name. We find that it does, and more. Users widely and systematically praise SLM-generated messages as being more clear, more concise, and more “couldn’t have said it better myself”. When presented with a human-written message versus a SLM rewrite, users say they’d rather receive the SLM rewrite. All these findings are statistically significant. Furthermore, when presented with human and SLM messages, users struggle to tell the difference, in one specification doing worse than random.

(Also summarized in What Can Copilot’s Earliest Users Teach Us About Generative AI at Work? at “Email effectiveness.” Also summarized in AI and Productivity Report at “Outlook Email Study.”)