WORKING PAPERS
Summarize or Tease: How Thumbnails Represent Videos and Shape Viewing Decisions (with Oded Netzer). [Paper]
Abstract: Thumbnails serve as representations of video content when viewers are ``in the dark" before watching. Yet, it remains unclear whether thumbnails should summarize the video, like a synopsis, or tease the video. We study how thumbnails, relative to the video content they represent, affect video viewing decisions. Using multimodal LLMs and computer vision thumbnail-video mining approach we transform unstructured thumbnails and video content into interpretable features, and then construct theory-based measures to characterize both the thumbnail itself and its relationship with the video it represents. We use secondary data from YouTube to document real-world relationships between thumbnails and video performance, and then experimental data from CTube, a video platform we built to randomize thumbnail exposure, to estimate a joint model of video choice and watchtime. Our results reveal a fundamental click--watchtime tradeoff. ``Teaser-style" thumbnails that are visually appealing and emotionally engaging increase clicks, but these features do not sustain viewing. Instead, thumbnails affect watchtime through how the video unfolds relative to the thumbnail: viewers are more likely to quit when video content diverges from the thumbnail, but this penalty weakens after viewers observe the moment revealed by the thumbnail in the video. Effective thumbnails require balancing visual impact with content alignment and timing, while tailoring each thumbnail to viewers and video content. We use these insights to develop optimal thumbnail selection at both the individual and aggregate level, leading up to around 10% improvement in aggregate watchtime relative to the original thumbnail selected by users.
The Impact of Banning Online Gambling Livestreams: Evidence from Twitch.tv (with Qifan Han and Andrey Simonov). [Paper] [Interactive Network]
Major revision at Marketing Science
Abstract: Can platforms curb harmful content by banning only part of it? We study the short-run effects of Twitch's October 2022 ban on livestreams of unlicensed gambling websites, using a novel panel of the top 6,000 streamers in which we detect banned content from video clips, stream titles, and in-stream chats. Difference-in-differences estimates show that streamers of banned content cut weekly gambling streaming hours by 65% and overall streaming by 45%. Streamers of unbanned content also reduced gambling and overall content, indicating spillovers to unbanned gambling content - a response concentrated among popular streamers, in line with reputational concerns. On the demand side, viewership, low-tier subscriptions, and engagement fell, while high-tier subscriptions remained unaffected. Gambling exposure also fell among viewers likely to be minors. A narrow ban thus reduced gambling content well beyond its formal scope, but at a cost in platform activity.
Collaboration Among Content Creators (with Qifan Han and Kinshuk Jerath). [Paper]
Abstract: We study content collaboration in the creator economy, in which competing creators mutually agree to collaborate on joint content and negotiate on content production and revenue sharing. Using a game theory model with creators competing for consumers on a Hotelling line, we show that collaboration allows creators to use the jointly-produced content to moderate competition, while using their individual content to expand into new audiences. This increases content diversity but also leads to increased monetizability of content. In general, collaboration among creators has an effect of increasing the profits of creators while reducing consumer surplus. When creators create content with heterogeneous entertainment values, the creator producing content of lower entertainment value has an incentive to free ride on the collaborative content. This free riding may increase surplus for consumers (who without collaboration would watch content of low entertainment value), thereby improving creators' profits as well as consumer surplus. Our results provide guidance to content creators, to platforms designing tools to facilitate collaborations, and to policy makers.
People See Creative, Not Ads: Generating Video Creative Insights at Scale with Multimodal AI (with Poppy Zhang and Shawndra Hill)
Abstract: Video advertising is central to digital marketing, yet marketers still have limited systematic knowledge about what makes video creative effective. A key challenge is that video creatives are high-dimensional, multimodal, and context-dependent: the same creative element may work differently across product subverticals and marketing objectives. We introduce a framework for generating video creative insights across millions of ads. The framework conceptualizes video creative as a set of interpretable elements that capture what an ad shows, how it communicates, and how it delivers brand and product information. We operationalize the framework using multimodal large language models, guided by advertising domain knowledge, to transform unstructured video creatives into structured measures of creative content, execution, and messaging. We then apply the framework to millions of video ads to study two core creative strategy problems. First, we show how creative effectiveness varies across subverticals, allowing marketers to generate subvertical-specific creative insights rather than relying on one-size-fits-all best practices. Second, we examine how creative elements differentially predict direct-response versus brand outcomes among ads that resonate with users, revealing when creative strategies that drive immediate action diverge from those that build brand value. Together, the findings show that effective video creative depends on both market context and marketing objective, and there is no universal prescription for effective creative. More broadly, this research demonstrates how multimodal AI can serve as a scalable measurement layer for creative intelligence, enabling advertisers, researchers and managers to systematically understand and evaluate video creative designs.
Quantifying Video Storyline: A Multimodal LLM Approach Informed by The Theory of Narratology (with Kumara Kahatapitiya, Poppy Zhang, Fanny Yang and Shawndra Hill)
Abstract: The narrative structure of video advertising creatives lies along a continuum, with completely non-narrative ads on one end, and well-developed, moving stories on the other end. Given the complexity of analyzing video creatives at scale, there is limited understanding of whether and how video narrative design affects ad effectiveness. We introduce a video creative narrative framework and a pipeline for its measurement using multimodal large language models guided by advertising domain knowledge and narrative theory. The framework is centered on the core idea that ad creatives can be structured by functional intent, with each video scene unit serving a distinct, communicative function while delivering product and brand information within seconds. We develop a video advertising-specific functional role taxonomy and propose a two-step, recurrent, probability-based pipeline to scalably discover data-driven video ad story structures. We demonstrate the utility of our framework with an application to 48 million video creatives from a large social media platform. We show that story-based creatives significantly improve video ad performance and provide actionable design recommendations to guide advertisers and creative strategists in design experimentation. Our framework demonstrates the value of using multimodal LLM for scalable video insight generation, offering a versatile tool for systematic video storyline understanding.
The Value Exchange: What the Best Ads Get Right (with Gil Chaimovsky, Steve Golub, Derek Scott, Poppy Zhang, Amel Awadelkarim, Kumara Kahatapitiya, Fanny Yang and Shawndra Hill)
Abstract: Advertising performance metrics reveal whether an advertisement succeeded but provide limited information about the creative decisions associated with that success. We study approximately 370,000 video advertisements on a large social media platform and use multimodal AI to measure video hooks, storyline arcs, and a broad set of full-video creative features. We focus on video ads that combine high platform-assessed user value with strong performance across direct response, sustained attention, and active engagement, and examine which creative characteristics distinguish these ads across outcomes. Several recurring patterns emerge. Creative decisions are important for video attention and engagement but are generally hard to move conversions. In particular, human-centered, affective, and story-driven creative is consistently associated with stronger attention; as viewing deepens, product demonstration and coherent storyline structure become increasingly prominent. Social response is associated with a different set of characteristics, including culturally specific language, recognizable settings, audience cues, and direct address. We also find that several creative forms associated with stronger attention, including personal storytelling and structured storyline arcs, are comparatively underused. The resulting evidence provides a basis for generating targeted creative alternatives, such as new hooks, storyline strucutres, and product presentations, that selectively vary either one part of the ad or whole ad for experimentation.
SELECTED WORK IN PROGRESS
Thumbnails as Visual Expectations: A Bayesian Learning Approach (with Oded Netzer).
Abstract: We build a Bayesian learning model to investigate how thumbnails, previews of video content, affect consumers’ reactions to videos in terms of video choice and watchtime via two roles: thumbnails as expectation-based reference points that shape viewers’ expectations of the video content before they start watching a video, and thumbnails as informational reference points that build viewers’ anticipation for a video. We model consumers' decisions to click on a video and continue watching the video as based on their priors (the thumbnail) and updated beliefs of the video content (the video's frames, characterized as multi-dimensional and correlated video topic proportions). We create a novel video streaming platform called "CTube" (a simplified version of YouTube) to conduct a novel field-framed online experiment by randomizing thumbnails viewers see to start a video. Leveraging the high-frequency clickstream data tracked by our video platform, we estimate the bayesian learning model by restoring from clickstream data individuals’ exposed video thumbnails, video choice and watchtime decisions. Our results suggest that viewers overall prefer watching videos longer when there is a higher disconfirmation between their initial content beliefs based on the thumbnail and updated beliefs based on the observed video scenes (signals). In addition, viewers prefer less content disconfirmation before observing the thumbnail, highlighting that the role of disconfirmation may change before and after viewers observe the moment highlighted by the thumbnail. Based on the model's estimates, we then run a series of counterfactual analyses to propose optimal thumbnails and compare them with current practices of thumbnail recommendation to guide creators and platforms in thumbnail selection.