First Advisor

Wu-chi Feng

Term of Graduation

Spring 2026

Date of Publication

6-3-2026

Document Type

Dissertation

Degree Name

Doctor of Philosophy (Ph.D.) in Computer Science

Department

Computer Science

Language

English

Subjects

Video Compression, Video Frame Interpolation

Physical Description

1 online resource (xi, 86 pages)

Abstract

Video has become the dominant form of online content, driven by the rapid growth of streaming, real-time communication, and immersive applications such as virtual and augmented reality. As resolutions and frame rates continue to increase, the demand for efficient video compression is steadily rising. Given the scale of video traffic, even modest improvements in compression efficiency can lead to significant gains in storage and transmission.

Traditional video codecs have achieved remarkable rate--distortion performance through decades of engineering and optimization. However, these methods rely on complex, hand-crafted designs and increasingly sophisticated prediction structures, which introduce substantial computational complexity and limit their adaptability to evolving content and application requirements. Recent neural video compression approaches attempt to learn spatio-temporal representations directly from data, but often suffer from limited flexibility across bitrate ranges and increased model size, making practical deployment challenging.

Video frame interpolation (VFI) provides an alternative perspective by synthesizing intermediate frames between sparsely transmitted keyframes, potentially reducing the overall bitrate. While modern CNN-based VFI methods demonstrate strong motion modeling capabilities, they remain sensitive to challenging scenarios such as occlusion, motion ambiguity, and large temporal gaps, especially when operating on compressed inputs.

In this thesis, we propose a hybrid compression framework that integrates the efficiency of traditional codecs with the motion modeling capability of neural interpolation. Specifically, keyframes are encoded using a conventional codec, while intermediate frames are reconstructed using a CNN-based VFI model guided by low-resolution, compressed representations of the target frame, referred to as hints. These hints provide additional structural information that helps resolve motion ambiguity and improves robustness under challenging conditions, including compression artifacts and complex motion patterns. The proposed framework enables more efficient compression while maintaining high visual fidelity, making it well suited for modern video applications.

Rights

© 2026 Pan Tan

In Copyright. URI: http://rightsstatements.org/vocab/InC/1.0/ This Item is protected by copyright and/or related rights. You are free to use this Item in any way that is permitted by the copyright and related rights legislation that applies to your use. For other uses you need to obtain permission from the rights-holder(s).

Persistent Identifier

https://archives.pdx.edu/ds/psu/44980

Available for download on Thursday, June 03, 2027

Share

COinS