arxiv:2409.08248

TextBoost: Towards One-Shot Personalization of Text-to-Image Models via Fine-tuning Text Encoder

Published on Sep 12

· Submitted by

akhaliq on Sep 13

Upvote

Authors:

NaHyeon Park ,

Kunhee Kim ,

Abstract

Recent breakthroughs in text-to-image models have opened up promising research avenues in personalized image generation, enabling users to create diverse images of a specific subject using natural language prompts. However, existing methods often suffer from performance degradation when given only a single reference image. They tend to overfit the input, producing highly similar outputs regardless of the text prompt. This paper addresses the challenge of one-shot personalization by mitigating overfitting, enabling the creation of controllable images through text prompts. Specifically, we propose a selective fine-tuning strategy that focuses on the text encoder. Furthermore, we introduce three key techniques to enhance personalization performance: (1) augmentation tokens to encourage feature disentanglement and alleviate overfitting, (2) a knowledge-preservation loss to reduce language drift and promote generalizability across diverse prompts, and (3) SNR-weighted sampling for efficient training. Extensive experiments demonstrate that our approach efficiently generates high-quality, diverse images using only a single reference image while significantly reducing memory and storage requirements.

View arXiv page View PDF Add to collection

Community

akhaliq

Paper submitter 7 days ago

https://textboost.github.io/

jackyhate

6 days ago

Very interesting job. Does the training process involve different data augmentation methods? If so, do they correspond to different enhanced pseudo-words A*?

nahyeonkaty

Paper author 6 days ago

Thank you for your interest in our work :) We indeed applied various types of augmentations, including a range of geometric and color transformations (Figure 10). As you kindly mentioned, each augmentation corresponds to a specific A*. For more details on the technical implementation, please feel free to visit our GitHub repository!