PaddleOCR/StyleText/README.md

English | [简体中文](README_ch.md)

## Style Text

### Contents
- [1. Introduction](#Introduction)
- [2. Preparation](#Preparation)
- [3. Quick Start](#Quick_Start)
- [4. Applications](#Applications)
- [5. Code Structure](#Code_structure)


<a name="Introduction"></a>
### Introduction

<div align="center">
    <img src="doc/images/3.png" width="800">
</div>

<div align="center">
    <img src="doc/images/9.png" width="600">
</div>


The Style-Text data synthesis tool is a tool based on Baidu and HUST cooperation research work, "Editing Text in the Wild" [https://arxiv.org/abs/1908.03047](https://arxiv.org/abs/1908.03047).

Different from the commonly used GAN-based data synthesis tools, the main framework of Style-Text includes:
* (1) Text foreground style transfer module.
* (2) Background extraction module.
* (3) Fusion module.

After these three steps, you can quickly realize the image text style transfer. The following figure is some results of the data synthesis tool.

<div align="center">
    <img src="doc/images/10.png" width="1000">
</div>


<a name="Preparation"></a>
#### Preparation

1. Please refer the [QUICK INSTALLATION](../doc/doc_en/installation_en.md) to install PaddlePaddle. Python3 environment is strongly recommended.
2. Download the pretrained models and unzip:

```bash
cd StyleText
wget https://paddleocr.bj.bcebos.com/dygraph_v2.0/style_text/style_text_models.zip
unzip style_text_models.zip
```

If you save the model in another location, please modify the address of the model file in `configs/config.yml`, and you need to modify these three configurations at the same time:

```
bg_generator:
  pretrain: style_text_models/bg_generator
...
text_generator:
  pretrain: style_text_models/text_generator
...
fusion_generator:
  pretrain: style_text_models/fusion_generator
```

<a name="Quick_Start"></a>
### Quick Start

#### Synthesis single image

1. You can run `tools/synth_image` and generate the demo image, which is saved in the current folder.

```python
python3 tools/synth_image.py -c configs/config.yml --style_image examples/style_images/2.jpg --text_corpus PaddleOCR --language en
```

* Note 1: The language options is correspond to the corpus. Currently, the tool only supports English(en), Simplified Chinese(ch) and Korean(ko).
* Note 2: Synth-Text is mainly used to generate images for OCR recognition models.
  So the height of style images should be around 32 pixels. Images in other sizes may behave poorly.
* Note 3: You can modify `use_gpu` in `configs/config.yml` to determine whether to use GPU for prediction.


For example, enter the following image and corpus `PaddleOCR`.

<div align="center">
    <img src="examples/style_images/2.jpg" width="300">
</div>

The result `fake_fusion.jpg` will be generated.

<div align="center">
    <img src="doc/images/4.jpg" width="300">
</div>

What's more, the medium result `fake_bg.jpg` will also be saved, which is the background output.

<div align="center">
    <img src="doc/images/7.jpg" width="300">
</div>


`fake_text.jpg` is the generated image with the same font style as `Style Input`.


<div align="center">
    <img src="doc/images/8.jpg" width="300">
</div>


#### Batch synthesis

In actual application scenarios, it is often necessary to synthesize pictures in batches and add them to the training set. StyleText can use a batch of style pictures and corpus to synthesize data in batches. The synthesis process is as follows:

1. The referenced dataset can be specifed in `configs/dataset_config.yml`:

   * `Global`：
     * `output_dir:`：Output synthesis data path.
   * `StyleSampler`：
     * `image_home`：style images' folder.
     * `label_file`：Style images' file list. If label is provided, then it is the label file path.
     * `with_label`：Whether the `label_file` is label file list.
   * `CorpusGenerator`：
     * `method`：Method of CorpusGenerator，supports `FileCorpus` and `EnNumCorpus`. If `EnNumCorpus` is used，No other configuration is needed，otherwise you need to set `corpus_file` and `language`.
     * `language`：Language of the corpus. Currently, the tool only supports English(en), Simplified Chinese(ch) and Korean(ko). 
     * `corpus_file`: Filepath of the corpus. Corpus file should be a text file which will be split by line-endings（'\n'）. Corpus generator samples one line each time.


Example of corpus file:
```
PaddleOCR
飞桨文字识别
StyleText
风格文本图像数据合成
```

We provide a general dataset containing Chinese, English and Korean (50,000 images in all) for your trial ([download link](https://paddleocr.bj.bcebos.com/dygraph_v2.0/style_text/chkoen_5w.tar)), some examples are given below :

<div align="center">
     <img src="doc/images/5.png" width="800">
</div>

2. You can run the following command to start synthesis task:

   ``` bash
   python3 tools/synth_dataset.py -c configs/dataset_config.yml
   ```

We also provide example corpus and images in `examples` folder.
    <div align="center">
        <img src="examples/style_images/1.jpg" width="300">
        <img src="examples/style_images/2.jpg" width="300">
    </div>
If you run the code above directly, you will get example output data in `output_data` folder.
You will get synthesis images and labels as below:
   <div align="center">
       <img src="doc/images/12.png" width="800">
   </div>
There will be some cache under the `label` folder. If the program exit unexpectedly, you can find cached labels there.
When the program finish normally, you will find all the labels in `label.txt` which give the final results.

<a name="Applications"></a>
### Applications
We take two scenes as examples, which are metal surface English number recognition and general Korean recognition, to illustrate practical cases of using StyleText to synthesize data to improve text recognition. The following figure shows some examples of real scene images and composite images:

<div align="center">
    <img src="doc/images/11.png" width="800">
</div>


After adding the above synthetic data for training, the accuracy of the recognition model is improved, which is shown in the following table:


| Scenario | Characters | Raw Data | Test Data | Only Use Raw Data</br>Recognition Accuracy | New Synthetic Data | Simultaneous Use of Synthetic Data</br>Recognition Accuracy | Index Improvement |
| -------- | ---------- | -------- | -------- | -------------------------- | ------------ | ---------------------- | -------- |
| Metal surface | English and numbers | 2203     | 650      | 0.5938                     | 20000        | 0.7546                 | 16%      |
| Random background | Korean       | 5631     | 1230     | 0.3012                     | 100000       | 0.5057                 | 20%      |


<a name="Code_structure"></a>
### Code Structure

```
StyleText
|-- arch                        // Network module files.
|   |-- base_module.py
|   |-- decoder.py
|   |-- encoder.py
|   |-- spectral_norm.py
|   `-- style_text_rec.py
|-- configs                     // Config files.
|   |-- config.yml
|   `-- dataset_config.yml
|-- engine                      // Synthesis engines.
|   |-- corpus_generators.py    // Sample corpus from file or generate random corpus.
|   |-- predictors.py           // Predict using network.
|   |-- style_samplers.py       // Sample style images.
|   |-- synthesisers.py         // Manage other engines to synthesis images.
|   |-- text_drawers.py         // Generate standard input text images.
|   `-- writers.py              // Write synthesis images and labels into files.
|-- examples                    // Example files.
|   |-- corpus
|   |   `-- example.txt
|   |-- image_list.txt
|   `-- style_images
|       |-- 1.jpg
|       `-- 2.jpg
|-- fonts                       // Font files.
|   |-- ch_standard.ttf
|   |-- en_standard.ttf
|   `-- ko_standard.ttf
|-- tools                       // Program entrance.
|   |-- __init__.py
|   |-- synth_dataset.py        // Synthesis dataset.
|   `-- synth_image.py          // Synthesis image.
`-- utils                       // Module of basic functions.
    |-- config.py
    |-- load_params.py
    |-- logging.py
    |-- math_functions.py
    `-- sys_funcs.py
```
-												fix typo

											
										
										
											4 years ago
+								English | [简体中文](README_ch.md)
-												change config; add doc

											
										
										
											4 years ago
-												fix doc

											
										
										
											4 years ago
+								## Style Text
-												fix style exit readme

											
										
										
											4 years ago
+								### Contents
 								- [1. Introduction](#Introduction)
 								- [2. Preparation](#Preparation)
-												fix doc

											
										
										
											4 years ago
+								- [3. Quick Start](#Quick_Start)
 								- [4. Applications](#Applications)
-												fix style exit readme

											
										
										
											4 years ago
+								- [5. Code Structure](#Code_structure)
-												change config; add doc

											
										
										
											4 years ago
-												fix style exit readme

											
										
										
											4 years ago
+								<a name="Introduction"></a>
 								### Introduction
-												change config; add doc

											
										
										
											4 years ago
-												fix style exit readme

											
										
										
											4 years ago
+								<div align="center">
 								    <img src="doc/images/3.png" width="800">
 								</div>
 								<div align="center">
-												fix doc

											
										
										
											4 years ago
+								    <img src="doc/images/9.png" width="600">
-												fix style exit readme

											
										
										
											4 years ago
+								</div>
-												Update README.md
											
										
										
											4 years ago
+								The Style-Text data synthesis tool is a tool based on Baidu and HUST cooperation research work, "Editing Text in the Wild" [https://arxiv.org/abs/1908.03047](https://arxiv.org/abs/1908.03047).
-												fix style exit readme

											
										
										
											4 years ago
 								Different from the commonly used GAN-based data synthesis tools, the main framework of Style-Text includes:
 								* (1) Text foreground style transfer module.
 								* (2) Background extraction module.
 								* (3) Fusion module.
-												fix typo

											
										
										
											4 years ago
+								After these three steps, you can quickly realize the image text style transfer. The following figure is some results of the data synthesis tool.
-												fix style exit readme

											
										
										
											4 years ago
 								<div align="center">
-												fix doc

											
										
										
											4 years ago
+								    <img src="doc/images/10.png" width="1000">
-												fix style exit readme

											
										
										
											4 years ago
+								</div>
 								<a name="Preparation"></a>
-												change config; add doc

											
										
										
											4 years ago
+								#### Preparation
-												rename files

											
										
										
											4 years ago
+. Please refer the [QUICK INSTALLATION](../doc/doc_en/installation_en.md) to install PaddlePaddle. Python3 environment is strongly recommended.
-												change config; add doc

											
										
										
											4 years ago
+. Download the pretrained models and unzip:
 								```bash
-												fix style exit readme

											
										
										
											4 years ago
+								cd StyleText
 								wget https://paddleocr.bj.bcebos.com/dygraph_v2.0/style_text/style_text_models.zip
-												change config; add doc

											
										
										
											4 years ago
+								unzip style_text_models.zip
 								```
-												fix style exit readme

											
										
										
											4 years ago
+								If you save the model in another location, please modify the address of the model file in `configs/config.yml`, and you need to modify these three configurations at the same time:
-												change config; add doc

											
										
										
											4 years ago
 								```
 								bg_generator:
-												Update README.md (#1687)

Instructions for modifying the 'configs/config.yml' error:
pretrain: style_text_models/bg_generator
											
										
										
											4 years ago
+								  pretrain: style_text_models/bg_generator
-												change config; add doc

											
										
										
											4 years ago
+								...
 								text_generator:
 								  pretrain: style_text_models/text_generator
 								...
 								fusion_generator:
 								  pretrain: style_text_models/fusion_generator
 								```
-												fix doc

											
										
										
											4 years ago
+								<a name="Quick_Start"></a>
 								### Quick Start
-												change config; add doc

											
										
										
											4 years ago
-												fix style exit readme

											
										
										
											4 years ago
+								#### Synthesis single image
-												change config; add doc

											
										
										
											4 years ago
-												fix doc

											
										
										
											4 years ago
+. You can run `tools/synth_image` and generate the demo image, which is saved in the current folder.
-												change config; add doc

											
										
										
											4 years ago
-												fix style exit readme

											
										
										
											4 years ago
+								```python
-												add support for cpu infer (#1480)

* add support for cpu infer

* fix readme
											
										
										
											4 years ago
+								python3 tools/synth_image.py -c configs/config.yml --style_image examples/style_images/2.jpg --text_corpus PaddleOCR --language en
-												change config; add doc

											
										
										
											4 years ago
+								```
-												增加language选项的说明 (#1810)

* Update README.md

* Update README_ch.md

* Update README.md
											
										
										
											4 years ago
+								* Note 1: The language options is correspond to the corpus. Currently, the tool only supports English(en), Simplified Chinese(ch) and Korean(ko).
-												add support for cpu infer (#1480)

* add support for cpu infer

* fix readme
											
										
										
											4 years ago
+								* Note 2: Synth-Text is mainly used to generate images for OCR recognition models.
-												add notice in height of style images. (#1475)

* add examples in doc

* add files

* modify files

* dbg typo

* add notice in size of style images
											
										
										
											4 years ago
+								  So the height of style images should be around 32 pixels. Images in other sizes may behave poorly.
-												add support for cpu infer (#1480)

* add support for cpu infer

* fix readme
											
										
										
											4 years ago
+								* Note 3: You can modify `use_gpu` in `configs/config.yml` to determine whether to use GPU for prediction.
-												add notice in height of style images. (#1475)

* add examples in doc

* add files

* modify files

* dbg typo

* add notice in size of style images
											
										
										
											4 years ago
-												fix doc

											
										
										
											4 years ago
 								For example, enter the following image and corpus `PaddleOCR`.
 								<div align="center">
 								    <img src="examples/style_images/2.jpg" width="300">
 								</div>
 								The result `fake_fusion.jpg` will be generated.
 								<div align="center">
 								    <img src="doc/images/4.jpg" width="300">
 								</div>
 								What's more, the medium result `fake_bg.jpg` will also be saved, which is the background output.
 								<div align="center">
 								    <img src="doc/images/7.jpg" width="300">
 								</div>
-												Update README.md
											
										
										
											4 years ago
+								`fake_text.jpg` is the generated image with the same font style as `Style Input`.
-												change config; add doc

											
										
										
											4 years ago
-												fix doc

											
										
										
											4 years ago
+								<div align="center">
 								    <img src="doc/images/8.jpg" width="300">
 								</div>
-												change config; add doc

											
										
										
											4 years ago
-												fix style exit readme

											
										
										
											4 years ago
+								#### Batch synthesis
-												change config; add doc

											
										
										
											4 years ago
-												fix doc

											
										
										
											4 years ago
+								In actual application scenarios, it is often necessary to synthesize pictures in batches and add them to the training set. StyleText can use a batch of style pictures and corpus to synthesize data in batches. The synthesis process is as follows:
-												change config; add doc

											
										
										
											4 years ago
 . The referenced dataset can be specifed in `configs/dataset_config.yml`:
-												fix style exit readme

											
										
										
											4 years ago
-												fix doc

											
										
										
											4 years ago
+								   * `Global`：
 								     * `output_dir:`：Output synthesis data path.
 								   * `StyleSampler`：
 								     * `image_home`：style images' folder.
 								     * `label_file`：Style images' file list. If label is provided, then it is the label file path.
 								     * `with_label`：Whether the `label_file` is label file list.
 								   * `CorpusGenerator`：
 								     * `method`：Method of CorpusGenerator，supports `FileCorpus` and `EnNumCorpus`. If `EnNumCorpus` is used，No other configuration is needed，otherwise you need to set `corpus_file` and `language`.
-												增加language选项的说明 (#1810)

* Update README.md

* Update README_ch.md

* Update README.md
											
										
										
											4 years ago
+								     * `language`：Language of the corpus. Currently, the tool only supports English(en), Simplified Chinese(ch) and Korean(ko).
-												Update README.md
											
										
										
											4 years ago
+								     * `corpus_file`: Filepath of the corpus. Corpus file should be a text file which will be split by line-endings（'\n'）. Corpus generator samples one line each time.
-												Update README.md
											
										
										
											4 years ago
-												python to python3

											
										
										
											4 years ago
+								Example of corpus file:
-												Update README.md
											
										
										
											4 years ago
+								```
 								PaddleOCR
 								飞桨文字识别
-												Update README.md
											
										
										
											4 years ago
+								StyleText
 								风格文本图像数据合成
-												Update README.md
											
										
										
											4 years ago
+								```
-												change config; add doc

											
										
										
											4 years ago
-												fix typo

											
										
										
											4 years ago
+								We provide a general dataset containing Chinese, English and Korean (50,000 images in all) for your trial ([download link](https://paddleocr.bj.bcebos.com/dygraph_v2.0/style_text/chkoen_5w.tar)), some examples are given below :
-												fix style exit readme

											
										
										
											4 years ago
 								<div align="center">
 								     <img src="doc/images/5.png" width="800">
 								</div>
-												change config; add doc

											
										
										
											4 years ago
+. You can run the following command to start synthesis task:
 								   ``` bash
-												add support for cpu infer (#1480)

* add support for cpu infer

* fix readme
											
										
										
											4 years ago
+								   python3 tools/synth_dataset.py -c configs/dataset_config.yml
-												change config; add doc

											
										
										
											4 years ago
+								   ```
-												add support for cpu infer (#1480)

* add support for cpu infer

* fix readme
											
										
										
											4 years ago
+								We also provide example corpus and images in `examples` folder.
-												Add batch examples in StyleText doc (#1469)

* add examples in doc

* add files

* modify files

* dbg typo
											
										
										
											4 years ago
+								    <div align="center">
 								        <img src="examples/style_images/1.jpg" width="300">
 								        <img src="examples/style_images/2.jpg" width="300">
 								    </div>
 								If you run the code above directly, you will get example output data in `output_data` folder.
 								You will get synthesis images and labels as below:
 								   <div align="center">
 								       <img src="doc/images/12.png" width="800">
 								   </div>
 								There will be some cache under the `label` folder. If the program exit unexpectedly, you can find cached labels there.
 								When the program finish normally, you will find all the labels in `label.txt` which give the final results.
-												fix style exit readme

											
										
										
											4 years ago
-												fix doc

											
										
										
											4 years ago
+								<a name="Applications"></a>
 								### Applications
-												fix style exit readme

											
										
										
											4 years ago
+								We take two scenes as examples, which are metal surface English number recognition and general Korean recognition, to illustrate practical cases of using StyleText to synthesize data to improve text recognition. The following figure shows some examples of real scene images and composite images:
 								<div align="center">
-												fix doc

											
										
										
											4 years ago
+								    <img src="doc/images/11.png" width="800">
-												fix style exit readme

											
										
										
											4 years ago
+								</div>
 								After adding the above synthetic data for training, the accuracy of the recognition model is improved, which is shown in the following table:
-												fix table (#1454)


											
										
										
											4 years ago
-												fix style exit readme

											
										
										
											4 years ago
+								| Scenario | Characters | Raw Data | Test Data | Only Use Raw Data</br>Recognition Accuracy | New Synthetic Data | Simultaneous Use of Synthetic Data</br>Recognition Accuracy | Index Improvement |
-												fix table (#1454)


											
										
										
											4 years ago
+								| -------- | ---------- | -------- | -------- | -------------------------- | ------------ | ---------------------- | -------- |
 								| Metal surface | English and numbers | 2203     | 650      | 0.5938                     | 20000        | 0.7546                 | 16%      |
 								| Random background | Korean       | 5631     | 1230     | 0.3012                     | 100000       | 0.5057                 | 20%      |
-												fix style exit readme

											
										
										
											4 years ago
 								<a name="Code_structure"></a>
 								### Code Structure
-												fix doc

											
										
										
											4 years ago
-												fix style exit readme

											
										
										
											4 years ago
+								```
-												fix conflict

											
										
										
											4 years ago
+								StyleText
-												fix doc

											
										
										
											4 years ago
+								|-- arch                        // Network module files.
-												fix style exit readme

											
										
										
											4 years ago
+								|   |-- base_module.py
 								|   |-- decoder.py
 								|   |-- encoder.py
 								|   |-- spectral_norm.py
 								|   `-- style_text_rec.py
-												fix doc

											
										
										
											4 years ago
+								|-- configs                     // Config files.
-												fix style exit readme

											
										
										
											4 years ago
+								|   |-- config.yml
 								|   `-- dataset_config.yml
-												fix doc

											
										
										
											4 years ago
+								|-- engine                      // Synthesis engines.
 								|   |-- corpus_generators.py    // Sample corpus from file or generate random corpus.
 								|   |-- predictors.py           // Predict using network.
 								|   |-- style_samplers.py       // Sample style images.
 								|   |-- synthesisers.py         // Manage other engines to synthesis images.
 								|   |-- text_drawers.py         // Generate standard input text images.
 								|   `-- writers.py              // Write synthesis images and labels into files.
 								|-- examples                    // Example files.
-												fix style exit readme

											
										
										
											4 years ago
+								|   |-- corpus
 								|   |   `-- example.txt
 								|   |-- image_list.txt
 								|   `-- style_images
 								|       |-- 1.jpg
 								|       `-- 2.jpg
-												fix doc

											
										
										
											4 years ago
+								|-- fonts                       // Font files.
-												fix style exit readme

											
										
										
											4 years ago
+								|   |-- ch_standard.ttf
 								|   |-- en_standard.ttf
 								|   `-- ko_standard.ttf
-												fix doc

											
										
										
											4 years ago
+								|-- tools                       // Program entrance.
-												fix style exit readme

											
										
										
											4 years ago
+								|   |-- __init__.py
-												fix doc

											
										
										
											4 years ago
+								|   |-- synth_dataset.py        // Synthesis dataset.
 								|   `-- synth_image.py          // Synthesis image.
 								`-- utils                       // Module of basic functions.
-												fix style exit readme

											
										
										
											4 years ago
+								    |-- config.py
 								    |-- load_params.py
 								    |-- logging.py
 								    |-- math_functions.py
 								    `-- sys_funcs.py
 								```