Continual Learning of Neural Machine Translation within Low Forgetting Risk Regions

Gu, Shuhao; Hu, Bojie; Feng, Yang

Computer Science > Computation and Language

arXiv:2211.01542 (cs)

[Submitted on 3 Nov 2022 (v1), last revised 4 Nov 2022 (this version, v2)]

Title:Continual Learning of Neural Machine Translation within Low Forgetting Risk Regions

Authors:Shuhao Gu, Bojie Hu, Yang Feng

View PDF

Abstract:This paper considers continual learning of large-scale pretrained neural machine translation model without accessing the previous training data or introducing model separation. We argue that the widely used regularization-based methods, which perform multi-objective learning with an auxiliary loss, suffer from the misestimate problem and cannot always achieve a good balance between the previous and new tasks. To solve the problem, we propose a two-stage training method based on the local features of the real loss. We first search low forgetting risk regions, where the model can retain the performance on the previous task as the parameters are updated, to avoid the catastrophic forgetting problem. Then we can continually train the model within this region only with the new training data to fit the new task. Specifically, we propose two methods to search the low forgetting risk regions, which are based on the curvature of loss and the impacts of the parameters on the model output, respectively. We conduct experiments on domain adaptation and more challenging language adaptation tasks, and the experimental results show that our method can achieve significant improvements compared with several strong baselines.

Comments:	EMNLP 2022 Main Conference Long Paper
Subjects:	Computation and Language (cs.CL)
Cite as:	arXiv:2211.01542 [cs.CL]
	(or arXiv:2211.01542v2 [cs.CL] for this version)
	https://doi.org/10.48550/arXiv.2211.01542

Submission history

From: Shuhao Gu [view email]
[v1] Thu, 3 Nov 2022 01:21:10 UTC (5,546 KB)
[v2] Fri, 4 Nov 2022 02:10:02 UTC (5,546 KB)

Computer Science > Computation and Language

Title:Continual Learning of Neural Machine Translation within Low Forgetting Risk Regions

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computation and Language

Title:Continual Learning of Neural Machine Translation within Low Forgetting Risk Regions

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators