A survey on knowledge-enhanced multimodal learning

Lymperaiou, Maria; Stamou, Giorgos

Computer Science > Machine Learning

arXiv:2211.12328 (cs)

[Submitted on 19 Nov 2022 (v1), last revised 23 Mar 2024 (this version, v3)]

Title:A survey on knowledge-enhanced multimodal learning

Authors:Maria Lymperaiou, Giorgos Stamou

View PDF

Abstract:Multimodal learning has been a field of increasing interest, aiming to combine various modalities in a single joint representation. Especially in the area of visiolinguistic (VL) learning multiple models and techniques have been developed, targeting a variety of tasks that involve images and text. VL models have reached unprecedented performances by extending the idea of Transformers, so that both modalities can learn from each other. Massive pre-training procedures enable VL models to acquire a certain level of real-world understanding, although many gaps can be identified: the limited comprehension of commonsense, factual, temporal and other everyday knowledge aspects questions the extendability of VL tasks. Knowledge graphs and other knowledge sources can fill those gaps by explicitly providing missing information, unlocking novel capabilities of VL models. In the same time, knowledge graphs enhance explainability, fairness and validity of decision making, issues of outermost importance for such complex implementations. The current survey aims to unify the fields of VL representation learning and knowledge graphs, and provides a taxonomy and analysis of knowledge-enhanced VL models.

Subjects:	Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
Cite as:	arXiv:2211.12328 [cs.LG]
	(or arXiv:2211.12328v3 [cs.LG] for this version)
	https://doi.org/10.48550/arXiv.2211.12328

Submission history

From: Maria Lymperaiou [view email]
[v1] Sat, 19 Nov 2022 14:00:50 UTC (19,397 KB)
[v2] Mon, 6 Feb 2023 18:56:53 UTC (19,398 KB)
[v3] Sat, 23 Mar 2024 08:48:14 UTC (1,197 KB)

Computer Science > Machine Learning

Title:A survey on knowledge-enhanced multimodal learning

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Machine Learning

Title:A survey on knowledge-enhanced multimodal learning

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators