Toward extracting and exploiting generalizable knowledge of deep 2D transformations in computer vision

Jiachen Kang; Wenjing Jia; Xiangjian He

doi:10.1016/j.neucom.2023.126882

Toward extracting and exploiting generalizable knowledge of deep 2D transformations in computer vision

Jiachen Kang, Wenjing Jia, Xiangjian He

School of Computer Science

Research output: Journal Publication › Article › peer-review

4 Citations (Scopus)

Abstract

Existing deep learning models suffer from out-of-distribution (o.o.d.) performance drop in computer vision tasks. In comparison, humans have a remarkable ability to interpret images, even if the scenes in the images are rare, thanks to the generalizability of acquired knowledge. This work attempts to answer two research questions: (1) the acquisition and (2) the utilization of generalizable knowledge about 2D transformations. To answer the first question, we demonstrate that deep neural networks can learn generalizable knowledge with a new training methodology based on synthetic datasets. The generalizability is reflected in the results that, even when the knowledge is learned from random noise, the networks can still achieve stable performance in parameter estimation tasks. To answer the second question, a novel architecture called “InterpretNet” is devised to utilize the learned knowledge in image classification tasks. The architecture consists of an ESTIMATOR and an IDENTIFIER, in addition to a CLASSIFIER. By emulating the “hypothesis-verification” process in human visual perception, our InterpretNet improves classification accuracy by 21.1%.

Original language	English
Article number	126882
Journal	Neurocomputing
Volume	562
DOIs	https://doi.org/10.1016/j.neucom.2023.126882
Publication status	Published - 28 Dec 2023

Keywords

Computer vision
Deep learning
Explainability
Knowledge acquisition
O.O.D. generalization

ASJC Scopus subject areas

Computer Science Applications
Cognitive Neuroscience
Artificial Intelligence

Access to Document

10.1016/j.neucom.2023.126882

Cite this

@article{bed5728bb99944be9d6472effbfbeef8,

title = "Toward extracting and exploiting generalizable knowledge of deep 2D transformations in computer vision",

abstract = "Existing deep learning models suffer from out-of-distribution (o.o.d.) performance drop in computer vision tasks. In comparison, humans have a remarkable ability to interpret images, even if the scenes in the images are rare, thanks to the generalizability of acquired knowledge. This work attempts to answer two research questions: (1) the acquisition and (2) the utilization of generalizable knowledge about 2D transformations. To answer the first question, we demonstrate that deep neural networks can learn generalizable knowledge with a new training methodology based on synthetic datasets. The generalizability is reflected in the results that, even when the knowledge is learned from random noise, the networks can still achieve stable performance in parameter estimation tasks. To answer the second question, a novel architecture called “InterpretNet” is devised to utilize the learned knowledge in image classification tasks. The architecture consists of an ESTIMATOR and an IDENTIFIER, in addition to a CLASSIFIER. By emulating the “hypothesis-verification” process in human visual perception, our InterpretNet improves classification accuracy by 21.1%.",

keywords = "Computer vision, Deep learning, Explainability, Knowledge acquisition, O.O.D. generalization",

author = "Jiachen Kang and Wenjing Jia and Xiangjian He",

note = "Publisher Copyright: {\textcopyright} 2023 Elsevier B.V.",

year = "2023",

month = dec,

day = "28",

doi = "10.1016/j.neucom.2023.126882",

language = "English",

volume = "562",

journal = "Neurocomputing",

issn = "0925-2312",

publisher = "Elsevier B.V.",

}

TY - JOUR

T1 - Toward extracting and exploiting generalizable knowledge of deep 2D transformations in computer vision

AU - Kang, Jiachen

AU - Jia, Wenjing

AU - He, Xiangjian

PY - 2023/12/28

Y1 - 2023/12/28

N2 - Existing deep learning models suffer from out-of-distribution (o.o.d.) performance drop in computer vision tasks. In comparison, humans have a remarkable ability to interpret images, even if the scenes in the images are rare, thanks to the generalizability of acquired knowledge. This work attempts to answer two research questions: (1) the acquisition and (2) the utilization of generalizable knowledge about 2D transformations. To answer the first question, we demonstrate that deep neural networks can learn generalizable knowledge with a new training methodology based on synthetic datasets. The generalizability is reflected in the results that, even when the knowledge is learned from random noise, the networks can still achieve stable performance in parameter estimation tasks. To answer the second question, a novel architecture called “InterpretNet” is devised to utilize the learned knowledge in image classification tasks. The architecture consists of an ESTIMATOR and an IDENTIFIER, in addition to a CLASSIFIER. By emulating the “hypothesis-verification” process in human visual perception, our InterpretNet improves classification accuracy by 21.1%.

AB - Existing deep learning models suffer from out-of-distribution (o.o.d.) performance drop in computer vision tasks. In comparison, humans have a remarkable ability to interpret images, even if the scenes in the images are rare, thanks to the generalizability of acquired knowledge. This work attempts to answer two research questions: (1) the acquisition and (2) the utilization of generalizable knowledge about 2D transformations. To answer the first question, we demonstrate that deep neural networks can learn generalizable knowledge with a new training methodology based on synthetic datasets. The generalizability is reflected in the results that, even when the knowledge is learned from random noise, the networks can still achieve stable performance in parameter estimation tasks. To answer the second question, a novel architecture called “InterpretNet” is devised to utilize the learned knowledge in image classification tasks. The architecture consists of an ESTIMATOR and an IDENTIFIER, in addition to a CLASSIFIER. By emulating the “hypothesis-verification” process in human visual perception, our InterpretNet improves classification accuracy by 21.1%.

KW - Computer vision

KW - Deep learning

KW - Explainability

KW - Knowledge acquisition

KW - O.O.D. generalization

UR - http://www.scopus.com/inward/record.url?scp=85174733772&partnerID=8YFLogxK

U2 - 10.1016/j.neucom.2023.126882

DO - 10.1016/j.neucom.2023.126882

M3 - Article

AN - SCOPUS:85174733772

SN - 0925-2312

VL - 562

JO - Neurocomputing

JF - Neurocomputing

M1 - 126882

ER -

Toward extracting and exploiting generalizable knowledge of deep 2D transformations in computer vision

Abstract

Keywords

ASJC Scopus subject areas

Access to Document

Other files and links

Fingerprint

Cite this