<?xml version="1.0" encoding="UTF-8"?><?xml-stylesheet type="text/xsl" href="static/style.xsl"?><OAI-PMH xmlns="http://www.openarchives.org/OAI/2.0/" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xsi:schemaLocation="http://www.openarchives.org/OAI/2.0/ http://www.openarchives.org/OAI/2.0/OAI-PMH.xsd"><responseDate>2026-09-18T19:23:07Z</responseDate><request verb="GetRecord" identifier="oai:digital.library.adelaide.edu.au:2440/113587" metadataPrefix="dim">https://digital.library.adelaide.edu.au/server/oai/request</request><GetRecord><record><header><identifier>oai:digital.library.adelaide.edu.au:2440/113587</identifier><datestamp>2026-06-12T08:33:23Z</datestamp><setSpec>com_2440_14759</setSpec><setSpec>col_2440_14760</setSpec></header><metadata><dim:dim xmlns:dim="http://www.dspace.org/xmlns/dspace/dim" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns:doc="http://www.lyncode.com/xoai" xsi:schemaLocation="http://www.dspace.org/xmlns/dspace/dim http://www.dspace.org/schema/dim.xsd">
   <dim:field mdschema="dc" element="contributor" qualifier="advisor">Shen, Chunhua</dim:field>
   <dim:field mdschema="dc" element="contributor" qualifier="advisor">Liu, Lingqiao</dim:field>
   <dim:field mdschema="dc" element="contributor" qualifier="advisor">van den Hengel, Anton John</dim:field>
   <dim:field mdschema="dc" element="contributor" qualifier="author">Qiao, Ruizhi</dim:field>
   <dim:field mdschema="dc" element="contributor" qualifier="school" lang="en">School of Computer Science</dim:field>
   <dim:field mdschema="dc" element="date" qualifier="issued">2017</dim:field>
   <dim:field mdschema="dc" element="identifier" qualifier="uri">http://hdl.handle.net/2440/113587</dim:field>
   <dim:field mdschema="dc" element="description" qualifier="abstract" lang="en">Compared with low-level features, mid-level representations of visual objects contain&#xd;
more discriminative and interpretable information and are beneficial for improving&#xd;
performance of classification and sharing learned information across object&#xd;
categories. These benefits draw tremendous attention of the computer vision communities&#xd;
and lots of breakthroughs have been made for various computer vision&#xd;
tasks with mid-level representations. In this thesis, we focus on the following problems&#xd;
regarding mid-level representations: 1) How to extract discriminative mid-level&#xd;
representations from local features? 2) How to suppress noisy components from&#xd;
mid-level representations? 3) And how to address the issue of visual-semantic discrepancy&#xd;
in mid-level representations? We deal with the first problem in the task of&#xd;
action recognition and the other two problems in the task of zero-shot learning.&#xd;
For the first problem, we devise a representation suitable for characterising human&#xd;
actions on the basis of a sequence of pose estimates generated by an RGB-D&#xd;
sensor. We show that discriminate sequence of poses typically occur over a short&#xd;
time window, and thus we propose a simple-but-effective local descriptor called a&#xd;
trajectorylet to capture the static and kinematic information within this interval. We&#xd;
also show that state of the art recognition results can be achieved by encoding each&#xd;
trajectorylet using a discriminative trajectorylet detector set which is selected from a&#xd;
large number of candidate detectors trained through exemplar-SVMs. The mid-level&#xd;
representation is obtained by pooling trajectorylet encodings.&#xd;
For the second problem, we follow the attractive research topic zero-shot learning&#xd;
and focus on classifying a visual concept merely from its associated online textual&#xd;
source, such as a Wikipedia article. We go further to consider one important factor:&#xd;
the textual representation as a mid-level representation is usually too noisy for the&#xd;
zero-shot learning tasks. We design a simple yet effective zero-shot learning method&#xd;
that is capable of suppressing noise in the text. Specifically, we propose an l₂‚₁-norm&#xd;
based objective function which can simultaneously suppress the noisy signal in the&#xd;
text and learn a function to match the text document and visual features. We also&#xd;
develop an optimization algorithm to efficiently solve the resulting problem.&#xd;
For the third problem, we observe that distributed word embeddings, which become&#xd;
a popular mid-level representation for zero-shot learning due to their easy&#xd;
accessibility, are designed to reflect semantic similarity rather than visual similarity&#xd;
and thus using them in zero-shot learning often leads to inferior performance To overcome this visual-semantic discrepancy, we here re-align the distributed word&#xd;
embedding with visual information by learning a neural network to map it into a&#xd;
new representation called the visually aligned word embedding (VAWE). We further&#xd;
design an objective function to encourage the neighbourhood structure of VAWEs to&#xd;
mirror that in the visual domain. This strategy gives more freedom in learning the&#xd;
mapping function and allows the learned mapping function to generalize to zeroshot&#xd;
learning methods and different visual features.</dim:field>
   <dim:field mdschema="dc" element="description" qualifier="dissertation" lang="en">Thesis (Ph.D.) -- University of Adelaide, School of Computer Science, 2018</dim:field>
   <dim:field mdschema="dc" element="subject" lang="en">Image classification</dim:field>
   <dim:field mdschema="dc" element="subject" lang="en">action recognition</dim:field>
   <dim:field mdschema="dc" element="subject" lang="en">zero-shot learning</dim:field>
   <dim:field mdschema="dc" element="subject" lang="en">attribute</dim:field>
   <dim:field mdschema="dc" element="subject" lang="en">word embedding</dim:field>
   <dim:field mdschema="dc" element="title" lang="en">Mid-level representations for action recognition and zero-shot learning</dim:field>
   <dim:field mdschema="dc" element="type" lang="en">Theses</dim:field>
   <dim:field mdschema="dc" element="provenance" lang="en">This electronic version is made publicly available by the University of Adelaide in accordance with its open access policy for student theses. Copyright in this thesis remains with the author. This thesis may incorporate third party material which has been used by the author pursuant to Fair Dealing exceptions. If you are the owner of any included third party copyright material you wish to be removed from this electronic version, please complete the take down form located at http://www.adelaide.edu.au/legals</dim:field>open.access</dim:dim></metadata></record></GetRecord></OAI-PMH>