RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    Spatialtemporal Local Directional Patterns for Facial Expression Recognition

    한글로보기

    https://www.riss.kr/link?id=T13307938

    • 저자
    • 발행사항

      용인 : 경희대학교, 2013

    • 학위논문사항

      학위논문(박사) -- 경희대학교 대학원 , 컴퓨터공학과 , 2013. 8

    • 발행연도

      2013

    • 작성언어

      영어

    • DDC

      004 판사항(20)

    • 발행국(도시)

      경기도

    • 형태사항

      171p. : 삽도 ; 26cm

    • 일반주기명

      경희대학교 논문은 저작권에 의해 보호받습니다.
      지도교수:Oksam Chae
      참고문헌 : p.150-168

    • 소장기관
      • 경희대학교 국제캠퍼스 도서관 소장기관정보
      • 경희대학교 중앙도서관 소장기관정보
      • 국립중앙도서관 국립중앙도서관 우편복사 서비스
    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
    • 오류접수
    인용문이 복사되었습니다.

    부가정보

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    The number of smart devices, computers and other personal electronic devices that we own and interact with is increasing everyday. Moreover, we need smarter devices that understand human behavior to facilitate our interaction with them. That is, as we are surrounded by more devices in our daily lives, those devices should be aware of their environment and their users, and should assess the emotional state and responses of the human user to adjust their answers and actions. Given that facial expressions and corporal language are the most emotion-related signals that humans emit, it is natural to incorporate human expression recognition capabilities into smart devices. We have evolved to understand and interact through corporal language, such as facial expressions, to the point of inferring another person's emotional state based on these cues. However, few advances have been made towards a model of human emotions understandable by a computer. Thus, the first step in that direction is the recognition and classification of human facial expressions as basic constituents to infer more complex human emotional states. Therefore, we need robust automatic face analysis algorithms to be included in our smart devices.

    To automatically recognize facial expressions we need a robust description of each expression---or in general, a robust face or image descriptor. Furthermore, that description should be general enough to accommodate the different ways in which each expression can be performed, while maintaining the discrimination among the different expressions. For example, there are many facial configurations that resemble a smile, and, even more, people may smile in different ways; nevertheless, the descriptor should be similar for all these cases. Simultaneously, when we extract that descriptor from anger expressions, it should be different from the smile one. Moreover, there are other challenges that a robust descriptor should overcome. For example, for daily-use devices, such as smart phones, the environment in which the pictures are captured varies immensely, e.g., we may have non-constant illumination, noise, rotation and background changes, among others. Thus, the challenge of creating an image descriptor is to overcome the changes in environmental conditions as well as inter- and intra-class variations.

    Thus, in this thesis we develop and analyze several robust facial descriptors that are discriminative between classes, yet enclose the intra-class variations due to imaging conditions, appearance changes, noise, and other factors. To achieve these descriptors, we develop the directional number as a building block for micro-pattern code schemes that encapsulate local information of the images into single codes. In a nutshell, a directional number models the prominent directions of the micro patterns, i.e., the structure of the textures in the images, both static and dynamic. Furthermore, we analyze and explore the use of the directional numbers to encode different patterns to solve the face analysis problem, and combine them with other types of information, such as color, to increase the discrimination of our codes. Additionally, we show that our directional numbers have several advantages over existing appearance-based coding schemes. First, we overcome the limitation of the common bit-string marking representation, used by previous methods, by using the directional numbers to generate codes that correlate the similarity between the patterns being coded with the generated codes, i.e., similar patterns generate codes that are close in the code space. This proximity in the code space eases the code processing for higher level representation. Second, the use of the directional numbers produces more flexible code schemes in comparison to previous codes, as we can mix and extract other information to and from the codes easily.

    Furthermore, we extend the concept of the directional number to the spatiotemporal domain, and explore its use to represent dynamic micro-patterns that appear in video sequences of facial expressions. Thus, by using temporal information we can increase the accuracy of the expression recognition problem. In this thesis, we propose several combinations of the spatiotemporal information to produce robust code schemes that exploit the structure and motion of the micro-patterns. Additionally, we propose modeling techniques that incorporate the spatiotemporal information in early coding stages, which prove to be more robust in comparison to mixing this information later in the descriptor as current methods do. Moreover, we develop and analyze a novel spatiotemporal face descriptor based on the spatiotemporal directional numbers that takes advantage of the temporal information of the face videos. Additionally, we extend and explore the use of the proposed spatiotemporal directional numbers to describe and classify dynamic textures. Given that the facial expressions can be thought as a dynamic texture, we can easily extend and apply our proposed descriptors to a more general classification problem. Our experiments support the generalization possibilities of our proposed spatiotemporal code schemes and descriptors.

    We tested all the proposed algorithms using several public available databases and protocols. Finally, we classify the different image descriptors using well-known support vector machines which utilize one-versus-one technique. Hence, we validate and demonstrate the superiority of our proposed techniques based on directional numbers, which can be reliably applied to different problems, such as facial expressions recognition, both with static and dynamic data, dynamic texture recognition, and we also show the potential of some codes to be applied to face recognition.
    번역하기

    The number of smart devices, computers and other personal electronic devices that we own and interact with is increasing everyday. Moreover, we need smarter devices that understand human behavior to facilitate our interaction with them. That is, as we...

    The number of smart devices, computers and other personal electronic devices that we own and interact with is increasing everyday. Moreover, we need smarter devices that understand human behavior to facilitate our interaction with them. That is, as we are surrounded by more devices in our daily lives, those devices should be aware of their environment and their users, and should assess the emotional state and responses of the human user to adjust their answers and actions. Given that facial expressions and corporal language are the most emotion-related signals that humans emit, it is natural to incorporate human expression recognition capabilities into smart devices. We have evolved to understand and interact through corporal language, such as facial expressions, to the point of inferring another person's emotional state based on these cues. However, few advances have been made towards a model of human emotions understandable by a computer. Thus, the first step in that direction is the recognition and classification of human facial expressions as basic constituents to infer more complex human emotional states. Therefore, we need robust automatic face analysis algorithms to be included in our smart devices.

    To automatically recognize facial expressions we need a robust description of each expression---or in general, a robust face or image descriptor. Furthermore, that description should be general enough to accommodate the different ways in which each expression can be performed, while maintaining the discrimination among the different expressions. For example, there are many facial configurations that resemble a smile, and, even more, people may smile in different ways; nevertheless, the descriptor should be similar for all these cases. Simultaneously, when we extract that descriptor from anger expressions, it should be different from the smile one. Moreover, there are other challenges that a robust descriptor should overcome. For example, for daily-use devices, such as smart phones, the environment in which the pictures are captured varies immensely, e.g., we may have non-constant illumination, noise, rotation and background changes, among others. Thus, the challenge of creating an image descriptor is to overcome the changes in environmental conditions as well as inter- and intra-class variations.

    Thus, in this thesis we develop and analyze several robust facial descriptors that are discriminative between classes, yet enclose the intra-class variations due to imaging conditions, appearance changes, noise, and other factors. To achieve these descriptors, we develop the directional number as a building block for micro-pattern code schemes that encapsulate local information of the images into single codes. In a nutshell, a directional number models the prominent directions of the micro patterns, i.e., the structure of the textures in the images, both static and dynamic. Furthermore, we analyze and explore the use of the directional numbers to encode different patterns to solve the face analysis problem, and combine them with other types of information, such as color, to increase the discrimination of our codes. Additionally, we show that our directional numbers have several advantages over existing appearance-based coding schemes. First, we overcome the limitation of the common bit-string marking representation, used by previous methods, by using the directional numbers to generate codes that correlate the similarity between the patterns being coded with the generated codes, i.e., similar patterns generate codes that are close in the code space. This proximity in the code space eases the code processing for higher level representation. Second, the use of the directional numbers produces more flexible code schemes in comparison to previous codes, as we can mix and extract other information to and from the codes easily.

    Furthermore, we extend the concept of the directional number to the spatiotemporal domain, and explore its use to represent dynamic micro-patterns that appear in video sequences of facial expressions. Thus, by using temporal information we can increase the accuracy of the expression recognition problem. In this thesis, we propose several combinations of the spatiotemporal information to produce robust code schemes that exploit the structure and motion of the micro-patterns. Additionally, we propose modeling techniques that incorporate the spatiotemporal information in early coding stages, which prove to be more robust in comparison to mixing this information later in the descriptor as current methods do. Moreover, we develop and analyze a novel spatiotemporal face descriptor based on the spatiotemporal directional numbers that takes advantage of the temporal information of the face videos. Additionally, we extend and explore the use of the proposed spatiotemporal directional numbers to describe and classify dynamic textures. Given that the facial expressions can be thought as a dynamic texture, we can easily extend and apply our proposed descriptors to a more general classification problem. Our experiments support the generalization possibilities of our proposed spatiotemporal code schemes and descriptors.

    We tested all the proposed algorithms using several public available databases and protocols. Finally, we classify the different image descriptors using well-known support vector machines which utilize one-versus-one technique. Hence, we validate and demonstrate the superiority of our proposed techniques based on directional numbers, which can be reliably applied to different problems, such as facial expressions recognition, both with static and dynamic data, dynamic texture recognition, and we also show the potential of some codes to be applied to face recognition.

    더보기

    목차 (Table of Contents)

    • I. INTRODUCTION 1
    • II. BACKGROUND STUDY 8
    • III. DIRECTIONAL NUMBERS 30
    • IV. FACIAL ANALYSIS 64
    • V. CONCLUSION 144
    • I. INTRODUCTION 1
    • II. BACKGROUND STUDY 8
    • III. DIRECTIONAL NUMBERS 30
    • IV. FACIAL ANALYSIS 64
    • V. CONCLUSION 144
    • BIBLIOGRAPHY 150
    • LIST OF PUBLICATIONS 169
    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼