San Francisco Bay Area
1K followers 500+ connections

Join to view profile

About

Research Interests: Machine Learning, Data Analytics, Cyber Security and Blockchain…

Articles by De

Activity

1K followers

See all activities

Experience & Education

  • Uniphore

View De’s full experience

See their title, tenure and more.

or

By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.

Licenses & Certifications

Publications

  • REX: Rapid Ensemble Classification System for Landslide Detection using Social Media

    Distributed Computing Systems (ICDCS), 2017 IEEE 37th International Conference on

    We study the problem of using Social Media to detect natural disasters, of which we are interested in a special kind, namely landslides. Employing information from Social Media presents unique research challenges, as there exists a considerable amount of noise due to multiple meanings of the search keywords, such as "landslide" and "mudslide". To tackle these challenges, we propose REX, a rapid ensemble classification system which can filter out noisy information by implementing two key ideas:…

    We study the problem of using Social Media to detect natural disasters, of which we are interested in a special kind, namely landslides. Employing information from Social Media presents unique research challenges, as there exists a considerable amount of noise due to multiple meanings of the search keywords, such as "landslide" and "mudslide". To tackle these challenges, we propose REX, a rapid ensemble classification system which can filter out noisy information by implementing two key ideas: (I) a new method for constructing independent classifiers that can be used for rapid ensemble classification of Social Media texts, where each classifier is built using randomized Explicit Semantic Analysis; and (II) a self-correction approach which takes advantage of the observation that the majority label assigned to Social Media texts belonging to a large event is highly accurate. We perform experiments using real data from Twitter over 1.5 years to show that REX classification achieves 0.98 in F-measure, which outperforms the standard Bag-of-Words algorithm by an average of 0.14 and the state-of-the-art Word2Vec algorithm by 0.04. We also release the annotated datasets used in the experiments as a contribution to the research community containing 282k labeled items.

    See publication
  • Information Diffusion Analysis of Rumor Dynamics over a Social-Interaction Based Model

    Collaboration and Internet Computing (CIC), 2016 IEEE 2nd International Conference on

    Rumors may potentially cause undesirable effect such as the widespread panic in the general public. Especially, with the unprecedented growth of different types of social and enterprise networks, rumors could reach a larger audience than before. Many researchers have proposed different approaches to analyze and detect rumors in social networks. However, most of them either study on theoretical models without real data experiments or use content-based analysis and limited information diffusion…

    Rumors may potentially cause undesirable effect such as the widespread panic in the general public. Especially, with the unprecedented growth of different types of social and enterprise networks, rumors could reach a larger audience than before. Many researchers have proposed different approaches to analyze and detect rumors in social networks. However, most of them either study on theoretical models without real data experiments or use content-based analysis and limited information diffusion analysis without fully considering social interactions. In this paper, we propose a social interaction based model FAST by taking four major properties of social interactions into account including familiarity, activeness, similarity, and trustworthiness. Also, we evaluate our model on real data from Sina Weibo (Twitter-like social network in China), which contains around 200 million tweets and 14 million Weibo users. Based on our model, we create a new metrics Fractional Directed Power Community Index (FD-PCI) derived from μ-PCI to identify the influential spreaders in social networks. FD-PCI shows better performance than conventional metrics such as K-core index and PageRank. Moreover, we obtain interesting influential features to detect rumors by the comparison between rumor and real news dynamics.

    See publication
  • Multi-hazard detection by integrating social media and physical sensors

    Social Media for Government Services, Springer

    Disaster Management is one of the most important functions of the government. FEMA and CDC are two examples of government agencies directly charged with handling disasters, whereas USGS is a scientific agency oriented towards disaster research. But regardless of the type or purpose, each of the mentioned agencies utilizes Social Media as part of its activities. One of the uses of Social Media is in detection of disasters, such as earthquakes. But disasters may lead to other kinds of disasters…

    Disaster Management is one of the most important functions of the government. FEMA and CDC are two examples of government agencies directly charged with handling disasters, whereas USGS is a scientific agency oriented towards disaster research. But regardless of the type or purpose, each of the mentioned agencies utilizes Social Media as part of its activities. One of the uses of Social Media is in detection of disasters, such as earthquakes. But disasters may lead to other kinds of disasters, forming multi-hazards such as landslides. Effective detection and management of multi-hazards cannot rely only on one information source. In this chapter, we describe and evaluate a prototype implementation of a landslide detection system LITMUS, which combines multiple physical sensors and Social Media to handle the inherent varied origins and composition of multi-hazards. Our results demonstrate that LITMUS detects more landslides than the ones reported by an authoritative source.

    See publication
  • LITMUS: a Multi-Service Composition System for Landslide Detection

    IEEE Transactions on Services Computing

    Landslides are an illustrative example of multi-hazards, which can be caused by earthquakes, rainfalls and human activity among other reasons. Detection of landslides presents a significant challenge, since there are no physical sensors that would detect landslides directly. A more recent approach in detection of natural hazards, such as earthquakes, involves the use of social media. We propose a multi-service composition approach and describe LITMUS, which is a landslide detection service that…

    Landslides are an illustrative example of multi-hazards, which can be caused by earthquakes, rainfalls and human activity among other reasons. Detection of landslides presents a significant challenge, since there are no physical sensors that would detect landslides directly. A more recent approach in detection of natural hazards, such as earthquakes, involves the use of social media. We propose a multi-service composition approach and describe LITMUS, which is a landslide detection service that combines data from both physical and social information services by filtering and then joining the information flow from those services based on their spatiotemporal features. Our results show that with such approach LITMUS detects 25 out of 27 landslides reported by USGS in December 2013 and 40 more landslide locations unreported by USGS during this period. LITMUS is a prototype tool that is used to investigate and implement research ideas in the area of disaster detection. We list some of the current work being done on refining the system that allows us to identify 137 landslide locations unreported by USGS during a more recent period of September 2014. Finally, we describe a live demonstration that displays landslide detection results on a web map in real-time.

    See publication
  • SPADE: A Social-spam Analytics and Detection Framework

    Social Network Analysis and Mining, Springer

    Social media such as Facebook, MySpace, and Twitter have become increasingly important for attracting millions of users. Consequently, spammers are increasing using such networks for propagating spam. Although existing filtering techniques such as collaborative filters and behavioral analysis filters are able to significantly reduce spam, each social network needs to build its own independent spam filter and support a spam team to keep spam prevention techniques current. To alleviate those…

    Social media such as Facebook, MySpace, and Twitter have become increasingly important for attracting millions of users. Consequently, spammers are increasing using such networks for propagating spam. Although existing filtering techniques such as collaborative filters and behavioral analysis filters are able to significantly reduce spam, each social network needs to build its own independent spam filter and support a spam team to keep spam prevention techniques current. To alleviate those problems, we propose a framework for spam analytics and detection which can be used across all social network sites. Specifically, the proposed framework SPADE has numerous benefits including (1) new spam detected on one social network can quickly be identified across social networks; (2) accuracy of spam detection will be improved through cross-domain classification and associative classification; (3) other techniques (such as blacklists and message shingling) can be integrated and centralized; (4) new social networks can plug into the system easily, preventing spam at an early stage. In SPADE, we present a uniform schema model to allow cross-social network integration. In this paper, we define the user, message, and web page model. Moreover, we provide an experimental study of real datasets from social networks to demonstrate the flexibility and feasibility of our framework. We extensively evaluated two major classification approaches in SPADE: cross-domain classification and associative classification. In cross-domain classification, SPADE achieved over 0.92 F-measure and over 91 % detection accuracy on web page model using Naïve Bayes classifier. In associative classification, SPADE also achieved 0.89 F-measure on message model and 0.87 F-measure on user profile model, respectively. Both detection accuracies are beyond 85 %. Based on those results, our SPADE has been demonstrated to be a competitive spam detection solution to social media.

    See publication
  • A Perspective of Evolution after Five Years: A Large-scale Study of Web Spam Evolution

    International Journal of Cooperative Information Systems (IJCIS)

    Identifying and detecting web spam is an ongoing battle between spam-researchers and spammers which has been going on since search engines allowed searching of web pages to the modern sharing of web links via social networks. A common challenge faced by spam-researchers is the fact that new techniques depend on requiring a corpus of legitimate and spam web pages. Although large corpora of legitimate web pages are available to researchers, the same cannot be said about web spam or spam web…

    Identifying and detecting web spam is an ongoing battle between spam-researchers and spammers which has been going on since search engines allowed searching of web pages to the modern sharing of web links via social networks. A common challenge faced by spam-researchers is the fact that new techniques depend on requiring a corpus of legitimate and spam web pages. Although large corpora of legitimate web pages are available to researchers, the same cannot be said about web spam or spam web pages. In this paper, we introduce the Webb Spam Corpus 2011 — a corpus of approximately 330,000 spam web pages — which we make available to researchers in the fight against spam. By having a standard corpus available, researchers can collaborate better on developing and reporting results of spam filtering techniques. The corpus contains web pages crawled from links found in over 6.3 million spam emails. We analyze multiple aspects of this corpus including redirection, HTTP headers, web page content, and classification evaluation. We also provide insights into changes in web spam since the last Webb Spam Corpus was released in 2006. These insights include: (1) spammers manipulate social media in spreading spam; (2) HTTP headers and content also change over time; (3) spammers have evolved and adopted new techniques to avoid the detection based on HTTP header information.


    Read More: https://www.worldscientific.com/doi/abs/10.1142/S0218843014410019

    See publication
  • Analysis and Detection of Low Quality Information in Social Networks

    Proc. of Ph.D. Symposium at 30th IEEE International Conference on Data Engineering (ICDE 2014)

    With social networks like Facebook, Twitter and Google+ attracting audiences of millions of users, they have been an important communication platform in daily life. This in turn attracts malicious users to the social networks as well, causing an increase in the incidence of low quality information. Low quality information such as spam and rumors is a nuisance to people and hinders them from consuming information that is pertinent to them or that they are looking for. Although individual social…

    With social networks like Facebook, Twitter and Google+ attracting audiences of millions of users, they have been an important communication platform in daily life. This in turn attracts malicious users to the social networks as well, causing an increase in the incidence of low quality information. Low quality information such as spam and rumors is a nuisance to people and hinders them from consuming information that is pertinent to them or that they are looking for. Although individual social networks are capable of filtering a significant amount of low quality information they receive, they usually require large amounts of resources (e.g, personnel) and incur a delay before detecting new types of low quality information. Also the evolution of various low quality information posts lots of challenges to defensive techniques. My PhD thesis work focuses on the analysis and detection of low quality information in social networks. We introduce social spam analytics and detection framework SPADE across multiple social networks showing the efficiency and flexibility of cross-domain classification and associative classification. For evolutionary study of low quality information, we present the results on large-scale study on Web spam and email spam over a long period of time. Furthermore, we provide activity-based detection approaches to filter out low quality information in social networks: click traffic analysis of short URL spam, behavior analysis of URL spam and information diffusion analysis of rumor. Our framework and detection techniques show promising results in analyzing and detecting low quality information in social networks.

    See publication
  • Click Traffic Analysis of Short URL Spam on Twitter

    Proc. of 9th IEEE International Conference on Collaborative Computing: Networking, Applications and Worksharing (CollaborateCom 2013)

    With an average of 80% length reduction, the URL shorteners have become the norm for sharing URLs on Twitter, mainly due to the 140-character limit per message. Unfortunately, spammers have also adopted the URL shorteners to camouflage and improve the user click-through of their spam URLs. In this paper, we measure the misuse of the short URLs and analyze the characteristics of the spam and non-spam short URLs. We utilize these measurements to enable the detection of spam short URLs. To achieve…

    With an average of 80% length reduction, the URL shorteners have become the norm for sharing URLs on Twitter, mainly due to the 140-character limit per message. Unfortunately, spammers have also adopted the URL shorteners to camouflage and improve the user click-through of their spam URLs. In this paper, we measure the misuse of the short URLs and analyze the characteristics of the spam and non-spam short URLs. We utilize these measurements to enable the detection of spam short URLs. To achieve this, we collected short URLs from Twitter and retrieved their click traffic data from Bitly, a popular URL shortening system. We first investigate the creators of over 600,000 Bitly short URLs to characterize short URL spammers. We then analyze the click traffic generated from various countries and referrers, and determine the top click sources for spam and non-spam short URLs. Our results show that the majority of the clicks are from direct sources and that the spammers utilize popular websites to attract more attention by cross-posting the links. We then use the click traffic data to classify the short
    URLs into spam vs. non-spam and compare the performance of the selected classifiers on the dataset. We determine that the Random Tree algorithm achieves the best performance with an accuracy of 90.81% and an F-measure value of 0.913.

    See publication
  • Evolutionary Study of Web Spam: Webb Spam Corpus 2011 versus Webb Spam Corpus 2006

    Proc. of 8th IEEE International Conference on Collaborative Computing: Networking, Applications and Worksharing (CollaborateCom 2012)

    With over 2:5 hours a day spent browsing websites online [1] and with over a billion pages [2], identifying and detecting web spam is an important problem. Although large corpora of legitimate web pages are available to researchers, the same cannot be said about web spam or spam web pages. We introduce the Webb Spam Corpus 2011 — a corpus of approximately 330; 000 spam web pages — which we make available to researchers in the fight against spam. By having a standard corpus available, researchers…

    With over 2:5 hours a day spent browsing websites online [1] and with over a billion pages [2], identifying and detecting web spam is an important problem. Although large corpora of legitimate web pages are available to researchers, the same cannot be said about web spam or spam web pages. We introduce the Webb Spam Corpus 2011 — a corpus of approximately 330; 000 spam web pages — which we make available to researchers in the fight against spam. By having a standard corpus available, researchers can collaborate better on developing and reporting results of spam filtering techniques. The corpus contains web pages crawled from links found in over 6:3 million spam emails. We analyze multiple aspects of this corpus including redirection, HTTP headers and web page content. We also provide insights into changes in web spam since the last Webb Spam Corpus was released in 2006. These insights include: 1) spammers manipulate social media in spreading spam; 2) HTTP headers also change over time (e.g. hosting IP addresses of web spam appear in more IP ranges); 3) Web spam content has evolved but the majority of content is still scam.

    Other authors
    See publication
  • A Social-Spam Detection Framework

    8th Annual Collaboration, Electronic messaging, Anti-Abuse and Spam Conference (CEAS 2011)

    Social networks such as Facebook, MySpace, and Twitter have become increasingly important for reaching millions of users. Consequently, spammers are increasing using such networks for propagating spam. Existing filtering techniques such as collaborative filters and behavioral analysis filters are able to significantly reduce spam, each social network needs to build its own independent spam filter and support a spam team to keep spam prevention techniques current. We propose a framework for spam…

    Social networks such as Facebook, MySpace, and Twitter have become increasingly important for reaching millions of users. Consequently, spammers are increasing using such networks for propagating spam. Existing filtering techniques such as collaborative filters and behavioral analysis filters are able to significantly reduce spam, each social network needs to build its own independent spam filter and support a spam team to keep spam prevention techniques current. We propose a framework for spam detection which can be used across all social network sites. There are numerous benefits of the framework including: 1) new spam detected on one social network, can quickly be identified across social networks; 2) accuracy of spam detection will improve with a large amount of data from across social networks; 3) other techniques (such as blacklists and message shingling) can be integrated and centralized; 4) new social networks can plug into the system easily, preventing spam at an early stage. We provide an experimental study of real datasets from social networks to demonstrate the flexibility and feasibility of our framework.

    Other authors
    See publication
Join now to see all publications

Courses

  • Advanced Operating Systems

    CS 6210

  • Computability, Algorithms, and Complexity

    CS 6505

  • Database System Concept& Design

    CS 6400

  • Internet Arch& Protocols

    CS 7260

  • Introduction to Graduate Study

    CS 7001

  • Legal Issues in Technology Transfer

    MGT 6799

  • Machine Learning (Prof. Andrew Ng. Standford University online course on Coursera.com website)

    -

  • Managing Resources of the Technological Firm

    MGT 6772

  • Network Security

    CS 6262

  • Principles of Management for Engineers

    MGT 6753

  • Real-time Systems

    CS 6235

  • Social Network Analysis (Prof. Lada Adamic. University of Michigan online course on Coursera.com website)

    -

Honors & Awards

  • Best Paper Award

    CollaborateCom 2013

    We received the best paper award for our paper "Is Email Spam Business Dying?: A Study on Evolution of Email Spam over Fifteen Years" on CollaborateCom 2013.

  • Student Travel Grant

    CollaborateCom 2012

  • Jinan Education Award

    Jinan University

    2 Master students per year
    Top 0.1%

  • The Second Place Team of IBM “Elite” Solution Design Contest 2009

    IBM SCT Guangzhou Region

  • Nanyue Excellent Graduate Student

    Department of Education of Guangdong Province

    Top 1%

  • The Second Prize of the 5th China Graduate Mathematical Contest in Modeling

    5th China Graduate Mathematical Contest in Modeling (CGMCM)

    Top 10%-30%

  • Excellent Graduate

    Jinan University

    Top 5%

  • Meritorious winner

    MCM Competition Organizing Committee

    Top 2%-11%

  • Honorable Mention (International Level) and the Third Prize (National Level) of ACM Asia Programming Contest Chengdu Site2005

    ACM/ICPC

  • The 34th Place (International Level) and the Third Prize (National Level) of ACM Asia Programming Contest Beijing Site2005

    ACM/ICPC

Languages

  • Chinese

    Native or bilingual proficiency

  • English

    Professional working proficiency

View De’s full profile

  • See who you know in common
  • Get introduced
  • Contact De directly
Join to view full profile

Other similar profiles

Explore top content on LinkedIn

Find curated posts and insights for relevant topics all in one place.

View top content

Add new skills with these courses