The digital revolution has rapidly changed the way we live. A huge amount of digital data is created from every aspect of our daily lives due to the development of the IT technology, advent of the Internet, and proliferation of the mobile environment ...
The digital revolution has rapidly changed the way we live. A huge amount of digital data is created from every aspect of our daily lives due to the development of the IT technology, advent of the Internet, and proliferation of the mobile environment that is represented by the smart-phone. The large quantity of digital data created is called “big data” and much interest is given to its use.
Although data accumulation is easier than ever, the question of changing the data into useful information through manipulation, processing, and analysis remains a difficult one to solve. In particular, the need to identify and make use of one’s opinions, feelings, interests, preferences, expectations, and complaints is rising because informal text data on the Internet, and via mobile communication and SNS, explosively reflects them directly and indirectly. However, detailed research results are not yet satisfactory due to the difficulty in analyzing and identifying what the data means. This is due to the diversity of analysis techniques that is organized according to the data source, type, contents, and various characteristics of the language. It is also a fact that technology and the ideas to make the most of this data are not yet advanced enough.
The news we watch every day is one example of informal data. Tens of thousands of news reports are created and digitalized every day, then distributed throughout the world. In that sense, news is the typical information text that makes up “big data” because its output is so enormous. The massive amount of news produced affects politics, economy, society, and is closely related to the stock price, which is the core index of economic activities. As a result, it is believed that the news has a close relationship with the stock price, and one can expect to find an investment opportunity and make a profit. Ultimately, we can predict the stock price and expect to create economic benefits if we can distinguish between favorable and unfavorable factors by properly analyzing the news.
Therefore, many studies over the years have proved that the news is closely related to the stock price, and that the stock price fluctuates as a result of the news. Studies also have been conducted to forecast the stock price using the news based on this relationship between the news and stock price. However, past research targeted a particular kind of news or the news in a particular area. As a result, immediate reflection and analysis was insufficient as in the real world a massive amount of news is created and dispersed in real time every day. Also, recent research using machine learning has been limited by the fact that personal judgment acts in forecasting the stock price of a particular company or identifying the sentiment polarity of the glossary for news text analysis.
This study presented an intelligent investment decision-making model based on opinion mining that extract opinions about stock price index increase/decrease, by taking massive news data as the “big data” composed of the informal text, scrapping and parsing the news automatically, and tagging the emotional word. In addition, the subject-oriented sentiment dictionary for the stock domain, not the general purpose sentiment dictionary, was directly implemented and applied by the prototype experiment for sentiment polarity tagging to analyze the informal text data. Afterwards, two spots with similar media characteristics were selected for the experiment, and the accuracy of stock price index increase/decrease was compared through learning and verification data split test.
The data (July 2011 to September 2011) for the study(Kim, 2012) using the general purpose sentiment dictionary was used for mutual comparison when conducting the prototype experiment that built up and verified the subject-oriented sentiment direction of the stock domain. The experiment results confirmed the importance of the subject-oriented sentiment dictionary, which identifies the news opinion index by demonstrating high forecast efficiency compared to the general purpose sentiment dictionary in the particular threshold value section when attempting to forecast the stock price without data split. In addition, a prediction capability similar to the general purpose sentiment direction was demonstrated in the verification data set for the learning/verification data split experiment that was attempted to resolve the over-fitting problem. Therefore, it was found that the accuracy of stock price increase/decrease could be improved sufficiently if the optimized condition, opinion, and threshold value can be extracted from the learning data.
The experiment was conducted for 80,000 news items about two companies, M and H, which were posted on the Stock section of the Naver portal site from January 2011 to December 2011. Company M, a relatively new online media company specialist, claims that they are more specialized in the stock exchange area, whereas Company H is a leading media company in the economy that was established with the motto of a “global comprehensive economic media company.” These two media companies have contrasting media characteristics. Consequently, the results of the experiment are different to some extent. In conclusion, Company M’s news opinion has shown a higher level of stock price prediction as well as in the quality of predictability in both the learning and verification data sets.
On the other hand, it was found that analyzing two media company’s news as a single data set produced a poorer result than individual analysis because two lots of opinions are mixed. In the end, there was some difference in forecast accuracy and quality in the two media outlets, but the price index of stocks could be estimated using news opinion. It was also found that the opinion index and optimal threshold value learned through opinion mining can be useful in forecasting the actual increase/decrease of the stock price inde