Sunday, July 22, 2018

Our Goal and Challenge : Data Lake with AI Implementation


Continuing on setting our goals, as i mentioned in my previous post combining artificial intelligence and blockchain are undiscovered area. We still don't really understand how to outcome with a good business application. What I mean by good application is, to find most effective and efficient way to use both technologies in the same environment. As far as, we know that both systems are data-driven in their own way.  

Looking forward, we are pretty much set with our goal for future. Our challenge is to design and develop a Data Lake. Here is the reason that we wanted to take this challenge; when we look at components in our development environment our foundation architecturally similar to a data-lake.
While developing this massive system here are the challenges that we are expecting to see and with today's data warehouses can impede usage and prevent users from maximizing their analytics;

Timeliness. Introducing new content to the enterprise data warehouse can be a time-consuming and cumbersome process. When users need immediate access to data, even short processing delays can be frustrating and cause users to bypass the proper processes in favor of getting the data quickly themselves. Users also may waste valuable time and resources to pull the data from operational systems, store and manage it themselves, and then analyze it.

Flexibility. Users not only lack on-demand access to any data they may need at any time, but also the ability to use the tools of their choice to analyze the data and derive critical insights. Additionally, current data warehousing solutions often store one type of data, while today’s users need to be able to analyze and aggregate data across many different formats.

Quality. Users may view the current data warehouse with suspicion. If where the data originated and how it has been acted on are unclear, users may not trust the data. Also, if users worry that the data in the data warehouse is missing or inaccurate, they may circumvent the warehouse in favor of getting the data themselves directly from other internal or external sources, potentially leading to multiple, conflicting instances of the same data.

Findability. With many current data warehousing solutions, users do not have a function to rapidly and easily search for and find the data they need when they need it. Inability to find data also limits the users’ ability to leverage and build on existing data analyses.

Advanced analytics users require a data storage solution based on an IT “push” model (not driven by specific analytics projects). Unlike existing solutions, which are specific to one or a small family of use cases, what is needed is a storage solution that enables multiple, varied use cases across the enterprise.

This new solution needs to support multiple reporting tools in a self-serve capacity, to allow rapid ingestion of new datasets without extensive modeling, and to scale large datasets while delivering performance. It should support advanced analytics, like machine learning and text analytics, and allow users to cleanse and process the data iteratively and to track lineage of data for compliance. Users should be able to easily search and explore structured, unstructured, internal, and external data from multiple sources in one secure place.


Traditionally Data Lake is a data-centered architecture featuring a repository capable of storing vast quantities of data in various formats; however, in our case we are using Blockchain therefore the information become decentralized. Data from webserver logs, data bases, social media, and third-party data is ingested into the Data Lake, in our challenge we are streaming data from stock markets such as NYSE, Nasdaq S&P and Down Jones. Curation takes place through capturing metadata and lineage and making it available in the data catalog (Datapedia). Security policies, including entitlements, also are applied.

Data can flow into the Data Lake by either batch processing or real-time processing of streaming data. Additionally, data itself is no longer restrained by initial schema decisions, and can be exploited more freely by the enterprise. Rising above this repository is a set of capabilities that allow IT to provide Data and Analytics as a Service (DAaaS), in a supply-demand model. IT takes the role of the data provider (supplier), while business users (data scientists, business analysts) are consumers.

The DAaaS model enables users to self-serve their data and analytic needs. Users browse the lake’s data catalog (a Datapedia) to find and select the available data and fill a metaphorical “shopping cart” (effectively an analytics sandbox) with data to work with. Once access is provisioned, users can use the analytics tools of their choice to develop models and gain insights. Subsequently, users can publish analytical models or push refined or transformed data back into the Data Lake to share with the larger community.

Although provisioning an analytic sandbox is a primary use, the Data Lake also has other applications. For example, the Data Lake can also be used to ingest raw data, curate the data, and apply ETL. This data can then be loaded to an Enterprise Data Warehouse. To take advantage of the flexibility provided by the Data Lake, organizations need to customize and configure the Data Lake to their specific requirements and domains.

The Data Lake can be an effective data management solution for advanced analytics experts and business users alike. A Data Lake allows users to analyze a large variety and volume when and how they want. Following a Data and Analytics as a Service (DAaaS) model provides users with on-demand, self-serve data.

However, to be successful, a Data Lake needs to leverage a multitude of products while being tailored to the industry and providing users with extensive, scalable customization. Knowledgent’s Informationists provide the blend of technical expertise and business acumen to help organizations design and implement their perfect Data Lake.






Sadik Erisen









Saturday, July 14, 2018

New Objectives with AI and integration of Blockchain Technology

Artificial intelligence and Blockchain technology are hottest two topics in computer science field nowadays. However they are both different components of computer science field and they both sub-serve for different purposes in other word to say they are both solving different problems. Although, there are many problems that we might know or not, these two technologies can help us to solve it. 

At this point, we need to look for list of problems and solutions and study whether both can be applied into same architecture or not. Therefore problem has to be clearly understood and re-structured. At most asking questions is profoundly important,  such as what do know? what do we need? what are the sub-problems ? what type of methods that need to apply to get what we want?

These pre-steps would help us to observe and comprehend the problem, hence we can create our own concept to approach to problematic. Later it would be easier to place each opponent (AI and Blockchain) into the architecture.  

In fact finding a problem can be profound but  knowing where to apply both technologies can facilitate our problematic; for example, among the parties blockchain would create an efficient, secure and reliable bridge and AI would focus on specific routine of tasks such as detecting and analyzing the problem. 

Combining of artificial intelligence and blockchain is still a largely undiscovered area. Even though the convergence of the two technologies has recieved its fair share of scholarly attention, many projects devoted to this groundbreaking combination are still scarce.  
Putting the two technologies together has the potential to use data in ways never before thought possible.

Data is the key ingredient for the development and enhancement of AI algorithms, and blockchain secures this data, allows us to audit all intermediary steps AI takes to draw conclusions from the data, and allows individuals to monetize their produced data.

AI can be incredibly revolutionary, but it must be designed with utmost precautions – blockchain can greatly assist in this. How the interplay between the two technologies will progress is anyone’s guess. However, its potential for true disruption is clearly there, and rapidly developing.


Sadik Erisen

Simplifying and Redefining Blockchain

Week 7-8 (7/14/2018)

The purpose of this article is to briefly explaining and re-defining the objective of the blockchain technology. However, i also think that explaining the origin of the blockchain technology would help us to gain a deeper understanding.

Most of the computer scientists and software engineers know that a system should be efficient, reliable and secure as well as cost-effective for conducting and recording transactions.

Throughout history, instruments of trust, such as minted coins, paper money, letters of credit, and banking systems, have emerged to facilitate the exchange of value and protect both parties.
Important innovations, including telephone lines, credit card systems, the Internet, and mobile technologies have improved the convenience, speed, and efficiency of transactions while shrinking and sometimes virtually eliminating the distance between buyers and sellers. Still, many business transactions remain inefficient, expensive, and vulnerable, suffering from the following limitations;
  • Cash is useful only in local transactions and in relatively small amounts. 
  • The time between transaction and settlement can be long. 
  • Duplication of effort and the need for third-party validation and/or the presence of intermediaries add to the inefficiencies.
  • Fraud, cyberattacks, and even simple mistakes add to the cost and complexity of doing business, and they expose all participants in the network to risk if a central system, such as a bank, is compromised. 
  • Credit card organizations have essentially created walled gardens with a high price of entry. Merchants must pay the high costs of on-boarding, which often involves considerable paperwork and a time-consuming vetting process. 
  • Half of the people in the world don’t have access to a bank account and have had to develop parallel payment systems to conduct transactions.
Transaction volumes worldwide are growing exponentially and will surely magnify the complexities, vulnerabilities, inefficiencies, and costs of current transaction systems. The growth of e-commerce, online banking, and in-app purchases, and the increasing mobility of people around the world have fueled the growth of transaction volumes. And transaction volumes will explode with the rise of Internet of Things (IoT) — autonomous objects, such as refrigerators that buy groceries when supplies are running low and cars that deliver themselves to your door, stopping for fuel along the way. To address these challenges and others, the world needs payment networks that are fast and that provide a mechanism that establishes trust, requires no specialized equipment, has no chargebacks or monthly fees, and provides a collective bookkeeping solution for ensuring transparency and trust.

Going back to explanation, blockchain is a shared, distributed ledger that facilitates the process of recording transactions and tracking assets in a business network. An asset can be tangible — a house, a car, cash, land  — or intangible like intellectual property, such as patents, copyrights, or branding. Virtually anything of value can be tracked and traded on a blockchain network, reducing risk and cutting costs for all involved. 

Basically, blockchain, most simply defined as a shared, immutable ledger, has the potential to be the technology that redefines those processes and many others. Blockchain allows increased trust and efficiency in the exchange of almost anything.

Although, there are some down sides of this technology such as transactions are irreversible, hacks and manipulation still occur and more transactions happen, the system generates more nodes and this would create time delay during the decrypting the chains.

As a team, we are still looking forward to figure out whether this technology is perfect fit for our system.

Sadik Erisen

Tuesday, June 26, 2018

Week 6 (6/25/2018)

Monday 6/25

Sadik briefed me and Rahi on the updates made on Grace through Gitlab. With that, I updated the raspberry pi and "git pulled" ($ git pull origin master) to merge the Gitlab repository with the pi's repository. I then looked into the dataset in Gitlab and researched into it.

Tuesday 6/26

Today was spent getting more familiar with Python, mainly focusing with the basics of Input/Output. I also updated the data flow chart for Grace. Tomorrow, we are planing to be able to extract words specifically from their part of speech. 

Saturday, June 23, 2018

Conversations

Statement and Response Relationship:

Grace stores information from the conversations as statements. Each user's input (statement) get tagged with any number (please see the : NLTK Tagging methods) and mapped to possible responses.
Therefore each statement object has a reference, which links the user's input to a number of other input statements.

The response object has counter which looks for the terms of frequency (please see the : TF-IDF Algorithm you can also visit grace search engine TF-IDF implementation/ Ranking section ) this attribute indicates the number of times that statement has been given to response. This makes it possible for the bot to determine if a particular  response is more commonly used than another.

Sadik Erisen