-->

Welcome to our Coding with python Page!!! hier you find various code with PHP, Python, AI, Cyber, etc ... Electricity, Energy, Nuclear Power

Showing posts with label ArtificialIntelligence. Show all posts
Showing posts with label ArtificialIntelligence. Show all posts

Friday, 15 October 2021

EVOLUTION OF MODELOPS: NOW A MORE ADVANCED ARTIFICIAL INTELLIGENCE

All about Agile, Ansible, DevOps, Docker, EXIN, Git, ICT, Jenkins, Kubernetes, Puppet, Selenium, Python, etc

ModelOps is a set of automated practices and tools that help deploy, manage, monitor, and improve models in production. The approach is designed to be model-centric, which means everything is instrumented around the model, from deployment to governance to inference and monitoring to scale.

Far and wide, investment in artificial intelligence and machine learning are drastically increasing and new data science projects are underway to build predictive and analytical models for various purposes. However, while companies plan to scale up sophisticated Artificial Intelligence solutions in a reasonable time, the harsh reality is that the adoption of these solutions is often stalled because companies generally focus more on development than on the operationalization of the models. On that note, ModelOps comes to the rescue bringing advancements in AI.

 

ModelOps Tools

Since the ModelOps approach brings all the players together, several emerging start-ups, as well as enterprise companies, offer ModelOps solutions to orchestrate these components collectively in an end-to-end fully automated model life cycle. Let us have a look at the figure below showing how by managing a platform enterprises can govern and scale any AI initiatives.

Powerful platforms like ModelOp center typically integrate with development platforms, IT systems, and enterprise applications so that businesses can leverage and extend ongoing investments in AI and IT. In this way, data scientists can work at scale using the tools they know best.

 

ModelOps-

  • is primarily focused on the governance and life cycle management of AI and decision models (including machine learning, knowledge graphs, rules, optimization, linguistic and agent-based models). Core capabilities include the management of model development environments, model repository, champion-challenger testing, model rollout/rollback, and CI/CD (continuous implementation/continuous delivery) integration
  • enables the retuning, retraining, or rebuilding of AI models, providing an uninterrupted flow between the development, operationalization, and maintenance of models within AI-based systems
  • provides business domain experts autonomy to assess the quality (interpret the outcomes and validate KPIs) of AI models in production and facilitates the ability to promote or demote AI models for inferencing without a full dependency on data scientists or ML engineers.

 

A More Advanced AI

AI answers Distress and Help-calls

Emergency relief services are flooded with distress and help calls in the event of any emergency. Managing such a huge number of calls is time-consuming and expensive when done manually. The chances of critical information being lost or unobserved are also a possibility. In such cases, AI can work as a 24/7 dispatcher. AI systems and voice assistants can analyze massive amounts of calls, determine what type of incident occurred and verify the location. They can not only interact with callers naturally and process those calls, but can also instantly transcribe and translate languages. AI systems can analyze the tone of voice for urgency, filtering redundant or less urgent calls and prioritizing them based on the emergency.

 

Predictive Analytics for Proactive Disaster Management

Machine learning and other data science approaches are not limited to assisting the on-ground relief teams or assisting only after the actual emergency. Machine learning approaches such as predictive analytics can also analyze past events to identify and extract patterns and populations vulnerable to natural calamities. A large number of supervised and unsupervised learning approaches are used to identify at-risk areas and improve predictions of future events. For instance, clustering algorithms can classify disaster data based on severity. They can identify and segregate climatic patterns which may cause local storms with the cloud conditions which may lead to a widespread cyclone.

Predictive machine learning models can also help officials distribute supplies to where people are going, rather than where they were by analyzing real-time behavior and movement of people.

In addition, predictive analytics techniques can also provide insight for understanding the economic and human impact of natural calamities. Artificial neural networks take in information such as region, country, and natural disaster type to predict the potential monetary impact of natural disasters.

Recent advances in cloud technologies and numerous open-source tools have enabled predictive analytics with almost no initial infrastructure investment. So agencies with limited resources can also build systems based on data science and develop more sophisticated models to analyze disasters.

As with every progressing technology, AI will also build on its existing capabilities. It has the potential to eliminate outages before they are detected and give disaster response leaders an informed, clearer picture of the disaster area, ultimately saving lives.



3 steps businesses can take to reduce bias in AI systems

Experts have proposed the prioritization of humans faced with technological advancement by working on three key areas.


  • Artificial intelligence constitutes one of the most impactful developments for businesses and organizations in general.
  • However, this fast-paced and unstoppable trend raises ethical issues.
  • It can be challenging to ensure that AI development is fair when the algorithms at its core are designed with racist, sexist, or other biases which are often unconscious.
  • Below, Lorena Blasco-Arcas and Hsin-Hsuan Meg Lee propose a human-centred view for the design of specific frameworks and regulatory systems.

“Okay, Google, what’s the weather today?” “Sorry, I don’t understand.”

Does the experience—interacting with smart machines that don’t respond to orders—sound familiar? This failure may leave people feeling dumbfounded, as if their intelligence were not on the same wavelength as the machines’. While this is not the intention of AI development (to interact selectively), such incidents are likely more frequent for “minorities” in the tech world.

The global artificial intelligence (AI) software market is forecast to boom in the coming years, reaching around 126 billion US dollars by 2025. The success of AI technology is forcing many existing companies to transform their business model and shift to AI. However, along with the advance, there is an increasing worry about the biases in the algorithm development of all these tools.

How AI flaws become apparent

Algorithmic bias is nothing new. However, to date, engineers have focused more on developing AI algorithms to solve complex problems than on monitoring and reporting the potential issues these technological advances bring. We have already seen examples of failures by technology with the rise of discriminatory practices. For instance, in 2016, Microsoft released its self-learning chatbot Tay on Twitter. It was supposed to be an experiment in “conversational understanding.” The AI tool could learn language fundamentals, and, over time, it could participate in a conversation by itself. However, the bot ended up developing racist and sexist traits on social media. Another example occurred at MIT when while working on facial recognition, Joy Buolamwini conducted a discriminatory experiment without knowing it. As a dark-skinned woman, she was not recognised by the AI as precisely as her white friend. The results completely missed the point of the experience, and she found out that 99% of white women were identified by the computer, compared to 65% of black women.

Was human intention behind the AI’s behaviour? Maybe, maybe not. These examples do not mean that the AI tools were fundamentally flawed or designed to be racist. Nevertheless, their design was biased, and they were not controlled enough before going public. Data biases can lead to discriminatory practices stemming from human intention or an unintended act, perpetuating generations’ bias(es). Worse still, because the resulting discrimination is almost always an unintentional emergent property of the algorithm’s use rather than a conscious choice by its programmers, it is tough to identify the source of the problem or explain it to a court. Machines tend to give the – false – impression that they are neutral.

How to develop an ethical and non-biased AI application in an undoubtedly biased and unbalanced society? Can AI be the holy grail by developing more balanced societies that overcome traditional inequality and exclusion? It is too early to say, and it seems apparent that we will witness many trial-and-error phases before achieving a consensus on what and how AI might be used ethically in our societies. Much like institutional racism, which requires fundamental shifts in the overall ecosystem, the problems in AI development also call for a similar change to create better output. To solve this issue, we propose to prioritise humans faced with technological advancement by working on three areas:

1. Unbiasing (biased) human beings

Behind the development and implementation of algorithms, there are developers and specific people in power positions. As seen in the data, the developer’s professional world is far from being diverse today, which explains some of the thinking logics that foster biases. Increasing the diversity of and access to developer positions in the big companies that dominate the industry would offer a more critical perspective on how algorithms are developed. This would increase human inclusion rather than the opposite. Suppose we understand algorithmic bias as imposing specific ideas using computers and math as an alibi. In that case, we can question the institutional logic behind the perpetuation of bias and discriminatory practices.

There is a need to increase control, monitoring systems, regulation and common ethical frameworks to ensure that human bias does not permeate the creation and development of algorithms. We echo the view of professors Ayanna Howard and Charles Isbell at Georgia Tech that recognising the importance of diversity in terms of data and leadership, and demanding accountability in certain decisions are essential guiding principles toward achieving a more just development and implementation of AI in the future.

2. Data for good instead of data for bias

Vital initiatives are developing that might help solve historical dataset biases, such as the one carried out by a researcher at the University of Ontario, who used the MNIST dataset and distilled that database of 60K images down to only 5 to train an AI model. Should these procedures be successfully applied to different contexts, they will make AI more accessible to companies that may not afford massive databases. It will also improve data privacy and data collection, as less information from individuals will be required to train relevant models.

3. Educating citizens in the advantages and risks of AI applications

AI development poses diverse and notable challenges concerning understanding societies, politics, business and even our daily lives as citizens As AI becomes increasingly present in business processes affecting individuals’ choices and possibilities, more education is needed to raise awareness and understanding of these topics.

The technology readiness of citizens will improve AI adoption and positively impact the critical assessment of AI implementation and its effects. A more aware citizen will be less tolerant of manipulation and acceptance of biased or unfair applications of AI tech, such as those related to surveillance that might conflict with civil liberties and rights.

Minimising bias in AI is essential to building trust.
Image: McKinsey&Company

Making machines more human, or even suppressing human intelligence, has often been treated as one of the ultimate goals of technological advancement. Human-centred technology development implies that the developers and companies using the machines should not only aim for innovation but also pay attention to their potential impact on society. Humans are flawed, meaning our society is naturally full of biases that are systematic and institutional, and we are not always aware of them. But we should avoid replicating the same issues in the machines we build.



Monday, 11 October 2021

Top Python Packages for Data Science and How to Best Use Them

Finding the top python packages and libraries that aren't only popular, but get the job done isn't easy. Here's a list to help you out.

Out of all the Python scientific libraries and packages available, which ones are not only popular but the most useful in getting the job done?

To help you filter down a list of libraries and packages worth adding to your data science toolbox, we have compiled our top picks for aspiring and practicing data scientists. But you’ll also want to know how to best use these tools for tricky, real-world data problems. So instead of leaving you with yet another top choice list among a quintillion lists, we explain how to make the most of these libraries using real-world examples.

You can learn more about how these packages fit into data science with Data Science Dojo's introduction to Python course.

Data Manipulation

pandas

There’s a reason why pandas consistently tops published ranks on data science related libraries in Python. The library can help you with a variety of tasks, but it is particularly useful for data manipulation or data wrangling. It can save you a lot of leg work in not only your typical rudimentary data manipulation tasks, but in handling some pretty tricky problems you might encounter when slicing and filtering.

Multi-indexed data can be one of these tricky tasks. The library pandas takes care of advanced indexing, including multi-indexing, where you might need to work with higher-dimensional data or multiple index levels. For example, number of user interactions might be indexed by 1) product category, 2) time of day user interacted with the product, and 3) location of the user. Instead of your typical table of rows and columns to represent the data, you might find it better to organize the number of user interactions into all cases that fall under x product category, with y time of day, and z location. This way you can easily see user interactions across each condition of product category, time of day, and user location. This saves you from having to apply a filter or group for all combinations of conditions in your traditional row-and-table structure.

Here is one way to multi-index data in pandas. With less than a few lines of code, pandas makes this easy to implement in Python:

import pandas as pd

data_multi_indx = table_data.set_index(['Product', 'Day of Week'])

print(data_multi_indx)
'''
Output:
Location Num User Interactions
Product Day of Week
Product 1 Morning A 3
Morning B 90
Morning C 7
Afternoon A 17
Afternoon B 1
Afternoon C 82
Product 2 Morning A 27
Morning B 70
Morning C 3
Afternoon A 1
Afternoon B 1
Afternoon C 98
Product 3 Morning A 94
Morning B 5
Morning C 1
Afternoon A 0
Afternoon B 7
Afternoon C 93
'''

For the more rudimentary data manipulation tasks, pandas doesn’t require much effort on your part. You can simply use the functions available for imputing missing values, one-hot encoding, dropping columns and rows, and so on.

Here are a few example classes and functions in pandas that make rudimentary data manipulation easy in a few lines of code, at most.

For more lessons with Pandas, visit Data Independent.

FeatureDescription
fillna(value)Fill in missing values on a column or the whole data frame with a value such as the mean, median, or mode.
isna(data)/isnull(data)Check for missing values.
get_dummies(data_frame['Column'])Apply one-hot encoding on a column.
to_numeric(data_frame['Column'])Convert a column of values from strings to numeric values.
to_string(data_frame['Column'])Convert a column of values from numeric values to strings.
to_datetime(data_frame['Column'])Convert a column of datetimes in string format to standard datetime format.
drop(columns=['Column0','Column1'])Drop specific columns or useless columns in your data frame.
drop(data.frame.index[[rownum0,rownum1]])Drop specific rows or useless rows in your data frame.


numpy

Another library that keeps topping the ranks is numpy. This library can handle many tasks, but it is particularly useful when working with multi-dimensional arrays and performing calculations on these arrays. This can be tricky to do in more conventional ways, where you need to find the index of a value or certain values inside another index, with multiple indices.

This is where numpy shows its strength. Its array() function means standard arrays can be simply added and nicely bundled into a multi-dimensional array. Calculations on these arrays can also be easily implemented using numpy’s vast array (pun intended) of mathematical functions.

Let’s picture an example where numpy’s multi-dimensional arrays are useful. A company tracks or records if a user was/was not shown a mobile product in the morning, afternoon, and night, delivered through a mobile notification. Based on the level of user interaction with the shown product, the company also records a user engagement score. Data points on each user’s shown product and engagement score are stored inside an array; each array stores these values for each user. The company would like to quickly and simply bundle all user arrays.

In addition to this, using engagement score and purchase history, the company would like to calculate and identify the minimum distance (or difference) across all users’ data points so that users who follow a similar pattern can be categorized and targeted accordingly.

numpy’s array() makes it easy to bundle user arrays into a multi-dimensional array and argmin() and linalg.norm() find the min Euclidean distance between users, as an example of the kinds of calculations that can be done on a multi-dimensional array:

import numpy as np

# Records tracking whether user was/was not shown product during
# morning, afternoon, and night, and user engagement score
user_0 = [0,0,1,0.7]
user_1 = [0,1,0,0.4]
user_2 = [1,0,0,0.0]
user_3 = [0,0,1,0.9]
user_4 = [0,1,0,0.3]
user_5 = [1,0,0,0.0]

# Create a multi-dimensional array to bundle all users

# Can use arrays with mixed data types by specifying
# the object data type in numpy multi-dimensional arrays
users_multi_dim = np.array([user_0,user_1,user_2,user_3,user_4,user_5],dtype=object)

print(users_multi_dim)
'''
Output:
[[0 0 1 0.7]
[0 1 0 0.4]
[1 0 0 0.0]
[0 0 1 0.9]
[0 1 0 0.3]
[1 0 0 0.0]]
'''

# To view which user was/was not shown the product
# either morning, afternoon or night, pandas easily
# allows you to index and label the data
row_names = [_ for _ in ['User 0','User 1','User 2','User 3','User 4','User 5']]
col_names = [_ for _ in ['Product Shown Morning','Product Shown Afternoon',
'Product Shown Night','User Engagement Score']]

users_df_indexed = pd.DataFrame(users_multi_dim,index=row_names,columns=col_names)

print(users_df_indexed)
'''
Output:
Product Shown Morning Product Shown Afternoon Product Shown Night User Engagement Score
User 0 0 0 1 0.7
User 1 0 1 0 0.4
User 2 1 0 0 0
User 3 0 0 1 0.9
User 4 0 1 0 0.3
User 5 1 0 0 0
'''

# Find which existing user is closest to the engagement
# and purchase behavior of a new user by calculating the
# min Euclidean distance on a numpy multi-dimensional array

user_0 = [0.7,51.90,2]
user_1 = [0.4,25.95,1]
user_2 = [0.0,0.00,0]
user_3 = [0.9,77.85,3]
user_4 = [0.3,25.95,1]
user_5 = [0.0,0.00,0]

users_multi_dim = np.array([user_0,user_1,user_2,user_3,user_4,user_5])
new_user = np.array([0.8,77.85,3])

closest_to_new = np.argmin(np.linalg.norm(users_multi_dim-new_user,axis=1))

print('User', closest_to_new, 'is closest to the new user')
'''
Output:
User 3 is closest to the new user
'''

Data Modeling

statsmodels

The main strength of statsmodels is its focus on statistics, going beyond the ‘machine learning out-of-the-box’ approach. This makes it a popular choice for data scientists. Conducting statistical tests to find significantly different variables, checking for normality in your data, checking the standard errors, and so on, cannot be underestimated when trying to build the most effective model you can build. Your model is only as good as your inputs, and statsmodels is designed to help you better understand and customize your inputs.

The library also covers an exhaustive list of predictive models to choose from, depending on your predictors and outcome variable(s). It covers your classic Linear Regression models (including ordinary least squares, weighted least squares, recursive least squares, and more), Generalized Linear models, Linear Mixed Effects models, Binomial and Poisson Bayesian models, Logit and Probit models, Time Series models (including autoregressive integrated moving average, dynamic factor, unobserved component,and more), Hidden Markov models, Principal Components and other techniques for Multivariate models, Kernel Density estimators, and lots more.

Here are the classes and functions in statsmodels that cover the main modeling techniques useful for many prediction tasks.

FeatureDescription
sm.OLS()sm.WLS()sm.GLS()sm.RecursiveLS()Linear Regression models: ordinary least squares, weight least squares, generalized least squares, recursive least squares
sm.GLM()Generalized Linear models. You can specify the model family – e.g. sm.families.Gamma()
sm.mixedlm()Linear Mixed Effects models
BinomialBayesMixedGLM()PoissonBayesMixedGLM()Binomial and Poisson Bayesian models
sm.Logit()sm.Probit()Logit and Probit discrete models
sm.tsa.ARMA()sm.tsa.ARIMA()sm.tsa.SARIMAX()sm.tsa.VARMAX()sm.tsa.SimpleExpSmoothing()sm.tsa.UnobservedComponents()sm.tsa.DynamicFactor()Time series models: autoregressive moving average, autoregressive integrated moving average, seasonal autoregressive integrated moving average, vector autoregression moving average, simple exponential smoothing, unobserved components, dynamic
factor
sm.tsa.MarkovRegression()sm.tsa.MarkovAutoregression()Univariate and multivariate kernel density estimators for nonparametric stats

scikit-learn

Any library that makes machine learning more accessible and easier to implement is bound to make the top choice list among aspiring and practicing data scientists. The library scikit-learn not only allows models to be easily implemented out-of-the-box but also offers some auto fine tuning.

Finding the best possible combination of model parameters is a key example of fine tuning. The library offers a few good ways to search for the optimal set of parameters, given the algorithm and problem to solve. The grid search and random search algorithms in scikit-learn evaluate different combinations of parameters until they find the best combo that results in the best outcome, or a better performing model. The grid search goes through every possible combination, whereas the random search randomly samples the parameters over a fixed number of times/iterations. Cross validating your model on many subsets of data is also easy to implement using scikit-learn. With this kind of automation, the library offers data scientists a massive time saver when building models.

The library also covers all the essential machine learning models from classification (including Support Vector Machine, Random Forest, etc), to regression (including Ridge Regression, Lasso Regression, etc), and clustering (including k-Means, Mean Shift, etc).

Here are the classes and functions in scikit-learn that cover the main modeling techniques useful for many prediction tasks.

FeatureDescription
SVC()GaussianNB()LogisticRegression()DecisionTreeClassifier()RandomForestClassifier()SGDClassifier()MLPClassifier()Classification models: Support Vector Machine, Gaussian Naïve Bayes, Logistic Regression, Decision Tree, Random Forest, Stochastic Gradient Descent, Multi-Layer Perceptron
linear_model.Ridge()linear_model.Lasso()SVR()DecisionTreeRegressor()RandomForestRegressor()SGDRegressorMLPRegressor()Regression models: Ridge Regression, Lasso Regression, Support Vector Machine, Decision Tree, Random Forest, Stochastic Gradient Descent, Multi-Layer Perceptron
KMeans()AffinityPropagation()MeanShift()AgglomerativeClusteringClustering models: k-Means, Affinity Propagation, Mean Shift, Agglomerative Hierarchical Clustering

Data Visualization

plotly

The libraries matplotlib and seaborn will easily take care of your basic static plot functions, which are important for your own internal exploration or understanding of the data. But when presenting visual insights to business folks or users, interactivity is where we are headed these days.

Using JavaScript functionality, plotly renders interactive graphs in the form of zooming in and panning out of the graph panel, hovering over objects for more information, and dragging objects into position to further explore relationships in the data. Graphs can be customized to your heart’s content.

Here are just a few of many tricks that plotly offers:

FeatureDescription
hovermodehoverinfoControls the mode and text when a user hovers over an object.
on_selection()on_click()Allows a user to select or click on an object and have that selected object change color, for example.
updateModifies a graph’s layout and data such as titles and annotations.
animateCreates an animated graph.

bokeh

Much like plotlybokeh also offers interactive graphs. But one feature that stands out in bokeh is linked interactions. This is useful when keeping separate graphs in unison, where the user interacts with one graph and needs to compare with the other while they are in sync. For example, a user zooms into a graph, effectively changing the range of the graph, and then would like to compare with the second graph. The second graph would need to automatically update its range so that both graphs can be easily compared like-for-like.

Here are some key tricks that bokeh offers:

FeatureDescription
figure()Creates a new plot and allows linking to the range of another plot.
HoverTool()hover_glyphAllows user to hover over an object for more information.
selection_glyphSelects a particular glyph object for styling.
Slider()Creates a slider to dynamically update the plot based on the slide range.

Rank

seo