-->

Welcome to our Coding with python Page!!! hier you find various code with PHP, Python, AI, Cyber, etc ... Electricity, Energy, Nuclear Power

Showing posts with label #DevOps. Show all posts
Showing posts with label #DevOps. Show all posts

Tuesday, 15 February 2022

Download Python Crash Course: A Hands-On, Project-Based Introduction to Programming, 2nd Edition

 Python Crash Course, 2nd Edition: A Hands-On, Project-Based Introduction to Programming

Python Crash Course, 2nd Edition: A Hands-On, Project-Based Introduction to Programming


Link - https://fr.b-ok.africa/dl/5416472/bee132

#Python, #100DaysOfCode, #CodeNewbies, #WomenWhoCode, #DevOps, #code, #Coding, #LearnToCode, #DataAnalytics, #DataScience, #MachineLearning, #AI, #programming, 

Monday, 25 October 2021

Using Data To Make Better Decisions

All about Agile, Ansible, DevOps, Docker, EXIN, Git, ICT, Jenkins, Kubernetes, Puppet, Selenium, Python, etc


How many tennis balls could you fit in all of the skyscrapers in New York City? How many gummy bears could you fit in an airplane? How long would it take to fill the Mariana Trench with peanut butter? 

Like most people, you’d probably need some time and maybe a whiteboard to guess these answers. Not only are the questions difficult to imagine and rationalize, but also there’s a lot of relevant background information needed to be able to make an accurate guess.

Without collecting all relevant data, any answer is a guess.

Imagine watching any sports game where you could only see one team and no scoreboard. You would be able to guess how each team is doing from positioning and reactions, but it would be difficult to say confidently which team is winning or who is more likely to win the entire game. It would be even more difficult to win bets against someone else watching normally on TV being spoon-fed a plethora of statistics and hearing the opinions of professional commentators. 

Despite the unfairness of this competition, this is how many people invest in stocks. Dark Pools can consist of more than half of the overall trade volume for a stock, and yet despite this huge imbalance, many people don’t know they exist let alone monitor their activities. Gaining access and utilizing all the information available before making critical decisions can level the playing field and facilitate educated decisions.
Data can help us make sense of big numbers and simplify complex ideas. Pluto is roughly 3 billion miles away, so how long would it take to walk there? About a billion hours, which is longer than watching all the content on all major streaming sites back to back 20,000 times. While the number and the analogy mean the same thing, one is substantially easier to understand intuitively than the other. 

Data explanation and visualization is a crucial component of understanding complex data points.


Financial data is infamously one of the largest data sets in the world and endlessly complicated, so seeing it in easy-to-process graphs instead of raw metrics helps elucidate. It’s difficult to visualize two numbers of orders of magnitude apart, but it’s easy to see how big one circle is against another.
Likewise, it’s almost impossible to rationalize numbers without context for what they mean. Financial data is rife with jargon and acronyms that require a dictionary to read, let alone understand. Being able to process this information is knowing not just what the words mean, but also how it affects a company and how it compares to other similar companies. 

One of the best ways to use data is not as a standalone item in a complex sheet but as a living, breathing, dynamic guide for making better decisions. Data in isolation is hard to understand intuitively and even worse to try to act upon. Combined with visualizations, it can form a complete picture for faster understanding and superior intuitive answers.

Not all data is created equal.


This paradigm is far from original, with many of the most prominent data sources using citations and best practices to try to eliminate the uncertainty. One of the largest crypto data providers revealed that 65%-95% of all their data was inaccurate and untrustworthy. In the wake of the LIBOR scandal, cracks in the financial system were exposed and the underbelly of market manipulation was revealed. Recently, payment for order flow was popularized, selling people’s trades to big institutions and allowing them to take profit away from investors. Far from novel, these are just several examples of data being difficult to trust. 

The commonality from all three of these is a fundamental agency problem; each of these groups stood to gain financially from manipulating their data or failing to correct incorrect data. The solution is to find data sources that are fundamentally incentivized to provide accurate data. Most brokers provide access to financial data, but often there is a conflict of interest, such as a broker selling their own stock, or fees on trading either explicit or invisible through slippage. 

For this reason, it’s essential to carefully evaluate the trustworthiness of data sources and not take information at face value, especially when critical decisions are being made from it. Healthy skepticism and asking if there is a conflict of interest can expose early on whether a data set is purely analytical or might be inaccurate and skewed.

Data is one of the most valuable resources in the world.


Almost half of the top seven biggest companies in the world use data as their primary product for good reason. Making informed, objective decisions can eliminate uncertainty on correct choices and often guarantee the best possible outcomes. Learning to make these decisions off of comprehensive, well understood and trustworthy data can drastically increase the effectiveness of any decision and bring order to an otherwise chaotic world.

Wednesday, 20 October 2021

ADVANTAGES OF HIGH-QUALITY PERIMETER SECURITY SYSTEMS

All about Agile, Ansible, DevOps, Docker, EXIN, Git, ICT, Jenkins, Kubernetes, Puppet, Selenium, Python, etc

When browsing the market for security systems which work to protect the entire perimeter or your business premises – look no further than perimeter security systems.

With a range of types to choose from, including security bollards, high security gates, vehicle barriers, you can be sure your premises will be protected from unauthorised visitors and thieves. This will ensure your business is safe from theft and intrusion or damage from trespassers. These perimeter security systems are often favoured by large companies with many access points to their sites. Hörmann, one of Europe’s leading manufacturers of door solutions also offering perimeter security systems, have identified the advantages of perimeter security systems. Highlighting why you should invest in high-quality solutions to protect your site. 

A Wide Choice of Security Solutions to Choose From 

Security_Solutions.jpeg

When searching for your perfect perimeter security system, you will notice just how much choice you have! Which is great, considering no two premises are the same. 

Below are just some solutions you can choose from: 

  • High security barriers 
  • Security bollards 
  • CCTV solutions 
  • Mobile vehicle barriers 
  • Automated sliding gates

Have Control Over Who Enters Your Premises 

The main advantage of installing perimeter security systems is the ability it provides you to control who enters your premises. Giving you the choice to permit access to only those who have been authorised and preventing unknown visitors from gaining access without approval. It is beneficial to cover the entire perimeter of your premises to ensure every area is covered. You could do this using a variety of types of perimeter security, determining your main entrance and smaller side entrances. In most cases, companies will choose to prioritise their main entrances with the highest levels of security. Followed by smaller systems in place for side entrances where footfall and traffic levels are low. 

Protect your Premises, Property and Staff 

This may seem obvious, but many people underestimate the potential for theft on their premises. Particularly if they believe their location to be safe. Despite where you are located as a company and if you have or have not been targeted by thieves before, perimeter security systems are paramount in providing protection for your premises, property, and staff. Investing in these systems in the best way to prevent theft, which could be both costly and detrimental to your company’s operations. 

These measures will also ensure your staff and visitors feel safe on site, providing peace of mind. This is particularly important if you have staff on site overnight when your premises are much more likely to be targeted. 

Perimeter Security Systems Can Be Budget Friendly 

When considering security systems most people will assume this is going to take a large chunk out their budget. Think again! Perimeter security systems can be customised to meet both your needs and your budget. Working with manufacturers you can determine the features which are most important to you and keep within your budget. If you need a low-cost solution to perimeter security systems, many will opt for a simple barrier or manual sliding gates. 

Control the Flow of Moving Traffic 

Security_Parameters.jpeg

If you manage a busy site with large amounts of both traffic and footfall, security systems like bollards and barriers are great for controlling this flow of moving traffic. This can become useful in other areas, like public spaces for events, where large gatherings are predicted to take place. Installing bollards or turnstiles can help to prevent dangerous rushes amongst crowds. Working to move people quickly but orderly through access points. 

#BigData #Analytics #frontend #MachineLearning #CyberSecurity #Python #RStats #TensorFlow #JavaScript #CloudComputing #Serverless #Linux #database #DataScience #100DaysofCode #ML #data #streaming #devops #php #css #java #AI #digital #opensource

WHAT HAPPENS WHEN BLOCKCHAIN, IOT, AND AI CONVERGE?

All about Agile, Ansible, DevOps, Docker, EXIN, Git, ICT, Jenkins, Kubernetes, Puppet, Selenium, Python, etc

Blockchain, internet of things (IoT), and artificial intelligence (AI) are driving the digital transformation of the post pandemic era.

The inevitable convergence of blockchain, IoT, and AI can form an impactful combination of security, interconnectivity, and autonomy to revolutionize the way things are done.

Synergy’ is a word that refers to a collaboration of entities that creates an impact much greater than that possible by the individual entities on their own. The convergence of blockchain, IoT, and AI is just that - a combination of technologies that have the potential to redefine the way businesses, industries, and even economies function, way more than they are already doing. A few applications and concepts have already seen an overlap between these technologies, with promising results. An example of this is the combination of AI and blockchain to manage UAV air traffic, making mass autonomous flight safer. This application alone, upon materialization, will redefine quite a number of industries like aviation and logistics. Thus, it’s not hard to imagine the impact that the combination of blockchain, IoT, and AI can have on the world. To understand the implications of the coalition of these technologies, it is first important to understand what they individually bring to the table.

Blockchain, IoT and AI Assessed Individually

Blockchain_IoT_and_AI_Assessed_Individually.png

Blockchain, the foundational technology for the cryptocurrency Bitcoin, enables a large number of computers to collectively perform computational tasks together and store information in a decentralized, immutable, and universally accessible manner. Various organizations, consortia, as well as industries, are considering using blockchain as a platform for consensus to democratize the decision-making processes among partners. 

The internet of things (IoT) provides seamless interconnectivity among different everyday objects armed with sensors, microprocessors, and transducers to create a network that can independently sense and gather data, continually analyze it, and perform programmed tasks when the situation calls for it. 

Artificial intelligence (AI) gives full autonomy of data analysis, decision-making and action to computers or other smart devices. It can replicate and even exceed human computational and cognitive capabilities under certain conditions, businesses are using AI to automate mainly repetitive, routine processes that may require heavy amounts of data processing and quick decision-making based on pure logic.

Combining Blockchain, IoT, and AI

The convergence of blockchain, IoT, and AI can enable organizations to maximize the benefits of each of these technologies while minimizing the risks and limitations associated with them. As IoT networks comprise a myriad of connected devices, there are numerous vulnerabilities in the network, leaving it prone to hacker attacks, fraud, and data theft. To prevent security issues, AI powered by machine learning can proactively defend against malware and hacker attacks. The security of the network and the data can be further enhanced by blockchain, which can limit illicit access and modification of the data on the network. AI can also enhance the functional capability of the IoT network by making it smarter and more autonomous. 



A potential case that demonstrates the convergence of blockchain, IoT, and AI is Fujitsu’s algorithm to measure workers’ heat stress levels. The algorithm constantly monitors workers’ physiological data such as temperature, humidity, activity levels, pulse, etc. using IoT wearables and sensors to track the correlation between different factors with workers’ health. The analysis can help the company improve working conditions for employees, and prevent them from having severe work-induced health issues. The addition of blockchain to this system can help to keep track of more personalized data by ensuring privacy or can help disburse health insurance amounts using smart contracts. Although the expected impact of the convergence of blockchain, IoT, and AI is exciting to think about, existing applications of these technologies are far from perfect. While 61% of companies claim to have adopted AI, their applications are still in the early stages and nowhere near the levels of sophistication required to achieve real transformation.

The same can be said for IoT and blockchain. However, with increased interest, investment, and innovation, the convergence of blockchain, IoT, and AI will eventually be a reality.

#ArtificialIntelligence #technology #innovation #5G 

#Python #coding #BigData #coding #CyberSecurity #Future #devops #java #CloudComputing #django   #100DaysOfCode #bot #TensorFlow #COVID19 #RHOP #ITC

Machine Learning For Absolute Beginners: A Plain English Introduction

All about Agile, Ansible, DevOps, Docker, EXIN, Git, ICT, Jenkins, Kubernetes, Puppet, Selenium, Python, etc
machine-learning-absolute-beginners-introduction-2nd.pdf

Machine Learning For Absolute Beginners: A Plain English Introduction (Second Edition) (Machine Learning From Scratch Book 1) Kindle Edition

Featured by Tableau as the first of "7 Books About Machine Learning for Beginners."

Ready to spin up a virtual GPU instance and smash through petabytes of data? Want to add 'Machine Learning' to your LinkedIn profile?

Well, hold on there...

Before you embark on your epic journey, there are some high-level theory and statistical principles to weave through first.
But rather than spend $30-$50 USD on a dense long textbook, you may want to read this book first. As a clear and concise alternative to a textbook, this book provides a practical and high-level introduction to machine learning.

Machine Learning for Absolute Beginners Second Edition has been written and designed for absolute beginners. This means plain English explanations and no coding experience required. Where core algorithms are introduced, clear explanations and visual examples are added to make it easy and engaging to follow along at home.

New Updated Edition
This major new edition features many topics not covered in the First Edition, including Cross Validation, Ensemble Modeling, Grid Search, Feature Engineering, and One-hot Encoding. Please note that this book is not a sequel to the First Edition but rather a restructured and revamped version of the First Edition. Readers of the First Edition should not feel compelled to purchase this Second Edition.

Disclaimer: If you have passed the 'beginner' stage in your study of machine learning and are ready to tackle coding and deep learning, you would be well served with a long-format textbook. If, however, you are yet to reach that Lion King moment - as a fully grown Simba looking over the Pride Lands of Africa - then this is the book to gently hoist you up and offer you a clear lay of the land.

In This Step-By-Step Guide You Will Learn:

• How to download free datasets
• What tools and machine learning libraries you need
• Data scrubbing techniques, including one-hot encodingbinning and dealing with missing data
• Preparing data for analysis, including k-fold Validation
• Regression analysis to create trend lines
• Clustering, including k-means clustering, to find new relationships
• The basics of Neural Networks
• Bias/Variance to improve your machine learning model
• Decision Trees to decode classification
• How to build your first Machine Learning Model to predict house values using Python

Frequently Asked Questions

Q: Do I need programming experience to complete this e-book?
A: This e-book is designed for absolute beginners, so no programming experience is required. However, two of the later chapters introduce Python to demonstrate an actual machine learning model, so you will see programming language used in this book.

Q: I have already purchased the First Edition of Machine Learning for Absolute Beginners, should I purchase this Second Edition?
A: As many of the topics from the First Edition are covered in the Second Edition, you may be better served reading a more advanced title on machine learning.

Q: Does this book include everything I need to become a machine learning expert?
A: Unfortunately, no. This book is designed for readers taking their first steps in machine learning and further learning will be required beyond this book to master machine learning.

Please feel welcome to join this introductory course by downloading a copy, or sending a free sample to your chosen device.
#MachineLearning #ML #100DaysOfCode #CodeNewbies #WomenWhoCode #DEVCommunity #DevOps #code #Coding #LearnToCode #Python #Robotics #DataScience #AI #programming

Monday, 11 October 2021

What are the 7 Terrifying Use Cases of #ArtificialIntelligence

Many advances in artificial intelligence (AI) are fascinating, but some recent use cases are downright creepy.

AI is being increasingly used to make important decisions. Major tech companies are the primary ones driving AI advances, and their algorithms impact billions of people. Unfortunately, these companies have zero accountability. 

3_Types_of_AI.jpeg

Source: Great Learning

YouTube (owned by Google) is helping to radicalize people into white supremacy. Google allowed advertisers to target people who search racist phrases like “black people ruin neighborhoods” and Facebook allowed advertisers to target groups like “jew haters”. Amazon’s facial recognition technology misidentified 28 members of congress as criminals, yet it is already in use by police departments. The newsfeed/timeline/recommendation algorithms of all the major platforms tend to reward incendiary content, prioritizing it for users.

While it can be easy to focus on regulations that are misguided or ineffective, we often take for granted safety standards and regulations that have largely worked well. 

Here are 7 terrifying artificial intelligence use cases:

1. Advanced Military Robots

Military_Robots.jpg

Source: Army.ml

Military robots are autonomous robots designed for military applications from transport to search & rescue and attack. One of the scariest potential uses of AI and robotics is the development of a robot soldier.  Although many have moved to ban the use of so-called "killer robots," the fact that the technology could potentially power those types of robots soon is upsetting, to say the least. In an experiment conducted by the scientists of Intelligent Systems in Switzerland, robots were made to compete for a food source in a single area. The robots could communicate by emitting light and, after they found the food source, they began turning their lights off or using them to steer competitors away from the food source.

Robots_Army.jpeg

Source: 123RF

2. Artificial Intelligence Police

Police_2025.jpeg

Source: National Police Chiefs' Council

Artificial intelligence could change policing. Police in certain cities around the US are experimenting with an AI algorithm that predicts which citizens are most likely to commit a crime in the future. Hitachi announced a similar system back in 2015. Maybe the film Minority Report wasn't completely off base in its representation of the future?

3. Robot Doctors 

Robot_Doctor_BBC.jpeg

Source: BBC

One of the biggest industries that AI could potentially benefit is healthcare. AI is already in use in many fields of medicine, even helping doctors decide on treatment. But, what if that AI system misses a critical aspect of your medical history or makes the wrong recommendation? Cases of AI being discriminatory in healthcare have been frequently documented in the past. Currently, AI-based suggestions are not considered to be superior to an experienced physician’s calls. However, it is very likely that, at some point in the future, AI will be the standard-setter in terms of human bodily diagnostics, treatment suggestions, and implementation. 

4. Schizophrenic AI Robots

Schrizophrennic_Robot.jpeg

Source: Daily Bits

Researchers at the University of Texas at Austin and Yale University used a neural network called DISCERN to teach the system certain stories. To simulate an excess of dopamine and a process called hyperlearning, they told the system to not forget as many details. The results were that the system displayed schizophrenic-like symptoms and began inserting itself into the stories. It even claimed responsibility for a terrorist bombing in one of the stories.

5. Manipulative AI Robots 

Manipulative_Robots.jpeg

Source: Autonomous Motion

In many cases, robots and AI systems seem inherently trustworthy--why would they have any reason to lie to or deceive others? Well, what if they were trained to do just that? Researchers at Georgia Tech have used the actions of squirrels and birds to teach robots how to hide from and deceive one another. The military has reportedly shown interest in the technology.

6. Artificial Intelligence Powered Robot Lovers

Sex_Robots.jpeg 

Source: The Sun

Among the many ethical concerns posed by robots and the AI systems that power them is the idea that humans could love, or at least copulate with, a robot companion. Companies are already trying to make "sex robots" a reality, and opponents are campaigning against it fervently. A machine will never replace genuine humanfeelings and emotions even if it is a sophisticated one. Building a relationship with a robot is weird. The choices we make, the actions we take, and the perceptions we have are all influenced by the emotions we are experiencing at any given moment.

7. Self Learning AI Supercomputers That Can Predict The Future

Robots_predicting_the_future.jpeg

Source: Times Higher Education

Nautilus is a supercomputer that can predict the future based on news articles. It is a self-learning supercomputer that was given information from millions of articles, dating back to the 1940s. It was able to locate Osama Bin Laden within 200km. Now, scientists are trying to see if it can predict actual future events, not ones that have already occurred.

Conclusion

Artificial intelligence (AI) is a great technological advancement and promises to eventually dominate almost every domain. AI is seeping into several industries, revolutionizing the way they conduct their business and manage their workforce. 

It may undoubtedly prove beneficial for the future but a complete AI takeover is also highly likely, if due measures aren’t taken now. Today, the use of AI has spread to almost every field of human activity. AI is like a mischievous kid that can’t be left alone and constantly needs adult supervision, and carelessness in controlling the technology can lead to an AI takeover. 

Tuesday, 5 October 2021

MLOps essentials: four pillars for Machine Learning Operations on AWS

When we approach modern Machine Learning problems in an AWS environment, there is more than traditional data preparation, model training, and final inferences to consider. Also, pure computing power is not the only concern we must deal with in creating an ML solution.

There is a substantial difference between creating and testing a Machine Learning model inside a Jupyter Notebook locally and releasing it on a production infrastructure capable of generating business value. 

The complexities of going live with a Machine Learning workflow in the Cloud are called a deployment gap and we will see together through this article how to tackle it by combining speed and agility in modeling and training with criteria of solidity, scalability, and resilience required by production environments.

The procedure we’ll dive into is similar to what happened with the DevOps model for "traditional" software development, and the MLOps paradigm, this is how we call it, is commonly proposed as "an end-to-end process to design, create and manage Machine Learning applications in a reproducible, testable and evolutionary way".

So as we will guide you through the following paragraphs, we will dive deep into the reasons and principles behind the MLOps paradigm and how it easily relates to the AWS ecosystem and the best practices of the AWS Well-Architected Framework.

Let’s start!

Why do we need MLOps?

As said before, Machine Learning workloads can be essentially seen as complex pieces of software, so we can still apply "traditional" software practices. Nonetheless, due to its experimental nature, Machine Learning brings to the table some essential differences, which require a lifecycle management paradigm tailored to their needs. 

These differences occur at all the various steps of a workload and contribute significantly to the deployment gap we talked about, so a description is obliged:

Code

Managing code in Machine Learning appliances is a complex matter. Let’s see why!

Collaboration on model experiments among data scientists is not as easy as sharing traditional code files: Jupyter Notebooks allow for writing and executing code, resulting in more difficult git chores to keep code synchronized between users, with frequent merge conflicts.

Developers must code on different sub-projects: ETL jobsmodel logictraining and validationinference logic, and Infrastructure-as-Code templates. All of these separate projects must be centrally managed and adequately versioned!

For modern software applications, there are many consolidated Version Control procedures like conventional commit, feature branching, squash and rebase, and continuous integration

These techniques however, are not always applicable to Jupyter Notebooks since, as stated before, they are not simple text files.

Development

Data scientists need to try many combinations of datasets, features, modeling techniques, algorithms, and parameter configurations to find the solution which best extracts business value

The key point is finding ways to track both succeeded and failed experiments while maintaining reproducibility and code reusability. Pursuing this goal means having instruments to allow for quick rollbacks and efficient monitoring of results, better if with visual tools.

Testing

Testing a Machine Learning workload is more complex than testing traditional software. 

Dataset requires continuous validation. Models developed by data scientists require ongoing quality evaluation, training validation, and performance checks

All these checks add to the typical unit and integration testing, defining the concept of Continuous Training, which is required to avoid model aging and concept drift

Unique to Machine Learning workflows, its purpose is to trigger retraining and serving the models automatically.

Deployment

Deployment of Machine Learning models in the Cloud is a challenging task. It typically requires creating various multi-step pipelines which serve to retrain and deploy the models automatically. 

This approach adds complexity to the solution and requires automating steps done manually by data scientists when training and validating new models in a project's experimental phase. 

It is crucial to create efficient retrain procedures!

Monitoring in Production

Machine Learning models are prone to decay much faster than "traditional" software. They can have reduced performances due to suboptimal coding, incorrect hardware choices in training and inference phases, and evolving data sets.

A proper methodology must take this degradation into account; therefore, we need a tracking mechanism to summarize workload statistics, monitor performancesand send alarm notifications

All of these procedures must be automated and are called Continuous Monitoring, which also has the added benefit of enabling Continuous Training by measuring meaningful thresholds.

We also want to apply rollbacks when a model inference deviates from selected scoring thresholds as quickly as possible to try new feature combinations.

Continuous Integration and Continuous Deployment

Machine Learning shares similar approaches to standard CI/CD pipelines of modern software applications: source control, unit testing, integration testing, continuous delivery of packages. 

Nonetheless, models and data sets require particular interventions.

Continuous integration now also requires, as said before, testing and validating data, data schemas, and models.

In this context, continuous delivery must be designed as an ML training pipeline capable of automatically deploying the inference as a reachable service.

As you can see, there is much on the table that makes structuring a Machine Learning project a very complex task. 

Before introducing the reader to the MLOps methodology, which puts all these crucial aspects under its umbrella, we will see how a typical Machine Learning workflow is structured, keeping into account what we have said until now.

Let’s go on!

A typical Machine Learning workflow in the Cloud

A Machine Learning workflow is not meant to be linear, just like traditional software. It is mainly composed of three distinct layers: datamodel, and code, and one will continuously give and retrieve feedback from others

So while with traditional software, we can say that each step that composes a workflow can be atomic and somehow isolated, in Machine Learning, this is not entirely true as the layers are deeply intertwined

A typical example is when changes to the data set require retraining or re-thinking a model. A different model also usually needs modifications to the code that runs it.

Let’s see together what every Layer is composed of and how it works.

The Data layer

The Data layer comprises all the tasks needed to manipulate data and make it available for model design and training: data ingestiondata inspection, cleaning, and finally, data preprocessing.

Data for real-world problems can be in the numbers of GB or even TB, continuously increasing, so we need proper storage for handling massive data lakes. 

The storage must be robust, allow efficient parallel processing, and integrate easily with tools for ETL jobs.

This layer is the most crucial, representing 80% of the work done in a Machine Learning workflow; two famous quotes state this fact: "garbage in, garbage out" and "your model is only as good as your data.” 

Most of these concepts are the prerogative of a Data Analytics practice, deeply entangled with Machine Learning, and we will analyze them in detail later on in this article.

The Model layer

The Model layer contains all the operations to designexperimenttrain, and validate one or more Machine Learning models. ML practitioners conduct trials on data in this layer, try algorithms on different hardware solutions, and do Hyperparameters tuning.

This layer is typically subject to frequent changes due to updates on both Data and Code, necessary to avoid concept drift. To properly handle its lifecycle management at scale, we must define automatic procedures for retraining and validation.

The Model layer is also a stage where discussions occur, between data scientists and stakeholders, about model validation, conceptual soundness, and biases on expected results.

The Code layer

In the Code layer, we define a set of procedures to put a model in production, manage inferences requests, store a model's metadata, analyze overall performancesmonitor the workflow (debugging, logging, auditing), and orchestrate CI/CD/CT/CM automatisms.

A good Code layer allows for a continuous feedback model, where the model evolves in time, taking into account the results of ongoing inferences.

All these three layers are managed by "sub-pipelines," which add up to each other to form a "macro-pipeline" known as Machine Learning Pipeline

Automatically designing, building, and running this Pipeline while reducing the deployment gap in the process is the core of the MLOps paradigm. 

MLOps on AWS: the four pillars

MLOps aims to make developing and maintaining Machine Learning workflows seamless and efficient. The data science community generally agrees that it is not a single technical solution, yet a series of best practices and guiding principles around Machine Learning.

An MLOps approach involves operations, techniques, and tools, which we can group into four main pillarsCollaborationReproducibilityContinuity, and Monitoring

We will now focus on each one, giving multiple practical examples that show how AWS, with many of its services, can be an invaluable tool to develop solutions that adhere to the paradigm’s best practices.

Collaboration

A good Machine Learning workflow should be collaborative, and collaboration occurs on all the ML pipelines.

Starting from Data Layer, we need a shared infrastructure, which means a distributed data lake. AWS offers several different storage solutions for this purpose, like Amazon Redshift, which is best for Data Warehousing, or Amazon FSx for Lustre, perfect as a distributed file system. Still, the most common service used for data lake creation is Amazon S3

To properly maintain a data lake, we need to regularly ingest data from different sources and manage shared access between collaborators, ensuring data is always up-to-date

This is not an easy task, and for that, we can take advantage of S3 LakeFormation, a managed service that helps in creating and maintaining a data lake, by working as a wrapper around AWS Glue and Glue Studio, in particular simplifying Glue’s Crawler set-up and maintenance.

S3 LakeFormation can also take care of data and collaborators' permission rules by managing users and roles underneath AWS Glue Catalog. This feature is crucial as collaboration also means maintaining governance over the data lake, avoiding unintended data manipulation by allowing or denying access to specific resources inside a catalog.

For the model layer, data scientists need a tool for collaborative design and coding of Machine Learning models. It must allow multiple users to work on the same experiment, quickly show the results of each collaborator, grant real-time pair programming, and avoid code regressions and merge conflicts as much as possible.

SageMaker is the all-in-one framework of choice for doing Machine Learning on AWS, and Amazon SageMaker Studio is a unique IDE explicitly developed for working with Jupyter Notebooks having collaboration in mind.

SageMaker Studio allows sharing a dedicated EC2 instance between different registered users, in which it is possible to save all the experiments done while developing a Machine Learning model. This instance can host Jupyter Notebooks directly or receive results, attachments, and graphics via API from other Notebook instances. 

SageMaker Studio is also directly integrated with SageMaker Experiments and SageMaker Feature Store.

The first one is a set of API that allows data scientists to record and archive a model trial, from tuning to validation, and report the results in the IDE console. The latter is a purpose-built managed store for sharing up-to-date parameters for different model trials.

SageMaker Feature Store represents a considerable step forward in maintaining governance over data parameters across different teams, mainly because it avoids a typical misused pattern of having different sets of parameters for training and inference. It is also a perfect solution to ensure that every data scientist working on a project has complete labeling visibility.

Reproducibility

To be robust, fault-tolerant, and scale properly, just like "traditional" software applications, a Machine Learning workflow must be reproducible.

One crucial point we must address with care, as we said before, is Version Control: we must ensure code, data, model metadata, and features are appropriately versioned. 

For Jupyter Notebooks, Git or AWS CodeCommit are natural choices, but managing the information of different trials, especially model metadata, requires some considerations.

We can use SageMaker Feature Store for metadata and features. It allows us to store data directly online in a managed store or integrate with AWS Glue (and S3 LakeFomation). It also enables data encryption using AWS KMS and can be controlled via API or inside SageMaker Studio.

When you want a workflow to be reproducible, you also mean experimenting on a larger scale, even in parallel, in a quick, predictable, and automatic way.

SageMaker offers different ways to mix and match different Machine Learning algorithms, and AWS allows for three possible approaches for executing a model.

Managed Algorithm: SageMaker offers up to 13 managed algorithms for common ML scenarios, and for each one, detailed documentation describes software and hardware specifications.

Bring your own algorithmdata scientists can quickly introduce custom logic on notebooks, as long as the model respects SageMaker fit() requirements.

Bring your own Containerparticular models such as DBScan require custom Kernels for running the algorithm, so SageMaker allows registering a custom container with the special Kernel and the code for running the model.

Data Scientists can tackle all these approaches together. 

SageMaker gives the possibility to define the hardware on which running a model training or validation by selecting the Instance Type and the Instance Size in the model properties, which is extremely important as different algorithms require CPU or GPU optimized machines. 

To fine-tune a model, SageMaker can run different Hyperparameter Tuning StrategiesRandom Search and Bayesian Search. These two strategies are entirely automatic, granting a way to test a more significant number of trial combinations in a fraction of time.

To enhancing the repeatability of experiments, we also need to manage different ways of doing data preprocessing (different data sets applied to the same model). For this, we have AWS Data Wrangler, which contains over 300 built-in data transformations to quickly normalize, transform, and combine features without having to write any code.

AWS Data Wrangler can be a good choice when the ML problem you’re addressing is somehow standardized, but for most cases, the datasets are extremely diverse, which means tackling ETL jobs on your own. 

For custom ETL jobs, AWS Glue is still the way to go, as it also allows saving Job Crawlers and Glue Catalogs (for repeatability). Along with AWS Glue and AWS Glue Studio, we have also tried AWS Glue Elastic Views, a new service to help to manage different data sources together.

Continuity

To make our Machine Learning workflow continuous, we must use pipeline automation as much as possible to manage its entire lifecycle.

We can break the entire ML workflow into three significant pipelines, one for each Machine Learning Layer.

Data engineering pipeline

The Data pipeline is composed of IngestionExplorationValidationCleaning, and Splitting phases. 

The Ingestion phase on AWS typically means bringing raw data to S3, using any available tool and technology: direct-API access, custom Lambda crawlers, S3 LakeFormation, or Amazon Kinesis Firehose

Then we have a preprocessing ETL phase, which is always required

AWS Glue is the most versatile among all the available tools for ETL, as it allows reading and aggregating information from all the previous services by using Glue Crawlers. These routines can poll from different data sources for new data.

We can manage Exploration, Validation, and Cleaning steps by creating custom scripts in a language of choice (e.g., Python) or using Jupyter Notebook, both orchestrated via AWS Step Functions

AWS Data Wrangler represents another viable solution, as it can automatically take care of all the steps and connect directly to Amazon SageMaker Pipelines.

Model pipeline

The Model pipeline consists of TrainingEvaluationTesting, and Packaging phases.

These phases can be managed directly from Jupyter Notebook files and integrated into a pipeline using AWS StepFunctions SageMaker SDK, which allows calling SageMaker functions inside a StepFunction script.

This exploit gives extreme flexibility as it allows to:

  1. Quickly start SageMaker training jobs with all the configured parameters.
  2. Evaluate models using SageMaker pre-build evaluation scores.
  3. Run multiple automated tests directly from code.
  4. Record all the steps in SageMaker Experiments.

Having the logic of this Pipeline on Jupyter Notebooks has the added benefit of having everything versioned and easily testable.

Packaging can be managed through Elastic Container Registry APIs, directly from a Jupyter Notebook or an external script. 

Deployment pipeline

The Deployment Pipeline runs the CI/CD part and is responsible for taking models online during the TrainingTesting, and Production phases. A key aspect during this Pipeline is that the demand for computational resources is different for all three stages and changes over time.

For example, training will require more resources than testing and production at first, but later on, as the demand for inferences will grow, production requirements will be higher (Dynamic Deployment).

We can apply Advanced deployment strategies typical of "traditional" software development to tackle ML workflows, including A/B testing, canary deployments, and blue/green deployments.

Every aspect of deployment can benefit from Infrastructure as Code techniques and a combination of AWS services like AWS CodePipeline, CloudFormation, and AWS StepFunctions.

Monitoring

Finally, good Machine Learning workflows must be monitorable, and monitoring occurs at various stages.

We have performance monitoring, which allows understanding how a model behaves in time. By continuously having feedback based on new inferences, we can avoid model aging (overfitting) and concept drift.

SageMaker Model Monitor helps during this phase as it can do real-time monitoring, detecting biases and divergences using Anomaly Detection techniques, and sending alerts to apply immediate remediation. 

When a model starts performing lower than the predefined threshold, our pipeline will begin a retraining process with an augmented data set, consisting of new information from predictions, different Hyperparameters combinations, or applying re-labeling on the data set features.

SageMaker Clarify is another service that we can exploit in the monitoring process. It detects potential bias during data preparation, model training, and production for selected critical features in the data set. 

For example, it can check for bias related to age in the initial dataset or in a trained model and generates detailed reports that quantify different types of possible bias. SageMaker Clarify also includes feature importance graphs for explaining model predictions.

Debugging a Machine Learning model, as we can see, is a long, complex, and costly process! There is another useful AWS service: SageMaker Debugger; it captures training metrics in real-time, such as data loss during regression, and sends alerts when anomalies are detected

SageMaker Debugger is great for immediately rectifying wrong model predictions.

Logging on AWS can be managed on the totality of the Pipeline using Amazon CloudWatch, which is available with all the services presented. Cloudwatch can be further enhanced using Kibana through ElasticSearch to have an easy way to explore log data.

We can also use CloudWatch to trigger automatic rollback procedures in case of alarms on some key metrics. Rollback is also triggered by failed deployments.

Finally, the reproducibility, continuity, and monitoring of an ML workload enables the cost/performance fine-tuning process, which happens cyclically across all the workload lifecycle. 

Sum Up

In this article, we’ve dived into the characteristics of the MLOps paradigm, showing how it took concepts and practices from its DevOps counterpart to allow Machine Learning to scale up to real-world problems and solve the so-called deployment gap.

We’ve shown that, while traditional software workloads have more linear lifecycles, Machine Learning problems are based on three macro-areas: Data, Model, and Code which are deeply interconnected and provide continuous feedback to each other.

We’ve seen how to tackle these particular workflows and how MLOps can manage some unique aspects like complexities in managing model’s code in Jupyter Notebooks, exploring datasets efficiently with correct ETL jobs, and providing fast and flexible feedback loops based on production metrics.

Models are the second most crucial thing after data. We’ve learned some strategies to avoid concept drift and model aging in time, such as Continuous Training, which requires a proper monitoring solution to provide quality metrics over inferences and an adequate pipeline to invoke new model analysis.

AWS provides some managed services to help with model training and pipelines in general, like SageMaker AutoPilot and SageMaker Pipelines.

We have also seen that AWS allows for multiple ways of creating and deploying models for inference, such as using pre-constructed models or bringing your container with custom code and algorithms. All images are saved and retrieved from Elastic Container Registry.

We’ve talked about how collaboration is critical due to the experimental nature of Machine Learning problems and how AWS helps by providing an all-in-one managed IDE called SageMaker Studio.

We have features like SageMaker Experiments for managing multiple experiments, SageMaker Feature Store for efficiently collecting and transforming data labels, or SageMaker Model Monitoring and SageMaker Debugger for checking model correctness and find eventual bugs.

We’ve also discussed techniques to make our Machine Learning infrastructure solid, repeatable, and flexible, easy to scale on-demand based on requirements evolving in time. 

Such methods involve using AWS Cloudformation templates to take advantage of Infrastructure as Code for repeatability, AWS Step Functions for structuring a state-machine to manage all the macro-areas, and tools like AWS CodeBuild, CodeDeploy, and CodePipeline to design proper CI/CD flows. 

We hope you’ve enjoyed your time reading this article and hopefully learned a few tricks to manage your Machine Learning workflows better.

As said before, if Machine Learning is your thing, we encourage again having a look at our articles with use-cases and analysis on what AWS offers to tackle ML problems here on Proud2beCloud!
As always, feel free to comment in the section below, and reach us for any doubt, question or idea! See you on #Proud2beCloud in a couple of weeks for another exciting story!

#MachineLearning #DataScience #Python #AI #100DaysOfCode #DEVCommunity #IoT #flutter #javascript #Serverless #womenintech #cybersecurity #RStats #technology  #WomenWhoCode #DeepLearning #data #MLOps #DevOps







Rank

seo