Python for Data Science

Looking at more resources online for Python for Data Science.

There are many good resources available.

Of course the main tools are: NumpyPandasMathPlotLibSkiKit-Learn has some amazing tools.

Kaggle for instance has Data Science contents, but good to install a local system like the Jupyter Notebook to speed things up as the Kaggle editor can lag and take some time to run on small data-sets.

The newer DataCamp has some neat tutorials on it and simple App to do daily exercises on your mobile device.

Here is the Python DataScience Handbook. Really useful.

A short tutorial: Learn Python for Data Science, a fun read.

A list of cool DataSci tutorials is here, and another how to get started with Python for DS.

Will add more later.


My Growing Collection of Tech Notes: Drupal, PHP, Linux, Symfony, and More

I’ve been keeping a running set of technical notes on Tumblr as I work through different web development projects. Over time it has turned into a personal reference library that covers Drupal development, PHP programming, Linux server setup, and the Symfony framework. Most of these notes come from real problems I’ve solved while building or maintaining websites, so the collection keeps expanding as I learn new tools and techniques.

A large portion of my notes focuses on Drupal because I spend a lot of time working with Drupal 7, Drupal 8, and the transition toward Drupal 9. I’ve documented everything from module development and data migration to caching, performance optimization, Varnish configuration, and headless Drupal workflows. Since Drupal 8 and Drupal 9 are built on Symfony, I also keep notes on Symfony concepts and PHP best practices that help improve development speed and code quality.

I also write down what I learn while setting up and managing Linux servers. Many of these entries involve Ubuntu 16.04, including installing essential software, configuring GNOME Flashback, setting up Webmin, enabling SSL on Apache, and improving performance with Memcached and PHP OpCode caching. As I explore more DevOps tools, I’ve added notes on Docker, Composer, Drush, and other utilities that make modern development smoother.

This list keeps growing as I continue learning about backend development, server optimization, and emerging technologies like augmented reality toolkits. Here are the topics I’ve documented so far:




Catch up on Drupal and Ubuntu Linux Posts

I’ve been catching up on my Ubuntu 16.04 Linux setup notes along with several Drupal posts. Below is a small collection of documentation, tutorials, and helpful threads that cover a range of web development topics. These notes focus on Drupal development, PHP programming, and building a reliable Linux server environment for web projects.There are a few entries I still need to pull over from my Tumblr archive, and I’ll add those when I have more time. As I continue working with Drupal, PHP, and Linux, this list will keep growing and improving.


Getting back into parallel computing with Apache Spark

Returning to parallel computing with Apache Spark has been insightful, especially observing the increasing mainstream adoption of the McColl and Valiant BSP (Bulk Synchronous Parallel) model beyond GPUs. This structured approach to parallel computation, with its emphasis on synchronized supersteps, offers a practical framework for diverse parallel architectures.While setting up Spark on clusters can involve effort and introduce overhead, ongoing optimizations are expected to enhance its efficiency over time. Improvements in data handling, memory management, and query execution aim to streamline parallel processing.A GitHub repository for Spark snippets has been created as a resource for practical examples. As Apache Spark continues to evolve in parallel with the HDFS (Hadoop Distributed File System), this repository intends to showcase solutions leveraging their combined strengths for scalable data processing.