Technical notes about past publications and work by Darrell Ulm including Apache Spark, software development work, computer programming, Parallel Computing, Algorithms, Koha, and Drupal. Source code snippets, like in Python for Spark. Retrospective of projects.
Python for Data Science
There are many good resources available.
Of course the main tools are: Numpy, Pandas, MathPlotLib, SkiKit-Learn has some amazing tools.
Kaggle for instance has Data Science contents, but good to install a local system like the Jupyter Notebook to speed things up as the Kaggle editor can lag and take some time to run on small data-sets.
The newer DataCamp has some neat tutorials on it and simple App to do daily exercises on your mobile device.
Here is the Python DataScience Handbook. Really useful.
A short tutorial: Learn Python for Data Science, a fun read.
A list of cool DataSci tutorials is here, and another how to get started with Python for DS.
Will add more later.
My Growing Collection of Tech Notes: Drupal, PHP, Linux, Symfony, and More
I’ve been keeping a running set of technical notes on Tumblr as I work through different web development projects. Over time it has turned into a personal reference library that covers Drupal development, PHP programming, Linux server setup, and the Symfony framework. Most of these notes come from real problems I’ve solved while building or maintaining websites, so the collection keeps expanding as I learn new tools and techniques.
A large portion of my notes focuses on Drupal because I spend a lot of time working with Drupal 7, Drupal 8, and the transition toward Drupal 9. I’ve documented everything from module development and data migration to caching, performance optimization, Varnish configuration, and headless Drupal workflows. Since Drupal 8 and Drupal 9 are built on Symfony, I also keep notes on Symfony concepts and PHP best practices that help improve development speed and code quality.
I also write down what I learn while setting up and managing Linux servers. Many of these entries involve Ubuntu 16.04, including installing essential software, configuring GNOME Flashback, setting up Webmin, enabling SSL on Apache, and improving performance with Memcached and PHP OpCode caching. As I explore more DevOps tools, I’ve added notes on Docker, Composer, Drush, and other utilities that make modern development smoother.
This list keeps growing as I continue learning about backend development, server optimization, and emerging technologies like augmented reality toolkits. Here are the topics I’ve documented so far:
- Drupal 8 alpha release for Google Books Text Filter Module
- Ubuntu 16.04 Setup Gnome Flashback
- Gnome Flashback Move Menu Bar to Bottom of Screen
- Install Webmin to Ubuntu 16.04
- Install Google Chrome amd64 on Ubuntu 16.04 Linux
- Install PHP on Linux
- Install Drush for Drupal via Composer
- Install Composer on Linux
- Download and Install Drupal 8 with Details
- Install MySQL Workbench on Ubuntu 16.04
- Turn on the PHP 7 OpCode Cache
- Nice tutorial for SSL for Apache2 on Ubuntu 16.04
- Memcached on Ubuntu 16.04
- Drupal 7 Performance Optimizations
- Varnish Setup Instructions for Drupal 7 and more Performance Links
- PHP Versions Fast Switch Between
- Drupal 7 CKEditor Module with Simple Image Upload
- Drupal 7 Block Caching API, Configurations and Modules
- Drupal 8 Data Migration Information
- Drupal 8 API : Custom Module Development in PHP
- HTML Archive Methods and Software
- PHP 5.6 to PHP 7.0 Upgrade Compatibility Check
- Kalabox, 1 click local development, Drupal, WordPress
- Installing Docker on Linux
- Drupal AdvAgg Advanced Aggregation Module
- Symfony PHP Framework Tutorials
- Drupal Permissions and Access Control
- Headless Drupal
- Twig Theme Coding with Drupal 8
- Drupal 8 and Backwards Compatibility in Drupal 9
- Augmented Reality AR Tool-kits
- Drupal 7 Varnish and Page Cache
- PHP Programming Optimization Methods
Catch up on Drupal and Ubuntu Linux Posts
- Drupal 8 Development in PHP
- Migration Tutorials for Drupal 8 (from Drupal 7 primarily or other systems)
- Technical Notes for Config of Drupal 7
- For Ubuntu 16 Setup Notes for Web Development System
Getting back into parallel computing with Apache Spark
Returning to parallel computing with Apache Spark has been insightful, especially observing the increasing mainstream adoption of the McColl and Valiant BSP (Bulk Synchronous Parallel) model beyond GPUs. This structured approach to parallel computation, with its emphasis on synchronized supersteps, offers a practical framework for diverse parallel architectures.While setting up Spark on clusters can involve effort and introduce overhead, ongoing optimizations are expected to enhance its efficiency over time. Improvements in data handling, memory management, and query execution aim to streamline parallel processing.A GitHub repository for Spark snippets has been created as a resource for practical examples. As Apache Spark continues to evolve in parallel with the HDFS (Hadoop Distributed File System), this repository intends to showcase solutions leveraging their combined strengths for scalable data processing.