Saturday, 22 January 2022

Running ZEUS Analysis Code


Although I built CERNLib a long time ago, apart from running PAW and looking at some old Ntuples I didn't get round to building my old ZEUS analysis code.

Setup the Environment

CERNLib is available as a package on Ubuntu.  However I've seen issues around running this on 64bit Linux. So for practicality I decided to setup a 32bit Linux VM. As noted before I opted for Linux Mint on VirtualBox. I picked the 32bit version of Mint 19.3 Tricia. 


Code Changes

So the Fortran compiler has moved on a bit since October 2000 which was the last time I ran this code. I needed to make a few changes to successfully compile my source. Fortunately nothing too significant was required.

.EQ. to .EQV.

I was using .EQ. for logical comparison of a LOGICAL variable, for example:
        IF (Q2WTFIRST.EQ..TRUE.) THEN
However this should be .EQV. :
        IF (Q2WTFIRST.EQV..TRUE.) THEN

Using 1 and 0 for True and False

Related to the above, a compiler error was thrown when using .eq. and 1 or 0 to represent true or false

tltcut = 0
... Change tltcut to 1 if the condition is met
if (tltcut.eq.1) then

Fix was to change the condition to:
        if (tltcut) then

DFLIB

At the time Bristol University were trying out Visual Fortran (DEC now Intel Visual Fortran) . We used this library in a Math error handler routine. DFLIB has quite a history but the upshot is it's only supported on Windows. I just took out the routine using it!

Makefile

Using Visual Fortran the build and linker was controlled by a. To build using gfortran in Linux I had to create a Makefile. This is an art in itself! The main issue I had to overcome was linking the CERNLIB libraries. The required libraries below, mathlib, packlib and kernlib were available to the linker with the standard path. pdflib however was in a separate folder that needed including with the -L flag. Some looking around to find the correct name was required, in the end I needed pdflib804.

-L /usr/lib/i386-linux-gnu -lmathlib -lpacklib -lkernlib -lpdflib804

.inc Files

Include files in Fortran are inserted into the source program. This is useful if you have a common block that you don't want to write over and over, for example defining variables. 


This was referenced in the source files with :
        #include "common.inc"

Incidentally, the # indicated the file was processed by a C preprocessor rather than the standard Fortran include.

My inc files were mixed in with the Fortran source files. The compiler didn't like this so I updated the folder structure to use separate src/inc/exe folders. Funnily enough this reflected the original folder structure when I ran the code on Unix boxes before we switched to Visual Fortran. The circle is complete!

How to run the Analysis code

This took a bit of remembering. The code was driven by a number of input files. 
  • Steering cards. This contains a list of cuts made to select the events. Different cards would define different cuts and therefore allow me to analyse systematic errors
  • File listing input Data .rz ntuples
  • File listing input Background .rz ntuples
  • File listing input MonteCarlo .rz ntuples
Of course, naming convention wasn't particularly useful, for a run on 1997 data there would be:
  • stc97_n for the steering cards
  • fort.45 for the Data file list
  • fort.46 for the Background file list
  • fort.47 for the MonteCarlo file list

I should make clear at this point that the files I'm running against are not the "raw" ZEUS datafiles. A preliminary job was run against the full set of data on tape to load a cutdown set of data that passed some loose cuts and create an ntuple with the necessary fields for this particular analysis. My ntuples were saved after this first step.

Results

I ran the code against a single 1997 data, background and MonteCarlo ntuple.

This created, amongst some other files, an israll97.hbook file. By opening PAW and entering hi/file 1 israll.hbook I was able to load this. hi/li listed the available histograms, this looked about right and when I opened one of them with hi/pl 105 success! I had created a histogram of the E-Pz (Total Energy - Momentum in the z, or beampipe direction) from my saved data files. I used this originally to help select the population of events to use in my analysis.



I've put the code up on GitHub:

The README.md describes the list of input and output files.

Next steps...
  • Run the PAW macro (kumac) files I used after this analysis step and add to GitHub.
  • Run against the full data set. I want to see how quickly this runs on current hardware. To run the full 96/97 data set against all steering cards would take about 20 hours!
  • Create a GitHub release to include sample .rz files. Hopefully to allow anyone to run this!
  • Try and build the 64bit version of CERNLIB so I don't have to run on a 32bit VM.

Wednesday, 12 January 2022

Adding Azure Pipeline status to GitHub README.md

I noticed in some GitHub projects a build status badge in the README.md display. As I have Azure Pipeline builds setup for two of my projects, CSVComparer and SuperMarketPlanner, I thought it would a good idea to add this.

Handily, there is a REST API interface available for the Azure Pipeline build!

The link for the badge is of the form:

https://dev.azure.com/{Organisation}/{Project}/_apis/build/status/{pipelinename}

For CSVComparer that's: 

https://dev.azure.com/jonathanscott80/CSVComparer/_apis/build/status/jscott7.CSVComparer

Then I can add a hyperlink to the badge to navigate to the latest build page when clicked. This is of the form:

https://dev.azure.com/{Organisation}/{Project}/_build/latest?definitionId={id}

Again, for CSVComparer that's: 

https://dev.azure.com/jonathanscott80/CSVComparer/_build/latest?definitionId=2

The definitionId is the ID of the Pipeline assigned by Azure. You can find this by navigating through to the Pipeline page via https://dev.azure.com/{Organisation}/

(Useful link for the official docs here : Azure Pipelines documentation | Microsoft Docs)

Finally, to put it all together add this to the README.md file:

[![Build Status](https://dev.azure.com/jonathanscott80/CSVComparer/_apis/build/status/jscott7.CSVComparer)](https://dev.azure.com/jonathanscott80/CSVComparer/_build/latest?definitionId=2)

Here is what it looks like:


And clicking on the badge takes us here:



Friday, 12 November 2021

Setting up a new Laptop (2021 Edition)

It's been 6 years since I last setup a dual booting laptop. That particular machine is starting to show its age now, particularly with a spinning disk hard-drive.

Before doing anything I created a Recovery USB. Search Recovery Drive on the task bar to open Recovery Media Creator. I followed the instructions to setup a recovery drive on a 32GB USB drive.

The new machine has a 1TB SSD and I wanted to really keep as much as possible for Windows. So, I checked how much space I used on my existing Linux install. The command df . -BG is useful for this. giving space used in GB. It turned out I had only used 75 GB. 

I downloaded the Ubuntu 21.10 ISO and used Startup Disk Creator to flash it to an old 4GB USB stick

To free up space for the Linux installation I needed to shrink the Windows disk partition. I used the Windows tool for this. Search for "Create and format hard disk partitions". This opened the Disk Management Tool:


I right-clicked on the main Windows Partition and selected Shrink... I then shrank by 150GB.

Once everything was in place. I inserted the Ubuntu USB and rebooted. For this machine I needed to explicitly tell it to boot from the USB by hitting F12 during startup to be able to change the boot order.

A welcome screen should pop up with options to Try Ubuntu or Install Ubuntu. Select the latter and follow the next screens, selecting language, keyboard and Wi-Fi when asked. 

  • When you get to Installation Type select "Something else, you can create or resize partitions yourself"
This will open a window listing the partitions. The empty 150GB partition created earlier is visible here. Select it and click on + to create a new partition. Do this twice:
  1. Create a Swap Partition. Select Logical Partition, beginning of this space and choose a size (I used 4GB) and under Mount point choose Swap
  2. Create a Root Partition. Select Primary Partition, beginning of this space and let it take the remaining space. Under Mount point choose /
Previously I created a home partition but decided I didn't really need it this time.

After going through the remaining screens the installation proceeded. 

Jittery screen
The Ubuntu installation was a success but the screen did have jitters, especially duing start up. I found the solution to this by editing /etc/default/grub and changing the GRUB_CMDLINE_LINUX_DEFAULT line to:

GRUB_CMDLINE_LINUX_DEFAULT="quiet splash i915.enable_psr=0"

After editing enter sudo update-grub and reboot

PSR is a Power saving setting (Panel Self Refresh) used to optimise power consumed by buffering some frames to the display when the view is static.

Bitlocker
After rebooting I now got the Grub menu to be able to select Ubuntu or Windows. When I selected Windows however I got a warning screen popup

"The system boot information has changed since BitLocker was enabled. You must supply a BitLocker recovery password to start this system."

I was able to log onto my Microsoft Account to obtain the recovery key: aka.ms/myrecoverykey

The machine can now boot easily to Windows or Ubuntu.





Monday, 16 August 2021

Python Range Bar Plots

I have some data in a json file that contains a list of timings for different tests. Each test has a name, start timestamp and duration in milliseconds. It looks something like this:

Now, I love the Matplotlib Python library. I used it a lot when learning Machine Learning during the 2020 lockdowns. It is an incredibly rich and powerful tool  for creating professional data visualizations. So, as I'm new to it I thought I'd see if I could create a Range Bar plot of these data. 

In essence I want to show each test as a individual bar, the size of which corresponds to the duration and each bar's left most edge corresponding to the start time.

Here is the script:

It's very simple. Most of the script is concerned with reading the data from the file and generating the data structures. I thought it would be nice to highlight the longest duration test in red. This would be useful for seeing if the times for different tests were close.
 
And the output:


Very simple. There are of course many options to polish this. I'll update the blog as I refine them.

Monday, 17 May 2021

.NET Versions and Transport Layer Security (TLS)

This little blog came about from looking into a strange problem where a query to a rest service was failing with :

System.Net.WebException: The underlying connection was closed: An unexpected error occurred on a send.

System.IO.Exception: Unable to read data from the transport connection: An Existing connection was forcibly closed by the remote host

The cause of this is from an inconsistent security protocol being used by the client and server. In this case the server is using TLSv1.2 and the client SSLv3.

You can get or set the SecurityProtocol in code:

I'm running this on .NET framework 4.8. So I should be using the SystemDefault protocol for the OS ( Windows 10) which is TLSv1.2. But when I printed out the output from the SecurityProtocol above I got SSLv3 | TLSv1.1 So what's going on?

Supported Runtime

In my app.config file I had the following startup entry:

Documentation here The versions are checked in order with the first one that matches a version installed on your being taken. If that version is present on your computer that will be used.

So with 4.5.2 the default version of SSL3 TLS1 and TLS 1.1 are used. 


There's a few fixes or workarounds for this. First build the code and retarget .NET 4.7.2 or later. Unfortunately this wasn't an option because I didn't own the executable code. This also ruled out explicitly setting the SecurityProtocol in code although this isn't recommended as it makes future changes difficult. You can also edit registry keys (see the link below) to change the default behaviour but I thought this would be too intrusive.

The solution was to add the following to the app.config:


For 4.5.2 this also requires the latest patches and registry keys . In an ideal word you would of course update. But software development rarely works in an ideal world.




Monday, 22 March 2021

Publishing .NETCore Application and running on Linux

Over the last year or so I've been writing a simple tool in .NET Core to compare two CSV files. It can be found in on GitHub here. I won't go into details about the tool in this post but being a .NET Core application I've been able to develop it on Windows and also deploy and run on Linux. This post describes how I published and deployed.

I've been using Visual Studio. When I want to create a new deployment I right-click on the Project and select Publish...

I have setup a profile, called "FolderProfile" (Very imaginative!)

  • TargetLocation : Folder name
  • Configuration : Release | Any CPU (from my solution)
  • Deployment mode : Framework-dependent
  • TargetFramework : netcoreapp3.0
  • TargetRuntime : Portable

(I used Framework-dependent deployment for cross-platform binaries. This creates a dll which can be run with dotnet <filename.dll> on any platform. Of course this requires the .NET runtime to be installed on the Linux machine.)
 
Then click on Publish.
 
This created a folder under Release/netcoreapp3.0/publish
 
I copied these files onto my Linux machine. Here are the important ones:
total 68
drwxr-xr-x  2 jonathan jonathan  4096 Mar 12 19:27 .
drwxr-xr-x 64 jonathan jonathan  4096 Mar 12 19:35 ..
-rwx------  1 jonathan jonathan   431 Jul  2  2020 CSVComparison.deps.json
-rwx------  1 jonathan jonathan 14336 Jul  2  2020 CSVComparison.dll
-rwx------  1 jonathan jonathan  4860 Jul  2  2020 CSVComparison.pdb
-rwx------  1 jonathan jonathan   154 Jul  2  2020 CSVComparison.runtimeconfig.json 
 

I've previously installed the .NET Core runtime on Ubuntu. As of the time of writing the installed version (dotnet --version) is 3.1.404.

The IMDB/movie_data.csv file is a 50000 row CSV I generated from the IMDB movie dataset. I had originally set this up for a machine learning exercise but it is also a good sized dataset to checkout the tool. I made two copies and edited one of the comments on the candidate file. Then, to run the CSV Comparison, I navigated to the folder containing the binaries and ran with this command-line:

dotnet ./CSVComparison.dll ../IMDB/movie_data2.csv ../IMDB/movie_data2.csv ../IMDB/movie_config.xml ./output

The output:

Reference: ./movie_data.csv
Candidate: ./movie_data2.csv
Saving results to output/ComparisonResults.csv
Finished. Comparison took 10517ms

And the result:

?xml version="1.0" encoding="utf-8"?>
<ComparisonDefinition xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns:xsd="http://www.w3.org/20
01/XMLSchema">
  <Delimiter>,</Delimiter>
  <KeyColumns>
    <Column>row</Column>
  </KeyColumns>
  <HeaderRowIndex>0</HeaderRowIndex>
  <ToleranceValue>0.1</ToleranceValue>
  <IgnoreInvalidRows>false</IgnoreInvalidRows>
  <ToleranceType>Relative</ToleranceType>
  <ExcludedColumns />
</ComparisonDefinition>

Date run: 22/03/2021 18:40:12
Reference: ./movie_data.csv
Candidate: ./movie_data2.csv
Number of Reference rows: 50001
Number of Candidate rows: 50001
Comparison took 10517ms
Number of breaks 1

Break Type,Key,Column Name,Reference Row, Reference Value, Candidate Row, Candidate Value
ValueMismatch,7230,review,96,"Exceptional movie that handles a theme of

Tuesday, 2 March 2021

Time for some Maths

While tidying up my old University Maths notes the other day I came across this problem. (I didn't get it right first time round!)

It's quite interesting and I thought I'd also use the opportunity to setup nice looking maths on the blog. So, the question:

Find the first two terms of the series for tanx in powers of x.

Now, 

$$tanx=\frac{sinx}{cosx}$$

and the power series expansions of sinx and cosx begin as:

$$sinx=x-\frac{x^3}{3!}+\frac{x^5}{5!}+...$$

$$cosx=1-\frac{x^2}{2!}+\frac{x^4}{4!}+...$$

Put this together:

$$tanx=\frac{x-\frac{x^3}{6}+...}{1-\frac{x^2}{2}+...}$$

Multiply by $\frac{1+\frac{x^2}{2}}{1+\frac{x^2}{2}}$ (This is just multiplying by 1 and chosen so we can deal with the $x^2$ term in the denominator), so:

$$tanx=\frac{(1+\frac{x^2}{2})(x-\frac{x^3}{6}+...)}{(1+\frac{x^2}{2})(1-\frac{x^2}{2}+...)}$$

If we now just multiply out the first term we get

$$tanx=\frac{x+\frac{x^3}{2}-\frac{x^3}{6}+...}{1+\frac{x^2}{2}-\frac{x^2}{2}+...}$$

$$tanx=\frac{x+\frac{x^3}{3}+...}{1+...}$$

The other terms in $x^4$ and $x^5$ can be ignored, giving:

$$tanx=x+\frac{x^3}{3}+...$$


To display the formulae I've added MathJax to the blog. 

Go to Theme then from the Customise dropdown select Edit HTML. Add the following just below the <head> element:



Finally, here's a useful link: MathJax basic tutorial and quick reference - Mathematics Meta Stack Exchange