jonny.donker@gmail.com

Qlik Replicate: Stripping latency out of log files

We have been doing a bit of “Stress and Volume” testing in Qlik Replicate over the past few days; investigating how my latency is introduced to a MS SQL server task if we run it through a log stream.

If you’re not aware – you can get minute latecy from Qlik Replicate by turning the “Performance” logs up to “Trace” or higher:

This will result in messages getting produced in the task’s log file like:

00011044: 2025-08-04T14:57:56 [PERFORMANCE     ]T:  Source latency 1.23 seconds, Target latency 2.42 seconds, Handling latency 1.19 seconds  (replicationtask.c:3879)
00011044: 2025-08-04T14:58:26 [PERFORMANCE     ]T:  Source latency 1.10 seconds, Target latency 2.27 seconds, Handling latency 1.16 seconds  (replicationtask.c:3879)
00009024: 2025-08-04T14:58:30 [SOURCE_CAPTURE  ]I:  Throughput monitor: Last DB time scanned: 2025-08-04T14:58:30.700. Last LSN scanned: 008050a1:00002ce8:0004. #scanned events: 492815.   (sqlserver_log_utils.c:5000)
00011044: 2025-08-04T14:58:57 [PERFORMANCE     ]T:  Source latency 0.61 seconds, Target latency 1.87 seconds, Handling latency 1.25 seconds  (replicationtask.c:3879)
00011044: 2025-08-04T14:59:27 [PERFORMANCE     ]T:  Source latency 1.10 seconds, Target latency 1.55 seconds, Handling latency 0.45 seconds  (replicationtask.c:3879)

I created a simple python script to parse a folder of log files and output it in a excel. Since the data ends up in a panda data frame; it would be easy to manipulate the data and output it in a specific way:

import os
import pandas as pd
from datetime import datetime


def strip_seconds(inString):
    
    index = inString.find(" ")  
    returnString = float(inString[0:index])
    return(returnString)


def format_timestamp(in_timestamp):
    
    # Capture dates like "2025-08-01T10:42:36" and add a microsecond section
    if len(in_timestamp) == 19:
        in_timestamp = in_timestamp + ":000000"

    date_format = "%Y-%m-%dT%H:%M:%S:%f"

    # Converts date string to a date object    
    date_obj = datetime.strptime(in_timestamp, date_format)

    return date_obj

        
def process_file(in_file_path):

    return_array = []
    
    with open(in_file_path, "r") as in_file:
    
        for line in in_file:

            upper_case = line.upper().strip()

            if upper_case.find('[PERFORMANCE     ]') >= 0:
                timestamp_temp = upper_case[10:37]
                timestamp = timestamp_temp.split(" ")[0]
                
                split_string = upper_case.split("LATENCY ")
                
                if len(split_string) == 4:
                
                    source_latency = strip_seconds(split_string[1])
                    target_latency = strip_seconds(split_string[2])
                    handling_latency = strip_seconds(split_string[3])

                    # Makes the date compatible with Excel
                    date_obj = format_timestamp(timestamp)
                    excel_datetime = date_obj.strftime("%Y-%m-%d %H:%M:%S")
                    
                    # If you're outputting to standard out
                    #print(f"{in_file_path}\t{time_stamp}\t{source_latency}\t{target_latency}\t{handling_latency}\n")

                    return_array.append([in_file_path, timestamp, excel_datetime, source_latency, target_latency, handling_latency])
                     
    return return_array
    

if __name__ == '__main__':

    log_folder = "/path/to/logfile/dir"
    out_excel = "OutLatency.xlsx"
    
    latency_data = []

    # Loops through files in log_folder
    for file_name in os.listdir(log_folder):

        focus_file = os.path.join(log_folder, file_name )

        if os.path.isfile(focus_file):
            filename, file_extension = os.path.splitext(focus_file)

            if file_extension.upper().endswith("LOG"):
                print(f"Processing file: {focus_file}")
                return_array = process_file(focus_file)

                latency_data += return_array
                

    df = pd.DataFrame(latency_data, columns=["File Name", "Time Stamp", "Excel Timestamp", "Source Latency", "Target Latency", "Handling Latency"])
    df.info()

    # Dump file to Excel; but you can dump to other formats like text etc
    df.to_excel(out_excel)

And voilà – we have an output to Excel to quickly create statistics on latency for a given task(s).

What statistics to use?

Latency is something you can analyse in different ways, depending on what you’re trying to answer. It is also important to pair latency statistics with the change volume that is coming through.

Has the latency jumped at a specific time because the source database is processing a daily batch? Is there a spike of latency around Christmas time where there are more financial transactions compared to a benign day in February?

Generally, if the testers are processing the latency data they provide back:

Average target latency
90^th percentile target latency
95^th percentile target latency
Min target latency
Max target latency

This is what we use to compare two runs to each other when the source load is consistent and we’re changing a Qlik Replicate task

August 6, 2025 by jonny.donker@gmail.com Python Qlik Replicate 0

Qlik Replicate – MS SQL dates to Confluent Avro

I am just writing a brief post about a conversation I was asked to join.

The question paraphrased:

For dates in an Microsoft SQL Server, when they get passed through Qlik Replicate to Kafka Avro – how are they interpreted? Epoch to utc? Or to the current time zone?

I didn’t know the answer so I created the following simple test:


CREATE TABLE dbo.JD_DATES_TEST
(
 ID INT IDENTITY(1,1) PRIMARY KEY,
 TEST_SMALLDATETIME SMALLDATETIME,
 TEST_DATE date,
 TEST_DATETIME datetime,
 TEST_DATETIME2 datetime2,
 TEST_DATETIMEOFFSET datetimeoffset
);

GO

INSERT INTO dbo.JD_DATES_TEST VALUES(current_timestamp, current_timestamp, current_timestamp, current_timestamp, current_timestamp);

SELECT * FROM dbo.JD_DATES_TEST;
/*
ID          TEST_SMALLDATETIME      TEST_DATE  TEST_DATETIME           TEST_DATETIME2              TEST_DATETIMEOFFSET
----------- ----------------------- ---------- ----------------------- --------------------------- ----------------------------------
1           2025-06-12 12:04:00     2025-06-12 2025-06-12 12:04:16.650 2025-06-12 12:04:16.6500000 2025-06-12 12:04:16.6500000 +00:00

(1 row(s) affected)


*/

On the other side after passing through Qlik Replicate and then onto Kafka in Avro format; we got:

{
  "data": {
    "ID": {
      "int": 1
    },
    "TEST_SMALLDATETIME": {
      "long": 1749729840000000
    },
    "TEST_DATE": {
      "int": 20251
    },
    "TEST_DATETIME": {
      "long": 1749729856650000
    },
    "TEST_DATETIME2": {
      "long": 1749729856650000
    },
    "TEST_DATETIMEOFFSET": {
      "string": "2025-06-12 12:04:16.6500000 +00:00"
    },
    "x_y": {
      "string": "1.0.0"
    }
  },
  "beforeData": null,
  "headers": {
    "operation": "REFRESH",
    "changeSequence": "",
    "timestamp": "",
    "streamPosition": "",
    "transactionId": "",
    "changeMask": null,
    "columnMask": null,
    "transactionEventCounter": null,
    "transactionLastEvent": null
  }
}

So 1749729856650000 equals Thursday, June 12, 2025 12:04:16.650 PM – which is local time.

June 12, 2025 by jonny.donker@gmail.com Qlik Replicate SQL 0

Borderline illegal Mac and Cheese

This is a Mac + Cheese recipe that I learnt from the OMG BBQ course.

It is truly frightening how decadent it is; I have never seen so much cheese go into one recipe. And that’s not even considering the amount of cream to add to it.

This is the recipe dictated by the owner so I haven’t tried cooking it myself; but the results were the most gooey, artery clogging, heart stopping Mac n Cheese I have ever tasted.

The Legend behind the recipe

The legend where this recipe originates from is from a “gangster” living in America; involved in all sorts of violence and drugs. One day an epiphany came to him that he was heading towards a very early grave.

So he got out of crime, went to rehab to got clean and opened up a BBQ restaurant/food van.

But he was looking for ideas to expand his business.

Using his former life’s intuition and the opportunities that some US states has leagalised cannabis; he realised there was a market for good “munchies” food at cannabis festivals that are regularly held.

The problem with some food truck’s menu choice; hot chips are a bottleneck to the process. Quite often people are waiting for chips to be cooked.

He was searching for something that he can serve with BBQ meats, very quick and easy to make, not a bottleneck in a food truck business and satisfies the festival goers particular needs.

This is when he came up with this recipe for Mac and Cheese.

The Recipe

Guesstimate serves 8 people

Ingredients

500g Macaroni pasta shapes
1kg Tasty cheese – grated
50gm Parmesan cheese – grated
2 large eggs
1L Thicken cream
2 tsp ground Black pepper
2 tsp Garlic Powder (Guesstimate – chef did it by feel)
2 tsp Onion powder (Guesstimate – chef did it by feel)
2 tsp Smoked Paprika (Guesstimate – chef did it by feel)
Extra Paprika for dusting on top

Method

Pre-heat an oven to 160°c .
Cook the pasta to packet instructions. Drain the pasta and allow it to cool slightly.
Reserve a cup of grated cheese.
Meanwhile in a large bowl; add in remaining cheese, eggs, thickened cream and seasonings. Gently mix to combine.
Once the pasta is cool enough; add into the bowl and stir to combine everything.
Pour into a baking dish. Sprinkle over reserved cheese and dust evenly with paprika.
Place on a tray to catch the drippings and bake for 25min until lightly brown on top.
Rest for 5min and then serve

Notes

While I don’t think one’s cardiologist would recommend eating this every day; I do see some potential as using this recipe as a base. This would be a show stopping comfort food on a cold winter’s night; either served as a side, or as a main with some nice bread.

One thing I would change from the base recipe that we were taught is to under cook the pasta by a couple of minutes. This will prevent it from going too mushy in the final product.

Other things I was thinking:

Don’t use pre-grated cheese – the additives in it don’t make the cheese melt as well.
Add some creamy compatible vegetables like pumpkin, broccoli, mushrooms and/or corn to it.
Instead of smoked paprika, use nutmeg
Experiment with a different combination of cheeses to add a more complex flavour.

June 9, 2025 by jonny.donker@gmail.com Cook 0

Nancy’s Christmas Noodle salad – Feast for the masses

My introduction to my in-law’s Christmas traditions was “Nancy’s Noodle Salad”; a recipe handed down to my to-be wife that she would cook for an Australian Christmas lunch.

The whole trouble is with her excitement for Christmas; she would quite often cook way too much; and we’re eating “Nancy’s Noodle Salad” for a week afterwards.

But this salad is important; nostalgias of times gone by and remembering passed on love ones. And someone on this vast World Wide Web might pick up Nancy’s Christmas Noodle salad and start their own tradition with it.

Anyway – I can’t read the original recipe and I have to get my wife to interpret the hand writing.

The Recipe

Ingredients

Makes a side for about 4 – my wife quadruples it for Christmas

Main Salad component

250g of Short pasta shapes
1 stick of Celery – diced
Half a of a Red and Green capsicums – diced
Small can of corn
1/2 cp Sultanas
300g Bacon – diced

Dressing

1/2 cp Salad oil (I just use canola oil)
1/2 cp Sugar
1/2 cp Vinegar
1 or 2 tsp of Keen’s curry powder

Method

Combine all dressing ingredients into a jar with a tight lid. Shake to combine and then set aside
Cook the pasta to packet instructions. Once cook; rinse under cold water and drain well and place in a large serving bowl.
Add the bacon to a cold non stick frypan. Turn to medium-high heat, stir occasionally until while foam appears around the cooked bacon piece. Remove bacon from pan and drain on paper towels.
In the serving bowl, add in the diced celery, diced capsicums, the small can of corn, sultanas and bacon. This can be wrapped up and placed in the fridge ready to be served.
When ready to serve, mix the salad ingredients together. Drizzle over salad dressing to desired level and mix into the salad.

June 3, 2025 by jonny.donker@gmail.com Cook Salads 0

“Idiots doing Idiot things” – The Top Eight Worst SQL Queries Ever

For over fifteen years I have worked with a Microsoft SQL Data Warehouse that is over twenty-five years old. Although it is ancient in this world of huge cloud-based Data Vaults – it churns out an abundant of value to the organisation I work for.

The problem with this value – it is opened to anyone who wants to access to build their queries off. This means we get a range of SQL queries running on the database. We have queries ranging from precise and efficient built queries that uses every trick in optimisation – all the way to “WTF are you trying to do?!” queries.

The second category is the one we have the most problem with. With limited resources of an OnPrem SQL server; some uses write atrocious queries without a perceived care for other people using the database. There is nothing more frustrating getting a support call out because of a failed data load; only to find the table is locked by someone running a query over the past six hours.

To counter this problem; I created a python script that checks what is running on the database every fifteen minutes. If a user has a query running for over half an hour, I get an alert to show what they are running.

This allows our team to:

Work out if someone is running a query that is blocking or taking resources away from our critical processes – resulting in timeouts
With our experience help users optimise poorly running queries so they have a better outcome
To stop any stupid queries from running

Since the queries are logged into a DB – I have four years of history of 14,000 unique queries running. This gives me a lot of learning experiences of what users are running and what problems they face.

From this learning – this is the Top 10 most horrible queries I see running on our database.

8. The “Discovery” Query

This query come from new and experience users exploring the data of the database. They will run something like this:

SELECT *
FROM dbo.A_VERY_LARGE_TABLE;

Look at the data, grab what they need and then just leave the query running in the background in their Server Management Studio. When we contact them an hour later; they act surprise that the query is still running in the background – returning millions upon millions of rows to their client.

This is an easy fix with training. Using the TOP command to limit the number of rows returned, use SP_HELP to return a schema of an object, provide a sandbox for people to explore in are all relatively simple fixes for this problem

7. The “Application Preview” Query

Our SAS application have a nasty default setting that if a user previews a data in a table; it will try and return the whole table – with the user completely unaware that this is happening.

Even worse; if the application locks up returning the data; the user kills the application from task manager, open it back up and performs the same steps. Since the first query was not killed, it will still be running on the database. So quite often you will see half a dozen of the same queries running on the database, all starting at different times.

To prevent this – it is important to assess new DB client software to see what it is trying to do in the background of the database. What does “preview” mean? Return 10, 100 or all the rows? Is there a setting that restricts the number of rows returned; or a timeout function that prevents complex views running forever just returning a small set of data?

If these features are available – it is important to either set this as default in the application; or part of the initial user setup / training program.

6. The “I am going to copy this to Excel and Analyse the Data” Query

This is like exploratory queries that come up. The user is running a query like:

SELECT *
FROM dbo.A_VERY_LARGE_TABLE;

“Why are you running this query?” I politely enquire.
“Oh – I am going to copy the data into excel and search/pivot/do fancy things with it”
“Ummm – this table has 8 billion rows in it…”

I can sympathise with queries like these.

The user thinks, “Why run the same query over and over again to get the data that I want to explore and present in different ways. Let just get one large cut and then manipulate it on the client.”

What users don’t realise that dealing with a huge amount of data on the client side is difficult. Exporting data out of the query client, into an application like excel can be frustrating with memory and CPU limitations compared to a spec’d-out server. Plus, Excel is not going to handle 8 billion rows.

The requirements of queries like these needs to be analysed – do the users need that atomic level data? Can strategic aggregation tables be provided to the user, so they have more manageable data to return. Can different technology like OLAP cubes be built on top of large tables to answer the user’s needs?

5. The “I got duplicates so I am going to get rid of them with a DISTINCT” Query

This is a pet peeve of mine.

The user has duplicate in their results and don’t know why. Instead of investigating why there is duplication (usually result of a poor join), they just slap a DISTINCT at the top of the query and call it done. Here is a simplified example that I regularly come across:

DECLARE @ACCOUNTS TABLE
(
	ORG INT,
	ACCOUNT_ID INT,
	BRANCH INT,
	BALANCE NUMERIC(18,2),
	PRIMARY KEY (ORG, ACCOUNT_ID)
);

INSERT INTO @ACCOUNTS VALUES(1, 18, 60, 50);
INSERT INTO @ACCOUNTS VALUES(1, 19, 60, 150);

DECLARE @BRANCHES TABLE
(
	ORG INT,
	BRANCH INT,
	BRANCH_NAME VARCHAR(20),
	BRANCH_STATE VARCHAR(3),
	PRIMARY KEY (ORG, BRANCH)
);

INSERT INTO @BRANCHES VALUES(1, 60, 'Branch of ORG 1', 'VIC');
INSERT INTO @BRANCHES VALUES(2, 60, 'Branch of ORG 2', 'VIC');


SELECT 
	A.ORG,
	A.ACCOUNT_ID, 
	A.BRANCH,
	A.BALANCE,
	B.BRANCH_STATE
FROM @ACCOUNTS A
JOIN @BRANCHES B
ON	-- B.ORG = A.ORG	-- User forgot this predicate in the join
	B.BRANCH = A.BRANCH;

Results:

ORG	ACCOUNT_ID	BRANCH	BALANCE	BRANCH_STATE
1	18	60	50.00	VIC
1	18	60	50.00	VIC
1	19	60	150.00	VIC
1	18	60	150.00	VIC

The user looks – “Oh Dear – I’ve got duplicates! Let’s get rid of them.”

SELECT DISTINCT
	A.ORG,
	A.ACCOUNT_ID, 
	A.BRANCH,
	A.BALANCE,
	B.BRANCH_STATE
FROM @ACCOUNTS A
JOIN @BRANCHES B
ON	-- B.ORG = A.ORG	-- User forgot this predicate in the join
	B.BRANCH = A.BRANCH;

This poses many problems:

The Database must work harder using incomplete joins on indexes to bring across the data; and then work harder supressing the duplicates
If the field list change; then the distinct might not work anymore.

For example, with the above query – if the user brings in the column “BRANCH_NAME” then the duplicates return:

SELECT DISTINCT
	A.ORG,
	A.ACCOUNT_ID, 
	A.BRANCH,
	A.BALANCE,
	--------- Add in new column ---------
	B.BRANCH_NAME,
	-------------------------------------
	B.BRANCH_STATE
FROM @ACCOUNTS A
JOIN @BRANCHES B
ON	-- B.ORG = A.ORG	-- User forgot this predicate in the join
	B.BRANCH = A.BRANCH;

ORG	ACCOUNT_ID	BRANCH	BALANCE	BRANCH_NAME	BRANCH_STATE
1	18	60	50.00	Branch of ORG 1	VIC
1	18	60	50.00	Branch of ORG 2	VIC
1	19	60	150.00	Branch of ORG 1	VIC
1	18	60	150.00	Branch of ORG 2	VIC

Preventing queries like this comes down to user experience and training. If they see duplicates in their data – their initial thoughts should be “Why do I have duplicates?”

Duplicates might be legitimate and a DISINCT might be OK – but quite often it is because of a bad join. For an inexperience user – it might be overwhelming to break down a large table to find where the duplicates are coming from, and they might need help from a more experienced user. This is where internal data forums are useful – where people can help each other with their problems.

4. The “Ever lengthening” Query

This is a query I see scheduled daily for users getting trends over time for a BI tool:

SELECT *
FROM dbo.A_TABLE A
JOIN dbo.B_TABLE B
ON	B.KEY_1 = A.KEY_1
WHERE
	--------------- Start Date ---------------
	A.BUSINESS_DATE >= '2025-01-01' AND
	------------------------------------------
	A.PREDICATE_1 = 'X' AND
	B.PREDICATE_2 = 'Y';

Initially it starts off OK – running fast and the user is happy. But as time goes by; the query gets slower and slower; returning more and more data. Eventually it gets to the point the query goes off into the nether and never returns. Since it is always starting the from the same point; it is constantly re-querying the same data over an over again

There are a couple of options you can handle queries like these:

Really examine the requirements of the user’s needs. How much data is really relevant for their trending report? In 2027; will data from 2025 be useful for the observer of the report?
Create a table specific for this report and append a timeframe of data (eg daily) onto it. The report than can just query this specific table and not having to regenerate complex joins over and over again for previous timeframes.

3. The “Linked Database” Query

With the growth of our Data Warehouse, other database on different servers started building process flows off our database; sometimes without out knowledge that they are doing this until we see a linked server connection coming up.

What pattern they use our database is where the problems lie. One downstream database runs this query every day:

SELECT *
FROM REMOTE_SERVER.SOME_DATABASE.dbo.A_BIG_HISTORICAL_TABLE;

So, they are truncating a landing table on their database and grabbing the whole table. Hundreds of millions of rows transferring across a slow linked server; taking hours and hours. And with time; this will get slower and slower as the source table grows its history.

This is a tricky one to tackle with the downstream users. Odds are if they design a cumbersome process like this their appetite for change to a faster (yet more complex solution) might not be high; especially if their process ‘works’.

If you can get around the political hurdle and convince them the process must change; there are many options to improve the performance.

Use a ETL tool to transfer the data across. Even a simple python script to copy batch by bath files across will be quicker than using a linked server.
Assess the downstream use cases for the data. Do they need the whole table with the complete history to satisfy their requirements? Do they need all the columns; or can some be trimmed off to reduce the amount of data getting transferred?
If the source table a historical table; only bring across the period deltas. With the downstream process I am having trouble with now; they are bringing across twenty-one million rows daily. If they only bring across the deltas from the night loads it brings across a whopping five thousand rows. That is 0.02% of the total row count. Add in a row count check between the two systems to have confidence that the two tables are in sync and the process will be astronomically quicker.

2. The “I don’t know where that query comes from” Query.

We have a continuing problem from a Microsoft Power BI report that runs a query that looks at data from the beginning of the month to the current date. It was a poorly written query and therefore as the month went along its performance got worse and worse.

Since the username associated with the query was from the Power BI server’s account – we had no idea who the query belonged to as a couple of simple fixes could drastically improve the performance.

We contacted the team that manages Power BI server and asked them to find the query and update the sql code.

They said it was completed, and the report sql was updated.

But soon the sql was back – running for hours and hours.

We contacted the team again –

“Hey that query is back.”

They try updating the report again; but soon the query was back again.

So, either someone was constantly uploading an original copy – or another copy of the report was buried somewhere on the server that the admins could not find.

Since we do not have access to their system, it hard to determine what is happening.

The barbaric solution would be to block the Power BI user from our system; but goodness knows how many critical business processes that will disrupt.

The best solution at the minute we are doing is just constantly killing the running SQL code on the database with the hope that someone will identify the constantly failing report. This can be tricky as well if the Power BI server automatically tries rerunning failed reports.

1. The “I don’t listen to your advice” Query.

Disclaimer – this is a frustration rant that I have to get off my chest.

We have a user that constantly runs a long running crook query. As in when I collated all the long running queries in research for this post – his was at the top by a significant margin.

The problem with the query itself is a join using an uncommonly used business key in a table that is not indexed. The fix itself is quite simple fix; use the primary keys in join.

But he has been running the same code for months – locked our nightly loads several times and caused incidents.

We tried the carrot approach. “Hey here is optimised code. It will make your query run quicker and be less burden on our database.” Got indifferent replies.

More locked loads we included his manager into correspondence, but she did not seem to want to be involved.

With my carrot approach communication, I gained the impression that he had a self-righteous personality and thinks his role and report is above question.

Enough was enough – it was the stick approach time.

We raised an operational risk against his activity.

The Ops risk manager asks, “Do you have evidence that you attempted to reason and help?”

“Yep,” I replied, “Here is a zipped-up folder of dozens of emails and chat messages that I sent him.”

Senior managers were involved; with my senior manager commenting to my manager in the background “Is that guy all there?”

Anyway, my manager wrote a very terse email saying to correct their query and if they crash the loads again; they will get their access removed from the database.

No committal reply. No apology. And to this day they are still running the same query; but just under the radar that it is not locking our loads.

I am watching him like a hawk – waiting for the day that I can cut his access from the database.

It is a pity that our Database does not have a charge back capability of processes used. I bet if his manager got a bill of $$$ from one person running one query; she would be more proactive in this problem.

In retrospect, I would have campaigned to have his access cut right away and make him justify why he should have it back. When there is hundreds of other people doing the right thing on the database; it is not fair that their deliverables are getting impacted by the indifference attitude of one user.

June 2, 2025 by jonny.donker@gmail.com Code SQL 0

Installing Animal Shelter Manager 3 on Docker – So close

Intro – Do Unto what the Sister In-law commands

My sister in-law runs a dog training business in Queensland. Part of her business is to track records of her canine clients – especially notes, vaccinations when they’re due, medical records and certificates.

In a previous job she had experience with Animal Shelter Manager (ASM3). She’s familiar with the features and interface to know it will cover her needs.

Her business is not big enough to justify the price of the SaaS version of ASM3; so being the tech savvy (debatable) one the family – it was my task getting it up and running for her.

This lead to several nights of struggling, annoyance and failure.

Fitting the pieces together

To start off I wanted to get a demo version running so I can see what I need to do to deploy it for her.

I checked out the ASM3 github repository and yes! There is a Dockerfile and a docker-compose.yml.

But no – it is six years old and didn’t work when I tried to build it. I tried a couple of other miscellaneous sites offering hope; but to no avail.

In the end after many google searches; I stumbled across https://lesbianunix.dev/about with the following guide:

Apart from a domain name that was sure to set off all the “appropriate content” filters at work; with a few modifications I could get it to work. Looking at the instructions from the author Ræn; it is substantially different to the old Dockerfile and the instructions on the ASM3 home page.

Let’s build it

With my base working version – I cobbled up some Dockerfiles for ASM3 and postres and a docker compose file to tie them together:

https://github.com/jon-donker/asm3_docker

(Note that this is not a production version and have to obscure passwords etc in the final version)

The containers build just fine and fire up with no problem.

Vising the website

http://localhost/

ASM3 redirects and builds the database – but then goes to a login page. I enter the username and password; but it loops back to the login page.

I think the problem lies with the base_url and service_url in asm3.conf; possibly with http-asm3.conf settings.

Anyway – I logged a issue with ASM3 see if it is something simple that I missed; or maybe I have to start pulling apart of source code to find what it is trying to do.

I’ll update this post when I find something.

April 25, 2025 by jonny.donker@gmail.com Docker 0

Qlik Replicate to AWS Postgres (Why are you so slow?!)

Wait – Performance problems writing to an AWS RDS Postgres database?

Haven’t we been here before?

Yes, we had, and I wrote quite a few posts of the trials and tribulations that we went through.

But this is a new problem that we came across that resulted in several messages backwards and forwards between us and Qlik before we worked out the problem.

Core System Migration.

Our organisation has several core systems; making maintenance of them expensive and holding us back in using the data in these systems in modern day tools like advance AI. Over the past several years – various projects are running to consolidate the systems together.

This is all fun and games for the downstream consumers – as they have lots of migration data coming down the pipelines. For instance, shell accounts getting created on the main system from the sacrificial system.

One downstream system wanted to exclude migration data from their downstream data and branch the data into another database so they can manipulate the migrated data to fit it into their pipeline.

I created the Qlik Replicate task to capture the migration data. It was a simple task to create. Unusually, the downstream users created their own tables that they want me top pipe the data into. In the past, we let Qlik Replicate create the table in the lower environment, copy the schema and use that schema going forwards.

Ready to go we fired up the task in testing to capture the test run through of the migration.

Slow. So Slow.

The test migration ran on the main core system, and we were ready to capture the changes under the user account running the migration process.

It was running slow. So slow.

We knew data was getting loaded as we were periodically running a SELECT COUNT(*) on the destination table. But we were running at less than 20tps.

Things we checked:

The source and target databases were not under duress.
The QR server (although a busy server) CPU and Memory wasn’t maxed out.
There were no critical errors in the error log.
Records were not getting written to the attrep_apply_exceptions table.
There were no triggers built off the landing table that might be slowing down the process
We knew from previous testing that we could get a higher tps.

I bumped up the logging on “Target Apply” to see if we can capture more details on the problem.

One by One.

After searching the log files, we came across an interesting message:

00007508: 2025-02-20T08:33:24:595067 [TARGET_APPLY    ]I:  Bulk apply operation failed. Trying to execute bulk statements in 'one-by-one' mode  (bulk_apply.c:2430)

00007508: 2025-02-20T08:33:24:814779 [TARGET_APPLY    ]I:  Applying INSERTS one-by-one for table 'dbo'.'DESTINATION_TABLE' (4)  (bulk_apply.c:4849)

For some reason instead of using a bulk load operation – QR was loading the records one by one. This accounted for the slow performance.

But why was it switching to this one-by-one mode? What caused the main import of bulk insert to fail – but one-by-still works.

Truncating data types

First, I suspected that the column names might be mismatched. I got the source and destination schema out and compared the two.

All the column names aligned correctly.

Turning up the logging we got the following message:

00003568: 2025-02-19T15:24:50 [TARGET_APPLY    ]T:  Error code 3 is identified as a data error  (csv_target.c:1013)

00003568: 2025-02-19T15:24:50 [TARGET_APPLY    ]T:  Command failed to load data with exit error code 3, Command output: psql:C:/Program Files/Attunity/Replicate/data/tasks/MY_CDC_TASK_NAME/data_files/0/LOAD00000001.csv.sql:1: ERROR:  value too long for type character varying(20)

CONTEXT:  COPY attrep_changes472AD6934FE46504, line 1, column col12: "2014-08-10 18:33:52.883" [1020417]  (csv_target.c:1087)

All the varchar fields between the source and target align correctly.

Then I noticed the column MODIFIED_DATE. On the source it is a datetime; while on the postgres target it is just a date.

My theory was that the bulk copy could not handle the conversion – but in the one-by-one; it could truncate the time component off the date and successfully load.

The downstream team changed the field from a date to a timestamp and I reloaded the data. With this fix the task went blindingly quick; from hours for just a couple of thousands of records to all done within minutes.

Conclusion

I suppose the main conclusion from this exercise is that “Qlik Replicate Knows best.”

Unless you have a very accurate mapping process from the source to the target; let QR create the destination table in a lower environment. Use this table as a source and build on it.

It will save a lot time and heartache later on.

March 17, 2025 by jonny.donker@gmail.com Qlik Replicate 0

Postgres: EBCDIC decoding through a JavaScript Function

EBCDIC? Didn’t that die out with punch cards and the Dinosaurs?

EBCDIC (Extended Binary Coded Decimal Interchange Code) is an eight-bit character encoding that was created by IBM in the ’60s.

While the rest of the world went on with ASCII and UTF-8; we still find fields in our DB2 database encoded in EBCDIC 037 just to make our lives miserable.

Qlik Replicate when replicating from these fields on its default settings; brings it across as a normal “string” and becomes quite unusable when loaded into a destination system.

Decoding EBCDIC in Postgres

To have the flexibility to decode particular fields in EBCDIC; we need to bring the fields across as BYTES instead of that QR suggests. This can be done in the Table Settings for the table in question:

On the destination Postgres database; load the table into a bytea field.

Now with a udf function in Postgres; we can decode the EBCDIC bytes fields into something readable:

CREATE OR REPLACE FUNCTION public.fn_convert_bytes2_037(
    in_bytes bytea)
    RETURNS character varying
    LANGUAGE 'plv8'
    COST 100
    VOLATILE PARALLEL UNSAFE
AS $BODY$
    const hex_037 = new Map([
        ["40", " ",],
        ["41", " ",],
        ["42", "â",],
        ["43", "ä",],
        ["44", "à",],
        ["45", "á",],
        ["46", "ã",],
        ["47", "å",],
        ["48", "ç",],
        ["49", "ñ",],
        ["4a", "¢",],
        ["4b", ".",],
        ["4c", "<",],
        ["4d", "(",],
        ["4e", "+",],
        ["4f", "|",],
        ["50", "&",],
        ["51", "é",],
        ["52", "ê",],
        ["53", "ë",],
        ["54", "è",],
        ["55", "í",],
        ["56", "î",],
        ["57", "ï",],
        ["58", "ì",],
        ["59", "ß",],
        ["5a", "!",],
        ["5b", "$",],
        ["5c", "*",],
        ["5d", ")",],
        ["5e", ";",],
        ["5f", "¬",],
        ["60", "-",],
        ["61", "/",],
        ["62", "Â",],
        ["63", "Ä",],
        ["64", "À",],
        ["65", "Á",],
        ["66", "Ã",],
        ["67", "Å",],
        ["68", "Ç",],
        ["69", "Ñ",],
        ["6a", "¦",],
        ["6b", ",",],
        ["6c", "%",],
        ["6d", "_",],
        ["6e", ">",],
        ["6f", "?",],
        ["70", "ø",],
        ["71", "É",],
        ["72", "Ê",],
        ["73", "Ë",],
        ["74", "È",],
        ["75", "Í",],
        ["76", "Î",],
        ["77", "Ï",],
        ["78", "Ì",],
        ["79", "`",],
        ["7a", ":",],
        ["7b", "#",],
        ["7c", "@",],
        ["7d", "'",],
        ["7e", "=",],
        ["7f", ","],
        ["80", "Ø",],
        ["81", "a",],
        ["82", "b",],
        ["83", "c",],
        ["84", "d",],
        ["85", "e",],
        ["86", "f",],
        ["87", "g",],
        ["88", "h",],
        ["89", "i",],
        ["8a", "«",],
        ["8b", "»",],
        ["8c", "ð",],
        ["8d", "ý",],
        ["8e", "þ",],
        ["8f", "±",],
        ["90", "°",],
        ["91", "j",],
        ["92", "k",],
        ["93", "l",],
        ["94", "m",],
        ["95", "n",],
        ["96", "o",],
        ["97", "p",],
        ["98", "q",],
        ["99", "r",],
        ["9a", "ª",],
        ["9b", "º",],
        ["9c", "æ",],
        ["9d", "¸",],
        ["9e", "Æ",],
        ["9f", "¤",],
        ["a0", "µ",],
        ["a1", "~",],
        ["a2", "s",],
        ["a3", "t",],
        ["a4", "u",],
        ["a5", "v",],
        ["a6", "w",],
        ["a7", "x",],
        ["a8", "y",],
        ["a9", "z",],
        ["aa", "¡",],
        ["ab", "¿",],
        ["ac", "Ð",],
        ["ad", "Ý",],
        ["ae", "Þ",],
        ["af", "®",],
        ["b0", "^",],
        ["b1", "£",],
        ["b2", "¥",],
        ["b3", "·",],
        ["b4", "©",],
        ["b5", "§",],
        ["b6", "¶",],
        ["b7", "¼",],
        ["b8", "½",],
        ["b9", "¾",],
        ["ba", "[",],
        ["bb", "]",],
        ["bc", "¯",],
        ["bd", "¨",],
        ["be", "´",],
        ["bf", "×",],
        ["c0", "{",],
        ["c1", "A",],
        ["c2", "B",],
        ["c3", "C",],
        ["c4", "D",],
        ["c5", "E",],
        ["c6", "F",],
        ["c7", "G",],
        ["c8", "H",],
        ["c9", "I",],
        ["ca", "",],
        ["cb", "ô",],
        ["cc", "ö",],
        ["cd", "ò",],
        ["ce", "ó",],
        ["cf", "õ",],
        ["d0", "}",],
        ["d1", "J",],
        ["d2", "K",],
        ["d3", "L",],
        ["d4", "M",],
        ["d5", "N",],
        ["d6", "O",],
        ["d7", "P",],
        ["d8", "Q",],
        ["d9", "R",],
        ["da", "¹",],
        ["db", "û",],
        ["dc", "ü",],
        ["dd", "ù",],
        ["de", "ú",],
        ["df", "ÿ",],
        ["e0", "\\",],
        ["e1", "÷",],
        ["e2", "S",],
        ["e3", "T",],
        ["e4", "U",],
        ["e5", "V",],
        ["e6", "W",],
        ["e7", "X",],
        ["e8", "Y",],
        ["e9", "Z",],
        ["ea", "²",],
        ["eb", "Ô",],
        ["ec", "Ö",],
        ["ed", "Ò",],
        ["ee", "Ó",],
        ["ef", "Õ",],
        ["f0", "0",],
        ["f1", "1",],
        ["f2", "2",],
        ["f3", "3",],
        ["f4", "4",],
        ["f5", "5",],
        ["f6", "6",],
        ["f7", "7",],
        ["f8", "8",],
        ["f9", "9",],
        ["fa", "³",],
        ["fb", "Û",],
        ["fc", "Ü",],
        ["fd", "Ù",],
        ["fe", "Ú"]
    ]);
 
    let in_varchar = "";
    let build_string = "";
     
    for (var loop_bytes = 0; loop_bytes < in_bytes.length; loop_bytes++)
    {
        /* Converts a byte character to a hex representation*/
        let focus_char = ('0' + (in_bytes[loop_bytes] & 0xFF).toString(16)).slice(-2); 
        let return_value = hex_037.get(focus_char.toLowerCase());
 
        /* If no mapping found - replace the character with a space */
        if(return_value === undefined)
        {
            return_value = " ";
        }
 
        build_string = build_string.concat(return_value)
    }
 
    return build_string
$BODY$;

The function can now be used in SQL:

SELECT public.fn_convert_bytes2_037(my_EBCDIC_byte_column)
FROM public.foo;

Reference

JavaScript bytes to HEX string function: Code Shock – How to Convert Between Hexadecimal Strings and Byte Arrays in JavaScript

January 9, 2025 by jonny.donker@gmail.com Postgres Qlik Replicate 0

Postgres JavaScript missing variables (But it is #$%^ there!)

It’s OK

I only cried and contemplated quitting working in IT and becoming a Nomad for a couple of hours.

But I got there in the end; but the following error message will probably plague my nightmares for a couple of weeks:

ERROR:  ReferenceError: inNumber1 is not defined
CONTEXT:  fn_js_number_adder() LINE 2: 	let total = inNumber1 + inNumber2 

SQL state: XX000

JavaScript: When in Rome – Do what the Romans do

My job today was to write a JavaScript function in Postgres to convert byte hex values to EBCDIC 037. The aim is to decommission some duplicate pipelines coming from our DB2 database by Qlik Replicate that deliver ASCII converted fields as well as the EBCDIC version.

I haven’t worked in JavaScript since my Uni days and well entrenched in the Python world for my day to day job. Over the past years I have converted using naming conventions in code from camelCase to under_score to match Python’s standard.

So going back to JavaScript – I knew that camelCase is the expected format. Since I didn’t know where my code was going to end up; I wanted it to look professional as it is a reflection on me.

So I wrote a JavaScript function paraphrased as:

CREATE OR REPLACE FUNCTION fn_js_number_adder(inNumber1 numeric, inNumber2 numeric)
RETURNS numeric
as
$$
	let total = inNumber1 + inNumber2

	return total
$$
LANGUAGE plv8;

Looks good – compiles with no errors.

But when I went to test it; I get the following error:

The error drove me crazy! It’s THERE! The variable is THERE!

The original function was a lot more extensive than above so I cut as much out of it as possible in case something else was causing the variable not to be recognised.

Still no luck.

I went to the functions section in pgAdmin as I wanted to compare it against an existing function I created to see what the difference was.

Interesting…

The function’s parameters have changed from inNumber1 and inNumber2 to innumber1 and innumber2.

Scripting out the function I got:

CREATE OR REPLACE FUNCTION public.fn_js_number_adder(
	innumber1 numeric,
	innumber2 numeric)
    RETURNS numeric
    LANGUAGE 'plv8'
    COST 100
    VOLATILE PARALLEL UNSAFE
AS $BODY$
	let total = inNumber1 + inNumber2

	return total
$BODY$;

So; either postgres or pgAdmin changed the case of the parameters from camelCase to lower case. This caused the variable not to be found later in the code.

The fix – Back to under_scores we go

My fix for this instance (whether standard or not) is to go back to under_scores:

CREATE OR REPLACE FUNCTION fn_js_number_adder(in_number1 numeric, in_number2 numeric)
RETURNS numeric
as
$$
	let total = in_number1 + in_number2

	return total
$$
LANGUAGE plv8;

This works and I could run the function

With the naming conventions; I suppose using under_score isn’t too much of a sin since it is a standard on databases. If you want to stay true to camelCase; the parameters can just be in under_score and the rest of the variables be in camelCase.

At lest it is working…now onto EBCDIC conversion.

January 7, 2025 by jonny.donker@gmail.com Postgres 0

Qlik Replicate: Oh Oracle – you’re a fussy beast

It’s all fun and games – until Qlik Replicate must copy 6 billion rows from a very wide Oracle table to GCS…

…in a small time window

…with the project not wanting to perform Stress and Volume testing

Oh boy.

Our Dev environment had 108 milling rows to play with, which ran “quick” in relationship the amount of data it had to copy. But being 33 times smaller; even if it takes an hour in Dev – extrapolating the time out will relate to over 30 hours of run time.

The project forged ahead in the implementation and QR only processed 2% of the changes before we ran out of the time window.

The QR servers didn’t seem stressed performance wise; had plenty of CPU and RAM. I suspect the bottle neck was in the bandwidth going out to GCS; but there was no way to monitor how much of the connection has been used.

When in doubt – change the file type

After the failed implementation, we tried to work out how we can improve the throughput to GCS in our dev environment.

I thought changing the destination’s file type might be a way. JSON is a chunky file format, and my hypothesis was if the JSON was compressed it would transfer to GCS quicker. We tested out a NULL connector, raw JSON, GZIP JSON and Parquet. As a test using Dev – we let a test task run for 20min to see how much data is copied across.

Full Load Tuning:

Transaction consistency timeout (seconds): 600
Commit rate during full load: 100000

Endpoint settings:

Maximum file size(KB): 1000000 KB (1GB)

Results

Unfortunately, my hypothesis on compressed JSON was incorrect. We speculated that compressing the JSON might have been taking up as much time as transferring it. I would have like to test this theory on a quieter QR server, but time is of the essence.

Parquet seemed to be the winner with the limited testing offering a nice little throughput boost over the JSON formats. But it wasn’t the silver bullet to our throughput problems. Added onto this; the downstream users would need to spend time modifying their ingestion pipelines.

Divide and conquer – until Oracle says no.

The next stage was to look if we could divide the table up into batches and transfer across section at a time. Looking at the primary key; it was an identity column that had little meaningful relation to easily divide up into batches.

There was another indexed column called RUN_DATE; which is a date relation to when the record was entered.

OK – let’s turn on Passthrough filtering and test it out.

First of all to test the syntax out in SQL Developer

SELECT COUNT(*)
FROM xxxxx.TRANSACTIONS
WHERE
    RUN_DATE >= '01/Jan/2023' AND
    RUN_DATE < '01/Jan/2024';

The query ran fine meaning that the date syntax was right.

Looking good – let’s add the filter to the Full Load Passthru Filter

But when running the task; it goes into “recoverable error” mode.

Looking into the logs:

00014204: 2024-11-26T08:38:26:184700 [SOURCE_UNLOAD   ]T:  Select statement for UNLOAD is 'SELECT "PK_ID","RUN_DATE", "LOTS", "OF", "OTHER, "COLUMNS"  FROM "xxxxx"."TRANSACTIONS" WHERE (RUN_DATE >= '01/Jan/2023' AND RUN_DATE &lt; '01/Jan/2024')'  (oracle_endpoint_utils.c:1941)
00014204: 2024-11-26T08:38:26:215961 [SOURCE_UNLOAD   ]T:  ORA-01858: a non-numeric character was found where a numeric was expected  [1020417]  (oracle_endpoint_unload.c:175)
00014204: 2024-11-26T08:38:26:215961 [SOURCE_UNLOAD   ]T:  Failed to init unloading table 'xxxxx'.'TRANSACTIONS' [1020417]  (oracle_endpoint_unload.c:385)
00014204: 2024-11-26T08:38:26:215961 [SOURCE_UNLOAD   ]E:  ORA-01858: a non-numeric character was found where a numeric was expected  [1020417]  (oracle_endpoint_unload.c:175)
00014204: 2024-11-26T08:38:26:215961 [SOURCE_UNLOAD   ]E:  Failed to init unloading table 'xxxxx'.'TRANSACTIONS' [1020417]  (oracle_endpoint_unload.c:385)
00014204: 2024-11-26T08:38:26:215961 [SOURCE_UNLOAD   ]T:  Error executing source loop [1020417]  (streamcomponent.c:1942)
00014204: 2024-11-26T08:38:26:215961 [SOURCE_UNLOAD   ]T:  Stream component 'st_1_SRC_DEV_B1_xxxxx' terminated [1020417]  (subtask.c:1643)
00014204: 2024-11-26T08:38:26:215961 [SOURCE_UNLOAD   ]T:  Free component st_1_SRC_DEV_B1_xxxxx  (oracle_endpoint.c:51)
00011868: 2024-11-26T08:38:26:215961 [TASK_MANAGER    ]I:  Task error notification received from subtask 1, thread 0, status 1020417  (replicationtask.c:3603)
00014204: 2024-11-26T08:38:26:215961 [SOURCE_UNLOAD   ]E:  Error executing source loop [1020417]  (streamcomponent.c:1942)
00014204: 2024-11-26T08:38:26:215961 [TASK_MANAGER    ]E:  Stream component failed at subtask 1, component st_1_SRC_DEV_B1_xxxxx  [1020417]  (subtask.c:1474)
00014204: 2024-11-26T08:38:26:215961 [SOURCE_UNLOAD   ]E:  Stream component 'st_1_SRC_DEV_B1_xxxxx' terminated [1020417]  (subtask.c:1643)
00011868: 2024-11-26T08:38:26:231570 [TASK_MANAGER    ]W:  Task 'TEST_xxxxx' encountered a recoverable error  (repository.c:6200)

Error code ORA-01858 seems to be the key to the problem. As an experiment I copied out the select code and ran it into SQL Developer.

Works fine 🙁

OK – maybe it is a quirk of SQL Developer?

Using sqlplus I ran the same code from the command line.

Again works fine 🙁

Resorting to good old Google – I searched ORA-01858.

The top hit was this article from Stack Overflow that recommended confirming the format of the date with the TO_DATE function.

OK Oracle; if you want to be fussy with your dates – let’s explicitly define the date format with TO_DATE in the Full Load Passthru Filter.

RUN_DATE >= TO_DATE('01/Jan/2023','DD/Mon/YYYY') AND RUN_DATE < TO_DATE('01/Jan/2024','DD/Mon/YYYY')

Ahhhh – that works better and Qlik Replicate now runs successfully with the passthrough filter.

Conclusion

I tried a different set of date formats; including an ISO date format and Oracle spat them all out. So using TO_DATE is the simplest way to avoid the ORA-01858 error. I can understand Oracle refusing to run on a date like 03/02/2024; I mean is it the 3rd of Feb 2024; or for Americans the 2nd of Mar 2024? But surprised something very clear like an ISO date format; or 03/Feb/2024 did not work.

Maybe how SQL Developer and SQLplus interacts with the database is different than QR that leads to the different in behaviour of how filters on dates work.

November 25, 2024 by jonny.donker@gmail.com Oracle Qlik Replicate 0

jonny.donker@gmail.com

What statistics to use?

The Legend behind the recipe

The Recipe

Ingredients

Method

Notes

The Recipe

Ingredients

Method

8. The “Discovery” Query

7. The “Application Preview” Query

6. The “I am going to copy this to Excel and Analyse the Data” Query

5. The “I got duplicates so I am going to get rid of them with a DISTINCT” Query

4. The “Ever lengthening” Query

3. The “Linked Database” Query

2. The “I don’t know where that query comes from” Query.

1. The “I don’t listen to your advice” Query.

Intro – Do Unto what the Sister In-law commands

Fitting the pieces together

Let’s build it

Core System Migration.

Slow. So Slow.

One by One.

Truncating data types

Conclusion

EBCDIC? Didn’t that die out with punch cards and the Dinosaurs?

Decoding EBCDIC in Postgres

Reference

JavaScript: When in Rome – Do what the Romans do

The fix – Back to under_scores we go

When in doubt – change the file type

Results

Divide and conquer – until Oracle says no.

Conclusion

If you found this useful

Recent Posts

Categories

Archives