- I went to Western Ghats near Ettina Bhuja: an overnight stay in a homestay, and easy hike early in the morning. It was a long weekend so a few friends also joined. Kaalu enjoyed hiking in the mountains.

- My on-boarding has started at Progress.

Spooky action at distance when using async Rust.
Read the following code and guess the output.
It has two concurrent tasks. The first task sets a cancellation token after 150ms. The second task accepts a variable initialized to 0, increments it twice with 100ms sleeps in between, and finally resets it to 0.
Then we have a tokio::select! that returns the first branch that completes, and cancels the second branch.
use std::sync::{Arc, Mutex};use tokio_util::sync::CancellationToken;async fn main() { let cancellation_token = CancellationToken::new(); let token = cancellation_token.clone(); let cancelled = tokio::spawn(async move { // X1 tokio::time::sleep(tokio::time::Duration::from_millis(150)).await; token.cancel(); }); let a = Arc::new(Mutex::new(0)); tokio::select! { _ = cancelled => { println!("cancellation token is set."); } _ = long_task(a.clone()) => { } }; println!("a ={}", a.lock().unwrap());}async fn long_task(state: Arc<Mutex<i32>>) { println!("long task started..."); { let mut lock = state.lock().unwrap(); *lock += 1; } // Y1 tokio::time::sleep(tokio::time::Duration::from_millis(100)).await; { let mut lock = state.lock().unwrap(); *lock += 1; } // Y2 tokio::time::sleep(tokio::time::Duration::from_millis(100)).await; { // clean up (reset to 0) let mut lock = state.lock().unwrap(); *lock = 0; } println!("long_task has ended.");}
If you guessed an answer other than 2, you need to read about cancellation safety. The Tokio documentation also talks about it extensively.
Because at 150ms, the spawned task completes, causing tokio::select! to select that branch and immediately drop the long_task future before it can reach the cleanup step. Note that long_task is dropped not because it checks the CancellationToken, but because tokio::select! automatically drops all non-winning futures.
This behavior can be dangerous: what if the remaining code was cleaning up a resource (like running a database cleanup query or releasing a lock) rather than resetting a variable? You could leave your system in a broken state.
long_task starts, sets a = 1, and yields at Y1 (sleep 100ms).Y1 finishes. long_task sets a = 2 and yields at Y2 (sleep 100ms).X1 finishes. The cancelled task completes, and its JoinHandle resolves.tokio::select! receives the completed branch and drops long_task while it is sleeping at Y2. The final cleanup block is never executed, leaving a = 2.Async rust has a few parts that doesn’t feel ‘rusty’ at all. Rust is pretty good at “forcing” local reasoning but async cancellation (drop of Future etc.) leads to non-local reasoning which leads to hard to follow sequence of events that leads to subtle bugs. I recently learnt about futurelock from this excellent blog post.
The RFD (Request For Discussion) from Oxide that describe futurelock (https://rfd.shared.oxide.computer/rfd/0609) is easy to read by an intermediate Rust programmer. Reading this RFD made me a little bit nervous about async which I though I knew decently well.
I created this diagram that summarizes the sequence of events in the RFD that eventually leads to the deadlock/futurelock. You can refer to this diagram when re-reading the RFD. It helps a lot.

I also dug a little deeper into the mechanism of Mutex waking up the relevant tasks when it is unlocked. In the past, I’ve written state machines with callbacks and I think about async in state-machine terms. Future and Waker works together to implment state-machine with callback.
Here is another diagram which shows how Waker is used to implement callback like mechanism for the example in the RFD.

I am going to do something terrible β store videos in a SQL database!
I have some requirements that makes it an acceptable plan. I am building a βstream storeβ S. Once videos are stored in S, users should be able to fetch a video segment (in color or grayscale) between two given timestamps at a given FPS. One should also be able to annotate frame later e.g., βthis frame has a face in itβ. Iβll read the incoming video stream and save data as frames to a database (in addition to storing raw videos on a backup server).
The primary job of S is to provide frames for CV analysis, so serving pixel-perfect video stream is not a requirement.
I am going to extract frames from video as JPEG β a lossy compression format. But what should be the quality of JPEG?
I downloaded a sample mkv file β 1080p at 30 FPS to do simple analysis.
https://filesamples.com/samples/video/mkv/sample_1280x720.mkv
| FPS | 23.976 |
|---|---|
| Duration | 28.237 |
| Size on disk | 16.63 MB |
I wrote a script that extract frames from the mkv using ffmpeg. The argument -qscale:v set the quality of frames: 2 is the best and 32 is the worst.
ffmpeg -i ../sample_1280x720.mkv '%04d.png' generates PNG folder which is 1.3 GB, almost 78x of original size. PNG is a lossless format. This is as bad as its get.ffmpeg -i ../sample_1280x720.mkv '%04d.jpg' generates JPEG with default quality picked by ffmpeg. It generates 33 MB of data. Almost 1.94x more. Great!ffmpeg -i ../sample_1280x720.mkv -qscale:v 2 '%04d.jpg' generates JPEG with best possible quality. The generated size is 188 MB (almost 11x more!).I did the same on a different file recorded at 60fps (original size . The data is below.
A few things to note
I think qscale:v=20 is a good default for my use case. Also I donβt have to extract frames at the same rate as they are recording. I am interested in events at the timescale of ~100ms and anything faster than 20FPS is overkill. i can just extract at 30 fps.
ffmpeg has a handy cli option -filter:v "fps=30" to fix the extraction fps to 30. Here is bonus rust code that does this. Donβt copy-paste blindly, it may not work.
/// Explode the given video into JPEG frames.
///
/// - *path*: Path of video file
/// - *fps*: Extract these many frames per seconds. The video may contain
/// more or less frames in a second.
pub fn extract_jpegs<P: AsRef<Path> + std::fmt::Debug>(
path: P,
fps: u16,
recording_start_timestamp_ms: i64
) -> anyhow::Result<usize> {
anyhow::ensure!(path.as_ref().exists(), "{:?} does not exists", path);
tracing::debug!(
"Extracting frames from {:?} for fps={fps} and jpg qscale {:?}.",
path.as_ref(),
self.qscale
);
let inst = Instant::now();
let mut cmd = std::process::Command::new(&self.ffmpeg_bin_path);
cmd.arg("-i").arg(path.as_ref());
// Extract at a given fps. Thanks <https://askubuntu.com/a/1019417/39035>.
cmd.arg("-filter:v").arg(format!("fps={fps}"));
if let Some(qscale) = self.qscale {
tracing::debug!("Setting qscale to {qscale}");
cmd.arg("-quality:v");
cmd.arg(qscale.to_string());
}
cmd.arg(self.frame_directory.join("%05d.jpg"));
let output = cmd.output()?;
anyhow::ensure!(
output.status.success(),
"Command failed\n.{}\n{}",
String::from_utf8_lossy(&output.stdout),
String::from_utf8_lossy(&output.stderr)
);
tracing::debug!(
"Extraction to {:?} is complete, took {:?}.",
&self.frame_directory,
inst.elapsed()
);
}

– I wrote another small utility to remind me that a LWN article has finally become open. I am not able to pay for LWN subscription due to HDFC Bank related issues and I forget to revisit the link when it is open. Perhaps I can rent a VPS in this money and read the article two weeks later?! LWN is a great resource and I feel bad for not paying for it though!
– My Wallet from Budgetbakers is no longer syncing with HDFC Bank. Their support is working on the issue. HDFC seem to have changed their login flow again! My another bank, DBS Bank, doesnβt have saving account APIπ€£. I opened account here thinking that they are βtech-savvyβ!
– Iβve been thinking about hiring a lot these days. At my current company which is an early stage startup, they have been struggling to hire a dev for the last 5 months! I was part of a few interviews β some went well and most were meh, but not able to hire for 5 months feels a bit extreme!
– Both mango trees in my street has mangoes this year! Here is my daughter Ookie playing with her friend. Fortunately, like many streets in Bengaluru, this street is a dead end and have no traffic.
hplip is not always up to date. It is nice to take a printout of an important document such as resume and read with attention they deserve. Many candidates spent days on them!I needed to add more derive traits on the enum generated by openapi-generate-cli e.g. strum::EnumIter.
openapi-generator-cli author template -g rust . This will write templates to the directory out.out directory. I had to add strum to Cargo.mutache and strum::EnumIter inside derive of all pub enum inside model.mustache file.openapi-generator-cli generate -t out -i openapi.json -g rust -o api-client-rsAnd you have updated client.
Here are some diffs
PS: I could not get the local version to run on my system. I used the following docker command
podman run --rm \
-v $PWD:/local openapitools/openapi-generator-cli generate \
-i https://beta.oneapi.dognosis.link/api/v1/openapi.json \
-t /local/templates \
-g rust \
-o /local/oneapi-client-rs
rust-projects/leptos at main Β· dilawar/rust-projects


There is also a QR scanner that may or may not work in your browser. I used a third party crate that I patched to compile with Leptos 0.7 but did not test is thoroughly.
leptos-use (https://leptos-use.rs/) makes it easy to access browserβs API such a audio/video streams. The audio recorder records the audio in chunk of a few seconds and plot the raw values on the canvas. The plotting is very basic here!
I tried using https://thawui.vercel.app/ components but could not make it work with my slightly complicated hooks that update the local storage. I reverted back to standard leptos input component.
## [component]
pub fn Form() -> impl IntoView {
let storage_key = RwSignal::new("".to_string());
let (state, set_state, _) = use_local_storage::<KeyVal, JsonSerdeCodec>(storage_key);
let upload_patient_consent_form = move |file_list: FileList| {
let len = file_list.length();
for i in 0..len {
if let Some(file) = file_list.get(i) {
tracing::info!("File to upload: {}", file.name());
}
}
};
view! {
<h5>"Form"</h5>
<Space vertical=true class=styles::ehr_list>
// Everything starts with this key
<ListItem label="Code".to_string()>
<input bind:value=storage_key />
</ListItem>
// Patient
<InputWithLabel key="phone".to_string() state set_state></InputWithLabel>
<InputWithLabel key="name".to_string() state set_state></InputWithLabel>
<SelectWithLabel
key="gender".to_string()
options=Gender::iter().map(|x| x.to_string()).collect()
state
set_state
></SelectWithLabel>
<InputWithLabel key="extra".to_string() state set_state></InputWithLabel>
</Space>
}
}
## [derive(Debug, strum::EnumIter, strum::Display)]
enum Gender {
Male,
Female,
Other,
}
## [component]
pub fn InputWithLabel(
key: String,
state: Signal<KeyVal>,
set_state: WriteSignal<KeyVal>
) -> impl IntoView {
let label = key.split("_").join(" ");
let key1 = key.to_string();
view! {
<Flex>
<Label>{label}</Label>
// Could not get thaw::Input to change when value in parent changes.
<input
prop:value=move || {
state.get().0.get(&key1).map(|x| x.to_string()).unwrap_or_default()
}
on:input=move |e| {
set_state
.update(|s| {
s.0.insert(key.to_string(), event_target_value(&e));
})
}
/>
</Flex>
}
}
// SelectWithLabel is not shown. See the linked repo for updated code.
To make this problem concrete, Iβve downloaded database of all CVEs from https://github.com/CVEProject/cvelistV5 using git clone. I want to create API that can query this repository. I can think of the following available options.
grep based tool e.g. git grep , rg , grep etc.SQL either using a third party tool or using a custom solution.Iβd have preferred 1 or 2 if I was writing a cli application. Since I am building a RESTful API, Iβd prefer 3 since queries will be easier to write and integrate into other applications. Custom solution is also not a bad idea if existing tooling is not good enough or there are other constraints. Most databases like PostgreSQL and sqlite3 allows JSON to be inserted and queried.
While searching for such plugins, I came across https://duckdb.org/ which seems to support this use cases. See for example https://duckdb.org/docs/data/multiple_files/overview π―.
For this exercise, I am not interested in performance. I can use caching to improve performance drastically later is need arise. I am looking for a solution that has the best DX β and if I am lucky doesnβt easily allow stupid mistakes!
The installation was a breeze. Single binary! It can also be used as Rust/Python library and third party integration for PHP are also available π―.
$ wget https://github.com/duckdb/duckdb/releases/download/v1.1.3/duckdb_cli-linux-amd64.zip
$ unzip duckdb_cli-linux-amd64.zip
Archive: duckdb_cli-linux-amd64.zip
inflating: duckdb
(PY311) [dilawar@rasmalai cve_database_git (master)]$ ./duckdb
v1.1.3 19864453f7
Enter ".help" for usage hints.
Connected to a transient in-memory database.
Use ".open FILENAME" to reopen on a persistent database.
D
Great, seems to work!
For a sanity check, I did a βloopbackβ β print what you read. Just to make sure that engine is parsing JSON files. I passed a glob that matches all JSON files. Loading all files together raised an error β mismatch in schema.
D SELECT * FROM 'cves/**/*.json';
Invalid Input Error: JSON transform error in file "cves/2001/1xxx/CVE-2001-1517.json", in record/value 1: Object {"affected":[{"product":"n/a","vendor":"n/a","vers... has unknown key "tags"
Try increasing 'sample_size', reducing 'maximum_depth', specifying 'columns', 'format' or 'records' manually, setting 'ignore_errors' to true, or setting 'union_by_name' to true when reading multiple files with a different structure.
D
Fair enough warning about inconsistent schema across files and it suggested what I should explore. I like tools that hints at what to do in case of errorπ―.
I am not sure which one is the best option here though: ignore_errors or perhaps union_by_name is a better ideaπ€? I am going ahead with union_by_name. I had to tweak the query little bit: SELECT * FROM read_json('cves/2000/**/*.json', union_by_name = true);
Nice!
Note that we are not yet inserting these JSON record into to a SQL table but rather processing them on the fly. I say this because I donβt see the opened database test.duck.db size to increase at all!
To insert these JSON records into SQL table that can be queries later, we use the following
D CREATE TABLE cve AS SELECT * FROM read_json_auto('./cves/2003/*/*.json', union_by_name = true);
100% ββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
D select COUNT(*) FROM cve;
ββββββββββββββββ
β count_star() β
β int64 β
ββββββββββββββββ€
β 1553 β
ββββββββββββββββ
I am ready to append to this table and write queries.
I did the same exercise as before using duckdb python module.
import duckdb
git_repo = _cve_repo_dir()
logger.info(f"Syching duckdb from {git_repo}...")
json_files = list(git_repo.glob("cves/*/**/*.json"))
con = duckdb.connect("cve.duck.db")
\# create table if it doesn't exists using a JSON.
first_json = json_files[-1]
logger.info(f"> Creating table using {first_json}")
con.sql(
f"CREATE TABLE IF NOT EXISTS cves AS SELECT * FROM read_json('{str(first_json)}', union_by_name = True)"
)
s = con.sql("SELECT * FROM cves;")
print(s)
2024-11-30 11:46:24.723 | INFO | utils:sync_duckdb:40 - > Creating table using /home/dilawar/Work/SUBCOM/keeda/keeda-py/cve_json_db.git/cves/2024/9xxx/CVE-2024-9999.json
βββββββββββββ_βββββββββββββ_ββββββββββββββββββββββ_βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β dataType β dataVersion β cveMetadata β containers β
β varchar β varchar β struct(cveid varch_ β struct(cna struct(affected struct(defaultstatus varchar, platforms varchar[], product va_ β
ββββββββββββββΌββββββββββββββΌβββββββββββββββββββββββΌβββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ€
β CVE_RECORD β 5.1 β {'cveId': CVE-2024_ β {'cna': {'affected': [{'defaultStatus': unaffected, 'platforms': [Windows], 'product': W_ β
ββββββββββββββ΄ββββββββββββββ΄βββββββββββββββββββββββ΄βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
Great, once I have table initialized with schema, I can insert entries from other files: INSERT INTO cves SELECT * FROM read_json('{str(file)}'). One has to ensure that other files have same schema. If not, then we get error while inserting.
TypeMismatchException: Mismatch Type Error: Type STRUCT(cveId VARCHAR, assignerOrgId UUID, state VARCHAR, dateReserved VARCHAR, dateUpdated
VARCHAR, dateRejected VARCHAR, assignerShortName VARCHAR) does not match with STRUCT(cveId VARCHAR, assignerOrgId UUID, state VARCHAR,
assignerShortName VARCHAR, dateReserved VARCHAR, datePublished VARCHAR, dateUpdated VARCHAR). Cannot cast STRUCTs - element "dateRejected" in
source struct was not found in target struct
In this case, the other files have some fields added over time and it make sense to add those columns/fields to the schema on the fly (ref: https://duckdb.org/docs/sql/statements/update.html#update-from-other-table).
But I just couldnβt figure out how to the resolve the following error. Looks like updating an existing table to support the schema of a new file is not supported out of the box.
TypeMismatchException: Mismatch Type Error: Type STRUCT(defaultStatus VARCHAR, platforms VARCHAR[], product VARCHAR, vendor VARCHAR, versions
STRUCT(lessThan VARCHAR, status VARCHAR, "version" VARCHAR, versionType VARCHAR)[]) does not match with STRUCT(collectionURL VARCHAR,
defaultStatus VARCHAR, packageName VARCHAR, product VARCHAR, versions STRUCT(lessThanOrEqual VARCHAR, status VARCHAR, "version" VARCHAR,
versionType VARCHAR)[], vendor VARCHAR). Cannot cast STRUCTs of different size
The solution to this problem in this case was simple. Each sub-directory contains files with the same schema. I created a table for each sub-directory!
2.7GB worth of CVEs were stored in 456MB of database file.
[dilawar@khaja keeda-py (download_cve_database)]$ du -sh cve_json_db.git/cves/
2.7G cve_json_db.git/cves/
[dilawar@khaja keeda-py (download_cve_database)]$ ls -ltrh cve.duck.db
-rw-r--r-- 1 dilawar dilawar 456M Nov 30 20:56 cve.duck.db
[dilawar@khaja keeda-py (download_cve_database)]$
Now the querying part. I want to query CVEs that affect a product called wpzoom. Thanks https://github.com/duckdb/duckdb/issues/9901#issuecomment-1843178389
D select * from cve_in_year_2024_9xxx where len(list_filter(containers->>'$..affected[*].vendor', x-> x == 'wpzoom'));
βββββββββββββ_βββββββββββββ_ββββββββββββββββββββββ_βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β dataType β dataVersion β cveMetadata β containers β
β varchar β varchar β struct(cveid varch_ β struct(cna struct(providermetadata struct(orgid uuid, shortname varchar, dateupd_ β
ββββββββββββββΌββββββββββββββΌβββββββββββββββββββββββΌβββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ€
β CVE_RECORD β 5.1 β {'cveId': CVE-2024_ β {'cna': {'providerMetadata': {'orgId': b15e7b5b-3da4-40ae-a43c-f7aa60e62599, 'sh_ β
ββββββββββββββ΄ββββββββββββββ΄βββββββββββββββββββββββ΄βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
D
I have somethat that may look like the output of tree like command. I want it a plain list wihtout nesting for further processing. I want to convert the following on the left to the one on the right.
{
"/opt": [
"a",
"b",
"c",
],
"/root": [
"a",
"b",
"c",
],
}
[
"/opt/a",
"/opt/b",
"/opt/c",
"/root/a",
"/root/b",
"/root/c",
]
The following gist shows how to do it. You can play with code on the playground here https://play.rust-lang.org/?version=stable&mode=debug&edition=2021&gist=dfdbc1036a4bc502488598a742552fba.
Line 12 converts HashMap to a vector of vector where inner most vector is join of key with individual values. flatten converts vector of vector to a vector (akin to concat).
https://gist.github.com/rust-play/3a11b9e4d16e992e995c8aa0d26c6141
First the rant!
For last 3 months, my Zoho Mail client lost an essential features: desktop notification when a calendar event is about to occur. It sends in-app notification but it is useless since I donβt pay attention to that. Calendar notifications are much more important since most of them are social contracts and one must not mixed them with mere notification. I missed many meetings or got late by a few minutes. A developer easily lose track of time when working!
I wrote to Zoho support and they told me that they are reimplementing it a brand new feature that will enable this again. Note to product manager β donβt break a working feature unless the implementation is ready. And they claimed they have enabled the desktop notification just for me. To this day, I am yet to see a desktop notification from Zoho Email client. I double, triple check the settings and after 10 years of experience with software development, I canβt figure is how to enable this settings, this tool is not for me!
So I wrote a small tool that does exactly this.
https://github.com/dilawar/ical-desktop-notification
Given calendar iCal url, it sends you desktop notification if an event is about to occur. At the time of this writing, βabout to occurβ means in 3 minutes. You can tweak this and perhaps send me a PR if you make it configurable from the cli.
How to find an iCal url? Most calendar providers should have it enabled in settings. Here is a screenshot from google calendar.
You should use the private URL if you also want to see the title of the event.
On windows, you can add this tool to Task Schedular.
I canβt believe I am using Windows β things you do to make a living!
/// Add to to PATH environment variable. If `append` is false, add to the
/// front of list.
pub fn add_to_path(path: &str, append: bool) -> anyhow::Result<()> {
use anyhow::Context;
use std::env;
use std::path::PathBuf;
let paths = env::var_os("PATH").context("empty PATH")?;
let mut paths = env::split_paths(&paths).collect::<Vec<_>>();
if append {
paths.push(PathBuf::from(path));
} else {
paths.insert(0, PathBuf::from(path));
}
let new_path = env::join_paths(paths)?;
env::set_var("PATH", new_path);
Ok(())
}
PWSTR β PCWSTR | let s: PWSTR=...; let r:PCWSTR = PCWSTR(s.0) |
String β PCWSTR | let s: String=...; let r:PCWSTR = PCWSTR(HSTRING::from(s).as_ptr()) |
&str β PCWSTR | let r:PCWSTR = w!("example") |
//! Trait for converting &str/String into PWSTR.//! Thanks https://github.com/microsoft/windows-rs/issues/973#issue-942298423## ![cfg(windows)]use windows::core::PWSTR;pub trait IntoPWSTR { fn into_pwstr(self) -> (PWSTR, Vec<u16>);}impl IntoPWSTR for &str { fn into_pwstr(self) -> (PWSTR, Vec<u16>) { let mut encoded = self.encode_utf16().chain([0u16]).collect::<Vec<u16>>(); (PWSTR(encoded.as_mut_ptr()), encoded) }}impl IntoPWSTR for String { fn into_pwstr(self) -> (PWSTR, Vec<u16>) { let mut encoded = self.encode_utf16().chain([0u16]).collect::<Vec<u16>>(); (PWSTR(encoded.as_mut_ptr()), encoded) }}
References https://github.com/microsoft/windows-rs/issues/973