poniedziałek, 24 października 2011

How to import content from Blosxom to google Blogger

I have decided to give a try to Google Blogger service. I am an old dinosaur used to command line and tired with mouse and menus but as there is GoogleCL I am not scare. The problem is with my old posts---there is no way to post backdated blog entries with GoogleCL. A problem...

Fortunately there is export/import features on Blogger: one can backup blog content and/or upload it back to Google. In particular to import posts (and comments) into a blog, one have to click Import Blog from the blog's Settings. Next one have to select appropriate file and fill out the word verification beneath. The Blogger data format is Atom. So, to successfully import my old Blosxom entries I have to convert them to Atom.

I have made a few test entries and export them to check how the data looks like. Pretty wired but most of the content is irrelevant as it is concerned with formatting (css styles and such stuff is included). Also as I had comments disabled at my previous blog the problem is further simplified.

I have consulted Atom schema and tried with the following:


<?xml version="1.0" encoding="UTF-8"?>
<feed xmlns="http://www.w3.org/2005/Atom"
xmlns:openSearch="http://a9.com/-/spec/opensearchrss/1.0/"
xmlns:georss="http://www.georss.org/georss"
xmlns:gd="http://schemas.google.com/g/2005"
xmlns:thr="http://purl.org/syndication/thread/1.0">';

<id>tag:blogger.com,1999:blog-1928418645181504144.archive</id>
<updated>2011-10-22T12:34:14.746-07:00</updated>
<title type='text'>pinkaccordions.blogspot.com</title>
<generator version='7.00' uri='http://www.blogger.com'>Blogger</generator>

The meaning of the elements should be obvious. The last element (generator) is required by Blogger import facility, otherwise error message is returned.

According to the schema inside feed element there is zero or more entry elements:


<entry>
<id>ID</id>
<published>DATE</published>
<updated>DATE</updated>
<category scheme="http://schemas.google.com/g/2005#kind" term="http://schemas.google.com/blogger/2008/kind#post"/>

<!-- tags, each as value of attribute `term' of element category -->
<category scheme='http://www.blogger.com/atom/ns#' term='tag1'/>
<category scheme='http://www.blogger.com/atom/ns#' term='tag2'/>
<title type='text'>title</title>
<content type='html'>post content ... </content>
</entry>

There is a final </feed> to guarantee that XML file is well formatted.

I have assumed the only important feature of id element is that it's content should be unique. I have decided to use MD5sum of the post content as IDs to guarantee that.

Finally, my old Blosxom-compatible entries looks similar to the example below:


<?xml version='1.0' encoding='iso-8859-2' ?>
<html xmlns="http://www.w3.org/1999/xhtml"><head>
<title>Przed finałami RWC 2011</title>
<!-- Tags: rwc2011,rugby,francja,polsat-->
</head><body><!-- ##Published : 2011-10-20T07:20:26CEST ##-->


<p>W RWC 2011 zostały już tylko dwa mecze: jutro (piątek), o trzecie miejsce oraz

So it was extremly easy to extract title, tags and publication date and format Atom-compliant XML file with the following Perl script:


#!/usr/bin/perl
# Variant of Blosxom to Blogger conversion
# 2011/10 t.przechlewski
#
use Digest::MD5 qw(md5_hex);

print '<?xml version="1.0" encoding="UTF-8"?>
<!-- id, title/updated jest wymagane w elementach feed/entry reszta opcjonalna -->
<!-- wyglada na minimalne oznakowanie -->
<feed xmlns="http://www.w3.org/2005/Atom"
xmlns:openSearch="http://a9.com/-/spec/opensearchrss/1.0/"
xmlns:georss="http://www.georss.org/georss"
xmlns:gd="http://schemas.google.com/g/2005"
xmlns:thr="http://purl.org/syndication/thread/1.0">';

print "<id>tag:blogger.com,1999:blog-1928418645181504144.archive</id>";
print "<updated>2011-10-22T12:34:14.746-07:00</updated>";
print "<title type='text'>pinkaccordions.blogspot.com</title>";
print "<generator version='7.00' uri='http://www.blogger.com'>Blogger</generator>\n";

foreach $post_file (@ARGV) {

my $post_title = $post_content = $md5sum = $published = '';
my @post_kws = ();
my $body = $in_pre = 0;
my $rel_URLs = 0;

print STDERR "\n$post_file opened!\n";
open POST, "$post_file" || die "*** cannot open $post_file ***\n";

while (<POST>) {
chomp();

if (/<title>(.+)<\/title>/) {$post_title = $1 ; next ; }
if (/<!--[ \t]*Tags:[ \t]*(.+)[ \t]*-->/) {$tags = $1 ; next ; }

if (/<\/head><body>/) {
$body = 1 ;
## </head><body><!-- ##Published : 2011-10-20T07:20:26CEST ##-->
if (/##Published[ \t]+:[ \t]+([0-9T\-\:]+).+##/) { $published = $1; }
print STDERR "Published: $published\n";
next;
}

if (/<\/body><\/html>/) { $body = 0 ; next }

if ( $body ) {
## sprawdzam `przenosnosc URLi':
if (/src[ \t]*=/) {
if (/pinkaccordions.homelinux.org/ || !(/http:\/\// ) ) { $rel_URLs = 1; }
}
## zawartość pre nie powinna być składana w jednym wierszu:
if (/<pre>/) { $in_pre = 1; $post_content .= "$_\n"; next ; }
if (/<\/pre>/) { $in_pre = 0; $post_content .= "$_ "; next ; }
if ( $in_pre ) { $post_content .= "$_\n"; }
else {
$post_content .= "$_ "; # ** musi być ze spacją **
}
}
}

### ### ###

if ($published eq '') {
warn "*** something wrong with: $post_file. Not published? Skipping....\n" ;
close(POST);
next ;
}
if ( $tags eq '' || $post_title eq '' ) {
die "*** something wrong with: $post_file (tags: $tags/title: $post_title)\n"; }
if ($rel_URLs) { die "*** suspicious relative URIs: $post_file\n"; }

$post_content =~ s/\&/&amp;/g;
$post_content =~ s/</&lt;/g;
$post_content =~ s/>/&gt;/g;

print STDERR "Title: $post_title Tags: $tags\n";

@post_kws = split /,/, $tags;
$md5sum = md5_hex($post_content);
print STDERR "MD5sum: $md5sum\n";

print "<entry>";
print "<id>tag:blogger.com,1999:post-$md5sum</id>";
print "<published>$published</published>";
print "<updated>$published</updated>";
print '<category scheme="http://schemas.google.com/g/2005#kind" term="http://schemas.google.com/blogger/2008/kind#post"/>';

## tags:
foreach $k (@post_kws) { print "<category scheme='http://www.blogger.com/atom/ns#' term='$k'/>"; }

print "<title type='text'>$post_title</title>";
print "<content type='html'>$post_content</content></entry>";

close(POST);

}

print "</feed>";

The minor problem was the default formatting of <pre>...</pre> which I use to show code snippets. I have to preserve line breaks (cf. $in_pre in the above Perl script) of pre element content as well as have to add the following to the default CSS styles (it is possible to modify CSS via Project →Template Designer →Advanced →Add CSS1)


pre { white-space:nowrap; font-size: 80%; }

To convert simply run script as follows:


perl blogspot-import.pl post1 post2 post3.... > converted-posts.xml

The above described script can be downloaded from here.

1In Polish: Projekt →Projektant szablonów →Advanced →Dodaj Arkusz CSS

czwartek, 20 października 2011

Przed finałami RWC 2011

W RWC 2011 zostały już tylko dwa mecze: jutro (piątek), o trzecie miejsce oraz w niedzielę finał Nowa Zelandia vs Francja. Oba na kanałach sportowych Polsatu. A więc finał jak w pierwszym turnieju i wynik też zapowiada się podobny.

Półfinał Francja vs Walia to było shame on Wales. Byli lepsi za wyjątkiem kopania z karnych. Z bodajże 5 kopnęli 1 (a przegrali 8:9) trzy razy zmieniając kopacza. Temu trzeciemu w pierwszej i ostatniej jego próbie w 75 min z 50 metrów prawie się udało (piłka przeleciała ciut pod poprzeczką).

Do tego sędzia usunął Walijczyka za tzw. tip tackle (przykład, po polsku wysoka szarża) już w 20 min, tj. prawie godzinę Walia grała w 14 przeciw 15. Nb. Francuz spadł na łeb, ale w dużym stopniu było to unintentional więc wielu ocenia decyzję sędziego jako pochopną... Pozostał niesmak... Francja nie zagrała jednego dobrego meczu a jest w finale.

środa, 19 października 2011

Completely plain emails in Thunderbird

To switch Thunderbird (version 7.*) to compose mails in plain text (completely plain emails) go to Edit → Preferences → Advanced → Config Editor (in Polish: Edycja → Preferencje → Zaawansowane → Edytor Ustawień) and set both mail.html_compose and mail.identity.default.compose_html to false.

BTW according to the Thunderbird documentation: When you write a message in Thunderbird, you use either an HTML editor or a plain text editor. If you use an HTML editor, then you can send the message in HTML, plain text, or both. If you use a plain text editor, then you can only send the message in plain text. You usually specify which type of editor to use by identity. In your account settings, on the Composition and Addressing page, either check or clear the checkbox: "Compose messages in HTML format". Unfortunately there is no such a checkbox in version 7 of TB, or it is hidden perfectly...

In order to temporarily shift between plain-text and HTML editors hold the SHIFT key while clicking the Compose (or Replay) buttons.

wtorek, 18 października 2011

Fragment ścieżki dźwiękowej filmiku z YT

Poniższe operacje miały na celu wykonanie (oryginalnego) dzwonka do telefonu.

Po pierwsze, zamieniłem plik z YToube na mp3 korzystając z serwisu www.flv2mp3.com. Rezultat konwersji ściągnąłem ,,do siebie''.

Po drugie, za pomocą ffmpeg wyciąłem pierwsze 8 sekund zaczynając od 2 sekundy:


ffmpeg -ss 00:00:02 -t 8 -i plik.mp3 -acodec copy -vcodec copy plik-8s.mp3

koniec

piątek, 14 października 2011

Zmarł Dennis Ritchie

Umarł Dennis Ritchie (9 października). Współtwórca systemu Unix i języka C. Parę dni wcześniej, jak umarł Steve Jobs, który na dobrą sprawę nic nie wymyślił -- był zdolnym managerem w stylu Gatesa czy Forda (tego antysemity od samochodów), to cały świat płakał...1

W takiej sytuacji przypomniała mi się historia Biblii Gutenberga, która powinna się nazywać Gutenberga-Fusta, a nawet Fusta, bo przecież wykolegował on był Gutenberga. A jednak nie, wielcy sprzedawcy nie przechodzą do historii...

1No prawie cały... RMS nie płakał.

czwartek, 13 października 2011

Wycieczka z Tczewa do Kwidzyna przez Gniew

Z Tczewa do Kwidzyna przez Walichnowy i Gniew. W Gniewie promem na drugi brzeg Wisły. Asfalt kiepski (ale nie bardzo kiepski), do tego kawałek drogi szutrowej ale takiej, że na upartego to nawet na oponach 23mm dałoby się jechać. Bardzo mały ruch samochodowy... Ogólnie mogę polecić. Zamiast do Kwidzyna można wykręcić zurik na Tczew przez Piekło-Białą Górę.

Ślad jest tutaj.

Dictionary-style headers in LaTeX

It is pretty simple to define dictionary-style headers in LaTeX. Look at the example below:


\documentclass[twoside]{report}
\usepackage{multicol}

%%% \leftmark/\rightmark inserts appropriate \mark{...}
\usepackage{fancyhdr}
\pagestyle{fancyplain}
\fancyhead{}
\fancyhead[LE]{\normalfont \small\itshape \rightmark }
\fancyhead[RE]{\normalfont \small\itshape \leftmark }
\fancyhead[RO]{\normalfont \small\itshape \leftmark }
\fancyhead[LO]{\normalfont \small\itshape \rightmark }
\fancyfoot[CO,CE]{\thepage}

\title{Vocabulary layout -- example}

\newcommand{\Entry}[2]{\par \leavevmode
\ignorespaces\markboth{#1}{#1} #1 #2}

\begin{document}
\maketitle

\begin{multicols}{2}

\Entry{AA BB CC DD00001}{aa bb cc dd ee ff gg hh ii jj mm nn oo pp qq rr ss tt uu vx 00001 zz.}
\Entry{AA BB CC DD00002}{aa bb cc dd ee ff gg hh ii jj mm nn oo pp qq rr ss tt uu vx 00002 zz.}
\Entry{AA BB CC DD00003}{aa bb cc dd ee ff gg hh ii jj mm nn oo pp qq rr ss tt uu vx 00003 zz.}

... ... ...

\Entry{AA BB CC DD00299}{aa bb cc dd ee ff gg hh ii jj mm nn oo pp qq rr ss tt uu vx 00299 zz.}
\Entry{AA BB CC DD00300}{aa bb cc dd ee ff gg hh ii jj mm nn oo pp qq rr ss tt uu vx 00300 zz.}


\end{multicols}

\end{document}

Crucial commands are \markboth, \leftmark and \rightmark. See also Typesetting a Dictionary with LaTeX.

piątek, 7 października 2011

How to Delete Duplicate Files in Linux

There is a dedicated tool in Linux to finds duplicate files in a given set of directories.


yum install fdupes

Now, if---for example---one wants to find duplicates in the /cmn and /home/tomek directories, one have to enter the following command:


fdupes -r /cmn /home/tomek
## or prompt the user for files to preserve, deleting all others:
fdupes -r -d /cmn /home/tomek
## or preserve the first file in each set of duplicates and delete
## the others without prompting the user (terrifying dangerous:-):
fdupes -r -d -N /cmn /home/tomek

Note -r for recursive. Without -r subdirectories are not scanned. One have to remember not to remove duplicates one needs.

Fdupes searches the given paths for duplicate files. Such files are found by comparing file sizes and MD5 signatures, followed by a byte-by-byte comparison.

środa, 28 września 2011

Perl, MySQL i UTF-8

Strasznie dużo czasu zmarnowałem usiłując zmienić kodowanie w bazie MySQL na UTF i dopasować skrypty Perla do tej zmiany.

Przestawienie MySQLa na UTF-8 jest proste:


mysql -u www -p --default-character-set=utf8
CREATE database kpma CHARACTER SET utf8 COLLATE utf8_bin;

Baza kpma (użytkownika www) będzie kodowana w UTF-8. Można też określić domyślne kodowanie wszystkich baz w pliku konfiguracji MySQLa, tj. w pliku /etc/mysql/my.cnf (Debian):


[mysqld]
...
default-character-set = utf8

Trochę diagnostyki:


use kpma;
show variables like 'char%';
show table status;

select tytul from Utwor;
## jest OK -- na konsoli widać poprawne różne znaki diakrytyczne

Prawdziwa męka to było zmuszenie Perla do poprawnego traktowania danych UTF.

Trzy kluczowe dla poprawnego przetwarzania UTF sprawy to: 1) klauzula binmod (patrz poniżej); 2) klauzula use utf8 (jeżeli skrypt zawiera napisy kodowane w UTF); 3) wpisy mysql_enable_utf8/SET NAMES utf8 dotyczące MySQLa.

Szkielet skryptu Perla wygląda następująco:


#!/usr/bin/perl -w
# -*- coding: utf-8 -*- --
#
use strict;
use utf8; ## skrypt zawiera napisy kodowane UTF
use CGI qw(:standard);
use DBI;
binmode(STDOUT, ":utf8"); ## bez tego problemy z UTF

my $dbname = 'kpma'; ## Nazwa bazy
my $dbuser = 'www'; ## Nazwa użytkownika MySQL
my $dbpasswd = '??????'; ## Hasło dla $dbuser

my $dsn = "dbi:mysql:$dbname:localhost:3306";

my $dbh = DBI->connect($dsn, "$dbuser", "$dbpasswd", { ChopBlanks => 1 });
$dbh->{'mysql_enable_utf8'} = 1;
$dbh->do('SET NAMES utf8');

my $SQL = "SELECT tytul FROM Utwor WHERE id_kompozytor1 = 59 ORDER BY rok ";
##my $SQL = "SELECT nazwisko FROM Kompozytor ";

my $sth = $dbh->prepare($SQL);

$sth ->execute();

while ( my @piece = $sth->fetchrow_array() ) { print ">> @piece\n"; }

$dbh->disconnect || warn "Nie mogę zamknąć bazy $dbname\n";

Jeżeli skrypt korzysta (pobiera dane) z param() to koniecznie należy zastosować funkcję decode_utf8:


use Encode; ## param() trzeba dekodować

if (param()) {# -- Wypełniono formularz --
## http://ahinea.com/en/tech/perl-unicode-struggle.html
my $who = Encode::decode_utf8(param("kto"));

Działa nawet z dość starym Perlem:


$perl --version

This is perl, v5.10.0 built for arm-linux-gnueabi-thread-multi

Copyright 1987-2007, Larry Wall

środa, 14 września 2011

Wycieczka rowerowa do Brodnicy

W poprzednią sobotę/niedzielę (2/3 09. br.) pojechałem z Miśkiem na rowerach do Brodnicy (za Kartuzami) i z powrotem. Ślady są tutaj i tutaj.

Wojna za progiem...

10 maja 1941 roku Rudolf Hoess (formalnie zastępca Adolfa Hitlera, tj. człowiek który zostałby fuererem gdyby, przykładowo wodzowi coś zaszkodziło) wystartował z Augsburga a następnie wyskoczył ze spadochronem niedaleko Eaglesham w Szkocji. Został aresztowany i resztę wojny spędził w więzieniu.

Propaganda 3Rz tłumaczyła prostemu ludowi powyższy -- przyznajmy szokujący incydent (wiceprezydent Biden wyskakuje na spadochronie z F16 pod Teheranem) -- iż towarzysz partyjny Hoess zwariował (Goebbels uważał takie tłumaczenie za prostackie i niewiarygodne).

Minister finansów RP 14. 09 br (albo dzień wcześniej) palnął publicznie, ich zitiere: Rostowski przywołał na koniec swą niedawną prywatną rozmowę z ,,prezesem wielkiego polskiego banku'', pracującym w ministerstwie za czasów transformacji w Polsce. Miał on mu powiedzieć, że po takich wstrząsach gospodarczych i politycznych, jakie teraz dotykają Europę, ,,rzadko się zdarza, by po 10 latach nie było także katastrofy wojennej''. I -- jak mówił minister -- ten prezes banku dodał, że ,,poważnie się zastanawia nad tym, by uzyskać dla dzieci zieloną kartę w USA''.

I co teraz na to Ministerstwo Propagandy, Oświecenia publicznego i Informacji?